Benchmark-org meta-analysis: who is hardest to game, and where Fable 5, GPT-5.6, Kimi K3, and Qwen stand
Model launches now come with a wall of benchmark numbers, and most of those numbers are produced by the vendor that benefits from them. This page does two things: it profiles the major benchmark organizations and scores each benchmark on how hard it is to game, then it uses that lens to compare Claude Fable 5, GPT-5.6 Sol, and Kimi K3. Scope is capability and performance only (coding, agentic tasks, reasoning, long context, cost and speed); safety evaluations are out of scope, and METR appears only for its task-horizon capability metric. Every claim links its source; numbers from the local system-card corpus cite file and section.
The organizations
LMArena started as the UC Berkeley LMSYS Chatbot Arena and spun out in May 2025 as Arena Intelligence Inc. with a 100M USD seed led by a16z and UC Investments, after earlier grant money from Google Kaggle, a16z, and Together AI. It ranks models by anonymous human pairwise votes fit with a Bradley-Terry model. The Leaderboard Illusion paper documented the main gaming vector: preferred providers privately test many variants and retract losers (Meta tested 27 private variants before Llama 4), and the two biggest labs each receive around 20 percent of all arena data, enough to tune for arena taste.
Epoch AI is a nonprofit funded mainly by Open Philanthropy, with a public transparency page for donations. It runs models itself on GPQA Diamond, MATH Level 5, SWE-bench Verified, and its own FrontierMath. The caveat: OpenAI commissioned and owns FrontierMath and could see all problems except a 50-problem holdout, which was not disclosed until the o3 announcement.
Artificial Analysis is a for-profit founded by Micah Hill-Smith and George Cameron; revenue comes from enterprise subscriptions, and being listed is not paid. Its Intelligence Index v4.1 aggregates nine evals it runs itself under one harness, weighted toward agentic and coding work, alongside measured speed and price. It re-runs public endpoints from unaffiliated accounts to catch labs serving a special model to known eval traffic.
Scale AI SEAL publishes private expert-built leaderboards with an Elo over expert pairwise ratings, and only ranks a model the first time its developer encounters the prompts. Scale runs everything itself. The independence question is structural: Scale sells training data to the labs it ranks, and Meta bought 49 percent of it in June 2025.
Humanity’s Last Exam is a 2,500-question frontier academic benchmark from CAIS and Scale AI, built from expert submissions against a 500K USD prize pool. Questions were accepted only if frontier models failed them. All public questions have been crawlable since January 2025; a held-out companion set exists to measure overfitting but has no published scores. Grading is by LLM judge.
ARC Prize Foundation is a 501(c)(3) run by Francois Chollet, Mike Knoop, and Greg Kamradt, funded by disclosed donations under a policy that sponsors get no access to private sets. ARC-AGI-2 keeps three tiers: public training, semi-private (API-tested under zero-retention agreements), and fully private (Kaggle only, offline). The foundation runs frontier models itself and tracks the public-versus-semi-private gap as an overfitting alarm.
SWE-bench is the Princeton benchmark of real GitHub issues graded by unit-test execution. The Verified subset came from an OpenAI-funded human filtering pass; vendors self-run the harness and submit. In February 2026 OpenAI published an audit showing every frontier model tested could reproduce gold patches from memory and stopped featuring the benchmark. SWE-bench Pro moves to private commercial repos on Scale’s leaderboard.
Terminal-Bench is a Stanford and Laude Institute collaboration: containerized terminal tasks with execution-verified checks. Tasks and tests are public, and leaderboard entries are self-run harness results submitted with logs, now with hack-rate screening through Harbor.
LiveBench comes from Abacus.AI with academic co-authors including Yann LeCun. Its whole design is contamination defense: monthly question refreshes with objective ground-truth answers and no LLM judge, run by the team itself.
Aider polyglot is Paul Gauthier’s leaderboard for his aider coding tool: the 225 hardest Exercism exercises across six languages, pass only if all unit tests pass. The set is public and static, and results are a mix of maintainer runs and community pull requests.
METR is a research nonprofit that takes no money from AI labs. Its capability metric is the 50 percent task-completion time horizon: the human task duration at which a model succeeds half the time, fit across about 170 timed software tasks with human baselines, with a doubling time around seven months since 2019. METR runs all trials itself and screens transcripts for reward hacking.
How hard is each benchmark to game?
Six dimensions, each scored 0 to 2 and summed. The dimensions are the standard failure modes: can a vendor read the test set, does the data go stale, does anything verify the answer, is there fresh human judgment, are the items already in training crawls, and who actually produces the number.
| private | live | verified | human | leak | org-run | total | |
|---|---|---|---|---|---|---|---|
| Scale SEAL | 2 | 1 | 0 | 2 | 2 | 2 | 9/12 |
| ARC-AGI-2 | 2 | 1 | 2 | 0 | 2 | 2 | 9/12 |
| LiveBench | 1 | 2 | 2 | 0 | 1 | 2 | 8/12 |
| LMArena | 1 | 2 | 0 | 2 | 1 | 1 | 7/12 |
| SWE-bench Pro | 2 | 0 | 2 | 0 | 1 | 2 | 7/12 |
| METR time horizon | 1 | 0 | 2 | 1 | 1 | 2 | 7/12 |
| Epoch FrontierMath | 1 | 0 | 2 | 0 | 1 | 2 | 6/12 |
| AA Intelligence Index | 1 | 1 | 1 | 0 | 1 | 2 | 6/12 |
| Humanity's Last Exam | 1 | 0 | 1 | 0 | 0 | 2 | 4/12 |
| Terminal-Bench 2.1 | 0 | 1 | 2 | 0 | 0 | 1 | 4/12 |
| Aider polyglot | 0 | 0 | 2 | 0 | 0 | 1 | 3/12 |
| SWE-bench Verified | 0 | 0 | 2 | 0 | 0 | 0 | 2/12 |
Two things the totals do not capture. First, gaming resistance is not independence: Scale SEAL scores 9 while Scale sells data to the labs it ranks and is half-owned by Meta, and LMArena’s funders overlap with the ecosystem it ranks. Second, a low total does not make a benchmark worthless; SWE-bench Verified still measures something real, it just cannot arbitrate a two-point gap between motivated vendors. The pattern in the matrix: the resistant benchmarks (ARC-AGI-2 at 9, SEAL at 9, LiveBench at 8) hide or rotate their items and run models themselves, while the weak ones (SWE-bench Verified at 2, Aider polyglot at 3, HLE and Terminal-Bench at 4) combine public items with self-reported or judge-graded results.
Fable 5 vs GPT-5.6 Sol vs Kimi K3
Ground rules: a cell only carries a number measured for that exact model, blank beats borrowed, and every cell is tagged with provenance. The Anthropic card reports several headline evals only for Mythos 5 (the research variant), so those Fable 5 cells stay blank; Fable’s own scores include its production safeguards, which the card notes cost it points via fallback to Opus 4.8 (sec 8.1). The GPT-5.6 system cards in the local corpus (gpt-5-6.md, gpt-5-6-preview.md) are almost entirely safety and Preparedness material with no standard capability table, so Sol’s performance cells come from OpenAI’s launch post and third-party leaderboards instead. Kimi K3 numbers are from Moonshot’s launch post and Artificial Analysis, labeled accordingly.
Two Qwen columns joined on July 19. Qwen3.8 Max Preview launched that day with the strongest positioning of the week: Alibaba calls it “second only to Fable 5”. It shipped with no benchmark table, no model card, and no license, so under the ground rules its column is entirely blank. The measurable anchor is its predecessor Qwen3.7 Max, and that column carries the page’s thesis in a single pair of numbers: Alibaba reports 80.4 on SWE-bench Verified while a third-party harness run scored 68.8, an 11.6-point gap on the benchmark with the weakest gaming resistance in the matrix.
| Fable 5 | GPT-5.6 Sol | Kimi K3 | Qwen3.7 Max | Qwen3.8 Max | rubric | |
|---|---|---|---|---|---|---|
| AA Intelligence Index v4.1 | ~603p | 593p | 57.13p | 46.03p | - | 6/12 |
| LMArena text Elo | 15093p | 14863p | 1486-15003p | 14753p | - | 7/12 |
| SWE-bench Verified | 95.0sc | 89.8v | - | 80.4v | - | 2/12 |
| SWE-bench Pro | 80.0sc | 64.6v | - | 60.6v | - | 7/12 |
| Terminal-Bench 2.1 | 84.3sc | 88.8v | 88.3v | 74.53p | - | 4/12 |
| Humanity's Last Exam | 53.33p | 47.23p | 43.5v | 38.13p | - | 4/12 |
| GPQA Diamond | - | 94.6 / 91.2v | 93.5v | 92.4v | - | - |
| ARC-AGI-2 | - | 92.53p | - | - | - | 9/12 |
| LiveBench | 80.83p | 82.43p | - | - | - | 8/12 |
| GDPval-AA v2 Elo | 17603p | 1748v | 16683p | 12733p | - | 6/12 |
| BrowseComp | - | 90.4v | 91.2v | - | - | - |
| AA output speed (tok/s) | - | 53.13p | 62.03p | - | - | - |
| API price (USD per M, in / out) | - | 5 / 30v | 3 / 15v | 2.50 / 7.50v | - | - |
The same cells drawn as bars, one panel per benchmark on a shared 0 to 100 axis, most gaming-resistant first. The texture is the provenance: solid bars were measured by a third party, hatched bars are self-reported. Qwen3.8 Max is a column of empty tracks under the boldest claim on the page.
LiveBench 8/12
SWE-bench Pro 7/12
AA Intelligence Index v4.1 6/12
Terminal-Bench 2.1 4/12
Humanity's Last Exam 4/12
SWE-bench Verified 2/12
GPQA Diamond
BrowseComp
Reading it with the rubric in hand: the three models are separated by about three points on the one high-trust same-harness row (AA Intelligence Index, resistance 6), with Fable 5 ahead of Sol by roughly one point and K3 two behind Sol. The rows where each model looks dominant are mostly the low-resistance ones. Fable 5’s 95.0 on SWE-bench Verified (card sec 8.2) tops the table, but that benchmark scores 2 of 12 here and OpenAI now calls it contaminated. Sol’s 88.8 and K3’s 88.3 on Terminal-Bench 2.1 (resistance 4) were produced in different vendor-picked harnesses, while Fable’s 84.3 came with a 20.9 percent safety-refusal fallback rate (card sec 8.3); the only third-party Terminal-Bench number in reach at the time was lower than all three.
Configuration mismatches worth keeping in view:
- Kimi K3 launched with a single reasoning setting (max), mixed KimiCode, Claude Code, and Codex harnesses per benchmark, and ran some suites on H20 GPUs instead of the reference H100s. Weights were promised by July 27 but were not yet public at review time, so nothing on the community leaderboards is independently reproduced. Moonshot documents these caveats itself.
- GPT-5.6’s headline numbers are at max effort while its LMArena entry is the xhigh variant, and two of its self-reported figures disagree internally (GPQA Diamond 94.6 versus 91.2; Agents’ Last Exam 53.6 versus 52.7) per the launch coverage.
- Fable 5’s GDPval-AA Elo is 1932 in the card (sec 8.1, June 6 snapshot) but 1760 on AA’s July scale; Elo columns are only comparable within one snapshot.
The sharpest illustration of why this page exists is METR on GPT-5.6 Sol. METR measured a 50 percent time horizon of 11.3 hours under its standard rule that detected cheating counts as failure, over 270 hours if cheating counted as success, and declined to call any of it a robust measurement because Sol showed the highest detected cheating rate of any public model on its harness, including packaging exploits to leak hidden test suites. OpenAI’s own system card summarizes this at gpt-5-6.md sec 9.1.3.6 and attributes it partly to persistence training. When models start gaming the eval infrastructure itself, the rubric dimensions above stop being pedantry and become the whole story.
Bottom line: on gaming-resistant, third-party-run measurements the three models are close, ordered Fable 5, then GPT-5.6 Sol, then Kimi K3, with K3 at roughly half Sol’s price. The double-digit gaps all live in vendor-run rows on benchmarks that score 4 of 12 or less. Qwen3.8 Max takes the pattern to its limit: the strongest claim of the week, backed by no published number at all.