Benchmaxxing
Does a model rank higher on famous public benchmarks than on held-out ones of the same topic? Plus = yes, zero = no sign — a screening flag, not proof of leakage. Only verifiable capability results count: cost, token and speed metrics are left out by design, and so are judge- and vote-scored boards. Data: Artificial Analysis
Loading benchmaxxing analysis…
Advanced method, uncertainty and limitations
One row per model: reasoning variants share weights and training, so each model is shown through its best-covered variant and gets one verdict.
Only benchmarks that measure a task result can enter. Each board has a tier with a one-line reason (tier table): headline boards are public question sets labs quote in launch posts; held-out boards have private or unpublished questions, or tasks newer than most models. Indexes built from other boards, judge- or vote-graded boards, vertical specialist boards (legal, finance, medical, spreadsheets) and rarely quoted or trait boards are not used, and a board without a tier is not used either. Cost, token and speed boards were never in.
Pairs: every headline board is paired with every held-out board of the same topic. The model is ranked on both boards among the models measured on both (at least 10 models from three families), so boards that test different fields are put on the same footing. The gap is the headline percentile minus the held-out percentile; the raw score is the plain mean of all the model's gaps. Pairs never cross topics: being strong at coding and weaker at maths is specialisation, not Benchmaxxing.
Coverage and shrinkage: n is the number of distinct headline boards plus held-out boards in the model's pairs, minus one. A model is scored with n ≥ 6 and pairs in at least 2 topics. Every score is pulled toward zero by n / (n + k), where k = 6.0 is estimated from how much score variance falls as n grows (clamped to 6–50). There is no adjustment for how strong the model is.
Jaggedness (Florian, 17 Sep 2026): the published score adds a second part — 0.3 × (the model's within-topic jaggedness − the catalog average, 13.0 points today). Jaggedness is the mean distance between the model's ranks on two boards of the same topic, in percentile points, with the topics weighted by how many independent comparisons each contributes — the unevenness the radar below makes visible. Distance between topics is still never counted. In a simulated catalog where nobody targets benchmarks and all of today's unevenness is assumed to be noise, the gap alone would tag 17 % of models and the two parts together 20 %, the extra tags falling on mid-table models, whose ranks have the most room to spread; that is why the second part is weighted at 0.3. Each model's two parts are printed in its report below.
Tags: three levels on the score as published (one decimal) — light from +3.0, medium from +6.0 and very strong from +12.0. The level follows that score alone: every scored model reaching a threshold carries its tag (CR-77.1, Florian 2026-09-17). Confidence is shown, not used to hide a tag — a tag built on fewer than 10 comparisons, or whose 80 % bootstrap interval (its headline boards and its held-out boards resampled separately, 400 times) reaches below zero, is marked ◔ uncertain and names the reason in its tooltip and in the report.
Known limitations: a positive gap fits benchmark-targeted training, but it also fits a model that is simply weaker at long agent work, which several held-out boards lean toward. Tier labels are editorial judgements about benchmarks, never about models; model names, labs, openness and prices are not inputs. Models covered mostly by one publisher's boards rest on that publisher's held-out tests.
This analysis cannot establish leakage, contamination, or intent. It is a descriptive screening signal that should prompt users to inspect sources, benchmark design and per-axis evidence.