Benchmark Heaven
Price & provider filters · adjusted costs
Global
Applies to price views & model offers; benchmark evidence stays unfiltered

MODEL COMPARISON

Find strengths. Understand tradeoffs.

Compare up to four exact model configurations across the full benchmark collection, with the evidence beside every score.

Composite retains Coding Agent v1.4 · source 2026-09-09

The five Composite inputs remain unchanged. Its Coding Agent input is the median across complete harness results in the retained v1.4 snapshot from 2026-09-09. Artificial Analysis now publishes v1.5, with different components. Explore current v1.5 separately. Dated snapshot labels on other indices identify unversioned source captures, not a verified semantic version.

BUILD YOUR COMPARISON

Choose up to four models

2 / 4 selected

Exact reasoning configurations. All catalog models are available; price and provider filters do not hide benchmark evidence.

2 models selected. 23 evaluation rows in the full comparison.
Radar axes · 6 / 8 selected

Choose 3–8 axes. Every option names one benchmark version and evaluation group. Unmeasured selections remain empty.

RELATIVE PERFORMANCE

Benchmark radar

Independent measurements

Each version uses its measured catalog minimum → 0 and maximum → 100; lower-is-better axes are reversed. Missing scores leave gaps. A zero is a measured minimum, not missing.

  • A · GPT-5.6 Sol (high)
  • B · Claude Sonnet 5 (Adaptive Reasoning, High Effort)

Scroll the chart horizontally. All values are also in the table below.

2550751001. GPQA Diamondv: 2026-09-10 snapshot2. Humanity's Last Examv: 2026-09-10 snapshot3. AA-LCR v1.1v: 1.14. SciCodev: 1.0.15. AA Coding Indexv: 2026-09-10 snapshot6. AA Intelligence Indexv: 2026-09-10 snapshotGPT-5.6 Sol (high) — GPQA Diamond (AA), snapshot-2026-09-10: 0.928282828282828 fraction; normalized 96.028GPT-5.6 Sol (high) — Humanity's Last Exam (AA text-only), snapshot-2026-09-10: 0.46014828544949 fraction; normalized 77.414GPT-5.6 Sol (high) — AA-LCR v1.1, 1.1: 0.816666666666667 fraction; normalized 92.105GPT-5.6 Sol (high) — SciCode (AA subproblems), 1.0.1: 0.577546296296296 fraction; normalized 89.074GPT-5.6 Sol (high) — AA Coding Index, snapshot-2026-09-10 (unversioned): 77.2 points; normalized 94.608GPT-5.6 Sol (high) — AA Intelligence Index, snapshot-2026-09-10 (unversioned): 42.5 points; normalized 78.024Claude Sonnet 5 (Adaptive Reasoning, High Effort) — Humanity's Last Exam (AA text-only), snapshot-2026-09-10: 0.357275254865616 fraction; normalized 59.697Claude Sonnet 5 (Adaptive Reasoning, High Effort) — AA-LCR v1.1, 1.1: 0.766666666666667 fraction; normalized 86.466Claude Sonnet 5 (Adaptive Reasoning, High Effort) — SciCode (AA subproblems), 1.0.1: 0.542824074074074 fraction; normalized 81.948
Radar values · normalized / 100, rounded to 3 decimals (exact native score). Missing or uninformative axes are never filled.
Axis / versionA · GPT-5.6 Sol (high)B · Claude Sonnet 5 (Adaptive Reasoning, High Effort)Normalization range
1. GPQA Diamond (AA)
Version snapshot-2026-09-10 · Published board
96.028 (0.928282828282828 fraction)No measured result612 measured configurations · 0.097979797979798–0.962626262626263 fraction · higher better
2. Humanity's Last Exam (AA text-only)
Version snapshot-2026-09-10 · Published board
77.414 (0.46014828544949 fraction)59.697 (0.357275254865616 fraction)607 measured configurations · 0.0106580166821131–0.591288229842447 fraction · higher better
3. AA-LCR v1.1
Version 1.1 · Published board
92.105 (0.816666666666667 fraction)86.466 (0.766666666666667 fraction)517 measured configurations · 0–0.886666666666667 fraction · higher better
4. SciCode (AA subproblems)
Version 1.0.1 · Published board
89.074 (0.577546296296296 fraction)81.948 (0.542824074074074 fraction)167 measured configurations · 0.143518518518519–0.630787037037037 fraction · higher better
5. AA Coding Index
Version snapshot-2026-09-10 (unversioned) · Published board
94.608 (77.2 points)No measured result256 measured configurations · 0–81.6 points · higher better
6. AA Intelligence Index
Version snapshot-2026-09-10 (unversioned) · Published board
78.024 (42.5 points)No measured result633 measured configurations · 3.8–53.4 points · higher better
How to read this chart

Formula: 100 × (value − minimum) / (maximum − minimum), or 100 minus that value for lower-is-better axes. Peers are independently measured, exactly matched catalog configurations; repeated source identities count once. Fewer than two peers, a constant range or unknown direction suppresses plotting. DesignArena results below 200 battles are excluded. Axes show version and published harness/configuration. Model prompts and test conditions can still differ; inspect the source before interpreting small differences. Ranges stay fixed when model selection changes. This chart does not change Composite.

EVERY COLLECTED BENCHMARK

Full benchmark comparison

Native units; versions and evaluation groups remain separate. Measured results take priority over vendor claims, then the latest observation. Small differences are not evidence of significance.

All matching benchmark versions and model scores with evidence
Benchmark / versionA · GPT-5.6 Sol (high)B · Claude Sonnet 5 (Adaptive Reasoning, High Effort)
Agentic
AA-Briefcase
Version snapshot-2026-09-10
Published board
Elo · higher better
1360.88 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Briefcase · snapshot-2026-09-10 · Published board
Exact value: 1360.88 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:briefcaseBreakdown.overall.elo
1177.84 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Briefcase · snapshot-2026-09-10 · Published board
Exact value: 1177.84 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:briefcaseBreakdown.overall.elo
GDPval-AA v2
Version 2
Published board
Elo · higher better
1524.01 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDPval-AA v2 · 2 · Published board
Exact value: 1524.01 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:gdpval
1308.45 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDPval-AA v2 · 2 · Published board
Exact value: 1308.45 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:gdpval
Terminal-Bench Hard (AA)
Version 74221fb
Published board
fraction · higher better
0.62121 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench Hard (AA) · 74221fb · Published board
Exact value: 0.621212121212121 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:terminalbenchHard
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Terminal-Bench v2.1 (AA)
Version 2.1
Published board
fraction · higher better
0.87266 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v2.1 (AA) · 2.1 · Published board
Exact value: 0.872659176029963 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:terminalbenchV21
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Terminal-Bench v4.0 (AA)
Version 4.0
Published board
fraction · higher better
0.20707 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v4.0 (AA) · 4.0 · Published board
Exact value: 0.207070707070707 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:terminalbenchV40
0.05051 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v4.0 (AA) · 4.0 · Published board
Exact value: 0.0505050505050505 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:terminalbenchV40
Coding
Artificial Analysis Coding Agent Index v1.4
Version 1.4
Codex
fraction · higher better
0.64116 fractionmeasuredobserved 2026-09-09artificialanalysis.ai
Evidence
Axis: Artificial Analysis Coding Agent Index v1.4 · 1.4 · Codex
Exact value: 0.6411644486641197 fraction
Observed: 2026-09-09 · publication date: not recorded
Observation id: aa-coding:1.4:35
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

SciCode (AA subproblems)
Version 1.0.1
Published board
fraction · higher better
0.57755 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: SciCode (AA subproblems) · 1.0.1 · Published board
Exact value: 0.577546296296296 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:scicode
0.54282 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: SciCode (AA subproblems) · 1.0.1 · Published board
Exact value: 0.542824074074074 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:scicode
AA Coding Index
Version snapshot-2026-09-10 (unversioned)
Published board
points · higher better
77.2 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Coding Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 77.2 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_coding_index:gpt-5.6-sol::high
Protocol: inspect · file data/raw/artificialanalysis.json
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

DesignArena Frontend
Version snapshot-2026-09-10 (unversioned)
Published board
Elo · higher better
1209 Elomeasured2255 battlesobserved 2026-09-10designarena.ai
Evidence
Axis: DesignArena Frontend · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 1209 Elo
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:frontend:gpt-5.6-sol::high
Protocol: inspect · file data/raw/designarena.json
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

DesignArena Full-Stack
Version snapshot-2026-09-10 (unversioned)
Published board
Elo · higher better
1168 Elomeasured1611 battlesobserved 2026-09-10designarena.ai
Evidence
Axis: DesignArena Full-Stack · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 1168 Elo
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:fullstack:gpt-5.6-sol::high
Protocol: inspect · file data/raw/designarena.json
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Instruction-following
IFBench (AA single-turn)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.69184 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: IFBench (AA single-turn) · snapshot-2026-09-10 · Published board
Exact value: 0.691836734693878 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:ifbench
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Knowledge
Humanity's Last Exam (AA text-only)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.46015 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Humanity's Last Exam (AA text-only) · snapshot-2026-09-10 · Published board
Exact value: 0.46014828544949 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:hle
0.35728 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Humanity's Last Exam (AA text-only) · snapshot-2026-09-10 · Published board
Exact value: 0.357275254865616 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:hle
AA-Omniscience Index
Version snapshot-2026-09-10
Published board
points · higher better
20.36667 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Omniscience Index · snapshot-2026-09-10 · Published board
Exact value: 20.3666666666667 points
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:omniscience
-3.68333 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Omniscience Index · snapshot-2026-09-10 · Published board
Exact value: -3.68333333333333 points
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:omniscience
Long-context
GDP.pdf (AA)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.278 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDP.pdf (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.278 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:gdpPdfAllPass
0.094 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDP.pdf (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.094 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:gdpPdfAllPass
AA-LCR v1.1
Version 1.1
Published board
fraction · higher better
0.81667 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-LCR v1.1 · 1.1 · Published board
Exact value: 0.816666666666667 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:lcr
0.76667 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-LCR v1.1 · 1.1 · Published board
Exact value: 0.766666666666667 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:lcr
MLCR-AA
Version snapshot-2026-09-10
Published board
fraction · higher better
0.17222 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: MLCR-AA · snapshot-2026-09-10 · Published board
Exact value: 0.172222222222222 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:mlcrOverall
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Reasoning
AA Intelligence Index
Version snapshot-2026-09-10 (unversioned)
Published board
points · higher better
42.5 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Intelligence Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 42.5 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_intelligence_index:gpt-5.6-sol::high
Protocol: inspect · file data/raw/artificialanalysis.json
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Science
CritPt (AA)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.25714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.257142857142857 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:critpt
0.15143 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.151428571428571 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:critpt
GPQA Diamond (AA)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.92828 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GPQA Diamond (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.928282828282828 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:gpqa
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Tool-use
AutomationBench-AA
Version 1.0.6
Published board
fraction · higher better
0.55312 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AutomationBench-AA · 1.0.6 · Published board
Exact value: 0.5531220357393729 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:automationBenchPartialScore
0.32097 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AutomationBench-AA · 1.0.6 · Published board
Exact value: 0.32096638814380396 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:ba0224cf-0351-4f56-8508-b3f1a740ae4a:automationBenchPartialScore
τ²-Bench Telecom (AA)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.83333 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: τ²-Bench Telecom (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.833333333333333 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:tau2
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

τ³-Banking (AA)
Version 1.0.1
Published board
fraction · higher better
0.36701 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: τ³-Banking (AA) · 1.0.1 · Published board
Exact value: 0.36701030927835 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:tauBanking
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

Vision
MMMU Pro (AA)
Version snapshot-2026-09-10
Published board
fraction · higher better
0.8185 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: MMMU Pro (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.81849710982659 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:mmmuPro
Unknown / no published match

No observation is attached to this exact catalog configuration. This does not establish that the model was never tested.

GPT-5.6 Sol (high)

Profile signals

1 threshold-crossing signal flagged · 21 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually strong CritPt (AA) snapshot-2026-09-10
    Why
    Observed: 0.25714 fraction
    Peer mean 0.03942 · peer sd 0.0762 · n 521 · families 367
    Directed z 2.857 · baseline z 1.225 · gap 1.632 · profile n 20
    0.25714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
    Exact value: 0.257142857142857 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:8afc250d-b538-45a2-812a-4605f4ffd87e:critpt

Protocol-compatible measured/vendor divergences

No verified protocol-compatible vendor/measured pair is available for this model; agreement cannot be assessed.

Claude Sonnet 5 (Adaptive Reasoning, High Effort)

Profile signals

No threshold-crossing signal · 10 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

Protocol-compatible measured/vendor divergences

No verified protocol-compatible vendor/measured pair is available for this model; agreement cannot be assessed.

Price & provider comparison · separate two-model selection

This price view has its own A/B selection and uses the global price and provider filters. Benchmark exploration above includes the full catalog.

Loading price and provider evidence…