← All models

Claude Opus 4.6 (Adaptive Reasoning, Max Effort)

deprecated by benchmark source
Anthropic · released 2026-02-05 · 11 offers

Context 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Amazon Bedrock
2Google
3Google Vertex AI
4Azure
5AWS Bedrock

Composite

5 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

78.6

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 89Epoch ECI: percentile 67Software ECI: percentile 66DesignArena: percentile 52

AA Coding Coding Agent v1.4 AA Intelligence 31.9AA Agentic Epoch ECI 155.3Software ECI 156.4DesignArena 1219/1228

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

23 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

Coding

  • DesignArena Web Apps (agentic)581,219 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack671,228 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.18378.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • FrontierMath Tiers 1–3 v2 (Epoch AI)5366.0%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)2726.8%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • Chess Puzzles (Epoch AI)2714.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)3525.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index8931.9

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

Vision

Unusual results

2 threshold-crossing signals flagged · 23 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak GPQA Diamond (OpenRouter run) — measured cost per task published 2026-09-15
    Why
    Observed: 0.74594 USD
    Peer mean 0.07408 · peer sd 0.13073 · n 124 · families 124
    Directed z -5.139 · baseline z 0.552 · gap -5.692 · profile n 22
    0.74594 USDmeasured· 1188 tasksobserved 2026-09-15openrouter.ai
    Evidence
    Axis: GPQA Diamond (OpenRouter run) — measured cost per task · published 2026-09-15 · Published board
    Exact value: 0.7459393897306398 USD
    Tasks evaluated: 1188
    Observed: 2026-09-15T23:29:03.209Z · publication date: not recorded
    Observation id: openrouter-cost:anthropic/claude-4.6-opus-20260205|gpqa_diamond
  • unusually strong Humanity's Last Exam (AA text-only) published 2026-09-10
    Why
    Observed: 0.39944 fraction
    Peer mean 0.15283 · peer sd 0.14006 · n 608 · families 449
    Directed z 1.761 · baseline z 0.239 · gap 1.522 · profile n 22
    0.39944 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Humanity's Last Exam (AA text-only) · published 2026-09-10 · Published board
    Exact value: 0.399443929564411 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:53c98840-47af-49aa-94e6-469fb17e9a1b:hle
Missing coverage · 119 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)31.9
Claude Opus 4.6 (Non-reasoning, High Effort)26.4
Token offers by platform · 8 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

Amazon Bedrock
Google
Azure
Google

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock

T-Systems LLM Hub (1)

T-Systems LLM Hub
Claude Opus 4.6 (Adaptive Reasoning, Max Effort) — benchmarks & cost | Benchmark Heaven