← All models

Claude Opus 4.5 (Reasoning)

deprecated by benchmark source
Anthropic · released 2025-11-24 · 10 offers

Context 200K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Azure
2Google
3Amazon Bedrock
4Google Vertex AI
5AWS Bedrock

Composite

5 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

69.9

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 86Epoch ECI: percentile 54Software ECI: percentile 51DesignArena: percentile 36

AA Coding Coding Agent v1.4 AA Intelligence 29.1AA Agentic Epoch ECI 150.1Software ECI 152.7DesignArena 1174/1166

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

14 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

Coding

  • DesignArena Web Apps (agentic)361,174 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack411,166 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.18277.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v20259391.3%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Science

Tool-use

Vision

Unusual results

2 threshold-crossing signals flagged · 17 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak GPQA Diamond (OpenRouter run) — measured cost per task published 2026-09-15
    Why
    Observed: 0.84288 USD
    Peer mean 0.07408 · peer sd 0.13073 · n 124 · families 124
    Directed z -5.881 · baseline z 0.671 · gap -6.552 · profile n 16
    0.84288 USDmeasured· 1188 tasksobserved 2026-09-15openrouter.ai
    Evidence
    Axis: GPQA Diamond (OpenRouter run) — measured cost per task · published 2026-09-15 · Published board
    Exact value: 0.8428752398989899 USD
    Tasks evaluated: 1188
    Observed: 2026-09-15T23:29:03.209Z · publication date: not recorded
    Observation id: openrouter-cost:anthropic/claude-4.5-opus-20251124|gpqa_diamond
  • unusually weak τ²-Bench Airline (OpenRouter run) — measured cost per task published 2026-09-15
    Why
    Observed: 0.53354 USD
    Peer mean 0.14769 · peer sd 0.2495 · n 118 · families 118
    Directed z -1.546 · baseline z 0.4 · gap -1.946 · profile n 16
    0.53354 USDmeasured· 300 tasksobserved 2026-09-15openrouter.ai
    Evidence
    Axis: τ²-Bench Airline (OpenRouter run) — measured cost per task · published 2026-09-15 · Published board
    Exact value: 0.5335405666666666 USD
    Tasks evaluated: 300
    Observed: 2026-09-15T23:29:03.209Z · publication date: not recorded
    Observation id: openrouter-cost:anthropic/claude-4.5-opus-20251124|tau_bench_verified_airline
Missing coverage · 128 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude Opus 4.5 (Reasoning)29.1
Claude Opus 4.5 (Non-reasoning)23.7
Token offers by platform · 7 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

7 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

Azure
Google
Amazon Bedrock
Amazon Bedrock

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock
Claude Opus 4.5 (Reasoning) — benchmarks & cost | Benchmark Heaven