← All models

Claude 4.5 Sonnet (Reasoning)

deprecated by benchmark source
Anthropic · released 2025-09-29 · 11 offers

Context 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Amazon Bedrock
2Google
3Google Vertex AI
4AWS Bedrock
5Azure

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

51.2

includes −2.5 for its Benchmaxxing signal (from 53.8; why, switch off in Options)

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 59AA Intelligence: percentile 72AA Agentic: percentile 51Epoch ECI: percentile 44Software ECI: percentile 31DesignArena: percentile 11

AA Coding 52.1Coding Agent v1.4 AA Intelligence 21.2AA Agentic 17.5Epoch ECI 146.8Software ECI 147.7DesignArena 1075/1060

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

23 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-Briefcase32712 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v247989 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench Hard (AA)8235.6%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)5555.8%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)180.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • τ²-Bench Airline (OpenRouter run)6173.0%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index5117.5

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.14645.7%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • AA Coding Index5952.1

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)91,075 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack171,060 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)315.2%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.17272.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v20258688.0%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Science

Tool-use

  • AutomationBench-AA v1.0.64214.1%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • τ²-Bench Telecom (AA)6878.1%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.16024.5%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Unusual results

2 threshold-crossing signals flagged · 26 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak GPQA Diamond (OpenRouter run) — measured cost per task published 2026-09-15
    Why
    Observed: 0.31611 USD
    Peer mean 0.07408 · peer sd 0.13073 · n 124 · families 124
    Directed z -1.851 · baseline z 0.122 · gap -1.973 · profile n 25
    0.31611 USDmeasured· 1386 tasksobserved 2026-09-15openrouter.ai
    Evidence
    Axis: GPQA Diamond (OpenRouter run) — measured cost per task · published 2026-09-15 · Published board
    Exact value: 0.31610685642135644 USD
    Tasks evaluated: 1386
    Observed: 2026-09-15T23:29:03.209Z · publication date: not recorded
    Observation id: openrouter-cost:anthropic/claude-4.5-sonnet-20250929|gpqa_diamond
  • unusually weak DesignArena Web Apps (agentic) published 2026-09-19
    Why
    Observed: 1075 Elo
    Peer mean 1198.82222 · peer sd 77.91656 · n 45 · families 44
    Directed z -1.589 · baseline z 0.111 · gap -1.701 · profile n 25
    1075 Elomeasured1503 battlesobserved 2026-09-19designarena.ai
    Evidence
    Axis: DesignArena Web Apps (agentic) · published 2026-09-19 · Published board
    Exact value: 1075 Elo
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:frontend:claude-sonnet-4.5::reasoning
    Protocol: inspect · file data/raw/designarena.json
Missing coverage · 118 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude 4.5 Sonnet (Reasoning)52.121.2
Claude 4.5 Sonnet (Non-reasoning)19.3
Token offers by platform · 8 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (5)

Amazon Bedrock
Google
Amazon Bedrock
Google
Azure

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock
Claude 4.5 Sonnet (Reasoning) — benchmarks & cost | Benchmark Heaven