← All models

Claude Opus 4.7 (Adaptive Reasoning, Max Effort)

deprecated by benchmark source
Anthropic · released 2026-04-16 · 12 offers

Context 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Google
2Amazon Bedrock
3Google Vertex AI
4Azure
5AWS Bedrock

Composite

5 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

82.2

dominance-adjusted from 84.1: a better-measured model that is at least as good on each of these inputs ranks above it

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 87Coding Agent v1.4: percentile 36AA Intelligence: percentile 95Epoch ECI: percentile 74Software ECI: percentile 68

AA Coding 73.6Coding Agent v1.4 51.6AA Intelligence 40.7AA Agentic Epoch ECI 156.3Software ECI 158.1DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

39 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

  • AA-AnalystAgent6843.8%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase691,252 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2771,396 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • ITBench-AA9146.7%

    ITBench-AA published 2026-09-10 · Published board — Tests root-cause diagnosis from offline Kubernetes incident snapshots.

  • Terminal-Bench Hard (AA)9551.5%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)8583.1%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • ApprenticeBench CUA (NeoCognition)157.0%

    ApprenticeBench CUA (NeoCognition) published 2026-09-14 · Claude Code — A computer-use agent operates the Odoo ERP through its screens across 100 sequential accounts-payable tasks with diminishing mentoring.

  • ApprenticeBench API (NeoCognition)814.0%

    ApprenticeBench API (NeoCognition) published 2026-09-14 · Claude Code — An agent completes the same 100 sequential accounts-payable tasks in the Odoo ERP with diminishing mentoring, using dedicated application API calls instead of the screen interface.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)536.7%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • EBR-bench (Earthborne Rangers, Epoch AI)3819.0%

    EBR-bench (Earthborne Rangers, Epoch AI) published 2026-09-18 · Published board — Learning-capability test: models repeatedly play the obscure campaign board game Earthborne Rangers with note-taking, and the score measures whether results improve across playthroughs.

Coding

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.18578.7%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • FrontierMath Tiers 1–3 v2 (Epoch AI)6470.2%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)4531.7%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • Chess Puzzles (Epoch AI)177.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)4528.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index9540.7

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Safety/Alignment

Science

Tool-use

Vision

Missing coverage · 104 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)73.640.7
Claude Opus 4.7 (Non-reasoning, High Effort)30.9
Claude Opus 4.7 (Adaptive Reasoning, Medium Effort)

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $5.00 / 1M
Cached input $0.500 / 1M
Cache write $6.25 / 1M
Output $25.00 / 1M
Status GA

Legacy annual Pro/Pro+ request billing only.

Multiplier 27.00×
Effective cost $1.080 / request

Legacy annual-plan multiplier.

Token offers by platform · 9 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

9 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (6)

Google
Google
Amazon Bedrock
Azure
Google
Amazon Bedrock

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock
Claude Opus 4.7 (Adaptive Reasoning, Max Effort) — benchmarks & cost | Benchmark Heaven