← All models

Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)

★ featured
Anthropic · released 2026-09-22 · 12 offers

Context 1M tokens

Top 3 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Amazon Bedrock
2Azure
3Google

Composite

1 of 7 inputs6 radar axes: DesignArena's two boards share one

100.0

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 100

AA Coding Coding Agent v1.4 AA Intelligence 57.6AA Agentic Epoch ECI Software ECI DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

16 of 265 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

16 of 17 values are Anthropic's own claims (†), not independent measurements; a matching independent result replaces a claim as soon as one exists.

Compare this model →

Agentic

  • Terminal-Bench 4.0 v4.0developer's claim66.4% self-reported by the developer

    Terminal-Bench 4.0 v4.0 · Published board — Terminal-Bench 4.0 result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

  • GDPval-AA v2.1developer's claim1,846 Elo self-reported by the developer

    GDPval-AA v2.1 v2.1 · Published board — GDPval-AA v2.1 result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic states the evaluation was run independently by Artificial Analysis; the value is published by Anthropic and kept as the vendor's claim, not as an independent measurement.

  • AutomationBenchdeveloper's claim40.0% self-reported by the developer

    AutomationBench published 2026-09-22 · Published board — AutomationBench result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic states the run was performed and reported by Zapier; the value is published by Anthropic and kept as the vendor's claim, not as an independent measurement.

  • OSWorld 2.0 (partial score) v2.0developer's claim81.8% self-reported by the developer

    OSWorld 2.0 (partial score) v2.0 · Published board — OSWorld 2.0 (partial score) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

  • OSWorld 2.0 (strict pass rate) v2.0developer's claim48.7% self-reported by the developer

    OSWorld 2.0 (strict pass rate) v2.0 · Published board — OSWorld 2.0 (strict pass rate) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

  • AA-Briefcase v1.1developer's claim1,822 Elo self-reported by the developer

    AA-Briefcase v1.1 v1.1 · Published board — AA-Briefcase v1.1 result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic states the evaluation was run independently by Artificial Analysis; the value is published by Anthropic and kept as the vendor's claim, not as an independent measurement.

Coding

  • FrontierCode v1.1 (Main)developer's claim54.4% self-reported by the developer

    FrontierCode v1.1 (Main) v1.1 · Published board — FrontierCode v1.1 (Main) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

  • CursorBench 4.0 v4.0developer's claim57.8% self-reported by the developer

    CursorBench 4.0 v4.0 · Published board — CursorBench 4.0 result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

  • SWE-bench Prodeveloper's claim89.9% self-reported by the developer

    SWE-bench Pro published 2026-09-22 · Published board — SWE-bench Pro result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

  • SWE-bench Multilingualdeveloper's claim93.9% self-reported by the developer

    SWE-bench Multilingual published 2026-09-22 · Published board — SWE-bench Multilingual result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

  • SWE-bench Multimodaldeveloper's claim61.4% self-reported by the developer

    SWE-bench Multimodal published 2026-09-22 · Published board — SWE-bench Multimodal result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

Knowledge

  • Humanity's Last Exam (with tools)developer's claim67.7% self-reported by the developer

    Humanity's Last Exam (with tools) published 2026-09-22 · Published board — Humanity's Last Exam (with tools) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

  • Humanity's Last Exam (no tools)developer's claim64.4% self-reported by the developer

    Humanity's Last Exam (no tools) published 2026-09-22 · Published board — Humanity's Last Exam (no tools) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

  • HealthBench Professionaldeveloper's claim65.6% self-reported by the developer

    HealthBench Professional published 2026-09-22 · Published board — HealthBench Professional result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 System Card (PDF, text layer); Anthropic's own run, not an independent measurement.

Reasoning

Science

  • Terminal-Bench-Science 0.1 v0.1developer's claim58.7% self-reported by the developer

    Terminal-Bench-Science 0.1 v0.1 · Published board — Terminal-Bench-Science 0.1 result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

Vision

  • Chartography (with tools)developer's claim89.0% self-reported by the developer

    Chartography (with tools) published 2026-09-22 · Published board — Chartography (with tools) result reported by Anthropic for Claude Opus 5.5 in its 2026-09-22 Claude Opus 5.5 launch post; Anthropic's own run, not an independent measurement.

Missing coverage · 253 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-22 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)57.6
Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)56.0
Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)53.6
Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)51.2
Claude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)42.3
Token offers by platform · 8 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (8)

Amazon Bedrock
Azure
Google
Amazon Bedrock
Amazon Bedrock
Azure
Google
Google
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) — benchmarks & cost | Benchmark Heaven