Benchmark Heaven
Price & provider filters · adjusted costs
Global
Applies to price views & model offers; benchmark evidence stays unfiltered
← All models

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

★ featured
Anthropic · released 2026-09-01 · 8 recorded token offers

Top 5 cheapest providers (Adjusted $/task)

Within the active global provider, residency and confidentiality filters.

#ProviderPlatformRaw input $/1MRaw output $/1MAdjusted $/task
1Amazon BedrockOpenRouter$10.00$50.00
2AnthropicOpenRouter$10.00$50.00
3GoogleOpenRouter$10.00$50.00
4AzureOpenRouter$10.00$50.00
5AWS BedrockAWS Bedrock$10.00$50.00

Benchmarks

Composite100.0
Composite evidence2/5
AA Coding · unversioned snapshot 2026-09-1081.6
AA Coding Agent · v1.4 · 2026-09-09
AA Intelligence · unversioned snapshot 2026-09-1053.4
DesignArena Frontend · unversioned snapshot 2026-09-10
DesignArena Full-Stack · unversioned snapshot 2026-09-10
Output speed (t/s)68

MODEL EVIDENCE

Benchmark sheet

16 / 73 registered benchmark versions covered · 18 observations including existing index snapshots.

Compare this model ↗
Composite retains Coding Agent v1.4 · source 2026-09-09

The five Composite inputs remain unchanged. Its Coding Agent input is the median across complete harness results in the retained v1.4 snapshot from 2026-09-09. Artificial Analysis now publishes v1.5, with different components. Explore current v1.5 separately. Dated snapshot labels on other indices identify unversioned source captures, not a verified semantic version.

Profile signals

1 threshold-crossing signal flagged · 17 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually strong CritPt (AA) snapshot-2026-09-10
    Why
    Observed: 0.29714 fraction
    Peer mean 0.03942 · peer sd 0.0762 · n 521 · families 367
    Directed z 3.382 · baseline z 1.882 · gap 1.5 · profile n 16
    0.29714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
    Exact value: 0.297142857142857 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:critpt

Protocol-compatible measured/vendor divergences

No verified protocol-compatible vendor/measured pair is available for this model; agreement cannot be assessed.

Agentic

Agentic: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AA-AnalystAgent

Version snapshot-2026-09-10

Published board

Tests analyst tasks using agentic Python execution across fourteen domains.

0.575 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-AnalystAgent · snapshot-2026-09-10 · Published board
Exact value: 0.575 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:analystAgent
AA-Briefcase

Version snapshot-2026-09-10

Published board

Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

1661.82 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Briefcase · snapshot-2026-09-10 · Published board
Exact value: 1661.82 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:briefcaseBreakdown.overall.elo
GDPval-AA v2

Version 2

Published board

Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

1763.64 Elomeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDPval-AA v2 · 2 · Published board
Exact value: 1763.64 Elo
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:gdpval
Harvey LAB-AA

Version snapshot-2026-09-10

Published board

Tests legal-work deliverables across practice areas on Harvey's private task set.

0.93024 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Harvey LAB-AA · snapshot-2026-09-10 · Published board
Exact value: 0.930235924156897 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:harveyLab
Terminal-Bench v2.1 (AA)

Version 2.1

Published board

Tests terminal-based work on the 89-task verified refresh using Terminus 2.

0.91386 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v2.1 (AA) · 2.1 · Published board
Exact value: 0.913857677902622 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:terminalbenchV21
Terminal-Bench v4.0 (AA)

Version 4.0

Published board

Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

0.5202 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Terminal-Bench v4.0 (AA) · 4.0 · Published board
Exact value: 0.52020202020202 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:terminalbenchV40

Coding

Coding: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
SciCode (AA subproblems)

Version 1.0.1

Published board

Tests scientific Python programming with scientist-annotated background information.

0.63079 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: SciCode (AA subproblems) · 1.0.1 · Published board
Exact value: 0.630787037037037 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:scicode
AA Coding Index

Version snapshot-2026-09-10 (unversioned)

Published board

Published AA index from the retained API snapshot. No verified semantic version was supplied.

81.6 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Coding Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 81.6 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_coding_index:claude-fable-5.1::max
Protocol: inspect · file data/raw/artificialanalysis.json

Knowledge

Knowledge: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
Humanity's Last Exam (AA text-only)

Version snapshot-2026-09-10

Published board

Tests expert-level knowledge on AA's text-only Humanity's Last Exam subset.

0.59129 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: Humanity's Last Exam (AA text-only) · snapshot-2026-09-10 · Published board
Exact value: 0.591288229842447 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:hle
AA-Omniscience Index

Version snapshot-2026-09-10

Published board

Tests factual reliability while rewarding correct answers and penalizing hallucinations.

43.45 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-Omniscience Index · snapshot-2026-09-10 · Published board
Exact value: 43.45 points
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:omniscience

Long-context

Long-context: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
GDP.pdf (AA)

Version snapshot-2026-09-10

Published board

Tests professional reasoning over long PDFs with AA document preparation and grading.

0.262 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GDP.pdf (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.262 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:gdpPdfAllPass
AA-LCR v1.1

Version 1.1

Published board

Tests reasoning across multiple long documents with corrected answer keys and grading.

0.85333 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA-LCR v1.1 · 1.1 · Published board
Exact value: 0.853333333333333 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:lcr
MLCR-AA

Version snapshot-2026-09-10

Published board

Tests medical-record synthesis and reasoning across long, fragmented case documents.

0.71111 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: MLCR-AA · snapshot-2026-09-10 · Published board
Exact value: 0.711111111111111 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:mlcrOverall

Reasoning

Reasoning: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AA Intelligence Index

Version snapshot-2026-09-10 (unversioned)

Published board

Published AA index from the retained API snapshot. No verified semantic version was supplied.

53.4 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AA Intelligence Index · snapshot-2026-09-10 (unversioned) · Published board
Exact value: 53.4 points
Observed: 2026-09-10 · publication date: not recorded
Observation id: legacy:aa_intelligence_index:claude-fable-5.1::max
Protocol: inspect · file data/raw/artificialanalysis.json

Science

Science: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
CritPt (AA)

Version snapshot-2026-09-10

Published board

Tests research-level physics reasoning with Python, symbolic and numerical answers.

0.29714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: CritPt (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.297142857142857 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:critpt
GPQA Diamond (AA)

Version snapshot-2026-09-10

Published board

Tests graduate-level biology, physics and chemistry knowledge on the Diamond subset.

0.93737 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: GPQA Diamond (AA) · snapshot-2026-09-10 · Published board
Exact value: 0.937373737373737 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:gpqa

Tool-use

Tool-use: every observation for this exact model configuration
Benchmark / versionEvaluation groupScore / provenance
AutomationBench-AA

Version 1.0.6

Published board

Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

0.59376 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: AutomationBench-AA · 1.0.6 · Published board
Exact value: 0.5937591715646424 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:automationBenchPartialScore
τ³-Banking (AA)

Version 1.0.1

Published board

Tests banking support agents that retrieve policies and change account state through tools.

0.47216 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
Evidence
Axis: τ³-Banking (AA) · 1.0.1 · Published board
Exact value: 0.472164948453608 fraction
Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:tauBanking
Missing coverage · 59 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

VariantCoding · snapshot 2026-09-10Intelligence · snapshot 2026-09-10
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)81.653.4
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)80.753.2
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)79.151.2
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)77.149.1
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)75.247.0

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $10.00 / 1M
Cached input $0.250 / 1M
Cache write $12.50 / 1M
Output $50.00 / 1M
Status GA

Token offers by platform — Adjusted $/task

Adjusted costs are modeled USD per task: AA output tokens × OpenRouter usage I/O (Chutes global fallback). These are general usage and benchmark proxies for coding-agent work. Missing AA data assumes 1,000 output tokens/task; unknown cache hit assumes 0%; unmeasured additional cache writes assume 0 tokens. Click any underlined price for exact inputs, dates and assumptions. Raw list prices use the selected fixed input/output blend, in USD per million tokens.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

Amazon Bedrockglobalamazon-bedrock$10.00 raw in $/1M$50.00 raw out $/1M
Anthropicglobalanthropic$10.00 raw in $/1M$50.00 raw out $/1M
Googleglobalgoogle-vertex/global$10.00 raw in $/1M$50.00 raw out $/1M
Azureglobalazure$10.00 raw in $/1M$50.00 raw out $/1M

AWS Bedrock (1)

AWS Bedrockglobal$10.00 raw in $/1M$50.00 raw out $/1M

Google Vertex AI (2)

Google Vertex AIglobal$10.00 raw in $/1M$50.00 raw out $/1M
Google Vertex AIeuEU$11.00 raw in $/1M$55.00 raw out $/1M

Anthropic (1)

Anthropicglobal$10.00 raw in $/1M$50.00 raw out $/1M
Benchmark Heaven — Model benchmarks & costs