← All models

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

★ featured
Anthropic · released 2026-09-01 · 8 offers

Output 71 tokens/sFirst token 146 sContext 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Amazon Bedrock
2Google
3Azure
4AWS Bedrock
5Google Vertex AI

Composite

7 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

98.5

includes −0.5 for its Benchmaxxing signal (from 99.0; why, switch off in Options)

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 100Coding Agent v1.4: percentile 100AA Intelligence: percentile 100Epoch ECI: percentile 97Software ECI: percentile 99DesignArena: percentile 96

AA Coding 81.6Coding Agent v1.4 70.4AA Intelligence 53.4AA Agentic Epoch ECI 164.5Software ECI 167.4DesignArena 1336/1336

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

52 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-AnalystAgent9757.5%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase1001,662 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v21001,764 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA8893.0%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • Terminal-Bench v2.1 (AA)10091.4%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)9752.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • ApprenticeBench CUA (NeoCognition)10072.0%

    ApprenticeBench CUA (NeoCognition) published 2026-09-14 · Claude Code — A computer-use agent operates the Odoo ERP through its screens across 100 sequential accounts-payable tasks with diminishing mentoring.

  • ApprenticeBench API (NeoCognition)10070.0%

    ApprenticeBench API (NeoCognition) published 2026-09-14 · Claude Code — An agent completes the same 100 sequential accounts-payable tasks in the Odoo ERP with diminishing mentoring, using dedicated application API calls instead of the screen interface.

  • Vals Index v2 (Vals AI)10068.8%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)8358.9%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)536.7%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • EBR-bench (Earthborne Rangers, Epoch AI)9257.1%

    EBR-bench (Earthborne Rangers, Epoch AI) published 2026-09-18 · Published board — Learning-capability test: models repeatedly play the obscure campaign board game Earthborne Rangers with note-taking, and the score measures whether results improve across playthroughs.

Coding

  • Artificial Analysis Coding Agent Index v1.410070.4%

    Artificial Analysis Coding Agent Index v1.4 v1.4 · Claude Code — The retained AA Coding Agent Index measures coding-agent systems using the earlier three-component implementation.

  • Artificial Analysis Coding Agent Index v1.510062.2%

    Artificial Analysis Coding Agent Index v1.5 v1.5 · Claude Code — Measures coding-agent systems on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA.

  • SciCode (AA subproblems) v1.0.110063.1%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • SWE-bench Multilingualno percentile89.1%

    SWE-bench Multilingual published 2026-09-10 · Published board — A 300-instance SWE-bench variant with tasks from 42 repositories across 9 programming languages.

  • SWE-bench Multimodalno percentile54.7%

    SWE-bench Multimodal published 2026-09-10 · Published board — A 480-instance SWE-bench variant whose issue descriptions include visual elements.

  • CursorBench 4.0 (Cursor) v4.0no percentile51.8%

    CursorBench 4.0 (Cursor) v4.0 · Published board — Agent evaluation on ambiguous, multi-file tasks drawn from real Cursor sessions.

  • Vibe Code Bench (Vals Index v2)9790.3%

    Vibe Code Bench (Vals Index v2) v2 · Published board — End-to-end app-building tasks.

  • Code Migration (Vals Index v2 subset)9760.0%

    Code Migration (Vals Index v2 subset) v2 · Published board — Porting projects to another language, including COBOL modernization.

  • FrontierCode 1.1 Main (Cognition) v1.1no percentile50.3%

    FrontierCode 1.1 Main (Cognition) v1.1 · claude-code — Whether a maintainer would merge the agent's pull request, on tasks crafted by open-source maintainers and graded with tests, rubrics and verifiers.

  • VulcanBench Frontier v4no percentile91.8%

    VulcanBench Frontier v4 v4 · Claude Code — 23 behavioural-reconstruction tasks: the model must repair a replacement implementation of a legacy program until hidden tests confirm it reproduces the program’s real drift from its written spec; the combined score weights functional correctness 50%, lint/complexity 8.5%, security 8.5% and judged Code quality 33%.

  • KernelBench-CUDA: GLM-5.2 Fused MoE (RTX PRO 6000) vrtx-pro-60006710.2

    KernelBench-CUDA: GLM-5.2 Fused MoE (RTX PRO 6000) vrtx-pro-6000 · Published board — Fuse GLM-5.2’s MoE forward pass (256+shared experts, top-8 routing) as one CUDA kernel over a frozen shape sweep; Triton, vLLM and other DSLs are banned — CUDA/PTX/CUTLASS only.

  • KernelBench-CUDA: DeepSeek NSA (RTX PRO 6000) vrtx-pro-6000100106.3

    KernelBench-CUDA: DeepSeek NSA (RTX PRO 6000) vrtx-pro-6000 · Published board — Implement DeepSeek’s Native Sparse Attention block (block top-n routing + sparse attention) as a CUDA kernel over a frozen shape sweep.

  • KernelBench-CUDA: MegaQwen Decode (RTX PRO 6000) vrtx-pro-6000806.43

    KernelBench-CUDA: MegaQwen Decode (RTX PRO 6000) vrtx-pro-6000 · Published board — Improve the known MegaQwen CUDA megakernel geometry for decode-only token throughput at context lengths 2k–128k; prefill is untimed.

  • KernelBench-CUDA: Grid MinGRU SPS (RTX PRO 6000) vrtx-pro-60006770.9

    KernelBench-CUDA: Grid MinGRU SPS (RTX PRO 6000) vrtx-pro-6000 · Published board — Non-LLM RL simulation: maximise simulation steps per second for a grid world with a MinGRU agent (roofline anchored at 150M peak SPS); fusion optional.

  • FrontierSWE v210056.3%

    FrontierSWE v2 v2 · proximus — 34 hand-written real-world software tasks: each model attempts every task in its own CLI agent harness for 5 trials per task under a 20-hour budget, scored by the site’s own review protocol; the headline value is mean@5 in percent with best@5/worst@5 bounds.

  • AA Coding Index10081.6

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Efficiency

Knowledge

Long-context

  • GDP.pdf (AA)9126.2%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.110085.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA10071.1%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • ArXivMath 06/2026 (MathArena) v2026-069091.0%

    ArXivMath 06/2026 (MathArena) v2026-06 · Published board — Research-level math problems with a checkable final answer, drawn from arXiv papers submitted in June 2026, so they postdate most training data.

  • BrokenArXiv 06/2026 (MathArena) v2026-068084.7%

    BrokenArXiv 06/2026 (MathArena) v2026-06 · Published board — Plausible but false proof statements taken from June 2026 arXiv papers; a model scores by refusing to prove them and saying the statement is false as written.

  • FrontierMath Tiers 1–3 v2 (Epoch AI)9790.2%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)8387.8%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • Chess Puzzles (Epoch AI)8847.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)9358.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index10053.4

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Safety/Alignment

Science

Tool-use

  • AutomationBench-AA v1.0.69059.4%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • τ³-Banking (AA) v1.0.19647.2%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

  • Excel Modeling Benchmark (Vals Index v2)10076.7%

    Excel Modeling Benchmark (Vals Index v2) v2 · Published board — Building and editing financial models in spreadsheets.

Unusual results

5 threshold-crossing signals flagged · 33 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak Vals Index v2 cost per test (Vals AI)
    Why
    Observed: 28.91624 USD
    Peer mean 7.03053 · peer sd 7.84824 · n 31 · families 30
    Directed z -2.789 · baseline z 1.605 · gap -4.393 · profile n 32
    28.91624 USDmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Vals Index v2 cost per test (Vals AI) · v2 · Published board
    Exact value: 28.916238 USD
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:04c8469c757a0176ac395d99
  • unusually strong CritPt (AA) published 2026-09-10
    Why
    Observed: 0.29714 fraction
    Peer mean 0.03928 · peer sd 0.07596 · n 522 · families 369
    Directed z 3.395 · baseline z 1.412 · gap 1.983 · profile n 32
    0.29714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · published 2026-09-10 · Published board
    Exact value: 0.297142857142857 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:critpt
  • unusually strong AA Intelligence Index published 2026-09-19
    Why
    Observed: 53.4 points
    Peer mean 15.94317 · peer sd 11.48893 · n 644 · families 483
    Directed z 3.26 · baseline z 1.416 · gap 1.844 · profile n 32
    53.4 pointsmeasuredobserved 2026-09-19artificialanalysis.ai
    Evidence
    Axis: AA Intelligence Index · published 2026-09-19 · Published board
    Exact value: 53.4 points
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:aa_intelligence_index:claude-fable-5.1::max
    Protocol: inspect · file data/raw/artificialanalysis.json
  • unusually strong Humanity's Last Exam (AA text-only) published 2026-09-10
    Why
    Observed: 0.59129 fraction
    Peer mean 0.15283 · peer sd 0.14006 · n 608 · families 449
    Directed z 3.13 · baseline z 1.42 · gap 1.711 · profile n 32
    0.59129 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Humanity's Last Exam (AA text-only) · published 2026-09-10 · Published board
    Exact value: 0.591288229842447 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:hle
  • unusually strong MLCR-AA published 2026-09-10
    Why
    Observed: 0.71111 fraction
    Peer mean 0.18737 · peer sd 0.17142 · n 84 · families 61
    Directed z 3.055 · baseline z 1.422 · gap 1.633 · profile n 32
    0.71111 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: MLCR-AA · published 2026-09-10 · Published board
    Exact value: 0.711111111111111 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:3e87c73e-a257-495e-9730-367a66229811:mlcrOverall
Missing coverage · 91 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)81.653.4
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)80.753.2
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)79.151.2
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)77.149.1
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)75.247.0

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $10.00 / 1M
Cached input $0.250 / 1M
Cache write $12.50 / 1M
Output $50.00 / 1M
Status GA
Token offers by platform · 6 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

6 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (3)

Amazon Bedrock
Google
Azure

AWS Bedrock (1)

AWS Bedrock

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) — benchmarks & cost | Benchmark Heaven