← All models

Claude 4.5 Haiku (Reasoning)

Anthropic · released 2025-10-15 · 13 offers

Output 108 tokens/sFirst token 16 sContext 200K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Google
2Amazon Bedrock
3Google Vertex AI
4Azure
5AWS Bedrock

Composite

4 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

41.2

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 49AA Intelligence: percentile 66AA Agentic: percentile 42Epoch ECI: percentile 27Software ECI: percentile 21

AA Coding 43.9Coding Agent v1.4 AA Intelligence 17.6AA Agentic 10.3Epoch ECI 142.4Software ECI 145.0DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

28 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

  • AA-AnalystAgent2715.0%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase27614 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v240854 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA1461.1%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • ITBench-AA3127.3%

    ITBench-AA published 2026-09-10 · Published board — Tests root-cause diagnosis from offline Kubernetes incident snapshots.

  • Terminal-Bench Hard (AA)6927.3%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)4444.2%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)180.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • τ²-Bench Airline (OpenRouter run)4267.8%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index4210.3

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.13442.2%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • AA Coding Index4943.9

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)263.8%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.17674.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA226.1%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • AIME 2025 (AA) v20258183.7%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Science

Tool-use

  • AutomationBench-AA v1.0.6293.2%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA1729.4%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ²-Bench Telecom (AA)5554.7%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.1279.3%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Missing coverage · 115 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude 4.5 Haiku (Reasoning)43.917.6
Claude 4.5 Haiku (Non-reasoning)15.4

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $1.00 / 1M
Cached input $0.100 / 1M
Cache write $1.25 / 1M
Output $5.00 / 1M
Status GA

Legacy annual Pro/Pro+ request billing only.

Multiplier 0.33×
Effective cost $0.013 / request

Legacy annual-plan multiplier.

Token offers by platform · 11 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

11 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (7)

Google
Amazon Bedrock
Google
Amazon Bedrock
Google
Azure
Amazon Bedrock

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock

T-Systems LLM Hub (1)

T-Systems LLM Hub
Claude 4.5 Haiku (Reasoning) — benchmarks & cost | Benchmark Heaven