← All models

Claude 4.5 Haiku (Non-reasoning)

Anthropic · released 2025-10-15 · 13 offers

Output 97 tokens/sFirst token 0.5 sContext 200K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Google
2Amazon Bedrock
3Google Vertex AI
4Azure
5AWS Bedrock

Composite

3 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

35.4

dominance-adjusted from 37.3: a better-measured model that is at least as good on each of these inputs ranks above it

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 62Epoch ECI: percentile 27Software ECI: percentile 21

AA Coding Coding Agent v1.4 AA Intelligence 15.4AA Agentic Epoch ECI 142.4Software ECI 145.0DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

13 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.14749.7%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA184.4%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • AIME 2025 (AA) v20254039.0%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Safety/Alignment

Science

Tool-use

Vision

Unusual results

1 threshold-crossing signal flagged · 14 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

Missing coverage · 131 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Claude 4.5 Haiku (Reasoning)43.917.6
Claude 4.5 Haiku (Non-reasoning)15.4

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $1.00 / 1M
Cached input $0.100 / 1M
Cache write $1.25 / 1M
Output $5.00 / 1M
Status GA

Legacy annual Pro/Pro+ request billing only.

Multiplier 0.33×
Effective cost $0.013 / request

Legacy annual-plan multiplier.

Token offers by platform · 11 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

11 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (7)

Google
Amazon Bedrock
Google
Amazon Bedrock
Google
Azure
Amazon Bedrock

Google Vertex AI (2)

Google Vertex AI
Google Vertex AI

AWS Bedrock (1)

AWS Bedrock

T-Systems LLM Hub (1)

T-Systems LLM Hub
Claude 4.5 Haiku (Non-reasoning) — benchmarks & cost | Benchmark Heaven