← All models

Kimi K2.5 (Reasoning)

open weightsdeprecated by benchmark source
Moonshot AI · released 2026-01-27 · 7 offers

Context 256K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1SiliconFlow
2Novita
3Venice
4Azure AI Foundry
5Amazon Bedrock

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

54.2

includes −5.4 for its Benchmaxxing signal (from 59.6; why, switch off in Options)

Benchmaxxing signal +5.4, light Benchmaxxing tag, uncertain: Based on only 6 comparisons — treat this tag as uncertain · light report →Based on only 6 comparisons — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 54AA Intelligence: percentile 78Epoch ECI: percentile 47Software ECI: percentile 30DesignArena: percentile 32

AA Coding 46.8Coding Agent v1.4 AA Intelligence 23.5AA Agentic Epoch ECI 148.0Software ECI 146.8DesignArena 1151/1123

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

17 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • APEX-Agents-AA1711.5%

    APEX-Agents-AA published 2026-09-10 · Published board — Tests professional-service tasks that require agents to produce locally graded deliverables.

  • GDPval-AA v245936 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench Hard (AA)8034.8%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)4645.7%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • τ²-Bench Airline (OpenRouter run)5571.2%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

Coding

  • AA Coding Index5446.8

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)321,151 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack331,123 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.18378.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Science

Tool-use

  • τ²-Bench Telecom (AA)9695.9%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.14114.2%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Missing coverage · 124 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Kimi K2.5 (Reasoning)46.823.5
Kimi K2.5 (Non-reasoning)19.4
Token offers by platform · 6 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

6 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

SiliconFlow
Venice
Amazon Bedrock
Novita

Azure AI Foundry (1)

Azure AI Foundry

AWS Bedrock (1)

AWS Bedrock
Kimi K2.5 (Reasoning) — benchmarks & cost | Benchmark Heaven