← All models

Kimi K2 Thinking

open weightsdeprecated by benchmark source
Moonshot AI · released 2025-11-06 · 4 offers

Context 256K tokens

Top 4 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Novita
2Google Vertex AI
3Google
4AWS Bedrock

Composite

3 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

45.3

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 74Epoch ECI: percentile 36Software ECI: percentile 19

AA Coding Coding Agent v1.4 AA Intelligence 22.0AA Agentic Epoch ECI 145.8Software ECI 144.3DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

21 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

Coding

  • SWE-bench Verifiedno percentile63.4%

    SWE-bench Verified published 2026-09-10 · mini-SWE-agent — A 500-instance human-filtered subset of SWE-bench created with OpenAI, served as the default Verified leaderboard view.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.17172.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v20259694.7%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

  • SimpleBench6939.6%

    SimpleBench published 2026-09-10 · Published board — A multiple-choice text benchmark of over 200 questions covering spatio-temporal reasoning, social intelligence, and linguistic adversarial robustness, on which a non-specialized human baseline outperforms every tested LLM.

  • AA Intelligence Index7422.0

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

Uncensored

  • UGI Willingness101.80

    UGI Willingness published 2026-09-10 · Published board — A private-question evaluation of how far a model follows challenging instructions before refusing or deviating.

  • UGI9352.7

    UGI published 2026-09-10 · Published board — A private-question suite combining sensitive-topic knowledge with willingness to follow controversial instructions.

Writing

  • EQ-Bench Creative Writing v3911,628 Elo

    EQ-Bench Creative Writing v3 v3 · Published board — A LLM-judged creative writing benchmark with vocabulary and GPT-slop controls.

  • EQ-Bench Longform Creative Writing v1.119269.0

    EQ-Bench Longform Creative Writing v1.11 · Published board — LLM-judged benchmark in which models plan and write a short story/novella over 8x 1000-word turns, scored 0-100 across 14 rubric dimensions.

  • UGI Writing10059.3

    UGI Writing published 2026-09-10 · Published board — A maintainer-owned writing evaluation considering intelligence, style, repetition and length adherence, informed by human preferences.

Unusual results

2 threshold-crossing signals flagged · 20 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

Missing coverage · 123 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Token offers by platform · 4 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

4 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (2)

Novita
Google

Google Vertex AI (1)

Google Vertex AI

AWS Bedrock (1)

AWS Bedrock
Kimi K2 Thinking — benchmarks & cost | Benchmark Heaven