← All models

Kimi K2.7 Code

open weights
Moonshot AI · released 2026-06-12 · 20 offers

Output 60 tokens/sFirst token 1.2 sContext 256K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1DeepInfra
2CoreWeave
3Inceptron
4Inceptron
5Venice

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

67.0

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 70AA Intelligence: percentile 82AA Agentic: percentile 59Epoch ECI: percentile 54Software ECI: percentile 34DesignArena: percentile 39

AA Coding 60.8Coding Agent v1.4 AA Intelligence 26.3AA Agentic 22.5Epoch ECI 150.1Software ECI 148.7DesignArena 1190/1185

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

31 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-Briefcase41856 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2571,114 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA5085.0%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • Terminal-Bench Hard (AA)9244.7%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)6767.4%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)471.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • τ²-Bench Airline (OpenRouter run)5771.8%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index5922.5

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.15147.8%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • DeepSWE (Datacurve, via Epoch AI)830.5%

    DeepSWE (Datacurve, via Epoch AI) published 2026-09-15 · mini-swe-agent — Pass rate of coding agents on original, long-horizon software engineering tasks, run by Datacurve with the mini-swe-agent harness.

  • AA Coding Index7060.8

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)401,190 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack481,185 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)5111.2%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.18779.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA5816.1%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • FrontierMath Tiers 1–3 v2 (Epoch AI)2154.0%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)312.2%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • SimpleBench8857.9%

    SimpleBench published 2026-09-10 · Published board — A multiple-choice text benchmark of over 200 questions covering spatio-temporal reasoning, social intelligence, and linguistic adversarial robustness, on which a non-specialized human baseline outperforms every tested LLM.

  • Chess Puzzles (Epoch AI)4621.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • AA Intelligence Index8226.3

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

  • AutomationBench-AA v1.0.65124.5%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA4540.2%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ²-Bench Telecom (AA)8590.1%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.15320.2%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Unusual results

1 threshold-crossing signal flagged · 33 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak DeepSWE (Datacurve, via Epoch AI) published 2026-09-15
    Why
    Observed: 0.30531 fraction
    Peer mean 0.56136 · peer sd 0.161 · n 63 · families 21
    Directed z -1.59 · baseline z 0.311 · gap -1.902 · profile n 32
    0.30531 fractionmeasuredobserved 2026-09-15epoch.ai
    Evidence
    Axis: DeepSWE (Datacurve, via Epoch AI) · published 2026-09-15 · mini-swe-agent
    Exact value: 0.3053097345132743 fraction
    Evaluation harness: mini-swe-agent
    Observed: 2026-09-15T09:05:24Z · publication date: not recorded
    Observation id: public:ed4d65bff33890a60159da84
Missing coverage · 110 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $0.950 / 1M
Cached input $0.190 / 1M
Output $4.00 / 1M
Status GA
Token offers by platform · 16 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

16 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (11)

DeepInfra
CoreWeave
Inceptron
Venice
ModelRun
SiliconFlow
Novita
BaseTen
Fireworks
Moonshot AI
Moonshot AI

Inceptron (1)

Inceptron

TensorX (1)

TensorX

Azure AI Foundry (2)

Azure AI Foundry
Azure AI Foundry

Nebius (1)

Nebius
Kimi K2.7 Code — benchmarks & cost | Benchmark Heaven