← All models

Kimi K2.6

open weightsdeprecated by benchmark source
Moonshot AI · released 2026-04-20 · 25 offers

Context 256K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Inceptron
2Inceptron
3Chutes
4Chutes
5Decart

Composite

6 of 7 inputs · 3 from the model family6 radar axes: DesignArena's two boards share one

57.5

includes −9.4 for its Benchmaxxing signal (from 67.0; why, switch off in Options)

Benchmaxxing signal +9.4, medium Benchmaxxing tag · medium report →

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 72Coding Agent v1.4: percentile 8AA Intelligence: percentile 84AA Agentic: percentile 58Epoch ECI: percentile 56Software ECI: percentile 40DesignArena: percentile 37

AA Coding 61.8Coding Agent v1.4 33.9AA Intelligence 27.5AA Agentic 22.1Epoch ECI 151.0Software ECI 149.9DesignArena —/1165

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

29 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • APEX-Agents-AA6628.5%

    APEX-Agents-AA published 2026-09-10 · Published board — Tests professional-service tasks that require agents to produce locally graded deliverables.

  • AA-Briefcase39820 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2581,115 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA4384.1%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • ITBench-AA4131.2%

    ITBench-AA published 2026-09-10 · Published board — Tests root-cause diagnosis from offline Kubernetes incident snapshots.

  • Terminal-Bench Hard (AA)9143.9%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)6665.9%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • τ²-Bench Airline (OpenRouter run)6773.7%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index5822.1

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • Artificial Analysis Coding Agent Index v1.4433.9%

    Artificial Analysis Coding Agent Index v1.4 v1.4 · Claude Code — The retained AA Coding Agent Index measures coding-agent systems using the earlier three-component implementation.

  • SciCode (AA subproblems) v1.0.16551.5%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • AA Coding Index7261.8

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Full-Stack391,165 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)5913.0%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19381.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Roleplay

  • Japanese-RP-Bench v2 (tegnike fork)10095.1

    Japanese-RP-Bench v2 (tegnike fork) v2 · Published board — Fork of Japanese-RP-Bench adding v2 metrics for role fidelity, conversation quality, persona stability, robustness, and recovery in Japanese roleplay.

Science

Tool-use

  • AutomationBench-AA v1.0.64113.0%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA4038.5%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ²-Bench Telecom (AA)9695.9%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.15823.3%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Writing

  • LLM Creative Story-Writing Benchmark860.30

    LLM Creative Story-Writing Benchmark published 2026-09-10 · Published board — Pairwise comparison of short stories written to the same constrained creative briefs, with LLM evaluator choices combined into a relative comparison score.

  • The Slop Index7221.1

    The Slop Index published 2026-09-10 · Published board — Leaderboard scoring how much each model's writing reads like chatbot slop versus pre-AI human writing, blending live arena votes with four mechanical axes.

Unusual results

1 threshold-crossing signal flagged · 29 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak Artificial Analysis Coding Agent Index v1.4
    Why
    Observed: 0.33889 fraction
    Peer mean 0.55251 · peer sd 0.11622 · n 26 · families 12
    Directed z -1.838 · baseline z 0.563 · gap -2.401 · profile n 28
    0.33889 fractionmeasuredobserved 2026-09-09artificialanalysis.ai
    Evidence
    Axis: Artificial Analysis Coding Agent Index v1.4 · v1.4 · Claude Code
    Exact value: 0.3388925953605663 fraction
    Evaluation harness: Claude Code
    Observed: 2026-09-09 · publication date: not recorded
    Observation id: aa-coding:1.4:48
Missing coverage · 113 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Kimi K2.661.827.5
Kimi K2.6 (Non-reasoning)23.6
Token offers by platform · 20 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

20 offers within the active global filters; “—” means the catalog is active but no public token price is available.

Inceptron (1)

Inceptron

OpenRouter (15)

Inceptron
Chutes
DigitalOcean
CoreWeave
SiliconFlow
DeepInfra
Parasail
Novita
Moonshot AI
BaseTen
Fireworks
Phala
Decart
Venice
Crusoe

Chutes (1)

Chutes

Azure AI Foundry (2)

Azure AI Foundry
Azure AI Foundry

Nebius (1)

Nebius
Kimi K2.6 — benchmarks & cost | Benchmark Heaven