← All models

MiniMax-M3

open weights
MiniMax · released 2026-06-01 · 16 offers

Output 120 tokens/sFirst token 1.0 sContext 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1CoreWeave
2DeepInfra
3Together
4Novita
5Venice

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

59.7

includes −6.9 for its Benchmaxxing signal (from 66.6; why, switch off in Options)

Benchmaxxing signal +6.9, medium Benchmaxxing tag · medium report →

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 67AA Intelligence: percentile 86AA Agentic: percentile 73Epoch ECI: percentile 42Software ECI: percentile 35DesignArena: percentile 45

AA Coding 58.6Coding Agent v1.4 AA Intelligence 29.6AA Agentic 30.8Epoch ECI 146.5Software ECI 149.3DesignArena 1215/1205

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

45 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-AnalystAgent1610.0%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase611,096 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2721,304 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA6788.4%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • Terminal-Bench Hard (AA)8942.4%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)6565.2%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)532.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • Vals Index v2 (Vals AI)1742.7%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)2048.3%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • Terminal-Bench 2.1 (Vals Index v2)1053.6%

    Terminal-Bench 2.1 (Vals Index v2) v2 · Published board — Command-line interface problem solving, as run by Vals AI.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)384.2%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • τ²-Bench Airline (OpenRouter run)5371.0%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • OSWorld 2.0, June 2026 task release (XLANG Lab) v2026.06.24no percentile4.6%

    OSWorld 2.0, June 2026 task release (XLANG Lab) v2026.06.24 · standard — A computer-use agent completes 108 long-horizon, real-world workflows across 31 self-hosted websites and desktop applications, each taking a person about 1.6 hours.

  • AA Agentic Index7330.8

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.15047.1%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • Vibe Code Bench (Vals Index v2)1647.6%

    Vibe Code Bench (Vals Index v2) v2 · Published board — End-to-end app-building tasks.

  • Code Migration (Vals Index v2 subset)2220.9%

    Code Migration (Vals Index v2 subset) v2 · Published board — Porting projects to another language, including COBOL modernization.

  • FrontierCode 1.1 Main (Cognition) v1.1no percentile14.7%

    FrontierCode 1.1 Main (Cognition) v1.1 · mini-swe-agent — Whether a maintainer would merge the agent's pull request, on tasks crafted by open-source maintainers and graded with tests, rubrics and verifiers.

  • SWE-rebench, issues 15 May – 1 Jul 2026 (Nebius) v2026-05-15..2026-07-012947.2%

    SWE-rebench, issues 15 May – 1 Jul 2026 (Nebius) v2026-05-15..2026-07-01 · SWE-rebench fixed ReAct scaffold, tool mode — A model resolves 111 fresh GitHub issues from 65 repositories, created between 15 May and 1 July 2026, inside one fixed minimal agent scaffold, run five times.

  • AA Coding Index6758.6

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)521,215 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack541,205 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)449.8%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19783.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA6117.2%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Reasoning

  • LiveBench 2026-06-25no percentile67.3

    LiveBench 2026-06-25 · Published board — A periodically refreshed suite of objectively graded tasks with separate category scores and an overall category mean.

  • Chess Puzzles (Epoch AI)2714.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)37.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index8629.6

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Roleplay

  • EQ-Bench 4 v401,150 Elo

    EQ-Bench 4 v4 · Published board — Measures active emotional and social intelligence of LLMs in multi-turn roleplay chats with simulated personas, ranked by an LLM-judge panel Elo.

Science

Tool-use

Vision

Writing

Unusual results

2 threshold-crossing signals flagged · 41 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually strong IFBench (AA single-turn) published 2026-09-10
    Why
    Observed: 0.82857 fraction
    Peer mean 0.48006 · peer sd 0.17139 · n 450 · families 331
    Directed z 2.033 · baseline z 0.138 · gap 1.895 · profile n 40
    0.82857 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: IFBench (AA single-turn) · published 2026-09-10 · Published board
    Exact value: 0.828571428571429 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:277f939a-985b-4b37-859d-b3eabc7c0b26:ifbench
  • unusually strong Humanity's Last Exam (AA text-only) published 2026-09-10
    Why
    Observed: 0.38971 fraction
    Peer mean 0.15283 · peer sd 0.14006 · n 608 · families 449
    Directed z 1.691 · baseline z 0.147 · gap 1.545 · profile n 40
    0.38971 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Humanity's Last Exam (AA text-only) · published 2026-09-10 · Published board
    Exact value: 0.389712696941613 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:277f939a-985b-4b37-859d-b3eabc7c0b26:hle
Missing coverage · 96 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Token offers by platform · 12 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

12 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (9)

DeepInfra
Together
Novita
ModelRun
Mara
SambaNova
CoreWeave
Venice
Parasail

TensorX (1)

TensorX

Nebius (1)

Nebius

Azure AI Foundry (1)

Azure AI Foundry
MiniMax-M3 — benchmarks & cost | Benchmark Heaven