← All models

DeepSeek V4 Pro 0813 (Reasoning, Max Effort)

open weights★ featured
DeepSeek · released 2026-08-13 · 24 offers

Output 73 tokens/sFirst token 1.1 sContext 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Sail Research
2Ionstream
3Novita
4NextBit
5DeepInfra

Composite

4 of 7 inputs · 1 from the model family6 radar axes: DesignArena's two boards share one

78.3

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 78Coding Agent v1.4: percentile 21AA Intelligence: percentile 92AA Agentic: percentile 82Epoch ECI: percentile 68

AA Coding 68.8Coding Agent v1.4 42.8AA Intelligence 36.3AA Agentic 42.3Epoch ECI 155.5Software ECI DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

39 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
Compare this model →

Agentic

  • AA-Briefcase711,265 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2871,493 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench v2.1 (AA)7978.7%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)7614.1%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • Vals Index v2 (Vals AI)3352.4%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)3050.4%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • Terminal-Bench 2.1 (Vals Index v2)1354.7%

    Terminal-Bench 2.1 (Vals Index v2) v2 · Published board — Command-line interface problem solving, as run by Vals AI.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)637.5%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • τ²-Bench Airline (OpenRouter run)9577.7%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index8242.3

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

Efficiency

Knowledge

Long-context

  • GDP.pdf (AA)5211.4%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19180.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA6317.8%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • FrontierMath Tiers 1–3 v2 (Epoch AI)4464.6%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)2726.8%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • Chess Puzzles (Epoch AI)8847.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)7843.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index9236.3

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Safety/Alignment

Science

Tool-use

  • AutomationBench-AA v1.0.68556.7%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA9449.6%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ³-Banking (AA) v1.0.18539.6%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

  • Excel Modeling Benchmark (Vals Index v2)1952.8%

    Excel Modeling Benchmark (Vals Index v2) v2 · Published board — Building and editing financial models in spreadsheets.

Missing coverage · 104 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Token offers by platform · 17 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

17 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (15)

Sail Research
Sail Research
Ionstream
Novita
NextBit
DeepInfra
CoreWeave
Parasail
DigitalOcean
SiliconFlow
Fireworks
Together
BaseTen
Phala
Venice

TensorX (1)

TensorX

Nebius (1)

Nebius
DeepSeek V4 Pro 0813 (Reasoning, Max Effort) — benchmarks & cost | Benchmark Heaven