← All models

DeepSeek V4.1 Flash (Reasoning, Max Effort)

open weights★ featured
DeepSeek · released 2026-09-10 · 24 offers

Output 221 tokens/sFirst token 0.8 sContext 1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1DeepInfra
2Relace
3Morph
4TrustedTokens
5Wafer

Composite

1 of 7 inputs6 radar axes: DesignArena's two boards share one

71.3

includes −6.7 for its Benchmaxxing signal (from 78.0; why, switch off in Options)

dominance-adjusted from 87.4: a better-measured model that is at least as good on each of these inputs ranks above it

Benchmaxxing signal +6.7, medium Benchmaxxing tag, uncertain: Based on only 7 comparisons and the 80 % interval reaches below zero — treat this tag as uncertain · medium report →Based on only 7 comparisons and the 80 % interval reaches below zero — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 94

AA Coding Coding Agent v1.4 AA Intelligence 39.5AA Agentic Epoch ECI Software ECI DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

19 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Compare this model →

Agentic

  • AA-Briefcase821,424 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2951,632 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench v4.0 (AA)8626.8%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • τ²-Bench Airline (OpenRouter run)8776.7%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

Coding

  • SciCode (AA subproblems) v1.0.16851.9%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

Efficiency

Knowledge

Long-context

  • GDP.pdf (AA)5812.8%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19984.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • ArXivMath 08/2026 (MathArena) v2026-08043.9%

    ArXivMath 08/2026 (MathArena) v2026-08 · Published board — Research-level math problems with a checkable final answer, drawn from arXiv papers submitted in August 2026, so they postdate most training data.

  • BrokenArXiv 08/2026 (MathArena) v2026-08054.8%

    BrokenArXiv 08/2026 (MathArena) v2026-08 · Published board — Plausible but false statements taken from August 2026 arXiv papers; a model scores by refusing to prove them and saying the statement is false as written.

Reasoning

Safety/Alignment

Science

  • CritPt (AA)8914.3%

    CritPt (AA) published 2026-09-10 · Published board — Tests research-level physics reasoning with Python, symbolic and numerical answers.

  • GPQA Diamond (OpenRouter run)8990.2%

    GPQA Diamond (OpenRouter run) published 2026-09-15 · Published board — Graduate-level science questions that resist retrieval and reward careful reasoning. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

Tool-use

  • AutomationBench-AA v1.0.610068.9%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

Vision

Missing coverage · 125 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Token offers by platform · 19 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

19 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (16)

DeepInfra
Morph
Wafer
Fireworks
Sail Research
Novita
Makora
DigitalOcean
Together
SiliconFlow
BaseTen
Modal
Phala
Venice
Relace
Parasail

TrustedTokens (1)

TrustedTokens

TensorX (1)

TensorX

Nebius (1)

Nebius
DeepSeek V4.1 Flash (Reasoning, Max Effort) — benchmarks & cost | Benchmark Heaven