← All models

Gemma 4 31B (Reasoning)

open weights
Google · released 2026-04-02 · 19 offers

Output 35 tokens/sFirst token 0.9 sContext 256K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Chutes
2Chutes
3DeepInfra
4Parasail
5Venice

Composite

4 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

34.5

includes −5.3 for its Benchmaxxing signal (from 39.8; why, switch off in Options)

Benchmaxxing signal +5.3, light Benchmaxxing tag, uncertain: The 80 % interval reaches below zero — treat this tag as uncertain · light report →The 80 % interval reaches below zero — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 48AA Intelligence: percentile 62AA Agentic: percentile 36Epoch ECI: percentile 29Software ECI: percentile 19

AA Coding 43.4Coding Agent v1.4 AA Intelligence 15.4AA Agentic 6.7Epoch ECI 142.7Software ECI 144.1DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

32 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

  • AA-Briefcase17377 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v236755 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA547.2%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • ITBench-AA5937.3%

    ITBench-AA published 2026-09-10 · Published board — Tests root-cause diagnosis from offline Kubernetes incident snapshots.

  • Terminal-Bench Hard (AA)8336.4%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)4443.4%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)180.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • τ²-Bench Airline (OpenRouter run)8376.1%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index366.70

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.14545.5%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • AA Coding Index4843.4

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)356.0%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.16669.7%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Science

Tool-use

  • AutomationBench-AA v1.0.6334.9%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA1128.3%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ²-Bench Telecom (AA)5759.9%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.14214.8%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Uncensored

  • UGI Willingness272.50

    UGI Willingness published 2026-09-10 · Published board — A private-question evaluation of how far a model follows challenging instructions before refusing or deviating.

  • UGI2721.5

    UGI published 2026-09-10 · Published board — A private-question suite combining sensitive-topic knowledge with willingness to follow controversial instructions.

Vision

Writing

  • EQ-Bench Creative Writing v3571,366 Elo

    EQ-Bench Creative Writing v3 v3 · Published board — A LLM-judged creative writing benchmark with vocabulary and GPT-slop controls.

  • Judgemark v410072.3

    Judgemark v4 v4 · Published board — Meta-evaluation that rates LLM judges on how discriminatively they score fixed creative-writing samples.

  • EQ-Bench Longform Creative Writing v1.115656.5

    EQ-Bench Longform Creative Writing v1.11 · Published board — LLM-judged benchmark in which models plan and write a short story/novella over 8x 1000-word turns, scored 0-100 across 14 rubric dimensions.

  • UGI Writing7738.6

    UGI Writing published 2026-09-10 · Published board — A maintainer-owned writing evaluation considering intelligence, style, repetition and length adherence, informed by human preferences.

Unusual results

2 threshold-crossing signals flagged · 32 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak Harvey LAB-AA published 2026-09-10
    Why
    Observed: 0.47228 fraction
    Peer mean 0.79402 · peer sd 0.17289 · n 43 · families 38
    Directed z -1.861 · baseline z 0.074 · gap -1.935 · profile n 31
    0.47228 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Harvey LAB-AA · published 2026-09-10 · Published board
    Exact value: 0.47228253003329 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:cd26a386-4873-46ff-b853-d239050025a2:harveyLab
  • unusually strong IFBench (AA single-turn) published 2026-09-10
    Why
    Observed: 0.75578 fraction
    Peer mean 0.48006 · peer sd 0.17139 · n 450 · families 331
    Directed z 1.609 · baseline z -0.038 · gap 1.646 · profile n 31
    0.75578 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: IFBench (AA single-turn) · published 2026-09-10 · Published board
    Exact value: 0.75578231292517 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:cd26a386-4873-46ff-b853-d239050025a2:ifbench
Missing coverage · 111 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Gemma 4 31B (Reasoning)43.415.4
Gemma 4 31B (Non-reasoning)33.213.9
Token offers by platform · 17 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

17 offers within the active global filters; “—” means the catalog is active but no public token price is available.

Chutes (1)

Chutes

OpenRouter (13)

Chutes
DeepInfra
Parasail
Venice
CoreWeave
DeepInfra
Novita
Together
ModelRun
Crusoe
DeepInfra
SambaNova
SiliconFlow

TrustedTokens (1)

TrustedTokens

AWS Bedrock (1)

AWS Bedrock

STACKIT (1)

STACKIT
Gemma 4 31B (Reasoning) — benchmarks & cost | Benchmark Heaven