← All models

gpt-oss-20b (low)

open weights
OpenAI · released 2025-08-05 · 17 offers

Output 232 tokens/sFirst token 0.5 sContext 131K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1AkashML
2Parasail
3DekaLLM
4CoreWeave
5DeepInfra

Composite

3 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

18.2

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 45Epoch ECI: percentile 15Software ECI: percentile 4

AA Coding Coding Agent v1.4 AA Intelligence 10.0AA Agentic Epoch ECI 137.8Software ECI 136.8DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

11 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.13231.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v20255862.3%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

  • Chess Puzzles (Epoch AI)112.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • AA Intelligence Index4510.0

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

Missing coverage · 133 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
gpt-oss-20b (high)20.79.0
gpt-oss-20b (low)10.0
Token offers by platform · 16 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

16 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (12)

AkashML
Parasail
DekaLLM
CoreWeave
DeepInfra
Novita
SiliconFlow
Together
Groq
Amazon Bedrock
Amazon Bedrock
Google

OVHcloud (1)

OVHcloud

Google Vertex AI (1)

Google Vertex AI

AWS Bedrock (1)

AWS Bedrock

STACKIT (1)

STACKIT
gpt-oss-20b (low) — benchmarks & cost | Benchmark Heaven