← All models

gpt-oss-120b (low)

open weights
OpenAI · released 2025-08-05 · 33 offers

Output 172 tokens/sFirst token 0.5 sContext 131K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1AkashML
2CoreWeave
3DekaLLM
4DeepInfra
5DigitalOcean

Composite

4 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

20.7

includes −1.3 for its Benchmaxxing signal (from 22.0; why, switch off in Options)

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 24AA Intelligence: percentile 46Epoch ECI: percentile 22Software ECI: percentile 7

AA Coding 21.2Coding Agent v1.4 AA Intelligence 10.2AA Agentic Epoch ECI 140.1Software ECI 138.9DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

14 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
Compare this model →

Agentic

  • GDPval-AA v214364 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench Hard (AA)325.3%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)2113.9%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

Coding

  • AA Coding Index2421.2

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.14546.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v20256266.7%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Safety/Alignment

Science

Tool-use

  • τ²-Bench Telecom (AA)4945.0%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.122.9%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Missing coverage · 129 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
gpt-oss-120b (high)30.412.3
gpt-oss-120b (low)21.210.2
Token offers by platform · 33 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

33 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (23)

AkashML
CoreWeave
DekaLLM
DeepInfra
DigitalOcean
Crusoe
Novita
Mancer 2
Parasail
BaseTen
SiliconFlow
Groq
Amazon Bedrock
Nebius
Amazon Bedrock
DeepInfra
Phala
SambaNova
DeepInfra
Cerebras
Google
Together
Mara

Google Vertex AI (1)

Google Vertex AI

OVHcloud (1)

OVHcloud

T-Systems LLM Hub (1)

T-Systems LLM Hub

TrustedTokens (1)

TrustedTokens

Azure AI Foundry (1)

Azure AI Foundry

Nebius (1)

Nebius

Scaleway (1)

Scaleway

IONOS (1)

IONOS

AWS Bedrock (1)

AWS Bedrock

STACKIT (1)

STACKIT
gpt-oss-120b (low) — benchmarks & cost | Benchmark Heaven