← All models

GPT-5.3 Codex (xhigh)

OpenAI · released 2026-02-05 · 4 offers

Output 134 tokens/sFirst token 40 sContext 400K tokens

Top 2 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Azure
2Azure AI Foundry

Composite

5 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

63.6

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Intelligence: percentile 89Epoch ECI: percentile 78Software ECI: percentile 66DesignArena: percentile 0

AA Coding Coding Agent v1.4 AA Intelligence 32.5AA Agentic Epoch ECI 156.6Software ECI 156.9DesignArena 1048/993

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

12 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

Coding

  • DesignArena Web Apps (agentic)21,048 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack0993 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.19883.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Safety/Alignment

Science

Tool-use

Vision

Unusual results

3 threshold-crossing signals flagged · 15 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak DesignArena Web Apps (agentic) published 2026-09-19
    Why
    Observed: 1048 Elo
    Peer mean 1198.82222 · peer sd 77.91656 · n 45 · families 44
    Directed z -1.936 · baseline z 0.774 · gap -2.709 · profile n 14
    1048 Elomeasured2243 battlesobserved 2026-09-19designarena.ai
    Evidence
    Axis: DesignArena Web Apps (agentic) · published 2026-09-19 · Published board
    Exact value: 1048 Elo
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:frontend:gpt-5.3-codex::xhigh
    Protocol: inspect · file data/raw/designarena.json
  • unusually weak DesignArena Full-Stack published 2026-09-19
    Why
    Observed: 993 Elo
    Peer mean 1176.70213 · peer sd 102.68003 · n 47 · families 46
    Directed z -1.789 · baseline z 0.763 · gap -2.552 · profile n 14
    993 Elomeasured1109 battlesobserved 2026-09-19designarena.ai
    Evidence
    Axis: DesignArena Full-Stack · published 2026-09-19 · Published board
    Exact value: 993 Elo
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:fullstack:gpt-5.3-codex::xhigh
    Protocol: inspect · file data/raw/designarena.json
  • unusually strong Terminal-Bench Hard (AA)
    Why
    Observed: 0.5303 fraction
    Peer mean 0.18472 · peer sd 0.16951 · n 432 · families 315
    Directed z 2.039 · baseline z 0.49 · gap 1.549 · profile n 14
    0.5303 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Terminal-Bench Hard (AA) · pinned revision 74221fb · Published board
    Exact value: 0.53030303030303 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:59b5b14b-5365-4ee7-824a-18a8e6309644:terminalbenchHard
Missing coverage · 130 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $1.75 / 1M
Cached input $0.175 / 1M
Output $14.00 / 1M
Status GA

Legacy annual Pro/Pro+ request billing only.

Multiplier 6.00×
Effective cost $0.240 / request

Legacy annual-plan multiplier.

Token offers by platform · 2 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

2 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (1)

Azure

Azure AI Foundry (1)

Azure AI Foundry
GPT-5.3 Codex (xhigh) — benchmarks & cost | Benchmark Heaven