← All models

GPT-5.4 (xhigh)

deprecated by benchmark source
OpenAI · released 2026-03-05 · 10 offers

Context 1.1M tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Azure
2Amazon Bedrock
3Azure AI Foundry
4AWS Bedrock
5T-Systems LLM Hub

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

71.3

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 82AA Intelligence: percentile 93Epoch ECI: percentile 79Software ECI: percentile 64DesignArena: percentile 6

AA Coding 71.1Coding Agent v1.4 AA Intelligence 39.0AA Agentic Epoch ECI 156.9Software ECI 156.1DesignArena 1074/1027

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

24 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

Coding

  • SWE-Bench Pro Publicno percentile59.1%

    SWE-Bench Pro Public published 2026-09-10 · Published board — A software-engineering evaluation of long-horizon tasks in public repositories, scored by patches passing both new and regression tests.

  • SWE Atlas Refactoring (Scale AI)2044.3%

    SWE Atlas Refactoring (Scale AI) published 2026-09-15 · Published board — Whether a coding agent restructures production code while preserving its behaviour, graded by tests and rubrics.

  • GSO software optimization, Opt@1 (UC Berkeley) vopt1-1026731.4%

    GSO software optimization, Opt@1 (UC Berkeley) vopt1-102 · OpenHands — An agent gets a real codebase and a performance test and must make the code as fast as an expert developer's optimization; 102 tasks across 10 codebases and 5 languages.

  • AA Coding Index8271.1

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)71,074 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack91,027 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.19582.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Safety/Alignment

Science

Tool-use

  • τ²-Bench Telecom (AA)8187.1%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.18539.6%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Unusual results

3 threshold-crossing signals flagged · 23 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak DesignArena Web Apps (agentic) published 2026-09-19
    Why
    Observed: 1074 Elo
    Peer mean 1198.82222 · peer sd 77.91656 · n 45 · families 44
    Directed z -1.602 · baseline z 0.904 · gap -2.506 · profile n 22
    1074 Elomeasured3317 battlesobserved 2026-09-19designarena.ai
    Evidence
    Axis: DesignArena Web Apps (agentic) · published 2026-09-19 · Published board
    Exact value: 1074 Elo
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:frontend:gpt-5.4::xhigh
    Protocol: inspect · file data/raw/designarena.json
  • unusually strong CritPt (AA) published 2026-09-10
    Why
    Observed: 0.23429 fraction
    Peer mean 0.03928 · peer sd 0.07596 · n 522 · families 369
    Directed z 2.567 · baseline z 0.714 · gap 1.853 · profile n 22
    0.23429 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · published 2026-09-10 · Published board
    Exact value: 0.2342857143 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:a89c4b28-2d8c-456e-88ea-255fb51fd2b6:critpt
  • unusually strong Terminal-Bench Hard (AA)
    Why
    Observed: 0.57576 fraction
    Peer mean 0.18472 · peer sd 0.16951 · n 432 · families 315
    Directed z 2.307 · baseline z 0.726 · gap 1.581 · profile n 22
    0.57576 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Terminal-Bench Hard (AA) · pinned revision 74221fb · Published board
    Exact value: 0.575757575757576 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:a89c4b28-2d8c-456e-88ea-255fb51fd2b6:terminalbenchHard
Missing coverage · 117 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
GPT-5.4 (xhigh)71.139.0
GPT-5.4 (low)27.6
GPT-5.4 (Non-reasoning)18.2

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $2.50 / 1M
Cached input $0.250 / 1M
Output $15.00 / 1M
Status GA

Legacy annual Pro/Pro+ request billing only.

Multiplier 6.00×
Effective cost $0.240 / request

Legacy annual-plan multiplier.

Token offers by platform · 8 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

8 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (4)

Azure
Azure
Azure
Amazon Bedrock

Azure AI Foundry (2)

Azure AI Foundry
Azure AI Foundry

AWS Bedrock (1)

AWS Bedrock

T-Systems LLM Hub (1)

T-Systems LLM Hub
GPT-5.4 (xhigh) — benchmarks & cost | Benchmark Heaven