← All models

GPT-6 Astra (max)

★ featured
OpenAI · released 2026-09-03 · 5 offers

Output 56 tokens/sFirst token 199 sContext 1M tokens

Top 2 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Azure
2Azure AI Foundry

Composite

7 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

97.9

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 96Coding Agent v1.4: percentile 93AA Intelligence: percentile 100Epoch ECI: percentile 99Software ECI: percentile 89DesignArena: percentile 96

AA Coding 76.9Coding Agent v1.4 67.0AA Intelligence 52.8AA Agentic Epoch ECI 166.3Software ECI 163.6DesignArena 1333/1363

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

48 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-AnalystAgent9051.2%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase951,562 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2911,580 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench v2.1 (AA)9688.4%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)9959.1%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • ApprenticeBench CUA (NeoCognition)10068.0%

    ApprenticeBench CUA (NeoCognition) published 2026-09-14 · Codex — A computer-use agent operates the Odoo ERP through its screens across 100 sequential accounts-payable tasks with diminishing mentoring.

  • ApprenticeBench API (NeoCognition)10065.0%

    ApprenticeBench API (NeoCognition) published 2026-09-14 · Codex — An agent completes the same 100 sequential accounts-payable tasks in the Odoo ERP with diminishing mentoring, using dedicated application API calls instead of the screen interface.

  • Vals Index v2 (Vals AI)9366.6%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)4053.5%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • Terminal-Bench 2.1 (Vals Index v2)10087.3%

    Terminal-Bench 2.1 (Vals Index v2) v2 · Published board — Command-line interface problem solving, as run by Vals AI.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)475.4%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • EBR-bench (Earthborne Rangers, Epoch AI)10076.2%

    EBR-bench (Earthborne Rangers, Epoch AI) published 2026-09-18 · Published board — Learning-capability test: models repeatedly play the obscure campaign board game Earthborne Rangers with note-taking, and the score measures whether results improve across playthroughs.

Coding

  • Artificial Analysis Coding Agent Index v1.410067.0%

    Artificial Analysis Coding Agent Index v1.4 v1.4 · Codex — The retained AA Coding Agent Index measures coding-agent systems using the earlier three-component implementation.

  • Artificial Analysis Coding Agent Index v1.510061.6%

    Artificial Analysis Coding Agent Index v1.5 v1.5 · Codex — Measures coding-agent systems on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA.

  • SciCode (AA subproblems) v1.0.18856.5%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • Vibe Code Bench (Vals Index v2)9489.6%

    Vibe Code Bench (Vals Index v2) v2 · Published board — End-to-end app-building tasks.

  • Code Migration (Vals Index v2 subset)10068.5%

    Code Migration (Vals Index v2 subset) v2 · Published board — Porting projects to another language, including COBOL modernization.

  • FrontierCode 1.1 Main (Cognition) v1.1no percentile53.3%

    FrontierCode 1.1 Main (Cognition) v1.1 · codex — Whether a maintainer would merge the agent's pull request, on tasks crafted by open-source maintainers and graded with tests, rubrics and verifiers.

  • DeepSWE (Datacurve, via Epoch AI)9473.2%

    DeepSWE (Datacurve, via Epoch AI) published 2026-09-15 · mini-swe-agent — Pass rate of coding agents on original, long-horizon software engineering tasks, run by Datacurve with the mini-swe-agent harness.

  • VulcanBench Frontier v410089.3%

    VulcanBench Frontier v4 v4 · Codex — 23 behavioural-reconstruction tasks: the model must repair a replacement implementation of a legacy program until hidden tests confirm it reproduces the program’s real drift from its written spec; the combined score weights functional correctness 50%, lint/complexity 8.5%, security 8.5% and judged Code quality 33%.

  • AA Coding Index9676.9

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Efficiency

Knowledge

Long-context

  • GDP.pdf (AA)9931.0%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19280.7%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA8435.0%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • ArXivMath 06/2026 (MathArena) v2026-0610094.4%

    ArXivMath 06/2026 (MathArena) v2026-06 · Published board — Research-level math problems with a checkable final answer, drawn from arXiv papers submitted in June 2026, so they postdate most training data.

  • BrokenArXiv 06/2026 (MathArena) v2026-0610099.1%

    BrokenArXiv 06/2026 (MathArena) v2026-06 · Published board — Plausible but false proof statements taken from June 2026 arXiv papers; a model scores by refusing to prove them and saying the statement is false as written.

  • FrontierMath Tiers 1–3 v2 (Epoch AI)10093.7%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)9797.6%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

  • ArXivMath 08/2026 (MathArena) v2026-0810088.6%

    ArXivMath 08/2026 (MathArena) v2026-08 · Published board — Research-level math problems with a checkable final answer, drawn from arXiv papers submitted in August 2026, so they postdate most training data.

  • BrokenArXiv 08/2026 (MathArena) v2026-0810081.9%

    BrokenArXiv 08/2026 (MathArena) v2026-08 · Published board — Plausible but false statements taken from August 2026 arXiv papers; a model scores by refusing to prove them and saying the statement is false as written.

Reasoning

  • Chess Puzzles (Epoch AI)10072.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • Mystery Game Puzzles (Epoch AI)10084.0%

    Mystery Game Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 mid-game positions of a well-known game whose identity Epoch deliberately keeps undisclosed, generated programmatically like Chess Puzzles.

  • AA Intelligence Index10052.8

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Safety/Alignment

Science

Tool-use

  • AutomationBench-AA v1.0.69968.5%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • τ³-Banking (AA) v1.0.18841.4%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

  • Excel Modeling Benchmark (Vals Index v2)8471.7%

    Excel Modeling Benchmark (Vals Index v2) v2 · Published board — Building and editing financial models in spreadsheets.

Vision

Unusual results

5 threshold-crossing signals flagged · 35 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak Vals Index v2 cost per test (Vals AI)
    Why
    Observed: 19.08878 USD
    Peer mean 7.03053 · peer sd 7.84824 · n 31 · families 30
    Directed z -1.536 · baseline z 1.526 · gap -3.062 · profile n 34
    19.08878 USDmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Vals Index v2 cost per test (Vals AI) · v2 · Published board
    Exact value: 19.088776 USD
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:951ce6690d69cd50f95fe408
  • unusually strong CritPt (AA) published 2026-09-10
    Why
    Observed: 0.31714 fraction
    Peer mean 0.03928 · peer sd 0.07596 · n 522 · families 369
    Directed z 3.658 · baseline z 1.373 · gap 2.285 · profile n 34
    0.31714 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: CritPt (AA) · published 2026-09-10 · Published board
    Exact value: 0.317142857142857 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:2f339a97-9a0d-499a-9cb5-e0db665bfa25:critpt
  • unusually strong AA Intelligence Index published 2026-09-19
    Why
    Observed: 52.8 points
    Peer mean 15.94317 · peer sd 11.48893 · n 644 · families 483
    Directed z 3.208 · baseline z 1.386 · gap 1.822 · profile n 34
    52.8 pointsmeasuredobserved 2026-09-19artificialanalysis.ai
    Evidence
    Axis: AA Intelligence Index · published 2026-09-19 · Published board
    Exact value: 52.8 points
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:aa_intelligence_index:gpt-6-astra::max
    Protocol: inspect · file data/raw/artificialanalysis.json
  • unusually strong Terminal-Bench v4.0 (AA)
    Why
    Observed: 0.59091 fraction
    Peer mean 0.10006 · peer sd 0.15678 · n 149 · families 109
    Directed z 3.131 · baseline z 1.415 · gap 1.716 · profile n 34
    0.59091 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Terminal-Bench v4.0 (AA) · v4.0 · Published board
    Exact value: 0.590909090909091 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:2f339a97-9a0d-499a-9cb5-e0db665bfa25:terminalbenchV40
  • unusually strong Mystery Game Puzzles (Epoch AI) published 2026-09-18
    Why
    Observed: 0.84 fraction
    Peer mean 0.31109 · peer sd 0.17549 · n 38 · families 29
    Directed z 3.014 · baseline z 1.392 · gap 1.622 · profile n 34
    0.84 fractionmeasuredobserved 2026-09-18epoch.ai
    Evidence
    Axis: Mystery Game Puzzles (Epoch AI) · published 2026-09-18 · Published board
    Exact value: 0.84 fraction
    Observed: 2026-09-18T23:49:00Z · publication date: not recorded
    Observation id: public:55a84368f006dde9d6900b85
Missing coverage · 95 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
GPT-6 Astra (high)77.151.0
GPT-6 Astra (max)76.952.8
GPT-6 Astra (medium)76.749.7
GPT-6 Astra (xhigh)75.952.5
GPT-6 Astra (low)75.746.0

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $10.00 / 1M
Cached input $1.00 / 1M
Cache write $12.50 / 1M
Output $50.00 / 1M
Status GA
Token offers by platform · 3 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

3 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (2)

Azure
Azure

Azure AI Foundry (1)

Azure AI Foundry
GPT-6 Astra (max) — benchmarks & cost | Benchmark Heaven