← All models

Grok 4.3 (high)

deprecated by benchmark source
xAI · released 2026-04-30 · 5 offers

Context 1M tokens

Top 3 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1AWS Bedrock
2Azure AI Foundry
3Google Vertex AI

Composite

4 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

46.7

includes −8.1 for its Benchmaxxing signal (from 54.9; why, switch off in Options)

Benchmaxxing signal +8.1, medium Benchmaxxing tag, uncertain: The 80 % interval reaches below zero — treat this tag as uncertain · medium report →The 80 % interval reaches below zero — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 47AA Intelligence: percentile 81AA Agentic: percentile 50DesignArena: percentile 16

AA Coding 42.2Coding Agent v1.4 AA Intelligence 25.4AA Agentic 17.2Epoch ECI Software ECI DesignArena 1144/1017

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

38 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-AnalystAgent118.8%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • APEX-Agents-AA4117.0%

    APEX-Agents-AA published 2026-09-10 · Published board — Tests professional-service tasks that require agents to produce locally graded deliverables.

  • AA-Briefcase37762 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2491,018 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Harvey LAB-AA1768.4%

    Harvey LAB-AA published 2026-09-10 · Published board — Tests legal-work deliverables across practice areas on Harvey's private task set. Artificial Analysis' run, graded by one LLM judge against task rubrics — not the same run or scale as Vals AI's HLAB row.

  • ITBench-AA4732.7%

    ITBench-AA published 2026-09-10 · Published board — Tests root-cause diagnosis from offline Kubernetes incident snapshots.

  • Terminal-Bench Hard (AA)8537.9%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)4139.7%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)180.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • Vals Index v2 (Vals AI)024.3%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)737.7%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • Terminal-Bench 2.1 (Vals Index v2)041.9%

    Terminal-Bench 2.1 (Vals Index v2) v2 · Published board — Command-line interface problem solving, as run by Vals AI.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)30.4%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • AA Agentic Index5017.2

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

Efficiency

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)345.8%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.17373.0%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA3510.0%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Math

  • FrontierMath Tiers 1–3 v2 (Epoch AI)1542.8%

    FrontierMath Tiers 1–3 v2 (Epoch AI) v2 · Published board — Unpublished, expert-written mathematics problems from undergraduate to research level with automatically checkable answers, run by Epoch AI on its private v2 set.

  • FrontierMath Tier 4 v2 (Epoch AI)614.6%

    FrontierMath Tier 4 v2 (Epoch AI) v2 · Published board — The hardest, research-level tier of Epoch AI's unpublished FrontierMath problems, run by Epoch AI on its private v2 set.

Reasoning

  • Chess Puzzles (Epoch AI)5325.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • AA Intelligence Index8125.4

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

Vision

Unusual results

8 threshold-crossing signals flagged · 41 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak Excel Modeling Benchmark (Vals Index v2)
    Why
    Observed: 18.394 percent
    Peer mean 59.33688 · peer sd 13.75792 · n 33 · families 32
    Directed z -2.976 · baseline z -0.32 · gap -2.656 · profile n 40
    18.394 percentmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Excel Modeling Benchmark (Vals Index v2) · v2 · Published board
    Exact value: 18.394 percent
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:d33d0caa222b3557aa7be3b6
  • unusually strong IFBench (AA single-turn) published 2026-09-10
    Why
    Observed: 0.81293 fraction
    Peer mean 0.48006 · peer sd 0.17139 · n 450 · families 331
    Directed z 1.942 · baseline z -0.443 · gap 2.385 · profile n 40
    0.81293 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: IFBench (AA single-turn) · published 2026-09-10 · Published board
    Exact value: 0.812925170068027 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:948892b5-db03-4118-a4a8-ccd51ed871ea:ifbench
  • unusually weak Vals Index v2 (Vals AI)
    Why
    Observed: 24.294 percent
    Peer mean 53.34655 · peer sd 11.98829 · n 31 · families 30
    Directed z -2.423 · baseline z -0.334 · gap -2.089 · profile n 40
    24.294 percentmeasured· 0.68 USD per rolloutobserved 2026-09-18vals.ai
    Evidence
    Axis: Vals Index v2 (Vals AI) · v2 · Published board
    Exact value: 24.294 percent
    Published mean cost: 0.676033 USD per rollout
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:17280b6726e7194773be8958
  • unusually strong AA-Omniscience Index published 2026-09-10
    Why
    Observed: 17.98333 points
    Peer mean -30.70814 · peer sd 31.08454 · n 520 · families 366
    Directed z 1.566 · baseline z -0.434 · gap 2 · profile n 40
    17.98333 pointsmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: AA-Omniscience Index · published 2026-09-10 · Published board
    Exact value: 17.9833333333333 points
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:948892b5-db03-4118-a4a8-ccd51ed871ea:omniscience
  • unusually strong Humanity's Last Exam (AA text-only) published 2026-09-10
    Why
    Observed: 0.3721 fraction
    Peer mean 0.15283 · peer sd 0.14006 · n 608 · families 449
    Directed z 1.566 · baseline z -0.434 · gap 1.999 · profile n 40
    0.3721 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: Humanity's Last Exam (AA text-only) · published 2026-09-10 · Published board
    Exact value: 0.372103799814643 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:948892b5-db03-4118-a4a8-ccd51ed871ea:hle
  • unusually weak Terminal-Bench 2.1 (Vals Index v2)
    Why
    Observed: 41.948 percent
    Peer mean 69.13129 · peer sd 11.92395 · n 31 · families 30
    Directed z -2.28 · baseline z -0.338 · gap -1.942 · profile n 40
    41.948 percentmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Terminal-Bench 2.1 (Vals Index v2) · v2 · Published board
    Exact value: 41.948 percent
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:87aeccfc7a3a40345d137406
  • unusually weak Vibe Code Bench (Vals Index v2)
    Why
    Observed: 19.403 percent
    Peer mean 65.32897 · peer sd 21.57746 · n 32 · families 31
    Directed z -2.128 · baseline z -0.341 · gap -1.787 · profile n 40
    19.403 percentmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Vibe Code Bench (Vals Index v2) · v2 · Published board
    Exact value: 19.403 percent
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:916dbc188e572c2cb9048bf8
  • unusually weak Code Migration (Vals Index v2 subset)
    Why
    Observed: 5.923 percent
    Peer mean 35.5897 · peer sd 15.81654 · n 33 · families 32
    Directed z -1.876 · baseline z -0.348 · gap -1.528 · profile n 40
    5.923 percentmeasuredobserved 2026-09-18vals.ai
    Evidence
    Axis: Code Migration (Vals Index v2 subset) · v2 · Published board
    Exact value: 5.923 percent
    Observed: 2026-09-18T13:00:43.826192+00:00 · publication date: not recorded
    Observation id: public:81478c858172353b8b0a2dd2
Missing coverage · 103 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Grok 4.3 (high)42.225.4
Grok 4.3 (Non-reasoning)35.214.5
Grok 4.3 (medium)24.8
Grok 4.3 (low)24.3
Token offers by platform · 3 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

3 offers within the active global filters; “—” means the catalog is active but no public token price is available.

AWS Bedrock (1)

AWS Bedrock

Azure AI Foundry (1)

Azure AI Foundry

Google Vertex AI (1)

Google Vertex AI
Grok 4.3 (high) — benchmarks & cost | Benchmark Heaven