← All models

Grok 4.3 (Non-reasoning)

deprecated by benchmark source
xAI · released 2026-04-30 · 5 offers

Output 100 tokens/sFirst token 0.6 sContext 1M tokens

Top 3 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1AWS Bedrock
2Azure AI Foundry
3Google Vertex AI

Composite

4 of 7 inputs · 2 from the model family6 radar axes: DesignArena's two boards share one

37.8

includes −8.1 for its Benchmaxxing signal (from 45.9; why, switch off in Options)

Benchmaxxing signal +8.1, medium Benchmaxxing tag, uncertain: The 80 % interval reaches below zero — treat this tag as uncertain · medium report →The 80 % interval reaches below zero — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 41AA Intelligence: percentile 60DesignArena: percentile 16

AA Coding 35.2Coding Agent v1.4 AA Intelligence 14.5AA Agentic Epoch ECI Software ECI DesignArena 1144/1017

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

17 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-Briefcase34724 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2491,031 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench Hard (AA)5918.9%

    Terminal-Bench Hard (AA) pinned revision 74221fb · Published board — Tests a pinned 44-task hard subset of terminal-based work using Terminus 2.

  • Terminal-Bench v2.1 (AA)3734.1%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)180.0%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

Coding

  • SciCode (AA subproblems) v1.0.12739.4%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • AA Coding Index4135.2

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Instruction-following

Knowledge

Long-context

  • GDP.pdf (AA)242.8%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.13432.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Reasoning

Science

Tool-use

  • AutomationBench-AA v1.0.6160.9%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • τ²-Bench Telecom (AA)6065.8%

    τ²-Bench Telecom (AA) published 2026-09-10 · Published board — Tests dual-control telecom agents that coordinate tool use with a simulated user.

  • τ³-Banking (AA) v1.0.1238.0%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

Vision

Missing coverage · 126 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Grok 4.3 (high)42.225.4
Grok 4.3 (Non-reasoning)35.214.5
Grok 4.3 (medium)24.8
Grok 4.3 (low)24.3
Token offers by platform · 3 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

3 offers within the active global filters; “—” means the catalog is active but no public token price is available.

AWS Bedrock (1)

AWS Bedrock

Azure AI Foundry (1)

Azure AI Foundry

Google Vertex AI (1)

Google Vertex AI
Grok 4.3 (Non-reasoning) — benchmarks & cost | Benchmark Heaven