← All models

Grok 4.6 (high)

★ featured
xAI · released 2026-08-12 · 5 offers

Output 55 tokens/sFirst token 32 sContext 500K tokens

Top 3 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1Amazon Bedrock
2Google Vertex AI
3AWS Bedrock

Composite

6 of 7 inputs · 4 from the model family6 radar axes: DesignArena's two boards share one

84.3

includes −5.6 for its Benchmaxxing signal (from 89.9; why, switch off in Options)

Benchmaxxing signal +5.6, light Benchmaxxing tag, uncertain: The 80 % interval reaches below zero — treat this tag as uncertain · light report →The 80 % interval reaches below zero — treat this tag as uncertain.

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 96AA Intelligence: percentile 97AA Agentic: percentile 97Epoch ECI: percentile 76Software ECI: percentile 74DesignArena: percentile 74

AA Coding 76.8Coding Agent v1.4 AA Intelligence 44.4AA Agentic 53.4Epoch ECI 156.3Software ECI 159.6DesignArena 1264/1273

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

38 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Composite attachments (used in the score, not counted as exact benchmarks):
  • Epoch ECI · attached
  • Software ECI · attached
  • DesignArena Web Apps (agentic) · attached
  • DesignArena Full-Stack · attached
Compare this model →

Agentic

  • AA-AnalystAgent6341.3%

    AA-AnalystAgent published 2026-09-10 · Published board — Tests analyst tasks using agentic Python execution across fourteen domains.

  • AA-Briefcase931,534 Elo

    AA-Briefcase published 2026-09-10 · Published board — Tests multi-week professional knowledge-work projects with linked tasks and large source collections.

  • GDPval-AA v2951,643 Elo

    GDPval-AA v2 v2 · Published board — Tests professional knowledge-work deliverables across occupations using AA's Stirrup harness.

  • Terminal-Bench v2.1 (AA)9688.4%

    Terminal-Bench v2.1 (AA) v2.1 · Published board — Tests terminal-based work on the 89-task verified refresh using Terminus 2.

  • Terminal-Bench v4.0 (AA)8321.2%

    Terminal-Bench v4.0 (AA) v4.0 · Published board — Tests terminal-based work on the 66-task release using mini-SWE-agent v2.4.6.

  • Vals Index v2 (Vals AI)5759.2%

    Vals Index v2 (Vals AI) v2 · Published board — GDP-weighted average of agentic model performance across finance, coding and legal tasks.

  • Finance Agent v2 (Vals Index v2)4353.7%

    Finance Agent v2 (Vals Index v2) v2 · Published board — Multi-step financial reasoning tasks.

  • Terminal-Bench 2.1 (Vals Index v2)7778.3%

    Terminal-Bench 2.1 (Vals Index v2) v2 · Published board — Command-line interface problem solving, as run by Vals AI.

  • HLAB — Harvey's Legal Agent Benchmark (Vals Index v2)8815.8%

    HLAB — Harvey's Legal Agent Benchmark (Vals Index v2) v2 · Published board — Long-horizon legal work product creation. Vals AI's run of Harvey's Legal Agent Benchmark, reported as accuracy — not comparable with Artificial Analysis' Harvey LAB-AA row.

  • τ²-Bench Airline (OpenRouter run)8176.0%

    τ²-Bench Airline (OpenRouter run) published 2026-09-15 · Published board — Multi-turn service agents making tool calls under strict policy constraints. (OpenRouter's own reproducible run, as published on openrouter.ai/benchmarks.)

  • AA Agentic Index9753.4

    AA Agentic Index published 2026-09-18 · Published board — Artificial Analysis publishes one Agentic Index per model, measured on its primary configuration and relayed by OpenRouter's Benchmarks API; the value is attached at family scope on the deterministic representative. Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Coding

  • SciCode (AA subproblems) v1.0.18856.5%

    SciCode (AA subproblems) v1.0.1 · Published board — Tests scientific Python programming with scientist-annotated background information.

  • CursorBench 4.0 (Cursor) v4.0no percentile40.4%

    CursorBench 4.0 (Cursor) v4.0 · Published board — Agent evaluation on ambiguous, multi-file tasks drawn from real Cursor sessions.

  • Vibe Code Bench (Vals Index v2)5876.2%

    Vibe Code Bench (Vals Index v2) v2 · Published board — End-to-end app-building tasks.

  • Code Migration (Vals Index v2 subset)6944.6%

    Code Migration (Vals Index v2 subset) v2 · Published board — Porting projects to another language, including COBOL modernization.

  • FrontierCode 1.1 Main (Cognition) v1.1no percentile48.0%

    FrontierCode 1.1 Main (Cognition) v1.1 · grok-build — Whether a maintainer would merge the agent's pull request, on tasks crafted by open-source maintainers and graded with tests, rubrics and verifiers.

  • DeepSWE (Datacurve, via Epoch AI)5665.2%

    DeepSWE (Datacurve, via Epoch AI) published 2026-09-15 · mini-swe-agent — Pass rate of coding agents on original, long-horizon software engineering tasks, run by Datacurve with the mini-swe-agent harness.

  • AA Coding Index9676.8

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

  • DesignArena Web Apps (agentic)811,264 Elo

    DesignArena Web Apps (agentic) published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

  • DesignArena Full-Stack831,273 Elo

    DesignArena Full-Stack published 2026-09-19 · Published board — Published Elo. Exact source attachment is retained; fewer than 200 battles excludes a row from radar normalization and anomaly peers.

Efficiency

Knowledge

Long-context

  • GDP.pdf (AA)6717.0%

    GDP.pdf (AA) published 2026-09-10 · Published board — Tests professional reasoning over long PDFs with AA document preparation and grading.

  • AA-LCR v1.19180.3%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

  • MLCR-AA4312.2%

    MLCR-AA published 2026-09-10 · Published board — Tests medical-record synthesis and reasoning across long, fragmented case documents.

Reasoning

  • Chess Puzzles (Epoch AI)7640.0%

    Chess Puzzles (Epoch AI) published 2026-09-18 · Published board — Best-move selection on 100 novel chess positions generated programmatically by Epoch AI, each with a single Stockfish-verified best move; probes spatial reasoning and planning.

  • AA Intelligence Index9744.4

    AA Intelligence Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Science

Tool-use

  • AutomationBench-AA v1.0.69766.7%

    AutomationBench-AA v1.0.6 · Published board — Tests multi-app SaaS workflows through REST tools on a held-out AutomationBench split.

  • EnterpriseOps-Gym-AA9148.3%

    EnterpriseOps-Gym-AA published 2026-09-10 · Published board — Tests enterprise workflows through MCP tools against resettable application environments.

  • τ³-Banking (AA) v1.0.110050.7%

    τ³-Banking (AA) v1.0.1 · Published board — Tests banking support agents that retrieve policies and change account state through tools.

  • Excel Modeling Benchmark (Vals Index v2)4762.7%

    Excel Modeling Benchmark (Vals Index v2) v2 · Published board — Building and editing financial models in spreadsheets.

Unusual results

1 threshold-crossing signal flagged · 37 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually strong AA Intelligence Index published 2026-09-19
    Why
    Observed: 44.4 points
    Peer mean 15.94317 · peer sd 11.48893 · n 644 · families 483
    Directed z 2.477 · baseline z 0.886 · gap 1.591 · profile n 36
    44.4 pointsmeasuredobserved 2026-09-19artificialanalysis.ai
    Evidence
    Axis: AA Intelligence Index · published 2026-09-19 · Published board
    Exact value: 44.4 points
    Observed: 2026-09-19 · publication date: not recorded
    Observation id: legacy:aa_intelligence_index:grok-4.6::high
    Protocol: inspect · file data/raw/artificialanalysis.json
Missing coverage · 103 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Variants / reasoning settings

Artificial Analysis snapshot 2026-09-19 · Data: Artificial Analysis · Data: Epoch AI (CC BY)

VariantAA CodingAA Intelligence
Grok 4.6 (high)76.844.4
Grok 4.6 (xhigh)75.944.3
Grok 4.6 (medium)74.443.0
Grok 4.6 (low)66.335.4

SUBSCRIPTION PLAN

GitHub Copilot

Current usage-based billing · model token cost is converted to AI Credits at 1 credit = $0.01.

Input $2.00 / 1M
Cached input $0.500 / 1M
Output $6.00 / 1M
Status GA
Token offers by platform · 3 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

3 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (1)

Amazon Bedrock

Google Vertex AI (1)

Google Vertex AI

AWS Bedrock (1)

AWS Bedrock
Grok 4.6 (high) — benchmarks & cost | Benchmark Heaven