← All models

Llama 3.3 Instruct 70B

open weights
Meta · released 2024-12-06 · 14 offers

Output 90 tokens/sFirst token 0.6 sContext 128K tokens

Top 5 cheapest providers (Adjusted $/task)

The same list price can give a different adjusted $/task (caching, token efficiency) — click a price for its inputs.

Within the active global provider, residency and confidentiality filters.

#ProviderAdjusted $/task
1DeepInfra
2Novita
3AkashML
4Parasail
5Groq

Composite

2 of 7 inputs6 radar axes: DesignArena's two boards share one

10.2

AA CodingCoding Agent v1.4AA IntelligenceAA AgenticEpoch ECISoftware ECIDesignArenaAA Coding: percentile 12AA Intelligence: percentile 32

AA Coding 11.9Coding Agent v1.4 AA Intelligence 7.7AA Agentic Epoch ECI Software ECI DesignArena —/—

Radar: percentile among all models measured on each input; a gap means not measured.

Benchmark sheet

20 of 140 registered benchmark versions · bars show the percentile among all models measured on each benchmark.

Compare this model →

Agentic

Coding

  • AA Coding Index1211.9

    AA Coding Index published 2026-09-19 · Published board — Artificial Analysis publishes this board without a version number; we keep the date each result was retained.

Efficiency

Instruction-following

Knowledge

Long-context

  • AA-LCR v1.12015.7%

    AA-LCR v1.1 v1.1 · Published board — Tests reasoning across multiple long documents with corrected answer keys and grading.

Math

  • AIME 2025 (AA) v2025127.7%

    AIME 2025 (AA) v2025 · Published board — Advanced mathematical problem solving on AIME I and II 2025.

Reasoning

Science

Tool-use

Uncensored

  • UGI Willingness694.00

    UGI Willingness published 2026-09-10 · Published board — A private-question evaluation of how far a model follows challenging instructions before refusing or deviating.

  • UGI7337.9

    UGI published 2026-09-10 · Published board — A private-question suite combining sensitive-topic knowledge with willingness to follow controversial instructions.

Writing

  • UGI Writing3126.2

    UGI Writing published 2026-09-10 · Published board — A maintainer-owned writing evaluation considering intelligence, style, repetition and length adherence, informed by human preferences.

Unusual results

1 threshold-crossing signal flagged · 22 eligible benchmark families

How flags are calculated

Heuristic screen, not statistical significance: benchmark families are correlated and source uncertainty is unknown. Peer evidence needs ≥ 20 independently measured matched configurations from ≥ 10 distinct model families. A flag needs a directed population z-score of magnitude ≥ 1.5 and a gap of ≥ 1.5 from the leave-one-benchmark-family-out mean z in the same direction, over ≥ 5 other benchmark families.

  • unusually weak ITBench-AA published 2026-09-10
    Why
    Observed: 0.00565 fraction
    Peer mean 0.31758 · peer sd 0.13369 · n 33 · families 31
    Directed z -2.333 · baseline z -0.649 · gap -1.684 · profile n 21
    0.00565 fractionmeasuredobserved 2026-09-10artificialanalysis.ai
    Evidence
    Axis: ITBench-AA · published 2026-09-10 · Published board
    Exact value: 0.00564971751412429 fraction
    Observed: 2026-09-10T21:47:16.627Z · publication date: not recorded
    Observation id: aa:976cc8ad-7904-4056-83c5-960181f47d5f:itBenchSre
Missing coverage · 123 benchmark versions

No result does not mean a zero, or that the model was never tested. Collection failures and disputed versions retain their distinct status.

Token offers by platform · 13 offers (Adjusted $/task)

Click any underlined price to see how it is estimated and where each input comes from. How we calculate adjusted cost.

13 offers within the active global filters; “—” means the catalog is active but no public token price is available.

OpenRouter (10)

DeepInfra
Novita
AkashML
Parasail
Groq
SambaNova
CoreWeave
Google
Google
Together

STACKIT (1)

STACKIT

OVHcloud (1)

OVHcloud

Scaleway (1)

Scaleway
Llama 3.3 Instruct 70B — benchmarks & cost | Benchmark Heaven