OpenAI Decisions (gpt-6-luna)

Data: JevBench v1.6.1, published 2026-10-06

Capability 73.5 · API offering, ranked on the API leaderboard · Capability rank #6 of 14 within the Jev-class caps

Measured 2026-10-06 on A5 ∪ P, 600 items, equated (+5.78 I / +6.15 C); round score-a5-1.

hosted API

Base model: undisclosed

intelligence

56.9

calibration

90.1

speed

89.3

cost

48.6

Composite (secondary)

62.5 · Composite rank #8 of 27 on the API leaderboard

Cost and measurement conditions

$0.052 per 1,000 decisions (public tariff)

measured: token usage of the 6 Oct 2026 A5 u P run x tariff USD 0.1/M input, 0.0/M output (spec openai-decisions.json; OpenAI Decisions API (POST /v1/decisions, gpt-6-luna) list price USD 0.10 per 1M input tokens, no output/cache charges; source https://developers.openai.com/api/docs/guides/decisions (checked 2026-10-06); actual usage), per 1,000 decisions on the v1.6.1 common cost basis (592 of 600 items; the 8 A5 u P items outside the Jev reference's accepted input range are excluded, cost-basis-a5.json). List price as of 6 Oct 2026 (launch day); re-checked at each revision.

p50 latency: 0.3 s

none (API)

Breakdowns (not part of the score)

Chance-corrected competence (0 = chance, 100 = perfect) on A5 ∪ P (600 items: 300 open + 300 sealed; fewer sealed items than the full set, so wider uncertainty). Raw values, not equated. Category cells use A5+P+L1+L2+L3; per-type/tier cells use the original basis above. Radars against Jev.

Request typeOpenSealed
choice72.669.2
noul48.920.6
score52.742.6

Capability by subject topic

Rules, policy & law45.2n=1154
Coding & software65.1n=167
Math & numbers1.9n=286
Finance & commerce43.7n=181
Support & operations50.5n=108
Everyday language75.2n=120
Safety & security47.1n=89

Use cases (TypeSafe categories)

Model routing63.7n=212
Legal & compliance54.3n=287
Customer support48.1n=155
E-commerce37.5n=92
Insurance claims34.7n=94
Risk assessment49.6n=89
Financial crime24.2n=75
Feature extraction32.3n=87
Lead generation58.0n=92
Recruiting19.0n=83
Knowledge graphs42.9n=82
LLM guardrails52.8n=84
Moderation67.6n=73
Code linting57.6n=82
Search & retrieval77.5n=82
Science24.5n=80
Advertising12.1n=85
Gaming8.5n=85
Demand forecasting0.0n=75

Cells need at least 15 answered items; † fewer than 30 (indicative only, listed below the radar in the compare view).