OpenAI Decisions (gpt-6-luna)
Data: JevBench v1.6.1, published 2026-10-06
Capability 73.5 · API offering, ranked on the API leaderboard · Capability rank #6 of 14 within the Jev-class caps
Measured 2026-10-06 on A5 ∪ P, 600 items, equated (+5.78 I / +6.15 C); round score-a5-1.
hosted API
intelligence
56.9
calibration
90.1
speed
89.3
cost
48.6
Composite (secondary)
62.5 · Composite rank #8 of 27 on the API leaderboard
Cost and measurement conditions
$0.052 per 1,000 decisions (public tariff)
measured: token usage of the 6 Oct 2026 A5 u P run x tariff USD 0.1/M input, 0.0/M output (spec openai-decisions.json; OpenAI Decisions API (POST /v1/decisions, gpt-6-luna) list price USD 0.10 per 1M input tokens, no output/cache charges; source https://developers.openai.com/api/docs/guides/decisions (checked 2026-10-06); actual usage), per 1,000 decisions on the v1.6.1 common cost basis (592 of 600 items; the 8 A5 u P items outside the Jev reference's accepted input range are excluded, cost-basis-a5.json). List price as of 6 Oct 2026 (launch day); re-checked at each revision.
p50 latency: 0.3 s
none (API)
Breakdowns (not part of the score)
Chance-corrected competence (0 = chance, 100 = perfect) on A5 ∪ P (600 items: 300 open + 300 sealed; fewer sealed items than the full set, so wider uncertainty). Raw values, not equated. Category cells use A5+P+L1+L2+L3; per-type/tier cells use the original basis above. Radars against Jev.
| Request type | Open | Sealed |
|---|---|---|
| choice | 72.6 | 69.2 |
| noul | 48.9 | 20.6 |
| score | 52.7 | 42.6 |
Capability by subject topic
| Rules, policy & law | 45.2 | n=1154 |
|---|---|---|
| Coding & software | 65.1 | n=167 |
| Math & numbers | 1.9 | n=286 |
| Finance & commerce | 43.7 | n=181 |
| Support & operations | 50.5 | n=108 |
| Everyday language | 75.2 | n=120 |
| Safety & security | 47.1 | n=89 |
Use cases (TypeSafe categories)
| Model routing | 63.7 | n=212 |
|---|---|---|
| Legal & compliance | 54.3 | n=287 |
| Customer support | 48.1 | n=155 |
| E-commerce | 37.5 | n=92 |
| Insurance claims | 34.7 | n=94 |
| Risk assessment | 49.6 | n=89 |
| Financial crime | 24.2 | n=75 |
| Feature extraction | 32.3 | n=87 |
| Lead generation | 58.0 | n=92 |
| Recruiting | 19.0 | n=83 |
| Knowledge graphs | 42.9 | n=82 |
| LLM guardrails | 52.8 | n=84 |
| Moderation | 67.6 | n=73 |
| Code linting | 57.6 | n=82 |
| Search & retrieval | 77.5 | n=82 |
| Science | 24.5 | n=80 |
| Advertising | 12.1 | n=85 |
| Gaming | 8.5 | n=85 |
| Demand forecasting | 0.0 | n=75 |
Cells need at least 15 answered items; † fewer than 30 (indicative only, listed below the radar in the compare view).