JevBench · board v1.7.43 · updated 2026-10-10

JevBench — hosted decision API benchmark

A benchmark for AI decision models: state and rubric in, typed answer out. We compare accuracy, calibration, latency and cost, independently of TypeSafe AI.

27 ranked API offerings · 1,500 decisions each (600 for 19 equated API re-runs) · How it works · Release history · Methodology · Data & JSON/CSV · aggregate JSON · Share · Other decision benchmarks · Explore ImageJevBench v0.3.0

Show all mixedSpeed and cost are compared within a group: open weights on the same GPU, APIs as sold by each provider.

JevBench board · headline

JevBench Composite Score (API offerings): 27 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

Weights:
Adjust weights ↓

View by:Capability ↓

The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis.

Greener = stronger within its column.

27 of 27 systems, sorted by official rank, #1 first.

  1. 1Sage 1.3.0APInewClosed APIBase model: undisclosed74.0I 66C 92S 94K 58est.$0.025
  2. 2d1APInewClosed APIBase model: undisclosed73.0I 62C 87S 89K 63$0.017
  3. 3Mercury DecideAPInewSystemBase model: undisclosedsource72.4I 66C 81S 85K 62est.$0.018
  4. 4Jev 1.13.0APIJev referenceBase model: undisclosed71.5I 64C 91S 91K 55$0.032
  5. 5wity-1APInewClosed APIBase model: undisclosed70.8I 70C 88S 72K 59$0.024
  6. 6Microsoft-Decision-1APInewClosed APIBase model: Qwen3.5-9Bsource*169.1I 57C 84S 82K 61$0.019
  7. 7SPX-CD FlashAPInewSystemBase model: undisclosedsource66.7I 58C 90S 79K 52est.$0.039
  8. 8OpenAI DecisionsAPInewClosed APIBase model: undisclosed62.5I 57C 90S 89K 49$0.052
  9. 9InstinctAPIClosed APIBase model: Qwen3.8-27Bsource*262.2I 48C 93S 87K 64$0.015
  10. 10SPX-CD ProAPInewSystemBase model: undisclosedsource47.7I 65C 92S 78K 43est.$0.079
  11. 11Vansa-3.4APIClosed APIBase model: Qwen3.5-4B (self-reported)*342.6I 41C 87S 95K 62$0.019
  12. 12GPT-6 Luna (low reasoning effort)APIClosed APIBase model: undisclosed40.5I 97C 99S 72K 39$0.108
  13. 13Autoloops – Gemma 4 31B ITAPIClosed APIBase model: undisclosed40.2I 58C 87S 85K 40$0.096
  14. 14GPT-6 Luna (default medium reasoning effort)APIClosed APIBase model: undisclosed39.2I 97C 100S 72K 38$0.113
  15. 15Fastino GLiDEAPInewClosed APIBase model: undisclosed30.6I 66C 88S 77K 36$0.137
  16. 16SimpleJev Qwen3.6-35B-A3BAPI35B-A3BBase model: undisclosedsource27.6I 56C 80S 77K 35est.$0.145
  17. 17GPT-5.6 LunaAPIClosed APIBase model: undisclosed21.4I 94C 95S 73K 30$0.213
  18. 18Instinct Dual 4BAPIClosed APIBase model: Qwen3.5-4Bsource*418.0I 28C 86S 89K 75$0.0068
  19. 19Gemini 3.1 Flash-LiteAPIClosed APIBase model: undisclosed16.9I 59C 61S 79K 29$0.230
  20. 20system-one-openAPIadapter · 2BBase model: Gemma 4 E2Bsource16.4I 28C 76S 79K 68est.$0.011
Show all 27 systems (7 more ranked)
  1. 21openjev-sglangAPI35B-A3BBase model: Qwen/Qwen3.6-35B-A3Bsource15.4I 39C 79S 78K 36est.$0.140
  2. 22DeepSeek V4.1 FlashAPILLM decoderBase model: undisclosed7.1I 95C 100S 68K 20$0.480
  3. 23SimpleJev Qwen3.8-27BAPI27BBase model: Qwen/Qwen3.8-27Bsource3.2I 53C 90S 77K 15est.$0.687
  4. 24decision-machine-1APIClosed APIBase model: undisclosed2.4I 13C 79S 93K 58$0.024
  5. 25JevActAPIClosed APIBase model: undisclosedsource1.1I 10C 47S 76K 69est.$0.011
  6. 26Fastino GLiNER-2.5-DecideAPInewfine-tune · 340MBase model: fastino/gliner2-large-v1source 1source 2*50.3I 7C 62S 85K 46$0.064
  7. 27Qwen3.8 27BAPI27BBase model: undisclosed0.0I 96C 100S 57K 0est.$2.178

Base-model notes

  1. *1 Microsoft-Decision-1 (base model: Qwen3.5-9B): Base model disclosed by Microsoft; the Decision-1 endpoint is evaluated as a hosted API offering.
  2. *2 Instinct (base model: Qwen3.8-27B): Operator-reported; weights not publicly verifiable.
  3. *3 Vansa-3.4 (base model: Qwen3.5-4B (self-reported)): Developer-reported to Benchmark Heaven (private correspondence, 1 Oct 2026): built on Qwen3.5-4B via an open decision fine-tune, with Vansa's own adapters; the intermediate fine-tune is not named. Not publicly documented or independently verified.
  4. *4 Instinct Dual 4B (base model: Qwen3.5-4B): Operator-reported; weights not publicly verifiable.
  5. *5 Fastino GLiNER-2.5-Decide (base model: fastino/gliner2-large-v1): Measured via Fastino's hosted API. A 340M English classifier (8k context), not built for multi-step reasoning: near chance on Noul and Score items; 23 very long items exceed its context. Our request follows Fastino's documented format. The open checkpoint scores 11.2 on the community Decision Index (rank 53 of 70; Jev 57.9). Fastino's larger GLiDE is a separate model, ranked separately on the API board.
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev — reference (TypeSafe, closed) (1)
  • Closed API (weights not public) (16)
  • Open weights · LLM decoder (6)
  • Open weights · encoder / classifier (1)
  • System (router / cascade / ensemble) (3)
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

JevBench board

JevBench Capability Score (API offerings)

Capability ranking of Jev-class systems

Capability Score averages Intelligence and Calibration. Sage 1.3.0 leads the Jev-class systems with 78.6.

Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘

Adjust cost / latency caps · 2× official
Official
#SystemICCap.$/1k
  1. 1Sage 1.3.0API65.658.278.6$0.025*
  2. 2Jev 1.13.0API63.654.777.1$0.032
  3. 3d1API62.162.874.4$0.017
  4. 4SPX-CD FlashAPI58.452.174.3$0.039*
  5. 5Mercury DecideAPI65.962.173.6$0.018*
  6. 6OpenAI DecisionsAPI56.948.673.5$0.052
  7. 7Microsoft-Decision-1API57.261.570.8$0.019
  8. 8InstinctAPI47.864.470.5$0.015
  9. 9Vansa-3.4API40.961.963.7$0.019
  10. 10Instinct Dual 4BAPI28.375.057.2$0.0068
Show all 14 Jev-class systems (4 more)
  1. 11system-one-openAPI27.968.351.9$0.011*
  2. 12decision-machine-1API13.158.346.1$0.024
  3. 13Fastino GLiNER-2.5-DecideAPI7.045.834.6$0.064
  4. 14JevActAPI10.168.628.8$0.011*

Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.

Cost and latency correlate here (Spearman ρ = 0.66, n = 27).

  • Jev — reference (TypeSafe, closed) (1)
  • Closed API (weights not public) (16)
  • Open weights · LLM decoder (6)
  • Open weights · encoder / classifier (1)
  • System (router / cascade / ensemble) (3)
  • green: ≤ reference
  • amber: ≤ cap (2× reference)
  • red: > cap
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).

  1. –GPT-6 Luna (medium)API96.938.498.4$0.11Outside: cost 3.5× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
  2. –Qwen3.8 27BAPI96.40.098.2$2.18*Outside: cost 67.4× Jev (v1.5 reference), latency 11.5× Jev (v1.5 reference)
  3. –GPT-6 Luna (low)API96.639.097.8$0.11Outside: cost 3.3× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
  4. –DeepSeek V4.1 FlashAPI94.719.697.4$0.48Outside: cost 14.9× Jev (v1.5 reference), latency 3.3× Jev (v1.5 reference)
  5. –GPT-5.6 LunaAPI94.130.294.5$0.21Outside: cost 6.6× Jev (v1.5 reference), latency 2.6× Jev (v1.5 reference)
  6. –wity-1API70.358.879.1$0.024Outside: latency 2.6× Jev (v1.5 reference)
  7. –SPX-CD ProAPI64.943.178.6$0.079*Outside: cost 2.4× Jev (v1.5 reference)
  8. –Fastino GLiDEAPI66.135.977.0$0.14Outside: cost 4.3× Jev (v1.5 reference)
  9. –Autoloops – Gemma 4 31B ITAPI58.140.572.7$0.096Outside: cost 3.0× Jev (v1.5 reference)
  10. –SimpleJev Qwen3.8-27BAPI52.914.971.6$0.69*Outside: cost 21.3× Jev (v1.5 reference)
  11. –SimpleJev Qwen3.6-35B-A3BAPI56.335.168.3$0.15*Outside: cost 4.5× Jev (v1.5 reference), latency 2.03× Jev (v1.5 reference)
  12. –Gemini 3.1 Flash-LiteAPI58.629.259.8$0.23Outside: cost 7.1× Jev (v1.5 reference)
  13. –openjev-sglangAPI38.835.658.7$0.14*Outside: cost 4.3× Jev (v1.5 reference), latency 2.1× Jev (v1.5 reference)

Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 14 of 27 systems qualify; the other 13, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.

Capability against cost and speed (API offerings)

Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 14 systems. Upper right is best: more capable and cheaper.2030405060708090100$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev (v1.5 reference)← priciercheaper →1. Sage 1.3.02. Jev 1.13.03. d14. SPX-CD Flash5. Mercury Decide
14 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 14 systems. Upper right is best: more capable and faster.203040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev (v1.5 reference)← slowerfaster →1. Sage 1.3.02. Jev 1.13.03. d14. SPX-CD Flash5. Mercury Decide
14 systems. Tap a bubble for its values.
  • Jev — reference (TypeSafe, closed) (1)
  • Closed API (weights not public) (16)
  • Open weights · LLM decoder (6)
  • Open weights · encoder / classifier (1)
  • System (router / cascade / ensemble) (3)
  • faint = outside Jev-class

Pareto frontier: capability against cost and speed (API offerings)

The red line joins the systems nobody beats on both axes at once (provider list price, measured endpoint latency); hover or tap a point for the system that beats it. Shaded: beyond the Jev-class cap.

Capability vs cost ($ per 1,000 decisions)
Outside the Jev-class cap (2× Jev)2030405060708090100$0.001$0.003$0.01$0.03$0.1$0.3$1$3$10$ per 1,000 decisions (list price) (log scale) — further left is betterCapability ScoreGPT-6 Luna (medium)GPT-6 Luna (low)wity-1d1InstinctInstinct Dual 4B

6 of 27 systems on the frontier: no other system is both cheaper and more capable.

Capability vs speed (median latency)
Outside the Jev-class cap (2× Jev)2030405060708090100100 ms200 ms500 ms1.0 s2.0 s5.0 s10 sMedian endpoint latency (log scale) — further left is betterCapability ScoreGPT-6 Luna (medium)GPT-5.6 Lunawity-1SPX-CD ProSage 1.3.0Vansa-3.4

6 of 27 systems on the frontier: no other system is both faster and more capable.

Every API offering we have measured (31)

No API offering is left out. “API offering” means an endpoint we do not run ourselves: vendor APIs and author-hosted demos (tagged). Ranked rows outside the Jev-class cost or latency caps (2× Jev) keep their Composite rank and appear below the divider of the Capability ranking. Only rows measured on v1.6.1 are ranked; every offering is measured on the v1.6.1 scale, on the full 1,500-item set or on A4 ∪ P (600 items, equated); none keeps an older score.

All hosted API offerings measured on JevBench, by status
SystemStatusCompositeCapabilityCost / 1,000Median latency
Ranked on the v1.6.1 scale (full set, or A4 ∪ P re-run equated) (27)
Sage 1.3.0 (Levanto Labs)hosted APIComposite #1full set (1,500 items)74.078.6est.$0.02470.15 s
d1 (Liquid AI)hosted APIComposite #2full set (1,500 items)73.074.4$0.01730.27 s
Mercury Decide (Inception; System One decisions API, served free on OpenRouter as inception/mercury-decide:free)hosted APIComposite #3full set (1,500 items)72.473.6est.$0.01840.33 s
Jev 1.13.0 (TypeSafe AI)hosted APIComposite #4full set (1,500 items)71.577.1$0.03230.24 s
wity-1 (Wity, reasoning auto)hosted APIComposite #5full set (1,500 items)outside Jev-class caps: latency 2.6× Jev (v1.5 reference)70.879.1$0.02361.57 s
Microsoft-Decision-1 (Azure Foundry)hosted APIComposite #6full set (1,500 items)69.170.8$0.01930.46 s
SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)hosted APIComposite #7full set (1,500 items)66.774.3est.$0.03950.98 s
OpenAI Decisions (gpt-6-luna)hosted APIComposite #8A5 ∪ P (600 items, equated)62.573.5$0.05170.30 s
Instinct (ZooWork, Qwen3.8-27B)hosted APIComposite #9A4 ∪ P (600 items, equated)62.270.5$0.01540.29 s
SPX-CD Pro (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)hosted APIComposite #10full set (1,500 items)outside Jev-class caps: cost 2.4× Jev (v1.5 reference)47.778.6est.$0.07901.13 s
Vansa-3.4 (Vansa, hosted System One API)hosted APIComposite #11A4 ∪ P (600 items, equated)42.663.7$0.01860.15 s
GPT-6 Luna (low reasoning effort)hosted APIComposite #12A4 ∪ P (600 items, equated)outside Jev-class caps: cost 3.3× Jev (v1.5 reference); latency 3.0× Jev (v1.5 reference)40.597.8$0.10811.85 s
Autoloops – Gemma 4 31B IThosted APIComposite #13A4 ∪ P (600 items, equated)outside Jev-class caps: cost 3.0× Jev (v1.5 reference)40.272.7$0.09650.58 s
GPT-6 Luna (default medium reasoning effort)hosted APIComposite #14A4 ∪ P (600 items, equated)outside Jev-class caps: cost 3.5× Jev (v1.5 reference); latency 3.0× Jev (v1.5 reference)39.298.4$0.11301.82 s
Fastino GLiDEhosted APIComposite #15A4 ∪ P (600 items, equated)outside Jev-class caps: cost 4.3× Jev (v1.5 reference)30.677.0$0.13740.74 s
SimpleJev Qwen3.6-35B-A3Bauthor-hosted demoComposite #16A4 ∪ P (600 items, equated)outside Jev-class caps: cost 4.5× Jev (v1.5 reference); latency 2.03× Jev (v1.5 reference)27.668.3est.$0.14511.25 s
GPT-5.6 Luna (low reasoning effort)hosted APIComposite #17A4 ∪ P (600 items, equated)outside Jev-class caps: cost 6.6× Jev (v1.5 reference); latency 2.6× Jev (v1.5 reference)21.494.5$0.21271.59 s
Instinct Dual 4Bhosted APIComposite #18A4 ∪ P (600 items, equated)18.057.2$0.00680.23 s
Gemini 3.1 Flash-Litehosted APIComposite #19A4 ∪ P (600 items, equated)outside Jev-class caps: cost 7.1× Jev (v1.5 reference)16.959.8$0.22950.96 s
system-one-open (Gemma 4 E2B LoRA on an L4)author-hosted demoComposite #20A4 ∪ P (600 items, equated)16.451.9est.$0.01141.02 s
openjev-sglang (Qwen3.6-35B-A3B on SGLang)author-hosted demoComposite #21A4 ∪ P (600 items, equated)outside Jev-class caps: cost 4.3× Jev (v1.5 reference); latency 2.1× Jev (v1.5 reference)15.458.7est.$0.13981.26 s
DeepSeek V4.1 Flash (thinking default)hosted APIComposite #22A4 ∪ P (600 items, equated)outside Jev-class caps: cost 14.9× Jev (v1.5 reference); latency 3.3× Jev (v1.5 reference)7.197.4$0.48012.04 s
SimpleJev Qwen3.8-27Bauthor-hosted demoComposite #23A4 ∪ P (600 items, equated)outside Jev-class caps: cost 21.3× Jev (v1.5 reference)3.271.6est.$0.68681.21 s
decision-machine-1 (milliseconds.ai)hosted APIComposite #24A4 ∪ P (600 items, equated)2.446.1$0.02450.17 s
JevAct (einptein, jev1-2b-v2)author-hosted demoComposite #25A4 ∪ P (600 items, equated)1.128.8est.$0.01120.80 s
Fastino GLiNER-2.5-Decide (hosted API)hosted APIComposite #26full set (1,500 items)0.334.6$0.06400.44 s
Qwen3.8 27B (Chutes TEE)hosted APIComposite #27A4 ∪ P (600 items, equated)outside Jev-class caps: cost 67.4× Jev (v1.5 reference); latency 11.5× Jev (v1.5 reference)0.098.2est.$2.17847.09 s
Also measured on v1.6.1, listed, not ranked (3)
BB-Qwen3.5-4B-LoRA (Babak Barazandeh, LoRA on Qwen3.5-4B-Base, hosted /v1/systemone)hosted APIconfiguration variant · not rankedfull set (1,500 items)—60.4—0.70 s
wity-1 (Wity, reasoning always)hosted APIconfiguration variant · not rankedfull set (1,500 items)70.679.3$0.02361.69 s
wity-1 (Wity, reasoning off)hosted APIconfiguration variant · not rankedfull set (1,500 items)49.463.0$0.02360.28 s
Wrappers that serve Jev (listed, never ranked, not class-assessed) (1)
classifier.dev (fast tier)hosted APIwrapper (serves Jev) · not rankedA4 ∪ P (600 items, equated)71.678.2est.$0.02350.81 s
31 of 31 systems

Model kind

Jev-class

Release

Parameters

No system in this release reports an exact parameter count.

Developer/API price $/1k
Base-model price $/1k

No system in this release reports a base-model reference price.

Scored cost $/1k
Alternative price $/1k

No system in this release reports an alternative pricing scenario.

p50 latency (s)
p95 latency (s)

Compare two systems

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0REFERENCE · TypeSafe — Jev — reference (TypeSafe, closed) · Score 71.5 (#4)
  • B: Sage 1.3.0 — Closed API (weights not public) · Score 74.0 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Sage 1.3.0. Intelligence: 63.6 vs 65.6; Calibration: 90.6 vs 91.6; Speed: 91.5 vs 93.6; Cost: 54.7 vs 58.2.50100Intelligence: Jev 1.13.0: 63.6; Sage 1.3.0: 65.6Intelligence63.6 · 65.6Calibration: Jev 1.13.0: 90.6; Sage 1.3.0: 91.6Calibration90.6 · 91.6Speed: Jev 1.13.0: 91.5; Sage 1.3.0: 93.6Speed91.5 · 93.6Cost: Jev 1.13.0: 54.7; Sage 1.3.0: 58.2Cost54.7 · 58.2
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Jev 1.13.0 vs Sage 1.3.0. Rules, policy & law: 49.4 vs 57.9; Coding & software: 64.3 vs 77.3; Math & numbers: 10.4 vs 17.6; Finance & commerce: 40.9 vs 52.7; Support & operations: 48.8 vs 71.1; Everyday language: 77.5 vs 78.6; Safety & security: 52.2 vs 73.1. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Rules, policy & law: Jev 1.13.0: 49.4; Sage 1.3.0: 57.9 — Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open, 1465 sealed).Rules, policy& law49.4 · 57.9Coding & software: Jev 1.13.0: 64.3; Sage 1.3.0: 77.3 — Coding & software: code, SQL, repositories, developer tools and IT systems. 382 items (62 open, 320 sealed).Coding &software64.3 · 77.3Math & numbers: Jev 1.13.0: 10.4; Sage 1.3.0: 17.6 — Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open, 333 sealed).Math &numbers10.4 · 17.6Finance & commerce: Jev 1.13.0: 40.9; Sage 1.3.0: 52.7 — Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open, 211 sealed).Finance &commerce40.9 · 52.7Support & operations: Jev 1.13.0: 48.8; Sage 1.3.0: 71.1 — Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open, 139 sealed).Support &operations48.8 · 71.1Everyday language: Jev 1.13.0: 77.5; Sage 1.3.0: 78.6 — Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open, 134 sealed).Everydaylanguage77.5 · 78.6Safety & security: Jev 1.13.0: 52.2; Sage 1.3.0: 73.1 — Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open, 103 sealed).Safety &security52.2 · 73.1

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Measured subject-topic competence on this release’s item set. Chance-corrected competence per category (0 = at or below chance; negative averages are reported as 0, 100 = perfect), Jev 1.13.0: S+P+L1+L2+L3; Sage 1.3.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open / 1465 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 382 items (62 open / 320 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open / 333 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open / 211 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open / 139 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open / 134 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open / 103 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jev 1.13.0 vs Sage 1.3.0. Model routing: 73.0 vs 78.4; Legal & compliance: 68.6 vs 65.3; Customer support: 49.6 vs 61.4; Other: 40.1 vs 44.3; E-commerce: 40.5 vs 59.2; Insurance claims: 46.1 vs 51.4; Risk assessment: 32.8 vs 41.0; Financial crime: 35.8 vs 29.3; Feature extraction: 28.3 vs 40.8; Lead generation: 49.2 vs 63.8; Recruiting: 0.0 vs 19.9; Knowledge graphs: 28.8 vs 72.8; LLM guardrails: 68.0 vs 91.3; Moderation: 51.6 vs 66.4; Code linting: 29.1 vs 62.9; Search & retrieval: 80.9 vs 90.0; Science: 21.2 vs 35.2; Advertising: 16.4 vs 43.8; Gaming: 5.1 vs 16.8; Demand forecasting: 0.0 vs 0.0. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Model routing: Jev 1.13.0: 73.0; Sage 1.3.0: 78.4 — Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open, 438 sealed).Modelrouting73.0 · 78.4Legal & compliance: Jev 1.13.0: 68.6; Sage 1.3.0: 65.3 — Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open, 441 sealed).Legal &compliance68.6 · 65.3Customer support: Jev 1.13.0: 49.6; Sage 1.3.0: 61.4 — Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open, 203 sealed).Customersupport49.6 · 61.4Other: Jev 1.13.0: 40.1; Sage 1.3.0: 44.3 — Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open, 150 sealed).Other40.1 · 44.3E-commerce: Jev 1.13.0: 40.5; Sage 1.3.0: 59.2 — E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open, 110 sealed).E-commerce40.5 · 59.2Insurance claims: Jev 1.13.0: 46.1; Sage 1.3.0: 51.4 — Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open, 110 sealed).Insuranceclaims46.1 · 51.4Risk assessment: Jev 1.13.0: 32.8; Sage 1.3.0: 41.0 — Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open, 110 sealed).Riskassessment32.8 · 41.0Financial crime: Jev 1.13.0: 35.8; Sage 1.3.0: 29.3 — Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open, 92 sealed).Financialcrime35.8 · 29.3Feature extraction: Jev 1.13.0: 28.3; Sage 1.3.0: 40.8 — Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open, 94 sealed).Featureextraction28.3 · 40.8Lead generation: Jev 1.13.0: 49.2; Sage 1.3.0: 63.8 — Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open, 91 sealed).Leadgeneration49.2 · 63.8Recruiting: Jev 1.13.0: 0.0; Sage 1.3.0: 19.9 — Recruiting: resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open, 93 sealed).Recruiting0.0 · 19.9Knowledge graphs: Jev 1.13.0: 28.8; Sage 1.3.0: 72.8 — Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open, 91 sealed).Knowledgegraphs28.8 · 72.8LLM guardrails: Jev 1.13.0: 68.0; Sage 1.3.0: 91.3 — LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open, 88 sealed).LLMguardrails68.0 · 91.3Moderation: Jev 1.13.0: 51.6; Sage 1.3.0: 66.4 — Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open, 86 sealed).Moderation51.6 · 66.4Code linting: Jev 1.13.0: 29.1; Sage 1.3.0: 62.9 — Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open, 87 sealed).Code linting29.1 · 62.9Search & retrieval: Jev 1.13.0: 80.9; Sage 1.3.0: 90.0 — Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open, 85 sealed).Search &retrieval80.9 · 90.0Science: Jev 1.13.0: 21.2; Sage 1.3.0: 35.2 — Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open, 88 sealed).Science21.2 · 35.2Advertising: Jev 1.13.0: 16.4; Sage 1.3.0: 43.8 — Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open, 85 sealed).Advertising16.4 · 43.8Gaming: Jev 1.13.0: 5.1; Sage 1.3.0: 16.8 — Gaming: player reports, in-game chat, game support. 86 items (3 open, 83 sealed).Gaming5.1 · 16.8Demand forecasting: Jev 1.13.0: 0.0; Sage 1.3.0: 0.0 — Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open, 80 sealed).Demandforecasting0.0 · 0.0

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Measured application-category competence on this release’s item set. Chance-corrected competence per category (0 = at or below chance; negative averages are reported as 0, 100 = perfect), Jev 1.13.0: S+P+L1+L2+L3; Sage 1.3.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.A point on the bold 0 ring printed 0.0 is a measured value, not a gap: these published cells clip below-chance results to 0, so 0 means at or below chance. Their score markers cannot fall inside that ring.
What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open / 438 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open / 441 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open / 203 sealed)
  • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open / 150 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open / 110 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open / 110 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open / 110 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open / 92 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open / 94 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open / 91 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open / 93 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open / 91 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open / 88 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open / 86 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open / 87 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open / 85 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open / 88 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open / 85 sealed)
  • Gaming — player reports, in-game chat, game support. 86 items (3 open / 83 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open / 80 sealed)

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs Sage 1.3.0. Choice · open: 79.0 vs 85.1; Choice · sealed: 75.8 vs 79.7; Noul · open: 54.5 vs 54.3; Noul · sealed: 46.1 vs 56.4; Score · open: 63.3 vs 57.7; Score · sealed: 63.0 vs 60.5. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Choice · open: Jev 1.13.0: 79.0; Sage 1.3.0: 85.1Choice · open79.0 · 85.1Choice · sealed: Jev 1.13.0: 75.8; Sage 1.3.0: 79.7Choice ·sealed75.8 · 79.7Noul · open: Jev 1.13.0: 54.5; Sage 1.3.0: 54.3Noul · open54.5 · 54.3Noul · sealed: Jev 1.13.0: 46.1; Sage 1.3.0: 56.4Noul · sealed46.1 · 56.4Score · open: Jev 1.13.0: 63.3; Sage 1.3.0: 57.7Score · open63.3 · 57.7Score · sealed: Jev 1.13.0: 63.0; Sage 1.3.0: 60.5Score ·sealed63.0 · 60.5

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Chance-corrected competence for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs Sage 1.3.0. Easy: 93.9 vs 89.6; Standard: 63.2 vs 62.3; Judge: 67.7 vs 67.5; Hard: 65.6 vs 70.5. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Jev 1.13.0: 93.9; Sage 1.3.0: 89.6Easy93.9 · 89.6Standard: Jev 1.13.0: 63.2; Sage 1.3.0: 62.3Standard63.2 · 62.3Judge: Jev 1.13.0: 67.7; Sage 1.3.0: 67.5Judge67.7 · 67.5Hard: Jev 1.13.0: 65.6; Sage 1.3.0: 70.5Hard65.6 · 70.5

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs Sage 1.3.0. Easy: 78.4 vs 77.5; Standard: 68.7 vs 75.8; Judge: 66.9 vs 73.2; Hard: 58.9 vs 60.8. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Jev 1.13.0: 78.4; Sage 1.3.0: 77.5Easy78.4 · 77.5Standard: Jev 1.13.0: 68.7; Sage 1.3.0: 75.8Standard68.7 · 75.8Judge: Jev 1.13.0: 66.9; Sage 1.3.0: 73.2Judge66.9 · 73.2Hard: Jev 1.13.0: 58.9; Sage 1.3.0: 60.8Hard58.9 · 60.8

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

O1S chance-corrected competence: equal mean across the request types present, clipped to 0–100. Zero means the equal-weight mean across present request types was at or below baseline; negative means are clipped. It does not mean every type scored zero. Choice uses a random-option baseline; Noul uses a fixed 50% accuracy baseline (abstentions count as wrong); Score uses the midpoint-guess error (random-guess fallback when every gold is at the midpoint). Compare models within a category; row pools and supported types can differ.

Categories use ruled subject-topic labels and authoring use-case precedence across the answered pool union, with stable item identities counted once.

  • Official O1S; recorded failures count where supported; no missing response is imputed; stable identity dedupe parent first.
  • S-based category cells pool recorded supported responses from S+P and the answered L1/L2/L3 supplements. API overlay category cells keep their separately recorded A4/A5+P union. Headline and per-type/tier scores are unchanged.
All values as a table
SpokeA: Jev 1.13.0B: Sage 1.3.0
The four score axes
Intelligence63.665.6
Calibration90.691.6
Speed91.593.6
Cost54.758.2
Capability by subject topic
Rules, policy & law49.457.9
Coding & software64.377.3
Math & numbers10.417.6
Finance & commerce40.952.7
Support & operations48.871.1
Everyday language77.578.6
Safety & security52.273.1
Use cases (TypeSafe categories)
Model routing73.078.4
Legal & compliance68.665.3
Customer support49.661.4
Other40.144.3
E-commerce40.559.2
Insurance claims46.151.4
Risk assessment32.841.0
Financial crime35.829.3
Feature extraction28.340.8
Lead generation49.263.8
Recruiting0.019.9
Knowledge graphs28.872.8
LLM guardrails68.091.3
Moderation51.666.4
Code linting29.162.9
Search & retrieval80.990.0
Science21.235.2
Advertising16.443.8
Gaming5.116.8
Demand forecasting0.00.0
Competence per request type, open / sealed
Choice · open79.085.1
Choice · sealed75.879.7
Noul · open54.554.3
Noul · sealed46.156.4
Score · open63.357.7
Score · sealed63.060.5
Competence per tier — open set
Easy93.989.6
Standard63.262.3
Judge67.767.5
Hard65.670.5
Competence per tier — sealed set
Easy78.477.5
Standard68.775.8
Judge66.973.2
Hard58.960.8

Category radars count each answered item once from the pools named under each radar. API overlay rows use A4+P+L1+L2+L3; A5+P+L1+L2+L3; S+P+L1+L2+L3. Raw and unequated; cells under 15 answered items are omitted. Per-type and tier radars retain each row's original measurement pools: A4/A5 rows have 300 open plus 300 sealed items; full-set rows have S 1,200 plus P 300.

Languages

Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts. Cells under 15 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. The mixed-language group is listed first, then English (1,306 items); the other 21 languages and the mixed-language group share 1699 items.

JevBench v1.6.1 competence by item language and system
Systemmixed78en1306es95pt90pl88de84da83it83fr80ja80ko77sv75tr75cs75fi73uk73nl72hi72ar71id71zh70no69el65
Sage 1.3.0API · S+P+L1+L2+L38071485542494445575136444332264451656250593743
d1API · S+P+L1+L2+L38066325534443336644028304128143327485141471232
Mercury DecideAPI · S+P+L1+L2+L3Cells use S+P, L1, L2 and L3. L3 has 1,416 recorded observations on its 1,418 items: 1,412 answers, three input refusals and one rate-limit (429) error, an operational non-answer rather than an answer. The two missing records (a second spent 429 kept only as a transport failure, and one request stopped before sending) are not imputed; the normal 98% recorded-completeness gate applies without exception. Besides 868 original free-route records, L3 uses 447 retained and 101 newly issued paid decisions on the currently declared Inception Mercury Decide 2026-09-30 route. That route qualified on public items only: top labels agreed on 287 of 297 (96.63%), but probabilities differ (one bit-exact vector, maximum absolute difference 0.977543). Equality with the historically measured weights and calibrator, and with the original free responses, is unproven. Headline, rank and original row price are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure.7269495039654644675243424434425843605640463844
Jev 1.13.0API · S+P+L1+L2+L3766838442248353436363015431173420364323471921
wity-1 (Wity, reasoning auto)API · S+P+L1+L2+L38376596631603647384249414757372755595949494032
Microsoft-Decision-1API · S1200+P300+L1+L2+L3; full native Foundry union8857465937514945505533424339243846344457543333
SPX-CD FlashAPI · S+P+L1+L2+L3786338433640273449483122301783824375146501229
OpenAI DecisionsAPI · A5+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3.7659364841552737433724234637303327403339374025
InstinctAPI · A4+P+L1+L2+L3675822353228273152482018341923223939394243137
SPX-CD ProAPI · S+P+L1+L2+L38270445537445343625035474119224737485429522947
Vansa-3.4API · A4+P+L1+L2+L35447124228171210244311013191291726153335011
GPT-6 Luna (low reasoning effort)API · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3.9797919391941009410099989696939694979710098989997
Autoloops – Gemma 4 31B ITAPI · A4+P+L1+L2+L37367445044566142616043484848476050554945534151
GPT-6 Luna (default medium reasoning effort)API · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3.9797989691981009999979810096959896979995971009996
Fastino GLiDEAPI · A4+P+L1+L2+L36973405329424337715243334322333756534750473540
SimpleJev Qwen3.6-35B-A3BAPI · A4+P+L1+L2+L36563374226433041535424344417222134302034422520
GPT-5.6 LunaAPI · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3.9693869284869290879490799397948987848591898291
Instinct Dual 4BAPI · A4+P+L1+L2+L3363212500027132000001161201500
Gemini 3.1 Flash-LiteAPI · A4+P+L1+L2+L38468426457535846635842455551274340493554544944
system-one-openAPI · A4+P+L1+L2+L3233181491113167166000001910022502
openjev-sglangAPI · A4+P+L1+L2+L36047202919280192442191212139133224143728200
DeepSeek V4.1 FlashAPI · A4+P+L1+L2+L3979594999298991001009998961009810096981001009910098100
SimpleJev Qwen3.8-27BAPI · A4+P+L1+L2+L37063284038433322544323253417312144433731552927
decision-machine-1API · A4+P+L1+L2+L3104000001000050000004008
JevActAPI · A4+P+L1+L2+L312001000000000000001001400
Fastino GLiNER-2.5-DecideAPI · S+P+L1+L2+L3+L4110000000008000010000000
Qwen3.8 27BAPI · A4+P+L1+L2+L3100959592939699971009898991001001009696100999310097100
BB-Qwen3.5-4B-LoRAAPI · S+P+L1+L2+L35341271920151516201319017010181510222410
wity-1 (Wity, reasoning always)API · S+P+L1+L2+L37576576832534151524434462745433750526458463823
wity-1 (Wity, reasoning off)API · S+P+L1+L2+L367523848293621261013171417176112430313333261
Wrappers (listed, never ranked)
classifier.devAPI · A4+P+L1+L2+L3737223353643342833412315359183225444417511434

Intelligence gate and Noul decisiveness

The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.

SystemIntelligenceChoiceNoulScoreNoul decisive rateAccuracy among decisive
Sage 1.3.065.6nativenativenative83 %95 %
d162.1nativenativenative83 %93 %
Mercury Decide65.9nativenativenative90 %90 %
Jev 1.13.063.6nativenativenative79 %98 %
wity-170.3nativenativenative91 %94 %
Microsoft-Decision-157.2nativenativenative77 %93 %
SPX-CD Flash58.4nativenativenative81 %94 %
OpenAI Decisions56.9—————
Instinct47.8below gate—————
SPX-CD Pro64.9nativenativenative84 %95 %
Vansa-3.440.9below gate—————
GPT-6 Luna96.6—————
Autoloops – Gemma 4 31B IT58.1—————
GPT-6 Luna96.9—————
Fastino GLiDE66.1—————
SimpleJev Qwen3.6-35B-A3B56.3—————
GPT-5.6 Luna94.1—————
Instinct Dual 4B28.3below gate—————
Gemini 3.1 Flash-Lite58.6—————
system-one-open27.9below gate—————
openjev-sglang38.8below gate—————
DeepSeek V4.1 Flash94.7—————
SimpleJev Qwen3.8-27B52.9—————
decision-machine-113.1below gate—————
JevAct10.1below gate—————
Fastino GLiNER-2.5-Decide7.0below gateconfidenceconfidenceconfidence100 %54 %
Qwen3.8 27B96.4—————
All data (31 systems)

Not yet measured on v1.6 · 0 systems with a dated carried score

Every ranked system of the live v1.5.7 board that is not measured on the v1.6.0 pool keeps its last published score, marked with the release that first published that measurement and that release's publication day. Carried rows are listed separately and never ranked together with v1.6-measured rows. v1.5 protocol (1,624 decisions: 904 open + 720 sealed); scores are on the v1.5 scale and are not comparable with v1.6-measured rows.

Dates: . The date is the publication day of the release that first published the measurement, not a per-model measurement timestamp.

Carried JevBench v1.5.x results, not ranked with v1.6 measurements
SystemMeasured onCapability (v1.5 scale)v1.5 CompositeCost / 1,000Median latency

How it works

One merged board: how newer draws join it

Each system appears once, with its latest valid measurement; earlier systems are not re-measured. The reference scale is v1.6.1 (draw v1.6.0). Every draw shares the same 300 public items (P, manifest 6b2321f8...). A fresh draw is equated to the reference draw with fixed anchor systems measured on both draws: offset = median over anchors of (metric on the reference S u P) - (metric on the new draw S u P), for Intelligence and Calibration, with a bootstrap 95% interval over anchors and items. A row from that draw is shown as published plus the offset; Capability and Composite follow from the equated axes. Until the anchor runs of a draw are complete its offset is 0 and the draw is marked 'anchor runs pending'.

DrawDrawnReleasesAnchor offset (Intelligence / Calibration)Status
v1.6.0 (reference)2026-10-01v1.6.0, v1.6.10 / 0reference scale · G_med 2.57
Fast-lane draw · v1.6-fastlane-202610092026-10-09v1.6.2, v1.6.3, v1.6.4+0.00 / +0.00anchor runs pending: shown as published
Regular draw · v1.6-regular-20261010-a22026-10-10v1.6.5, v1.6.6, v1.6.7, v1.6.8+0.00 / +0.00anchor runs pending: shown as published
  • Anchor pool: hopper, kev-0.6b, kev-4b, kev-8b, localjev-qwen3.5-4b, malkuth-4b, metask-jev-4b, raw-phi-4-mini, raw-qwen3-1.7b, raw-qwen3-4b-instruct-2507, raw-qwen3-8b, typecastlm. The same 12 anchors answered two sealed draws of v1.6.0 (S and the A2 supplement): median difference -0.29 Intelligence and +1.56 Calibration points (v1.6.1 results, v16.equating_A2).
  • On the shared public items, systems on both new draws score 6 to 13 sealed points below the reference field's sealed-vs-public line (reference residual SD 3.8). Either the new sealed draws are harder or their cohorts are tuned on the public items; only anchor runs can separate the two. Until then their rows are, if anything, understated.
  • Rows from a newer draw keep the gap-penalty reference (G_med) their release was scored with; the row tooltip and tag name the draw and date. Their topic, use-case and language cells use their own draw’s item set and labels, so they are shown on their release page (linked from the draw tag in Release history), not in the radars and language table here.

Open weights and API offerings (board v1.7.43)

Why open weights and APIs are compared in separate groups
  • Fair cost and speed. We run every open-weights model on hardware we rent and operate, so cost and latency compare on the same terms. The GPU cost calculator prices your own setup: own hardware, on-demand or long-term rental.
  • API prices can change. An API price is the vendor's decision. It can be subsidised (for example on top-end GPUs) and raised later, and readers cannot reproduce it.
  • Different fairness needs. A hosted endpoint chooses its own hardware and sees the benchmark inputs, so API offerings are compared with each other.
  • Open-source focus. JevBench exists to make open decision models comparable and reproducible. Every score and the method are identical in every view; the All view ranks both groups together by Capability.
  • The main board at /jev-models ranks open-weights systems: published weights that we ran ourselves, on GPU or CPU machines we rent and operate. A system we measured through an endpoint we do not run (vendor API, author-hosted or third-party-hosted endpoint, for example Qwen3.8 27B via Chutes) is an API offering, even when its base weights are open; API offerings are ranked on this board.
  • Why separate boards: open weights can be compared on equal hosting terms, while an API price is a vendor decision that can be subsidised or raised later and is not reproducible by readers. Every score and measurement is the same on both boards; only the set of ranked rows differs, and ranks are the published order filtered to that set.
  • Jev 1.13.0 is a hosted API. It stays on the open-weights board as the reference row (it defines the Jev-class cost and latency caps) and is not ranked there; it is ranked on the API leaderboard.
  • Official cost basis is unchanged: the Cost axis keeps each row's documented reference price (see the cost notes below; APIs with a known base model are priced at the developer's own list price). The base-model reference price only applies to open-weights rows; API offerings are ranked at their own list price on this board. The GPU cost calculator on the main board is a What-If for your own hosting and never changes a score or rank.
  • Full API re-run (A4, v1.7.7). Most offerings on this board were measured on 6 Oct 2026 on a fresh sealed API set A4 (300 never-used sealed items) plus the same 300 public items every system answers, and equated to the v1.6.1 S ∪ P scale with the published A2/A3 supplement method (offset +2.57 Intelligence, +3.35 Calibration; pool of 8 ranked self-hosted systems re-run on A4 ∪ P). Jev, Sage, wity-1 and Fastino GLiNER-2.5-Decide keep their full-set S ∪ P scores; Liquid AI d1 answered the full S ∪ P set on 6 Oct 2026 (v1.7.8) and is scored like them, without equating. The pool is mid-strength, so for the strongest LLM rows the offset is an extrapolation (likely within ±2 Intelligence points; the 95% intervals include it). Nine A4 ∪ P items of about 77,000–82,000 input tokens, beyond the Jev reference's accepted input range, count against Intelligence when refused but are left out of cost. Rows without a public tariff keep their documented v1.5 cost estimate. A4 is now retired for everyone. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
  • Language and use-case cells (v1.7.12). These breakdowns add two sealed supplements (354 language items, 33 use-case items, drawn and reviewed on 6 Oct 2026), so every language and use case has at least 30 items. Headline scores are unchanged; rows not yet run on the supplements are tagged S+P in the language table.
  • Every API offering we reached is measured on v1.6.1: either on the full 1,500-item set or on A4 ∪ P (600 items, equated). No preliminary public-set rows remain.

Revision history

  • v1.7.43 (2026-10-10): Display only: a Pareto frontier section after the two scatter charts on both boards — Capability against cost per 1,000 decisions and against median latency (log axes; list price and measured endpoint latency on the API board). The red line joins the systems no other system beats on both axes. No score or rank changed.
  • v1.7.42 (2026-10-10): One merged board. The version tabs are gone: /jev-models is always the current board with every system’s latest measurement, and releases are listed under Release history (old version pages keep their URLs). The 16 systems measured on the two fresh draws after v1.6.1 (fast-lane draw v1.6.2–v1.6.4, regular draw v1.6.5–v1.6.7) join the boards with their published scores and a draw tag; each draw is put on the v1.6.1 scale by an anchor offset that stays 0 until the anchor systems have run on it (see How it works). A filter on top switches between Open weights, API and All (both groups, ranked by Capability). Header text is shortened; the details moved into How it works. Jeff 1.0 Large (fast-lane draw) enters the open-weights Capability top 5 at #2.
  • v1.7.41 (2026-10-10): Needle 3 now has language and category cells on the current pools. It keeps its published v1.5 headline. With the benchmark owner's approval its closed engine was sent all supplement items (L1, L2, L3) once more on an isolated machine, plus 128 never-asked S+P items chosen by their use-case label to complete the last radar spoke (these also count in the other cells, so its S+P part is not a random sample); together with an earlier partial run both radars reach at least 30 answered items per spoke (52 per subject topic, 43 per use case). The engine abstains on about 30 % of items and returns invalid UTF-8 on about 7 %; both count as unanswered, so 22 languages stay below the 60-item target (19 to 59 answered items). Its answers to the same public items differ between machines on about 10 %. Its chance-corrected competence is below zero in every cell, shown as 0. The v1.7.40 note for Needle 3 (options as tools) is corrected the same way. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.40 (2026-10-10): Needle 3, options as tools, now has language and category cells on the current pools. It keeps its published v1.5 headline. With the benchmark owner's approval its closed engine ran once more on an isolated machine; the run covered most of L3 and half of L1 before the machine's time limit and is combined with an earlier partial run. Both radars are complete (at least 58 per subject topic and 43 per use case). The engine abstains on about 9 % of items and returns invalid UTF-8 on about 7 % (mostly non-Latin scripts); both count as unanswered, so 18 languages stay below the 60-item target (24 to 59 answered items; columns under 30 carry the low-n mark). Its chance-corrected competence is below zero in every cell, shown as 0. ClassOne Gemma 4 E2B keeps its exception with a corrected reason: its weights were identical, but its server gives different answers after each restart. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.39 (2026-10-10): GLiNER2 large now has language and category cells on the current pools: at least 62 completed responses in every language, 98 in every subject topic and 80 in every use case. It keeps its published v1.5 headline; the cells come from a new run with the original pinned model on an isolated machine. GLiNER2 gained 180 more answered items that had been reserved earlier but never sent: its languages now rest on 50 to 59 answered items where they are below the 60-item target (17 of 23), and its radars on at least 69 per topic and 56 per use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.38 (2026-10-10): Three more historical rows now have language and category cells on the current pools: reflex-27b, Surogate Rune 26B-A4B v3 (RTX PRO 6000) and jeff. They keep their published v1.5 headlines. reflex-27b completed its partly answered S+P set and the supplements; Surogate Rune adds the supplements to its 6 October S+P run; jeff (CPU) answered all pools. Their repeated answers to the 300 public items agree 100 %. Each row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.37 (2026-10-10): Three more historical rows now have language and category cells on the current pools: Von, GLiNER2 and GLiNER2.5 multi. They keep their published v1.5 headlines; the cells come from new runs with the original pinned model and setup on isolated machines. Von and GLiNER2.5 multi reach at least 61 completed responses in every language, 100 in every subject topic and 80 in every use case. GLiNER2 fills both radars (at least 57 per topic, 45 per use case) and every language column, but about 1,000 of its issued items were lost to out-of-memory kills and a machine shutdown; they stay counted as spent, so 22 languages rest on 42 to 59 answered items, below the 60-item target. Three rows that cannot be measured now give the concrete reason: Aplomb 1 (pinned revision deleted upstream), ClassOne Gemma 4 E2B (re-run did not reproduce its original answers) and swanOne (weights return HTTP 401). L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.36 (2026-10-10): Nine more rows now cover S+P+L1+L2+L3 in the language table and both radars: Diffusion Jev, SPX-CD-Omni, Seb-9B, APUS-OpenJev-v1-9B, APUS-OpenJev-v1-35B-A3B, Jobe Qwen3.5-4B, AutoJev-27B (RTX PRO 6000), NInfer Qwen3.8-27B NVFP4 (T=1.5) and OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). The last four are historical rows: they keep their published v1.5 headlines, and their language and category cells come from a new run on the current pools with the original pinned model and serving setup. Diffusion Jev, SPX-CD-Omni, Seb-9B and both APUS rows match their original public answers on at least 97 %. The razorback16 OpenJev model samples stochastically, so its repeated passes agree on 86 to 89 % of the public items; its note says so. Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.35 (2026-10-10): Seven more rows now cover S+P+L1+L2+L3 in the language table and both radars: René-1 31B FP8, decisio v0.8.0 on gemma-4-12B-it, Hopper 12B trained, decider-12b v2, decider-12b v1, Bobcat Flash 1.2 and Mica v0.1 4B. Each answered all 1,505 supplemental items with its original pinned model and serving setup on an isolated machine; answers to the 300 public items match the original runs (at least 99 %). Mica keeps its published v1.5 headline; its language and category cells come from a new run on the current pools. Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The language table now lists the mixed-language group first. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.34 (2026-10-10): Seven more self-hosted rows now cover S+P+L1+L2+L3 in the language table and both radars: TypeCastLM 1.4.0, CoCo-Decision-4B, ClassOne Qwen 3.5 9B, EXAONE-4.0-1.2B-JEV v0.3, jul fast, WaterSheep and Tacet Sonata. Each answered all 1,505 supplemental items with its original pinned model and serving setup on an isolated machine; their answers to the 300 public items match the original runs (at least 97 %). Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. Long-input refusals stay scored as before. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
  • v1.7.33 (2026-10-10): Messier One v0.2 now covers S+P+L1+L2+L3, with all 1,505 supplemental items answered natively on the original image with the network disabled: at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The original P300 answers recur in each supplement and count once; the original S+P measurement had no HTTP errors. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and original row pricing are unchanged.
  • v1.7.31 (2026-10-10): Microsoft-Decision-1 joins the hosted API board after a native Azure Foundry measurement on 1,200 sealed plus 300 public observations: Capability 70.78, Calibration 84.40 and Composite A 69.11. Its 21 context-limit failures remain in the scored denominator. Native median latency is 459 ms on the 124-request speed subset. This new full-set O1S run uses its completed-field median gap reference of 7.088 score points; historical rows retain their own dates, pools and references rather than being remeasured or converted. The existing top five on Composite A and capped Capability are unchanged. All 50 topic/use-case/language cells are recorded; thin cells stay suppressed or table-only, and qualified supplemental coverage remains pending.
  • v1.7.30 (2026-10-10): decisio v0.8.0 on gemma-4-31B-it now covers S+P+L1+L2+L3, with all 1,505 supplemental items answered natively: at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The original P300 answers recur in each supplement and count once. The original 23 HTTP 500 errors remain operational non-answers and are not imputed. L3 items were reviewed by an OpenAI model, so its cells based on L3 carry that exposure; this row is the original locally served Gemma offering. Headline scores, Composite, ranks and original row pricing are unchanged.
  • v1.7.29 (2026-10-09): Mercury Decide now has at least 64 completed observations in every language, 109 in every subject topic and 83 in every use case, using S+P+L1+L2+L3. L3 retains 1,416 records on its full 1,418-item denominator, including original refusals and operational failures; two spent failures remain missing. Paid supplemental decisions use the currently declared Inception 2026-09-30 offering and the original native protocol. Public top labels agree on 96.63%, with differing probabilities and unproven historical weights/calibrator equality; the row coverage note gives the comparison. Headline scores, Composite, ranks and original row pricing are unchanged.
  • v1.7.28 (2026-10-09): GLiNER 2.5 Small now covers P300 + L1 + L2 + L3 with 1,802 completed native responses: at least 60 in each of 23 languages, 78 in every subject topic and 52 in every use case. Its original offline CPU/fp32 profile is unchanged. Three earlier public-item OOM outcomes remain missing and contribute no completed coverage; no requests were replayed. These cells exclude historical S answers and stay outside headline scores. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.27 (2026-10-09): OpenJev DeBERTa v3 Large adds 60 new L5 items: 20 each in Arabic, Hindi and Greek. It now has at least 63 completed responses in every language (Arabic 70, Hindi 73, Greek 69). The original native context was checked before keeper selection; all 60 requests succeeded, and earlier failures remain scored. Other rows and category views retain their existing measurements. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.26 (2026-10-08): BB-Qwen3.5-4B-LoRA now covers S+P+L1+L2+L3, with at least 65 completed responses in each of 23 languages and at least 83 on every capability and use-case spoke. Two DNS failures remain scored and do not count toward completed coverage. Its reference cost stays unknown; these view values do not impute a price. OpenJev DeBERTa v3 Large now includes its completed L1/L2 runs; Arabic (50), Hindi (53) and Greek (49) still need more completed responses. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.25 (2026-10-08): Watt Flash 0.1 and OpenJev Verdict completed their language supplements. Both now cover S+P+L1+L2+L3, with at least 65 completed responses in each of 23 languages and at least 83 in each capability and use-case category. Other rows retain their existing measurements, including Fastino’s Hindi L4 supplement. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.24 (2026-10-08): Fastino GLiNER 2.5 Decide now has 69 completed Hindi responses after a new 20-item Hindi supplement. Seventeen requests succeeded; three HTTP 400 failures remain scored and do not count toward completed coverage. L4 is used only for this offering’s Hindi language cell, with OpenAI author and independent Anthropic reviewer provenance disclosed. The mghafiri Qwen3.5 0.8B row also completed L3, expanding its language and category cells. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.23 (2026-10-08): The Languages table includes every listed model, including historical and catalogue entries. These rows use their stored language measurements where available and show pending cells where measurements are unfinished. Historical headline scores are not used as language values. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.22 (2026-10-08): Languages and radars now include completed L1/L2/L3 runs for SPX CD Flash, SPX CD Pro and Decisio Gemma 4 12B v0.9. Coverage counts completed supported responses and recorded input refusals; authentication, rate-limit, service and transport errors do not satisfy reporting thresholds. Competence keeps its original scored observations. 3 listed rows still have incomplete radar coverage and show the reason. Every listed row remains in the completion check. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.21 (2026-10-08): Languages and radars: more rows completed their L1/L2/L3 supplement runs, among them Instinct (its API is back), Surogate Rune 26B-A4B v3, Xor 26B-A4B, JADE, Jebadiah 27B, Kev 27B, Decision 2.0 Vega 27B and the self-hosted SimpleJev Qwen3.8-27B. 3 rows still show a reason instead of a full radar: newly listed rows whose supplement runs are not done yet, rows not measured on this item pool, and one model whose pinned weights could not be accessed. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.20 (2026-10-08): Correction and progress for the language and radar views. The self-hosted row SimpleJev Qwen3.8-27B wrongly counted the hosted SimpleJev API row's L1/L2/L3 supplement answers in its language and category cells (v1.7.18 and v1.7.19), because both share a file name. Its cells now use only its own S + P answers until its own supplement runs finish, and the cell builder no longer lets a hosted-API run count for a self-hosted row. More open-weights rows completed their L1/L2/L3 runs. 3 rows still show a reason instead of a full radar. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.19 (2026-10-08): Languages and radars: the open-weights rows were run on the L3 supplement overnight, and most rows that still lacked the earlier L1/L2 supplements got them too. Their language cells and topic/use-case radars now count every item they answered (S + P + L1 + L2 + L3), and Qwen3.8 27B (Chutes) finished its L3 run. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows carry that exposure; a note next to those rows says so, and headline scores do not use L3. 3 rows still show a reason instead of a full radar, for example runs still in progress, a provider that is down, withdrawn weights or rows not measured on this pool. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.18 (2026-10-08): Languages and radars for every row. A new sealed supplement, L3 (1,118 items drawn on 7 Oct 2026: items written natively in 21 languages plus a mixed-language group, and English items for thin radar categories such as everyday language, safety and the "other" use case), gives every language at least 60 items on P ∪ L1 ∪ L2 ∪ L3 (before: as few as 4). The hosted API rows answered it: their language cells and their use-case and topic radars now rest on about 2,100 answered items (3,000 for rows that answered the full sealed set), with at least 69 on every spoke. Open-weights rows are being run on it in the current GPU wave and update as they finish. 3 rows that do not reach 30 items on every spoke yet show the reason next to the row. Headline scores, Capability, Composite and every rank are unchanged.
  • v1.7.17 (2026-10-07): Languages now includes every measured row on both boards, with wrappers listed below the models. Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts. 19 API rows re-run on A4/A5 have answered L1/L2 supplements; their category radars count those items too. L3 coverage will follow when measured.
  • v1.7.16 (2026-10-07): Compare view, API board: subject-topic and use-case radars for OpenAI Decisions, the 17 API offerings re-run on A4 ∪ P and the classifier.dev wrapper. Their 600 new sealed items were labelled with the same recipe as every other item (Winnow-12B Q8 on our own GPU pod; uc1 items keep their authoring use case); values are raw and rest on 600 items (300 sealed), so more categories fall under the 30-item spoke minimum and are listed below the radar. Each of these rows also gets a breakdown section on its own page. Every live row now has all breakdown views, and a release check keeps it that way. No score, axis, cost or rank changed.
  • v1.7.15 (2026-10-07): Compare view, API board: OpenAI Decisions, the 17 API offerings re-run on A4 ∪ P and the classifier.dev wrapper now show competence per request type and per tier on the open and sealed items they answered (300 open + 300 sealed: fewer sealed items than the full set, so wider uncertainty; the view note says so), and Liquid AI d1 shows its subject-topic and use-case radars. Values are raw, computed from the stored per-item results with the same scorer as every other row. Subject-topic and use-case radars for the 600-item rows follow once their 600 new sealed items are labelled. No score, axis, cost or rank changed.
  • v1.7.14 (2026-10-07): Open-weights board: H2O-Lightning-4B v1.1 (H2O.ai) and Kahn1 4B (Okura66), two Qwen3.5-4B fine-tunes submitted as JevBench add-requests, measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a8 of v1.6.1). Both are priced with labelled cost estimates like every other Qwen3.5-4B row; Kahn1's server reports no token usage, so its estimate is our own count of its input tokens (three passes per decision, its server default). Both rank inside the Jev-class caps: H2O-Lightning-4B at #3 of the Capability Score and #1 of the Composite, Kahn1 at #23 of the Capability Score. No published score, axis or cost changed.
  • v1.7.13 (2026-10-07): Open-weights board: Surogate Rune 26B-A4B v3 (measured on the v1.6 pool for the first time; its dated v1.5 row stays in the carried list) and Xor 26B-A4B (Juspay), both from the top 20 of the community Jev Decision Index on Hugging Face, measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a7 of v1.6.1). Both rank inside the Jev-class caps, at #3 and #4 of the Capability Score. No published score, axis or cost changed.
  • v1.7.12 (2026-10-07): Languages and use cases: the per-language and per-use-case cells now also count two sealed supplements, a language supplement L1 (354 items) and a use-case supplement L2 (33 items), drawn and reviewed on 6 Oct 2026. Every one of the 22 languages (plus the mixed-language group) and every one of the 20 use cases now has at least 30 items (before: as few as 5). 82 rows answered both supplements; rows not run on them yet keep their S + P cells and are tagged S+P (or S+P+L1). Headline scores, Capability, Composite and every rank are unchanged on both boards.
  • v1.7.11 (2026-10-07): Open-weights board: 12 new self-hosted Jev-compatible rows from the top 20 of the community Jev Decision Index on Hugging Face (Perplexity Decider v1.1 27B, torchcast-decision-27b, Kev 27B, decider chat on Gemma-4-31B-it, GEV-26B-Decide, Decision 2.0 Vega 27B, JEV-27B, Jebadiah 27B, JADE, Bespoke Nimble 9B v3, SimpleJev Qwen3.8-27B self-hosted, JPT-35B-A3B), measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a6 of v1.6.1). No published score, axis or cost changed; ranks move only where a new row places above. Rows priced above the Jev-class cost or latency caps are listed under the limits section, not in the Capability ranking.
  • v1.7.10 (2026-10-06): API board: OpenAI Decisions (POST /v1/decisions with gpt-6-luna) replaces its preliminary public-set row with its official result. It answered a fresh sealed API set A5 (300 never-used sealed items, since A4 had already been sent to OpenAI) plus the 300 public items (598 of 600 answered) and is equated to the v1.6.1 scale with the same method and pool as the A4 rows (offsets +5.78 Intelligence, +6.15 Calibration). It is ranked on the API board: Composite 62.5 (95% interval 51.6–64.6), Capability 73.5; list price USD 0.10 per 1M input tokens as of 6 Oct 2026 (launch day), re-checked at each revision. It enters the API Composite top 5 at #5, statistically tied with Instinct (62.2, now #6), and the Jev-class Capability top 5 at #4 (Instinct moves to #5, Vansa-3.4 to #6). Every other score is unchanged; the open-weights board is unchanged.
  • v1.7.9 (2026-10-06): API board: the OpenAI Decisions API (POST /v1/decisions with gpt-6-luna, opened on 6 Oct 2026) is added as a preliminary, unranked row from the 300 public v1.6 items (298 answered; Capability 68.8, Composite 37.2 on that set; list price USD 0.10 per 1M input tokens). Its official sealed run waits for the next fresh sealed API draw. No score or rank changed on either board.
  • v1.7.8 (2026-10-06): API board: Liquid AI d1 is added as a ranked API offering. It answered the full v1.6.1 set (1,200 sealed + 300 public items, all 1,500 answered) on 6 Oct 2026 and is scored exactly like the other full-set API rows, without equating (Composite 73.0, Capability 74.4; tariff USD 0.04 per 1M input tokens). It enters the API board top 5 in Composite and in Jev-class Capability. Every other score is unchanged; the open-weights board is unchanged.
  • v1.7.7 (2026-10-06): API board: 17 API offerings re-measured in full on a fresh sealed API set (A4, 300 never-used sealed items plus the 300 public items, 6 Oct 2026) and equated to the v1.6.1 scale with the published A2/A3 supplement method. They replace the hatched preliminary rows and are ranked on the API board: Instinct and Vansa-3.4 enter its Composite top 5, and Instinct, Vansa-3.4 and Instinct Dual 4B join Sage and Jev in its Jev-class Capability top 5; GPT-6 Luna, Qwen3.8 27B, GPT-6 Luna low, DeepSeek Flash and GPT-5.6 Luna have the highest raw Capability (98.4, 98.2, 97.8, 97.4 and 94.5) but sit outside the Jev-class caps. Qwen3.8 27B (Chutes) is ranked as well (Composite 0.0: no published tariff, so its documented USD 2.18 estimate per 1,000 decisions gives Cost 0). Instinct and Instinct Dual 4B are now treated as production APIs (public tariff), so no demo-endpoint latency adjustment applies to them. classifier.dev (fast tier) is listed as a wrapper, not ranked. No preliminary rows remain. The open-weights board is unchanged.
  • v1.7.6 (2026-10-06): API board: Vansa-3.4 and Instinct Dual 4B now have v1.6 public-set figures and appear as preliminary rows; Autoloops stays pending while its full run on the fresh sealed set is in progress. Very long items that exceed a provider’s context window count as wrong without ending the run. No rank changed.
  • v1.7.5 (2026-10-06): API board: every API offering with a v1.6 public-set figure is drawn in the Composite and Capability charts as a hatched, unranked “preliminary” row at its score position (public set of 300 items; full sealed re-evaluation running); offerings without any v1.6 figure are greyed “pending” rows. Adds Qwen3.8 27B (Chutes) and Fastino GLiDE to the public-set table. The four ranked rows and every rank are unchanged; the open-weights board is unchanged.
  • v1.7.4 (2026-10-06): API board: every reachable API offering that was not yet re-measured on v1.6 now has a dated public-set figure (the 300 public v1.6 items, no sealed items), shown next to its older v1.5 score with Jev and other ranked APIs on the same 300 items as anchors. Not ranked and not comparable with the 1,500-item headline; no score or rank of either board changed.
  • v1.7.3 (2026-10-06): Display only, no score or rank changed. The API board ranks every offering at its own list price and no longer shows a base-model reference price (that comparison only matters against open weights). The Jev reference row and, with “Show API offerings” on, every API offering now sit at their score position in each ranking section (Composite and Capability) instead of the folded end of the list.
  • v1.7.2 (2026-10-06): Display only. The API roster says why carried API rows were not re-run on v1.6 yet (exposure cadence, retired v1.6.0 sealed set) and that no v1.5 to v1.6 conversion is applied.
  • v1.7.1 (2026-10-06): Display only, no score or rank changed. Headings name the board (open weights / API offerings); the API board leads with the Composite Score and lists every API offering we measured, including mode variants, carried v1.5.x rows and wrappers; the Jev reference row reads “Not ranked, only shown as a reference to compare with” on the open-weights board; long base-model notes became numbered footnotes under the ranking; a short expandable note explains the split.
  • v1.7.0 (2026-10-06): Leaderboard split. /jev-models ranks open-weights systems we ran on our own hardware, with Jev 1.13.0 as an unranked reference row and a “Show API offerings” switch; hosted API offerings are ranked on the new /jev-models/api board. No score was recomputed: ranks are the published order filtered to each board. Adds the GPU cost What-If.
  • v1.6.1 (2026-10-05): Hosted APIs (fastino-gliner-2-5-decide, jev-1.13.0, sage-1.3.0, wity-1, wity-1-always, wity-1-off) answer the full 1,500-item set instead of the 600-item API subset and are no longer equated; self-hosted rows unchanged; whole v1.6.0 sealed draw retired. Cost per 1,000 decisions is measured on one common item set (all items except the 23 outside the Jev reference input range) for every token-priced row.
  • v1.6.0 (2026-10-05): Rotating item draw (1,200 sealed + 300 public); hosted APIs on API subsets equated to the self-hosted scale; /jev-models/v1.6.0 stays available unchanged.
  • Earlier releases: v1.5.7, v1.5.6, v1.5.5 and the historical boards below.

System types (colours)

  • Jev — reference (TypeSafe, closed) — The closed system JevBench is named after, shown as the reference.
  • Closed API (weights not public) — Available through a hosted API; the weights cannot be downloaded.
  • Open weights · LLM decoder — An autoregressive language model with public weights, including fine-tunes, merges and Jev rebuilds.
  • Open weights · diffusion LM — A language model that generates by iterative denoising instead of token by token.
  • Open weights · encoder / classifier — BERT-style encoders, NLI zero-shot classifiers and GLiNER-type models.
  • Open weights · reranker — A cross-encoder or LLM reranker that scores options against the input.
  • Base model control (no decision fine-tune, raw logits) — An official open checkpoint without any decision fine-tune, read out from raw logits, used as a floor.
  • System (router / cascade / ensemble) — Several models combined at inference time.

Colours show architecture only; they never change a score or rank. Each class is assigned from cited evidence (config.json, model card, provider docs, or our own run record for hosted APIs).

Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts.

Method amendments

  • 5 Oct 2026, v1.6.1 (Florian's decision): hosted/API systems now answer the same full item set as self-hosted systems (1,200 sealed + 300 public). Their earlier answers on the API subsets were reused and only the missing sealed items were sent; every item was logged in the exposure ledger before it was sent. API rows are no longer equated. Consequence: every v1.6.0 sealed item has now been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases; v1.6.2 onward draws fresh items from the reserve.
  • 5 Oct 2026, v1.6.1 cost rule (Florian's decision): cost per 1,000 decisions is now measured on one common item set for every system: all items except the 23 very long items that lie outside the Jev reference's accepted input range. Before, a system that processed those items paid for their tokens while a system that refused them did not. Refusals still count against Intelligence exactly as before. Hosted APIs with a per-token tariff are priced from their measured tokens on this common set; carried per-decision prices of self-hosted systems are unchanged. Effect: Sage 1.3.0 USD 0.0247 per 1,000 decisions on the common set (0.0650 on all 1,500 items); Jev 1.13.0 0.0323 (unchanged).

The v1.6.1 amendments above say hosted APIs are no longer equated. Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).

Rotating item sets

Each release draws fresh sealed decisions from a larger reserve. Self-hosted open-weights models (run offline on our own GPU pods or Sandy) answer S and P; externally hosted models answer the same full set (S and P) as the self-hosted systems since v1.6.1, so hosted APIs are no longer equated.

SetItemsChoiceNoulScoreAnswered by
S · Sealed release draw1,200600300300self-hosted systems (run offline on our own GPU pods or Sandy) and, from v1.6.1, externally hosted APIs
A · API subset (part of S; used by the v1.6.0 hosted-API measurements)3001507575every system (inside S since v1.6.1); retired for future draws
P · Public set3001507575every system

Not in the table above (frozen release data): API set A4 · 300 never-used sealed items, answered together with P by the 18 API endpoints re-run on 6 Oct 2026 (the 17 ranked offerings and the classifier.dev wrapper); retired after that run.

  • Self-hosted systems: 1,500 items (S 1,200 + P 300). Hosted APIs: 1,500 items (S 1,200 + P 300), the same as self-hosted systems; A (300) is the part of S that earlier API measurements used. Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
  • A sealed item is scored in at most three releases, then retired. An item used in one release is not drawn again in the next. If coverage minimums cannot be met, the release waits for newly reviewed items.
  • The selection seed is committed (SHA-256) before any inference, and the draw is a deterministic function of policy, seed and item id.

API-exposure rule

  • An external exposure is any sealed item sent to an endpoint we do not control: closed APIs, and open-weights models reached through third-party hosts, routers or a submitter's endpoint. Timeouts count as exposure.
  • From v1.6.1 hosted models receive the full item set (S plus P); items they had already answered were reused and only the missing sealed items were sent. Hosted models are re-measured at most once every three refresh releases unless a verified new model version ships. Between measurements they keep their last score with its measurement date.
  • Every externally sent sealed item is logged before dispatch. An item exposed to a provider is never scored again for that provider; once two different providers have received it, it retires for everyone. In v1.6.0 Jev 1.13.0 and Fastino GLiNER-2.5-Decide both received the same A, so A retires globally. With the v1.6.1 amendment every v1.6.0 sealed item has been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases.

Public-versus-sealed gap penalty

Intelligence is half public, half sealed. A system whose public score exceeds its sealed score by more than the field-median gap (G_med = 2.6 points) plus 8 points loses one Intelligence point per excess point. Hosted APIs are compared on P versus A against the same self-hosted systems' P-versus-A gap (-1.2 points).

Hosted APIs (not equated since v1.6.1)

Hosted APIs that answered the full set are scored exactly like self-hosted systems; no equating offset is applied to them. Category and language values stay raw. Exception on the live boards: API rows re-run on the fresh API set A4 (v1.7.7) answered A4 ∪ P (600 items) and are equated with the A2/A3 supplement method; OpenAI Decisions the same way on A5 ∪ P (v1.7.10; see Open weights and API offerings above).

Headline and Composite

Capability = mean(Intelligence, Calibration) for systems within twice the Jev 1.13.0 cost and median latency (the official caps; the sliders change only your view). Composite (option A): Weighted harmonic mean of Intelligence, Calibration, Speed and Cost, multiplied by squared penalties when Intelligence is below its configured floor, or Speed or Cost is below 50. Equal 25% weights do not mean an arithmetic average. It leads the API board and is secondary on open weights. Request types Choice, Noul and Score weigh equally; tiers weigh easy 0.1, standard 0.2, hard 0.4, judge 0.3.

Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1,000 decisions); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s.

Costs of v1.6-measured systems carry each system's published v1.5.4 cost per 1,000 decisions (pricing rules unchanged; v1.6 item lengths differ) unless the row says otherwise; Fastino's is an estimate from its published tariff and measured tokens.

Noul decisiveness and Score baseline (addendum B)

Scored with method option B (scorer setting O1S), selected on 3 Oct 2026 after the v1.6 scores were known and disclosed as a post-results change. Each split × type competence is clipped at 0 before the type weighting, so a type answered no better than chance counts as chance instead of negative. The Score chance baseline is the error of always predicting the mid-scale level, so a flat know-nothing distribution earns about 0. A Score cell whose golds all sit at mid-scale keeps the v1.5 random-level baseline; this only occurs in small breakdown and bootstrap cells. Calibration is unchanged. This run: G_med = 2.57 points.

Failed requests and very long items

23 items of this draw are very long (about 77,000 to 81,000 input tokens). Systems with a shorter context window refuse them; Jev-Omni runs out of GPU memory on them on its listed RTX 6000 (48 GB) recipe; decider-4b v2's server truncates them to about 32,800 tokens and answers; Plumb-4B reads them in full. Each outcome is the system's own recipe and is scored as such. A failed, refused or unparseable answer counts as wrong for Intelligence and stays in the denominator; it does not enter Calibration (the v1.5 rule, applied to every system). Failed answers per system: Sage 1.3.0 0/1500 · Mercury Decide 23/1500 · Jev 1.13.0 23/1500 · wity-1 (Wity, reasoning auto) 0/1500 · SPX-CD Flash 23/1500 · SPX-CD Pro 23/1500 · Fastino GLiNER-2.5-Decide 31/1500 · BB-Qwen3.5-4B-LoRA 0/1500 · wity-1 (Wity, reasoning always) 0/1500 · wity-1 (Wity, reasoning off) 0/1500 · classifier.dev 9/600 · Instinct 9/600 · Vansa-3.4 9/600 · GPT-6 Luna (low reasoning effort) 0/600 · Autoloops – Gemma 4 31B IT 9/600 · GPT-6 Luna (default medium reasoning effort) 4/600 · Fastino GLiDE 9/600 · SimpleJev Qwen3.6-35B-A3B 9/600 · GPT-5.6 Luna 0/600 · Instinct Dual 4B 0/600 · Gemini 3.1 Flash-Lite 1/600 · system-one-open 0/600 · openjev-sglang 9/600 · DeepSeek V4.1 Flash 5/600 · SimpleJev Qwen3.8-27B 9/600 · decision-machine-1 9/600 · JevAct 21/600 · Qwen3.8 27B 8/600 · OpenAI Decisions 2/600 · d1 0/1500 · Microsoft-Decision-1 0/0.

Calibration basis

Systems that return a full probability distribution are calibrated on all components (top-label error, plus distribution distance for Choice and ranked-probability error for Score). Fastino GLiNER-2.5-Decide returns a single confidence value, so its Calibration is the top-label error only and is not like-for-like with full-distribution systems.

Overnight full re-measure (4–5 Oct 2026)

Every system with a reproducible recipe was re-run on the v1.6.0 pool overnight with the same pinned inputs and scorer (method option B / O1S). This page uses scoring round score-v161-3 (v1.6.1) (2026-10-06 00:39:36 UTC). Only complete runs (1,500 items self-hosted, the full API input for hosted APIs) are ranked; partial runs are never ranked, and systems not yet re-measured keep their dated v1.5.x score in the separate table.

  • Scores are the official v1.6.0 scorer (score_v16.py, method option B / O1S, bootstrap B = 1,000) over complete outputs only; each scoring round is kept separately.
  • Hosted APIs are scored on the same full item set as self-hosted systems (S 1,200 + P 300, v1.6.1) and are not equated.
  • GPU-class deviations: reproducible recipes used the hardware listed per row, including H100 for large fast-lane decoders and RTX 6000 for the baseline; hardware differences remain in measured latency. The standard x2 + 0.15 s self-hosted adjustment is an assumption, not a hardware normalization.
  • Wity auto remains outside the latency cap and stays ranked in Composite A. OFF and ALWAYS are unranked variants of the AUTO main row.
  • 6 further measured candidates await a separate publication decision.
  • Sage 1.3.0 (Levanto Labs) was measured on 5 Oct from Sandy (Helsinki), text only. Cost uses the Levanto list tariff (USD 0.05/M input, USD 10/M output; levanto.ai/pricing, read 5 Oct) over all 600 answered rows: USD 0.0766/1,000 decisions. Superseded by the v1.6.1 cost rule (common item set): USD 0.0247/1,000 decisions, inside the Jev-class cost cap.
  • A fresh Monday Jev 1.13.0 re-check on A3 measured p50 0.239 s and equated Capability 77.5 versus the official 76.5, within the confidence interval; the official v1.6.0 Jev row is retained.

Display note on the overnight text above (frozen release data). Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).

Supplementary API draws A2 and A3

History: hosted APIs were first measured on API subsets (A, then A2 and A3) and equated to the self-hosted scale in v1.6.0. From v1.6.1 they answer the same full set as self-hosted systems, so no equating is applied; the earlier subsets A, A2 and A3 are retired.

Per-model exposure counts (hosted and author-hosted endpoints)

System (provider)v1.6 sealed items sentScored sealed setStatus
wity-1-auto (wity via railway)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired
wity-1-off (wity via railway)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired
wity-1-always (wity via railway)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired
jev-1.13.0 (typesafe)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired
fastino-gliner-2-5-decide (fastino)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired
gpt-6-luna (openai)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
gpt-6-luna-low (openai)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
gpt-5.6-luna (openai)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
gemini-3.1-flash-lite (google)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
deepseek-flash (deepseek)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
qwen3.8-27b (chutes)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
decision-machine-1 (milliseconds.ai)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
classifier-dev-fast (classifier.dev)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
vansa-3.4 (vansa)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
kushal-gemma4-31b-it-autoloops (autoloops)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
system-one-open (modal via modal)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
openjev-sglang (modal via modal)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
instinct (zoowork)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
instinct-dual-4b (zoowork)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
simplejev-qwen3.8-27b (featherless)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
simplejev-qwen3.6-35b-a3b (featherless)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
jevact (jevact)300A4 (retired after this run)measured 6 Oct 2026 on A4 ∪ P (v1.7.7)
Sage 1.3.0 (Levanto Labs)1200S (1,200 sealed) + P (300 public)full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired

Provenance

Measured on the v1.6 pool (run completion day, UTC): 2026-10-05: Sage 1.3.0, Jev 1.13.0, Fastino GLiNER-2.5-Decide · 2026-10-06: wity-1 (Wity, reasoning auto), SPX-CD Flash, SPX-CD Pro, wity-1 (Wity, reasoning always), wity-1 (Wity, reasoning off), classifier.dev, Instinct, Vansa-3.4, GPT-6 Luna (low reasoning effort), Autoloops – Gemma 4 31B IT, GPT-6 Luna (default medium reasoning effort), Fastino GLiDE, SimpleJev Qwen3.6-35B-A3B, GPT-5.6 Luna, Instinct Dual 4B, Gemini 3.1 Flash-Lite, system-one-open, openjev-sglang, DeepSeek V4.1 Flash, SimpleJev Qwen3.8-27B, decision-machine-1, JevAct, Qwen3.8 27B, OpenAI Decisions, d1 · 2026-10-07: Mercury Decide, BB-Qwen3.5-4B-LoRA · 2026-10-10: Microsoft-Decision-1.

Aggregate files: results sha256 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1 · categories sha256 0159c4828b3e8620148af06848ef0a5ccfc7d7f0c11aec02460d844dc0c62689 · dated carry sha256 334e851944e92383ab3071e52b3012040c7023d0fe236395953169ebee8f0459 · live category union (data/raw/benchmarks/jevbench/v1.6/jevbench-v1.6.1-category-cells.json) sha256 3b491b5c29d9480cc11c12486c7f507140582f2445e23abcf236a33baa5d7eae. Scoring source sha256 e15347783b8bedeeddd9017a831608a957f5bf0386f05902c50161ce0468f911. The method, release data and carry artifact are independently hashable.

Release history

The board above is always the current merged board: every system with its latest measurement. Each release stays readable as published.

Fresh full-set API addenda

These separately sourced measurements use their own accepted sealed draw. Frozen historical exports remain unchanged. All 50 topic, use-case and language labels are retained; official cells need 15 items and numeric radar spokes need 30.

Valid answers and completed coverage are shown separately. An accepted model/input refusal can count toward coverage while remaining a failed scored answer; authentication, rate-limit and service failures do not.

Microsoft-Decision-1 (Azure Foundry) · all 50 cells

Measured 2026-10-10 · O1S gap reference G_med 7.0881 score points · Aggregate results

Historical comparators keep their own measurement dates, draws and reference values. This addition does not remeasure them on the new pool.

Languages

CategoryItemsValid answersCompleted coverageErrorsOfficial competenceSample statusTypes / pool
English1281126012812156.83radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
German979797051.36radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
French808080049.75radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Spanish949494046.21radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Italian797979045.37radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Portuguese858585058.75radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Dutch787878046.12radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Danish777777049.39radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Swedish737373042.01radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Norwegian646464032.59radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Finnish727272023.96radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Polish969696036.65radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Czech737373038.58radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Greek737373033.36radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Turkish747474042.68radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Ukrainian707070038.45radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Arabic767676044.21radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Hindi858585033.80radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Chinese696868153.79radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Japanese858585054.97radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Korean757575033.10radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Indonesian707070057.30radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Mixed-language797979088.32radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union

Subject topics

CategoryItemsValid answersCompleted coverageErrorsOfficial competenceSample statusTypes / pool
Math & numbers4594484591120.68radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Coding & software277277277057.43radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Rules, policy & law160015991599153.18radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Finance & commerce236236236056.25radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Support & operations1671571671056.25radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Everyday language152152152075.13radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Safety & security114114114053.57radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union

Use cases

CategoryItemsValid answersCompleted coverageErrorsOfficial competenceSample statusTypes / pool
Search and retrieval979797085.32radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Scientific discovery979797036.91radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Model routing400400400064.30radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
LLM guardrails101101101069.68radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Semantic code linting969696019.83radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Feature extraction for predictive modeling969696036.89radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Recruiting989898029.70radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Lead generation114114114063.85radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Customer support249240249960.37radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Insurance claims107107107042.83radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Financial crime116115115128.96radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Legal and compliance471471471063.42radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
E-commerce marketplaces114114114052.63radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Moderation and trust and safety989898067.15radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Advertising107107107037.00radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Gaming909090025.06radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Risk assessment126126126051.21radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Demand forecasting88888800.00radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Graphs and knowledge graphs989898035.49radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Other2422302421247.10radar sample minimum metchoice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.