JevBench — all decision models
A benchmark for AI decision models: state and rubric in, typed answer out. We compare accuracy, calibration, latency and cost, independently of TypeSafe AI.
169 ranked systems · 1,500 decisions each (600 for 19 equated API re-runs) · 14 newer ranked rows tagged with their draw · 13 older rows dated separately · How it works · Release history · Methodology · Data & JSON/CSV · aggregate JSON · Share · Other decision benchmarks · Explore ImageJevBench v0.3.0
JevBench board · headline
JevBench Capability Score (all groups)
Capability ranking of Jev-class systems
Capability Score averages Intelligence and Calibration. Quyet-1.0-Large leads the Jev-class systems with 81.7.
Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘
Adjust cost / latency caps · 2× official
ScoreCCost
ScoreCap.Capability
Score$/1k$/1k decisions
- 1Quyet-1.0-Large73.450.581.7$0.045*
- 2Jeff-1.0-Largev1.6.4 draw71.949.281.5$0.049*
- 3decisio v0.8.0 on gemma-4-31B-it70.451.779.6$0.041*
- 4Sage 1.3.0API65.658.278.6$0.025*
- 5deck-31B73.049.777.6$0.048*
- 6Jev 1.13.0API63.654.777.1$0.032
- 7René-1 31B FP861.746.076.0$0.063*
- 8H2O-Lightning-4B v1.160.060.375.0$0.021*
- 9d1API62.162.874.4$0.017
- 10SPX-CD FlashAPI58.452.174.3$0.039*
Show all 116 Jev-class systems (106 more)
- 11Mercury DecideAPI65.962.173.6$0.018*
- 12OpenAI DecisionsAPI56.948.673.5$0.052
- 13Bobcat Flash 1.265.849.773.5$0.048*
- 14Surogate Rune 26B-A4B v356.149.673.5$0.048*
- 15decider-12b v263.057.872.5$0.026*
- 16Xor 26B-A4B58.750.772.4$0.044*
- 17decider-12b v1, stock Gemma-4-12B-it60.657.872.2$0.026*
- 18torchcast-decision-12b60.556.471.7$0.028*
- 19Jev-Omni55.556.171.3$0.029*
- 20Winnow-12B Q859.556.671.2$0.028*
- 21Cygnet54.856.470.9$0.028*
- 22Microsoft-Decision-1API57.261.570.8$0.019
- 23decisio v0.8.0 on gemma-4-12B-it52.359.370.7$0.023*
- 24decider chat on Gemma-4-31B-it58.550.870.6$0.044*
- 25InstinctAPI47.864.470.5$0.015
- 26GEV-26B-Decide49.851.170.1$0.043*
- 27decisio v0.9.0 on gemma-4-12B-it52.359.370.0$0.023*
- 28Blink v0.3 26B-A4Bv1.6.7 draw49.249.669.9$0.048*
- 29Wald 4B v2.1v1.6.7 draw52.046.369.3$0.062*
- 30Hopper 12B trained51.756.268.1$0.029*
- 31SPX-CD-Omni48.858.368.1$0.025*
- 32Aplomb 145.857.967.5$0.025*
- 33Bespoke Nimble 9B v356.546.166.8$0.063*
- 34Metask rain 4Bv1.6.4 draw51.151.866.0$0.040*
- 35Plumb-4B43.063.165.4$0.017*
- 36decider-4b v240.164.564.9$0.015*
- 37JevK5 v0.338.563.164.2$0.017*
- 38Clef-Flash42.246.863.8$0.059*
- 39Vansa-3.4API40.961.963.7$0.019
- 40deck-4B v1.038.663.563.6$0.017*
- 41janus 4B37.664.462.9$0.015*
- 42JevK5 v0.2.036.363.162.8$0.017*
- 43Hopper34.762.362.4$0.018*
- 44Imajev-4B34.663.361.9$0.017*
- 45jqv33.951.261.9$0.042*
- 46metask-jev-4b39.257.761.8$0.026*
- 47Kahn1 4B39.746.461.7$0.061*
- 48lev41.161.861.7$0.019*
- 49Malkuth-4B33.855.861.1$0.030*
- 50Decision 4B v1.232.463.160.4$0.017*
- 51Wald 4Bv1.6.4 draw41.154.359.8$0.033*
- 52Quyet-1.0-Medium41.763.459.2$0.017*
- 53Decision-4B37.965.959.1$0.014*
- 54Manchego v2.132.564.358.3$0.015*
- 55Messier One v0.242.464.858.1$0.015
- 56Diffusion Jev54.949.658.1$0.048*
- 57jev-local41.958.957.5$0.024*
- 58Instinct Dual 4BAPI28.375.057.2$0.0068
- 59Decision 4B v1.129.163.157.2$0.017*
- 60JEV Qwen3.5-9B Base NVFP429.347.657.1$0.056*
- 61CoCo-Decision-4B32.765.256.9$0.014*
- 62APUS-OpenJev-v1-9B47.248.756.9$0.051*
- 63typecastlm29.864.056.5$0.016*
- 64TypeCastLM 1.4.035.464.856.4$0.015*
- 65SemIf27.163.155.4$0.017*
- 66local-jev Qwen3.5-4B24.159.355.0$0.023*
- 67Decision 2B22.966.454.8$0.013*
- 68spark-s1-4b-v643.260.454.0$0.021*
- 69system-one-openAPI27.968.351.9$0.011*
- 70ClassOne Qwen 3.5 9B42.249.350.5$0.049*
- 71Liquid AI d1-3Bv1.6.7 draw22.150.649.5$0.044*
- 72ZeroEntropy zerank-215.648.449.4$0.052*
- 73open-alternative-jev22.363.349.1$0.017*
- 74OpenSourceJev23.269.148.6$0.011*
- 75EXAONE-4.0-1.2B-JEV v0.314.857.947.7$0.025*
- 76Malkuth-2B15.965.647.6$0.014*
- 77Nemotron Diffusion 8B20.455.547.5$0.030*
- 78decision-machine-1API13.158.346.1$0.024
- 79decisor-4bv1.6.4 draw39.651.445.7$0.042*
- 80Qwen3-Reranker-4B16.348.444.7$0.052*
- 81decider-2b20.764.944.6$0.015*
- 82BAAI bge-reranker-v2-m30.159.344.1$0.023*
- 83kev 4B19.565.844.0$0.014*
- 84Mixedbread mxbai-rerank-base-v21.160.343.9$0.021*
- 85lev-350m4.880.043.3$0.0046*
- 86Certo v10.296.943.2$0.0013*
- 87smalljev semantic-v98.760.843.0$0.020*
- 88jul fast7.569.042.8$0.011*
- 89Alibaba GTE Reranker ModernBERT-base0.069.341.7$0.011*
- 90watt-flash-0.13.889.041.6$0.0023
- 91Decision Fast6.180.140.6$0.0046*
- 92openJev Verdict 1.45.086.640.4$0.0028*
- 93Quyet-1.0-Small-EN6.189.139.5$0.0023*
- 94OpenDecision7.479.138.0$0.0050*
- 95Vega 4Bv1.6.7 draw16.747.237.7$0.058*
- 96verdict-small2.3100.037.6$0.00087*
- 97kev 0.6B7.680.136.9$0.0046*
- 98WaterSheep5.368.735.6$0.011*
- 99Raw Phi-4 mini direct logits13.753.235.3$0.036*
- 100Fastino GLiNER-2.5-DecideAPI7.045.834.6$0.064
- 101Raw Qwen3 4B Instruct 2507 direct logits34.063.434.6$0.017*
- 102Vega 0.8Bv1.6.7 draw3.161.533.0$0.019*
- 103kev 0.5B4.680.132.9$0.0046*
- 104Liquid AI d1-omni-600Mv1.6.7 draw3.373.132.8$0.0079*
- 105Quyet-1.0-Small4.079.632.5$0.0048*
- 106Quyet-1.0-Tiny3.788.431.1$0.0024*
- 107Open Jev JSON Canvas61.249.330.6$0.049*
- 108JevActAPI10.168.628.8$0.011*
- 109openJev Verdict4.186.626.6$0.0028*
- 110Tacet Sonata2.289.024.3$0.0023*
- 111CLM-8B0.150.520.1$0.045*
- 112Deem 0.8B v16.580.318.1$0.0045*
- 113Laya multilingual0.382.217.0$0.0039*
- 114ClassOne Gemma 4 E2B7.870.215.2$0.0098*
- 115Raw Qwen3 1.7B direct logits5.168.58.3$0.011*
- 116Raw Qwen3 0.6B direct logits5.677.68.2$0.0056*
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency correlate here (Spearman ρ = 0.32, n = 169).
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (16)
- Open weights · LLM decoder (117)
- Open weights · diffusion LM (4)
- Open weights · encoder / classifier (18)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- System (router / cascade / ensemble) (3)
- green: ≤ reference
- amber: ≤ cap (2× reference)
- red: > cap
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).
- –GPT-6 Luna (medium)API96.938.498.4$0.11Outside: cost 3.5× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
- –Qwen3.8 27BAPI96.40.098.2$2.18*Outside: cost 67.4× Jev (v1.5 reference), latency 11.5× Jev (v1.5 reference)
- –GPT-6 Luna (low)API96.639.097.8$0.11Outside: cost 3.3× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
- –DeepSeek V4.1 FlashAPI94.719.697.4$0.48Outside: cost 14.9× Jev (v1.5 reference), latency 3.3× Jev (v1.5 reference)
- –GPT-5.6 LunaAPI94.130.294.5$0.21Outside: cost 6.6× Jev (v1.5 reference), latency 2.6× Jev (v1.5 reference)
- –Perplexity Decider v1.1 27B75.729.982.8$0.22*Outside: cost 6.7× Jev (v1.5 reference)
- –torchcast-decision-27b71.830.381.6$0.21*Outside: cost 6.5× Jev (v1.5 reference)
- –wity-1API70.358.879.1$0.024Outside: latency 2.6× Jev (v1.5 reference)API price · eligibility checked at the developer's list price; at base-model pricing it would exceed the cost cap (2.70× Jev)
- –Kev 27B64.218.378.9$0.53*Outside: cost 16.4× Jev (v1.5 reference)
- –SPX-CD ProAPI64.943.178.6$0.079*Outside: cost 2.4× Jev (v1.5 reference)
- –Eikos-27B64.428.777.2$0.24*Outside: cost 7.4× Jev (v1.5 reference)
- –Fastino GLiDEAPI66.135.977.0$0.14Outside: cost 4.3× Jev (v1.5 reference)
- –OpenJev70.328.575.9$0.24*Outside: cost 7.5× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
- –Clef59.128.175.3$0.25*Outside: cost 7.7× Jev (v1.5 reference)
- –SimpleJev Qwen3.8-27B57.915.174.9$0.68*Outside: cost 20.9× Jev (v1.5 reference), latency 2.3× Jev (v1.5 reference)
- –JEV-27B56.831.073.9$0.20*Outside: cost 6.2× Jev (v1.5 reference)
- –RSI-Jev v6.1-VL 27Bv1.6.7 draw61.631.173.5$0.20*Outside: cost 6.1× Jev (v1.5 reference)
- –Autoloops – Gemma 4 31B ITAPI58.140.572.7$0.096Outside: cost 3.0× Jev (v1.5 reference)
- –swanOne56.242.272.4$0.085*Outside: cost 2.6× Jev (v1.5 reference)
- –AutoJev-27B55.629.471.6$0.23*Outside: cost 7.0× Jev (v1.5 reference)
- –SimpleJev Qwen3.8-27BAPI52.914.971.6$0.69*Outside: cost 21.3× Jev (v1.5 reference)
- –Decision 2.0 Vega 27B60.429.671.3$0.22*Outside: cost 6.9× Jev (v1.5 reference)
- –JPT-35B-A3B51.433.670.7$0.16*Outside: cost 5.0× Jev (v1.5 reference)
- –Jebadiah 27B57.629.970.3$0.22*Outside: cost 6.7× Jev (v1.5 reference)
- –NInfer Qwen3.8-Flash-Next mixed49.042.569.7$0.082*Outside: cost 2.5× Jev (v1.5 reference)
- –JevOne46.039.868.9$0.10*Outside: cost 3.1× Jev (v1.5 reference)
- –SimpleJev Qwen3.6-35B-A3BAPI56.335.168.3$0.15*Outside: cost 4.5× Jev (v1.5 reference), latency 2.03× Jev (v1.5 reference)
- –NInfer Qwen3.8-27B NVFP451.829.168.0$0.23*Outside: cost 7.1× Jev (v1.5 reference)
- –Open-Jev 27B v1.152.514.468.0$0.72*Outside: cost 22.2× Jev (v1.5 reference), latency 2.8× Jev (v1.5 reference)
- –LitJev43.328.466.6$0.24*Outside: cost 7.6× Jev (v1.5 reference), latency 5.0× Jev (v1.5 reference)
- –decider-35b-a3b48.734.466.2$0.15*Outside: cost 4.8× Jev (v1.5 reference)
- –Bev / Bonsai 27B43.828.266.1$0.25*Outside: cost 7.6× Jev (v1.5 reference), latency 3.4× Jev (v1.5 reference)
- –Open-Jev 9B43.933.161.8$0.17*Outside: cost 5.3× Jev (v1.5 reference), latency 2.2× Jev (v1.5 reference)
- –Qwen3.5-9B Jev-like data-mix v241.945.760.9$0.065*Outside: cost 2.001× Jev (v1.5 reference)
- –Standard One 8B SHv1.6.7 draw34.543.260.0$0.078*Outside: cost 2.4× Jev (v1.5 reference)
- –Gemini 3.1 Flash-LiteAPI58.629.259.8$0.23Outside: cost 7.1× Jev (v1.5 reference)
- –reflex 4B30.963.159.0$0.017*Outside: latency 4.6× Jev (v1.5 reference)
- –openjev-sglangAPI38.835.658.7$0.14*Outside: cost 4.3× Jev (v1.5 reference), latency 2.1× Jev (v1.5 reference)
- –Standard One 8B32.843.358.4$0.078*Outside: cost 2.4× Jev (v1.5 reference)
- –Clef-omniv1.6.7 draw34.526.858.3$0.28*Outside: cost 8.5× Jev (v1.5 reference)
- –Bespoke Nimble 9B43.136.856.5$0.13*Outside: cost 4.0× Jev (v1.5 reference)
- –kev 8B29.340.450.0$0.097*Outside: cost 3.0× Jev (v1.5 reference)
- –JADE47.933.649.2$0.16*Outside: cost 5.1× Jev (v1.5 reference)
- –Open-Jev 2B21.833.147.7$0.17*Outside: cost 5.3× Jev (v1.5 reference)
- –Laya typed-decisions5.582.544.2$0.0038*Outside: latency 3.0× Jev (v1.5 reference)
- –Gutsy 0.8B v0.3v1.6.7 draw8.678.443.1$0.0053*Outside: latency 2.7× Jev (v1.5 reference)
- –Qwen3.5-0.8B Decision Model7.379.542.6$0.0048*Outside: latency 18.3× Jev (v1.5 reference)
- –open-jev-deberta-v3-large2.877.635.4$0.0056*Outside: latency 5.7× Jev (v1.5 reference)
- –Laya1.884.932.5$0.0032*Outside: latency 2.9× Jev (v1.5 reference)
- –SimpleJev13.070.331.6$0.0098*Outside: latency 14.3× Jev (v1.5 reference)
- –Raw Qwen3 8B direct logits26.545.530.1$0.065*Outside: cost 2.02× Jev (v1.5 reference)
- –system-one29.345.029.7$0.068*Outside: cost 2.1× Jev (v1.5 reference)
- –Mirror0.089.325.8$0.0023*Outside: latency 2.4× Jev (v1.5 reference)
Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 116 of 169 systems qualify; the other 53, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed (all groups)
Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (16)
- Open weights · LLM decoder (117)
- Open weights · diffusion LM (4)
- Open weights · encoder / classifier (18)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- System (router / cascade / ensemble) (3)
- faint = outside Jev-class
Pareto frontier: capability against cost and speed (all groups)
The red line joins the systems nobody beats on both axes at once; hover or tap a point for the system that beats it. Shaded: beyond the Jev-class cap.
14 of 169 systems on the frontier: no other system is both cheaper and more capable.
9 of 169 systems on the frontier: no other system is both faster and more capable.
Model kind
Jev-class
Release
No system in this release reports an exact parameter count.
JevBench board
JevBench Composite Score (all groups): 169 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Adjust weights ↓
The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis, and the column headings sort too.
Greener = stronger within its column.
169 of 169 systems, sorted by official rank, #1 first.
- 1Sage 1.3.0APInewClosed APIBase model: undisclosed74.0I 66C 92S 94K 58est.$0.025
- 2d1APInewClosed APIBase model: undisclosed73.0I 62C 87S 89K 63$0.017
- 3H2O-Lightning-4B v1.1newfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource72.5I 60C 90S 93K 60est.$0.021
- 4Mercury DecideAPInewSystemBase model: undisclosedsource72.4I 66C 81S 85K 62est.$0.018
- 5decisio v0.8.0 on gemma-4-31B-itnew31B · FP8 on loadBase model: google/gemma-4-31B-itsource71.7I 70C 89S 91K 52est.$0.041
- 6Jev 1.13.0APIJev referenceBase model: undisclosed71.5I 64C 91S 91K 55$0.032
- 7Quyet-1.0-Largenewadapter · 31BBase model: google/gemma-4-31B-itsource71.4I 73C 90S 87K 51est.$0.045
- 8decider-12b v2newfine-tune · 12BBase model: google/gemma-4-12B-itsource70.9I 63C 82S 90K 58est.$0.026
- 9wity-1APInewClosed APIBase model: undisclosedBase-model reference price (base undisclosed at the author's request): 44.0 (would be #41)70.8I 70C 88S 72K 59$0.024
- 10decider-12b v1, stock Gemma-4-12B-itnew12BBase model: google/gemma-4-12B-itsource70.4I 61C 84S 90K 58est.$0.026
- 11torchcast-decision-12bnewadapter · 12BBase model: google/gemma-4-12B-itsource69.9I 61C 83S 92K 56est.$0.028
- 12Microsoft-Decision-1APInewClosed APIBase model: Qwen3.5-9Bsource*169.1I 57C 84S 82K 61$0.019
- 13Winnow-12B Q8merge · 12B · GGUF Q8_0Base model: google/gemma-4-12B-itsource68.9I 60C 83S 87K 57est.$0.028
- 14deck-31Bnew31B · FP8Base model: google/gemma-4-31B-itsource68.7I 73C 82S 87K 50est.$0.048
- 15Jeff-1.0-Largenewfine-tune · 31BBase model: undisclosed68.6I 72C 91S 89K 49est.$0.049
- 16Cygnet12BBase model: google/gemma-4-12B-itsource68.6I 55C 87S 92K 56est.$0.028
- 17decisio v0.8.0 on gemma-4-12B-itnew12BBase model: google/gemma-4-12B-itsource68.2I 52C 89S 87K 59est.$0.023
- 18decisio v0.9.0 on gemma-4-12B-itnew12BBase model: google/gemma-4-12B-itsource67.8I 52C 88S 86K 59est.$0.023
- 19Jev-Omnifine-tuneBase model: google/gemma-4-12B-itsource67.7I 56C 87S 85K 56est.$0.029
- 20Xor 26B-A4Bnewmerge · 26B-A4BBase model: google/gemma-4-26B-A4B-itsource67.4I 59C 86S 91K 51est.$0.044
Show all 169 systems (149 more ranked)
- 21Bobcat Flash 1.2newfine-tune · 26B / 4B active · FP8 on loadBase model: google/gemma-4-26B-A4B-itsource67.3I 66C 81S 91K 50est.$0.048
- 22SPX-CD FlashAPInewSystemBase model: undisclosedsource66.7I 58C 90S 79K 52est.$0.039
- 23decider chat on Gemma-4-31B-itnew31BBase model: google/gemma-4-31B-itsource66.1I 59C 83S 86K 51est.$0.044
- 24Hopper 12B trainednewfine-tune · 12BBase model: google/gemma-4-12B-itsource65.8I 52C 85S 85K 56est.$0.029
- 25Surogate Rune 26B-A4B v3new26B-A4BBase model: google/gemma-4-26B-A4B-itsource*265.0I 56C 91S 87K 50est.$0.048
- 26GEV-26B-Decidenewdistilled · 26B-A4BBase model: google/gemma-4-26B-A4B-itsource64.2I 50C 90S 90K 51est.$0.043
- 27SPX-CD-Omninewfine-tune · 12BBase model: google/gemma-4-12B-itsource63.2I 49C 87S 89K 58est.$0.025
- 28Metask rain 4Bnewfine-tune · 4BBase model: undisclosed62.7I 51C 81S 80K 52est.$0.040
- 29OpenAI DecisionsAPInewClosed APIBase model: undisclosed62.5I 57C 90S 89K 49$0.052
- 30InstinctAPIClosed APIBase model: Qwen3.8-27Bsource*362.2I 48C 93S 87K 64$0.015
- 31Blink v0.3 26B-A4Bnewfine-tune · 26B-A4B · NVFP4Base model: undisclosed61.2I 49C 91S 91K 50est.$0.048
- 32Diffusion Jevnew26B / 4B activeBase model: google/diffusiongemma-26B-A4B-itsource59.3I 55C 61S 87K 50est.$0.048
- 33René-1 31B FP8newfine-tune · 31B · FP8Base model: google/gemma-4-31B-itsource55.8I 62C 90S 87K 46est.$0.063
- 34Aplomb 1newfine-tune · 5.3BBase model: Qwen/Qwen3.5-4Bsource54.6I 46C 89S 89K 58est.$0.025
- 35Wald 4B v2.1newfine-tune · 4BBase model: undisclosed54.3I 52C 87S 93K 46est.$0.062
- 36Bespoke Nimble 9B v3newadapter · 9BBase model: Qwen/Qwen3.5-9Bsource53.5I 57C 77S 89K 46est.$0.063
- 37APUS-OpenJev-v1-9Bnewfine-tune · 9BBase model: Qwen/Qwen3.5-9Bsource49.2I 47C 67S 82K 49est.$0.051
- 38Plumb-4Bfine-tune · 4BBase model: alibiserikbay/JevK5source48.3I 43C 88S 94K 63est.$0.017
- 39SPX-CD ProAPInewSystemBase model: undisclosedsource47.7I 65C 92S 78K 43est.$0.079
- 40Messier One v0.2newfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource45.3I 42C 74S 93K 65$0.015
- 41spark-s1-4b-v6adapter · 4BBase model: Qwen3.5-4Bsource44.9I 43C 65S 87K 60est.$0.021
- 42swanOneNVFP4Base model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source43.8I 56C 89S 83K 42est.$0.085
- 43Quyet-1.0-Mediumnewadapter · 4BBase model: Qwen/Qwen3.5-4Bsource43.6I 42C 77S 89K 63est.$0.017
- 44Vansa-3.4APIClosed APIBase model: Qwen3.5-4B (self-reported)*442.6I 41C 87S 95K 62$0.019
- 45levadapter · 4BBase model: Qwen/Qwen3.5-4Bsource42.2I 41C 82S 88K 62est.$0.019
- 46NInfer Qwen3.8-Flash-Next mixedLLM decoderBase model: Qwen3.8-Flash-Nextsource42.0I 49C 90S 89K 43est.$0.082
- 47jev-local9BBase model: Qwen/Qwen3.5-9Bsource41.3I 42C 73S 76K 59est.$0.024
- 48decider-4b v2adapter · 4BBase model: Qwen/Qwen3.5-4B-Basesource41.2I 40C 90S 92K 65est.$0.015
- 49GPT-6 Luna (low reasoning effort)APIClosed APIBase model: undisclosed40.5I 97C 99S 72K 39$0.108
- 50Autoloops – Gemma 4 31B ITAPIClosed APIBase model: undisclosed40.2I 58C 87S 85K 40$0.096
- 51Wald 4Bnewfine-tune · 4BBase model: undisclosed40.0I 41C 78S 82K 54est.$0.033
- 52GPT-6 Luna (default medium reasoning effort)APIClosed APIBase model: undisclosed39.2I 97C 100S 72K 38$0.113
- 53ClassOne Qwen 3.5 9Bnew9BBase model: Qwen/Qwen3.5-9Bsource38.5I 42C 59S 90K 49est.$0.049
- 54JevK5 v0.3adapterBase model: Qwen/Qwen3.5-4Bsource37.4I 39C 90S 94K 63est.$0.017
- 55metask-jev-4b4BBase model: Qwen3.5-4Bsource37.3I 39C 84S 91K 58est.$0.026
- 56deck-4B v1.0newadapter · FP8Base model: alibiserikbay/JevK5 (built on Qwen/Qwen3.5-4B)source 1source 237.3I 39C 89S 90K 63est.$0.017
- 57Clef-Flashfine-tuneBase model: Qwen/Qwen3.5-9BsourceWorkers AI price scenario (latency unmeasured): 39.2 (would be #52)36.6I 42C 85S 88K 47est.$0.059
- 58Decision-4BadapterBase model: Qwen/Qwen3.5-4Bsource35.4I 38C 80S 91K 66est.$0.014
- 59janus 4Bnewfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource35.2I 38C 88S 93K 64est.$0.015
- 60Qwen3.5-9B Jev-like data-mix v2adapter · 9BBase model: Qwen/Qwen3.5-9Bsource33.6I 42C 80S 85K 46est.$0.065
- 61decisor-4bnewfine-tune · 4B · FP8Base model: undisclosed33.5I 40C 52S 92K 51est.$0.042
- 62JevK5 v0.2.0adapter · 4BBase model: Qwen3.5-4Bsource32.1I 36C 89S 92K 63est.$0.017
- 63JevOneadapterBase model: Qwen/Qwen3.6-35B-A3Bsource31.1I 46C 92S 90K 40est.$0.101
- 64Kahn1 4Bnewfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource31.1I 40C 84S 90K 46est.$0.061
- 65Fastino GLiDEAPInewClosed APIBase model: undisclosed30.6I 66C 88S 77K 36$0.137
- 66TypeCastLM 1.4.0new3.8BBase model: Qwen/Qwen3.5-4Bsource29.8I 35C 77S 93K 65est.$0.015
- 67Hopperadapter · 4BBase model: Qwen/Qwen3.5-4Bsource28.9I 35C 90S 91K 62est.$0.018
- 68Imajev-4Badapter · 4BBase model: Qwen/Qwen3.5-4Bsource28.7I 35C 89S 91K 63est.$0.017
- 69SimpleJev Qwen3.6-35B-A3BAPI35B-A3BBase model: undisclosedsource27.6I 56C 80S 77K 35est.$0.145
- 70Malkuth-4Badapter · 4BBase model: Qwen/Qwen3.5-4B-Basesource26.2I 34C 88S 90K 56est.$0.030
- 71jqv32BBase model: Qwen3-32Bsource25.4I 34C 90S 83K 51est.$0.042
- 72JPT-35B-A3Bnewmerge · 35B-A3BBase model: Qwen/Qwen3.5-35B-A3Bsource25.4I 51C 90S 91K 34est.$0.163
- 73decider-35b-a3bfine-tune · 35B-A3BBase model: Qwen/Qwen3.5-35B-A3B-Basesource24.8I 49C 84S 91K 34est.$0.154
- 74CoCo-Decision-4Bnewfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource24.6I 33C 81S 91K 65est.$0.014
- 75Manchego v2.1adapter · 4BBase model: Qwen/Qwen3.5-4Bsource24.5I 33C 84S 91K 64est.$0.015
- 76Decision 4B v1.2adapter · 4BBase model: Qwen/Qwen3.5-4Bsource24.5I 32C 88S 94K 63est.$0.017
- 77Raw Qwen3 4B Instruct 2507 direct logitsoriginal · 4BBase model: Qwen/Qwen3-4B-Instruct-2507 (unmodified)source21.8I 34C 35S 91K 63est.$0.017
- 78RSI-Jev v6.1-VL 27Bnewmerge · 27BBase model: undisclosed21.6I 62C 85S 85K 31est.$0.197
- 79GPT-5.6 LunaAPIClosed APIBase model: undisclosed21.4I 94C 95S 73K 30$0.213
- 80JEV-27Bnewdistilled · 27BBase model: Qwen/Qwen3.8-27Bsource21.3I 57C 91S 87K 31est.$0.199
- 81torchcast-decision-27bnewmerge · 27BBase model: startlux-models/StartLux-Decision-27Bsource 1source 2*521.3I 72C 91S 88K 30est.$0.210
- 82Bespoke Nimble 9Badapter · 9BBase model: Qwen3.5-9Bsource21.1I 43C 70S 86K 37est.$0.128
- 83reflex 4Badapter · 4BBase model: Qwen/Qwen3.5-4Bsource20.6I 31C 87S 69K 63est.$0.017
- 84Perplexity Decider v1.1 27Bnewfine-tune · 27BBase model: Qwen/Qwen3.8-27Bsource20.6I 76C 90S 87K 30est.$0.217
- 85JADEnewadapter · 27BBase model: Qwen/Qwen3.8-27Bsource20.2I 48C 51S 86K 34est.$0.164
- 86typecastlm4BBase model: Qwen3.5-4Bsource19.7I 30C 83S 93K 64est.$0.016
- 87Jebadiah 27Bnewmerge · 27BBase model: Qwen/Qwen3.8-27Bsource19.3I 58C 83S 86K 30est.$0.217
- 88Decision 2.0 Vega 27Bnewadapter · 27BBase model: Qwen/Qwen3.8-27Bsource 1source 218.9I 60C 82S 84K 30est.$0.222
- 89Standard One 8B SHnewfine-tune · 8BBase model: undisclosed18.7I 34C 85S 84K 43est.$0.078
- 90Decision 4B v1.1adapter · 4BBase model: Qwen/Qwen3.5-4Bsource18.7I 29C 85S 94K 63est.$0.017
- 91AutoJev-27Bfine-tune · 27BBase model: Qwen/Qwen3.8-27Bsource18.5I 56C 88S 89K 29est.$0.226
- 92Eikos-27Badapter · 27BBase model: Qwen/Qwen3.8-27Bsource18.1I 64C 90S 88K 29est.$0.238
- 93Instinct Dual 4BAPIClosed APIBase model: Qwen3.5-4Bsource*618.0I 28C 86S 89K 75$0.0068
- 94NInfer Qwen3.8-27B NVFP427BBase model: Qwen3.8-27Bsource17.7I 52C 84S 90K 29est.$0.231
- 95OpenJev26B-A4BBase model: google/diffusiongemma-26B-A4B-itsource 1source 217.3I 70C 81S 73K 29est.$0.241
- 96Open-Jev 9Badapter · 9BBase model: Qwen/Qwen3.5-9Bsource17.0I 44C 80S 74K 33est.$0.170
- 97Gemini 3.1 Flash-LiteAPIClosed APIBase model: undisclosed16.9I 59C 61S 79K 29$0.230
- 98Standard One 8Badapter · 8BBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source16.9I 33C 84S 93K 43est.$0.078
- 99Cleffine-tuneBase model: Qwen/Qwen3.8-27BsourceWorkers AI price scenario (latency unmeasured): 29.4 (would be #67)16.8I 59C 92S 83K 28est.$0.249
- 100system-one-openAPIadapter · 2BBase model: Gemma 4 E2Bsource16.4I 28C 76S 79K 68est.$0.011
- 101JEV Qwen3.5-9B Base NVFP49B · NVFP4Base model: ig1/Qwen3.5-9B-NVFP4source16.0I 29C 85S 94K 48est.$0.056
- 102SemIf4BBase model: Qwen/Qwen3.5-4Bsource15.5I 27C 84S 91K 63est.$0.017
- 103openjev-sglangAPI35B-A3BBase model: Qwen/Qwen3.6-35B-A3Bsource15.4I 39C 79S 78K 36est.$0.140
- 104Bev / Bonsai 27B27B · GGUF PQ2_0Base model: Qwen/Qwen3.8-27Bsource11.7I 44C 89S 73K 28est.$0.247
- 105LitJev27BBase model: Qwen/Qwen3.8-27Bsource11.5I 43C 90S 68K 28est.$0.244
- 106local-jev Qwen3.5-4Bfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource11.4I 24C 86S 86K 59est.$0.023
- 107system-one8BBase model: Qwen3-8Bsource11.0I 29C 30S 91K 45est.$0.068
- 108kev 8Badapter · 8BBase model: Qwen/Qwen3-8B-Basesource10.6I 29C 71S 89K 40est.$0.097
- 109Decision 2Badapter · 2BBase model: openbmb/MiniCPM5-2Bsource10.3I 23C 87S 90K 66est.$0.013
- 110OpenSourceJevadapter · 4B · GGUF Q4_K_MBase model: Qwen3.5-4Bsource10.3I 23C 74S 77K 69est.$0.011
- 111open-alternative-jev4BBase model: Qwen3.5-4Bsource9.4I 22C 76S 91K 63est.$0.017
- 112Raw Qwen3 8B direct logitsoriginal · 8BBase model: Qwen/Qwen3-8B-Basesource9.3I 27C 34S 89K 46est.$0.065
- 113Liquid AI d1-3Bnewfine-tune · 3BBase model: undisclosed8.9I 22C 77S 94K 51est.$0.044
- 114decider-2badapter · 2BBase model: Qwen/Qwen3.5-2B-Basesource7.7I 21C 69S 95K 65est.$0.015
- 115Nemotron Diffusion 8B8BBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource7.3I 20C 75S 94K 55est.$0.030
- 116DeepSeek V4.1 FlashAPILLM decoderBase model: undisclosed7.1I 95C 100S 68K 20$0.480
- 117kev 4Badapter · 4BBase model: Qwen/Qwen3-4B-Basesource*76.6I 20C 68S 90K 66est.$0.014
- 118Clef-omninewfine-tune · 30B-A3BBase model: undisclosed6.1I 34C 82S 87K 27est.$0.276
- 119Kev 27Bnewfine-tune · 27BBase model: Qwen/Qwen3.8-27Bsource5.8I 64C 94S 84K 18est.$0.529
- 120Malkuth-2Badapter · 2BBase model: empero-ai/Qwen3.8-2B-Distillsource4.0I 16C 79S 92K 66est.$0.014
- 121Qwen3-Reranker-4BRerankerBase model: Qwen/Qwen3-4B-Basesource3.7I 16C 73S 82K 48est.$0.052
- 122Vega 4Bnewadapter · 4BBase model: undisclosed3.6I 17C 59S 84K 47est.$0.058
- 123SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool)new27BBase model: Qwen/Qwen3.8-27Bsource3.4I 58C 92S 76K 15est.$0.676
- 124ZeroEntropy zerank-2RerankerBase model: Qwen/Qwen3-4Bsource3.3I 16C 83S 83K 48est.$0.052
- 125EXAONE-4.0-1.2B-JEV v0.3new1.2BBase model: LGAI-EXAONE/EXAONE-4.0-1.2Bsource3.3I 15C 81S 94K 58est.$0.025
- 126Open-Jev 2Badapter · 2BBase model: Qwen/Qwen3.5-2Bsource3.2I 22C 74S 77K 33est.$0.170
- 127SimpleJev Qwen3.8-27BAPI27BBase model: Qwen/Qwen3.8-27Bsource3.2I 53C 90S 77K 15est.$0.687
- 128Open-Jev 27B v1.1adapter · 27BBase model: Qwen/Qwen3.8-27Bsource2.9I 52C 83S 73K 14est.$0.716
- 129Raw Phi-4 mini direct logitsoriginalBase model: microsoft/Phi-4-mini-instruct (unmodified)source2.5I 14C 57S 92K 53est.$0.036
- 130decision-machine-1APIClosed APIBase model: undisclosed2.4I 13C 79S 93K 58$0.024
- 131SimpleJev0.8BBase model: Qwen/Qwen3.5-0.8Bsource2.1I 13C 50S 59K 70est.$0.0098
- 132JevActAPIClosed APIBase model: undisclosedsource1.1I 10C 47S 76K 69est.$0.011
- 133smalljev semantic-v9adapter · 2BBase model: openbmb/MiniCPM5-2B-Basesource0.8I 9C 77S 90K 61est.$0.020
- 134Gutsy 0.8B v0.3newfine-tune · 0.8B · GGUF Q8_0Base model: undisclosed0.8I 9C 78S 73K 78est.$0.0053
- 135kev 0.6Badapter · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.5I 8C 66S 92K 80est.$0.0046
- 136jul fastnew2BBase model: openbmb/MiniCPM5-2Bsource0.5I 7C 78S 91K 69est.$0.011
- 137OpenDecisionEncoder / classifierBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.5I 7C 69S 87K 79est.$0.0050
- 138ClassOne Gemma 4 E2BnewE2BBase model: google/gemma-4-E2B-itsource0.5I 8C 23S 93K 70est.$0.0098
- 139Qwen3.5-0.8B Decision Modelfine-tune · 0.8BBase model: Qwen/Qwen3.5-0.8B-Basesource0.5I 7C 78S 56K 80est.$0.0048
- 140Fastino GLiNER-2.5-DecideAPInewfine-tune · 340MBase model: fastino/gliner2-large-v1source 1source 2*80.3I 7C 62S 85K 46$0.064
- 141Deem 0.8B v1LLM decoderBase model: Qwen/Qwen3.5-0.8Bsource0.3I 7C 30S 85K 80est.$0.0045
- 142Decision Fastadapter · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.3I 6C 75S 91K 80est.$0.0046
- 143Quyet-1.0-Small-ENnewfine-tune · 153MBase model: answerdotai/ModernBERT-basesource0.3I 6C 73S 95K 89est.$0.0023
- 144Laya typed-decisionsfine-tune · 421MBase model: ModernBERT-largesource0.2I 5C 83S 71K 83est.$0.0038
- 145WaterSheepnewEncoder / classifierBase model: answerdotai/ModernBERT-basesource0.2I 5C 66S 78K 69est.$0.011
- 146openJev Verdict 1.4Encoder / classifierBase model: knowledgator/gliclass-modern-base-v2.0source0.2I 5C 76S 81K 87est.$0.0028
- 147Raw Qwen3 0.6B direct logitsoriginal · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.2I 6C 11S 94K 78est.$0.0056
- 148lev-350madapterBase model: LiquidAI/LFM2.5-350Msource0.2I 5C 82S 94K 80est.$0.0046
- 149Raw Qwen3 1.7B direct logitsoriginal · 1.7BBase model: Qwen/Qwen3-1.7B-Basesource0.1I 5C 12S 93K 69est.$0.011
- 150kev 0.5Badapter · 0.5BBase model: Qwen/Qwen2.5-0.5Bsource0.1I 5C 61S 92K 80est.$0.0046
- 151openJev Verdictfine-tuneBase model: knowledgator/gliclass-modern-base-v2.0source 1source 20.1I 4C 49S 79K 87est.$0.0028
- 152Quyet-1.0-Smallnewfine-tune · 328MBase model: aisingapore/SEA-LION-ModernBERT-300Msource0.1I 4C 61S 95K 80est.$0.0048
- 153watt-flash-0.1newfine-tune · 140MBase model: jhu-clsp/mmBERT-smallsource0.1I 4C 79S 76K 89$0.0023
- 154Quyet-1.0-Tinynewdistilled · 183MBase model: jhu-clsp/mmBERT-smallsource0.1I 4C 58S 95K 88est.$0.0024
- 155Liquid AI d1-omni-600Mnewfine-tune · 600MBase model: undisclosed0.0I 3C 62S 94K 73est.$0.0079
- 156Vega 0.8Bnewadapter · 0.8BBase model: undisclosed0.0I 3C 63S 89K 61est.$0.019
- 157open-jev-deberta-v3-largeadapterBase model: microsoft/deberta-v3-largesource0.0I 3C 68S 68K 78est.$0.0056
- 158verdict-smallfine-tune · 118MBase model: intfloat/multilingual-e5-smallsource0.0I 2C 73S 89K 100est.$0.0009
- 159Tacet Sonatanewfine-tune · 144MBase model: jhu-clsp/mmBERT-smallsource0.0I 2C 46S 80K 89est.$0.0023
- 160Layafine-tune · 421MBase model: ModernBERT-largesource0.0I 2C 63S 73K 85est.$0.0032
- 161Mixedbread mxbai-rerank-base-v2RerankerBase model: Qwen2.5 (size not stated)source 1source 20.0I 1C 87S 93K 60est.$0.021
- 162Laya multilingualfine-tune · 322MBase model: mmBERT-basesource0.0I 0C 34S 78K 82est.$0.0039
- 163Certo v1Encoder / classifierBase model: ModernBERT-largesource0.0I 0C 86S 93K 97est.$0.0013
- 164CLM-8Bfine-tune · 8BBase model: Qwen/Qwen3-8Bsource0.0I 0C 40S 93K 51est.$0.045
- 165BAAI bge-reranker-v2-m3fine-tuneBase model: BAAI/bge-m3source0.0I 0C 88S 94K 59est.$0.023
- 166Alibaba GTE Reranker ModernBERT-basefine-tuneBase model: answerdotai/ModernBERT-basesource0.0I 0C 83S 94K 69est.$0.011
- 167MirrorEncoder / classifierBase model: undisclosed0.0I 0C 52S 73K 89est.$0.0023
- 168Open Jev JSON Canvas26B-A4BBase model: google/diffusiongemma-26B-A4B-itsource0.0I 61C 0S 86K 49est.$0.049
- 169Qwen3.8 27BAPI27BBase model: undisclosed0.0I 96C 100S 57K 0est.$2.178
Base-model notes
- *1 Microsoft-Decision-1 (base model: Qwen3.5-9B): Base model disclosed by Microsoft; the Decision-1 endpoint is evaluated as a hosted API offering.
- *2 Surogate Rune 26B-A4B v3 (base model: google/gemma-4-26B-A4B-it): This model also powers our System1 Models s1-pro service.
- *3 Instinct (base model: Qwen3.8-27B): Operator-reported; weights not publicly verifiable.
- *4 Vansa-3.4 (base model: Qwen3.5-4B (self-reported)): Developer-reported to Benchmark Heaven (private correspondence, 1 Oct 2026): built on Qwen3.5-4B via an open decision fine-tune, with Vansa's own adapters; the intermediate fine-tune is not named. Not publicly documented or independently verified.
- *5 torchcast-decision-27b (base model: startlux-models/StartLux-Decision-27B): Direct base is StartLux-Decision-27B (startlux-models, CC-BY-NC-4.0), itself a modified Alibaba Qwen model; its card does not name the Qwen version. Qwen3.8-27B is inferred from the identical text_config.
- *6 Instinct Dual 4B (base model: Qwen3.5-4B): Operator-reported; weights not publicly verifiable.
- *7 kev 4B (base model: Qwen/Qwen3-4B-Base): Evaluated Qwen3 variant; later releases use a different base.
- *8 Fastino GLiNER-2.5-Decide (base model: fastino/gliner2-large-v1): Measured via Fastino's hosted API. A 340M English classifier (8k context), not built for multi-step reasoning: near chance on Noul and Score items; 23 very long items exceed its context. Our request follows Fastino's documented format. The open checkpoint scores 11.2 on the community Decision Index (rank 53 of 70; Jev 57.9). Fastino's larger GLiDE is a separate model, ranked separately on the API board.
Adjust weights ↓
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (16)
- Open weights · LLM decoder (117)
- Open weights · diffusion LM (4)
- Open weights · encoder / classifier (18)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- System (router / cascade / ensemble) (3)
- Striped bar = same system under the labelled alternative price assumption
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
- A: Jev 1.13.0REFERENCE · TypeSafe — Jev — reference (TypeSafe, closed) · Score 71.5 (#6)
- B: Sage 1.3.0 — Closed API (weights not public) · Score 74.0 (#1)
The four score axes
Capability by subject topic
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
What each category means · items per category
- Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open / 1465 sealed)
- Coding & software — code, SQL, repositories, developer tools and IT systems. 382 items (62 open / 320 sealed)
- Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open / 333 sealed)
- Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open / 211 sealed)
- Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open / 139 sealed)
- Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open / 134 sealed)
- Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open / 103 sealed)
Use cases (TypeSafe categories)
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
What each category means · items per category
- Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open / 438 sealed)
- Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open / 441 sealed)
- Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open / 203 sealed)
- Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open / 150 sealed)
- E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open / 110 sealed)
- Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open / 110 sealed)
- Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open / 110 sealed)
- Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open / 92 sealed)
- Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open / 94 sealed)
- Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open / 91 sealed)
- Recruiting — resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open / 93 sealed)
- Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open / 91 sealed)
- LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open / 88 sealed)
- Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open / 86 sealed)
- Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open / 87 sealed)
- Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open / 85 sealed)
- Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open / 88 sealed)
- Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open / 85 sealed)
- Gaming — player reports, in-game chat, game support. 86 items (3 open / 83 sealed)
- Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open / 80 sealed)
Competence per request type, open / sealed
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — open set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — sealed set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
How the categories were made
O1S chance-corrected competence: equal mean across the request types present, clipped to 0–100. Zero means the equal-weight mean across present request types was at or below baseline; negative means are clipped. It does not mean every type scored zero. Choice uses a random-option baseline; Noul uses a fixed 50% accuracy baseline (abstentions count as wrong); Score uses the midpoint-guess error (random-guess fallback when every gold is at the midpoint). Compare models within a category; row pools and supported types can differ.
Categories use ruled subject-topic labels and authoring use-case precedence across the answered pool union, with stable item identities counted once.
- Official O1S; recorded failures count where supported; no missing response is imputed; stable identity dedupe parent first.
- S-based category cells pool recorded supported responses from S+P and the answered L1/L2/L3 supplements. API overlay category cells keep their separately recorded A4/A5+P union. Headline and per-type/tier scores are unchanged.
All values as a table
| Spoke | A: Jev 1.13.0 | B: Sage 1.3.0 |
|---|---|---|
| The four score axes | ||
| Intelligence | 63.6 | 65.6 |
| Calibration | 90.6 | 91.6 |
| Speed | 91.5 | 93.6 |
| Cost | 54.7 | 58.2 |
| Capability by subject topic | ||
| Rules, policy & law | 49.4 | 57.9 |
| Coding & software | 64.3 | 77.3 |
| Math & numbers | 10.4 | 17.6 |
| Finance & commerce | 40.9 | 52.7 |
| Support & operations | 48.8 | 71.1 |
| Everyday language | 77.5 | 78.6 |
| Safety & security | 52.2 | 73.1 |
| Use cases (TypeSafe categories) | ||
| Model routing | 73.0 | 78.4 |
| Legal & compliance | 68.6 | 65.3 |
| Customer support | 49.6 | 61.4 |
| Other | 40.1 | 44.3 |
| E-commerce | 40.5 | 59.2 |
| Insurance claims | 46.1 | 51.4 |
| Risk assessment | 32.8 | 41.0 |
| Financial crime | 35.8 | 29.3 |
| Feature extraction | 28.3 | 40.8 |
| Lead generation | 49.2 | 63.8 |
| Recruiting | 0.0 | 19.9 |
| Knowledge graphs | 28.8 | 72.8 |
| LLM guardrails | 68.0 | 91.3 |
| Moderation | 51.6 | 66.4 |
| Code linting | 29.1 | 62.9 |
| Search & retrieval | 80.9 | 90.0 |
| Science | 21.2 | 35.2 |
| Advertising | 16.4 | 43.8 |
| Gaming | 5.1 | 16.8 |
| Demand forecasting | 0.0 | 0.0 |
| Competence per request type, open / sealed | ||
| Choice · open | 79.0 | 85.1 |
| Choice · sealed | 75.8 | 79.7 |
| Noul · open | 54.5 | 54.3 |
| Noul · sealed | 46.1 | 56.4 |
| Score · open | 63.3 | 57.7 |
| Score · sealed | 63.0 | 60.5 |
| Competence per tier — open set | ||
| Easy | 93.9 | 89.6 |
| Standard | 63.2 | 62.3 |
| Judge | 67.7 | 67.5 |
| Hard | 65.6 | 70.5 |
| Competence per tier — sealed set | ||
| Easy | 78.4 | 77.5 |
| Standard | 68.7 | 75.8 |
| Judge | 66.9 | 73.2 |
| Hard | 58.9 | 60.8 |
Category radars count each answered item once from the pools named under each radar. API overlay rows use A4+P+L1+L2+L3; A5+P+L1+L2+L3; S+P+L1+L2+L3. Raw and unequated; cells under 15 answered items are omitted. Per-type and tier radars retain each row's original measurement pools: A4/A5 rows have 300 open plus 300 sealed items; full-set rows have S 1,200 plus P 300.
Languages
Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts. Cells under 15 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. The mixed-language group is listed first, then English (1,306 items); the other 21 languages and the mixed-language group share 1699 items.
| System | mixed78 | en1306 | es95 | pt90 | pl88 | de84 | da83 | it83 | fr80 | ja80 | ko77 | sv75 | tr75 | cs75 | fi73 | uk73 | nl72 | hi72 | ar71 | id71 | zh70 | no69 | el65 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sage 1.3.0API · S+P+L1+L2+L3 | 80 | 71 | 48 | 55 | 42 | 49 | 44 | 45 | 57 | 51 | 36 | 44 | 43 | 32 | 26 | 44 | 51 | 65 | 62 | 50 | 59 | 37 | 43 |
| d1API · S+P+L1+L2+L3 | 80 | 66 | 32 | 55 | 34 | 44 | 33 | 36 | 64 | 40 | 28 | 30 | 41 | 28 | 14 | 33 | 27 | 48 | 51 | 41 | 47 | 12 | 32 |
| H2O-Lightning-4B v1.1self-hosted · S+P+L1+L2+L3 | 85 | 63 | 55 | 59 | 60 | 54 | 49 | 38 | 54 | 58 | 41 | 37 | 34 | 46 | 26 | 49 | 46 | 39 | 48 | 55 | 47 | 31 | 23 |
| Mercury DecideAPI · S+P+L1+L2+L3Cells use S+P, L1, L2 and L3. L3 has 1,416 recorded observations on its 1,418 items: 1,412 answers, three input refusals and one rate-limit (429) error, an operational non-answer rather than an answer. The two missing records (a second spent 429 kept only as a transport failure, and one request stopped before sending) are not imputed; the normal 98% recorded-completeness gate applies without exception. Besides 868 original free-route records, L3 uses 447 retained and 101 newly issued paid decisions on the currently declared Inception Mercury Decide 2026-09-30 route. That route qualified on public items only: top labels agreed on 287 of 297 (96.63%), but probabilities differ (one bit-exact vector, maximum absolute difference 0.977543). Equality with the historically measured weights and calibrator, and with the original free responses, is unproven. Headline, rank and original row price are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure. | 72 | 69 | 49 | 50 | 39 | 65 | 46 | 44 | 67 | 52 | 43 | 42 | 44 | 34 | 42 | 58 | 43 | 60 | 56 | 40 | 46 | 38 | 44 |
| decisio v0.8.0 on gemma-4-31B-itself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and all 1,505 newly completed frozen supplemental items (354 L1, 33 L2, 1,118 L3). Each supplement also carries the same original P300 answers by stable item identity; they count once in the union. The original S+P measurement has 1,477 native answers and 23 HTTP 500 errors, which remain operational non-answers and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Decisio is the original locally served Gemma offering, not an OpenAI API offering. | 86 | 75 | 62 | 65 | 47 | 64 | 61 | 57 | 74 | 67 | 57 | 51 | 61 | 51 | 53 | 65 | 61 | 69 | 65 | 55 | 56 | 48 | 49 |
| Jev 1.13.0API · S+P+L1+L2+L3 | 76 | 68 | 38 | 44 | 22 | 48 | 35 | 34 | 36 | 36 | 30 | 15 | 43 | 11 | 7 | 34 | 20 | 36 | 43 | 23 | 47 | 19 | 21 |
| Quyet-1.0-Largeself-hosted · S+P+L1+L2+L3 | 84 | 76 | 47 | 57 | 56 | 66 | 54 | 40 | 66 | 55 | 43 | 52 | 59 | 43 | 29 | 64 | 59 | 67 | 67 | 46 | 55 | 45 | 64 |
| decider-12b v2self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 87 | 67 | 51 | 58 | 61 | 40 | 49 | 55 | 61 | 49 | 54 | 43 | 55 | 46 | 43 | 57 | 59 | 67 | 58 | 56 | 40 | 58 | 36 |
| wity-1 (Wity, reasoning auto)API · S+P+L1+L2+L3 | 83 | 76 | 59 | 66 | 31 | 60 | 36 | 47 | 38 | 42 | 49 | 41 | 47 | 57 | 37 | 27 | 55 | 59 | 59 | 49 | 49 | 40 | 32 |
| wity-1 (Wity, reasoning always)API · S+P+L1+L2+L3 | 75 | 76 | 57 | 68 | 32 | 53 | 41 | 51 | 52 | 44 | 34 | 46 | 27 | 45 | 43 | 37 | 50 | 52 | 64 | 58 | 46 | 38 | 23 |
| decider-12b v1, stock Gemma-4-12B-itself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 85 | 64 | 48 | 50 | 49 | 42 | 46 | 54 | 61 | 51 | 42 | 46 | 47 | 38 | 32 | 49 | 44 | 61 | 50 | 49 | 40 | 51 | 43 |
| torchcast-decision-12bself-hosted · S+P+L1+L2+L3 | 86 | 65 | 49 | 64 | 47 | 48 | 39 | 53 | 64 | 50 | 44 | 47 | 45 | 44 | 33 | 54 | 40 | 68 | 53 | 54 | 44 | 58 | 44 |
| Microsoft-Decision-1API · S1200+P300+L1+L2+L3; full native Foundry union | 88 | 57 | 46 | 59 | 37 | 51 | 49 | 45 | 50 | 55 | 33 | 42 | 43 | 39 | 24 | 38 | 46 | 34 | 44 | 57 | 54 | 33 | 33 |
| Winnow-12B Q8self-hosted · S+P+L1+L2+L3 | 78 | 64 | 40 | 59 | 39 | 52 | 39 | 54 | 53 | 45 | 41 | 29 | 44 | 33 | 41 | 46 | 36 | 57 | 55 | 34 | 38 | 44 | 40 |
| deck-31Bself-hosted · S+P+L1+L2+L3 | 91 | 78 | 60 | 69 | 55 | 68 | 65 | 60 | 74 | 71 | 56 | 58 | 71 | 50 | 53 | 68 | 60 | 70 | 61 | 59 | 55 | 45 | 60 |
| Jeff-1.0-Largeself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Cygnetself-hosted · S+P+L1+L2+L3 | 79 | 57 | 38 | 56 | 41 | 41 | 32 | 39 | 63 | 45 | 34 | 27 | 42 | 36 | 21 | 42 | 45 | 53 | 54 | 34 | 39 | 46 | 40 |
| decisio v0.8.0 on gemma-4-12B-itself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 79 | 56 | 39 | 55 | 39 | 42 | 33 | 38 | 63 | 41 | 29 | 25 | 34 | 35 | 25 | 37 | 44 | 48 | 56 | 38 | 40 | 42 | 27 |
| decisio v0.9.0 on gemma-4-12B-itself-hosted · S+P+L1+L2+L3 | 81 | 56 | 41 | 55 | 38 | 40 | 31 | 41 | 62 | 41 | 28 | 29 | 36 | 34 | 25 | 37 | 44 | 48 | 56 | 37 | 40 | 43 | 31 |
| Jev-Omniself-hosted · S+P+L1+L2+L3 | 86 | 56 | 38 | 58 | 59 | 44 | 45 | 37 | 58 | 50 | 38 | 27 | 49 | 28 | 25 | 50 | 38 | 57 | 55 | 47 | 48 | 46 | 37 |
| Xor 26B-A4Bself-hosted · S+P+L1+L2+L3 | 80 | 63 | 46 | 40 | 48 | 56 | 51 | 36 | 52 | 44 | 34 | 45 | 43 | 32 | 33 | 50 | 40 | 68 | 50 | 50 | 50 | 42 | 46 |
| Bobcat Flash 1.2self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 82 | 72 | 45 | 47 | 38 | 54 | 33 | 33 | 51 | 41 | 43 | 37 | 39 | 28 | 26 | 47 | 37 | 48 | 50 | 44 | 58 | 29 | 27 |
| SPX-CD FlashAPI · S+P+L1+L2+L3 | 78 | 63 | 38 | 43 | 36 | 40 | 27 | 34 | 49 | 48 | 31 | 22 | 30 | 17 | 8 | 38 | 24 | 37 | 51 | 46 | 50 | 12 | 29 |
| decider chat on Gemma-4-31B-itself-hosted · S+P+L1+L2+L3 | 75 | 64 | 58 | 54 | 34 | 53 | 50 | 45 | 56 | 59 | 45 | 39 | 52 | 45 | 32 | 43 | 44 | 51 | 44 | 43 | 44 | 44 | 49 |
| Hopper 12B trainedself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 79 | 54 | 50 | 51 | 42 | 48 | 33 | 39 | 54 | 42 | 32 | 23 | 38 | 28 | 16 | 36 | 32 | 48 | 47 | 34 | 32 | 24 | 34 |
| Surogate Rune 26B-A4B v3 (v1.6 pool)self-hosted · S+P+L1+L2+L3 | 75 | 59 | 47 | 43 | 36 | 47 | 37 | 37 | 45 | 44 | 33 | 37 | 45 | 38 | 19 | 52 | 37 | 55 | 43 | 40 | 50 | 23 | 33 |
| GEV-26B-Decideself-hosted · S+P+L1+L2+L3 | 78 | 52 | 33 | 37 | 31 | 42 | 37 | 42 | 37 | 28 | 31 | 23 | 30 | 28 | 21 | 35 | 29 | 42 | 21 | 40 | 28 | 27 | 25 |
| SPX-CD-Omniself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 80 | 53 | 35 | 40 | 46 | 36 | 29 | 32 | 50 | 44 | 41 | 34 | 44 | 26 | 31 | 48 | 22 | 44 | 56 | 35 | 38 | 24 | 34 |
| Metask rain 4Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| OpenAI DecisionsAPI · A5+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3. | 76 | 59 | 36 | 48 | 41 | 55 | 27 | 37 | 43 | 37 | 24 | 23 | 46 | 37 | 30 | 33 | 27 | 40 | 33 | 39 | 37 | 40 | 25 |
| InstinctAPI · A4+P+L1+L2+L3 | 67 | 58 | 22 | 35 | 32 | 28 | 27 | 31 | 52 | 48 | 20 | 18 | 34 | 19 | 23 | 22 | 39 | 39 | 39 | 42 | 43 | 13 | 7 |
| Blink v0.3 26B-A4Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Diffusion Jevself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 82 | 59 | 51 | 57 | 49 | 65 | 37 | 41 | 52 | 44 | 48 | 29 | 43 | 45 | 36 | 49 | 44 | 63 | 50 | 50 | 51 | 38 | 16 |
| René-1 31B FP8self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 75 | 67 | 45 | 39 | 42 | 46 | 40 | 31 | 55 | 37 | 35 | 37 | 54 | 13 | 25 | 44 | 33 | 47 | 43 | 42 | 46 | 20 | 43 |
| Aplomb 1self-hosted · S+PThe supplement run could not reproduce the measured model: the pinned Hugging Face revision of its code and weights was deleted upstream (the repository history was rewritten), so the same files cannot be fetched. The radar shows the existing S + P items. | — | 44 | 84 | 50† | 42 | 68† | 13† | 52† | 72† | 26† | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Wald 4B v2.1self-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Bespoke Nimble 9B v3self-hosted · S+P+L1+L2+L3 | 80 | 61 | 45 | 50 | 35 | 44 | 30 | 30 | 56 | 41 | 27 | 7 | 40 | 32 | 17 | 36 | 23 | 32 | 42 | 60 | 54 | 20 | 37 |
| wity-1 (Wity, reasoning off)API · S+P+L1+L2+L3 | 67 | 52 | 38 | 48 | 29 | 36 | 21 | 26 | 10 | 13 | 17 | 14 | 17 | 17 | 6 | 11 | 24 | 30 | 31 | 33 | 33 | 26 | 1 |
| APUS-OpenJev-v1-9Bself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 74 | 48 | 33 | 31 | 25 | 37 | 18 | 22 | 37 | 41 | 8 | 0 | 20 | 25 | 6 | 16 | 28 | 29 | 30 | 34 | 35 | 17 | 12 |
| Plumb-4Bself-hosted · S+P+L1+L2+L3 | 72 | 45 | 35 | 45 | 36 | 30 | 38 | 27 | 32 | 34 | 25 | 8 | 32 | 40 | 14 | 24 | 31 | 25 | 28 | 27 | 39 | 22 | 20 |
| SPX-CD ProAPI · S+P+L1+L2+L3 | 82 | 70 | 44 | 55 | 37 | 44 | 53 | 43 | 62 | 50 | 35 | 47 | 41 | 19 | 22 | 47 | 37 | 48 | 54 | 29 | 52 | 29 | 47 |
| Messier One v0.2self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and all 1,505 newly completed frozen supplemental items (354 L1, 33 L2, 1,118 L3), answered natively on the original Messier image with the network disabled. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. All 1,505 supplemental requests succeeded. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 76 | 45 | 39 | 45 | 35 | 37 | 27 | 27 | 31 | 58 | 34 | 14 | 26 | 25 | 15 | 23 | 33 | 32 | 27 | 47 | 33 | 28 | 20 |
| spark-s1-4b-v6self-hosted · S+P+L1+L2+L3 | 80 | 46 | 33 | 31 | 18 | 28 | 9 | 17 | 37 | 35 | 22 | 0 | 22 | 23 | 0 | 15 | 29 | 20 | 23 | 10 | 26 | 2 | 7 |
| swanOneself-hosted · S+PThe pinned model weights cannot currently be accessed (Hugging Face returned HTTP 401). Supplement measurements remain pending access recovery; the radar shows the existing S + P items. | — | 57 | 86 | 68† | 48 | 85† | 26† | 45† | 91† | 42† | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Quyet-1.0-Mediumself-hosted · S+P+L1+L2+L3 | 60 | 44 | 33 | 32 | 25 | 42 | 25 | 15 | 35 | 32 | 23 | 8 | 15 | 20 | 0 | 29 | 33 | 35 | 19 | 21 | 35 | 12 | 14 |
| Vansa-3.4API · A4+P+L1+L2+L3 | 54 | 47 | 12 | 42 | 28 | 17 | 12 | 10 | 24 | 43 | 11 | 0 | 13 | 19 | 12 | 9 | 17 | 26 | 15 | 33 | 35 | 0 | 11 |
| levself-hosted · S+P+L1+L2+L3 | 55 | 41 | 28 | 20 | 22 | 22 | 12 | 22 | 28 | 22 | 29 | 8 | 22 | 11 | 2 | 14 | 17 | 28 | 19 | 24 | 37 | 5 | 37 |
| NInfer Qwen3.8-Flash-Next mixedself-hosted · S+P+L1+L2+L3 | 72 | 53 | 41 | 50 | 24 | 33 | 30 | 33 | 55 | 33 | 28 | 16 | 43 | 20 | 27 | 36 | 35 | 54 | 31 | 38 | 43 | 25 | 20 |
| jev-localself-hosted · S+P+L1+L2+L3 | 71 | 46 | 43 | 31 | 21 | 31 | 21 | 22 | 33 | 32 | 17 | 7 | 20 | 12 | 7 | 21 | 29 | 10 | 34 | 25 | 29 | 11 | 5 |
| decider-4b v2self-hosted · S+P+L1+L2+L3 | 62 | 42 | 41 | 25 | 11 | 28 | 8 | 22 | 21 | 32 | 32 | 0 | 18 | 13 | 1 | 8 | 18 | 32 | 22 | 24 | 29 | 1 | 1 |
| GPT-6 Luna (low reasoning effort)API · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3. | 97 | 97 | 91 | 93 | 91 | 94 | 100 | 94 | 100 | 99 | 98 | 96 | 96 | 93 | 96 | 94 | 97 | 97 | 100 | 98 | 98 | 99 | 97 |
| Autoloops – Gemma 4 31B ITAPI · A4+P+L1+L2+L3 | 73 | 67 | 44 | 50 | 44 | 56 | 61 | 42 | 61 | 60 | 43 | 48 | 48 | 48 | 47 | 60 | 50 | 55 | 49 | 45 | 53 | 41 | 51 |
| Wald 4Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT-6 Luna (default medium reasoning effort)API · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3. | 97 | 97 | 98 | 96 | 91 | 98 | 100 | 99 | 99 | 97 | 98 | 100 | 96 | 95 | 98 | 96 | 97 | 99 | 95 | 97 | 100 | 99 | 96 |
| ClassOne Qwen 3.5 9Bself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 71 | 45 | 38 | 51 | 26 | 36 | 23 | 19 | 30 | 43 | 23 | 0 | 17 | 8 | 1 | 4 | 11 | 30 | 17 | 18 | 40 | 13 | 31 |
| JevK5 v0.3self-hosted · S+P+L1+L2+L3 | 67 | 43 | 40 | 42 | 29 | 19 | 20 | 13 | 28 | 30 | 18 | 0 | 28 | 27 | 8 | 16 | 28 | 17 | 22 | 34 | 29 | 3 | 18 |
| metask-jev-4bself-hosted · S+P+L1+L2+L3 | 73 | 40 | 30 | 34 | 15 | 21 | 16 | 19 | 28 | 24 | 12 | 9 | 15 | 4 | 0 | 9 | 11 | 23 | 16 | 12 | 26 | 4 | 0 |
| deck-4B v1.0self-hosted · S+P+L1+L2+L3 | 67 | 43 | 37 | 45 | 32 | 22 | 21 | 19 | 29 | 30 | 22 | 0 | 31 | 26 | 5 | 13 | 23 | 18 | 27 | 36 | 29 | 19 | 17 |
| Clef-Flashself-hosted · S+P+L1+L2+L3 | 69 | 44 | 40 | 31 | 8 | 34 | 9 | 19 | 41 | 29 | 23 | 18 | 10 | 2 | 0 | 11 | 12 | 23 | 23 | 23 | 36 | 13 | 11 |
| Decision-4Bself-hosted · S+P+L1+L2+L3 | 61 | 43 | 27 | 19 | 0 | 31 | 0 | 7 | 20 | 22 | 26 | 2 | 6 | 10 | 0 | 0 | 21 | 17 | 13 | 15 | 25 | 0 | 0 |
| janus 4Bself-hosted · S+P+L1+L2+L3 | 63 | 40 | 35 | 23 | 15 | 19 | 9 | 15 | 21 | 20 | 16 | 2 | 25 | 9 | 0 | 0 | 11 | 29 | 17 | 15 | 20 | 10 | 7 |
| Qwen3.5-9B Jev-like data-mix v2self-hosted · S+P+L1+L2+L3 | 71 | 46 | 29 | 24 | 18 | 37 | 8 | 13 | 35 | 31 | 12 | 0 | 25 | 3 | 0 | 9 | 18 | 30 | 38 | 24 | 24 | 3 | 21 |
| decisor-4bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| JevK5 v0.2.0self-hosted · S+P+L1+L2+L3 | 65 | 36 | 26 | 36 | 21 | 21 | 23 | 15 | 24 | 20 | 6 | 0 | 10 | 13 | 7 | 4 | 24 | 21 | 19 | 18 | 26 | 13 | 2 |
| JevOneself-hosted · S+P+L1+L2+L3 | 73 | 49 | 37 | 35 | 22 | 33 | 10 | 21 | 32 | 31 | 22 | 10 | 20 | 10 | 9 | 20 | 24 | 30 | 40 | 18 | 32 | 1 | 0 |
| Kahn1 4Bself-hosted · S+P+L1+L2+L3 | 67 | 40 | 36 | 46 | 28 | 35 | 25 | 30 | 29 | 45 | 25 | 5 | 35 | 25 | 21 | 14 | 33 | 19 | 28 | 41 | 30 | 20 | 20 |
| Fastino GLiDEAPI · A4+P+L1+L2+L3 | 69 | 73 | 40 | 53 | 29 | 42 | 43 | 37 | 71 | 52 | 43 | 33 | 43 | 22 | 33 | 37 | 56 | 53 | 47 | 50 | 47 | 35 | 40 |
| TypeCastLM 1.4.0self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 57 | 38 | 24 | 25 | 8 | 22 | 16 | 17 | 28 | 18 | 12 | 0 | 17 | 9 | 0 | 5 | 17 | 9 | 6 | 15 | 19 | 13 | 3 |
| Hopperself-hosted · S+P+L1+L2+L3 | 56 | 36 | 31 | 22 | 15 | 22 | 0 | 15 | 16 | 12 | 13 | 0 | 14 | 12 | 3 | 0 | 12 | 15 | 17 | 15 | 11 | 0 | 0 |
| Imajev-4Bself-hosted · S+P+L1+L2+L3 | 51 | 37 | 30 | 27 | 9 | 24 | 0 | 19 | 12 | 10 | 14 | 0 | 9 | 0 | 0 | 0 | 8 | 10 | 13 | 5 | 29 | 1 | 0 |
| SimpleJev Qwen3.6-35B-A3BAPI · A4+P+L1+L2+L3 | 65 | 63 | 37 | 42 | 26 | 43 | 30 | 41 | 53 | 54 | 24 | 34 | 44 | 17 | 22 | 21 | 34 | 30 | 20 | 34 | 42 | 25 | 20 |
| Malkuth-4Bself-hosted · S+P+L1+L2+L3 | 49 | 37 | 21 | 17 | 3 | 11 | 7 | 17 | 7 | 5 | 4 | 0 | 7 | 0 | 0 | 0 | 0 | 0 | 4 | 2 | 20 | 0 | 0 |
| jqvself-hosted · S+P+L1+L2+L3 | 70 | 39 | 16 | 24 | 18 | 29 | 8 | 17 | 15 | 17 | 21 | 0 | 10 | 5 | 0 | 18 | 7 | 27 | 16 | 25 | 29 | 0 | 3 |
| JPT-35B-A3Bself-hosted · S+P+L1+L2+L3 | 86 | 56 | 39 | 48 | 32 | 45 | 34 | 28 | 43 | 47 | 32 | 28 | 44 | 17 | 6 | 39 | 34 | 28 | 46 | 33 | 45 | 17 | 20 |
| decider-35b-a3bself-hosted · S+P+L1+L2+L3 | 77 | 52 | 31 | 29 | 16 | 45 | 23 | 23 | 37 | 24 | 29 | 16 | 44 | 9 | 5 | 24 | 22 | 41 | 36 | 24 | 38 | 14 | 28 |
| CoCo-Decision-4Bself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 51 | 36 | 26 | 23 | 21 | 17 | 18 | 22 | 28 | 17 | 16 | 0 | 14 | 9 | 0 | 2 | 7 | 24 | 16 | 20 | 25 | 12 | 1 |
| Manchego v2.1self-hosted · S+P+L1+L2+L3 | 50 | 33 | 25 | 11 | 0 | 15 | 0 | 19 | 12 | 8 | 17 | 0 | 2 | 0 | 8 | 0 | 21 | 21 | 9 | 6 | 4 | 7 | 3 |
| Decision 4B v1.2self-hosted · S+P+L1+L2+L3 | 65 | 36 | 36 | 32 | 16 | 25 | 16 | 22 | 27 | 34 | 21 | 0 | 24 | 32 | 22 | 27 | 28 | 20 | 35 | 34 | 25 | 12 | 14 |
| Raw Qwen3 4B Instruct 2507 direct logitsself-hosted · S+P+L1+L2+L3 | 59 | 34 | 17 | 31 | 5 | 32 | 0 | 17 | 24 | 4 | 17 | 0 | 10 | 3 | 19 | 2 | 14 | 25 | 19 | 21 | 31 | 0 | 4 |
| RSI-Jev v6.1-VL 27Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| GPT-5.6 LunaAPI · A4+P+L1+L2+L3L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of this row carry that exposure; headline scores do not use L3. | 96 | 93 | 86 | 92 | 84 | 86 | 92 | 90 | 87 | 94 | 90 | 79 | 93 | 97 | 94 | 89 | 87 | 84 | 85 | 91 | 89 | 82 | 91 |
| JEV-27Bself-hosted · S+P+L1+L2+L3 | 71 | 62 | 36 | 42 | 23 | 41 | 36 | 37 | 54 | 39 | 33 | 37 | 43 | 9 | 19 | 47 | 37 | 45 | 47 | 37 | 43 | 26 | 48 |
| torchcast-decision-27bself-hosted · S+P+L1+L2+L3 | 81 | 76 | 49 | 58 | 51 | 45 | 43 | 40 | 67 | 52 | 42 | 53 | 47 | 23 | 21 | 50 | 38 | 57 | 61 | 39 | 51 | 35 | 44 |
| Bespoke Nimble 9Bself-hosted · S+P+L1+L2+L3 | 66 | 45 | 41 | 36 | 35 | 41 | 10 | 21 | 37 | 26 | 19 | 4 | 23 | 9 | 12 | 6 | 25 | 26 | 39 | 40 | 32 | 14 | 11 |
| reflex 4Bself-hosted · S+P+L1+L2+L3 | 56 | 33 | 33 | 24 | 4 | 13 | 1 | 13 | 15 | 12 | 11 | 0 | 6 | 11 | 17 | 0 | 17 | 17 | 29 | 4 | 18 | 0 | 4 |
| Perplexity Decider v1.1 27Bself-hosted · S+P+L1+L2+L3 | 85 | 79 | 50 | 58 | 62 | 63 | 43 | 52 | 70 | 51 | 48 | 42 | 57 | 21 | 38 | 63 | 33 | 67 | 64 | 42 | 64 | 39 | 57 |
| JADEself-hosted · S+P+L1+L2+L3 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| typecastlmself-hosted · S+P+L1+L2+L3 | 54 | 29 | 30 | 19 | 6 | 24 | 6 | 11 | 24 | 23 | 13 | 0 | 5 | 7 | 1 | 0 | 8 | 15 | 14 | 8 | 20 | 13 | 3 |
| Jebadiah 27Bself-hosted · S+P+L1+L2+L3 | 65 | 64 | 36 | 47 | 30 | 34 | 22 | 39 | 51 | 41 | 29 | 18 | 37 | 18 | 11 | 31 | 33 | 49 | 46 | 29 | 44 | 17 | 26 |
| Decision 2.0 Vega 27Bself-hosted · S+P+L1+L2+L3 | 68 | 64 | 39 | 49 | 23 | 36 | 26 | 30 | 35 | 36 | 41 | 30 | 44 | 18 | 22 | 29 | 30 | 46 | 51 | 21 | 43 | 21 | 40 |
| Standard One 8B SHself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Decision 4B v1.1self-hosted · S+P+L1+L2+L3 | 59 | 33 | 29 | 39 | 12 | 24 | 10 | 21 | 24 | 41 | 17 | 0 | 15 | 19 | 11 | 18 | 22 | 7 | 27 | 28 | 22 | 5 | 5 |
| AutoJev-27B (denis-pplx, Qwen3.8-27B)self-hosted · S+P+L1+L2+L3 | 74 | 59 | 41 | 37 | 27 | 35 | 33 | 32 | 51 | 38 | 30 | 24 | 42 | 16 | 18 | 27 | 36 | 47 | 48 | 32 | 45 | 25 | 27 |
| Eikos-27Bself-hosted · S+P+L1+L2+L3 | 75 | 68 | 45 | 54 | 40 | 46 | 56 | 44 | 50 | 58 | 38 | 41 | 46 | 26 | 30 | 41 | 39 | 56 | 53 | 43 | 50 | 42 | 49 |
| Instinct Dual 4BAPI · A4+P+L1+L2+L3 | 36 | 32 | 12 | 5 | 0 | 0 | 0 | 2 | 7 | 13 | 2 | 0 | 0 | 0 | 0 | 0 | 11 | 6 | 12 | 0 | 15 | 0 | 0 |
| NInfer Qwen3.8-27B NVFP4self-hosted · S+P+L1+L2+L3 | 72 | 55 | 45 | 51 | 35 | 40 | 32 | 23 | 41 | 42 | 33 | 22 | 38 | 22 | 17 | 18 | 37 | 58 | 37 | 26 | 43 | 17 | 8 |
| OpenJev (thinking, BF16)self-hosted · S+P+L1+L2+L3 | 87 | 78 | 59 | 57 | 31 | 55 | 36 | 60 | 64 | 69 | 53 | 50 | 65 | 43 | 53 | 60 | 46 | 53 | 62 | 51 | 56 | 41 | 42 |
| Open-Jev 9Bself-hosted · S+P+L1+L2+L3 | 70 | 47 | 40 | 48 | 23 | 44 | 26 | 21 | 26 | 35 | 18 | 0 | 24 | 5 | 14 | 28 | 19 | 35 | 15 | 30 | 38 | 14 | 17 |
| Gemini 3.1 Flash-LiteAPI · A4+P+L1+L2+L3 | 84 | 68 | 42 | 64 | 57 | 53 | 58 | 46 | 63 | 58 | 42 | 45 | 55 | 51 | 27 | 43 | 40 | 49 | 35 | 54 | 54 | 49 | 44 |
| Standard One 8Bself-hosted · S+P+L1+L2+L3 | 62 | 38 | 22 | 17 | 20 | 17 | 1 | 24 | 26 | 19 | 7 | 0 | 12 | 3 | 0 | 4 | 0 | 23 | 14 | 23 | 34 | 6 | 10 |
| Clefself-hosted · S+P+L1+L2+L3 | 82 | 65 | 50 | 49 | 27 | 49 | 38 | 46 | 55 | 45 | 33 | 43 | 43 | 21 | 30 | 31 | 36 | 52 | 44 | 37 | 49 | 23 | 36 |
| system-one-openAPI · A4+P+L1+L2+L3 | 23 | 31 | 8 | 14 | 9 | 11 | 13 | 16 | 7 | 16 | 6 | 0 | 0 | 0 | 0 | 0 | 19 | 10 | 0 | 2 | 25 | 0 | 2 |
| JEV Qwen3.5-9B Base NVFP4self-hosted · S+P+L1+L2+L3 | 40 | 27 | 24 | 12 | 6 | 16 | 1 | 1 | 21 | 17 | 0 | 0 | 7 | 0 | 0 | 0 | 2 | 5 | 13 | 0 | 7 | 0 | 5 |
| SemIf, formerly OpenJevself-hosted · S+P+L1+L2+L3 | 58 | 30 | 15 | 16 | 2 | 22 | 5 | 10 | 14 | 15 | 9 | 0 | 19 | 0 | 0 | 0 | 11 | 15 | 7 | 11 | 20 | 0 | 0 |
| openjev-sglangAPI · A4+P+L1+L2+L3 | 60 | 47 | 20 | 29 | 19 | 28 | 0 | 19 | 24 | 42 | 19 | 12 | 12 | 13 | 9 | 13 | 32 | 24 | 14 | 37 | 28 | 20 | 0 |
| Bev / Bonsai 27Bself-hosted · S+P+L1+L2+L3 | 67 | 46 | 26 | 35 | 19 | 30 | 14 | 17 | 25 | 22 | 16 | 0 | 9 | 0 | 0 | 0 | 24 | 28 | 6 | 21 | 27 | 4 | 6 |
| LitJevself-hosted · S+P+L1+L2+L3 | 69 | 49 | 34 | 39 | 27 | 36 | 30 | 23 | 34 | 34 | 28 | 0 | 24 | 22 | 10 | 10 | 32 | 39 | 40 | 29 | 30 | 13 | 1 |
| local-jev Qwen3.5-4Bself-hosted · S+P+L1+L2+L3 | 35 | 15 | 11 | 11 | 0 | 5 | 0 | 2 | 8 | 0 | 3 | 0 | 0 | 0 | 0 | 0 | 7 | 8 | 0 | 0 | 11 | 0 | 0 |
| system-oneself-hosted · S+P+L1+L2+L3 | 51 | 31 | 6 | 27 | 4 | 36 | 7 | 17 | 19 | 6 | 0 | 0 | 7 | 0 | 0 | 10 | 16 | 18 | 0 | 11 | 3 | 0 | 0 |
| kev 8Bself-hosted · S+P+L1+L2+L3 | 51 | 35 | 24 | 26 | 9 | 22 | 0 | 8 | 7 | 0 | 7 | 0 | 0 | 4 | 0 | 0 | 0 | 2 | 0 | 9 | 27 | 0 | 0 |
| Decision 2Bself-hosted · S+P+L1+L2+L3 | 20 | 14 | 7 | 0 | 0 | 10 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 3 | 0 | 0 | 0 | 0 | 0 | 0 |
| OpenSourceJevself-hosted · S+P+L1+L2+L3 | 33 | 20 | 18 | 5 | 2 | 11 | 0 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10 | 0 | 0 |
| open-alternative-jevself-hosted · S+P+L1+L2+L3 | 43 | 24 | 20 | 3 | 0 | 14 | 3 | 7 | 3 | 9 | 11 | 0 | 0 | 3 | 0 | 0 | 7 | 2 | 7 | 4 | 4 | 0 | 1 |
| Raw Qwen3 8B direct logitsself-hosted · S+P+L1+L2+L3 | 54 | 30 | 7 | 15 | 0 | 30 | 0 | 11 | 17 | 7 | 18 | 8 | 16 | 0 | 0 | 0 | 1 | 22 | 7 | 5 | 14 | 0 | 0 |
| Liquid AI d1-3Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| decider-2bself-hosted · S+P+L1+L2+L3 | 42 | 25 | 7 | 7 | 0 | 14 | 2 | 4 | 3 | 2 | 8 | 0 | 23 | 0 | 0 | 5 | 8 | 12 | 0 | 0 | 7 | 0 | 0 |
| Nemotron Diffusion 8Bself-hosted · S+P+L1+L2+L3 | 30 | 16 | 0 | 0 | 16 | 19 | 0 | 0 | 6 | 0 | 0 | 0 | 4 | 0 | 12 | 0 | 0 | 18 | 0 | 0 | 3 | 0 | 0 |
| DeepSeek V4.1 FlashAPI · A4+P+L1+L2+L3 | 97 | 95 | 94 | 99 | 92 | 98 | 99 | 100 | 100 | 99 | 98 | 96 | 100 | 98 | 100 | 96 | 98 | 100 | 100 | 99 | 100 | 98 | 100 |
| kev 4Bself-hosted · S+P+L1+L2+L3 | 44 | 25 | 28 | 11 | 0 | 14 | 0 | 5 | 19 | 0 | 7 | 0 | 16 | 0 | 0 | 0 | 0 | 13 | 0 | 0 | 16 | 0 | 0 |
| Clef-omniself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Kev 27Bself-hosted · S+P+L1+L2+L3 | 77 | 69 | 47 | 50 | 38 | 42 | 42 | 44 | 64 | 44 | 41 | 45 | 44 | 20 | 26 | 50 | 31 | 51 | 55 | 35 | 48 | 22 | 33 |
| Malkuth-2Bself-hosted · S+P+L1+L2+L3 | 35 | 14 | 2 | 0 | 2 | 6 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4 | 0 | 0 |
| Qwen3-Reranker-4Bself-hosted · S+P+L1+L2+L3 | 4 | 0 | 0 | 0 | 0 | 3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Vega 4Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool)self-hosted · S+P+L1+L2+L3 | 80 | 62 | 44 | 52 | 34 | 41 | 43 | 29 | 61 | 41 | 33 | 24 | 39 | 25 | 29 | 32 | 41 | 49 | 50 | 31 | 50 | 29 | 26 |
| ZeroEntropy zerank-2self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| EXAONE-4.0-1.2B-JEV v0.3self-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 5 | 6 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Open-Jev 2Bself-hosted · S+P+L1+L2+L3 | 39 | 18 | 11 | 1 | 6 | 17 | 10 | 0 | 9 | 7 | 12 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 3 | 6 | 0 | 0 |
| SimpleJev Qwen3.8-27BAPI · A4+P+L1+L2+L3 | 70 | 63 | 28 | 40 | 38 | 43 | 33 | 22 | 54 | 43 | 23 | 25 | 34 | 17 | 31 | 21 | 44 | 43 | 37 | 31 | 55 | 29 | 27 |
| Open-Jev 27B v1.1self-hosted · S+P+L1+L2+L3 | 64 | 55 | 44 | 41 | 22 | 40 | 24 | 21 | 44 | 33 | 28 | 37 | 38 | 8 | 11 | 23 | 11 | 42 | 36 | 34 | 37 | 24 | 23 |
| Raw Phi-4 mini direct logitsself-hosted · S+P+L1+L2+L3 | 13 | 2 | 6 | 5 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 7 | 0 | 0 | 0 | 0 | 10 | 0 | 0 | 15 | 0 | 2 |
| decision-machine-1API · A4+P+L1+L2+L3 | 1 | 0 | 4 | 0 | 0 | 0 | 0 | 0 | 10 | 0 | 0 | 0 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 4 | 0 | 0 | 8 |
| SimpleJevself-hosted · S+P+L1+L2+L3 | 18 | 9 | 14 | 5 | 18 | 9 | 0 | 0 | 10 | 2 | 2 | 3 | 1 | 0 | 0 | 4 | 0 | 0 | 0 | 2 | 12 | 8 | 27 |
| JevActAPI · A4+P+L1+L2+L3 | 12 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10 | 0 | 14 | 0 | 0 |
| smalljev semantic-v9self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Gutsy 0.8B v0.3self-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| kev 0.6Bself-hosted · S+P+L1+L2+L3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| jul fastself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| OpenDecisionself-hosted · S+P+L1+L2+L3 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| ClassOne Gemma 4 E2Bself-hosted · S+PA supplement run was made with the same weights and serving code, but this model's server gives different answers after each restart (verified on identical weights), so the new run does not reproduce the row's original answers (87.7 % agreement on the 300 public items, below the 95 % bar). Its language and category supplements are not published; the radar shows the existing S + P items. | — | 0 | 0 | 17† | 15 | 21† | 0† | 0† | 2† | 0† | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Qwen3.5-0.8B Decision Modelself-hosted · S+P+L1+L2+L3 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Fastino GLiNER-2.5-DecideAPI · S+P+L1+L2+L3+L4 | 11 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 8 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Deem 0.8B v1self-hosted · S+P+L1+L2+L3 | 5 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 8 | 0 | 0 |
| Decision Fastself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Quyet-1.0-Small-ENself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Laya typed-decisionsself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| WaterSheepself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| openJev Verdict 1.4self-hosted · S+P+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Raw Qwen3 0.6B direct logitsself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10 | 0 | 0 | 0 | 0 | 0 | 0 | 13 | 0 | 4 | 0 | 0 | 8 | 0 | 0 |
| lev-350mself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Raw Qwen3 1.7B direct logitsself-hosted · S+P+L1+L2+L3 | 10 | 0 | 0 | 6 | 0 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| kev 0.5Bself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| openJev Verdictself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Quyet-1.0-Smallself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| watt-flash-0.1self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Quyet-1.0-Tinyself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| Liquid AI d1-omni-600Mself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Vega 0.8Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| open-jev-deberta-v3-largeself-hosted · S+P+L1+L2+L3+L5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| verdict-smallself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| Tacet Sonataself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Layaself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Mixedbread mxbai-rerank-base-v2self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Laya multilingualself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 5 | 1 | 4 | 12 | 3 | 0 | 0 | 0 | 0 |
| Certo v1self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| CLM-8Bself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| BAAI bge-reranker-v2-m3self-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Alibaba GTE Reranker ModernBERT-baseself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Mirrorself-hosted · S+P+L1+L2+L3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Open Jev JSON Canvasself-hosted · S+P+L1+L2+L3 | 76 | 65 | 55 | 42 | 47 | 56 | 37 | 55 | 60 | 55 | 51 | 40 | 34 | 46 | 44 | 33 | 41 | 50 | 47 | 54 | 51 | 49 | 53 |
| Qwen3.8 27BAPI · A4+P+L1+L2+L3 | 100 | 95 | 95 | 92 | 93 | 96 | 99 | 97 | 100 | 98 | 98 | 99 | 100 | 100 | 100 | 96 | 96 | 100 | 99 | 93 | 100 | 97 | 100 |
| APUS-OpenJev-v1-35B-A3Bself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 82 | 54 | 37 | 40 | 47 | 40 | 27 | 25 | 41 | 47 | 34 | 14 | 23 | 33 | 12 | 31 | 32 | 37 | 39 | 35 | 37 | 17 | 10 |
| BB-Qwen3.5-4B-LoRAAPI · S+P+L1+L2+L3 | 53 | 41 | 27 | 19 | 20 | 15 | 15 | 16 | 20 | 13 | 19 | 0 | 17 | 0 | 1 | 0 | 18 | 15 | 10 | 22 | 24 | 1 | 0 |
| Seb-9Bself-hosted · S+P+L1+L2+L3Language and category cells use the unchanged original S+P observations and the newly completed frozen supplemental items (L1, L2, L3), answered with the original model and serving setup on an isolated pod. Each supplement carries the same original P300 answers by stable item identity; they count once in the union. Failed requests stay scored as failures and are not imputed. Headline, rank, Composite and original pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 73 | 49 | 33 | 26 | 12 | 29 | 6 | 17 | 27 | 23 | 24 | 0 | 30 | 10 | 0 | 8 | 22 | 34 | 30 | 12 | 35 | 11 | 0 |
| AutoJev-27B (RTX PRO 6000)self-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 74 | 60 | 42 | 38 | 27 | 35 | 33 | 32 | 51 | 40 | 30 | 22 | 40 | 15 | 18 | 28 | 37 | 46 | 48 | 32 | 45 | 25 | 27 |
| Bosun v3.1 0.6Bself-hosted · S81+P299+L1+L2+L3Historical headline · measured language cells | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 7 | 0 | 0 |
| GLiNER2self-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. About 1,000 issued items were lost to out-of-memory kills and a machine shutdown; they stay counted as spent and are never re-asked, so 17 of 23 languages rest on 50 to 59 answered items instead of the 60-item target. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 11 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 14 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 12 | 0 | 0 | 10 |
| GLiNER2 largeself-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 7 | 9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 9 | 7 | 0 | 0 | 0 | 10 | 0 | 8 | 8 | 0 | 0 |
| GLiNER2.5 multiself-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4 | 0 | 0 |
| GLiNER2.5 smallself-hosted · P+L1+L2+L3Historical headline · measured language cells | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| jeffself-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Jobe Qwen3.5-4Bself-hosted · S+P+L1+L2+L3Catalogue entry · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 58 | 30 | 14 | 16 | 1 | 25 | 4 | 9 | 14 | 15 | 8 | 0 | 19 | 0 | 0 | 0 | 11 | 15 | 6 | 13 | 20 | 0 | 0 |
| mica-v01-4bself-hosted · S+P+L1+L2+L3Catalogue entry · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 55 | 38 | 32 | 22 | 3 | 26 | 0 | 22 | 18 | 26 | 4 | 1 | 25 | 4 | 4 | 13 | 9 | 22 | 30 | 12 | 21 | 0 | 0 |
| Needle 3self-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from new runs of the supplements L1, L2 and L3 (all items) on the original pinned closed engine, approved by the benchmark owner, on an isolated machine, plus 281 S+P items: 153 from an earlier partial run and 128 never-asked items chosen by their use-case label ('other') to complete that radar spoke (these items also count in the other cells, so the S+P part is not a random sample); the rest of S+P was not run. The engine abstains on about 30 % of items and returns invalid UTF-8 on about 7 % (mostly non-Latin scripts); both count as unanswered. Its chance-corrected competence is below zero in every cell, which the table shows as 0. Its answers to the same public items differ between machines for about 10 % of items. 22 of 23 languages rest on 19 to 59 answered items instead of the 60-item target; the radars are complete. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0† | 0 | 0 | 0 | 0 | 0† | 0 | 0† | 0 | 0 | 0† | 0 | 0† |
| Needle 3, options as toolsself-hosted · S+P+L1+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells combine an earlier partial run on the current pools (175 S+P items) with a new run of the supplements L3 (1,382 of 1,418 items) and L1 (329 of 654) on the original pinned closed engine, approved by the benchmark owner, on an isolated machine; L2 and the rest of S+P were not run before the machine's time limit. The engine abstains on about 9 % of items and returns invalid UTF-8 on about 7 % (mostly non-Latin scripts); both count as unanswered. Its chance-corrected competence is below zero in every cell, which the table shows as 0. Its answers to the same public items differ between machines for about 10 % of items. 18 of 23 languages rest on 24 to 59 answered items instead of the 60-item target; the radars are complete. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0† | 0 | 0 | 0 | 0 | 0 | 3 | 0† |
| NInfer Qwen3.8-27B NVFP4 (T=1.5)self-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 71 | 50 | 39 | 51 | 24 | 32 | 29 | 20 | 38 | 37 | 27 | 15 | 37 | 19 | 18 | 18 | 33 | 51 | 26 | 25 | 38 | 14 | 5 |
| OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)self-hosted · S+P+L1+L2+L3Catalogue entry · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. This model samples stochastically: on the 300 public items its answers in the supplement passes agree with its S+P pass on 86-89 %; those public items count once, from the S+P pass. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 60 | 54 | 46 | 41 | 49 | 46 | 29 | 37 | 42 | 44 | 37 | 30 | 30 | 23 | 27 | 31 | 28 | 33 | 27 | 35 | 31 | 33 | 34 |
| reflex-27bself-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 68 | 56 | 34 | 38 | 27 | 36 | 32 | 20 | 47 | 49 | 29 | 20 | 28 | 23 | 11 | 18 | 31 | 37 | 37 | 40 | 30 | 13 | 5 |
| Surogate Rune 26B-A4B v3 (RTX PRO 6000)self-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells combine its S+P run on the current pools (6 Oct 2026) with new L1, L2 and L3 runs on the same pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 75 | 59 | 47 | 43 | 36 | 47 | 37 | 37 | 45 | 44 | 33 | 37 | 45 | 38 | 19 | 52 | 37 | 55 | 43 | 40 | 50 | 23 | 33 |
| Vonself-hosted · S+P+L1+L2+L3Historical headline · measured language cellsThis historical row keeps its published v1.5 headline. Its language and category cells come from a new measurement on the current item pools (S+P, L1, L2, L3), run with the original pinned model and serving setup on an isolated machine; each item was sent once and failed requests stay scored as failures. Headline, rank, Composite and pricing are unchanged. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Wrappers (listed, never ranked) | |||||||||||||||||||||||
| classifier.devAPI · A4+P+L1+L2+L3 | 73 | 72 | 23 | 35 | 36 | 43 | 34 | 28 | 33 | 41 | 23 | 15 | 35 | 9 | 18 | 32 | 25 | 44 | 44 | 17 | 51 | 14 | 34 |
| metask-jev-rain-12Bself-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| ryotide_qwen9self-hosted · not measured | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Intelligence gate and Noul decisiveness
The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.
| System | Intelligence | Choice | Noul | Score | Noul decisive rate | Accuracy among decisive |
|---|---|---|---|---|---|---|
| Sage 1.3.0 | 65.6 | native | native | native | 83 % | 95 % |
| d1 | 62.1 | native | native | native | 83 % | 93 % |
| H2O-Lightning-4B v1.1 | 60.0 | native | native | native | 100 % | 85 % |
| Mercury Decide | 65.9 | native | native | native | 90 % | 90 % |
| decisio v0.8.0 on gemma-4-31B-it | 70.4 | native | native | native | 94 % | 96 % |
| Jev 1.13.0 | 63.6 | native | native | native | 79 % | 98 % |
| Quyet-1.0-Large | 73.4 | native | native | native | 91 % | 93 % |
| decider-12b v2 | 63.0 | native | native | native | 99 % | 83 % |
| wity-1 | 70.3 | native | native | native | 91 % | 94 % |
| decider-12b v1, stock Gemma-4-12B-it | 60.6 | native | native | native | 96 % | 87 % |
| torchcast-decision-12b | 60.5 | native | native | native | 96 % | 86 % |
| Microsoft-Decision-1 | 57.2 | native | native | native | 77 % | 93 % |
| Winnow-12B Q8 | 59.5 | native | native | native | 86 % | 92 % |
| deck-31B | 73.0 | native | native | native | 96 % | 94 % |
| Jeff-1.0-Large | 71.9 | native | native | native | not supported | — |
| Cygnet | 54.8 | native | native | native | 81 % | 92 % |
| decisio v0.8.0 on gemma-4-12B-it | 52.3 | native | native | native | 82 % | 91 % |
| decisio v0.9.0 on gemma-4-12B-it | 52.3 | native | native | native | 81 % | 91 % |
| Jev-Omni | 55.5 | native | native | native | 74 % | 94 % |
| Xor 26B-A4B | 58.7 | native | native | native | 80 % | 94 % |
| Bobcat Flash 1.2 | 65.8 | native | native | native | 89 % | 92 % |
| SPX-CD Flash | 58.4 | native | native | native | 81 % | 94 % |
| decider chat on Gemma-4-31B-it | 58.5 | native | native | native | 80 % | 98 % |
| Hopper 12B trained | 51.7 | native | native | native | 79 % | 93 % |
| Surogate Rune 26B-A4B v3 | 56.1 | native | native | native | 73 % | 97 % |
| GEV-26B-Decide | 49.8below gate | native | native | native | 70 % | 96 % |
| SPX-CD-Omni | 48.8below gate | native | native | native | 68 % | 93 % |
| Metask rain 4B | 51.1 | native | native | native | not supported | — |
| OpenAI Decisions | 56.9 | — | — | — | — | — |
| Instinct | 47.8below gate | — | — | — | — | — |
| Blink v0.3 26B-A4B | 49.2below gate | native | native | native | 67 % | 95 % |
| Diffusion Jev | 54.9 | native | native | native | 93 % | 85 % |
| René-1 31B FP8 | 61.7 | native | native | native | 78 % | 98 % |
| Aplomb 1 | 45.8below gate | native | native | native | 68 % | 93 % |
| Wald 4B v2.1 | 52.0 | native | native | native | 100 % | 78 % |
| Bespoke Nimble 9B v3 | 56.5 | native | native | native | 87 % | 89 % |
| APUS-OpenJev-v1-9B | 47.2below gate | native | native | native | 88 % | 85 % |
| Plumb-4B | 43.0below gate | native | native | native | 71 % | 91 % |
| SPX-CD Pro | 64.9 | native | native | native | 84 % | 95 % |
| Messier One v0.2 | 42.4below gate | native | native | native | 85 % | 85 % |
| spark-s1-4b-v6 | 43.2below gate | native | native | native | 84 % | 83 % |
| swanOne | 56.2 | native | native | native | 81 % | 93 % |
| Quyet-1.0-Medium | 41.7below gate | native | native | native | 76 % | 84 % |
| Vansa-3.4 | 40.9below gate | — | — | — | — | — |
| lev | 41.1below gate | native | native | native | 84 % | 83 % |
| NInfer Qwen3.8-Flash-Next mixed | 49.0below gate | native | native | native | 75 % | 93 % |
| jev-local | 41.9below gate | native | native | native | 85 % | 82 % |
| decider-4b v2 | 40.1below gate | native | native | native | 64 % | 92 % |
| GPT-6 Luna | 96.6 | — | — | — | — | — |
| Autoloops – Gemma 4 31B IT | 58.1 | — | — | — | — | — |
| Wald 4B | 41.1below gate | native | native | native | not supported | — |
| GPT-6 Luna | 96.9 | — | — | — | — | — |
| ClassOne Qwen 3.5 9B | 42.2below gate | native | native | native | 97 % | 79 % |
| JevK5 v0.3 | 38.5below gate | native | native | native | 66 % | 95 % |
| metask-jev-4b | 39.2below gate | native | native | native | 75 % | 88 % |
| deck-4B v1.0 | 38.6below gate | native | native | native | 68 % | 93 % |
| Clef-Flash | 42.2below gate | native | native | native | 68 % | 93 % |
| Decision-4B | 37.9below gate | native | native | native | 72 % | 88 % |
| janus 4B | 37.6below gate | native | native | native | 63 % | 94 % |
| Qwen3.5-9B Jev-like data-mix v2 | 41.9below gate | native | native | native | 71 % | 91 % |
| decisor-4b | 39.6below gate | native | native | native | not supported | — |
| JevK5 v0.2.0 | 36.3below gate | native | native | native | 62 % | 93 % |
| JevOne | 46.0below gate | native | native | native | 62 % | 96 % |
| Kahn1 4B | 39.7below gate | native | native | native | 70 % | 87 % |
| Fastino GLiDE | 66.1 | — | — | — | — | — |
| TypeCastLM 1.4.0 | 35.4below gate | native | native | native | 77 % | 81 % |
| Hopper | 34.7below gate | native | native | native | 56 % | 94 % |
| Imajev-4B | 34.6below gate | native | native | native | 52 % | 96 % |
| SimpleJev Qwen3.6-35B-A3B | 56.3 | — | — | — | — | — |
| Malkuth-4B | 33.8below gate | native | native | native | 65 % | 92 % |
| jqv | 33.9below gate | native | native | native | 66 % | 94 % |
| JPT-35B-A3B | 51.4 | native | native | native | 74 % | 94 % |
| decider-35b-a3b | 48.7below gate | native | native | native | 74 % | 91 % |
| CoCo-Decision-4B | 32.7below gate | native | native | native | 56 % | 96 % |
| Manchego v2.1 | 32.5below gate | native | native | native | 62 % | 90 % |
| Decision 4B v1.2 | 32.4below gate | native | native | native | 59 % | 94 % |
| Raw Qwen3 4B Instruct 2507 direct logits | 34.0below gate | native | native | native | 92 % | 73 % |
| RSI-Jev v6.1-VL 27B | 61.6 | native | native | native | 83 % | 91 % |
| GPT-5.6 Luna | 94.1 | — | — | — | — | — |
| JEV-27B | 56.8 | native | native | native | 79 % | 95 % |
| torchcast-decision-27b | 71.8 | native | native | native | 83 % | 97 % |
| Bespoke Nimble 9B | 43.1below gate | native | native | native | 77 % | 87 % |
| reflex 4B | 30.9below gate | native | native | native | 58 % | 92 % |
| Perplexity Decider v1.1 27B | 75.7 | native | native | native | 91 % | 95 % |
| JADE | 47.9below gate | native | native | native | 94 % | 91 % |
| typecastlm | 29.8below gate | native | native | native | 57 % | 90 % |
| Jebadiah 27B | 57.6 | native | native | native | 74 % | 98 % |
| Decision 2.0 Vega 27B | 60.4 | native | native | native | 73 % | 99 % |
| Standard One 8B SH | 34.5below gate | native | native | native | 61 % | 88 % |
| Decision 4B v1.1 | 29.1below gate | native | native | native | 57 % | 94 % |
| AutoJev-27B | 55.6 | native | native | native | 70 % | 99 % |
| Eikos-27B | 64.4 | native | native | native | 81 % | 98 % |
| Instinct Dual 4B | 28.3below gate | — | — | — | — | — |
| NInfer Qwen3.8-27B NVFP4 | 51.8 | native | native | native | 82 % | 89 % |
| OpenJev | 70.3 | native | native | native | 96 % | 93 % |
| Open-Jev 9B | 43.9below gate | native | native | native | 81 % | 86 % |
| Gemini 3.1 Flash-Lite | 58.6 | — | — | — | — | — |
| Standard One 8B | 32.8below gate | native | native | native | 68 % | 90 % |
| Clef | 59.1 | native | native | native | 82 % | 94 % |
| system-one-open | 27.9below gate | — | — | — | — | — |
| JEV Qwen3.5-9B Base NVFP4 | 29.3below gate | native | native | native | 53 % | 89 % |
| SemIf, formerly OpenJev | 27.1below gate | native | native | native | 62 % | 90 % |
| openjev-sglang | 38.8below gate | — | — | — | — | — |
| Bev / Bonsai 27B | 43.8below gate | native | native | native | 70 % | 91 % |
| LitJev | 43.3below gate | native | native | native | 66 % | 98 % |
| local-jev Qwen3.5-4B | 24.1below gate | native | native | native | 35 % | 99 % |
| system-one | 29.3below gate | native | native | native | 92 % | 73 % |
| kev 8B | 29.3below gate | native | native | native | 77 % | 81 % |
| Decision 2B | 22.9below gate | native | native | native | 39 % | 94 % |
| OpenSourceJev | 23.2below gate | native | native | native | 49 % | 95 % |
| open-alternative-jev | 22.3below gate | native | native | native | 60 % | 87 % |
| Raw Qwen3 8B direct logits | 26.5below gate | native | native | native | 92 % | 71 % |
| Liquid AI d1-3B | 22.1below gate | native | native | native | 49 % | 83 % |
| decider-2b | 20.7below gate | native | native | native | 68 % | 78 % |
| Nemotron Diffusion 8B | 20.4below gate | native | native | native | 61 % | 77 % |
| DeepSeek V4.1 Flash | 94.7 | — | — | — | — | — |
| kev 4B | 19.5below gate | native | native | native | 67 % | 82 % |
| Clef-omni | 34.5below gate | native | native | native | 65 % | 86 % |
| Kev 27B | 64.2 | native | native | native | 80 % | 97 % |
| Malkuth-2B | 15.9below gate | native | native | native | 53 % | 84 % |
| Qwen3-Reranker-4B | 16.3below gate | native | native | native | 25 % | 75 % |
| Vega 4B | 16.7below gate | native | native | native | 70 % | 68 % |
| SimpleJev Qwen3.8-27B | 57.9 | native | native | native | 78 % | 96 % |
| ZeroEntropy zerank-2 | 15.6below gate | native | native | native | 14 % | 87 % |
| EXAONE-4.0-1.2B-JEV v0.3 | 14.8below gate | native | native | native | 40 % | 85 % |
| Open-Jev 2B | 21.8below gate | native | native | native | 67 % | 75 % |
| SimpleJev Qwen3.8-27B | 52.9 | — | — | — | — | — |
| Open-Jev 27B v1.1 | 52.5 | native | native | native | 67 % | 94 % |
| Raw Phi-4 mini direct logits | 13.7below gate | native | native | native | 50 % | 61 % |
| decision-machine-1 | 13.1below gate | — | — | — | — | — |
| SimpleJev | 13.0below gate | native | native | native | 100 % | 58 % |
| JevAct | 10.1below gate | — | — | — | — | — |
| smalljev semantic-v9 | 8.7below gate | native | native | native | 23 % | 73 % |
| Gutsy 0.8B v0.3 | 8.6below gate | native | native | native | 29 % | 74 % |
| kev 0.6B | 7.6below gate | native | native | native | 43 % | 69 % |
| jul fast | 7.5below gate | native | native | native | 30 % | 89 % |
| OpenDecision | 7.4below gate | native | native | native | 51 % | 68 % |
| ClassOne Gemma 4 E2B | 7.8below gate | native | native | native | 83 % | 59 % |
| Qwen3.5-0.8B Decision Model | 7.3below gate | native | native | native | 22 % | 88 % |
| Fastino GLiNER-2.5-Decide | 7.0below gate | confidence | confidence | confidence | 100 % | 54 % |
| Deem 0.8B v1 | 6.5below gate | native | native | native | 77 % | 59 % |
| Decision Fast | 6.1below gate | native | native | native | 14 % | 90 % |
| Quyet-1.0-Small-EN | 6.1below gate | native | native | native | 28 % | 64 % |
| Laya typed-decisions | 5.5below gate | native | native | native | 1 % | 100 % |
| WaterSheep | 5.3below gate | native | native | native | 37 % | 69 % |
| openJev Verdict 1.4 | 5.0below gate | native | native | native | 0 % | — |
| Raw Qwen3 0.6B direct logits | 5.6below gate | native | native | native | 100 % | 42 % |
| lev-350m | 4.8below gate | native | native | native | 8 % | 69 % |
| Raw Qwen3 1.7B direct logits | 5.1below gate | native | native | native | 100 % | 42 % |
| kev 0.5B | 4.6below gate | native | native | native | 26 % | 56 % |
| openJev Verdict | 4.1below gate | native | native | native | 33 % | 60 % |
| Quyet-1.0-Small | 4.0below gate | native | native | native | 29 % | 71 % |
| watt-flash-0.1 | 3.8below gate | native | native | native | 1 % | 60 % |
| Quyet-1.0-Tiny | 3.7below gate | native | native | native | 20 % | 64 % |
| Liquid AI d1-omni-600M | 3.3below gate | native | native | native | 27 % | 56 % |
| Vega 0.8B | 3.1below gate | native | native | native | 32 % | 63 % |
| open-jev-deberta-v3-large | 2.8below gate | native | native | native | 15 % | 58 % |
| verdict-small | 2.3below gate | native | native | native | 45 % | 62 % |
| Tacet Sonata | 2.2below gate | native | native | native | 33 % | 56 % |
| Laya | 1.8below gate | native | native | native | 19 % | 61 % |
| Mixedbread mxbai-rerank-base-v2 | 1.1below gate | native | native | native | 0 % | — |
| Laya multilingual | 0.3below gate | native | native | native | 72 % | 50 % |
| Certo v1 | 0.2below gate | native | native | native | 1 % | 25 % |
| CLM-8B | 0.1below gate | native | native | native | 51 % | 69 % |
| BAAI bge-reranker-v2-m3 | 0.1below gate | native | native | native | 0 % | — |
| Alibaba GTE Reranker ModernBERT-base | 0.0below gate | native | native | native | 0 % | — |
| Mirror | 0.0below gate | native | native | native | 92 % | 60 % |
| Open Jev JSON Canvas | 61.2 | label | label | label | 100 % | 85 % |
| Qwen3.8 27B | 96.4 | — | — | — | — | — |
All data (193 systems)
Not yet measured on v1.6 · 13 systems with a dated carried score
Every ranked system of the live v1.5.7 board that is not measured on the v1.6.0 pool keeps its last published score, marked with the release that first published that measurement and that release's publication day. Carried rows are listed separately and never ranked together with v1.6-measured rows. v1.5 protocol (1,624 decisions: 904 open + 720 sealed); scores are on the v1.5 scale and are not comparable with v1.6-measured rows.
Dates: v1.5.0 published 2026-09-28 (12) · v1.5.3 published 2026-09-29 (1). The date is the publication day of the release that first published the measurement, not a per-model measurement timestamp.
| System | Measured on | Capability (v1.5 scale) | v1.5 Composite | Cost / 1,000 | Median latency |
|---|---|---|---|---|---|
| AutoJev-27B (RTX PRO 6000) | measured on v1.5.0 (2026-09-28) | 79.7 | 19.5 · was #56 on v1.5.7 | $0.2262estimate | 0.33 s |
| Surogate Rune 26B-A4B v3 (RTX PRO 6000)Not yet measured on the v1.6 pool. | measured on v1.5.0 (2026-09-28) | 79.0 | 66.5 · was #19 on v1.5.7 | $0.0502estimate | 0.35 s |
| reflex-27b (Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 74.4 | 13.2 · was #68 on v1.5.7 | $0.2973estimate | 2.75 s |
| NInfer Qwen3.8-27B NVFP4 (T=1.5) | measured on v1.5.0 (2026-09-28) | 73.8 | 18.5 · was #59 on v1.5.7 | $0.2305estimate | 0.25 s |
| jeff (Logan Markewich, GLiFormer 400M) | measured on v1.5.0 (2026-09-28) | 42.4 | 0.1 · was #89 on v1.5.7 | $0.0043estimate | 7.14 s |
| Von (wfzyx, Option-Marker 395M) | measured on v1.5.0 (2026-09-28) | 41.7 | 0.0 · was #110 on v1.5.7 | $0.0038estimate | 0.92 s |
| Bosun v3.1 0.6B | measured on v1.5.3 (2026-09-29) | 39.4 | 2.5 · was #78 on v1.5.7 | $0.0056estimate | 4.04 s |
| GLiNER2.5 multi (Fastino, 287M) | measured on v1.5.0 (2026-09-28) | 36.0 | 2.7 · was #77 on v1.5.7 | $0.0028estimate | 1.37 s |
| GLiNER2.5 small (Fastino, 74M) | measured on v1.5.0 (2026-09-28) | 32.5 | 0.9 · was #85 on v1.5.7 | $0.0028estimate | 0.47 s |
| GLiNER2 large (Fastino) | measured on v1.5.0 (2026-09-28) | 32.4 | 8.4 · was #70 on v1.5.7 | $0.0056estimate | 1.88 s |
| GLiNER2 (Fastino, gliner2.5-base) | measured on v1.5.0 (2026-09-28) | 24.5 | 2.3 · was #79 on v1.5.7 | $0.0028estimate | 1.00 s |
| Needle 3 (Cactus, 2-bit, local CPU) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 · was #102 on v1.5.7 | $0.0191estimate | 135.51 s |
| Needle 3, options as tools (post-hoc adapter mode) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 · was #103 on v1.5.7 | $0.0191estimate | 58.16 s |
How it works
One merged board: how newer draws join it
Each system appears once, with its latest valid measurement; earlier systems are not re-measured. The reference scale is v1.6.1 (draw v1.6.0). Every draw shares the same 300 public items (P, manifest 6b2321f8...). A fresh draw is equated to the reference draw with fixed anchor systems measured on both draws: offset = median over anchors of (metric on the reference S u P) - (metric on the new draw S u P), for Intelligence and Calibration, with a bootstrap 95% interval over anchors and items. A row from that draw is shown as published plus the offset; Capability and Composite follow from the equated axes. Until the anchor runs of a draw are complete its offset is 0 and the draw is marked 'anchor runs pending'.
| Draw | Drawn | Releases | Anchor offset (Intelligence / Calibration) | Status |
|---|---|---|---|---|
| v1.6.0 (reference) | 2026-10-01 | v1.6.0, v1.6.1 | 0 / 0 | reference scale · G_med 2.57 |
| Fast-lane draw · v1.6-fastlane-20261009 | 2026-10-09 | v1.6.2, v1.6.3, v1.6.4 | +0.00 / +0.00 | anchor runs pending: shown as published |
| Regular draw · v1.6-regular-20261010-a2 | 2026-10-10 | v1.6.5, v1.6.6, v1.6.7 | +0.00 / +0.00 | anchor runs pending: shown as published |
- Anchor pool: hopper, kev-0.6b, kev-4b, kev-8b, localjev-qwen3.5-4b, malkuth-4b, metask-jev-4b, raw-phi-4-mini, raw-qwen3-1.7b, raw-qwen3-4b-instruct-2507, raw-qwen3-8b, typecastlm. The same 12 anchors answered two sealed draws of v1.6.0 (S and the A2 supplement): median difference -0.29 Intelligence and +1.56 Calibration points (v1.6.1 results, v16.equating_A2).
- On the shared public items, systems on both new draws score 6 to 13 sealed points below the reference field's sealed-vs-public line (reference residual SD 3.8). Either the new sealed draws are harder or their cohorts are tuned on the public items; only anchor runs can separate the two. Until then their rows are, if anything, understated.
- Rows from a newer draw keep the gap-penalty reference (G_med) their release was scored with; the row tooltip and tag name the draw and date. Their topic, use-case and language cells use their own draw’s item set and labels, so they are shown on their release page (linked from the draw tag in Release history), not in the radars and language table here.
Open weights and API offerings (board v1.7.43)
Why open weights and APIs are compared in separate groups
- Fair cost and speed. We run every open-weights model on hardware we rent and operate, so cost and latency compare on the same terms. The GPU cost calculator prices your own setup: own hardware, on-demand or long-term rental.
- API prices can change. An API price is the vendor's decision. It can be subsidised (for example on top-end GPUs) and raised later, and readers cannot reproduce it.
- Different fairness needs. A hosted endpoint chooses its own hardware and sees the benchmark inputs, so API offerings are compared with each other.
- Open-source focus. JevBench exists to make open decision models comparable and reproducible. Every score and the method are identical in every view; the All view ranks both groups together by Capability.
- The main board at /jev-models ranks open-weights systems: published weights that we ran ourselves, on GPU or CPU machines we rent and operate. A system we measured through an endpoint we do not run (vendor API, author-hosted or third-party-hosted endpoint, for example Qwen3.8 27B via Chutes) is an API offering, even when its base weights are open; API offerings are ranked on the API leaderboard.
- Why separate boards: open weights can be compared on equal hosting terms, while an API price is a vendor decision that can be subsidised or raised later and is not reproducible by readers. Every score and measurement is the same on both boards; only the set of ranked rows differs, and ranks are the published order filtered to that set.
- Jev 1.13.0 is a hosted API. It stays on the open-weights board as the reference row (it defines the Jev-class cost and latency caps) and is not ranked there; it is ranked on the API leaderboard.
- Official cost basis is unchanged: the Cost axis keeps each row's documented reference price (see the cost notes below; APIs with a known base model are priced at the developer's own list price). The base-model reference price only applies to open-weights rows; API offerings are ranked at their own list price. The GPU cost calculator on the main board is a What-If for your own hosting and never changes a score or rank.
- Full API re-run (A4, v1.7.7). Most API offerings were measured on 6 Oct 2026 on a fresh sealed API set A4 (300 never-used sealed items) plus the same 300 public items every system answers, and equated to the v1.6.1 S ∪ P scale with the published A2/A3 supplement method (offset +2.57 Intelligence, +3.35 Calibration; pool of 8 ranked self-hosted systems re-run on A4 ∪ P). Jev, Sage, wity-1 and Fastino GLiNER-2.5-Decide keep their full-set S ∪ P scores; Liquid AI d1 answered the full S ∪ P set on 6 Oct 2026 (v1.7.8) and is scored like them, without equating. The pool is mid-strength, so for the strongest LLM rows the offset is an extrapolation (likely within ±2 Intelligence points; the 95% intervals include it). Nine A4 ∪ P items of about 77,000–82,000 input tokens, beyond the Jev reference's accepted input range, count against Intelligence when refused but are left out of cost. Rows without a public tariff keep their documented v1.5 cost estimate. A4 is now retired for everyone. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
- Language and use-case cells (v1.7.12). These breakdowns add two sealed supplements (354 language items, 33 use-case items, drawn and reviewed on 6 Oct 2026), so every language and use case has at least 30 items. Headline scores are unchanged; rows not yet run on the supplements are tagged S+P in the language table.
- Every API offering we reached is measured on v1.6.1: either on the full 1,500-item set or on A4 ∪ P (600 items, equated). No preliminary public-set rows remain.
Revision history
- v1.7.43 (2026-10-10): Display only: a Pareto frontier section after the two scatter charts on both boards — Capability against cost per 1,000 decisions and against median latency (log axes; list price and measured endpoint latency on the API board). The red line joins the systems no other system beats on both axes. No score or rank changed.
- v1.7.42 (2026-10-10): One merged board. The version tabs are gone: /jev-models is always the current board with every system’s latest measurement, and releases are listed under Release history (old version pages keep their URLs). The 16 systems measured on the two fresh draws after v1.6.1 (fast-lane draw v1.6.2–v1.6.4, regular draw v1.6.5–v1.6.7) join the boards with their published scores and a draw tag; each draw is put on the v1.6.1 scale by an anchor offset that stays 0 until the anchor systems have run on it (see How it works). A filter on top switches between Open weights, API and All (both groups, ranked by Capability). Header text is shortened; the details moved into How it works. Jeff 1.0 Large (fast-lane draw) enters the open-weights Capability top 5 at #2.
- v1.7.41 (2026-10-10): Needle 3 now has language and category cells on the current pools. It keeps its published v1.5 headline. With the benchmark owner's approval its closed engine was sent all supplement items (L1, L2, L3) once more on an isolated machine, plus 128 never-asked S+P items chosen by their use-case label to complete the last radar spoke (these also count in the other cells, so its S+P part is not a random sample); together with an earlier partial run both radars reach at least 30 answered items per spoke (52 per subject topic, 43 per use case). The engine abstains on about 30 % of items and returns invalid UTF-8 on about 7 %; both count as unanswered, so 22 languages stay below the 60-item target (19 to 59 answered items). Its answers to the same public items differ between machines on about 10 %. Its chance-corrected competence is below zero in every cell, shown as 0. The v1.7.40 note for Needle 3 (options as tools) is corrected the same way. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.40 (2026-10-10): Needle 3, options as tools, now has language and category cells on the current pools. It keeps its published v1.5 headline. With the benchmark owner's approval its closed engine ran once more on an isolated machine; the run covered most of L3 and half of L1 before the machine's time limit and is combined with an earlier partial run. Both radars are complete (at least 58 per subject topic and 43 per use case). The engine abstains on about 9 % of items and returns invalid UTF-8 on about 7 % (mostly non-Latin scripts); both count as unanswered, so 18 languages stay below the 60-item target (24 to 59 answered items; columns under 30 carry the low-n mark). Its chance-corrected competence is below zero in every cell, shown as 0. ClassOne Gemma 4 E2B keeps its exception with a corrected reason: its weights were identical, but its server gives different answers after each restart. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.39 (2026-10-10): GLiNER2 large now has language and category cells on the current pools: at least 62 completed responses in every language, 98 in every subject topic and 80 in every use case. It keeps its published v1.5 headline; the cells come from a new run with the original pinned model on an isolated machine. GLiNER2 gained 180 more answered items that had been reserved earlier but never sent: its languages now rest on 50 to 59 answered items where they are below the 60-item target (17 of 23), and its radars on at least 69 per topic and 56 per use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.38 (2026-10-10): Three more historical rows now have language and category cells on the current pools: reflex-27b, Surogate Rune 26B-A4B v3 (RTX PRO 6000) and jeff. They keep their published v1.5 headlines. reflex-27b completed its partly answered S+P set and the supplements; Surogate Rune adds the supplements to its 6 October S+P run; jeff (CPU) answered all pools. Their repeated answers to the 300 public items agree 100 %. Each row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.37 (2026-10-10): Three more historical rows now have language and category cells on the current pools: Von, GLiNER2 and GLiNER2.5 multi. They keep their published v1.5 headlines; the cells come from new runs with the original pinned model and setup on isolated machines. Von and GLiNER2.5 multi reach at least 61 completed responses in every language, 100 in every subject topic and 80 in every use case. GLiNER2 fills both radars (at least 57 per topic, 45 per use case) and every language column, but about 1,000 of its issued items were lost to out-of-memory kills and a machine shutdown; they stay counted as spent, so 22 languages rest on 42 to 59 answered items, below the 60-item target. Three rows that cannot be measured now give the concrete reason: Aplomb 1 (pinned revision deleted upstream), ClassOne Gemma 4 E2B (re-run did not reproduce its original answers) and swanOne (weights return HTTP 401). L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.36 (2026-10-10): Nine more rows now cover S+P+L1+L2+L3 in the language table and both radars: Diffusion Jev, SPX-CD-Omni, Seb-9B, APUS-OpenJev-v1-9B, APUS-OpenJev-v1-35B-A3B, Jobe Qwen3.5-4B, AutoJev-27B (RTX PRO 6000), NInfer Qwen3.8-27B NVFP4 (T=1.5) and OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). The last four are historical rows: they keep their published v1.5 headlines, and their language and category cells come from a new run on the current pools with the original pinned model and serving setup. Diffusion Jev, SPX-CD-Omni, Seb-9B and both APUS rows match their original public answers on at least 97 %. The razorback16 OpenJev model samples stochastically, so its repeated passes agree on 86 to 89 % of the public items; its note says so. Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.35 (2026-10-10): Seven more rows now cover S+P+L1+L2+L3 in the language table and both radars: René-1 31B FP8, decisio v0.8.0 on gemma-4-12B-it, Hopper 12B trained, decider-12b v2, decider-12b v1, Bobcat Flash 1.2 and Mica v0.1 4B. Each answered all 1,505 supplemental items with its original pinned model and serving setup on an isolated machine; answers to the 300 public items match the original runs (at least 99 %). Mica keeps its published v1.5 headline; its language and category cells come from a new run on the current pools. Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The language table now lists the mixed-language group first. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.34 (2026-10-10): Seven more self-hosted rows now cover S+P+L1+L2+L3 in the language table and both radars: TypeCastLM 1.4.0, CoCo-Decision-4B, ClassOne Qwen 3.5 9B, EXAONE-4.0-1.2B-JEV v0.3, jul fast, WaterSheep and Tacet Sonata. Each answered all 1,505 supplemental items with its original pinned model and serving setup on an isolated machine; their answers to the 300 public items match the original runs (at least 97 %). Every row has at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. Long-input refusals stay scored as before. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and pricing are unchanged.
- v1.7.33 (2026-10-10): Messier One v0.2 now covers S+P+L1+L2+L3, with all 1,505 supplemental items answered natively on the original image with the network disabled: at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The original P300 answers recur in each supplement and count once; the original S+P measurement had no HTTP errors. L3 items were reviewed by an OpenAI model, so cells based on L3 carry that exposure; headline scores do not use L3. Headline scores, Capability, Composite, ranks and original row pricing are unchanged.
- v1.7.31 (2026-10-10): Microsoft-Decision-1 joins the hosted API board after a native Azure Foundry measurement on 1,200 sealed plus 300 public observations: Capability 70.78, Calibration 84.40 and Composite A 69.11. Its 21 context-limit failures remain in the scored denominator. Native median latency is 459 ms on the 124-request speed subset. This new full-set O1S run uses its completed-field median gap reference of 7.088 score points; historical rows retain their own dates, pools and references rather than being remeasured or converted. The existing top five on Composite A and capped Capability are unchanged. All 50 topic/use-case/language cells are recorded; thin cells stay suppressed or table-only, and qualified supplemental coverage remains pending.
- v1.7.30 (2026-10-10): decisio v0.8.0 on gemma-4-31B-it now covers S+P+L1+L2+L3, with all 1,505 supplemental items answered natively: at least 65 completed responses in every language, 109 in every subject topic and 83 in every use case. The original P300 answers recur in each supplement and count once. The original 23 HTTP 500 errors remain operational non-answers and are not imputed. L3 items were reviewed by an OpenAI model, so its cells based on L3 carry that exposure; this row is the original locally served Gemma offering. Headline scores, Composite, ranks and original row pricing are unchanged.
- v1.7.29 (2026-10-09): Mercury Decide now has at least 64 completed observations in every language, 109 in every subject topic and 83 in every use case, using S+P+L1+L2+L3. L3 retains 1,416 records on its full 1,418-item denominator, including original refusals and operational failures; two spent failures remain missing. Paid supplemental decisions use the currently declared Inception 2026-09-30 offering and the original native protocol. Public top labels agree on 96.63%, with differing probabilities and unproven historical weights/calibrator equality; the row coverage note gives the comparison. Headline scores, Composite, ranks and original row pricing are unchanged.
- v1.7.28 (2026-10-09): GLiNER 2.5 Small now covers P300 + L1 + L2 + L3 with 1,802 completed native responses: at least 60 in each of 23 languages, 78 in every subject topic and 52 in every use case. Its original offline CPU/fp32 profile is unchanged. Three earlier public-item OOM outcomes remain missing and contribute no completed coverage; no requests were replayed. These cells exclude historical S answers and stay outside headline scores. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.27 (2026-10-09): OpenJev DeBERTa v3 Large adds 60 new L5 items: 20 each in Arabic, Hindi and Greek. It now has at least 63 completed responses in every language (Arabic 70, Hindi 73, Greek 69). The original native context was checked before keeper selection; all 60 requests succeeded, and earlier failures remain scored. Other rows and category views retain their existing measurements. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.26 (2026-10-08): BB-Qwen3.5-4B-LoRA now covers S+P+L1+L2+L3, with at least 65 completed responses in each of 23 languages and at least 83 on every capability and use-case spoke. Two DNS failures remain scored and do not count toward completed coverage. Its reference cost stays unknown; these view values do not impute a price. OpenJev DeBERTa v3 Large now includes its completed L1/L2 runs; Arabic (50), Hindi (53) and Greek (49) still need more completed responses. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.25 (2026-10-08): Watt Flash 0.1 and OpenJev Verdict completed their language supplements. Both now cover S+P+L1+L2+L3, with at least 65 completed responses in each of 23 languages and at least 83 in each capability and use-case category. Other rows retain their existing measurements, including Fastino’s Hindi L4 supplement. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.24 (2026-10-08): Fastino GLiNER 2.5 Decide now has 69 completed Hindi responses after a new 20-item Hindi supplement. Seventeen requests succeeded; three HTTP 400 failures remain scored and do not count toward completed coverage. L4 is used only for this offering’s Hindi language cell, with OpenAI author and independent Anthropic reviewer provenance disclosed. The mghafiri Qwen3.5 0.8B row also completed L3, expanding its language and category cells. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.23 (2026-10-08): The Languages table includes every listed model, including historical and catalogue entries. These rows use their stored language measurements where available and show pending cells where measurements are unfinished. Historical headline scores are not used as language values. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.22 (2026-10-08): Languages and radars now include completed L1/L2/L3 runs for SPX CD Flash, SPX CD Pro and Decisio Gemma 4 12B v0.9. Coverage counts completed supported responses and recorded input refusals; authentication, rate-limit, service and transport errors do not satisfy reporting thresholds. Competence keeps its original scored observations. 3 listed rows still have incomplete radar coverage and show the reason. Every listed row remains in the completion check. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.21 (2026-10-08): Languages and radars: more rows completed their L1/L2/L3 supplement runs, among them Instinct (its API is back), Surogate Rune 26B-A4B v3, Xor 26B-A4B, JADE, Jebadiah 27B, Kev 27B, Decision 2.0 Vega 27B and the self-hosted SimpleJev Qwen3.8-27B. 3 rows still show a reason instead of a full radar: newly listed rows whose supplement runs are not done yet, rows not measured on this item pool, and one model whose pinned weights could not be accessed. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.20 (2026-10-08): Correction and progress for the language and radar views. The self-hosted row SimpleJev Qwen3.8-27B wrongly counted the hosted SimpleJev API row's L1/L2/L3 supplement answers in its language and category cells (v1.7.18 and v1.7.19), because both share a file name. Its cells now use only its own S + P answers until its own supplement runs finish, and the cell builder no longer lets a hosted-API run count for a self-hosted row. More open-weights rows completed their L1/L2/L3 runs. 3 rows still show a reason instead of a full radar. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.19 (2026-10-08): Languages and radars: the open-weights rows were run on the L3 supplement overnight, and most rows that still lacked the earlier L1/L2 supplements got them too. Their language cells and topic/use-case radars now count every item they answered (S + P + L1 + L2 + L3), and Qwen3.8 27B (Chutes) finished its L3 run. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows carry that exposure; a note next to those rows says so, and headline scores do not use L3. 3 rows still show a reason instead of a full radar, for example runs still in progress, a provider that is down, withdrawn weights or rows not measured on this pool. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.18 (2026-10-08): Languages and radars for every row. A new sealed supplement, L3 (1,118 items drawn on 7 Oct 2026: items written natively in 21 languages plus a mixed-language group, and English items for thin radar categories such as everyday language, safety and the "other" use case), gives every language at least 60 items on P ∪ L1 ∪ L2 ∪ L3 (before: as few as 4). The hosted API rows answered it: their language cells and their use-case and topic radars now rest on about 2,100 answered items (3,000 for rows that answered the full sealed set), with at least 69 on every spoke. Open-weights rows are being run on it in the current GPU wave and update as they finish. 3 rows that do not reach 30 items on every spoke yet show the reason next to the row. Headline scores, Capability, Composite and every rank are unchanged.
- v1.7.17 (2026-10-07): Languages now includes every measured row on both boards, with wrappers listed below the models. Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts. 19 API rows re-run on A4/A5 have answered L1/L2 supplements; their category radars count those items too. L3 coverage will follow when measured.
- v1.7.16 (2026-10-07): Compare view, API board: subject-topic and use-case radars for OpenAI Decisions, the 17 API offerings re-run on A4 ∪ P and the classifier.dev wrapper. Their 600 new sealed items were labelled with the same recipe as every other item (Winnow-12B Q8 on our own GPU pod; uc1 items keep their authoring use case); values are raw and rest on 600 items (300 sealed), so more categories fall under the 30-item spoke minimum and are listed below the radar. Each of these rows also gets a breakdown section on its own page. Every live row now has all breakdown views, and a release check keeps it that way. No score, axis, cost or rank changed.
- v1.7.15 (2026-10-07): Compare view, API board: OpenAI Decisions, the 17 API offerings re-run on A4 ∪ P and the classifier.dev wrapper now show competence per request type and per tier on the open and sealed items they answered (300 open + 300 sealed: fewer sealed items than the full set, so wider uncertainty; the view note says so), and Liquid AI d1 shows its subject-topic and use-case radars. Values are raw, computed from the stored per-item results with the same scorer as every other row. Subject-topic and use-case radars for the 600-item rows follow once their 600 new sealed items are labelled. No score, axis, cost or rank changed.
- v1.7.14 (2026-10-07): Open-weights board: H2O-Lightning-4B v1.1 (H2O.ai) and Kahn1 4B (Okura66), two Qwen3.5-4B fine-tunes submitted as JevBench add-requests, measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a8 of v1.6.1). Both are priced with labelled cost estimates like every other Qwen3.5-4B row; Kahn1's server reports no token usage, so its estimate is our own count of its input tokens (three passes per decision, its server default). Both rank inside the Jev-class caps: H2O-Lightning-4B at #3 of the Capability Score and #1 of the Composite, Kahn1 at #23 of the Capability Score. No published score, axis or cost changed.
- v1.7.13 (2026-10-07): Open-weights board: Surogate Rune 26B-A4B v3 (measured on the v1.6 pool for the first time; its dated v1.5 row stays in the carried list) and Xor 26B-A4B (Juspay), both from the top 20 of the community Jev Decision Index on Hugging Face, measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a7 of v1.6.1). Both rank inside the Jev-class caps, at #3 and #4 of the Capability Score. No published score, axis or cost changed.
- v1.7.12 (2026-10-07): Languages and use cases: the per-language and per-use-case cells now also count two sealed supplements, a language supplement L1 (354 items) and a use-case supplement L2 (33 items), drawn and reviewed on 6 Oct 2026. Every one of the 22 languages (plus the mixed-language group) and every one of the 20 use cases now has at least 30 items (before: as few as 5). 82 rows answered both supplements; rows not run on them yet keep their S + P cells and are tagged S+P (or S+P+L1). Headline scores, Capability, Composite and every rank are unchanged on both boards.
- v1.7.11 (2026-10-07): Open-weights board: 12 new self-hosted Jev-compatible rows from the top 20 of the community Jev Decision Index on Hugging Face (Perplexity Decider v1.1 27B, torchcast-decision-27b, Kev 27B, decider chat on Gemma-4-31B-it, GEV-26B-Decide, Decision 2.0 Vega 27B, JEV-27B, Jebadiah 27B, JADE, Bespoke Nimble 9B v3, SimpleJev Qwen3.8-27B self-hosted, JPT-35B-A3B), measured on the same sealed v1.6 pool with the same scorer, G_med and cost basis as every other row (addendum a6 of v1.6.1). No published score, axis or cost changed; ranks move only where a new row places above. Rows priced above the Jev-class cost or latency caps are listed under the limits section, not in the Capability ranking.
- v1.7.10 (2026-10-06): API board: OpenAI Decisions (POST /v1/decisions with gpt-6-luna) replaces its preliminary public-set row with its official result. It answered a fresh sealed API set A5 (300 never-used sealed items, since A4 had already been sent to OpenAI) plus the 300 public items (598 of 600 answered) and is equated to the v1.6.1 scale with the same method and pool as the A4 rows (offsets +5.78 Intelligence, +6.15 Calibration). It is ranked on the API board: Composite 62.5 (95% interval 51.6–64.6), Capability 73.5; list price USD 0.10 per 1M input tokens as of 6 Oct 2026 (launch day), re-checked at each revision. It enters the API Composite top 5 at #5, statistically tied with Instinct (62.2, now #6), and the Jev-class Capability top 5 at #4 (Instinct moves to #5, Vansa-3.4 to #6). Every other score is unchanged; the open-weights board is unchanged.
- v1.7.9 (2026-10-06): API board: the OpenAI Decisions API (POST /v1/decisions with gpt-6-luna, opened on 6 Oct 2026) is added as a preliminary, unranked row from the 300 public v1.6 items (298 answered; Capability 68.8, Composite 37.2 on that set; list price USD 0.10 per 1M input tokens). Its official sealed run waits for the next fresh sealed API draw. No score or rank changed on either board.
- v1.7.8 (2026-10-06): API board: Liquid AI d1 is added as a ranked API offering. It answered the full v1.6.1 set (1,200 sealed + 300 public items, all 1,500 answered) on 6 Oct 2026 and is scored exactly like the other full-set API rows, without equating (Composite 73.0, Capability 74.4; tariff USD 0.04 per 1M input tokens). It enters the API board top 5 in Composite and in Jev-class Capability. Every other score is unchanged; the open-weights board is unchanged.
- v1.7.7 (2026-10-06): API board: 17 API offerings re-measured in full on a fresh sealed API set (A4, 300 never-used sealed items plus the 300 public items, 6 Oct 2026) and equated to the v1.6.1 scale with the published A2/A3 supplement method. They replace the hatched preliminary rows and are ranked on the API board: Instinct and Vansa-3.4 enter its Composite top 5, and Instinct, Vansa-3.4 and Instinct Dual 4B join Sage and Jev in its Jev-class Capability top 5; GPT-6 Luna, Qwen3.8 27B, GPT-6 Luna low, DeepSeek Flash and GPT-5.6 Luna have the highest raw Capability (98.4, 98.2, 97.8, 97.4 and 94.5) but sit outside the Jev-class caps. Qwen3.8 27B (Chutes) is ranked as well (Composite 0.0: no published tariff, so its documented USD 2.18 estimate per 1,000 decisions gives Cost 0). Instinct and Instinct Dual 4B are now treated as production APIs (public tariff), so no demo-endpoint latency adjustment applies to them. classifier.dev (fast tier) is listed as a wrapper, not ranked. No preliminary rows remain. The open-weights board is unchanged.
- v1.7.6 (2026-10-06): API board: Vansa-3.4 and Instinct Dual 4B now have v1.6 public-set figures and appear as preliminary rows; Autoloops stays pending while its full run on the fresh sealed set is in progress. Very long items that exceed a provider’s context window count as wrong without ending the run. No rank changed.
- v1.7.5 (2026-10-06): API board: every API offering with a v1.6 public-set figure is drawn in the Composite and Capability charts as a hatched, unranked “preliminary” row at its score position (public set of 300 items; full sealed re-evaluation running); offerings without any v1.6 figure are greyed “pending” rows. Adds Qwen3.8 27B (Chutes) and Fastino GLiDE to the public-set table. The four ranked rows and every rank are unchanged; the open-weights board is unchanged.
- v1.7.4 (2026-10-06): API board: every reachable API offering that was not yet re-measured on v1.6 now has a dated public-set figure (the 300 public v1.6 items, no sealed items), shown next to its older v1.5 score with Jev and other ranked APIs on the same 300 items as anchors. Not ranked and not comparable with the 1,500-item headline; no score or rank of either board changed.
- v1.7.3 (2026-10-06): Display only, no score or rank changed. The API board ranks every offering at its own list price and no longer shows a base-model reference price (that comparison only matters against open weights). The Jev reference row and, with “Show API offerings” on, every API offering now sit at their score position in each ranking section (Composite and Capability) instead of the folded end of the list.
- v1.7.2 (2026-10-06): Display only. The API roster says why carried API rows were not re-run on v1.6 yet (exposure cadence, retired v1.6.0 sealed set) and that no v1.5 to v1.6 conversion is applied.
- v1.7.1 (2026-10-06): Display only, no score or rank changed. Headings name the board (open weights / API offerings); the API board leads with the Composite Score and lists every API offering we measured, including mode variants, carried v1.5.x rows and wrappers; the Jev reference row reads “Not ranked, only shown as a reference to compare with” on the open-weights board; long base-model notes became numbered footnotes under the ranking; a short expandable note explains the split.
- v1.7.0 (2026-10-06): Leaderboard split. /jev-models ranks open-weights systems we ran on our own hardware, with Jev 1.13.0 as an unranked reference row and a “Show API offerings” switch; hosted API offerings are ranked on the new /jev-models/api board. No score was recomputed: ranks are the published order filtered to each board. Adds the GPU cost What-If.
- v1.6.1 (2026-10-05): Hosted APIs (fastino-gliner-2-5-decide, jev-1.13.0, sage-1.3.0, wity-1, wity-1-always, wity-1-off) answer the full 1,500-item set instead of the 600-item API subset and are no longer equated; self-hosted rows unchanged; whole v1.6.0 sealed draw retired. Cost per 1,000 decisions is measured on one common item set (all items except the 23 outside the Jev reference input range) for every token-priced row.
- v1.6.0 (2026-10-05): Rotating item draw (1,200 sealed + 300 public); hosted APIs on API subsets equated to the self-hosted scale; /jev-models/v1.6.0 stays available unchanged.
- Earlier releases: v1.5.7, v1.5.6, v1.5.5 and the historical boards below.
System types (colours)
- Jev — reference (TypeSafe, closed) — The closed system JevBench is named after, shown as the reference.
- Closed API (weights not public) — Available through a hosted API; the weights cannot be downloaded.
- Open weights · LLM decoder — An autoregressive language model with public weights, including fine-tunes, merges and Jev rebuilds.
- Open weights · diffusion LM — A language model that generates by iterative denoising instead of token by token.
- Open weights · encoder / classifier — BERT-style encoders, NLI zero-shot classifiers and GLiNER-type models.
- Open weights · reranker — A cross-encoder or LLM reranker that scores options against the input.
- Base model control (no decision fine-tune, raw logits) — An official open checkpoint without any decision fine-tune, read out from raw logits, used as a floor.
- System (router / cascade / ensemble) — Several models combined at inference time.
Colours show architecture only; they never change a score or rank. Each class is assigned from cited evidence (config.json, model card, provider docs, or our own run record for hosted APIs).
Self-hosted rows use S 1,200 + P 300 + L1 (354 items) + L2 (33 items) + L3 (1,118 items), where answered. API rows re-run on A4/A5 use their 300/300-item sealed subset instead of S. Every row’s tag lists the pools it actually answered. L3 is a sealed language supplement (1,118 items), drawn 2026-10-07. Every language has at least 60 items in P ∪ L1 ∪ L2 ∪ L3. Header counts show the full S + P + L1 + L2 + L3 pool. Headline scores are unchanged; these raw cells are unequated, scored for language/category views only and outside the Composite. L3 items were written natively by Claude Sonnet 5.5, each solved blind and language-checked by GPT-6.1 Sol, with gold kept only when both agree or a second review confirms; no gold comes from Jev or any measured API. L3 is API-facing by design and is excluded from future headline draws. Rows with unfinished runs retain their actual coverage tags. L3 items were reviewed by an OpenAI model, so the L3-based language and category cells of OpenAI rows (OpenAI Decisions, GPT-6 Luna, GPT-5.6 Luna) carry that exposure; headline scores do not use L3. C1 adds English items for thin radar categories (everyday language, safety and the other use case). L4 adds 20 new Hindi items for Fastino GLiNER 2.5 Decide only. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing. Seventeen L4 requests succeeded; three HTTP 400 failures remain in the scored observations and do not count as completed coverage. L4 is outside headline and category scores and the common header counts. L5 adds 60 new items for OpenJev DeBERTa v3 Large only: 20 each in Arabic, Hindi and Greek. They were authored with OpenAI and independently solved and language-reviewed with Anthropic before sealing, then checked against the original 512-token native context before keeper selection. All 60 native requests succeeded. Previous failures remain scored and do not count as completed coverage. L5 is outside headline and category scores and the common header counts.
Method amendments
- 5 Oct 2026, v1.6.1 (Florian's decision): hosted/API systems now answer the same full item set as self-hosted systems (1,200 sealed + 300 public). Their earlier answers on the API subsets were reused and only the missing sealed items were sent; every item was logged in the exposure ledger before it was sent. API rows are no longer equated. Consequence: every v1.6.0 sealed item has now been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases; v1.6.2 onward draws fresh items from the reserve.
- 5 Oct 2026, v1.6.1 cost rule (Florian's decision): cost per 1,000 decisions is now measured on one common item set for every system: all items except the 23 very long items that lie outside the Jev reference's accepted input range. Before, a system that processed those items paid for their tokens while a system that refused them did not. Refusals still count against Intelligence exactly as before. Hosted APIs with a per-token tariff are priced from their measured tokens on this common set; carried per-decision prices of self-hosted systems are unchanged. Effect: Sage 1.3.0 USD 0.0247 per 1,000 decisions on the common set (0.0650 on all 1,500 items); Jev 1.13.0 0.0323 (unchanged).
The v1.6.1 amendments above say hosted APIs are no longer equated. Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
Rotating item sets
Each release draws fresh sealed decisions from a larger reserve. Self-hosted open-weights models (run offline on our own GPU pods or Sandy) answer S and P; externally hosted models answer the same full set (S and P) as the self-hosted systems since v1.6.1, so hosted APIs are no longer equated.
| Set | Items | Choice | Noul | Score | Answered by |
|---|---|---|---|---|---|
| S · Sealed release draw | 1,200 | 600 | 300 | 300 | self-hosted systems (run offline on our own GPU pods or Sandy) and, from v1.6.1, externally hosted APIs |
| A · API subset (part of S; used by the v1.6.0 hosted-API measurements) | 300 | 150 | 75 | 75 | every system (inside S since v1.6.1); retired for future draws |
| P · Public set | 300 | 150 | 75 | 75 | every system |
Not in the table above (frozen release data): API set A4 · 300 never-used sealed items, answered together with P by the 18 API endpoints re-run on 6 Oct 2026 (the 17 ranked offerings and the classifier.dev wrapper); retired after that run.
- Self-hosted systems: 1,500 items (S 1,200 + P 300). Hosted APIs: 1,500 items (S 1,200 + P 300), the same as self-hosted systems; A (300) is the part of S that earlier API measurements used. Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
- A sealed item is scored in at most three releases, then retired. An item used in one release is not drawn again in the next. If coverage minimums cannot be met, the release waits for newly reviewed items.
- The selection seed is committed (SHA-256) before any inference, and the draw is a deterministic function of policy, seed and item id.
API-exposure rule
- An external exposure is any sealed item sent to an endpoint we do not control: closed APIs, and open-weights models reached through third-party hosts, routers or a submitter's endpoint. Timeouts count as exposure.
- From v1.6.1 hosted models receive the full item set (S plus P); items they had already answered were reused and only the missing sealed items were sent. Hosted models are re-measured at most once every three refresh releases unless a verified new model version ships. Between measurements they keep their last score with its measurement date.
- Every externally sent sealed item is logged before dispatch. An item exposed to a provider is never scored again for that provider; once two different providers have received it, it retires for everyone. In v1.6.0 Jev 1.13.0 and Fastino GLiNER-2.5-Decide both received the same A, so A retires globally. With the v1.6.1 amendment every v1.6.0 sealed item has been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases.
Public-versus-sealed gap penalty
Intelligence is half public, half sealed. A system whose public score exceeds its sealed score by more than the field-median gap (G_med = 2.6 points) plus 8 points loses one Intelligence point per excess point. Hosted APIs are compared on P versus A against the same self-hosted systems' P-versus-A gap (-1.2 points).
Hosted APIs (not equated since v1.6.1)
Hosted APIs that answered the full set are scored exactly like self-hosted systems; no equating offset is applied to them. Category and language values stay raw. Exception on the live boards: API rows re-run on the fresh API set A4 (v1.7.7) answered A4 ∪ P (600 items) and are equated with the A2/A3 supplement method; OpenAI Decisions the same way on A5 ∪ P (v1.7.10; see Open weights and API offerings above).
Headline and Composite
Capability = mean(Intelligence, Calibration) for systems within twice the Jev 1.13.0 cost and median latency (the official caps; the sliders change only your view). The Composite (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost with the v1.5 low-axis gates; it remains secondary. Request types Choice, Noul and Score weigh equally; tiers weigh easy 0.1, standard 0.2, hard 0.4, judge 0.3.
Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1,000 decisions); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s.
Costs of v1.6-measured systems carry each system's published v1.5.4 cost per 1,000 decisions (pricing rules unchanged; v1.6 item lengths differ) unless the row says otherwise; Fastino's is an estimate from its published tariff and measured tokens.
Noul decisiveness and Score baseline (addendum B)
Scored with method option B (scorer setting O1S), selected on 3 Oct 2026 after the v1.6 scores were known and disclosed as a post-results change. Each split × type competence is clipped at 0 before the type weighting, so a type answered no better than chance counts as chance instead of negative. The Score chance baseline is the error of always predicting the mid-scale level, so a flat know-nothing distribution earns about 0. A Score cell whose golds all sit at mid-scale keeps the v1.5 random-level baseline; this only occurs in small breakdown and bootstrap cells. Calibration is unchanged. This run: G_med = 2.57 points.
Failed requests and very long items
23 items of this draw are very long (about 77,000 to 81,000 input tokens). Systems with a shorter context window refuse them; Jev-Omni runs out of GPU memory on them on its listed RTX 6000 (48 GB) recipe; decider-4b v2's server truncates them to about 32,800 tokens and answers; Plumb-4B reads them in full. Each outcome is the system's own recipe and is scored as such. A failed, refused or unparseable answer counts as wrong for Intelligence and stays in the denominator; it does not enter Calibration (the v1.5 rule, applied to every system). Failed answers per system: Sage 1.3.0 0/1500 · H2O-Lightning-4B v1.1 0/1500 · Mercury Decide 23/1500 · decisio v0.8.0 on gemma-4-31B-it 23/1500 · Jev 1.13.0 23/1500 · Quyet-1.0-Large 0/1500 · decider-12b v2 0/1500 · wity-1 (Wity, reasoning auto) 0/1500 · decider-12b v1, stock Gemma-4-12B-it 0/1500 · torchcast-decision-12b 23/1500 · Winnow-12B Q8 23/1500 · deck-31B 23/1500 · Cygnet 23/1500 · decisio v0.8.0 on gemma-4-12B-it 23/1500 · decisio v0.9.0 on gemma-4-12B-it 23/1500 · Jev-Omni 23/1500 · Xor 26B-A4B 0/1500 · Bobcat Flash 1.2 0/1500 · SPX-CD Flash 23/1500 · decider chat on Gemma-4-31B-it 22/1500 · Hopper 12B trained 23/1500 · Surogate Rune 26B-A4B v3 23/1500 · GEV-26B-Decide 23/1500 · SPX-CD-Omni 23/1500 · Diffusion Jev 26/1500 · René-1 31B FP8 23/1500 · Aplomb 1 0/1500 · Bespoke Nimble 9B v3 23/1500 · APUS-OpenJev-v1-9B 23/1500 · Plumb-4B 0/1500 · SPX-CD Pro 23/1500 · Messier One v0.2 0/1500 · spark-s1-4b-v6 0/1500 · swanOne 23/1500 · Quyet-1.0-Medium 0/1500 · lev 25/1500 · NInfer Qwen3.8-Flash-Next mixed 23/1500 · jev-local 23/1500 · decider-4b v2 0/1500 · ClassOne Qwen 3.5 9B 23/1500 · JevK5 v0.3 23/1500 · metask-jev-4b 23/1500 · deck-4B v1.0 23/1500 · Clef-Flash 0/1500 · Decision-4B 23/1500 · janus 4B 23/1500 · Qwen3.5-9B Jev-like data-mix v2 23/1500 · JevK5 v0.2.0 23/1500 · JevOne 0/1500 · Kahn1 4B 23/1500 · TypeCastLM 1.4.0 0/1500 · Hopper 0/1500 · Imajev-4B 23/1500 · Malkuth-4B 23/1500 · jqv 23/1500 · JPT-35B-A3B 23/1500 · decider-35b-a3b 0/1500 · CoCo-Decision-4B 23/1500 · Manchego v2.1 23/1500 · Decision 4B v1.2 23/1500 · Raw Qwen3 4B Instruct 2507 direct logits 23/1500 · JEV-27B 23/1500 · torchcast-decision-27b 0/1500 · Bespoke Nimble 9B 23/1500 · reflex 4B 0/1500 · Perplexity Decider v1.1 27B 23/1500 · JADE 392/1500 · typecastlm 0/1500 · Jebadiah 27B 23/1500 · Decision 2.0 Vega 27B 23/1500 · Decision 4B v1.1 23/1500 · AutoJev-27B 23/1500 · Eikos-27B 23/1500 · NInfer Qwen3.8-27B NVFP4 23/1500 · OpenJev 23/1500 · Open-Jev 9B 23/1500 · Standard One 8B 23/1500 · Clef 0/1500 · JEV Qwen3.5-9B Base NVFP4 23/1500 · SemIf, formerly OpenJev 23/1500 · Bev / Bonsai 27B 23/1500 · LitJev 23/1500 · local-jev Qwen3.5-4B 0/1500 · system-one 23/1500 · kev 8B 25/1500 · Decision 2B 23/1500 · OpenSourceJev 23/1500 · open-alternative-jev 0/1500 · Raw Qwen3 8B direct logits 23/1500 · decider-2b 0/1500 · Nemotron Diffusion 8B 23/1500 · kev 4B 29/1500 · Kev 27B 23/1500 · Malkuth-2B 23/1500 · Qwen3-Reranker-4B 0/1500 · SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool) 23/1500 · ZeroEntropy zerank-2 0/1500 · EXAONE-4.0-1.2B-JEV v0.3 23/1500 · Open-Jev 2B 23/1500 · Open-Jev 27B v1.1 23/1500 · Raw Phi-4 mini direct logits 23/1500 · SimpleJev 23/1500 · smalljev semantic-v9 0/1500 · kev 0.6B 30/1500 · jul fast 2/1500 · OpenDecision 0/1500 · ClassOne Gemma 4 E2B 23/1500 · Qwen3.5-0.8B Decision Model 0/1500 · Fastino GLiNER-2.5-Decide 31/1500 · Deem 0.8B v1 23/1500 · Decision Fast 23/1500 · Quyet-1.0-Small-EN 0/1500 · Laya typed-decisions 0/1500 · WaterSheep 0/1500 · openJev Verdict 1.4 0/1500 · Raw Qwen3 0.6B direct logits 23/1500 · lev-350m 23/1500 · Raw Qwen3 1.7B direct logits 23/1500 · kev 0.5B 32/1500 · openJev Verdict 0/1500 · Quyet-1.0-Small 0/1500 · watt-flash-0.1 23/1500 · Quyet-1.0-Tiny 0/1500 · open-jev-deberta-v3-large 73/1500 · verdict-small 0/1500 · Tacet Sonata 0/1500 · Laya 0/1500 · Mixedbread mxbai-rerank-base-v2 0/1500 · Laya multilingual 0/1500 · Certo v1 0/1500 · CLM-8B 0/1500 · BAAI bge-reranker-v2-m3 0/1500 · Alibaba GTE Reranker ModernBERT-base 0/1500 · Mirror 592/1500 · Open Jev JSON Canvas 23/1500 · APUS-OpenJev-v1-35B-A3B 23/1500 · BB-Qwen3.5-4B-LoRA 0/1500 · Seb-9B 23/1500 · wity-1 (Wity, reasoning always) 0/1500 · wity-1 (Wity, reasoning off) 0/1500 · classifier.dev 9/600 · Instinct 9/600 · Vansa-3.4 9/600 · GPT-6 Luna (low reasoning effort) 0/600 · Autoloops – Gemma 4 31B IT 9/600 · GPT-6 Luna (default medium reasoning effort) 4/600 · Fastino GLiDE 9/600 · SimpleJev Qwen3.6-35B-A3B 9/600 · GPT-5.6 Luna 0/600 · Instinct Dual 4B 0/600 · Gemini 3.1 Flash-Lite 1/600 · system-one-open 0/600 · openjev-sglang 9/600 · DeepSeek V4.1 Flash 5/600 · SimpleJev Qwen3.8-27B 9/600 · decision-machine-1 9/600 · JevAct 21/600 · Qwen3.8 27B 8/600 · OpenAI Decisions 2/600 · d1 0/1500 · Microsoft-Decision-1 0/0 · decisor-4b 0/1500 · Jeff-1.0-Large 21/1500 · Metask rain 4B 0/1500 · metask-jev-rain-12B 0/1500 · ryotide_qwen9 0/1500 · Wald 4B 0/1500 · Blink v0.3 26B-A4B 29/1500 · Clef-omni 0/1500 · Gutsy 0.8B v0.3 29/1500 · Liquid AI d1-3B 0/1500 · Liquid AI d1-omni-600M 0/1500 · RSI-Jev v6.1-VL 27B 29/1500 · Standard One 8B SH 29/1500 · Vega 0.8B 0/1500 · Vega 4B 0/1500 · Wald 4B v2.1 0/1500.
Calibration basis
Systems that return a full probability distribution are calibrated on all components (top-label error, plus distribution distance for Choice and ranked-probability error for Score). Fastino GLiNER-2.5-Decide returns a single confidence value, so its Calibration is the top-label error only and is not like-for-like with full-distribution systems.
Overnight full re-measure (4–5 Oct 2026)
Every system with a reproducible recipe was re-run on the v1.6.0 pool overnight with the same pinned inputs and scorer (method option B / O1S). This page uses scoring round score-v161-3 (v1.6.1) (2026-10-06 00:39:36 UTC). Only complete runs (1,500 items self-hosted, the full API input for hosted APIs) are ranked; partial runs are never ranked, and systems not yet re-measured keep their dated v1.5.x score in the separate table.
- Scores are the official v1.6.0 scorer (score_v16.py, method option B / O1S, bootstrap B = 1,000) over complete outputs only; each scoring round is kept separately.
- Hosted APIs are scored on the same full item set as self-hosted systems (S 1,200 + P 300, v1.6.1) and are not equated.
- GPU-class deviations: reproducible recipes used the hardware listed per row, including H100 for large fast-lane decoders and RTX 6000 for the baseline; hardware differences remain in measured latency. The standard x2 + 0.15 s self-hosted adjustment is an assumption, not a hardware normalization.
- Wity auto remains outside the latency cap and stays ranked in Composite A. OFF and ALWAYS are unranked variants of the AUTO main row.
- 6 further measured candidates await a separate publication decision.
- Sage 1.3.0 (Levanto Labs) was measured on 5 Oct from Sandy (Helsinki), text only. Cost uses the Levanto list tariff (USD 0.05/M input, USD 10/M output; levanto.ai/pricing, read 5 Oct) over all 600 answered rows: USD 0.0766/1,000 decisions. Superseded by the v1.6.1 cost rule (common item set): USD 0.0247/1,000 decisions, inside the Jev-class cost cap.
- A fresh Monday Jev 1.13.0 re-check on A3 measured p50 0.239 s and equated Capability 77.5 versus the official 76.5, within the confidence interval; the official v1.6.0 Jev row is retained.
Display note on the overnight text above (frozen release data). Exception since v1.7.7: the 17 API rows re-run on A4 ∪ P (600 items) are equated (+2.57 Intelligence, +3.35 Calibration); see “Full API re-run (A4, v1.7.7)” under Open weights and API offerings. Full-set API rows (Jev, Sage, wity-1, Fastino GLiNER-2.5-Decide, d1 (Liquid AI)) are not equated. OpenAI Decisions (gpt-6-luna) ran later on the fresh sealed API set A5 ∪ P (600 items; A4 had already been sent to that provider) and is equated the same way with A5's own offsets (+5.78 Intelligence, +6.15 Calibration, same pool).
Supplementary API draws A2 and A3
History: hosted APIs were first measured on API subsets (A, then A2 and A3) and equated to the self-hosted scale in v1.6.0. From v1.6.1 they answer the same full set as self-hosted systems, so no equating is applied; the earlier subsets A, A2 and A3 are retired.
Per-model exposure counts (hosted and author-hosted endpoints)
| System (provider) | v1.6 sealed items sent | Scored sealed set | Status |
|---|---|---|---|
| wity-1-auto (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-off (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-always (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| jev-1.13.0 (typesafe) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| fastino-gliner-2-5-decide (fastino) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| gpt-6-luna (openai) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| gpt-6-luna-low (openai) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| gpt-5.6-luna (openai) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| gemini-3.1-flash-lite (google) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| deepseek-flash (deepseek) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| qwen3.8-27b (chutes) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| decision-machine-1 (milliseconds.ai) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| classifier-dev-fast (classifier.dev) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| vansa-3.4 (vansa) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| kushal-gemma4-31b-it-autoloops (autoloops) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| system-one-open (modal via modal) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| openjev-sglang (modal via modal) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| instinct (zoowork) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| instinct-dual-4b (zoowork) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| simplejev-qwen3.8-27b (featherless) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| simplejev-qwen3.6-35b-a3b (featherless) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| jevact (jevact) | 300 | A4 (retired after this run) | measured 6 Oct 2026 on A4 ∪ P (v1.7.7) |
| Sage 1.3.0 (Levanto Labs) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
Revision history
- v1.6.1 · 2026-10-05 · Hosted APIs (fastino-gliner-2-5-decide, jev-1.13.0, sage-1.3.0, wity-1, wity-1-always, wity-1-off) answer the full 1,500-item set instead of the 600-item API subset and are no longer equated; self-hosted rows unchanged; whole v1.6.0 sealed draw retired. Cost per 1,000 decisions is measured on one common item set (all items except the 23 outside the Jev reference input range) for every token-priced row.
- v1.6.0 · 2026-10-05 · Rotating item draw (1,200 sealed + 300 public); hosted APIs on API subsets equated to the self-hosted scale; /jev-models/v1.6.0 stays available unchanged.
Provenance
Measured on the v1.6 pool (run completion day, UTC): 2026-10-02: Winnow-12B Q8, Cygnet, Jev-Omni, Plumb-4B, swanOne, NInfer Qwen3.8-Flash-Next mixed, jev-local, decider-4b v2, JevK5 v0.3, metask-jev-4b, Qwen3.5-9B Jev-like data-mix v2, JevK5 v0.2.0, JevOne, Hopper, Imajev-4B, Malkuth-4B, jqv, decider-35b-a3b, Decision 4B v1.2, Raw Qwen3 4B Instruct 2507 direct logits, Bespoke Nimble 9B, reflex 4B, typecastlm, Decision 4B v1.1, AutoJev-27B, Eikos-27B, NInfer Qwen3.8-27B NVFP4, OpenJev, Open-Jev 9B, Standard One 8B, JEV Qwen3.5-9B Base NVFP4, SemIf, formerly OpenJev, LitJev, local-jev Qwen3.5-4B, system-one, kev 8B, OpenSourceJev, open-alternative-jev, Raw Qwen3 8B direct logits, decider-2b, kev 4B, Malkuth-2B, Qwen3-Reranker-4B, ZeroEntropy zerank-2, Open-Jev 2B, Raw Phi-4 mini direct logits, SimpleJev, smalljev semantic-v9, kev 0.6B, OpenDecision, Qwen3.5-0.8B Decision Model, openJev Verdict 1.4, Raw Qwen3 0.6B direct logits, Raw Qwen3 1.7B direct logits, kev 0.5B, openJev Verdict, open-jev-deberta-v3-large, verdict-small, Laya, Mixedbread mxbai-rerank-base-v2, Certo v1, CLM-8B, BAAI bge-reranker-v2-m3, Alibaba GTE Reranker ModernBERT-base, Mirror, Open Jev JSON Canvas · 2026-10-04: H2O-Lightning-4B v1.1, Quyet-1.0-Large, decider-12b v2, decider-12b v1, stock Gemma-4-12B-it, torchcast-decision-12b, Quyet-1.0-Medium, lev, Clef-Flash, Decision-4B, janus 4B, Manchego v2.1, Clef, Bev / Bonsai 27B, Nemotron Diffusion 8B, Open-Jev 27B v1.1, Deem 0.8B v1, Quyet-1.0-Small-EN, Laya typed-decisions, Quyet-1.0-Small, Quyet-1.0-Tiny, Laya multilingual · 2026-10-05: Sage 1.3.0, Jev 1.13.0, deck-31B, spark-s1-4b-v6, deck-4B v1.0, Decision 2B, Fastino GLiNER-2.5-Decide, Decision Fast, lev-350m · 2026-10-06: decisio v0.8.0 on gemma-4-31B-it, wity-1 (Wity, reasoning auto), decisio v0.8.0 on gemma-4-12B-it, Xor 26B-A4B, SPX-CD Flash, decider chat on Gemma-4-31B-it, Hopper 12B trained, Surogate Rune 26B-A4B v3, GEV-26B-Decide, SPX-CD-Omni, Diffusion Jev, Aplomb 1, Bespoke Nimble 9B v3, SPX-CD Pro, JPT-35B-A3B, JEV-27B, torchcast-decision-27b, Perplexity Decider v1.1 27B, JADE, Jebadiah 27B, Decision 2.0 Vega 27B, Kev 27B, SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool), wity-1 (Wity, reasoning always), wity-1 (Wity, reasoning off), classifier.dev, Instinct, Vansa-3.4, GPT-6 Luna (low reasoning effort), Autoloops – Gemma 4 31B IT, GPT-6 Luna (default medium reasoning effort), Fastino GLiDE, SimpleJev Qwen3.6-35B-A3B, GPT-5.6 Luna, Instinct Dual 4B, Gemini 3.1 Flash-Lite, system-one-open, openjev-sglang, DeepSeek V4.1 Flash, SimpleJev Qwen3.8-27B, decision-machine-1, JevAct, Qwen3.8 27B, OpenAI Decisions, d1 · 2026-10-07: Mercury Decide, decisio v0.9.0 on gemma-4-12B-it, Bobcat Flash 1.2, René-1 31B FP8, APUS-OpenJev-v1-9B, Messier One v0.2, ClassOne Qwen 3.5 9B, Kahn1 4B, TypeCastLM 1.4.0, CoCo-Decision-4B, EXAONE-4.0-1.2B-JEV v0.3, jul fast, ClassOne Gemma 4 E2B, WaterSheep, watt-flash-0.1, Tacet Sonata, APUS-OpenJev-v1-35B-A3B, BB-Qwen3.5-4B-LoRA, Seb-9B · 2026-10-09: decisor-4b, Jeff-1.0-Large, Metask rain 4B, metask-jev-rain-12B, Wald 4B · 2026-10-10: Microsoft-Decision-1, ryotide_qwen9, Blink v0.3 26B-A4B, Clef-omni, Gutsy 0.8B v0.3, Liquid AI d1-3B, Liquid AI d1-omni-600M, RSI-Jev v6.1-VL 27B, Standard One 8B SH, Vega 0.8B, Vega 4B, Wald 4B v2.1.
Aggregate files: results sha256 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1 · categories sha256 0159c4828b3e8620148af06848ef0a5ccfc7d7f0c11aec02460d844dc0c62689 · dated carry sha256 334e851944e92383ab3071e52b3012040c7023d0fe236395953169ebee8f0459 · live category union (data/raw/benchmarks/jevbench/v1.6/jevbench-v1.6.1-category-cells.json) sha256 3b491b5c29d9480cc11c12486c7f507140582f2445e23abcf236a33baa5d7eae. Scoring source sha256 e15347783b8bedeeddd9017a831608a957f5bf0386f05902c50161ce0468f911. The method, release data and carry artifact are independently hashable.
Release history
The board above is always the current merged board: every system with its latest measurement. Each release stays readable as published.
- JevBench v1.6.7 · 2026-10-10 · Regular draw, + Blink v0.3 26B-A4B (10 systems, same frozen field median). Merged into the live board.
- JevBench v1.6.6 · 2026-10-10 · Regular draw, + Clef-omni (9 systems); restated by v1.6.7.
- JevBench v1.6.5 · 2026-10-10 · Regular draw (v1.6-regular-20261010-a2): 8 new open-weights systems; restated by v1.6.6.
- JevBench v1.6.4 · 2026-10-10 · Fast-lane draw (v1.6-fastlane-20261009), completed: 6 systems incl. 2 wrappers. Merged into the live board.
- JevBench v1.6.3 · 2026-10-10 · Fast-lane draw, 5 systems; superseded by v1.6.4.
- JevBench v1.6.2 · 2026-10-09 · Fast-lane draw, first 4 systems; superseded by v1.6.4.
- JevBench v1.6.1 · 2026-10-06 · Full re-measure on the v1.6.0 draw: hosted APIs on the full set. The reference scale of the live board.
- JevBench v1.6.0 · 2026-10-05 · Rotating sealed item sets, API-exposure rule, Noul decisiveness, language view, dated carry.
- JevBench v1.5.7 · 2026-10-04 · Last v1.5 point release.
- JevBench v1.5.6 · 2026-10-03 · v1.5 point release.
- JevBench v1.5.5 · 2026-10-02 · v1.5 point release.
- JevBench v1.5.0–v1.5.4 · 2026-09-28 · v1.5 method (Composite, Capability Score, Jev-class caps); v1.5.1–v1.5.4 on 29 Sep.
- JevBench v1.4–v1.4.2.2 · 2026-09-23 · v1.4 series.
- JevBench v1.0 · 2026-09-19 · First JevBench board.
Fresh full-set API addenda
These separately sourced measurements use their own accepted sealed draw. Frozen historical exports remain unchanged. All 50 topic, use-case and language labels are retained; official cells need 15 items and numeric radar spokes need 30.
Valid answers and completed coverage are shown separately. An accepted model/input refusal can count toward coverage while remaining a failed scored answer; authentication, rate-limit and service failures do not.
Microsoft-Decision-1 (Azure Foundry) · all 50 cells
Measured 2026-10-10 · O1S gap reference G_med 7.0881 score points · Aggregate results
Historical comparators keep their own measurement dates, draws and reference values. This addition does not remeasure them on the new pool.
Languages
| Category | Items | Valid answers | Completed coverage | Errors | Official competence | Sample status | Types / pool |
|---|---|---|---|---|---|---|---|
| English | 1281 | 1260 | 1281 | 21 | 56.83 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| German | 97 | 97 | 97 | 0 | 51.36 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| French | 80 | 80 | 80 | 0 | 49.75 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Spanish | 94 | 94 | 94 | 0 | 46.21 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Italian | 79 | 79 | 79 | 0 | 45.37 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Portuguese | 85 | 85 | 85 | 0 | 58.75 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Dutch | 78 | 78 | 78 | 0 | 46.12 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Danish | 77 | 77 | 77 | 0 | 49.39 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Swedish | 73 | 73 | 73 | 0 | 42.01 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Norwegian | 64 | 64 | 64 | 0 | 32.59 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Finnish | 72 | 72 | 72 | 0 | 23.96 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Polish | 96 | 96 | 96 | 0 | 36.65 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Czech | 73 | 73 | 73 | 0 | 38.58 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Greek | 73 | 73 | 73 | 0 | 33.36 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Turkish | 74 | 74 | 74 | 0 | 42.68 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Ukrainian | 70 | 70 | 70 | 0 | 38.45 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Arabic | 76 | 76 | 76 | 0 | 44.21 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Hindi | 85 | 85 | 85 | 0 | 33.80 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Chinese | 69 | 68 | 68 | 1 | 53.79 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Japanese | 85 | 85 | 85 | 0 | 54.97 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Korean | 75 | 75 | 75 | 0 | 33.10 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Indonesian | 70 | 70 | 70 | 0 | 57.30 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Mixed-language | 79 | 79 | 79 | 0 | 88.32 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
Subject topics
| Category | Items | Valid answers | Completed coverage | Errors | Official competence | Sample status | Types / pool |
|---|---|---|---|---|---|---|---|
| Math & numbers | 459 | 448 | 459 | 11 | 20.68 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Coding & software | 277 | 277 | 277 | 0 | 57.43 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Rules, policy & law | 1600 | 1599 | 1599 | 1 | 53.18 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Finance & commerce | 236 | 236 | 236 | 0 | 56.25 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Support & operations | 167 | 157 | 167 | 10 | 56.25 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Everyday language | 152 | 152 | 152 | 0 | 75.13 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Safety & security | 114 | 114 | 114 | 0 | 53.57 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
Use cases
| Category | Items | Valid answers | Completed coverage | Errors | Official competence | Sample status | Types / pool |
|---|---|---|---|---|---|---|---|
| Search and retrieval | 97 | 97 | 97 | 0 | 85.32 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Scientific discovery | 97 | 97 | 97 | 0 | 36.91 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Model routing | 400 | 400 | 400 | 0 | 64.30 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| LLM guardrails | 101 | 101 | 101 | 0 | 69.68 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Semantic code linting | 96 | 96 | 96 | 0 | 19.83 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Feature extraction for predictive modeling | 96 | 96 | 96 | 0 | 36.89 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Recruiting | 98 | 98 | 98 | 0 | 29.70 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Lead generation | 114 | 114 | 114 | 0 | 63.85 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Customer support | 249 | 240 | 249 | 9 | 60.37 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Insurance claims | 107 | 107 | 107 | 0 | 42.83 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Financial crime | 116 | 115 | 115 | 1 | 28.96 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Legal and compliance | 471 | 471 | 471 | 0 | 63.42 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| E-commerce marketplaces | 114 | 114 | 114 | 0 | 52.63 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Moderation and trust and safety | 98 | 98 | 98 | 0 | 67.15 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Advertising | 107 | 107 | 107 | 0 | 37.00 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Gaming | 90 | 90 | 90 | 0 | 25.06 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Risk assessment | 126 | 126 | 126 | 0 | 51.21 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Demand forecasting | 88 | 88 | 88 | 0 | 0.00 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Graphs and knowledge graphs | 98 | 98 | 98 | 0 | 35.49 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
| Other | 242 | 230 | 242 | 12 | 47.10 | radar sample minimum met | choice, noul, score / S1200+P300+L1+L2+L3; full native Foundry union |
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
Open this section to load the earlier public-only board and diagnostics.