JevBench v1.6.1 — Jev alternatives ranking
JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
Release v1.6.1 · 1,500 decisions per self-hosted system and 1,500 per hosted API · 92 ranked systems · 28 systems retain a separately dated v1.5.x score · only system-level aggregates are published · aggregate results JSON · SHA-256 5d4567d5e5acd945d17dd082adbbd6188d38174ca523b0fa6b3b02c6cc2dc5b1
Share this version · View live board · API leaderboard · Previous release: JevBench v1.6.0
Making decisions from images? Explore Image JevBench v0.1.5 and compare its systems.
JevBench v1.6.1 · headline
JevBench Capability Score
Capability ranking of Jev-class systems
Capability Score averages Intelligence and Calibration. Quyet-1.0-Large leads the Jev-class systems with 81.7.
Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘
Adjust cost / latency caps · 2× official
ScoreCCost
ScoreCap.Capability
Score$/1k$/1k decisions
- 1Quyet-1.0-Large73.450.581.7$0.045*
- 2Sage 1.3.0API65.658.278.6$0.025*
- 3deck-31B73.049.777.6$0.048*
- 4Jev 1.13.0API63.654.777.1$0.032
- 5torchcast-decision-12b60.556.471.7$0.028*
- 6Jev-Omni55.556.171.3$0.029*
- 7Winnow-12B Q859.556.671.2$0.028*
- 8Cygnet54.856.470.9$0.028*
- 9Plumb-4B43.063.165.4$0.017*
- 10decider-4b v240.164.564.9$0.015*
Show all 64 Jev-class systems (54 more)
- 11JevK5 v0.338.563.164.2$0.017*
- 12Clef-Flash42.246.863.8$0.059*
- 13deck-4B v1.038.663.563.6$0.017*
- 14JevK5 v0.2.036.363.162.8$0.017*
- 15Hopper34.762.362.4$0.018*
- 16Imajev-4B34.663.361.9$0.017*
- 17jqv33.951.261.9$0.042*
- 18metask-jev-4b39.257.761.8$0.026*
- 19lev41.161.861.7$0.019*
- 20Malkuth-4B33.855.861.1$0.030*
- 21Decision 4B v1.232.463.160.4$0.017*
- 22Quyet-1.0-Medium41.763.459.2$0.017*
- 23Manchego v2.132.564.358.3$0.015*
- 24jev-local41.958.957.5$0.024*
- 25Decision 4B v1.129.163.157.2$0.017*
- 26JEV Qwen3.5-9B Base NVFP429.347.657.1$0.056*
- 27typecastlm29.864.056.5$0.016*
- 28SemIf27.163.155.4$0.017*
- 29local-jev Qwen3.5-4B24.159.355.0$0.023*
- 30Decision 2B22.966.454.8$0.013*
- 31spark-s1-4b-v643.260.454.0$0.021*
- 32ZeroEntropy zerank-215.648.449.4$0.052*
- 33open-alternative-jev22.363.349.1$0.017*
- 34OpenSourceJev23.269.148.6$0.011*
- 35Malkuth-2B15.965.647.6$0.014*
- 36Nemotron Diffusion 8B20.455.547.5$0.030*
- 37Qwen3-Reranker-4B16.348.444.7$0.052*
- 38decider-2b20.764.944.6$0.015*
- 39BAAI bge-reranker-v2-m30.159.344.1$0.023*
- 40kev 4B19.565.844.0$0.014*
- 41Mixedbread mxbai-rerank-base-v21.160.343.9$0.021*
- 42lev-350m4.880.043.3$0.0046*
- 43Certo v10.296.943.2$0.0013*
- 44smalljev semantic-v98.760.843.0$0.020*
- 45Alibaba GTE Reranker ModernBERT-base0.069.341.7$0.011*
- 46Decision Fast6.180.140.6$0.0046*
- 47openJev Verdict 1.45.086.640.4$0.0028*
- 48Quyet-1.0-Small-EN6.189.139.5$0.0023*
- 49OpenDecision7.479.138.0$0.0050*
- 50verdict-small2.3100.037.6$0.00087*
- 51kev 0.6B7.680.136.9$0.0046*
- 52Raw Phi-4 mini direct logits13.753.235.3$0.036*
- 53Fastino GLiNER-2.5-DecideAPI7.045.834.6$0.064
- 54Raw Qwen3 4B Instruct 2507 direct logits34.063.434.6$0.017*
- 55kev 0.5B4.680.132.9$0.0046*
- 56Quyet-1.0-Small4.079.632.5$0.0048*
- 57Quyet-1.0-Tiny3.788.431.1$0.0024*
- 58Open Jev JSON Canvas61.249.330.6$0.049*
- 59openJev Verdict4.186.626.6$0.0028*
- 60CLM-8B0.150.520.1$0.045*
- 61Deem 0.8B v16.580.318.1$0.0045*
- 62Laya multilingual0.382.217.0$0.0039*
- 63Raw Qwen3 1.7B direct logits5.168.58.3$0.011*
- 64Raw Qwen3 0.6B direct logits5.677.68.2$0.0056*
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = 0.20, n = 92).
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- Open weights · LLM decoder (62)
- Open weights · diffusion LM (3)
- Open weights · encoder / classifier (13)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- green: ≤ reference
- amber: ≤ cap (2× reference)
- red: > cap
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).
- –wity-1API70.358.879.1$0.024Outside: latency 2.6× Jev (v1.5 reference)API price · eligibility checked at the developer's list price; at base-model pricing it would exceed the cost cap (2.70× Jev)
- –Eikos-27B64.428.777.2$0.24*Outside: cost 7.4× Jev (v1.5 reference)
- –OpenJev70.328.575.9$0.24*Outside: cost 7.5× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
- –Clef59.128.175.3$0.25*Outside: cost 7.7× Jev (v1.5 reference)
- –swanOne56.242.272.4$0.085*Outside: cost 2.6× Jev (v1.5 reference)
- –AutoJev-27B55.629.471.6$0.23*Outside: cost 7.0× Jev (v1.5 reference)
- –NInfer Qwen3.8-Flash-Next mixed49.042.569.7$0.082*Outside: cost 2.5× Jev (v1.5 reference)
- –JevOne46.039.868.9$0.10*Outside: cost 3.1× Jev (v1.5 reference)
- –NInfer Qwen3.8-27B NVFP451.829.168.0$0.23*Outside: cost 7.1× Jev (v1.5 reference)
- –Open-Jev 27B v1.152.514.468.0$0.72*Outside: cost 22.2× Jev (v1.5 reference), latency 2.8× Jev (v1.5 reference)
- –LitJev43.328.466.6$0.24*Outside: cost 7.6× Jev (v1.5 reference), latency 5.0× Jev (v1.5 reference)
- –decider-35b-a3b48.734.466.2$0.15*Outside: cost 4.8× Jev (v1.5 reference)
- –Bev / Bonsai 27B43.828.266.1$0.25*Outside: cost 7.6× Jev (v1.5 reference), latency 3.4× Jev (v1.5 reference)
- –Open-Jev 9B43.933.161.8$0.17*Outside: cost 5.3× Jev (v1.5 reference), latency 2.2× Jev (v1.5 reference)
- –Qwen3.5-9B Jev-like data-mix v241.945.760.9$0.065*Outside: cost 2.0× Jev (v1.5 reference)
- –reflex 4B30.963.159.0$0.017*Outside: latency 4.6× Jev (v1.5 reference)
- –Standard One 8B32.843.358.4$0.078*Outside: cost 2.4× Jev (v1.5 reference)
- –Bespoke Nimble 9B43.136.856.5$0.13*Outside: cost 4.0× Jev (v1.5 reference)
- –kev 8B29.340.450.0$0.097*Outside: cost 3.0× Jev (v1.5 reference)
- –Open-Jev 2B21.833.147.7$0.17*Outside: cost 5.3× Jev (v1.5 reference)
- –Laya typed-decisions5.582.544.2$0.0038*Outside: latency 3.0× Jev (v1.5 reference)
- –Qwen3.5-0.8B Decision Model7.379.542.6$0.0048*Outside: latency 18.3× Jev (v1.5 reference)
- –open-jev-deberta-v3-large2.877.635.4$0.0056*Outside: latency 5.7× Jev (v1.5 reference)
- –Laya1.884.932.5$0.0032*Outside: latency 2.9× Jev (v1.5 reference)
- –SimpleJev13.070.331.6$0.0098*Outside: latency 14.3× Jev (v1.5 reference)
- –Raw Qwen3 8B direct logits26.545.530.1$0.065*Outside: cost 2.0× Jev (v1.5 reference)
- –system-one29.345.029.7$0.068*Outside: cost 2.1× Jev (v1.5 reference)
- –Mirror0.089.325.8$0.0023*Outside: latency 2.4× Jev (v1.5 reference)
Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 64 of 92 systems qualify; the other 28, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed
Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- Open weights · LLM decoder (62)
- Open weights · diffusion LM (3)
- Open weights · encoder / classifier (13)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- faint = outside Jev-class
Model kind
Jev-class
Release
No system in this release reports an exact parameter count.
JevBench v1.6.1
JevBench Composite Score: 92 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Adjust weights ↓
The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis, and the column headings sort too.
Greener = stronger within its column.
92 of 92 systems, sorted by official rank, #1 first.
- 1Sage 1.3.0APInewClosed APIBase model: undisclosed74.0I 66C 92S 94K 58est.$0.025
- 2Jev 1.13.0APIJev referenceBase model: undisclosed71.5I 64C 91S 91K 55$0.032
- 3Quyet-1.0-Largenewadapter · 31BBase model: google/gemma-4-31B-itsource71.4I 73C 90S 87K 51est.$0.045
- 4wity-1APInewClosed APIBase model: undisclosedBase-model reference price (base undisclosed at the author's request): 44.0 (would be #11)70.8I 70C 88S 72K 59$0.024
- 5torchcast-decision-12bnewadapter · 12BBase model: google/gemma-4-12B-itsource69.9I 61C 83S 92K 56est.$0.028
- 6Winnow-12B Q8merge · 12B · GGUF Q8_0Base model: google/gemma-4-12B-itsource68.9I 60C 83S 87K 57est.$0.028
- 7deck-31Bnew31B · FP8Base model: google/gemma-4-31B-itsource68.7I 73C 82S 87K 50est.$0.048
- 8Cygnet12BBase model: google/gemma-4-12B-itsource68.6I 55C 87S 92K 56est.$0.028
- 9Jev-Omnifine-tuneBase model: google/gemma-4-12B-itsource67.7I 56C 87S 85K 56est.$0.029
- 10Plumb-4Bfine-tune · 4BBase model: alibiserikbay/JevK5source48.3I 43C 88S 94K 63est.$0.017
- 11spark-s1-4b-v6adapter · 4BBase model: Qwen3.5-4Bsource44.9I 43C 65S 87K 60est.$0.021
- 12swanOneNVFP4Base model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source43.8I 56C 89S 83K 42est.$0.085
- 13Quyet-1.0-Mediumnewadapter · 4BBase model: Qwen/Qwen3.5-4Bsource43.6I 42C 77S 89K 63est.$0.017
- 14levadapter · 4BBase model: Qwen/Qwen3.5-4Bsource42.2I 41C 82S 88K 62est.$0.019
- 15NInfer Qwen3.8-Flash-Next mixedLLM decoderBase model: Qwen3.8-Flash-Nextsource42.0I 49C 90S 89K 43est.$0.082
- 16jev-local9BBase model: Qwen/Qwen3.5-9Bsource41.3I 42C 73S 76K 59est.$0.024
- 17decider-4b v2adapter · 4BBase model: Qwen/Qwen3.5-4B-Basesource41.2I 40C 90S 92K 65est.$0.015
- 18JevK5 v0.3adapterBase model: Qwen/Qwen3.5-4Bsource37.4I 39C 90S 94K 63est.$0.017
- 19metask-jev-4b4BBase model: Qwen3.5-4Bsource37.3I 39C 84S 91K 58est.$0.026
- 20deck-4B v1.0newadapter · FP8Base model: alibiserikbay/JevK5 (built on Qwen/Qwen3.5-4B)source 1source 237.3I 39C 89S 90K 63est.$0.017
Show all 92 systems (72 more ranked, 0 more not ranked)
- 21Clef-Flashfine-tuneBase model: Qwen/Qwen3.5-9BsourceWorkers AI price scenario (latency unmeasured): 39.2 (would be #18)36.6I 42C 85S 88K 47est.$0.059
- 22Qwen3.5-9B Jev-like data-mix v2adapter · 9BBase model: Qwen/Qwen3.5-9Bsource33.6I 42C 80S 85K 46est.$0.065
- 23JevK5 v0.2.0adapter · 4BBase model: Qwen3.5-4Bsource32.1I 36C 89S 92K 63est.$0.017
- 24JevOneadapterBase model: Qwen/Qwen3.6-35B-A3Bsource31.1I 46C 92S 90K 40est.$0.101
- 25Hopperadapter · 4BBase model: Qwen/Qwen3.5-4Bsource28.9I 35C 90S 91K 62est.$0.018
- 26Imajev-4Badapter · 4BBase model: Qwen/Qwen3.5-4Bsource28.7I 35C 89S 91K 63est.$0.017
- 27Malkuth-4Badapter · 4BBase model: Qwen/Qwen3.5-4B-Basesource26.2I 34C 88S 90K 56est.$0.030
- 28jqv32BBase model: Qwen3-32Bsource25.4I 34C 90S 83K 51est.$0.042
- 29decider-35b-a3bfine-tune · 35B-A3BBase model: Qwen/Qwen3.5-35B-A3B-Basesource24.8I 49C 84S 91K 34est.$0.154
- 30Manchego v2.1adapter · 4BBase model: Qwen/Qwen3.5-4Bsource24.5I 33C 84S 91K 64est.$0.015
- 31Decision 4B v1.2adapter · 4BBase model: Qwen/Qwen3.5-4Bsource24.5I 32C 88S 94K 63est.$0.017
- 32Raw Qwen3 4B Instruct 2507 direct logitsoriginal · 4BBase model: Qwen/Qwen3-4B-Instruct-2507 (unmodified)source21.8I 34C 35S 91K 63est.$0.017
- 33Bespoke Nimble 9Badapter · 9BBase model: Qwen3.5-9Bsource21.1I 43C 70S 86K 37est.$0.128
- 34reflex 4Badapter · 4BBase model: Qwen/Qwen3.5-4Bsource20.6I 31C 87S 69K 63est.$0.017
- 35typecastlm4BBase model: Qwen3.5-4Bsource19.7I 30C 83S 93K 64est.$0.016
- 36Decision 4B v1.1adapter · 4BBase model: Qwen/Qwen3.5-4Bsource18.7I 29C 85S 94K 63est.$0.017
- 37AutoJev-27Bfine-tune · 27BBase model: Qwen/Qwen3.8-27Bsource18.5I 56C 88S 89K 29est.$0.226
- 38Eikos-27Badapter · 27BBase model: Qwen/Qwen3.8-27Bsource18.1I 64C 90S 88K 29est.$0.238
- 39NInfer Qwen3.8-27B NVFP427BBase model: Qwen3.8-27Bsource17.7I 52C 84S 90K 29est.$0.231
- 40OpenJev26B-A4BBase model: google/diffusiongemma-26B-A4B-itsource 1source 217.3I 70C 81S 73K 29est.$0.241
- 41Open-Jev 9Badapter · 9BBase model: Qwen/Qwen3.5-9Bsource17.0I 44C 80S 74K 33est.$0.170
- 42Standard One 8Badapter · 8BBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source16.9I 33C 84S 93K 43est.$0.078
- 43Cleffine-tuneBase model: Qwen/Qwen3.8-27BsourceWorkers AI price scenario (latency unmeasured): 29.4 (would be #25)16.8I 59C 92S 83K 28est.$0.249
- 44JEV Qwen3.5-9B Base NVFP49B · NVFP4Base model: ig1/Qwen3.5-9B-NVFP4source16.0I 29C 85S 94K 48est.$0.056
- 45SemIf4BBase model: Qwen/Qwen3.5-4Bsource15.5I 27C 84S 91K 63est.$0.017
- 46Bev / Bonsai 27B27B · GGUF PQ2_0Base model: Qwen/Qwen3.8-27Bsource11.7I 44C 89S 73K 28est.$0.247
- 47LitJev27BBase model: Qwen/Qwen3.8-27Bsource11.5I 43C 90S 68K 28est.$0.244
- 48local-jev Qwen3.5-4Bfine-tune · 4BBase model: Qwen/Qwen3.5-4Bsource11.4I 24C 86S 86K 59est.$0.023
- 49system-one8BBase model: Qwen3-8Bsource11.0I 29C 30S 91K 45est.$0.068
- 50kev 8Badapter · 8BBase model: Qwen/Qwen3-8B-Basesource10.6I 29C 71S 89K 40est.$0.097
- 51Decision 2Badapter · 2BBase model: openbmb/MiniCPM5-2Bsource10.3I 23C 87S 90K 66est.$0.013
- 52OpenSourceJevadapter · 4B · GGUF Q4_K_MBase model: Qwen3.5-4Bsource10.3I 23C 74S 77K 69est.$0.011
- 53open-alternative-jev4BBase model: Qwen3.5-4Bsource9.4I 22C 76S 91K 63est.$0.017
- 54Raw Qwen3 8B direct logitsoriginal · 8BBase model: Qwen/Qwen3-8B-Basesource9.3I 27C 34S 89K 46est.$0.065
- 55decider-2badapter · 2BBase model: Qwen/Qwen3.5-2B-Basesource7.7I 21C 69S 95K 65est.$0.015
- 56Nemotron Diffusion 8B8BBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource7.3I 20C 75S 94K 55est.$0.030
- 57kev 4Badapter · 4BBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.6.6I 20C 68S 90K 66est.$0.014
- 58Malkuth-2Badapter · 2BBase model: empero-ai/Qwen3.8-2B-Distillsource4.0I 16C 79S 92K 66est.$0.014
- 59Qwen3-Reranker-4BRerankerBase model: Qwen/Qwen3-4B-Basesource3.7I 16C 73S 82K 48est.$0.052
- 60ZeroEntropy zerank-2RerankerBase model: Qwen/Qwen3-4Bsource3.3I 16C 83S 83K 48est.$0.052
- 61Open-Jev 2Badapter · 2BBase model: Qwen/Qwen3.5-2Bsource3.2I 22C 74S 77K 33est.$0.170
- 62Open-Jev 27B v1.1adapter · 27BBase model: Qwen/Qwen3.8-27Bsource2.9I 52C 83S 73K 14est.$0.716
- 63Raw Phi-4 mini direct logitsoriginalBase model: microsoft/Phi-4-mini-instruct (unmodified)source2.5I 14C 57S 92K 53est.$0.036
- 64SimpleJev0.8BBase model: Qwen/Qwen3.5-0.8Bsource2.1I 13C 50S 59K 70est.$0.0098
- 65smalljev semantic-v9adapter · 2BBase model: openbmb/MiniCPM5-2B-Basesource0.8I 9C 77S 90K 61est.$0.020
- 66kev 0.6Badapter · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.5I 8C 66S 92K 80est.$0.0046
- 67OpenDecisionEncoder / classifierBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.5I 7C 69S 87K 79est.$0.0050
- 68Qwen3.5-0.8B Decision Modelfine-tune · 0.8BBase model: Qwen/Qwen3.5-0.8B-Basesource0.5I 7C 78S 56K 80est.$0.0048
- 69Fastino GLiNER-2.5-DecideAPInewClosed APIBase model: undisclosed · Hosted Fastino API (model id fastino/GLiNER-2.5-Decide). Checked fastino.ai home page and the GLiNER2.5-Decide announcement blog: no statement of the hosted endpoint's base model. Open-weight HF cards fastino/GLiNER2.5-Decide (base_model fastino/gliner2-large-v1) and fastino/GLiNER2.5-Decide-1B (base_model fastino/gliner2-xl-0111) exist, but nothing public says which one (or a different build) the hosted API serves, so no base is claimed.0.3I 7C 62S 85K 46$0.064
- 70Deem 0.8B v1LLM decoderBase model: Qwen/Qwen3.5-0.8Bsource0.3I 7C 30S 85K 80est.$0.0045
- 71Decision Fastadapter · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.3I 6C 75S 91K 80est.$0.0046
- 72Quyet-1.0-Small-ENnewfine-tune · 153MBase model: answerdotai/ModernBERT-basesource0.3I 6C 73S 95K 89est.$0.0023
- 73Laya typed-decisionsfine-tune · 421MBase model: ModernBERT-largesource0.2I 5C 83S 71K 83est.$0.0038
- 74openJev Verdict 1.4Encoder / classifierBase model: knowledgator/gliclass-modern-base-v2.0source0.2I 5C 76S 81K 87est.$0.0028
- 75Raw Qwen3 0.6B direct logitsoriginal · 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.2I 6C 11S 94K 78est.$0.0056
- 76lev-350madapterBase model: LiquidAI/LFM2.5-350Msource0.2I 5C 82S 94K 80est.$0.0046
- 77Raw Qwen3 1.7B direct logitsoriginal · 1.7BBase model: Qwen/Qwen3-1.7B-Basesource0.1I 5C 12S 93K 69est.$0.011
- 78kev 0.5Badapter · 0.5BBase model: Qwen/Qwen2.5-0.5Bsource0.1I 5C 61S 92K 80est.$0.0046
- 79openJev Verdictfine-tuneBase model: knowledgator/gliclass-modern-base-v2.0source 1source 20.1I 4C 49S 79K 87est.$0.0028
- 80Quyet-1.0-Smallnewfine-tune · 328MBase model: aisingapore/SEA-LION-ModernBERT-300Msource0.1I 4C 61S 95K 80est.$0.0048
- 81Quyet-1.0-Tinynewdistilled · 183MBase model: jhu-clsp/mmBERT-smallsource0.1I 4C 58S 95K 88est.$0.0024
- 82open-jev-deberta-v3-largeadapterBase model: microsoft/deberta-v3-largesource0.0I 3C 68S 68K 78est.$0.0056
- 83verdict-smallfine-tune · 118MBase model: intfloat/multilingual-e5-smallsource0.0I 2C 73S 89K 100est.$0.0009
- 84Layafine-tune · 421MBase model: ModernBERT-largesource0.0I 2C 63S 73K 85est.$0.0032
- 85Mixedbread mxbai-rerank-base-v2RerankerBase model: Qwen2.5 (size not stated)source 1source 20.0I 1C 87S 93K 60est.$0.021
- 86Laya multilingualfine-tune · 322MBase model: mmBERT-basesource0.0I 0C 34S 78K 82est.$0.0039
- 87Certo v1Encoder / classifierBase model: ModernBERT-largesource0.0I 0C 86S 93K 97est.$0.0013
- 88CLM-8Bfine-tune · 8BBase model: Qwen/Qwen3-8Bsource0.0I 0C 40S 93K 51est.$0.045
- 89BAAI bge-reranker-v2-m3fine-tuneBase model: BAAI/bge-m3source0.0I 0C 88S 94K 59est.$0.023
- 90Alibaba GTE Reranker ModernBERT-basefine-tuneBase model: answerdotai/ModernBERT-basesource0.0I 0C 83S 94K 69est.$0.011
- 91MirrorEncoder / classifierBase model: undisclosed0.0I 0C 52S 73K 89est.$0.0023
- 92Open Jev JSON Canvas26B-A4BBase model: google/diffusiongemma-26B-A4B-itsource0.0I 61C 0S 86K 49est.$0.049
Adjust weights ↓
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- Open weights · LLM decoder (62)
- Open weights · diffusion LM (3)
- Open weights · encoder / classifier (13)
- Open weights · reranker (5)
- Base model control (no decision fine-tune, raw logits) (5)
- Striped bar = same system under the labelled alternative price assumption
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
- A: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (#2)
- B: Sage 1.3.0 — Closed API (weights not public) · Score 74.0 (#1)
The four score axes
Capability by subject topic
What each category means · items per category
- Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open / 612 sealed)
- Coding & software — code, SQL, repositories, developer tools and IT systems. 302 items (62 open / 240 sealed)
- Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open / 156 sealed)
- Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open / 72 sealed)
- Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open / 54 sealed)
- Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open / 35 sealed)
- Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open / 31 sealed)
Use cases (TypeSafe categories)
Low sample, n < 30 — indicative only
| Category (items) | A: Jev 1.13.0 | B: Sage 1.3.0 |
|---|---|---|
| Feature extraction for predictive modeling (19) | 22.5 n=19 | 34.4 n=19 |
| Graphs and knowledge graphs (18) | 49.8 n=18 | 74.1 n=18 |
| LLM guardrails (17) | 44.1 n=17 | 65.7 n=17 |
| Search and retrieval (17) | 80.5 n=17 | 83.3 n=17 |
| Lead generation (17) | 67.5 n=17 | 75.5 n=17 |
| Semantic code linting (16) | 12.2 n=16 | 22.9 n=16 |
| Demand forecasting (16) | 0.0 n=16 | 0.0 n=16 |
No published value for either system (under 15 answered items, or no per-category values): Gaming (14), Scientific discovery (14), Recruiting (13), Advertising (11).
Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.
What each category means · items per category
- Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open / 356 sealed)
- Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open / 313 sealed)
- Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open / 112 sealed)
- E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open / 39 sealed)
- Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open / 37 sealed)
- Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open / 32 sealed)
- Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open / 30 sealed)
- Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open / 25 sealed)
Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 145 items (24 open / 121 sealed) — not a use case of its own, so it is counted but not drawn.
Competence per request type, open / sealed
Competence per tier — open set
Competence per tier — sealed set
How the categories were made
Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).
Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 v1.6 items was also labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod (same model, questions and taxonomy as v1.5); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) were checked by hand; item-group rules fix the systematic misses (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems saw S u P (1,500 items); hosted API systems saw only A u P (600 items), so their cells cover fewer items.
- Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
- Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
| Spoke | A: Jev 1.13.0 | B: Sage 1.3.0 |
|---|---|---|
| The four score axes | ||
| Intelligence | 63.6 | 65.6 |
| Calibration | 90.6 | 91.6 |
| Speed | 91.5 | 93.6 |
| Cost | 54.7 | 58.2 |
| Capability by subject topic | ||
| Rules, policy & law | 70.6 | 68.1 |
| Coding & software | 74.5 | 80.6 |
| Math & numbers | 30.5 | 37.9 |
| Support & operations | 34.9 | 60.9 |
| Finance & commerce | 60.3 | 66.3 |
| Everyday language | 51.5 | 47.4 |
| Safety & security | 35.2 | 65.3 |
| Use cases (TypeSafe categories) | ||
| Model routing | 78.2 | 83.2 |
| Legal & compliance | 81.3 | 70.3 |
| Customer support | 47.6 | 63.3 |
| E-commerce | 44.6 | 66.5 |
| Risk assessment | 42.1 | 33.7 |
| Insurance claims | 76.5 | 65.1 |
| Financial crime | 36.0 | 35.0 |
| Moderation | 53.4 | 70.5 |
| Competence per request type, open / sealed | ||
| Choice · open | 79.0 | 85.1 |
| Choice · sealed | 75.8 | 79.7 |
| Noul · open | 54.5 | 54.3 |
| Noul · sealed | 46.1 | 56.4 |
| Score · open | 63.3 | 57.7 |
| Score · sealed | 63.0 | 60.5 |
| Competence per tier — open set | ||
| Easy | 93.9 | 89.6 |
| Standard | 63.2 | 62.3 |
| Judge | 67.7 | 67.5 |
| Hard | 65.6 | 70.5 |
| Competence per tier — sealed set | ||
| Easy | 78.4 | 77.5 |
| Standard | 68.7 | 75.8 |
| Judge | 66.9 | 73.2 |
| Hard | 58.9 | 60.8 |
Self-hosted cells and hosted-API cells pool S 1,200 + P 300 (1,500 items); raw and unequated. Cells under 15 items are omitted. Sealed counts in the compare view refer to the sealed set S (1,200), which hosted APIs now answer in full.
Languages
Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: self-hosted systems over S 1,200 + P 300, and every hosted API that answered the full set the same way (1,500 items). Unequated and outside the Composite. Cells under 15 items are left empty. A dagger (†) marks every displayed cell with fewer than 30 answered items. English (1,160 items) is listed first; the other 21 languages and the mixed-language group share 340 items. In 14 further groups no system reaches the 15-item reporting minimum, so they get no column (items in the pool shown): Mixed-language (14), Hindi (13), Chinese (13), Turkish (12), Arabic (12), Ukrainian (12), Korean (11), Dutch (11), Czech (11), Swedish (10), Norwegian (9), Greek (9), Finnish (8), Indonesian (5) — 150 items, scored like every other item. This is the v1.6.1 main-pool breakdown. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.1.
| System | en1160 | pl31 | es31 | pt†29 | de†27 | da†22 | it†17 | ja†17 | fr†16 |
|---|---|---|---|---|---|---|---|---|---|
| Sage 1.3.0 · API | 68 | 47 | 84 | 77† | 80† | 35† | 61† | 54† | 69† |
| Jev 1.13.0 · API | 67 | 28 | 80 | 76† | 72† | 48† | 48† | 33† | 76† |
| Quyet-1.0-Large · self-hosted | 74 | 63 | 87 | 84† | 94† | 64† | 66† | 68† | 74† |
| wity-1 · API | 75 | 55 | 78 | 79† | 85† | 44† | 67† | 30† | 45† |
| torchcast-decision-12b · self-hosted | 62 | 56 | 75 | 73† | 85† | 43† | 56† | 36† | 79† |
| Winnow-12B Q8 · self-hosted | 61 | 56 | 78 | 76† | 79† | 35† | 45† | 29† | 79† |
| deck-31B · self-hosted | 76 | 72 | 83 | 89† | 97† | 67† | 72† | 82† | 94† |
| Cygnet · self-hosted | 52 | 47 | 82 | 75† | 77† | 34† | 35† | 46† | 77† |
| Jev-Omni · self-hosted | 52 | 74 | 78 | 77† | 56† | 40† | 40† | 47† | 54† |
| Plumb-4B · self-hosted | 42 | 43 | 78 | 61† | 73† | 35† | 22† | 5† | 46† |
| spark-s1-4b-v6 · self-hosted | 42 | 29 | 79 | 70† | 71† | 18† | 24† | 33† | 70† |
| swanOne · self-hosted | 57 | 48 | 86 | 68† | 85† | 26† | 45† | 42† | 91† |
| Quyet-1.0-Medium · self-hosted | 42 | 21 | 77 | 50† | 72† | 39† | 25† | 18† | 31† |
| lev · self-hosted | 40 | 32 | 66 | 57† | 45† | 4† | 31† | 3† | 36† |
| NInfer Qwen3.8-Flash-Next mixed · self-hosted | 49 | 46 | 78 | 76† | 68† | 27† | 43† | 27† | 85† |
| jev-local · self-hosted | 43 | 24 | 71 | 34† | 49† | 24† | 49† | 15† | 60† |
| decider-4b v2 · self-hosted | 39 | 19 | 79 | 39† | 61† | 11† | 43† | 25† | 2† |
| JevK5 v0.3 · self-hosted | 40 | 26 | 79 | 42† | 60† | 20† | 13† | 0† | 56† |
| metask-jev-4b · self-hosted | 39 | 24 | 72 | 61† | 51† | 22† | 41† | 25† | 40† |
| deck-4B v1.0 · self-hosted | 40 | 26 | 79 | 57† | 60† | 18† | 13† | 0† | 57† |
| Clef-Flash · self-hosted | 42 | 28 | 73 | 66† | 64† | 1† | 48† | 0† | 69† |
| Qwen3.5-9B Jev-like data-mix v2 · self-hosted | 42 | 27 | 66 | 44† | 57† | 0† | 31† | 0† | 70† |
| JevK5 v0.2.0 · self-hosted | 33 | 32 | 77 | 49† | 70† | 32† | 18† | 26† | 25† |
| JevOne · self-hosted | 46 | 35 | 76 | 61† | 72† | 29† | 42† | 16† | 36† |
| Hopper · self-hosted | 33 | 23 | 72 | 49† | 55† | 5† | 24† | 0† | 34† |
| Imajev-4B · self-hosted | 35 | 20 | 70 | 41† | 58† | 1† | 40† | 7† | 12† |
| Malkuth-4B · self-hosted | 35 | 19 | 61 | 41† | 47† | 18† | 46† | 11† | 31† |
| jqv · self-hosted | 37 | 45 | 58 | 46† | 61† | 0† | 40† | 23† | 35† |
| decider-35b-a3b · self-hosted | 49 | 38 | 73 | 47† | 64† | 18† | 22† | 0† | 47† |
| Manchego v2.1 · self-hosted | 32 | 15 | 54 | 43† | 66† | 15† | 49† | 5† | 9† |
| Decision 4B v1.2 · self-hosted | 32 | 16 | 73 | 36† | 67† | 5† | 41† | 0† | 31† |
| Raw Qwen3 4B Instruct 2507 direct logits · self-hosted | 32 | 37 | 58 | 30† | 55† | 11† | 24† | 9† | 38† |
| Bespoke Nimble 9B · self-hosted | 42 | 52 | 74 | 48† | 53† | 3† | 46† | 9† | 55† |
| reflex 4B · self-hosted | 31 | 10 | 76 | 57† | 51† | 1† | 22† | 20† | 17† |
| typecastlm · self-hosted | 27 | 23 | 55 | 28† | 64† | 24† | 27† | 23† | 28† |
| Decision 4B v1.1 · self-hosted | 29 | 12 | 68 | 46† | 55† | 0† | 33† | 1† | 44† |
| AutoJev-27B · self-hosted | 56 | 45 | 74 | 54† | 83† | 40† | 50† | 30† | 75† |
| Eikos-27B · self-hosted | 66 | 47 | 73 | 69† | 78† | 45† | 57† | 48† | 65† |
| NInfer Qwen3.8-27B NVFP4 · self-hosted | 51 | 51 | 84 | 85† | 75† | 28† | 52† | 46† | 92† |
| OpenJev · self-hosted | 76 | 32 | 76 | 89† | 93† | 29† | 82† | 31† | 66† |
| Open-Jev 9B · self-hosted | 44 | 34 | 67 | 56† | 78† | 29† | 59† | 0† | 21† |
| Standard One 8B · self-hosted | 35 | 45 | 75 | 40† | 53† | 12† | 46† | 17† | 50† |
| Clef · self-hosted | 62 | 36 | 77 | 80† | 84† | 34† | 70† | 46† | 76† |
| JEV Qwen3.5-9B Base NVFP4 · self-hosted | 27 | 11 | 66 | 39† | 62† | 3† | 14† | 0† | 31† |
| SemIf, formerly OpenJev · self-hosted | 27 | 5 | 49 | 39† | 57† | 15† | 31† | 18† | 35† |
| Bev / Bonsai 27B · self-hosted | 44 | 30 | 55 | 61† | 72† | 24† | 19† | 43† | 20† |
| LitJev · self-hosted | 45 | 43 | 75 | 69† | 78† | 43† | 35† | 33† | 52† |
| local-jev Qwen3.5-4B · self-hosted | 12 | 0 | 42 | 27† | 57† | 0† | 16† | 0† | 34† |
| system-one · self-hosted | 30 | 14 | 44 | 38† | 48† | 25† | 9† | 8† | 27† |
| kev 8B · self-hosted | 34 | 27 | 69 | 46† | 38† | 0† | 9† | 0† | 23† |
| Decision 2B · self-hosted | 13 | 9 | 42 | 25† | 49† | 4† | 0† | 0† | 0† |
| OpenSourceJev · self-hosted | 19 | 23 | 60 | 32† | 54† | 14† | 30† | 0† | 3† |
| open-alternative-jev · self-hosted | 21 | 23 | 63 | 31† | 47† | 13† | 36† | 17† | 14† |
| Raw Qwen3 8B direct logits · self-hosted | 27 | 19 | 53 | 25† | 49† | 0† | 19† | 20† | 36† |
| decider-2b · self-hosted | 23 | 26 | 44 | 19† | 53† | 10† | 41† | 0† | 4† |
| Nemotron Diffusion 8B · self-hosted | 14 | 35 | 45 | 17† | 56† | 0† | 38† | 9† | 44† |
| kev 4B · self-hosted | 23 | 12 | 63 | 27† | 37† | 0† | 24† | 0† | 4† |
| Malkuth-2B · self-hosted | 13 | 22 | 40 | 19† | 41† | 0† | 8† | 0† | 0† |
| Qwen3-Reranker-4B · self-hosted | 0 | 0 | 12 | 0† | 3† | 0† | 0† | 0† | 0† |
| ZeroEntropy zerank-2 · self-hosted | 0 | 0 | 25 | 0† | 5† | 0† | 0† | 0† | 0† |
| Open-Jev 2B · self-hosted | 18 | 22 | 49 | 19† | 65† | 6† | 15† | 0† | 0† |
| Open-Jev 27B v1.1 · self-hosted | 53 | 33 | 69 | 64† | 73† | 25† | 34† | 11† | 46† |
| Raw Phi-4 mini direct logits · self-hosted | 0 | 0 | 18 | 1† | 4† | 0† | 23† | 0† | 14† |
| SimpleJev · self-hosted | 9 | 14 | 36 | 17† | 51† | 6† | 5† | 0† | 0† |
| smalljev semantic-v9 · self-hosted | 0 | 3 | 3 | 0† | 1† | 0† | 11† | 0† | 0† |
| kev 0.6B · self-hosted | 0 | 8 | 0 | 0† | 13† | 0† | 0† | 0† | 6† |
| OpenDecision · self-hosted | 0 | 0 | 6 | 0† | 21† | 10† | 0† | 0† | 6† |
| Qwen3.5-0.8B Decision Model · self-hosted | 0 | 1 | 0 | 0† | 16† | 0† | 0† | 0† | 0† |
| Fastino GLiNER-2.5-Decide · API | 0 | 0 | 20 | 0† | 7† | 0† | 0† | 0† | 0† |
| Deem 0.8B v1 · self-hosted | 0 | 12 | 0 | 0† | 37† | 0† | 0† | 1† | 0† |
| Decision Fast · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Quyet-1.0-Small-EN · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 7† | 0† | 0† |
| Laya typed-decisions · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| openJev Verdict 1.4 · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Raw Qwen3 0.6B direct logits · self-hosted | 0 | 0 | 0 | 0† | 0† | 9† | 0† | 0† | 23† |
| lev-350m · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Raw Qwen3 1.7B direct logits · self-hosted | 0 | 0 | 0 | 26† | 0† | 0† | 0† | 7† | 16† |
| kev 0.5B · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| openJev Verdict · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 2† | 0† |
| Quyet-1.0-Small · self-hosted | 0 | 0 | 0 | 0† | 5† | 0† | 0† | 2† | 0† |
| Quyet-1.0-Tiny · self-hosted | 0 | 0 | 0 | 0† | 7† | 0† | 0† | 9† | 0† |
| open-jev-deberta-v3-large · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| verdict-small · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 2† | 0† | 0† |
| Laya · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Mixedbread mxbai-rerank-base-v2 · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Laya multilingual · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Certo v1 · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| CLM-8B · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 3† | 0† |
| BAAI bge-reranker-v2-m3 · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Alibaba GTE Reranker ModernBERT-base · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Mirror · self-hosted | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Open Jev JSON Canvas · self-hosted | 62 | 64 | 72 | 74† | 78† | 25† | 45† | 62† | 94† |
Intelligence gate and Noul decisiveness
The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.
| System | Intelligence | Choice | Noul | Score | Noul decisive rate | Accuracy among decisive |
|---|---|---|---|---|---|---|
| Sage 1.3.0 | 65.6 | native | native | native | 83 % | 95 % |
| Jev 1.13.0 | 63.6 | native | native | native | 79 % | 98 % |
| Quyet-1.0-Large | 73.4 | native | native | native | 91 % | 93 % |
| wity-1 | 70.3 | native | native | native | 91 % | 94 % |
| torchcast-decision-12b | 60.5 | native | native | native | 96 % | 86 % |
| Winnow-12B Q8 | 59.5 | native | native | native | 86 % | 92 % |
| deck-31B | 73.0 | native | native | native | 96 % | 94 % |
| Cygnet | 54.8 | native | native | native | 81 % | 92 % |
| Jev-Omni | 55.5 | native | native | native | 74 % | 94 % |
| Plumb-4B | 43.0below gate | native | native | native | 71 % | 91 % |
| spark-s1-4b-v6 | 43.2below gate | native | native | native | 84 % | 83 % |
| swanOne | 56.2 | native | native | native | 81 % | 93 % |
| Quyet-1.0-Medium | 41.7below gate | native | native | native | 76 % | 84 % |
| lev | 41.1below gate | native | native | native | 84 % | 83 % |
| NInfer Qwen3.8-Flash-Next mixed | 49.0below gate | native | native | native | 75 % | 93 % |
| jev-local | 41.9below gate | native | native | native | 85 % | 82 % |
| decider-4b v2 | 40.1below gate | native | native | native | 64 % | 92 % |
| JevK5 v0.3 | 38.5below gate | native | native | native | 66 % | 95 % |
| metask-jev-4b | 39.2below gate | native | native | native | 75 % | 88 % |
| deck-4B v1.0 | 38.6below gate | native | native | native | 68 % | 93 % |
| Clef-Flash | 42.2below gate | native | native | native | 68 % | 93 % |
| Qwen3.5-9B Jev-like data-mix v2 | 41.9below gate | native | native | native | 71 % | 91 % |
| JevK5 v0.2.0 | 36.3below gate | native | native | native | 62 % | 93 % |
| JevOne | 46.0below gate | native | native | native | 62 % | 96 % |
| Hopper | 34.7below gate | native | native | native | 56 % | 94 % |
| Imajev-4B | 34.6below gate | native | native | native | 52 % | 96 % |
| Malkuth-4B | 33.8below gate | native | native | native | 65 % | 92 % |
| jqv | 33.9below gate | native | native | native | 66 % | 94 % |
| decider-35b-a3b | 48.7below gate | native | native | native | 74 % | 91 % |
| Manchego v2.1 | 32.5below gate | native | native | native | 62 % | 90 % |
| Decision 4B v1.2 | 32.4below gate | native | native | native | 59 % | 94 % |
| Raw Qwen3 4B Instruct 2507 direct logits | 34.0below gate | native | native | native | 92 % | 73 % |
| Bespoke Nimble 9B | 43.1below gate | native | native | native | 77 % | 87 % |
| reflex 4B | 30.9below gate | native | native | native | 58 % | 92 % |
| typecastlm | 29.8below gate | native | native | native | 57 % | 90 % |
| Decision 4B v1.1 | 29.1below gate | native | native | native | 57 % | 94 % |
| AutoJev-27B | 55.6 | native | native | native | 70 % | 99 % |
| Eikos-27B | 64.4 | native | native | native | 81 % | 98 % |
| NInfer Qwen3.8-27B NVFP4 | 51.8 | native | native | native | 82 % | 89 % |
| OpenJev | 70.3 | native | native | native | 96 % | 93 % |
| Open-Jev 9B | 43.9below gate | native | native | native | 81 % | 86 % |
| Standard One 8B | 32.8below gate | native | native | native | 68 % | 90 % |
| Clef | 59.1 | native | native | native | 82 % | 94 % |
| JEV Qwen3.5-9B Base NVFP4 | 29.3below gate | native | native | native | 53 % | 89 % |
| SemIf, formerly OpenJev | 27.1below gate | native | native | native | 62 % | 90 % |
| Bev / Bonsai 27B | 43.8below gate | native | native | native | 70 % | 91 % |
| LitJev | 43.3below gate | native | native | native | 66 % | 98 % |
| local-jev Qwen3.5-4B | 24.1below gate | native | native | native | 35 % | 99 % |
| system-one | 29.3below gate | native | native | native | 92 % | 73 % |
| kev 8B | 29.3below gate | native | native | native | 77 % | 81 % |
| Decision 2B | 22.9below gate | native | native | native | 39 % | 94 % |
| OpenSourceJev | 23.2below gate | native | native | native | 49 % | 95 % |
| open-alternative-jev | 22.3below gate | native | native | native | 60 % | 87 % |
| Raw Qwen3 8B direct logits | 26.5below gate | native | native | native | 92 % | 71 % |
| decider-2b | 20.7below gate | native | native | native | 68 % | 78 % |
| Nemotron Diffusion 8B | 20.4below gate | native | native | native | 61 % | 77 % |
| kev 4B | 19.5below gate | native | native | native | 67 % | 82 % |
| Malkuth-2B | 15.9below gate | native | native | native | 53 % | 84 % |
| Qwen3-Reranker-4B | 16.3below gate | native | native | native | 25 % | 75 % |
| ZeroEntropy zerank-2 | 15.6below gate | native | native | native | 14 % | 87 % |
| Open-Jev 2B | 21.8below gate | native | native | native | 67 % | 75 % |
| Open-Jev 27B v1.1 | 52.5 | native | native | native | 67 % | 94 % |
| Raw Phi-4 mini direct logits | 13.7below gate | native | native | native | 50 % | 61 % |
| SimpleJev | 13.0below gate | native | native | native | 100 % | 58 % |
| smalljev semantic-v9 | 8.7below gate | native | native | native | 23 % | 73 % |
| kev 0.6B | 7.6below gate | native | native | native | 43 % | 69 % |
| OpenDecision | 7.4below gate | native | native | native | 51 % | 68 % |
| Qwen3.5-0.8B Decision Model | 7.3below gate | native | native | native | 22 % | 88 % |
| Fastino GLiNER-2.5-Decide | 7.0below gate | confidence | confidence | confidence | 100 % | 54 % |
| Deem 0.8B v1 | 6.5below gate | native | native | native | 77 % | 59 % |
| Decision Fast | 6.1below gate | native | native | native | 14 % | 90 % |
| Quyet-1.0-Small-EN | 6.1below gate | native | native | native | 28 % | 64 % |
| Laya typed-decisions | 5.5below gate | native | native | native | 1 % | 100 % |
| openJev Verdict 1.4 | 5.0below gate | native | native | native | 0 % | — |
| Raw Qwen3 0.6B direct logits | 5.6below gate | native | native | native | 100 % | 42 % |
| lev-350m | 4.8below gate | native | native | native | 8 % | 69 % |
| Raw Qwen3 1.7B direct logits | 5.1below gate | native | native | native | 100 % | 42 % |
| kev 0.5B | 4.6below gate | native | native | native | 26 % | 56 % |
| openJev Verdict | 4.1below gate | native | native | native | 33 % | 60 % |
| Quyet-1.0-Small | 4.0below gate | native | native | native | 29 % | 71 % |
| Quyet-1.0-Tiny | 3.7below gate | native | native | native | 20 % | 64 % |
| open-jev-deberta-v3-large | 2.8below gate | native | native | native | 15 % | 58 % |
| verdict-small | 2.3below gate | native | native | native | 45 % | 62 % |
| Laya | 1.8below gate | native | native | native | 19 % | 61 % |
| Mixedbread mxbai-rerank-base-v2 | 1.1below gate | native | native | native | 0 % | — |
| Laya multilingual | 0.3below gate | native | native | native | 72 % | 50 % |
| Certo v1 | 0.2below gate | native | native | native | 1 % | 25 % |
| CLM-8B | 0.1below gate | native | native | native | 51 % | 69 % |
| BAAI bge-reranker-v2-m3 | 0.1below gate | native | native | native | 0 % | — |
| Alibaba GTE Reranker ModernBERT-base | 0.0below gate | native | native | native | 0 % | — |
| Mirror | 0.0below gate | native | native | native | 92 % | 60 % |
| Open Jev JSON Canvas | 61.2 | label | label | label | 100 % | 85 % |
All data (127 systems)
Not yet measured on v1.6 · 28 systems with a dated carried score
Every ranked system of the live v1.5.7 board that is not measured on the v1.6.0 pool keeps its last published score, marked with the release that first published that measurement and that release's publication day. Carried rows are listed separately and never ranked together with v1.6-measured rows. v1.5 protocol (1,624 decisions: 904 open + 720 sealed); scores are on the v1.5 scale and are not comparable with v1.6-measured rows.
Dates: v1.5.0 published 2026-09-28 (25) · v1.5.3 published 2026-09-29 (2) · v1.5.6 published 2026-10-03 (1). The date is the publication day of the release that first published the measurement, not a per-model measurement timestamp.
| System | Measured on | Capability (v1.5 scale) | v1.5 Composite | Cost / 1,000 | Median latency |
|---|---|---|---|---|---|
| Qwen3.8 27B (Chutes TEE) | measured on v1.5.0 (2026-09-28) | 96.8 | 0.0 · was #108 on v1.5.7 | $2.1784estimate | 6.49 s |
| GPT-6 Luna (default medium reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.9 | 38.8 · was #38 on v1.5.7 | $0.1138estimate | 1.56 s |
| DeepSeek V4.1 Flash (thinking default) | measured on v1.5.0 (2026-09-28) | 95.3 | 6.6 · was #72 on v1.5.7 | $0.4976estimate | 1.78 s |
| GPT-6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.1 | 40.5 · was #37 on v1.5.7 | $0.1075estimate | 1.58 s |
| GPT-5.6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 94.5 | 22.4 · was #51 on v1.5.7 | $0.2047tariff | 1.32 s |
| Autoloops – Gemma 4 31B IT | measured on v1.5.0 (2026-09-28) | 81.3 | 40.5 · was #36 on v1.5.7 | $0.1032tariff | 0.61 s |
| SimpleJev Qwen3.8-27B | measured on v1.5.0 (2026-09-28) | 80.0 | 3.4 · was #74 on v1.5.7 | $0.6868estimate | 1.68 s |
| AutoJev-27B (RTX PRO 6000) | measured on v1.5.0 (2026-09-28) | 79.7 | 19.5 · was #56 on v1.5.7 | $0.2262estimate | 0.33 s |
| Surogate Rune 26B-A4B v3 (RTX PRO 6000)Not yet measured on the v1.6 pool. | measured on v1.5.0 (2026-09-28) | 79.0 | 66.5 · was #19 on v1.5.7 | $0.0502estimate | 0.35 s |
| Gemini 3.1 Flash-Lite | measured on v1.5.0 (2026-09-28) | 76.1 | 19.6 · was #54 on v1.5.7 | $0.2194tariff | 0.86 s |
| reflex-27b (Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 74.4 | 13.2 · was #68 on v1.5.7 | $0.2973estimate | 2.75 s |
| Instinct (ZooWork, Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 73.9 | 18.3 · was #60 on v1.5.7 | $0.2298estimate | 0.52 s |
| NInfer Qwen3.8-27B NVFP4 (T=1.5) | measured on v1.5.0 (2026-09-28) | 73.8 | 18.5 · was #59 on v1.5.7 | $0.2305estimate | 0.25 s |
| Vansa-3.4 (Vansa, hosted System One API)Hosted API retained at its v1.5.6 score under the three-refresh exposure cadence. | measured on v1.5.6 (2026-10-03) | 72.8 | 71.6 · was #5 on v1.5.7 | $0.0193estimate | 0.14 s |
| openjev-sglang (Qwen3.6-35B-A3B on SGLang) | measured on v1.5.0 (2026-09-28) | 70.8 | 29.0 · was #45 on v1.5.7 | $0.1398estimate | 1.20 s |
| Instinct Dual 4B | measured on v1.5.3 (2026-09-29) | 65.7 | 47.0 · was #30 on v1.5.7 | $0.0212estimate | 0.51 s |
| system-one-open (Gemma 4 E2B LoRA on an L4) | measured on v1.5.0 (2026-09-28) | 56.6 | 42.4 · was #35 on v1.5.7 | $0.0114estimate | 1.15 s |
| decision-machine-1 (milliseconds.ai) | measured on v1.5.0 (2026-09-28) | 48.0 | 3.2 · was #75 on v1.5.7 | $0.0286estimate | 0.18 s |
| jeff (Logan Markewich, GLiFormer 400M) | measured on v1.5.0 (2026-09-28) | 42.4 | 0.1 · was #89 on v1.5.7 | $0.0043estimate | 7.14 s |
| Von (wfzyx, Option-Marker 395M) | measured on v1.5.0 (2026-09-28) | 41.7 | 0.0 · was #110 on v1.5.7 | $0.0038estimate | 0.92 s |
| Bosun v3.1 0.6B | measured on v1.5.3 (2026-09-29) | 39.4 | 2.5 · was #78 on v1.5.7 | $0.0056estimate | 4.04 s |
| JevAct (einptein, jev1-2b-v2) | measured on v1.5.0 (2026-09-28) | 36.9 | 1.5 · was #81 on v1.5.7 | $0.0112estimate | 0.80 s |
| GLiNER2.5 multi (Fastino, 287M) | measured on v1.5.0 (2026-09-28) | 36.0 | 2.7 · was #77 on v1.5.7 | $0.0028estimate | 1.37 s |
| GLiNER2.5 small (Fastino, 74M) | measured on v1.5.0 (2026-09-28) | 32.5 | 0.9 · was #85 on v1.5.7 | $0.0028estimate | 0.47 s |
| GLiNER2 large (Fastino) | measured on v1.5.0 (2026-09-28) | 32.4 | 8.4 · was #70 on v1.5.7 | $0.0056estimate | 1.88 s |
| GLiNER2 (Fastino, gliner2.5-base) | measured on v1.5.0 (2026-09-28) | 24.5 | 2.3 · was #79 on v1.5.7 | $0.0028estimate | 1.00 s |
| Needle 3 (Cactus, 2-bit, local CPU) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 · was #102 on v1.5.7 | $0.0191estimate | 135.51 s |
| Needle 3, options as tools (post-hoc adapter mode) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 · was #103 on v1.5.7 | $0.0191estimate | 58.16 s |
Method · v1.6.1
System types (colours)
- Jev — reference (TypeSafe, closed) — The closed system JevBench is named after, shown as the reference.
- Closed API (weights not public) — Available through a hosted API; the weights cannot be downloaded.
- Open weights · LLM decoder — An autoregressive language model with public weights, including fine-tunes, merges and Jev rebuilds.
- Open weights · diffusion LM — A language model that generates by iterative denoising instead of token by token.
- Open weights · encoder / classifier — BERT-style encoders, NLI zero-shot classifiers and GLiNER-type models.
- Open weights · reranker — A cross-encoder or LLM reranker that scores options against the input.
- Base model control (no decision fine-tune, raw logits) — An official open checkpoint without any decision fine-tune, read out from raw logits, used as a floor.
- System (router / cascade / ensemble) — Several models combined at inference time.
Colours show architecture only; they never change a score or rank. Each class is assigned from cited evidence (config.json, model card, provider docs, or our own run record for hosted APIs).
Method amendments
- 5 Oct 2026, v1.6.1 (Florian's decision): hosted/API systems now answer the same full item set as self-hosted systems (1,200 sealed + 300 public). Their earlier answers on the API subsets were reused and only the missing sealed items were sent; every item was logged in the exposure ledger before it was sent. API rows are no longer equated. Consequence: every v1.6.0 sealed item has now been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases; v1.6.2 onward draws fresh items from the reserve.
- 5 Oct 2026, v1.6.1 cost rule (Florian's decision): cost per 1,000 decisions is now measured on one common item set for every system: all items except the 23 very long items that lie outside the Jev reference's accepted input range. Before, a system that processed those items paid for their tokens while a system that refused them did not. Refusals still count against Intelligence exactly as before. Hosted APIs with a per-token tariff are priced from their measured tokens on this common set; carried per-decision prices of self-hosted systems are unchanged. Effect: Sage 1.3.0 USD 0.0247 per 1,000 decisions on the common set (0.0650 on all 1,500 items); Jev 1.13.0 0.0323 (unchanged).
Rotating item sets
Each release draws fresh sealed decisions from a larger reserve. Self-hosted open-weights models (run offline on our own GPU pods or Sandy) answer S and P; externally hosted models answer the same full set (S and P) as the self-hosted systems since v1.6.1, so hosted APIs are no longer equated.
| Set | Items | Choice | Noul | Score | Answered by |
|---|---|---|---|---|---|
| S · Sealed release draw | 1,200 | 600 | 300 | 300 | self-hosted systems (run offline on our own GPU pods or Sandy) and, from v1.6.1, externally hosted APIs |
| A · API subset (part of S; used by the v1.6.0 hosted-API measurements) | 300 | 150 | 75 | 75 | every system (inside S since v1.6.1); retired for future draws |
| P · Public set | 300 | 150 | 75 | 75 | every system |
- Self-hosted systems: 1,500 items (S 1,200 + P 300). Hosted APIs: 1,500 items (S 1,200 + P 300), the same as self-hosted systems; A (300) is the part of S that earlier API measurements used.
- A sealed item is scored in at most three releases, then retired. An item used in one release is not drawn again in the next. If coverage minimums cannot be met, the release waits for newly reviewed items.
- The selection seed is committed (SHA-256) before any inference, and the draw is a deterministic function of policy, seed and item id.
API-exposure rule
- An external exposure is any sealed item sent to an endpoint we do not control: closed APIs, and open-weights models reached through third-party hosts, routers or a submitter's endpoint. Timeouts count as exposure.
- From v1.6.1 hosted models receive the full item set (S plus P); items they had already answered were reused and only the missing sealed items were sent. Hosted models are re-measured at most once every three refresh releases unless a verified new model version ships. Between measurements they keep their last score with its measurement date.
- Every externally sent sealed item is logged before dispatch. An item exposed to a provider is never scored again for that provider; once two different providers have received it, it retires for everyone. In v1.6.0 Jev 1.13.0 and Fastino GLiNER-2.5-Decide both received the same A, so A retires globally. With the v1.6.1 amendment every v1.6.0 sealed item has been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases.
Public-versus-sealed gap penalty
Intelligence is half public, half sealed. A system whose public score exceeds its sealed score by more than the field-median gap (G_med = 2.6 points) plus 8 points loses one Intelligence point per excess point. Hosted APIs are compared on P versus A against the same self-hosted systems' P-versus-A gap (-1.2 points).
Hosted APIs (not equated since v1.6.1)
Hosted APIs that answered the full set are scored exactly like self-hosted systems; no equating offset is applied to them. Category and language values stay raw.
Headline and Composite
Capability = mean(Intelligence, Calibration) for systems within twice the Jev 1.13.0 cost and median latency (the official caps; the sliders change only your view). The Composite (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost with the v1.5 low-axis gates; it remains secondary. Request types Choice, Noul and Score weigh equally; tiers weigh easy 0.1, standard 0.2, hard 0.4, judge 0.3.
Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s.
Costs of v1.6-measured systems carry each system's published v1.5.4 cost per 1,000 decisions (pricing rules unchanged; v1.6 item lengths differ) unless the row says otherwise; Fastino's is an estimate from its published tariff and measured tokens.
Noul decisiveness and Score baseline (addendum B)
Scored with method option B (scorer setting O1S), selected on 3 Oct 2026 after the v1.6 scores were known and disclosed as a post-results change. Each split × type competence is clipped at 0 before the type weighting, so a type answered no better than chance counts as chance instead of negative. The Score chance baseline is the error of always predicting the mid-scale level, so a flat know-nothing distribution earns about 0. A Score cell whose golds all sit at mid-scale keeps the v1.5 random-level baseline; this only occurs in small breakdown and bootstrap cells. Calibration is unchanged. This run: G_med = 2.57 points.
Failed requests and very long items
23 items of this draw are very long (about 77,000 to 81,000 input tokens). Systems with a shorter context window refuse them; Jev-Omni runs out of GPU memory on them on its listed RTX 6000 (48 GB) recipe; decider-4b v2's server truncates them to about 32,800 tokens and answers; Plumb-4B reads them in full. Each outcome is the system's own recipe and is scored as such. A failed, refused or unparseable answer counts as wrong for Intelligence and stays in the denominator; it does not enter Calibration (the v1.5 rule, applied to every system). Failed answers per system: Sage 1.3.0 0/1500 · Jev 1.13.0 23/1500 · Quyet-1.0-Large 0/1500 · wity-1 0/1500 · torchcast-decision-12b 23/1500 · Winnow-12B Q8 23/1500 · deck-31B 23/1500 · Cygnet 23/1500 · Jev-Omni 23/1500 · Plumb-4B 0/1500 · spark-s1-4b-v6 0/1500 · swanOne 23/1500 · Quyet-1.0-Medium 0/1500 · lev 25/1500 · NInfer Qwen3.8-Flash-Next mixed 23/1500 · jev-local 23/1500 · decider-4b v2 0/1500 · JevK5 v0.3 23/1500 · metask-jev-4b 23/1500 · deck-4B v1.0 23/1500 · Clef-Flash 0/1500 · Qwen3.5-9B Jev-like data-mix v2 23/1500 · JevK5 v0.2.0 23/1500 · JevOne 0/1500 · Hopper 0/1500 · Imajev-4B 23/1500 · Malkuth-4B 23/1500 · jqv 23/1500 · decider-35b-a3b 0/1500 · Manchego v2.1 23/1500 · Decision 4B v1.2 23/1500 · Raw Qwen3 4B Instruct 2507 direct logits 23/1500 · Bespoke Nimble 9B 23/1500 · reflex 4B 0/1500 · typecastlm 0/1500 · Decision 4B v1.1 23/1500 · AutoJev-27B 23/1500 · Eikos-27B 23/1500 · NInfer Qwen3.8-27B NVFP4 23/1500 · OpenJev 23/1500 · Open-Jev 9B 23/1500 · Standard One 8B 23/1500 · Clef 0/1500 · JEV Qwen3.5-9B Base NVFP4 23/1500 · SemIf, formerly OpenJev 23/1500 · Bev / Bonsai 27B 23/1500 · LitJev 23/1500 · local-jev Qwen3.5-4B 0/1500 · system-one 23/1500 · kev 8B 25/1500 · Decision 2B 23/1500 · OpenSourceJev 23/1500 · open-alternative-jev 0/1500 · Raw Qwen3 8B direct logits 23/1500 · decider-2b 0/1500 · Nemotron Diffusion 8B 23/1500 · kev 4B 29/1500 · Malkuth-2B 23/1500 · Qwen3-Reranker-4B 0/1500 · ZeroEntropy zerank-2 0/1500 · Open-Jev 2B 23/1500 · Open-Jev 27B v1.1 23/1500 · Raw Phi-4 mini direct logits 23/1500 · SimpleJev 23/1500 · smalljev semantic-v9 0/1500 · kev 0.6B 30/1500 · OpenDecision 0/1500 · Qwen3.5-0.8B Decision Model 0/1500 · Fastino GLiNER-2.5-Decide 31/1500 · Deem 0.8B v1 23/1500 · Decision Fast 23/1500 · Quyet-1.0-Small-EN 0/1500 · Laya typed-decisions 0/1500 · openJev Verdict 1.4 0/1500 · Raw Qwen3 0.6B direct logits 23/1500 · lev-350m 23/1500 · Raw Qwen3 1.7B direct logits 23/1500 · kev 0.5B 32/1500 · openJev Verdict 0/1500 · Quyet-1.0-Small 0/1500 · Quyet-1.0-Tiny 0/1500 · open-jev-deberta-v3-large 73/1500 · verdict-small 0/1500 · Laya 0/1500 · Mixedbread mxbai-rerank-base-v2 0/1500 · Laya multilingual 0/1500 · Certo v1 0/1500 · CLM-8B 0/1500 · BAAI bge-reranker-v2-m3 0/1500 · Alibaba GTE Reranker ModernBERT-base 0/1500 · Mirror 592/1500 · Open Jev JSON Canvas 23/1500 · wity-1 0/1500 · wity-1 0/1500.
Calibration basis
Systems that return a full probability distribution are calibrated on all components (top-label error, plus distribution distance for Choice and ranked-probability error for Score). Fastino GLiNER-2.5-Decide returns a single confidence value, so its Calibration is the top-label error only and is not like-for-like with full-distribution systems.
Overnight full re-measure (4–5 Oct 2026)
Every system with a reproducible recipe was re-run on the v1.6.0 pool overnight with the same pinned inputs and scorer (method option B / O1S). This page uses scoring round score-v161-3 (v1.6.1) (2026-10-06 00:39:36 UTC). Only complete runs (1,500 items self-hosted, the full API input for hosted APIs) are ranked; partial runs are never ranked, and systems not yet re-measured keep their dated v1.5.x score in the separate table.
- Scores are the official v1.6.0 scorer (score_v16.py, method option B / O1S, bootstrap B = 1,000) over complete outputs only; each scoring round is kept separately.
- Hosted APIs are scored on the same full item set as self-hosted systems (S 1,200 + P 300, v1.6.1) and are not equated.
- GPU-class deviations: reproducible recipes used the hardware listed per row, including H100 for large fast-lane decoders and RTX 6000 for the baseline; hardware differences remain in measured latency. The standard x2 + 0.15 s self-hosted adjustment is an assumption, not a hardware normalization.
- Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s. Wity auto remains outside the latency cap and stays ranked in Composite A. OFF and ALWAYS are unranked variants of the AUTO main row.
- 6 further measured candidates await a separate publication decision.
- Sage 1.3.0 (Levanto Labs) was measured on 5 Oct from Sandy (Helsinki), text only. Cost uses the Levanto list tariff (USD 0.05/M input, USD 10/M output; levanto.ai/pricing, read 5 Oct) over all 600 answered rows: USD 0.0766/1,000 decisions. It exceeds the frozen Jev-class cost cap and is excluded from the Capability headline.
- A fresh Monday Jev 1.13.0 re-check on A3 measured p50 0.239 s and equated Capability 77.5 versus the official 76.5, within the confidence interval; the official v1.6.0 Jev row is retained.
Supplementary API draws A2 and A3
History: hosted APIs were first measured on API subsets (A, then A2 and A3) and equated to the self-hosted scale in v1.6.0. From v1.6.1 they answer the same full set as self-hosted systems, so no equating is applied; the earlier subsets A, A2 and A3 are retired.
Per-model exposure counts (hosted and author-hosted endpoints)
| System (provider) | v1.6 sealed items sent | Scored sealed set | Status |
|---|---|---|---|
| wity-1-auto (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-off (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-always (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| jev-1.13.0 (typesafe) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| fastino-gliner-2-5-decide (fastino) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| gpt-6-luna (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gpt-6-luna-low (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gpt-5.6-luna (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gemini-3.1-flash-lite (google) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| deepseek-flash (deepseek) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| qwen3.8-27b (chutes) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| decision-machine-1 (milliseconds.ai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| classifier-dev-fast (classifier.dev) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| vansa-3.4 (vansa) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| kushal-gemma4-31b-it-autoloops (autoloops) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| system-one-open (modal via modal) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| openjev-sglang (modal via modal) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| instinct (zoowork) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| instinct-dual-4b (zoowork) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| simplejev-qwen3.8-27b (featherless) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| simplejev-qwen3.6-35b-a3b (featherless) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| jevact (jevact) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| Sage 1.3.0 (Levanto Labs) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
Revision history
- v1.6.1 · 2026-10-05 · Hosted APIs (fastino-gliner-2-5-decide, jev-1.13.0, sage-1.3.0, wity-1, wity-1-always, wity-1-off) answer the full 1,500-item set instead of the 600-item API subset and are no longer equated; self-hosted rows unchanged; whole v1.6.0 sealed draw retired. Cost per 1,000 decisions is measured on one common item set (all items except the 23 outside the Jev reference input range) for every token-priced row.
- v1.6.0 · 2026-10-05 · Rotating item draw (1,200 sealed + 300 public); hosted APIs on API subsets equated to the self-hosted scale; /jev-models/v1.6.0 stays available unchanged.
Provenance
Measured on the v1.6 pool (run completion day, UTC): 2026-10-02: Winnow-12B Q8, Cygnet, Jev-Omni, Plumb-4B, swanOne, NInfer Qwen3.8-Flash-Next mixed, jev-local, decider-4b v2, JevK5 v0.3, metask-jev-4b, Qwen3.5-9B Jev-like data-mix v2, JevK5 v0.2.0, JevOne, Hopper, Imajev-4B, Malkuth-4B, jqv, decider-35b-a3b, Decision 4B v1.2, Raw Qwen3 4B Instruct 2507 direct logits, Bespoke Nimble 9B, reflex 4B, typecastlm, Decision 4B v1.1, AutoJev-27B, Eikos-27B, NInfer Qwen3.8-27B NVFP4, OpenJev, Open-Jev 9B, Standard One 8B, JEV Qwen3.5-9B Base NVFP4, SemIf, formerly OpenJev, LitJev, local-jev Qwen3.5-4B, system-one, kev 8B, OpenSourceJev, open-alternative-jev, Raw Qwen3 8B direct logits, decider-2b, kev 4B, Malkuth-2B, Qwen3-Reranker-4B, ZeroEntropy zerank-2, Open-Jev 2B, Raw Phi-4 mini direct logits, SimpleJev, smalljev semantic-v9, kev 0.6B, OpenDecision, Qwen3.5-0.8B Decision Model, openJev Verdict 1.4, Raw Qwen3 0.6B direct logits, Raw Qwen3 1.7B direct logits, kev 0.5B, openJev Verdict, open-jev-deberta-v3-large, verdict-small, Laya, Mixedbread mxbai-rerank-base-v2, Certo v1, CLM-8B, BAAI bge-reranker-v2-m3, Alibaba GTE Reranker ModernBERT-base, Mirror, Open Jev JSON Canvas · 2026-10-04: Quyet-1.0-Large, torchcast-decision-12b, Quyet-1.0-Medium, lev, Clef-Flash, Manchego v2.1, Clef, Bev / Bonsai 27B, Nemotron Diffusion 8B, Open-Jev 27B v1.1, Deem 0.8B v1, Quyet-1.0-Small-EN, Laya typed-decisions, Quyet-1.0-Small, Quyet-1.0-Tiny, Laya multilingual · 2026-10-05: Sage 1.3.0, Jev 1.13.0, deck-31B, spark-s1-4b-v6, deck-4B v1.0, Decision 2B, Fastino GLiNER-2.5-Decide, Decision Fast, lev-350m · 2026-10-06: wity-1, wity-1, wity-1.
Aggregate files: results sha256 5d4567d5e5acd945d17dd082adbbd6188d38174ca523b0fa6b3b02c6cc2dc5b1 · categories sha256 db333c130a7b16beae0aacc153d3dd69cf8e54f277e790dad41769208570d6db · dated carry sha256 334e851944e92383ab3071e52b3012040c7023d0fe236395953169ebee8f0459. Scoring source sha256 e15347783b8bedeeddd9017a831608a957f5bf0386f05902c50161ce0468f911. The method, release data and carry artifact are independently hashable.
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
Open this section to load the earlier public-only board and diagnostics.