Official JevBench v1.6.0

JevBench v1.6.0 — Jev alternatives ranking

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Release v1.6.0 · 1,500 decisions per self-hosted system and 600 per hosted API · 92 ranked systems · 28 systems retain a separately dated v1.5.x score · only system-level aggregates are published · aggregate results JSON · SHA-256 1161dbbfa1a3225008d14d42dc1ccabfd0de51be8bb23251ba929533b840dabb

Share this version · View live board · Previous release: JevBench v1.5.7

Making decisions from images? Explore Image JevBench v0.1.5 and compare its systems.

JevBench v1.6.0 · headline

JevBench Capability Score

Capability ranking of Jev-class systems

Capability Score averages Intelligence and Calibration. Quyet-1.0-Large leads the Jev-class systems with 81.7.

Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘

Adjust cost / latency caps · 2× official
Official
#SystemICCap.$/1k
  1. 1Quyet-1.0-Large73.450.581.7$0.045*
  2. 2deck-31B73.049.977.6$0.047*
  3. 3Jev 1.13.0API62.454.776.5$0.032*
  4. 4torchcast-decision-12b60.556.471.7$0.028*
  5. 5Jev-Omni55.556.171.3$0.029*
  6. 6Winnow-12B Q859.556.671.2$0.028*
  7. 7Cygnet54.856.470.9$0.028*
  8. 8Plumb-4B43.063.165.4$0.017*
  9. 9decider-4b v240.164.564.9$0.015*
  10. 10JevK5 v0.338.563.164.2$0.017*
Show all 63 Jev-class systems (53 more)
  1. 11Clef-Flash42.246.863.8$0.059*
  2. 12deck-4B v1.038.663.763.6$0.016*
  3. 13JevK5 v0.2.036.363.162.8$0.017*
  4. 14Hopper34.762.362.4$0.018*
  5. 15Imajev-4B34.663.361.9$0.017*
  6. 16jqv33.951.261.9$0.042*
  7. 17metask-jev-4b39.257.761.8$0.026*
  8. 18lev41.161.861.7$0.019*
  9. 19Malkuth-4B33.855.861.1$0.030*
  10. 20Decision 4B v1.232.463.160.4$0.017*
  11. 21Quyet-1.0-Medium41.763.459.2$0.017*
  12. 22Manchego v2.132.564.358.3$0.015*
  13. 23jev-local41.958.957.5$0.024*
  14. 24Decision 4B v1.129.163.157.2$0.017*
  15. 25JEV Qwen3.5-9B Base NVFP429.347.657.1$0.056*
  16. 26typecastlm29.864.056.5$0.016*
  17. 27SemIf27.163.155.4$0.017*
  18. 28local-jev Qwen3.5-4B24.159.355.0$0.023*
  19. 29Decision 2B22.966.454.8$0.013*
  20. 30spark-s1-4b-v643.260.454.0$0.021*
  21. 31ZeroEntropy zerank-215.648.449.4$0.052*
  22. 32open-alternative-jev22.363.349.1$0.017*
  23. 33OpenSourceJev23.269.148.6$0.011*
  24. 34Malkuth-2B15.965.647.6$0.014*
  25. 35Nemotron Diffusion 8B20.455.547.5$0.030*
  26. 36Qwen3-Reranker-4B16.348.444.7$0.052*
  27. 37decider-2b20.764.944.6$0.015*
  28. 38BAAI bge-reranker-v2-m30.159.344.1$0.023*
  29. 39kev 4B19.565.844.0$0.014*
  30. 40Mixedbread mxbai-rerank-base-v21.160.343.9$0.021*
  31. 41lev-350m4.880.043.3$0.0046*
  32. 42Certo v10.296.943.2$0.0013*
  33. 43smalljev semantic-v98.760.843.0$0.020*
  34. 44Alibaba GTE Reranker ModernBERT-base0.069.341.7$0.011*
  35. 45Decision Fast6.180.140.6$0.0046*
  36. 46openJev Verdict 1.45.086.640.4$0.0028*
  37. 47Quyet-1.0-Small-EN6.189.139.5$0.0023*
  38. 48OpenDecision7.479.138.0$0.0050*
  39. 49Fastino GLiNER-2.5-DecideAPI6.446.637.6$0.060*
  40. 50verdict-small2.3100.037.6$0.00087*
  41. 51kev 0.6B7.680.136.9$0.0046*
  42. 52Raw Phi-4 mini direct logits13.753.235.3$0.036*
  43. 53Raw Qwen3 4B Instruct 2507 direct logits34.063.434.6$0.017*
  44. 54kev 0.5B4.680.132.9$0.0046*
  45. 55Quyet-1.0-Small4.079.632.5$0.0048*
  46. 56Quyet-1.0-Tiny3.788.431.1$0.0024*
  47. 57Open Jev JSON Canvas61.249.330.6$0.049*
  48. 58openJev Verdict4.186.626.6$0.0028*
  49. 59CLM-8B0.150.520.1$0.045*
  50. 60Deem 0.8B v16.580.318.1$0.0045*
  51. 61Laya multilingual0.382.217.0$0.0039*
  52. 62Raw Qwen3 1.7B direct logits5.168.58.3$0.011*
  53. 63Raw Qwen3 0.6B direct logits5.677.68.2$0.0056*

Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.

Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = 0.18, n = 92).

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • Jev-compatible decision model
  • system-one-open
  • green: ≤ reference
  • amber: ≤ cap (2× reference)
  • red: > cap
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).

  1. –wity-1API71.658.479.8$0.024*Outside: latency 2.6× Jev (v1.5 reference)API price · eligibility checked at the developer's list price; at base-model pricing it would exceed the cost cap (2.70× Jev)
  2. –Eikos-27B64.428.777.2$0.24*Outside: cost 7.4× Jev (v1.5 reference)
  3. –Sage 1.3.0API64.343.577.1$0.077*Outside: cost 2.4× Jev (v1.5 reference)
  4. –OpenJev70.328.575.9$0.24*Outside: cost 7.5× Jev (v1.5 reference), latency 3.0× Jev (v1.5 reference)
  5. –Clef59.128.175.3$0.25*Outside: cost 7.7× Jev (v1.5 reference)
  6. –swanOne56.242.272.4$0.085*Outside: cost 2.6× Jev (v1.5 reference)
  7. –AutoJev-27B55.629.471.6$0.23*Outside: cost 7.0× Jev (v1.5 reference)
  8. –NInfer Qwen3.8-Flash-Next mixed49.042.569.7$0.082*Outside: cost 2.5× Jev (v1.5 reference)
  9. –JevOne46.039.868.9$0.10*Outside: cost 3.1× Jev (v1.5 reference)
  10. –NInfer Qwen3.8-27B NVFP451.829.168.0$0.23*Outside: cost 7.1× Jev (v1.5 reference)
  11. –Open-Jev 27B v1.152.514.468.0$0.72*Outside: cost 22.2× Jev (v1.5 reference), latency 2.8× Jev (v1.5 reference)
  12. –LitJev43.328.466.6$0.24*Outside: cost 7.6× Jev (v1.5 reference), latency 5.0× Jev (v1.5 reference)
  13. –decider-35b-a3b48.734.466.2$0.15*Outside: cost 4.8× Jev (v1.5 reference)
  14. –Bev / Bonsai 27B43.828.266.1$0.25*Outside: cost 7.6× Jev (v1.5 reference), latency 3.4× Jev (v1.5 reference)
  15. –Open-Jev 9B43.933.161.8$0.17*Outside: cost 5.3× Jev (v1.5 reference), latency 2.2× Jev (v1.5 reference)
  16. –Qwen3.5-9B Jev-like data-mix v241.945.760.9$0.065*Outside: cost 2.0× Jev (v1.5 reference)
  17. –reflex 4B30.963.159.0$0.017*Outside: latency 4.6× Jev (v1.5 reference)
  18. –Standard One 8B32.843.358.4$0.078*Outside: cost 2.4× Jev (v1.5 reference)
  19. –Bespoke Nimble 9B43.136.856.5$0.13*Outside: cost 4.0× Jev (v1.5 reference)
  20. –kev 8B29.340.450.0$0.097*Outside: cost 3.0× Jev (v1.5 reference)
  21. –Open-Jev 2B21.833.147.7$0.17*Outside: cost 5.3× Jev (v1.5 reference)
  22. –Laya typed-decisions5.582.544.2$0.0038*Outside: latency 3.0× Jev (v1.5 reference)
  23. –Qwen3.5-0.8B Decision Model7.379.542.6$0.0048*Outside: latency 18.3× Jev (v1.5 reference)
  24. –open-jev-deberta-v3-large2.877.635.4$0.0056*Outside: latency 5.7× Jev (v1.5 reference)
  25. –Laya1.884.932.5$0.0032*Outside: latency 2.9× Jev (v1.5 reference)
  26. –SimpleJev13.070.331.6$0.0098*Outside: latency 14.3× Jev (v1.5 reference)
  27. –Raw Qwen3 8B direct logits26.545.530.1$0.065*Outside: cost 2.0× Jev (v1.5 reference)
  28. –system-one29.345.029.7$0.068*Outside: cost 2.1× Jev (v1.5 reference)
  29. –Mirror0.089.325.8$0.0023*Outside: latency 2.4× Jev (v1.5 reference)

Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 63 of 92 systems qualify; the other 29, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.

Capability against cost and speed

Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 63 systems. Upper right is best: more capable and cheaper.01020304050607080100$0.0010$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev (v1.5 reference)← priciercheaper →1. Quyet-1.0-Large2. deck-31B3. Jev 1.13.04. torchcast-decision-12b5. Jev-Omni
63 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 63 systems. Upper right is best: more capable and faster.010203040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev (v1.5 reference)← slowerfaster →1. Quyet-1.0-Large2. deck-31B3. Jev 1.13.04. torchcast-decision-12b5. Jev-Omni
63 systems. Tap a bubble for its values.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • Jev-compatible decision model
  • system-one-open
  • faint = outside Jev-class
127 of 127 systems

Model kind

Jev-class

Release

Parameters

No system in this release reports an exact parameter count.

Developer/API price $/1k
Base-model price $/1k
Scored cost $/1k
Alternative price $/1k
p50 latency (s)
p95 latency (s)

JevBench v1.6.0

JevBench Composite Score: 92 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

Weights:
Adjust weights ↓

View by:Capability ↑

The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis.

Greener = stronger within its column.

92 of 92 systems, sorted by official rank, #1 first.

  1. 1Quyet-1.0-LargenewBase model: undisclosed71.4I 73C 90S 87K 51est.$0.045
  2. 2Jev 1.13.0APIBase model: undisclosed71.1I 62C 91S 91K 55est.$0.032
  3. 3wity-1APInewBase model: undisclosedQwen3.6-35B-A3B base-model reference price: 44.2 (would be #11)70.9I 72C 88S 72K 58est.$0.024
  4. 4torchcast-decision-12bnewBase model: undisclosed69.9I 61C 83S 92K 56est.$0.028
  5. 5deck-31BnewBase model: undisclosed69.3I 73C 82S 87K 50est.$0.047
  6. 6Winnow-12B Q8Base model: google/gemma-4-12B-itsource68.9I 60C 83S 87K 57est.$0.028
  7. 7CygnetBase model: google/gemma-4-12B-itsource68.6I 55C 87S 92K 56est.$0.028
  8. 8Jev-OmniBase model: google/gemma-4-12B-itsource67.7I 56C 87S 85K 56est.$0.029
  9. 9Sage 1.3.0APInewBase model: undisclosed50.1I 64C 90S 94K 43est.$0.077
  10. 10Plumb-4BBase model: alibiserikbay/JevK5source48.3I 43C 88S 94K 63est.$0.017
  11. 11spark-s1-4b-v6Base model: Qwen3.5-4Bsource44.9I 43C 65S 87K 60est.$0.021
  12. 12swanOneBase model: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4source43.8I 56C 89S 83K 42est.$0.085
  13. 13Quyet-1.0-MediumnewBase model: undisclosed43.6I 42C 77S 89K 63est.$0.017
  14. 14levBase model: undisclosed42.2I 41C 82S 88K 62est.$0.019
  15. 15NInfer Qwen3.8-Flash-Next mixedBase model: Qwen3.8-Flash-Nextsource42.0I 49C 90S 89K 43est.$0.082
  16. 16jev-localBase model: Qwen/Qwen3.5-9Bsource41.3I 42C 73S 76K 59est.$0.024
  17. 17decider-4b v2Base model: Qwen/Qwen3.5-4B-Basesource41.2I 40C 90S 92K 65est.$0.015
  18. 18JevK5 v0.3Base model: Qwen/Qwen3.5-4Bsource37.4I 39C 90S 94K 63est.$0.017
  19. 19metask-jev-4bBase model: Qwen3.5-4Bsource37.3I 39C 84S 91K 58est.$0.026
  20. 20deck-4B v1.0newBase model: undisclosed37.3I 39C 89S 90K 64est.$0.016
Show all 92 systems (72 more ranked, 0 more not ranked)
  1. 21Clef-FlashBase model: undisclosedWorkers AI price scenario (latency unmeasured): 39.2 (would be #18)36.6I 42C 85S 88K 47est.$0.059
  2. 22Qwen3.5-9B Jev-like data-mix v2Base model: Qwen/Qwen3.5-9Bsource33.6I 42C 80S 85K 46est.$0.065
  3. 23JevK5 v0.2.0Base model: Qwen3.5-4Bsource32.1I 36C 89S 92K 63est.$0.017
  4. 24JevOneBase model: Qwen/Qwen3.6-35B-A3Bsource31.1I 46C 92S 90K 40est.$0.101
  5. 25HopperBase model: Qwen/Qwen3.5-4Bsource28.9I 35C 90S 91K 62est.$0.018
  6. 26Imajev-4BBase model: Qwen/Qwen3.5-4Bsource28.7I 35C 89S 91K 63est.$0.017
  7. 27Malkuth-4BBase model: Qwen/Qwen3.5-4B-Basesource26.2I 34C 88S 90K 56est.$0.030
  8. 28jqvBase model: Qwen3-32Bsource25.4I 34C 90S 83K 51est.$0.042
  9. 29decider-35b-a3bBase model: Qwen/Qwen3.5-35B-A3B-Basesource24.8I 49C 84S 91K 34est.$0.154
  10. 30Manchego v2.1Base model: Qwen/Qwen3.5-4Bsource24.5I 33C 84S 91K 64est.$0.015
  11. 31Decision 4B v1.2Base model: undisclosed24.5I 32C 88S 94K 63est.$0.017
  12. 32Raw Qwen3 4B Instruct 2507 direct logitsBase model: undisclosedsource21.8I 34C 35S 91K 63est.$0.017
  13. 33Bespoke Nimble 9BBase model: Qwen3.5-9Bsource21.1I 43C 70S 86K 37est.$0.128
  14. 34reflex 4BBase model: Qwen/Qwen3.5-4Bsource20.6I 31C 87S 69K 63est.$0.017
  15. 35typecastlmBase model: Qwen3.5-4Bsource19.7I 30C 83S 93K 64est.$0.016
  16. 36Decision 4B v1.1Base model: undisclosed18.7I 29C 85S 94K 63est.$0.017
  17. 37AutoJev-27BBase model: Qwen/Qwen3.8-27Bsource18.5I 56C 88S 89K 29est.$0.226
  18. 38Eikos-27BBase model: Qwen/Qwen3.8-27Bsource18.1I 64C 90S 88K 29est.$0.238
  19. 39NInfer Qwen3.8-27B NVFP4Base model: Qwen3.8-27Bsource17.7I 52C 84S 90K 29est.$0.231
  20. 40OpenJevBase model: undisclosedsource17.3I 70C 81S 73K 29est.$0.241
  21. 41Open-Jev 9BBase model: Qwen/Qwen3.5-9Bsource17.0I 44C 80S 74K 33est.$0.170
  22. 42Standard One 8BBase model: mistralai/Ministral-3-8B-Instruct-2512-BF16source16.9I 33C 84S 93K 43est.$0.078
  23. 43ClefBase model: undisclosedWorkers AI price scenario (latency unmeasured): 29.4 (would be #25)16.8I 59C 92S 83K 28est.$0.249
  24. 44JEV Qwen3.5-9B Base NVFP4Base model: ig1/Qwen3.5-9B-NVFP4source16.0I 29C 85S 94K 48est.$0.056
  25. 45SemIfBase model: Qwen/Qwen3.5-4Bsource15.5I 27C 84S 91K 63est.$0.017
  26. 46Bev / Bonsai 27BBase model: Qwen/Qwen3.8-27Bsource11.7I 44C 89S 73K 28est.$0.247
  27. 47LitJevBase model: Qwen/Qwen3.8-27Bsource11.5I 43C 90S 68K 28est.$0.244
  28. 48local-jev Qwen3.5-4BBase model: Qwen/Qwen3.5-4Bsource11.4I 24C 86S 86K 59est.$0.023
  29. 49system-oneBase model: Qwen3-8Bsource11.0I 29C 30S 91K 45est.$0.068
  30. 50kev 8BBase model: Qwen/Qwen3-8B-Basesource10.6I 29C 71S 89K 40est.$0.097
  31. 51Decision 2BBase model: openbmb/MiniCPM5-2Bsource10.3I 23C 87S 90K 66est.$0.013
  32. 52OpenSourceJevBase model: Qwen3.5-4Bsource10.3I 23C 74S 77K 69est.$0.011
  33. 53open-alternative-jevBase model: Qwen3.5-4Bsource9.4I 22C 76S 91K 63est.$0.017
  34. 54Raw Qwen3 8B direct logitsBase model: Qwen/Qwen3-8B-Basesource9.3I 27C 34S 89K 46est.$0.065
  35. 55decider-2bBase model: Qwen/Qwen3.5-2B-Basesource7.7I 21C 69S 95K 65est.$0.015
  36. 56Nemotron Diffusion 8BBase model: nvidia/Nemotron-Labs-Diffusion-8Bsource7.3I 20C 75S 94K 55est.$0.030
  37. 57kev 4BBase model: Qwen/Qwen3-4B-Basesource · Evaluated Qwen3 variant; later releases use a different base.6.6I 20C 68S 90K 66est.$0.014
  38. 58Malkuth-2BBase model: empero-ai/Qwen3.8-2B-Distillsource4.0I 16C 79S 92K 66est.$0.014
  39. 59Qwen3-Reranker-4BBase model: Qwen/Qwen3-4B-Basesource3.7I 16C 73S 82K 48est.$0.052
  40. 60ZeroEntropy zerank-2Base model: Qwen/Qwen3-4Bsource3.3I 16C 83S 83K 48est.$0.052
  41. 61Open-Jev 2BBase model: Qwen/Qwen3.5-2Bsource3.2I 22C 74S 77K 33est.$0.170
  42. 62Open-Jev 27B v1.1Base model: Qwen/Qwen3.8-27Bsource2.9I 52C 83S 73K 14est.$0.716
  43. 63Raw Phi-4 mini direct logitsBase model: undisclosedsource2.5I 14C 57S 92K 53est.$0.036
  44. 64SimpleJevBase model: Qwen/Qwen3.5-0.8Bsource2.1I 13C 50S 59K 70est.$0.0098
  45. 65smalljev semantic-v9Base model: openbmb/MiniCPM5-2B-Basesource0.8I 9C 77S 90K 61est.$0.020
  46. 66kev 0.6BBase model: Qwen/Qwen3-0.6B-Basesource0.5I 8C 66S 92K 80est.$0.0046
  47. 67OpenDecisionBase model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0source0.5I 7C 69S 87K 79est.$0.0050
  48. 68Qwen3.5-0.8B Decision ModelBase model: Qwen/Qwen3.5-0.8B-Basesource0.5I 7C 78S 56K 80est.$0.0048
  49. 69Deem 0.8B v1Base model: Qwen/Qwen3.5-0.8Bsource0.3I 7C 30S 85K 80est.$0.0045
  50. 70Decision FastBase model: Qwen/Qwen3-0.6B-Basesource0.3I 6C 75S 91K 80est.$0.0046
  51. 71Quyet-1.0-Small-ENnewBase model: undisclosed0.3I 6C 73S 95K 89est.$0.0023
  52. 72Fastino GLiNER-2.5-DecideAPInewBase model: undisclosed0.3I 6C 69S 85K 47est.$0.060
  53. 73Laya typed-decisionsBase model: ModernBERT-largesource0.2I 5C 83S 71K 83est.$0.0038
  54. 74openJev Verdict 1.4Base model: knowledgator/gliclass-modern-base-v2.0source0.2I 5C 76S 81K 87est.$0.0028
  55. 75Raw Qwen3 0.6B direct logitsBase model: Qwen/Qwen3-0.6B-Basesource0.2I 6C 11S 94K 78est.$0.0056
  56. 76lev-350mBase model: LiquidAI/LFM2.5-350Msource0.2I 5C 82S 94K 80est.$0.0046
  57. 77Raw Qwen3 1.7B direct logitsBase model: Qwen/Qwen3-1.7B-Basesource0.1I 5C 12S 93K 69est.$0.011
  58. 78kev 0.5BBase model: Qwen/Qwen2.5-0.5Bsource0.1I 5C 61S 92K 80est.$0.0046
  59. 79openJev VerdictBase model: undisclosedsource0.1I 4C 49S 79K 87est.$0.0028
  60. 80Quyet-1.0-SmallnewBase model: undisclosed0.1I 4C 61S 95K 80est.$0.0048
  61. 81Quyet-1.0-TinynewBase model: undisclosed0.1I 4C 58S 95K 88est.$0.0024
  62. 82open-jev-deberta-v3-largeBase model: microsoft/deberta-v3-largesource0.0I 3C 68S 68K 78est.$0.0056
  63. 83verdict-smallBase model: intfloat/multilingual-e5-smallsource0.0I 2C 73S 89K 100est.$0.0009
  64. 84LayaBase model: ModernBERT-largesource0.0I 2C 63S 73K 85est.$0.0032
  65. 85Mixedbread mxbai-rerank-base-v2Base model: undisclosedsource0.0I 1C 87S 93K 60est.$0.021
  66. 86Laya multilingualBase model: mmBERT-basesource0.0I 0C 34S 78K 82est.$0.0039
  67. 87Certo v1Base model: ModernBERT-largesource0.0I 0C 86S 93K 97est.$0.0013
  68. 88CLM-8BBase model: Qwen/Qwen3-8Bsource0.0I 0C 40S 93K 51est.$0.045
  69. 89BAAI bge-reranker-v2-m3Base model: BAAI/bge-m3source0.0I 0C 88S 94K 59est.$0.023
  70. 90Alibaba GTE Reranker ModernBERT-baseBase model: answerdotai/ModernBERT-basesource0.0I 0C 83S 94K 69est.$0.011
  71. 91MirrorBase model: undisclosed0.0I 0C 52S 73K 89est.$0.0023
  72. 92Open Jev JSON CanvasBase model: google/diffusiongemma-26B-A4B-itsource0.0I 61C 0S 86K 49est.$0.049
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Unclassified
  • Jev-compatible decision model
  • system-one-open
  • Striped bar = same system under the labelled alternative price assumption
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

Compare two systems

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 71.1 (#2)
  • B: Quyet-1.0-Large — Jev-compatible decision model · Score 71.4 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Quyet-1.0-Large. Intelligence: 62.4 vs 73.4; Calibration: 90.6 vs 90.0; Speed: 91.5 vs 86.9; Cost: 54.7 vs 50.5.50100Intelligence62.4 · 73.4Calibration90.6 · 90.0Speed91.5 · 86.9Cost54.7 · 50.5
0–100, the values in the table. A system with no published axis draws at 0 and says so.

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Jev 1.13.0 vs Quyet-1.0-Large. Rules, policy & law: 72.0 vs 77.2; Coding & software: 76.5 vs 79.1; Math & numbers: 33.9 vs 50.7; Support & operations: 47.6 vs 66.2; Finance & commerce: 16.1 vs 78.2; Everyday language: 59.1 vs 57.2; Safety & security: 88.8 vs 52.6.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open, 612 sealed).Rules, policy& law72.0 · 77.2Coding & software: code, SQL, repositories, developer tools and IT systems. 302 items (62 open, 240 sealed).Coding &software76.5 · 79.1Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open, 156 sealed).Math &numbers33.9 · 50.7Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open, 72 sealed).Support &operations47.6 · 66.2Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open, 54 sealed).Finance &commerce16.1 · 78.2Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open, 35 sealed).Everydaylanguage59.1 · 57.2Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open, 31 sealed).Safety &security88.8 · 52.6
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open / 612 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 302 items (62 open / 240 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open / 156 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open / 72 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open / 54 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open / 35 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open / 31 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jev 1.13.0 vs Quyet-1.0-Large. Model routing: 81.0 vs 82.1; Legal & compliance: 78.3 vs 85.1; Customer support: 50.5 vs 56.2; E-commerce: — vs 65.0; Risk assessment: — vs 53.4; Insurance claims: — vs 79.8; Financial crime: 50.4 vs 50.0; Moderation: — vs 78.3; Feature extraction: — vs 55.9; Knowledge graphs: — vs 77.8; LLM guardrails: — vs 60.0; Search & retrieval: — vs 92.2; Lead generation: — vs 99.9; Code linting: — vs 17.6; Demand forecasting: — vs 0.9.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open, 356 sealed).Model routing81.0 · 82.1Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open, 313 sealed).Legal &compliance78.3 · 85.1Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open, 112 sealed).Customersupport50.5 · 56.2E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open, 39 sealed).E-commerce65.0Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open, 37 sealed).Riskassessment53.4Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open, 32 sealed).Insuranceclaims79.8Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open, 30 sealed).Financialcrime50.4 · 50.0Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open, 25 sealed).Moderation78.3Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 19 items (5 open, 14 sealed).Featureextraction55.9Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 18 items (3 open, 15 sealed).Knowledgegraphs77.8LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 17 items (5 open, 12 sealed).LLMguardrails60.0Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 17 items (5 open, 12 sealed).Search &retrieval92.2Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 17 items (4 open, 13 sealed).Leadgeneration99.9Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 16 items (3 open, 13 sealed).Code linting17.6Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 16 items (3 open, 13 sealed).Demandforecasting0.9
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; hover a category for its definition and item count.
What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open / 356 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open / 313 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open / 112 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open / 39 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open / 37 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open / 32 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open / 30 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open / 25 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 19 items (5 open / 14 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 18 items (3 open / 15 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 17 items (5 open / 12 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 17 items (5 open / 12 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 17 items (4 open / 13 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 16 items (3 open / 13 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 16 items (3 open / 13 sealed)

Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 145 items (24 open / 121 sealed) — not a use case of its own, so it is counted but not drawn.

Low n (under 15 items, not plotted): Gaming (14), Scientific discovery (14), Recruiting (13), Advertising (11).

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs Quyet-1.0-Large. Choice · open: 79.0 vs 85.0; Choice · sealed: 72.9 vs 82.6; Noul · open: 54.5 vs 71.4; Noul · sealed: 59.6 vs 66.2; Score · open: 63.3 vs 69.3; Score · sealed: 58.3 vs 65.9.50100Choice · open79.0 · 85.0Choice ·sealed72.9 · 82.6Noul · open54.5 · 71.4Noul · sealed59.6 · 66.2Score · open63.3 · 69.3Score ·sealed58.3 · 65.9
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs Quyet-1.0-Large. Easy: 93.9 vs 93.2; Standard: 63.2 vs 74.0; Judge: 67.7 vs 68.6; Hard: 65.6 vs 81.4.50100Easy93.9 · 93.2Standard63.2 · 74.0Judge67.7 · 68.6Hard65.6 · 81.4
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs Quyet-1.0-Large. Easy: 85.1 vs 80.2; Standard: 64.5 vs 79.1; Judge: 86.1 vs 76.6; Hard: 45.0 vs 69.6.50100Easy85.1 · 80.2Standard64.5 · 79.1Judge86.1 · 76.6Hard45.0 · 69.6
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).

Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 v1.6 items was also labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod (same model, questions and taxonomy as v1.5); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) were checked by hand; item-group rules fix the systematic misses (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems saw S u P (1,500 items); hosted API systems saw only A u P (600 items), so their cells cover fewer items.

  • Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
  • Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
SpokeA: Jev 1.13.0B: Quyet-1.0-Large
The four score axes
Intelligence62.473.4
Calibration90.690.0
Speed91.586.9
Cost54.750.5
Capability by subject topic
Rules, policy & law72.077.2
Coding & software76.579.1
Math & numbers33.950.7
Support & operations47.666.2
Finance & commerce16.178.2
Everyday language59.157.2
Safety & security88.852.6
Use cases (TypeSafe categories)
Model routing81.082.1
Legal & compliance78.385.1
Customer support50.556.2
E-commerce—65.0
Risk assessment—53.4
Insurance claims—79.8
Financial crime50.450.0
Moderation—78.3
Feature extraction—55.9
Knowledge graphs—77.8
LLM guardrails—60.0
Search & retrieval—92.2
Lead generation—99.9
Code linting—17.6
Demand forecasting—0.9
Competence per request type, open / sealed
Choice · open79.085.0
Choice · sealed72.982.6
Noul · open54.571.4
Noul · sealed59.666.2
Score · open63.369.3
Score · sealed58.365.9
Competence per tier — open set
Easy93.993.2
Standard63.274.0
Judge67.768.6
Hard65.681.4
Competence per tier — sealed set
Easy85.180.2
Standard64.579.1
Judge86.176.6
Hard45.069.6

Self-hosted cells pool S 1,200 + P 300 (1,500 items); hosted-API cells pool A 300 + P 300 (600 items). Raw and unequated: compare systems of the same lane within a category. Supplementary A2 topic/use-case cells cover P300 only; family/language cells cover A2+P600. Sage A3 topic, use-case, family and language cells cover public P300 only; cells under 15 items are omitted. Sealed counts in the compare view refer to self-hosted systems (S 1,200); hosted APIs answered A 300.

Languages

Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: self-hosted systems over S 1,200 + P 300, hosted APIs over their A or A2 subset + P (600 items). A2 topic/use-case cells cover P300 only. Unequated and outside the Composite. Cells under 15 items are left empty. A dagger (†) marks every displayed cell with fewer than 30 answered items. English (1,160 items) is listed first; the other 21 languages and the mixed-language column share 340 items. This is the v1.6.0 main-pool breakdown. The expanded uc1.1 multilingual pool is a candidate for a later release and is not part of v1.6.0.

JevBench v1.6.0 competence by item language and system
Systemen1160pl31es31pt†29de†27da†22it†17ja†17fr†16mixed†14hi†13zh†13tr†12ar†12uk†12ko†11nl†11cs†11sv†10no†9el†9fi†8id†5
Quyet-1.0-Large · self-hosted74638784†94†64†66†68†74†——————————————
Jev 1.13.0 · API68—81†—87†——————————————————
wity-1 · API7385†——91†——————————————————
torchcast-decision-12b · self-hosted62567573†85†43†56†36†79†——————————————
deck-31B · self-hosted76728389†97†67†72†82†94†——————————————
Winnow-12B Q8 · self-hosted61567876†79†35†45†29†79†——————————————
Cygnet · self-hosted52478275†77†34†35†46†77†——————————————
Jev-Omni · self-hosted52747877†56†40†40†47†54†——————————————
Sage 1.3.0 · API70——————————————————————
Plumb-4B · self-hosted42437861†73†35†22†5†46†——————————————
spark-s1-4b-v6 · self-hosted42297970†71†18†24†33†70†——————————————
swanOne · self-hosted57488668†85†26†45†42†91†——————————————
Quyet-1.0-Medium · self-hosted42217750†72†39†25†18†31†——————————————
lev · self-hosted40326657†45†4†31†3†36†——————————————
NInfer Qwen3.8-Flash-Next mixed · self-hosted49467876†68†27†43†27†85†——————————————
jev-local · self-hosted43247134†49†24†49†15†60†——————————————
decider-4b v2 · self-hosted39197939†61†11†43†25†2†——————————————
JevK5 v0.3 · self-hosted40267942†60†20†13†0†56†——————————————
metask-jev-4b · self-hosted39247261†51†22†41†25†40†——————————————
deck-4B v1.0 · self-hosted40267957†60†18†13†0†57†——————————————
Clef-Flash · self-hosted42287366†64†1†48†0†69†——————————————
Qwen3.5-9B Jev-like data-mix v2 · self-hosted42276644†57†0†31†0†70†——————————————
JevK5 v0.2.0 · self-hosted33327749†70†32†18†26†25†——————————————
JevOne · self-hosted46357661†72†29†42†16†36†——————————————
Hopper · self-hosted33237249†55†5†24†0†34†——————————————
Imajev-4B · self-hosted35207041†58†1†40†7†12†——————————————
Malkuth-4B · self-hosted35196141†47†18†46†11†31†——————————————
jqv · self-hosted37455846†61†0†40†23†35†——————————————
decider-35b-a3b · self-hosted49387347†64†18†22†0†47†——————————————
Manchego v2.1 · self-hosted32155443†66†15†49†5†9†——————————————
Decision 4B v1.2 · self-hosted32167336†67†5†41†0†31†——————————————
Raw Qwen3 4B Instruct 2507 direct logits · self-hosted32375830†55†11†24†9†38†——————————————
Bespoke Nimble 9B · self-hosted42527448†53†3†46†9†55†——————————————
reflex 4B · self-hosted31107657†51†1†22†20†17†——————————————
typecastlm · self-hosted27235528†64†24†27†23†28†——————————————
Decision 4B v1.1 · self-hosted29126846†55†0†33†1†44†——————————————
AutoJev-27B · self-hosted56457454†83†40†50†30†75†——————————————
Eikos-27B · self-hosted66477369†78†45†57†48†65†——————————————
NInfer Qwen3.8-27B NVFP4 · self-hosted51518485†75†28†52†46†92†——————————————
OpenJev · self-hosted76327689†93†29†82†31†66†——————————————
Open-Jev 9B · self-hosted44346756†78†29†59†0†21†——————————————
Standard One 8B · self-hosted35457540†53†12†46†17†50†——————————————
Clef · self-hosted62367780†84†34†70†46†76†——————————————
JEV Qwen3.5-9B Base NVFP4 · self-hosted27116639†62†3†14†0†31†——————————————
SemIf, formerly OpenJev · self-hosted2754939†57†15†31†18†35†——————————————
Bev / Bonsai 27B · self-hosted44305561†72†24†19†43†20†——————————————
LitJev · self-hosted45437569†78†43†35†33†52†——————————————
local-jev Qwen3.5-4B · self-hosted1204227†57†0†16†0†34†——————————————
system-one · self-hosted30144438†48†25†9†8†27†——————————————
kev 8B · self-hosted34276946†38†0†9†0†23†——————————————
Decision 2B · self-hosted1394225†49†4†0†0†0†——————————————
OpenSourceJev · self-hosted19236032†54†14†30†0†3†——————————————
open-alternative-jev · self-hosted21236331†47†13†36†17†14†——————————————
Raw Qwen3 8B direct logits · self-hosted27195325†49†0†19†20†36†——————————————
decider-2b · self-hosted23264419†53†10†41†0†4†——————————————
Nemotron Diffusion 8B · self-hosted14354517†56†0†38†9†44†——————————————
kev 4B · self-hosted23126327†37†0†24†0†4†——————————————
Malkuth-2B · self-hosted13224019†41†0†8†0†0†——————————————
Qwen3-Reranker-4B · self-hosted00120†3†0†0†0†0†——————————————
ZeroEntropy zerank-2 · self-hosted00250†5†0†0†0†0†——————————————
Open-Jev 2B · self-hosted18224919†65†6†15†0†0†——————————————
Open-Jev 27B v1.1 · self-hosted53336964†73†25†34†11†46†——————————————
Raw Phi-4 mini direct logits · self-hosted00181†4†0†23†0†14†——————————————
SimpleJev · self-hosted9143617†51†6†5†0†0†——————————————
smalljev semantic-v9 · self-hosted0330†1†0†11†0†0†——————————————
kev 0.6B · self-hosted0800†13†0†0†0†6†——————————————
OpenDecision · self-hosted0060†21†10†0†0†6†——————————————
Qwen3.5-0.8B Decision Model · self-hosted0100†16†0†0†0†0†——————————————
Deem 0.8B v1 · self-hosted01200†37†0†0†1†0†——————————————
Decision Fast · self-hosted0000†0†0†0†0†0†——————————————
Quyet-1.0-Small-EN · self-hosted0000†0†0†7†0†0†——————————————
Fastino GLiNER-2.5-Decide · API2—25†—28†——————————————————
Laya typed-decisions · self-hosted0000†0†0†0†0†0†——————————————
openJev Verdict 1.4 · self-hosted0000†0†0†0†0†0†——————————————
Raw Qwen3 0.6B direct logits · self-hosted0000†0†9†0†0†23†——————————————
lev-350m · self-hosted0000†0†0†0†0†0†——————————————
Raw Qwen3 1.7B direct logits · self-hosted00026†0†0†0†7†16†——————————————
kev 0.5B · self-hosted0000†0†0†0†0†0†——————————————
openJev Verdict · self-hosted0000†0†0†0†2†0†——————————————
Quyet-1.0-Small · self-hosted0000†5†0†0†2†0†——————————————
Quyet-1.0-Tiny · self-hosted0000†7†0†0†9†0†——————————————
open-jev-deberta-v3-large · self-hosted0000†0†0†0†0†0†——————————————
verdict-small · self-hosted0000†0†0†2†0†0†——————————————
Laya · self-hosted0000†0†0†0†0†0†——————————————
Mixedbread mxbai-rerank-base-v2 · self-hosted0000†0†0†0†0†0†——————————————
Laya multilingual · self-hosted0000†0†0†0†0†0†——————————————
Certo v1 · self-hosted0000†0†0†0†0†0†——————————————
CLM-8B · self-hosted0000†0†0†0†3†0†——————————————
BAAI bge-reranker-v2-m3 · self-hosted0000†0†0†0†0†0†——————————————
Alibaba GTE Reranker ModernBERT-base · self-hosted0000†0†0†0†0†0†——————————————
Mirror · self-hosted0000†0†0†0†0†0†——————————————
Open Jev JSON Canvas · self-hosted62647274†78†25†45†62†94†——————————————

Intelligence gate and Noul decisiveness

The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.

SystemIntelligenceChoiceNoulScoreNoul decisive rateAccuracy among decisive
Quyet-1.0-Large73.4nativenativenative91 %93 %
Jev 1.13.062.4nativenativenative83 %98 %
wity-171.6nativenativenative90 %92 %
torchcast-decision-12b60.5nativenativenative96 %86 %
deck-31B73.0nativenativenative96 %94 %
Winnow-12B Q859.5nativenativenative86 %92 %
Cygnet54.8nativenativenative81 %92 %
Jev-Omni55.5nativenativenative74 %94 %
Sage 1.3.064.3nativenativenative83 %93 %
Plumb-4B43.0below gatenativenativenative71 %91 %
spark-s1-4b-v643.2below gatenativenativenative84 %83 %
swanOne56.2nativenativenative81 %93 %
Quyet-1.0-Medium41.7below gatenativenativenative76 %84 %
lev41.1below gatenativenativenative84 %83 %
NInfer Qwen3.8-Flash-Next mixed49.0below gatenativenativenative75 %93 %
jev-local41.9below gatenativenativenative85 %82 %
decider-4b v240.1below gatenativenativenative64 %92 %
JevK5 v0.338.5below gatenativenativenative66 %95 %
metask-jev-4b39.2below gatenativenativenative75 %88 %
deck-4B v1.038.6below gatenativenativenative68 %93 %
Clef-Flash42.2below gatenativenativenative68 %93 %
Qwen3.5-9B Jev-like data-mix v241.9below gatenativenativenative71 %91 %
JevK5 v0.2.036.3below gatenativenativenative62 %93 %
JevOne46.0below gatenativenativenative62 %96 %
Hopper34.7below gatenativenativenative56 %94 %
Imajev-4B34.6below gatenativenativenative52 %96 %
Malkuth-4B33.8below gatenativenativenative65 %92 %
jqv33.9below gatenativenativenative66 %94 %
decider-35b-a3b48.7below gatenativenativenative74 %91 %
Manchego v2.132.5below gatenativenativenative62 %90 %
Decision 4B v1.232.4below gatenativenativenative59 %94 %
Raw Qwen3 4B Instruct 2507 direct logits34.0below gatenativenativenative92 %73 %
Bespoke Nimble 9B43.1below gatenativenativenative77 %87 %
reflex 4B30.9below gatenativenativenative58 %92 %
typecastlm29.8below gatenativenativenative57 %90 %
Decision 4B v1.129.1below gatenativenativenative57 %94 %
AutoJev-27B55.6nativenativenative70 %99 %
Eikos-27B64.4nativenativenative81 %98 %
NInfer Qwen3.8-27B NVFP451.8nativenativenative82 %89 %
OpenJev70.3nativenativenative96 %93 %
Open-Jev 9B43.9below gatenativenativenative81 %86 %
Standard One 8B32.8below gatenativenativenative68 %90 %
Clef59.1nativenativenative82 %94 %
JEV Qwen3.5-9B Base NVFP429.3below gatenativenativenative53 %89 %
SemIf, formerly OpenJev27.1below gatenativenativenative62 %90 %
Bev / Bonsai 27B43.8below gatenativenativenative70 %91 %
LitJev43.3below gatenativenativenative66 %98 %
local-jev Qwen3.5-4B24.1below gatenativenativenative35 %99 %
system-one29.3below gatenativenativenative92 %73 %
kev 8B29.3below gatenativenativenative77 %81 %
Decision 2B22.9below gatenativenativenative39 %94 %
OpenSourceJev23.2below gatenativenativenative49 %95 %
open-alternative-jev22.3below gatenativenativenative60 %87 %
Raw Qwen3 8B direct logits26.5below gatenativenativenative92 %71 %
decider-2b20.7below gatenativenativenative68 %78 %
Nemotron Diffusion 8B20.4below gatenativenativenative61 %77 %
kev 4B19.5below gatenativenativenative67 %82 %
Malkuth-2B15.9below gatenativenativenative53 %84 %
Qwen3-Reranker-4B16.3below gatenativenativenative25 %75 %
ZeroEntropy zerank-215.6below gatenativenativenative14 %87 %
Open-Jev 2B21.8below gatenativenativenative67 %75 %
Open-Jev 27B v1.152.5nativenativenative67 %94 %
Raw Phi-4 mini direct logits13.7below gatenativenativenative50 %61 %
SimpleJev13.0below gatenativenativenative100 %58 %
smalljev semantic-v98.7below gatenativenativenative23 %73 %
kev 0.6B7.6below gatenativenativenative43 %69 %
OpenDecision7.4below gatenativenativenative51 %68 %
Qwen3.5-0.8B Decision Model7.3below gatenativenativenative22 %88 %
Deem 0.8B v16.5below gatenativenativenative77 %59 %
Decision Fast6.1below gatenativenativenative14 %90 %
Quyet-1.0-Small-EN6.1below gatenativenativenative28 %64 %
Fastino GLiNER-2.5-Decide6.4below gateconfidenceconfidenceconfidence100 %52 %
Laya typed-decisions5.5below gatenativenativenative1 %100 %
openJev Verdict 1.45.0below gatenativenativenative0 %—
Raw Qwen3 0.6B direct logits5.6below gatenativenativenative100 %42 %
lev-350m4.8below gatenativenativenative8 %69 %
Raw Qwen3 1.7B direct logits5.1below gatenativenativenative100 %42 %
kev 0.5B4.6below gatenativenativenative26 %56 %
openJev Verdict4.1below gatenativenativenative33 %60 %
Quyet-1.0-Small4.0below gatenativenativenative29 %71 %
Quyet-1.0-Tiny3.7below gatenativenativenative20 %64 %
open-jev-deberta-v3-large2.8below gatenativenativenative15 %58 %
verdict-small2.3below gatenativenativenative45 %62 %
Laya1.8below gatenativenativenative19 %61 %
Mixedbread mxbai-rerank-base-v21.1below gatenativenativenative0 %—
Laya multilingual0.3below gatenativenativenative72 %50 %
Certo v10.2below gatenativenativenative1 %25 %
CLM-8B0.1below gatenativenativenative51 %69 %
BAAI bge-reranker-v2-m30.1below gatenativenativenative0 %—
Alibaba GTE Reranker ModernBERT-base0.0below gatenativenativenative0 %—
Mirror0.0below gatenativenativenative92 %60 %
Open Jev JSON Canvas61.2labellabellabel100 %85 %
All data (127 systems)

Not yet measured on v1.6 · 28 systems with a dated carried score

Every ranked system of the live v1.5.7 board that is not measured on the v1.6.0 pool keeps its last published score, marked with the release that first published that measurement and that release's publication day. Carried rows are listed separately and never ranked together with v1.6-measured rows. v1.5 protocol (1,624 decisions: 904 open + 720 sealed); scores are on the v1.5 scale and are not comparable with v1.6-measured rows.

Dates: v1.5.0 published 2026-09-28 (25) · v1.5.3 published 2026-09-29 (2) · v1.5.6 published 2026-10-03 (1). The date is the publication day of the release that first published the measurement, not a per-model measurement timestamp.

Carried JevBench v1.5.x results, not ranked with v1.6 measurements
SystemMeasured onCapability (v1.5 scale)v1.5 CompositeCost / 1,000Median latency
Qwen3.8 27B (Chutes TEE)measured on v1.5.0 (2026-09-28)96.80.0 · was #108 on v1.5.7$2.1784estimate6.49 s
GPT-6 Luna (default medium reasoning effort)measured on v1.5.0 (2026-09-28)95.938.8 · was #38 on v1.5.7$0.1138estimate1.56 s
DeepSeek V4.1 Flash (thinking default)measured on v1.5.0 (2026-09-28)95.36.6 · was #72 on v1.5.7$0.4976estimate1.78 s
GPT-6 Luna (low reasoning effort)measured on v1.5.0 (2026-09-28)95.140.5 · was #37 on v1.5.7$0.1075estimate1.58 s
GPT-5.6 Luna (low reasoning effort)measured on v1.5.0 (2026-09-28)94.522.4 · was #51 on v1.5.7$0.2047tariff1.32 s
Autoloops – Gemma 4 31B ITmeasured on v1.5.0 (2026-09-28)81.340.5 · was #36 on v1.5.7$0.1032tariff0.61 s
SimpleJev Qwen3.8-27Bmeasured on v1.5.0 (2026-09-28)80.03.4 · was #74 on v1.5.7$0.6868estimate1.68 s
AutoJev-27B (RTX PRO 6000)measured on v1.5.0 (2026-09-28)79.719.5 · was #56 on v1.5.7$0.2262estimate0.33 s
Surogate Rune 26B-A4B v3 (RTX PRO 6000)Not yet measured on the v1.6 pool.measured on v1.5.0 (2026-09-28)79.066.5 · was #19 on v1.5.7$0.0502estimate0.35 s
Gemini 3.1 Flash-Litemeasured on v1.5.0 (2026-09-28)76.119.6 · was #54 on v1.5.7$0.2194tariff0.86 s
reflex-27b (Qwen3.8-27B)measured on v1.5.0 (2026-09-28)74.413.2 · was #68 on v1.5.7$0.2973estimate2.75 s
Instinct (ZooWork, Qwen3.8-27B)measured on v1.5.0 (2026-09-28)73.918.3 · was #60 on v1.5.7$0.2298estimate0.52 s
NInfer Qwen3.8-27B NVFP4 (T=1.5)measured on v1.5.0 (2026-09-28)73.818.5 · was #59 on v1.5.7$0.2305estimate0.25 s
Vansa-3.4 (Vansa, hosted System One API)Hosted API retained at its v1.5.6 score under the three-refresh exposure cadence.measured on v1.5.6 (2026-10-03)72.871.6 · was #5 on v1.5.7$0.0193estimate0.14 s
openjev-sglang (Qwen3.6-35B-A3B on SGLang)measured on v1.5.0 (2026-09-28)70.829.0 · was #45 on v1.5.7$0.1398estimate1.20 s
Instinct Dual 4Bmeasured on v1.5.3 (2026-09-29)65.747.0 · was #30 on v1.5.7$0.0212estimate0.51 s
system-one-open (Gemma 4 E2B LoRA on an L4)measured on v1.5.0 (2026-09-28)56.642.4 · was #35 on v1.5.7$0.0114estimate1.15 s
decision-machine-1 (milliseconds.ai)measured on v1.5.0 (2026-09-28)48.03.2 · was #75 on v1.5.7$0.0286estimate0.18 s
jeff (Logan Markewich, GLiFormer 400M)measured on v1.5.0 (2026-09-28)42.40.1 · was #89 on v1.5.7$0.0043estimate7.14 s
Von (wfzyx, Option-Marker 395M)measured on v1.5.0 (2026-09-28)41.70.0 · was #110 on v1.5.7$0.0038estimate0.92 s
Bosun v3.1 0.6Bmeasured on v1.5.3 (2026-09-29)39.42.5 · was #78 on v1.5.7$0.0056estimate4.04 s
JevAct (einptein, jev1-2b-v2)measured on v1.5.0 (2026-09-28)36.91.5 · was #81 on v1.5.7$0.0112estimate0.80 s
GLiNER2.5 multi (Fastino, 287M)measured on v1.5.0 (2026-09-28)36.02.7 · was #77 on v1.5.7$0.0028estimate1.37 s
GLiNER2.5 small (Fastino, 74M)measured on v1.5.0 (2026-09-28)32.50.9 · was #85 on v1.5.7$0.0028estimate0.47 s
GLiNER2 large (Fastino)measured on v1.5.0 (2026-09-28)32.48.4 · was #70 on v1.5.7$0.0056estimate1.88 s
GLiNER2 (Fastino, gliner2.5-base)measured on v1.5.0 (2026-09-28)24.52.3 · was #79 on v1.5.7$0.0028estimate1.00 s
Needle 3 (Cactus, 2-bit, local CPU)measured on v1.5.0 (2026-09-28)0.00.0 · was #102 on v1.5.7$0.0191estimate135.51 s
Needle 3, options as tools (post-hoc adapter mode)measured on v1.5.0 (2026-09-28)0.00.0 · was #103 on v1.5.7$0.0191estimate58.16 s

Method · v1.6.0

Rotating item sets

Each release draws fresh sealed decisions from a larger reserve. Self-hosted open-weights models (run offline on our own GPU pods or Sandy) answer S and P; externally hosted models answer only the API subset A and P.

SetItemsChoiceNoulScoreAnswered by
S · Sealed release draw1,200600300300self-hosted systems only (run offline on our own GPU pods or Sandy)
A · API subset (part of S)3001507575self-hosted systems and externally hosted APIs
P · Public set3001507575every system
  • Self-hosted systems: 1,500 items (S 1,200 + P 300). Hosted APIs: 600 items (A 300 + P 300).
  • A sealed item is scored in at most three releases, then retired. An item used in one release is not drawn again in the next. If coverage minimums cannot be met, the release waits for newly reviewed items.
  • The selection seed is committed (SHA-256) before any inference, and the draw is a deterministic function of policy, seed and item id.

API-exposure rule

  • An external exposure is any sealed item sent to an endpoint we do not control: closed APIs, and open-weights models reached through third-party hosts, routers or a submitter's endpoint. Timeouts count as exposure.
  • Hosted models receive only the release's API subset A plus the public set P, and are re-measured at most once every three refresh releases unless a verified new model version ships. Between measurements they keep their last score with its measurement date.
  • Every externally sent sealed item is logged before dispatch. An item exposed to a provider is never scored again for that provider; once two different providers have received it, it retires for everyone. In v1.6.0 Jev 1.13.0 and Fastino GLiNER-2.5-Decide both received the same A, so A retires globally.

Public-versus-sealed gap penalty

Intelligence is half public, half sealed. A system whose public score exceeds its sealed score by more than the field-median gap (G_med = 2.6 points) plus 8 points loses one Intelligence point per excess point. Hosted APIs are compared on P versus A against the same self-hosted systems' P-versus-A gap (-1.2 points).

Comparable scores for hosted APIs

Median over complete full-coverage self-hosted systems of metric(S u P) - metric(A u P), per metric (Intelligence, Calibration); API value = clip(raw_A + offset). Offsets in this run: Intelligence -2.21, Calibration +2.34. Median over the complete full-coverage self-hosted systems of this run (at least five required). Category and language values stay raw and unequated.

Headline and Composite

Capability = mean(Intelligence, Calibration) for systems within twice the Jev 1.13.0 cost and median latency (the official caps; the sliders change only your view). The Composite (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost with the v1.5 low-axis gates; it remains secondary. Request types Choice, Noul and Score weigh equally; tiers weigh easy 0.1, standard 0.2, hard 0.4, judge 0.3.

Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s.

Costs of v1.6-measured systems carry each system's published v1.5.4 cost per 1,000 decisions (pricing rules unchanged; v1.6 item lengths differ) unless the row says otherwise; Fastino's is an estimate from its published tariff and measured tokens.

Noul decisiveness and Score baseline (addendum B)

Scored with method option B (scorer setting O1S), selected on 3 Oct 2026 after the v1.6 scores were known and disclosed as a post-results change. Each split × type competence is clipped at 0 before the type weighting, so a type answered no better than chance counts as chance instead of negative. The Score chance baseline is the error of always predicting the mid-scale level, so a flat know-nothing distribution earns about 0. A Score cell whose golds all sit at mid-scale keeps the v1.5 random-level baseline; this only occurs in small breakdown and bootstrap cells. Calibration is unchanged. This run: G_med = 2.57 points; hosted-API offsets Intelligence -2.21, Calibration +2.34.

Failed requests and very long items

23 items of this draw are very long (about 77,000 to 81,000 input tokens). Systems with a shorter context window refuse them; Jev-Omni runs out of GPU memory on them on its listed RTX 6000 (48 GB) recipe; decider-4b v2's server truncates them to about 32,800 tokens and answers; Plumb-4B reads them in full. Each outcome is the system's own recipe and is scored as such. A failed, refused or unparseable answer counts as wrong for Intelligence and stays in the denominator; it does not enter Calibration (the v1.5 rule, applied to every system). Failed answers per system: Quyet-1.0-Large 0/1500 · Jev 1.13.0 10/600 · wity-1 0/600 · torchcast-decision-12b 23/1500 · deck-31B 23/1500 · Winnow-12B Q8 23/1500 · Cygnet 23/1500 · Jev-Omni 23/1500 · Sage 1.3.0 0/600 · Plumb-4B 0/1500 · spark-s1-4b-v6 0/1500 · swanOne 23/1500 · Quyet-1.0-Medium 0/1500 · lev 25/1500 · NInfer Qwen3.8-Flash-Next mixed 23/1500 · jev-local 23/1500 · decider-4b v2 0/1500 · JevK5 v0.3 23/1500 · metask-jev-4b 23/1500 · deck-4B v1.0 23/1500 · Clef-Flash 0/1500 · Qwen3.5-9B Jev-like data-mix v2 23/1500 · JevK5 v0.2.0 23/1500 · JevOne 0/1500 · Hopper 0/1500 · Imajev-4B 23/1500 · Malkuth-4B 23/1500 · jqv 23/1500 · decider-35b-a3b 0/1500 · Manchego v2.1 23/1500 · Decision 4B v1.2 23/1500 · Raw Qwen3 4B Instruct 2507 direct logits 23/1500 · Bespoke Nimble 9B 23/1500 · reflex 4B 0/1500 · typecastlm 0/1500 · Decision 4B v1.1 23/1500 · AutoJev-27B 23/1500 · Eikos-27B 23/1500 · NInfer Qwen3.8-27B NVFP4 23/1500 · OpenJev 23/1500 · Open-Jev 9B 23/1500 · Standard One 8B 23/1500 · Clef 0/1500 · JEV Qwen3.5-9B Base NVFP4 23/1500 · SemIf, formerly OpenJev 23/1500 · Bev / Bonsai 27B 23/1500 · LitJev 23/1500 · local-jev Qwen3.5-4B 0/1500 · system-one 23/1500 · kev 8B 25/1500 · Decision 2B 23/1500 · OpenSourceJev 23/1500 · open-alternative-jev 0/1500 · Raw Qwen3 8B direct logits 23/1500 · decider-2b 0/1500 · Nemotron Diffusion 8B 23/1500 · kev 4B 29/1500 · Malkuth-2B 23/1500 · Qwen3-Reranker-4B 0/1500 · ZeroEntropy zerank-2 0/1500 · Open-Jev 2B 23/1500 · Open-Jev 27B v1.1 23/1500 · Raw Phi-4 mini direct logits 23/1500 · SimpleJev 23/1500 · smalljev semantic-v9 0/1500 · kev 0.6B 30/1500 · OpenDecision 0/1500 · Qwen3.5-0.8B Decision Model 0/1500 · Deem 0.8B v1 23/1500 · Decision Fast 23/1500 · Quyet-1.0-Small-EN 0/1500 · Fastino GLiNER-2.5-Decide 16/600 · Laya typed-decisions 0/1500 · openJev Verdict 1.4 0/1500 · Raw Qwen3 0.6B direct logits 23/1500 · lev-350m 23/1500 · Raw Qwen3 1.7B direct logits 23/1500 · kev 0.5B 32/1500 · openJev Verdict 0/1500 · Quyet-1.0-Small 0/1500 · Quyet-1.0-Tiny 0/1500 · open-jev-deberta-v3-large 73/1500 · verdict-small 0/1500 · Laya 0/1500 · Mixedbread mxbai-rerank-base-v2 0/1500 · Laya multilingual 0/1500 · Certo v1 0/1500 · CLM-8B 0/1500 · BAAI bge-reranker-v2-m3 0/1500 · Alibaba GTE Reranker ModernBERT-base 0/1500 · Mirror 592/1500 · Open Jev JSON Canvas 23/1500 · wity-1 0/600 · wity-1 0/600.

Calibration basis

Systems that return a full probability distribution are calibrated on all components (top-label error, plus distribution distance for Choice and ranked-probability error for Score). Fastino GLiNER-2.5-Decide returns a single confidence value, so its Calibration is the top-label error only and is not like-for-like with full-distribution systems.

Overnight full re-measure (4–5 Oct 2026)

Every system with a reproducible recipe was re-run on the v1.6.0 pool overnight with the same pinned inputs and scorer (method option B / O1S). This page uses scoring round score-overnight-12 (2026-10-05 09:06:16 UTC). Only complete runs (1,500 items self-hosted, the full API input for hosted APIs) are ranked; partial runs are never ranked, and systems not yet re-measured keep their dated v1.5.x score in the separate table.

  • Scores are the official v1.6.0 scorer (score_v16.py, method option B / O1S, bootstrap B = 1,000) over complete outputs only; each scoring round is kept separately.
  • Hosted APIs are scored on their API subset plus P and equated to the self-hosted scale; under the exposure rule the v1.6.0 subset A is retired, so newly measured APIs use a fresh supplementary draw.
  • GPU-class deviations: reproducible recipes used the hardware listed per row, including H100 for large fast-lane decoders and RTX 6000 for the baseline; hardware differences remain in measured latency. The standard x2 + 0.15 s self-hosted adjustment is an assumption, not a hardware normalization.
  • Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s. Wity auto remains outside the latency cap and stays ranked in Composite A. OFF and ALWAYS are unranked variants of the AUTO main row.
  • 6 further measured candidates await a separate publication decision.
  • Sage 1.3.0 (Levanto Labs) was measured on 5 Oct from Sandy (Helsinki), text only. Cost uses the Levanto list tariff (USD 0.05/M input, USD 10/M output; levanto.ai/pricing, read 5 Oct) over all 600 answered rows: USD 0.0766/1,000 decisions. It exceeds the frozen Jev-class cost cap and is excluded from the Capability headline.
  • A fresh Monday Jev 1.13.0 re-check on A3 measured p50 0.239 s and equated Capability 77.5 versus the official 76.5, within the confidence interval; the official v1.6.0 Jev row is retained.

Supplementary API draws A2 and A3

The original API subset A is retired after exposure to two providers. Newly measured hosted APIs use a fresh supplementary A2 draw of 300 sealed items (150 Choice, 75 Noul, 75 Score), disjoint from S and P, with the seed committed before the draw, plus the same public P300. A2 has reached four external entities, including hosting proxies, and is retired for future draws. API A2 scores are equated using the median S+P minus A2+P offset over 12 self-hosted systems: Intelligence -0.29, Calibration +1.56. That pool leans toward weaker systems; its main-A offset differs from the full-pool offset by about 1.8 Intelligence points, so strong API rows may have roughly ±2 points of equating bias. Bootstrap intervals cover pool resampling, not this selection effect. A2 topic and use-case radar cells cover P300 only because sealed A2 items have no topic labels; cells under 15 items are omitted. Family and language cells cover A2+P600. Sage 1.3.0 uses the fresh supplementary A3 subset: 300 never-exposed sealed items plus the same public P300, equated over 15 self-hosted systems (Intelligence offset +1.34, Calibration +3.72). A3 was sent to TypeSafe and Levanto and is now retired. Sage category and language cells cover public P300 only, with cells under 15 items omitted.

Per-model exposure counts (hosted and author-hosted endpoints)

System (provider)v1.6 sealed items sentScored sealed setStatus
wity-1-auto (wity via railway)300A2 (300 sealed) + P (300 public)measured on v1.6.0 A2 u P
wity-1-off (wity via railway)300A2 (300 sealed) + P (300 public)measured on v1.6.0 A2 u P
wity-1-always (wity via railway)300A2 (300 sealed) + P (300 public)measured on v1.6.0 A2 u P
jev-1.13.0 (typesafe)600A (300 sealed, official) + A3 (300 sealed, re-check) + P (300 public)official A measurement retained; 5 Oct A3 re-check confirmed it; A3 now retired
fastino-gliner-2-5-decide (fastino)300A (300 sealed) + P (300 public)measured on v1.6.0 A u P
gpt-6-luna (openai)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
gpt-6-luna-low (openai)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
gpt-5.6-luna (openai)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
gemini-3.1-flash-lite (google)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
deepseek-flash (deepseek)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
qwen3.8-27b (chutes)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
decision-machine-1 (milliseconds.ai)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
classifier-dev-fast (classifier.dev)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
vansa-3.4 (vansa)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
kushal-gemma4-31b-it-autoloops (autoloops)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
system-one-open (modal via modal)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
openjev-sglang (modal via modal)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
instinct (zoowork)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
instinct-dual-4b (zoowork)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
simplejev-qwen3.8-27b (featherless)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
simplejev-qwen3.6-35b-a3b (featherless)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
jevact (jevact)0v1.5 sealed-720 (retired for scoring)carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope)
Sage 1.3.0 (Levanto Labs)300A3 (300 sealed) + P (300 public)measured 5 Oct 2026; A3 exposed to TypeSafe and Levanto, now retired

Provenance

Measured on the v1.6 pool (run completion day, UTC): 2026-10-01: Jev 1.13.0 · 2026-10-02: Winnow-12B Q8, Cygnet, Jev-Omni, Plumb-4B, swanOne, NInfer Qwen3.8-Flash-Next mixed, jev-local, decider-4b v2, JevK5 v0.3, metask-jev-4b, Qwen3.5-9B Jev-like data-mix v2, JevK5 v0.2.0, JevOne, Hopper, Imajev-4B, Malkuth-4B, jqv, decider-35b-a3b, Decision 4B v1.2, Raw Qwen3 4B Instruct 2507 direct logits, Bespoke Nimble 9B, reflex 4B, typecastlm, Decision 4B v1.1, AutoJev-27B, Eikos-27B, NInfer Qwen3.8-27B NVFP4, OpenJev, Open-Jev 9B, Standard One 8B, JEV Qwen3.5-9B Base NVFP4, SemIf, formerly OpenJev, LitJev, local-jev Qwen3.5-4B, system-one, kev 8B, OpenSourceJev, open-alternative-jev, Raw Qwen3 8B direct logits, decider-2b, kev 4B, Malkuth-2B, Qwen3-Reranker-4B, ZeroEntropy zerank-2, Open-Jev 2B, Raw Phi-4 mini direct logits, SimpleJev, smalljev semantic-v9, kev 0.6B, OpenDecision, Qwen3.5-0.8B Decision Model, openJev Verdict 1.4, Raw Qwen3 0.6B direct logits, Raw Qwen3 1.7B direct logits, kev 0.5B, openJev Verdict, open-jev-deberta-v3-large, verdict-small, Laya, Mixedbread mxbai-rerank-base-v2, Certo v1, CLM-8B, BAAI bge-reranker-v2-m3, Alibaba GTE Reranker ModernBERT-base, Mirror, Open Jev JSON Canvas · 2026-10-03: Fastino GLiNER-2.5-Decide · 2026-10-04: Quyet-1.0-Large, torchcast-decision-12b, Quyet-1.0-Medium, lev, Clef-Flash, Manchego v2.1, Clef, Bev / Bonsai 27B, Nemotron Diffusion 8B, Open-Jev 27B v1.1, Deem 0.8B v1, Quyet-1.0-Small-EN, Laya typed-decisions, Quyet-1.0-Small, Quyet-1.0-Tiny, Laya multilingual · 2026-10-05: deck-31B, Sage 1.3.0, spark-s1-4b-v6, deck-4B v1.0, Decision 2B, Decision Fast, lev-350m.

Aggregate files: results sha256 1161dbbfa1a3225008d14d42dc1ccabfd0de51be8bb23251ba929533b840dabb · categories sha256 e1150bf0c6750ed1a500e04ac08f63b8dd0be89002250b73b206d5a5d1bfd295 · dated carry sha256 334e851944e92383ab3071e52b3012040c7023d0fe236395953169ebee8f0459. Scoring source sha256 04ad6eb3cfc06e78784a51629c8e2acab3e0acbe8134f3f00bd72bfcf61c949b. The method, release data and carry artifact are independently hashable.

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.