JevBench v1.4.2.2

JevBench by Benchmark Heaven

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Making decisions from images? Explore Image JevBench v0.1.2 and compare its systems.

Scored 27 Sept 2026 · protocol jevbench::v1.4 · 534 public + 308 sealed aggregate decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 7f39b2f742a6… · v1.0 results

Share this version · View current JevBench release

JevBench v1.4.2.2 · headline ranking

Capability ranking of Jev-class systems

Capability averages Intelligence and Calibration. Imajev-4B leads the Jev-class systems with 66.3.

Jev-class means at most 2× Jev's cost and median latency. How we choose ↘

#SystemICCap.$/1k
  1. 1Imajev-4B52.259.766.3$0.022*
  2. 2Jev 1.13.053.152.064.7$0.040Reference system for the Jev-class limits
  3. 3Plumb-4B53.055.864.2$0.030*
  4. 4Hopper48.058.763.5$0.024*
  5. 5Cygnet49.552.862.2$0.037*
  6. 6decider-4b v249.460.962.2$0.020*
  7. –classifier.dev51.684.362.0$0.0033*Not ranked in the official JevBench Score (honorable mention)
  8. 7JevK5 v0.2.048.959.561.7$0.022*
  9. 8ZeroEntropy zerank-242.149.858.9$0.047
  10. 9JEV Qwen3.5-9B Base NVFP446.843.357.3$0.077*
Show all 52 Jev-class systems (42 more)
  1. 10Winnow-12B Q848.352.956.6$0.037*
  2. 11Decision 2B38.862.556.4$0.018*
  3. 12decider-35b-a3b47.245.356.2$0.067*
  4. 13metask-jev-4b44.754.555.8$0.033*
  5. 14SemIf44.459.555.6$0.022*
  6. 15Jev-Omni46.853.055.4$0.037*
  7. 16Jobe Qwen3.5-4B44.159.555.1$0.022*
  8. 17Qwen3-Reranker-4B44.649.254.9$0.050
  9. 18decision-machine-141.353.754.8$0.035
  10. 19lev-350m34.876.152.7$0.0063*
  11. 20Malkuth-4B43.161.552.3$0.019*
  12. 21Decision Fast37.176.151.2$0.0063*
  13. 22djev47.057.651.2$0.026
  14. 23OpenSourceJev41.864.051.1$0.016*
  15. 24openJev Verdict 1.429.482.450.7$0.0039*
  16. 25Raw Phi-4 mini direct logits41.849.650.3$0.048*
  17. 26OpenJev45.445.550.2$0.066*
  18. 27system-one-open44.264.849.5$0.015*
  19. 28open-alternative-jev38.659.648.6$0.022*
  20. 29typecastlm34.060.248.0$0.021*
  21. 30Malkuth-2B41.361.547.4$0.019*
  22. 31spark-s1-4b-v645.157.946.3$0.025*
  23. 32Mixedbread mxbai-rerank-base-v26.867.945.5$0.012
  24. 33BAAI bge-reranker-v2-m35.073.444.6$0.0077
  25. 34OpenDecision31.875.344.4$0.0066*
  26. 35verdict-small18.196.642.5$0.0013*
  27. 36smalljev semantic-v925.757.942.4$0.025*
  28. 37JevAct29.364.342.2$0.015*
  29. 38Alibaba GTE Reranker ModernBERT-base4.869.641.9$0.010
  30. 39Certo v10.1100.041.5$0.00097*
  31. 40decider-2b38.561.041.0$0.020*
  32. 41kev 4B42.161.840.9$0.019*
  33. 42GLiNER2.5 multi23.182.440.1$0.0039*
  34. 43kev 0.5B30.576.140.1$0.0063*
  35. 44openJev Verdict30.083.138.5$0.0037*
  36. 45Raw Qwen3 4B Instruct 2507 direct logits46.459.737.7$0.022*
  37. 46GLiNER2.5 small20.582.435.6$0.0039*
  38. 47CLM-8B22.478.431.1$0.0052*
  39. 48Raw Qwen3 1.7B direct logits33.264.928.8$0.015*
  40. 49GLiNER227.483.126.3$0.0037*
  41. 50Open Jev JSON Canvas48.245.624.1$0.065*
  42. 51Raw Qwen3 0.6B direct logits22.873.921.8$0.0074*

Wide coloured bar = Capability (0–100). Thin red line = cost per 1,000 decisions; log scale, each gridline = 10×, shorter is cheaper. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ for the full values, median latency and cost relative to Jev.

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • system-one-open
  • Cost line
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability, not numbered. Each row says which limit it misses, measured against Jev 1.13.0 ($0.040 per 1,000 decisions, median 0.65 s).

  1. –GPT-6 Luna (medium)97.436.095.4$0.14Outside: cost 3.4× Jev, latency 2.3× Jev
  2. –DeepSeek V4.1 Flash94.016.894.7$0.59Outside: cost 14.9× Jev, latency 2.2× Jev
  3. –GPT-6 Luna (low)95.836.993.9$0.13Outside: cost 3.2× Jev, latency 2.2× Jev
  4. –GPT-5.6 Luna93.128.590.3$0.24Outside: cost 6.1× Jev
  5. –djev71.626.979.7$0.27*Outside: cost 6.9× Jev
  6. –Autoloops – Gemma 4 31B IT59.836.069.5$0.14Outside: cost 3.4× Jev
  7. –Qwen3.8 27B40.40.067.0$2.67*Outside: cost 66.9× Jev, latency 8.8× Jev
  8. –Instinct51.124.664.9$0.33*Outside: cost 8.2× Jev
  9. –NInfer Qwen3.8-Flash-Next mixed49.538.964.1$0.11*Outside: cost 2.7× Jev
  10. –NInfer Qwen3.8-27B NVFP451.535.263.7$0.14*Outside: cost 3.6× Jev
  11. –SimpleJev Qwen3.8-27B51.639.563.0$0.10*Outside: cost 2.6× Jev, latency 3.1× Jev
  12. –JevOne47.635.962.6$0.14*Outside: cost 3.4× Jev
  13. –reflex-27b46.432.361.8$0.18*Outside: cost 4.5× Jev, latency 6.0× Jev
  14. –swanOne52.638.661.8$0.11*Outside: cost 2.8× Jev
  15. –LitJev46.333.661.4$0.16*Outside: cost 4.1× Jev, latency 6.4× Jev
  16. –openjev-sglang49.436.559.3$0.13*Outside: cost 3.3× Jev, latency 2.1× Jev
  17. –NInfer Qwen3.8-27B NVFP451.535.259.3$0.14*Outside: cost 3.6× Jev
  18. –jqv46.447.559.0$0.056*Outside: latency 2.5× Jev
  19. –reflex 4B47.559.758.9$0.022*Outside: latency 5.7× Jev
  20. –local-jev Qwen3.5-4B44.455.858.9$0.030*Outside: latency 2.4× Jev
  21. –OpenJev58.127.858.1$0.25*Outside: cost 6.4× Jev
  22. –Gemini 3.1 Flash-Lite54.527.456.9$0.26Outside: cost 6.6× Jev
  23. –Standard One 8B46.239.456.0$0.10*Outside: cost 2.6× Jev
  24. –Von34.577.855.1$0.0055*Outside: Speed below the 2× latency line
  25. –jev-local45.243.354.7$0.077*Outside: latency 3.4× Jev
  26. –Qwen3.5-9B Jev-like data-mix v247.442.454.4$0.083*Outside: cost 2.1× Jev
  27. –Open-Jev 9B44.228.153.0$0.25*Outside: cost 6.2× Jev, latency 2.5× Jev
  28. –SimpleJev Qwen3.6-35B-A3B45.738.152.8$0.12*Outside: cost 2.9× Jev, latency 2.6× Jev
  29. –jeff36.876.652.3$0.0060*Outside: latency 3.1× Jev
  30. –Bespoke Nimble 9B46.333.451.4$0.17*Outside: cost 4.2× Jev
  31. –Laya36.186.249.9$0.0029*Outside: latency 2.6× Jev
  32. –Open-Jev 2B42.328.148.8$0.25*Outside: cost 6.2× Jev, latency 2.3× Jev
  33. –Qwen3.5-0.8B Decision Model28.175.748.2$0.0065*Outside: latency 22.1× Jev
  34. –open-jev-deberta-v3-large25.674.046.1$0.0073*Outside: latency 5.6× Jev
  35. –kev 0.6B34.276.142.1$0.0063*Outside: latency 2.0× Jev
  36. –kev 8B41.844.041.0$0.073*Outside: latency 2.0× Jev
  37. –system-one43.641.538.2$0.089*Outside: cost 2.2× Jev
  38. –SimpleJev21.568.335.3$0.011*Outside: latency 12.5× Jev
  39. –Raw Qwen3 8B direct logits45.741.934.9$0.087*Outside: cost 2.2× Jev
  40. –GLiNER2 large31.173.328.0$0.0077*Outside: latency 3.6× Jev
  41. –Mirror13.673.319.8$0.0077*Outside: latency 2.8× Jev

Jev-class = cost per decision at most 2× Jev 1.13.0's (≤ $0.080 per 1,000 decisions) and median latency at most 2× Jev 1.13.0's (≤ 1.30 s, the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score). 52 of 93 systems qualify; the other 41, including the general-purpose LLMs, are listed below the divider in the ranking. JevK5 v0.2.0, OpenSourceJev have no recorded median latency (carried from v1.3); for them the Speed axis decides, at the 2× latency equivalent (Speed ≥ 77.2). The charts below show speed and cost beside Capability; the official JevBench Score weighs all four axes.

Capability against cost and speed

Jev-class systems are shown by default. Bubble size follows the official JevBench Score. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 52 systems. Upper right is best: more capable and cheaper.2030405060708090100$0.0010$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev cost← priciercheaper →1. Imajev-4B2. Jev 1.13.03. Plumb-4B4. Hopper5. Cygnet
52 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 52 systems. Upper right is best: more capable and faster.203040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev latency← slowerfaster →1. Imajev-4B2. Jev 1.13.03. Plumb-4B4. Hopper5. Cygnet
52 systems. Tap a bubble for its values.

JevBench v1.4.2.2 ranking

JevBench v1.4.2.2

JevBench Composite Score: 91 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · What changed in v1.4 ↓

Weights:
Adjust weights ↓
View by:Capability ↑

Imajev-4B leads the JevBench Score at 67.37, ahead of Plumb-4B (65.84). The v1.4.2 scoring code and earlier measurement rows are unchanged.

Greener = stronger within its column.

95 of 95 systems, sorted by official rank, #1 first.

  1. 1Imajev-4Bnew67.4I 52C 80S 91K 60est.$0.022
  2. 2Plumb-4B65.8I 53C 75S 93K 56est.$0.030
  3. 3decider-4b v264.1I 49C 75S 93K 61est.$0.020
  4. 4Jev 1.13.0API63.3I 53C 76S 83K 52$0.040
  5. 5JevK5 v0.2.062.0I 49C 75S 91K 60est.$0.022
  6. 6Cygnet61.8I 50C 75S 91K 53est.$0.037
  7. 7Hopper59.4I 48C 79S 87K 59est.$0.024
  8. 8Winnow-12B Q855.6I 48C 65S 82K 53est.$0.037
  9. 9reflex 4B54.0I 47C 70S 68K 60est.$0.022
  10. 10djev (Maisa, diffusion-gemma)52.2I 47C 55S 91K 58ann.$0.026
  11. 11Jev-Omni51.3I 47C 64S 82K 53est.$0.037
  12. 12metask-jev-4b47.8I 45C 67S 89K 55est.$0.033
  13. 13SemIf47.7I 44C 67S 84K 59est.$0.022
  14. 14Jobe Qwen3.5-4B46.9I 44C 66S 86K 60est.$0.022
  15. 15local-jev Qwen3.5-4B46.8I 44C 73S 75K 56est.$0.030
  16. 16system-one-openAPI45.1I 44C 55S 77K 65est.$0.015
  17. 17spark-s1-4b-v644.6I 45C 48S 81K 58est.$0.025
  18. 18Malkuth-4B44.5I 43C 61S 88K 62est.$0.019
  19. 19jqv44.4I 46C 72S 75K 47est.$0.056
  20. 20Qwen3-Reranker-4B43.5I 45C 65S 79K 49$0.050
Show all 95 systems (71 more ranked, 4 more not ranked)
  1. 21decider-35b-a3b41.2I 47C 65S 81K 45est.$0.067
  2. 22Raw Qwen3 4B Instruct 2507 direct logits41.0I 46C 29S 88K 60est.$0.022
  3. 23OpenSourceJev40.9I 42C 60S 82K 64est.$0.016
  4. 24ZeroEntropy zerank-240.2I 42C 76S 79K 50$0.047
  5. 25decision-machine-1API39.9I 41C 68S 93K 54$0.035
  6. 26Malkuth-2B38.9I 41C 53S 91K 62est.$0.019
  7. 27Raw Phi-4 mini direct logits38.0I 42C 59S 89K 50est.$0.048
  8. 28JEV Qwen3.5-9B Base NVFP437.7I 47C 68S 93K 43est.$0.077
  9. 29OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)36.9I 45C 55S 83K 45est.$0.066
  10. 30kev 4B36.1I 42C 40S 76K 62est.$0.019
  11. 31Decision 2B35.8I 39C 74S 84K 63est.$0.018
  12. 32Qwen3.5-9B Jev-like data-mix v235.2I 47C 61S 82K 42est.$0.083
  13. 33GPT-6 Luna (low reasoning effort)API35.1I 96C 92S 74K 37$0.127
  14. 34SimpleJev Qwen3.8-27BAPI34.6I 52C 74S 71K 39est.$0.104
  15. 35NInfer Qwen3.8-Flash-Next mixed34.0I 50C 79S 88K 39est.$0.109
  16. 36swanOne33.6I 53C 71S 83K 39est.$0.111
  17. 37GPT-6 Luna (default medium reasoning effort)API33.3I 97C 93S 73K 36$0.135
  18. 38open-alternative-jev33.2I 39C 59S 83K 60est.$0.022
  19. 39jev-local32.5I 45C 64S 69K 43est.$0.077
  20. 40Decision Fast32.5I 37C 65S 82K 76est.$0.0063
  21. 41decider-2b30.7I 39C 43S 83K 61est.$0.020
  22. 42jeff30.6I 37C 68S 63K 77est.$0.0060
  23. 43Laya30.3I 36C 64S 71K 86est.$0.0029
  24. 44Autoloops – Gemma 4 31B ITAPI30.0I 60C 79S 84K 36$0.136
  25. 45Standard One 8B29.1I 46C 66S 92K 39est.$0.104
  26. 46lev-350m28.5I 35C 71S 85K 76est.$0.0063
  27. 47openjev-sglangAPI27.7I 49C 69S 77K 36est.$0.131
  28. 48Von27.5I 34C 76S 70K 78est.$0.0055
  29. 49NInfer Qwen3.8-27B NVFP4 (T=1.5)26.9I 51C 76S 80K 35est.$0.145
  30. 50NInfer Qwen3.8-27B NVFP426.3I 51C 67S 80K 35est.$0.145
  31. 51kev 8B25.6I 42C 40S 75K 44est.$0.073
  32. 52JevOne25.5I 48C 78S 88K 36est.$0.137
  33. 53typecastlm25.3I 34C 62S 93K 60est.$0.021
  34. 54SimpleJev Qwen3.6-35B-A3BAPI24.9I 46C 60S 75K 38est.$0.116
  35. 55kev 0.6B24.8I 34C 50S 76K 76est.$0.0063
  36. 56Raw Qwen3 8B direct logits23.7I 46C 24S 86K 42est.$0.087
  37. 57system-one23.4I 44C 33S 84K 41est.$0.089
  38. 58OpenDecision21.6I 32C 57S 80K 75est.$0.0066
  39. 59LitJev19.5I 46C 77S 67K 34est.$0.163
  40. 60openJev Verdict 1.419.0I 29C 72S 78K 82est.$0.0039
  41. 61kev 0.5B18.9I 31C 50S 77K 76est.$0.0063
  42. 62Bespoke Nimble 9B18.7I 46C 56S 79K 33est.$0.166
  43. 63GPT-5.6 LunaAPI18.5I 93C 87S 78K 28$0.242
  44. 64openJev Verdict18.1I 30C 47S 77K 83est.$0.0037
  45. 65Raw Qwen3 1.7B direct logits18.1I 33C 24S 90K 65est.$0.015
  46. 66reflex-27b17.8I 46C 77S 67K 32est.$0.181
  47. 67JevActAPI16.9I 29C 55S 76K 64est.$0.015
  48. 68djev (thinking)15.2I 72C 88S 75K 27est.$0.274
  49. 69GLiNER2 large15.1I 31C 25S 62K 73est.$0.0077
  50. 70OpenJev (thinking, BF16)14.8I 58C 58S 76K 28est.$0.255
  51. 71Qwen3.5-0.8B Decision Model14.5I 28C 68S 49K 76est.$0.0065
  52. 72Gemini 3.1 Flash-LiteAPI14.3I 54C 59S 82K 27$0.264
  53. 73open-jev-deberta-v3-large12.6I 26C 67S 66K 74est.$0.0073
  54. 74smalljev semantic-v912.3I 26C 59S 80K 58est.$0.025
  55. 75GLiNER211.8I 27C 25S 72K 83est.$0.0037
  56. 76InstinctAPI11.4I 51C 79S 84K 25est.$0.325
  57. 77Open-Jev 9B11.2I 44C 62S 72K 28est.$0.249
  58. 78Open-Jev 2B10.0I 42C 55S 73K 28est.$0.249
  59. 79GLiNER2.5 multi9.8I 23C 57S 68K 82est.$0.0039
  60. 80CLM-8B8.6I 22C 40S 94K 78est.$0.0052
  61. 81SimpleJev7.5I 21C 49S 57K 68est.$0.011
  62. 82GLiNER2.5 small7.2I 20C 51S 78K 82est.$0.0039
  63. 83Raw Qwen3 0.6B direct logits7.1I 23C 21S 90K 74est.$0.0074
  64. 84verdict-small5.7I 18C 67S 85K 97est.$0.0013
  65. 85DeepSeek V4.1 FlashAPI4.8I 94C 96S 72K 17$0.594
  66. 86Mirror2.1I 14C 26S 71K 73est.$0.0077
  67. 87Mixedbread mxbai-rerank-base-v20.4I 7C 84S 88K 68$0.012
  68. 88BAAI bge-reranker-v2-m30.2I 5C 84S 90K 73$0.0077
  69. 89Alibaba GTE Reranker ModernBERT-base0.2I 5C 79S 91K 70$0.010
  70. 90Certo v10.0I 0C 83S 94K 100est.$0.0010
  71. 91Open Jev JSON Canvas0.0I 48C 0S 84K 46est.$0.065
  72. classifier.dev (honorable mention)API70.8I 52C 72S 88K 84est.$0.0033
  73. Qwen3.8 27B (partial run)API0.0I 40C 94S 61K 0est.$2.669
  74. Needle 3, options as tools (partial run)—I –C –S –K –est.$0.014
  75. Needle 3 (partial run)—I –C –S –K –est.$0.024
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; an axis at 0 drops out together with its gate. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • system-one-open
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers; new = first listed in v1.4.2.2; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

Compare two systems

Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, accuracy by family on the current v1.4 question set (hard tier and sealed set together), and the sealed set alone. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#4)
  • B: Imajev-4B — Jev rebuild · Score 67.4 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Imajev-4B. Intelligence: 53.1 vs 52.2; Calibration: 76.3 vs 80.4; Speed: 83.3 vs 90.6; Cost: 52.0 vs 59.7.50100Intelligence53.1 · 52.2Calibration76.3 · 80.4Speed83.3 · 90.6Cost52.0 · 59.7
0–100, the values in the table. A label-only system has no calibration (counted as 0).

Accuracy per tier, incl. sealed

Radar: accuracy per tier, incl. sealed, two systemsAccuracy per tier, incl. sealed, Jev 1.13.0 vs Imajev-4B. Easy: 100% vs 100%; Standard: 99% vs 99%; Judge: 95% vs 89%; Hard: 74% vs —; Sealed: 37% vs 37%.50100Easy100% · 100%Standard99% · 99%Judge95% · 89%Hard74%Sealed37% · 37%
Share correct per tier; Sealed = the 308 private decisions, aggregate only.

Current question set by family (hard + sealed)

Imajev-4B has no published hard-tier family breakdown; families that need it are left out (—).

Radar: current question set by family (hard + sealed), two systemsCurrent question set by family (hard + sealed), Jev 1.13.0 vs Imajev-4B. Ambiguous / abstain: 43% vs —; Judge: 54% vs —; Long policy: 44% vs —; Multi-hop: 64% vs —; Probability: 63% vs —; Temporal / numeric: 28% vs —; Trade-off: 55% vs —; Routing: 100% vs —; Trap / adversarial: 83% vs —; Paraphrase: 64% vs 50%; Safety judge: 38% vs 44%.50100Ambiguous /abstain43%Judge54%Long policy44%Multi-hop64%Probability63%Temporal /numeric28%Trade-off55%Routing100%Trap /adversarial83%Paraphrase64% · 50%Safety judge38% · 44%
Share correct per family across the 220 hard-tier decisions (public and held out) and the 308 sealed decisions of v1.4, pooled; Routing is hard-tier only, Paraphrase and Safety judge sealed only.

Sealed set by family

Radar: sealed set by family, two systemsSealed set by family, Jev 1.13.0 vs Imajev-4B. Ambiguous / abstain: 30% vs 41%; Judge: 34% vs 44%; Long policy: 28% vs 33%; Multi-hop: 45% vs 34%; Paraphrase: 64% vs 50%; Probability: 50% vs 25%; Safety judge: 38% vs 44%; Temporal / numeric: 29% vs 30%; Trade-off: 38% vs 38%; Trap / adversarial: 42% vs 58%.50100Ambiguous /abstain30% · 41%Judge34% · 44%Long policy28% · 33%Multi-hop45% · 34%Paraphrase64% · 50%Probability50% · 25%Safety judge38% · 44%Temporal /numeric29% · 30%Trade-off38% · 38%Trap /adversarial42% · 58%
Share correct within each sealed family — system-level aggregates; the items stay private.
All values as a table
SpokeA: Jev 1.13.0B: Imajev-4B
The four score axes
Intelligence53.152.2
Calibration76.380.4
Speed83.390.6
Cost52.059.7
Accuracy per tier, incl. sealed
Easy100%100%
Standard99%99%
Judge95%89%
Hard74%—
Sealed37%37%
Current question set by family (hard + sealed)
Ambiguous / abstain43%—
Judge54%—
Long policy44%—
Multi-hop64%—
Probability63%—
Temporal / numeric28%—
Trade-off55%—
Routing100%—
Trap / adversarial83%—
Paraphrase64%50%
Safety judge38%44%
Sealed set by family
Ambiguous / abstain30%41%
Judge34%44%
Long policy28%33%
Multi-hop45%34%
Paraphrase64%50%
Probability50%25%
Safety judge38%44%
Temporal / numeric29%30%
Trade-off38%38%
Trap / adversarial42%58%

Axes, accuracy, latency and cost

Every system with its four axes, public and sealed accuracy and the gap between them. Click a column heading to sort; filter by name, type, openness, API flag or release. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.

Greener = stronger within its column; faster counts as stronger. Badges: API sealed text went to the operator's endpoint · est. estimated price · ann. announced price not yet bookable · not ranked partial run or honorable mention · new first listed in v1.4.2.2.

Endpoint
1
Imajev-4B
†Imajev-4B: requested adapter c9e5f132465da85d31735ec502d5557982671a7d and server a0134749e0900189c129cd6bb5000969f3b64bb5; one rotation with calibration.json. The optimized image included flash-linear-attention/fla-core 0.5.2 and causal-conv1d 1.7.0; CUDA profiling confirmed the pinned kernels ran. Cost is an estimate using the public DeepInfra Qwen/Qwen3.5-4B base-model reference price ($0.03/M input, $0.15/M output), with no generated output tokens; it is not a GPU bill.
newby mohit67890 · details
67.452.280.490.659.786.1%37.0%+49.1 ppest.$0.0220.04 sRunPod GPU
2
Plumb-4B
†Crh225/plumb-4b @ 55de037801a8a9b9de3db5c0e16cef86210c2186: merged bf16 weights, LoRA r16 on the attention projections of JevK5 v0.2 (itself Qwen3.5-4B), served by the author's documented command, jevk5 v0.2.0's own jevk5-serve (github.com/allebee/jevk5, Apache-2.0), unmodified. One forward pass per decision: softmax over the declared options' answer-letter logits at the last position divided by the package's own T = 2.07 (jevk5_config.json), up to 16 options, 0 generated tokens; inputs over 16,384 tokens are refused rather than truncated. Disclosed by the author: no JevBench item was trained or tuned on and every checkpoint and temperature choice came from his own held-out sets, but the 231 public items were scored after each of five training rounds and that aggregate feedback shaped the recipe (hard mining, long documents, document length) — development against the public distribution, stated plainly. His own overlap audit dropped 9 training items sharing more than two 8-word spans with a public item. Independent top-five gate (25 Sep 2026): LEGIT — cost basis (same as its JevK5 parent), public-versus-sealed gap against like-for-like rows, an 8-gram screen of its released training data (no row shares more than two 8-grams with a public item) and an exact composite recomputation. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium H100 80 GB (Hopper) pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate at the EmpirioLabs Qwen3.5-4B public pay-as-you-go list price ($0.04/M input, no output generated), applied to this system's measured tokens; it is not a GPU bill. 25 Sep 2026 pre-release correction: retired DeepInfra $0.03/M input reference replaced by bookable EmpirioLabs $0.04/M exact-base input reference; draft score 67.09 to 65.84.
by crh225 · crh225, JevK5 v0.2 + LoRA · details
65.853.075.593.555.889.6%38.0%+51.6 ppest.$0.0300.02 sRunPod GPU
3
decider-4b v2
†Decider-ai 1.2.2 (PyPI wheel identical to tag v1.2.2 abadc94), weights Mapika/decider-4b rev 7ab294cbdf6be6ac17fc818c10cdead744393d92 (decider_config version 4b-v2, T=1.935), uvicorn decider.serve:app. Disclosed by the author: 8,000 of the v2 LoRA rows come from generators written from the published names of the ten sealed families (no item read). Independent #1 gate (24 Sep): LEGIT. The author's private stage-2 training rows could not be audited for overlap with public items. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B dense size class, as decider-2b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Mapika · details
64.149.475.092.960.983.5%34.7%+48.8 ppest.$0.0200.02 sRunPod GPU
4
Jev 1.13.0APIby TypeSafe AI · details
63.353.176.383.352.086.6%36.7%+49.9 pp$0.0400.65 sAPI
5
JevK5 v0.2.0
†Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
by allebee · details
62.048.974.591.159.585.3%33.1%+52.2 ppest.$0.022—unknown
6
Cygnet
†Blockbrain-ai/cygnet-recipe 81974de: frozen google/gemma-4-12B-it rev 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 on unmodified vLLM 0.30.0 (image vllm/vllm-openai:v0.30.0@sha256:8a69ffad…), the author's shim: options as letters in the benchmark's label order, logits masked to the option letters, one calibration temperature T=3.4 fitted on the author's own items. The shim sends its own system prompt and renders structured state with json indent=1. Over-context/over-26-option inputs are HTTP 422 (a wrong answer). Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by blockbrain · blockbrain, frozen Gemma-4-12B-it · details
61.849.574.990.752.887.9%33.8%+54.1 ppest.$0.0370.04 sRunPod GPU
7
Hopper
†Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
by HopitAI · details
59.448.079.186.858.782.3%34.1%+48.2 ppest.$0.0240.13 sRunPod GPU
8
Winnow-12B Q8
†The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
by Eldan Ring · details
55.648.364.882.352.985.7%33.1%+52.6 ppest.$0.0370.23 sRunPod GPU
9
reflex 4B
†The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
by kshetrajna12 · details
54.047.570.468.059.779.2%28.2%+51.0 ppest.$0.0221.80 sRunPod GPU
10
djev
†The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
by Maisa (David Villalón) · Maisa, diffusion-gemma · details
52.247.055.491.457.684.0%29.9%+54.1 ppann.$0.0260.24 sAPI
11
Jev-Omni
†Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by akhilaaa3 · akhilaaa3, Gemma-4-12B merged · details
51.346.864.181.553.088.7%32.1%+56.6 ppest.$0.0370.22 sRunPod GPU
12
metask-jev-4b
†Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
by Wayfind (metask-ai) · details
47.844.766.989.154.579.7%27.6%+52.1 ppest.$0.0330.07 sRunPod GPU
13
SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ · details
47.744.466.883.759.581.0%26.3%+54.7 ppest.$0.0220.20 sRunPod GPU
14
Jobe Qwen3.5-4B
†No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
by MantisShrimpdev · frozen · details
46.944.166.185.659.581.0%25.6%+55.3 ppest.$0.0220.13 sRunPod GPU
15
local-jev Qwen3.5-4B
†Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
by Amith Chandrappa (amithgc) · details
46.844.473.375.055.880.5%26.0%+54.5 ppest.$0.0300.71 sRunPod GPU
16
system-one-openAPIby mithalouni · Gemma 4 E2B LoRA on an L4 · details
45.144.254.977.064.873.2%27.6%+45.6 ppest.$0.0150.65 sauthor demo
17
spark-s1-4b-v6
†Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085 · details
44.645.147.681.057.979.2%26.6%+52.6 ppest.$0.0250.31 sRunPod GPU
18
Malkuth-4B
†Dhtocks/malkuth-4b rev 11dc416995dab324803cb6c533f1d5c69e19d630 (rank-16 LoRA + pointer head over Qwen/Qwen3.5-4B-Base rev 1001bb4d826a52d1f399e183466143f4da7b741b), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb (the author's evaluation revision). The released head carries temperature 1.0 (no fitted temperature), although the card says calibrated. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B one-pass size class, as kev-4b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by newfull5 (dhtocks) · newfull5, Kev post-train · details
44.543.161.588.361.574.9%23.4%+51.5 ppest.$0.0190.09 sRunPod GPU
19
jqv
†A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
by hjmurmur (Octalab) · Qwen3-32B zero-shot · details
44.446.471.674.647.580.1%28.2%+51.8 ppest.$0.0560.75 sRunPod GPU
20
Qwen3-Reranker-4B
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Qwen · details
43.544.665.278.749.268.0%29.9%+38.1 pp$0.0500.13 sRunPod GPU
21
decider-35b-a3b
†The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
by Mapika · details
41.247.265.380.845.383.1%31.5%+51.6 ppest.$0.0670.29 sRunPod GPU
22
Raw Qwen3 4B Instruct 2507 direct logits
†Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction · details
41.046.429.187.659.769.7%27.3%+42.4 ppest.$0.0220.08 sRunPod GPU
23
OpenSourceJev
†DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp · details
40.941.860.382.064.078.4%26.3%+52.1 ppest.$0.016—unknown
24
ZeroEntropy zerank-2
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by ZeroEntropy · details
40.242.175.879.049.870.1%28.6%+41.6 pp$0.0470.13 sRunPod GPU
25
decision-machine-1
†A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
APIby milliseconds.ai (Baptiste Laget) · details
39.941.368.392.953.767.5%25.6%+41.9 pp$0.0350.17 sAPI
26
Malkuth-2B
†Dhtocks/malkuth-2b rev 401304b989451876070d83c488271b4f927d03ab (rank-16 LoRA + pointer head over empero-ai/Qwen3.8-2B-Distill rev e37a2dc4acc68ad75a91e07e63168cb04cc06345, fitted T=1.1755), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (no hosted ~2B listed; the 4B price errs high, the decider-2b precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by newfull5 (dhtocks) · newfull5, Kev post-train · details
38.941.353.591.461.569.7%24.4%+45.3 ppest.$0.0190.04 sRunPod GPU
27
Raw Phi-4 mini direct logits
†Neutral raw-logit control, not JevBench-directed.
by Microsoft / neutral reproduction · details
38.041.858.888.849.665.8%29.2%+36.6 ppest.$0.0480.06 sRunPod GPU
28
JEV Qwen3.5-9B Base NVFP4
†Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
by WilfLin · details
37.746.867.793.343.375.3%29.5%+45.8 ppest.$0.0770.02 sRunPod GPU
29
OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16 · details
36.945.455.083.245.581.8%28.6%+53.2 ppest.$0.0660.24 sRunPod GPU
30
kev 4B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview · details
36.142.139.675.761.866.2%22.4%+43.8 ppest.$0.0190.55 sRunPod GPU
31
Decision 2B
†Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v59 · details
35.838.874.184.362.575.3%26.0%+49.4 ppest.$0.0180.19 sRunPod GPU
32
Qwen3.5-9B Jev-like data-mix v2
†The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
by jsaurabh · details
35.247.461.382.042.478.4%29.2%+49.1 ppest.$0.083—unknown
33
GPT-6 Luna
†OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · low reasoning effort · details
35.195.892.073.736.999.1%92.9%+6.3 pp$0.1271.44 sAPI
34
SimpleJev Qwen3.8-27B
†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI · details
34.651.674.571.239.586.6%35.7%+50.9 ppest.$0.1041.01 sauthor demo
35
NInfer Qwen3.8-Flash-Next mixed
†Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
by Igor L. / NInfer contributors · details
34.049.578.688.238.989.6%34.1%+55.5 ppest.$0.1090.08 sRunPod GPU
36
swanOne
†Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. blockbrain-ai/swanone-recipe abcca2e789316472e58d4e824b61c04f419c3ba5: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4 rev 925d7be6c14c6c9442ef83e8f05b5a3c39304f69 on vllm/vllm-openai:qwen38-flash-next@sha256:0aea3024… with MiaAI Lab's nine patched vLLM files (hash-verified) and the author's shim (option letters in the benchmark's order, the model's own renormalised letter mass, no temperature), H100.md serve command, here on an RTX PRO 6000 Blackwell. The shim sends its own system prompt; structured state is rendered with json indent=1. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter qwen/qwen3.8-flash list price $0.15/M input (same underlying Flash-Next weights; the NInfer Flash-Next precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by blockbrain · blockbrain, Qwen3.8-Flash-Next NVFP4 · details
33.652.671.082.538.688.7%38.6%+50.1 ppest.$0.1110.28 sRunPod GPU
37
GPT-6 Luna
†OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · default medium reasoning effort · details
33.397.493.572.636.099.6%95.5%+4.1 pp$0.1351.48 sAPI
38
open-alternative-jev
†With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
by IkerMoel · Qwen3.5-4B, IkerMoel · details
33.238.658.783.559.674.0%24.4%+49.7 ppest.$0.0220.21 sRunPod GPU
39
jev-local
†The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
by us (GitHub) · Qwen3.5-9B · details
32.545.264.269.243.374.9%29.5%+45.3 ppest.$0.0771.05 sRunPod GPU
40
Decision Fast
†Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v53a · details
32.537.165.381.676.163.2%25.6%+37.6 ppest.$0.00630.24 sRunPod GPU
41
decider-2b
†The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
by Mapika · details
30.738.543.583.261.071.0%24.7%+46.3 ppest.$0.0200.26 sRunPod GPU
42
jeff
†Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
by Logan Markewich · Logan Markewich, GLiFormer 400M · details
30.636.867.963.576.662.8%33.1%+29.7 ppest.$0.00600.94 sCPU
43
Laya
†The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
by Convai Innovations · Convai Innovations, ModernBERT-large 421M · details
30.336.163.771.186.258.4%30.8%+27.6 ppest.$0.00290.79 sCPU
44
Autoloops – Gemma 4 31B IT
†Gemma 4 31B IT served through the Autoloops systemone API. Cost = the API's own token usage x Autoloops' published rates ($0.20/M input, $0.45/M output; no output tokens were used). API measurement: public and sealed item content reached the operator endpoint; no gold labels or answers were sent.
APIby Autoloops · details
30.059.879.284.036.092.8%45.8%+47.1 pp$0.1360.60 sAPI
45
Standard One 8B
†StandardThinking/StandardOne-8B rev 0f14d009a9400e55ea5a00a89b4d859882db704e (Ministral-3-8B-Instruct-2512 + LoRA, merged) on stock SGLang 0.5.20 behind the author's jev-adapter from server/ at the same revision, nominated configuration: --prompt-wording native --native-system-prompt none --default-temperature 1.35, context 8192. Disclosed by the author: benchmark-directed development; the native wording was selected by an ablation on the public hard tier; 181 of 359,497 training rows reuse one generic 58-character public-hard instruction. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter mistralai/ministral-8b-2512 list price $0.15/M input (= the author's proposed Mistral API price for the exact base), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Standard Thinking (myeongho12) · details
29.146.265.992.039.476.6%26.6%+50.0 ppest.$0.1040.02 sRunPod GPU
46
lev-350m
†Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M · details
28.534.870.685.376.158.4%25.0%+33.4 ppest.$0.00630.17 sRunPod GPU
47
openjev-sglangAPIby ekzhang · Qwen3.6-35B-A3B on SGLang · details
27.749.469.277.136.585.3%33.1%+52.2 ppest.$0.1310.68 sauthor demo
48
Von
†The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M · details
27.534.575.770.577.857.1%27.9%+29.2 ppest.$0.0055—unknown
49
NInfer Qwen3.8-27B NVFP4
†T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
by Igor L. / NInfer contributors · T=1.5 · details
26.951.576.080.135.283.1%33.1%+50.0 ppest.$0.1450.37 sRunPod GPU
50
NInfer Qwen3.8-27B NVFP4
†Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
by Igor L. / NInfer contributors · details
26.351.567.280.135.283.1%33.1%+50.0 ppest.$0.1450.37 sRunPod GPU
51
kev 8B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
by Jared Palmer · research preview · details
25.641.840.274.944.071.4%21.8%+49.7 ppest.$0.0730.59 sRunPod GPU
52
JevOne
†Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
by Juspay · details
25.547.677.688.535.989.6%33.8%+55.8 ppest.$0.1370.09 sRunPod GPU
53
typecastlm
†Typecastlm[server] 1.1.2 (wheel identical to git e95911e), checkpoint mihailgribov/typecastlm-qwen3.5-3.8b rev dcfecfdc28e44ef82631901628535b71574e27af: frozen Qwen/Qwen3.5-4B blocks 0-27 plus a computed 39-row head, one forward pass, four calibration temperatures by question type; transformers 5.3.0, flash-linear-attention 0.5.2. States over 32,768 tokens are folded in the middle, not refused. The server reads a noul question's meaning from the order of its two criteria (first = true), not from their keys. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (the exact base weights; one pass, no output), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Mikhail Gribov · Mikhail Gribov, Qwen3.5-4B computed head · details
25.334.062.192.660.277.9%29.9%+48.1 ppest.$0.0210.03 sRunPod GPU
54
SimpleJev Qwen3.6-35B-A3B
†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI · details
24.945.759.875.038.181.4%28.2%+53.1 ppest.$0.1160.85 sauthor demo
55
kev 0.6B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview · details
24.834.250.075.676.166.7%24.0%+42.6 ppest.$0.00630.59 sRunPod GPU
56
Raw Qwen3 8B direct logits
†Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction · details
23.745.724.186.341.968.4%26.3%+42.1 ppest.$0.0870.08 sRunPod GPU
57
system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke · details
23.443.632.884.441.571.9%24.4%+47.5 ppest.$0.0890.17 sRunPod GPU
58
OpenDecision
†A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
by Deepan Wadhwa · ModernBERT-large zero-shot · details
21.631.857.179.975.353.2%25.6%+27.6 ppest.$0.00660.34 sRunPod GPU
59
LitJev
†The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
by Zhengxu Yu · Qwen3.8-27B · details
19.546.376.666.733.686.1%30.8%+55.3 ppest.$0.1632.03 sRunPod GPU
60
openJev Verdict 1.4
†Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
by Hemant (heman10x) · details
19.029.472.078.182.457.6%27.9%+29.7 ppest.$0.00390.31 sCPU
61
kev 0.5B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · details
18.930.549.777.076.149.4%27.3%+22.1 ppest.$0.00630.43 sRunPod GPU
62
Bespoke Nimble 9B
†Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
by Bespoke Labs · details
18.746.356.478.733.479.7%28.9%+50.8 ppest.$0.1660.39 sRunPod GPU
63
GPT-5.6 LunaAPIby OpenAI · low reasoning effort · details
18.593.187.477.528.597.4%89.0%+8.4 pp$0.2420.97 sAPI
64
openJev Verdict
†The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
by Hemant (heman10x) · heman10x, ModernBERT-base 151M · details
18.130.047.076.783.155.4%24.7%+30.7 ppest.$0.00370.28 sCPU
65
Raw Qwen3 1.7B direct logits
†Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction · details
18.133.224.389.764.954.1%26.0%+28.1 ppest.$0.0150.07 sRunPod GPU
66
reflex-27b
†The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
by kshetrajna12 · Qwen3.8-27B · details
17.846.477.267.532.387.0%29.5%+57.5 ppest.$0.1811.89 sRunPod GPU
67
JevAct
†Requested in GitHub issue #66. The author published an endpoint and an evaluation client, not weights, and his API is not the TypeSafe wire format, so JevBench's own `jevact_api` adapter implements exactly the mapping table in his eval_code.zip: state to state, instructions to the question, the labels with their criteria text as the options in label order, noul as false/true, and results[0].options[i].probability read back as the probability of labels[i] by index. His server answers HTTP 400 with `all inference items exceeded max_context_tokens or contained reserved markers` for an input it cannot take, and his own client and test script record exactly that as one failed item and continue; the adapter therefore surfaces that one documented refusal as the harness's 422 refusal path, which scores it as a wrong answer and does not count toward the stop rule - the same treatment the swanOne and Cygnet packages get for their own 422. Any other 400 or an outage still stops the run. The endpoint reports no token accounting, so cost is the labelled 2B size-class estimate. The endpoint is one machine in China and the latency includes that distance; it is not a production API and gets the standard x2 self-host adjustment, without the +0.15 s that only our own servers carry. 237/308 sealed items answered validly (failures count as wrong)
APIby einptein · einptein, jev1-2b-v2 · details
16.929.355.176.564.361.9%23.4%+38.5 ppest.$0.0150.40 sauthor demo
68
djev
†Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
by David Villalon / Maisa · thinking · details
15.271.687.875.226.987.4%60.1%+27.4 ppest.$0.2740.43 sRunPod GPU
69
GLiNER2 large
†The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino · details
15.131.124.861.773.356.7%28.6%+28.1 ppest.$0.00771.10 sCPU
70
OpenJev
†OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
by razorback16 · thinking, BF16 · details
14.858.158.176.127.888.7%42.2%+46.5 ppest.$0.2550.46 sRunPod GPU
71
Qwen3.5-0.8B Decision Model
†JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
by Mourad Ghafiri · details
14.528.168.249.275.759.3%34.7%+24.6 ppest.$0.00657.15 sCPU
72
Gemini 3.1 Flash-Lite
†307/308 sealed items answered validly (failures count as wrong)
APIby Google · details
14.354.559.381.827.487.0%38.6%+48.4 pp$0.2640.76 sAPI
73
open-jev-deberta-v3-large
†297/308 sealed items answered validly (failures count as wrong)
by Kotoba Labs · local CPU · details
12.625.666.666.074.052.4%29.5%+22.8 ppest.$0.00731.77 sCPU
74
smalljev semantic-v9
†The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
by Aditya (isHeSatoshi) · details
12.325.759.279.857.960.6%26.9%+33.7 ppest.$0.0250.41 sRunPod GPU
75
GLiNER2
†A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
by Fastino · Fastino, gliner2.5-base · details
11.827.425.271.883.158.0%29.2%+28.8 ppest.$0.00370.31 sCPU
76
Instinct
†Requested in GitHub issue #69. A free evaluation/demo endpoint that speaks TypeSafe's /v1/systemone wire format, so the unchanged typesafe adapter ran it and no mapping of ours was involved; we registered the account and created the key ourselves in their console. The author states frozen Qwen3.8-27B base weights with no fine-tuning, read in one forward pass per question with no autoregressive decoding; the serving stack is not public, so nothing about it could be reviewed and the row rests on the API's own behaviour. Cost is an ESTIMATE: ZooWork publishes no bookable price, so the base model's public reference price is used (OpenRouter model-level qwen/qwen3.8-27b, $0.42/M input, zero output for a direct-logit readout); the author-announced tariff is not used. Classified as a free evaluation/demo endpoint (no SLA, no status page, no terms, pricing 'to be announced'), so the x2 latency adjustment applies. Re-scored when a bookable price is published.
APIby rayrain-srp (ZooWork) · ZooWork, Qwen3.8-27B · details
11.451.178.783.924.686.6%35.1%+51.5 ppest.$0.3250.26 sauthor demo
77
Open-Jev 9B
†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai) · details
11.244.261.872.028.177.5%29.9%+47.6 ppest.$0.2490.75 sRunPod GPU
78
Open-Jev 2B
†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai) · details
10.042.355.373.528.164.5%26.3%+38.2 ppest.$0.2490.66 sRunPod GPU
79
GLiNER2.5 multi
†The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
by Fastino · Fastino, 287M · details
9.823.157.267.882.448.9%32.8%+16.1 ppest.$0.00390.43 sCPU
80
CLM-8B
†Contrastive-LM/CLM commit cca045ffdb07b3ebcfe6938537cdeac5e14899c9, head Contrastive-LM/CLM-v0.1-8B rev 87655cb835bd76fd66c2da78e1e3709f7fa11a94 (clm-latest), Qwen3-8B rev b968826d9c46dd6066d109eabc6255188de91218 last-token pooling via vLLM. Authors' documented system_one path: state + instructions as state text, each option description as a candidate action, softmax over contrastive scores at temperature 1.0. Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod; no operator endpoint; no golds were exposed. Latency is in-process Engine.answer time on the serial standard+judge items with the self-hosted adjustment. Cost is estimated at the Qwen3-Embedding-8B hosted list price ($0.01/M input, same-size 8B pooling encoder) over CLM's measured encoder tokens; it is not a GPU bill.
by Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini) · Contrastive-LM, clm-latest · details
8.622.439.893.678.440.7%24.0%+16.7 ppest.$0.00520.02 sRunPod GPU
81
SimpleJev
†143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU · details
7.521.549.157.568.354.5%34.7%+19.8 ppest.$0.0113.99 sCPU
82
GLiNER2.5 small
†The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino · Fastino, 74M · details
7.220.550.777.882.445.9%28.6%+17.3 ppest.$0.00390.11 sCPU
83
Raw Qwen3 0.6B direct logits
†Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction · details
7.122.820.789.973.948.1%25.6%+22.4 ppest.$0.00740.07 sRunPod GPU
84
verdict-small
†Requested in GitHub issue #73. Run through the author's own `verdict serve` on our CPU, which speaks TypeSafe's /v1/systemone wire format, so JevBench's unchanged typesafe adapter ran it and no mapping of ours was involved. A 118M multilingual bi-encoder: every option is scored against the rendered state by cosine similarity, at the model's own scale (temperature 1.0, no calibrator fitted on JevBench items, as the author states). Structured state is rendered as `key: value` lines by his own code. Code review before the run: the only network call is the Hugging Face download of his own checkpoint, no telemetry, no key, no rule written against public items. The `usage.input_tokens` his server reports is a word count and not a tokeniser count, so the cost is the labelled size-class estimate rather than a measured token price. Self-host latency gets the standard x2 + 0.15 s adjustment. The author's own public-set figures were easy 0.938, standard 0.486, hard 0.396 on an Apple M5 CPU. Offline measurement of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) on our own CPU through the author's server; no operator endpoint.
by Manavarya09 (Manav) · Manavarya09, multilingual-e5-small 118M · details
5.718.166.985.096.653.7%28.9%+24.8 ppest.$0.00130.05 sCPU
85
DeepSeek V4.1 Flash
†298/308 sealed items answered validly (failures count as wrong)
APIby DeepSeek · thinking default · details
4.894.095.571.616.897.8%94.8%+3.0 pp$0.5941.42 sAPI
86
Mirror
†171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
by Bluusun · details
2.113.626.070.873.341.1%9.1%+32.0 ppest.$0.00770.90 sauthor demo
87
Mixedbread mxbai-rerank-base-v2
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Mixedbread · details
0.46.884.187.567.937.2%34.4%+2.8 pp$0.0120.07 sRunPod GPU
88
BAAI bge-reranker-v2-m3
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by BAAI · details
0.25.084.289.573.439.4%27.9%+11.5 pp$0.00770.03 sRunPod GPU
89
Alibaba GTE Reranker ModernBERT-base
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Alibaba-NLP · details
0.24.878.990.669.633.8%33.4%+0.3 pp$0.0100.05 sRunPod GPU
90
Certo v1
†The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
by AltSlate Labs · details
0.00.183.094.0100.031.6%29.5%+2.1 ppest.$0.00100.02 sRunPod GPU
91
Open Jev JSON Canvas
†Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
by JoshuaSP · details
0.048.20.084.145.684.4%31.2%+53.2 ppest.$0.0650.22 sRunPod GPU
—
classifier.dev
†Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
APIby mrmps (@michael_chomsky) · fast tier · details
honorable mention · not ranked
70.851.672.487.684.385.3%34.4%+50.9 ppest.$0.00330.39 sAPI
—
Qwen3.8 27B
†69/308 sealed items answered validly (failures count as wrong); partial: Chutes rate limit stopped the run after 81/308 items; unranked as in v1.3
APIby Qwen / Chutes · Chutes TEE · details
partial · not ranked
0.040.493.661.30.071.9%21.8%+50.1 ppest.$2.6695.75 sAPI
—
Needle 3, options as tools
†V1.4: Not re-measured: same as needle-3.
by Cactus Compute · post-hoc adapter mode · details
partial · not ranked
—————22.1%——est.$0.0143.78 sCPU
—
Needle 3
†V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.
by Cactus Compute · Cactus, 2-bit, local CPU · details
partial · not ranked
—————22.5%——est.$0.0241.69 sCPU

95 of 95 systems, sorted by official rank, #1 first.

API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. The sealed text and answers remain private; only system-level aggregates appear here. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.

All 88 system notes and disclosures
  • † Imajev-4B: Imajev-4B: requested adapter c9e5f132465da85d31735ec502d5557982671a7d and server a0134749e0900189c129cd6bb5000969f3b64bb5; one rotation with calibration.json. The optimized image included flash-linear-attention/fla-core 0.5.2 and causal-conv1d 1.7.0; CUDA profiling confirmed the pinned kernels ran. Cost is an estimate using the public DeepInfra Qwen/Qwen3.5-4B base-model reference price ($0.03/M input, $0.15/M output), with no generated output tokens; it is not a GPU bill.
  • † Plumb-4B (crh225, JevK5 v0.2 + LoRA): Crh225/plumb-4b @ 55de037801a8a9b9de3db5c0e16cef86210c2186: merged bf16 weights, LoRA r16 on the attention projections of JevK5 v0.2 (itself Qwen3.5-4B), served by the author's documented command, jevk5 v0.2.0's own jevk5-serve (github.com/allebee/jevk5, Apache-2.0), unmodified. One forward pass per decision: softmax over the declared options' answer-letter logits at the last position divided by the package's own T = 2.07 (jevk5_config.json), up to 16 options, 0 generated tokens; inputs over 16,384 tokens are refused rather than truncated. Disclosed by the author: no JevBench item was trained or tuned on and every checkpoint and temperature choice came from his own held-out sets, but the 231 public items were scored after each of five training rounds and that aggregate feedback shaped the recipe (hard mining, long documents, document length) — development against the public distribution, stated plainly. His own overlap audit dropped 9 training items sharing more than two 8-word spans with a public item. Independent top-five gate (25 Sep 2026): LEGIT — cost basis (same as its JevK5 parent), public-versus-sealed gap against like-for-like rows, an 8-gram screen of its released training data (no row shares more than two 8-grams with a public item) and an exact composite recomputation. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium H100 80 GB (Hopper) pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate at the EmpirioLabs Qwen3.5-4B public pay-as-you-go list price ($0.04/M input, no output generated), applied to this system's measured tokens; it is not a GPU bill. 25 Sep 2026 pre-release correction: retired DeepInfra $0.03/M input reference replaced by bookable EmpirioLabs $0.04/M exact-base input reference; draft score 67.09 to 65.84.
  • † decider-4b v2 (Mapika): Decider-ai 1.2.2 (PyPI wheel identical to tag v1.2.2 abadc94), weights Mapika/decider-4b rev 7ab294cbdf6be6ac17fc818c10cdead744393d92 (decider_config version 4b-v2, T=1.935), uvicorn decider.serve:app. Disclosed by the author: 8,000 of the v2 LoRA rows come from generators written from the published names of the ten sealed families (no item read). Independent #1 gate (24 Sep): LEGIT. The author's private stage-2 training rows could not be audited for overlap with public items. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B dense size class, as decider-2b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
  • † Cygnet (blockbrain, frozen Gemma-4-12B-it): Blockbrain-ai/cygnet-recipe 81974de: frozen google/gemma-4-12B-it rev 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 on unmodified vLLM 0.30.0 (image vllm/vllm-openai:v0.30.0@sha256:8a69ffad…), the author's shim: options as letters in the benchmark's label order, logits masked to the option letters, one calibration temperature T=3.4 fitted on the author's own items. The shim sends its own system prompt and renders structured state with json indent=1. Over-context/over-26-option inputs are HTTP 422 (a wrong answer). Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
  • † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • † reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • † djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
  • † Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
  • † Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
  • † local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
  • † spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Malkuth-4B (newfull5, Kev post-train): Dhtocks/malkuth-4b rev 11dc416995dab324803cb6c533f1d5c69e19d630 (rank-16 LoRA + pointer head over Qwen/Qwen3.5-4B-Base rev 1001bb4d826a52d1f399e183466143f4da7b741b), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb (the author's evaluation revision). The released head carries temperature 1.0 (no fitted temperature), although the card says calibrated. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B one-pass size class, as kev-4b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • † Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
  • † OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
  • † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • † Malkuth-2B (newfull5, Kev post-train): Dhtocks/malkuth-2b rev 401304b989451876070d83c488271b4f927d03ab (rank-16 LoRA + pointer head over empero-ai/Qwen3.8-2B-Distill rev e37a2dc4acc68ad75a91e07e63168cb04cc06345, fitted T=1.1755), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (no hosted ~2B listed; the 4B price errs high, the decider-2b precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
  • † JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
  • † kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
  • † Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
  • † GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
  • † swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4): Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. blockbrain-ai/swanone-recipe abcca2e789316472e58d4e824b61c04f419c3ba5: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4 rev 925d7be6c14c6c9442ef83e8f05b5a3c39304f69 on vllm/vllm-openai:qwen38-flash-next@sha256:0aea3024… with MiaAI Lab's nine patched vLLM files (hash-verified) and the author's shim (option letters in the benchmark's order, the model's own renormalised letter mass, no temperature), H100.md serve command, here on an RTX PRO 6000 Blackwell. The shim sends its own system prompt; structured state is rendered with json indent=1. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter qwen/qwen3.8-flash list price $0.15/M input (same underlying Flash-Next weights; the NInfer Flash-Next precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • † jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • † Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • † jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • † Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • † Autoloops – Gemma 4 31B IT: Gemma 4 31B IT served through the Autoloops systemone API. Cost = the API's own token usage x Autoloops' published rates ($0.20/M input, $0.45/M output; no output tokens were used). API measurement: public and sealed item content reached the operator endpoint; no gold labels or answers were sent.
  • † Standard One 8B (Standard Thinking): StandardThinking/StandardOne-8B rev 0f14d009a9400e55ea5a00a89b4d859882db704e (Ministral-3-8B-Instruct-2512 + LoRA, merged) on stock SGLang 0.5.20 behind the author's jev-adapter from server/ at the same revision, nominated configuration: --prompt-wording native --native-system-prompt none --default-temperature 1.35, context 8192. Disclosed by the author: benchmark-directed development; the native wording was selected by an ablation on the public hard tier; 181 of 359,497 training rows reuse one generic 58-character public-hard instruction. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter mistralai/ministral-8b-2512 list price $0.15/M input (= the author's proposed Mistral API price for the exact base), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
  • † NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
  • † NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
  • † kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
  • † typecastlm (Mikhail Gribov, Qwen3.5-4B computed head): Typecastlm[server] 1.1.2 (wheel identical to git e95911e), checkpoint mihailgribov/typecastlm-qwen3.5-3.8b rev dcfecfdc28e44ef82631901628535b71574e27af: frozen Qwen/Qwen3.5-4B blocks 0-27 plus a computed 39-row head, one forward pass, four calibration temperatures by question type; transformers 5.3.0, flash-linear-attention 0.5.2. States over 32,768 tokens are folded in the middle, not refused. The server reads a noul question's meaning from the order of its two criteria (first = true), not from their keys. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (the exact base weights; one pass, no output), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
  • † Raw Qwen3 8B direct logits: Neutral raw-logit control.
  • † OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • † LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
  • † Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • † openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • † Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
  • † reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • † JevAct (einptein, jev1-2b-v2): Requested in GitHub issue #66. The author published an endpoint and an evaluation client, not weights, and his API is not the TypeSafe wire format, so JevBench's own `jevact_api` adapter implements exactly the mapping table in his eval_code.zip: state to state, instructions to the question, the labels with their criteria text as the options in label order, noul as false/true, and results[0].options[i].probability read back as the probability of labels[i] by index. His server answers HTTP 400 with `all inference items exceeded max_context_tokens or contained reserved markers` for an input it cannot take, and his own client and test script record exactly that as one failed item and continue; the adapter therefore surfaces that one documented refusal as the harness's 422 refusal path, which scores it as a wrong answer and does not count toward the stop rule - the same treatment the swanOne and Cygnet packages get for their own 422. Any other 400 or an outage still stops the run. The endpoint reports no token accounting, so cost is the labelled 2B size-class estimate. The endpoint is one machine in China and the latency includes that distance; it is not a production API and gets the standard x2 self-host adjustment, without the +0.15 s that only our own servers carry. 237/308 sealed items answered validly (failures count as wrong)
  • † djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
  • † GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • † Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
  • † Gemini 3.1 Flash-Lite: 307/308 sealed items answered validly (failures count as wrong)
  • † open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
  • † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • † GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • † Instinct (ZooWork, Qwen3.8-27B): Requested in GitHub issue #69. A free evaluation/demo endpoint that speaks TypeSafe's /v1/systemone wire format, so the unchanged typesafe adapter ran it and no mapping of ours was involved; we registered the account and created the key ourselves in their console. The author states frozen Qwen3.8-27B base weights with no fine-tuning, read in one forward pass per question with no autoregressive decoding; the serving stack is not public, so nothing about it could be reviewed and the row rests on the API's own behaviour. Cost is an ESTIMATE: ZooWork publishes no bookable price, so the base model's public reference price is used (OpenRouter model-level qwen/qwen3.8-27b, $0.42/M input, zero output for a direct-logit readout); the author-announced tariff is not used. Classified as a free evaluation/demo endpoint (no SLA, no status page, no terms, pricing 'to be announced'), so the x2 latency adjustment applies. Re-scored when a bookable price is published.
  • † Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • † CLM-8B (Contrastive-LM, clm-latest): Contrastive-LM/CLM commit cca045ffdb07b3ebcfe6938537cdeac5e14899c9, head Contrastive-LM/CLM-v0.1-8B rev 87655cb835bd76fd66c2da78e1e3709f7fa11a94 (clm-latest), Qwen3-8B rev b968826d9c46dd6066d109eabc6255188de91218 last-token pooling via vLLM. Authors' documented system_one path: state + instructions as state text, each option description as a candidate action, softmax over contrastive scores at temperature 1.0. Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod; no operator endpoint; no golds were exposed. Latency is in-process Engine.answer time on the serial standard+judge items with the self-hosted adjustment. Cost is estimated at the Qwen3-Embedding-8B hosted list price ($0.01/M input, same-size 8B pooling encoder) over CLM's measured encoder tokens; it is not a GPU bill.
  • † SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
  • † GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
  • † verdict-small (Manavarya09, multilingual-e5-small 118M): Requested in GitHub issue #73. Run through the author's own `verdict serve` on our CPU, which speaks TypeSafe's /v1/systemone wire format, so JevBench's unchanged typesafe adapter ran it and no mapping of ours was involved. A 118M multilingual bi-encoder: every option is scored against the rendered state by cosine similarity, at the model's own scale (temperature 1.0, no calibrator fitted on JevBench items, as the author states). Structured state is rendered as `key: value` lines by his own code. Code review before the run: the only network call is the Hugging Face download of his own checkpoint, no telemetry, no key, no rule written against public items. The `usage.input_tokens` his server reports is a word count and not a tokeniser count, so the cost is the labelled size-class estimate rather than a measured token price. Self-host latency gets the standard x2 + 0.15 s adjustment. The author's own public-set figures were easy 0.938, standard 0.486, hard 0.396 on an Apple M5 CPU. Offline measurement of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) on our own CPU through the author's server; no operator endpoint.
  • † DeepSeek V4.1 Flash (thinking default): 298/308 sealed items answered validly (failures count as wrong)
  • † Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
  • † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • † Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
  • † classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
  • † Qwen3.8 27B (Chutes TEE): 69/308 sealed items answered validly (failures count as wrong); partial: Chutes rate limit stopped the run after 81/308 items; unranked as in v1.3
  • † Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
  • † Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.

Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.

Artifact: v1.4.2.2 results JSON · SHA-256 7f39b2f742a6… · JevBench v1.4.2.2 release and method

What changed in v1.4

  • Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence: I = 0.8 × I_v1.3 + 0.2 × I_sealed, where I_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates.
  • Calibration blends toward the sealed-inclusive result at the approved weight: C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35).
  • The k = 1 generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points: I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half.
  • The four axes use an equal-weight harmonic mean (p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0.
  • The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.

What the run says

Jev alternatives, open source and self-hosting

The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are Imajev-4B (#1, 67.4), Plumb-4B (#2, 65.8), Hopper (#7, 59.4), Winnow-12B Q8 (#8, 55.6). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost. Version v1.4.2.2 measures 534 public and 308 sealed decisions: 20% of Intelligence comes from the sealed set, a public-minus-sealed gap above 25 points costs Intelligence, and Intelligence, Speed or Cost below 50 each pull the score down quadratically. What changed in v1.4 · method and tiers.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.

What a decision costs

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.

How costs are estimated

One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

Price rules from v1.4.3 and v1.5 on (disclosed before any v1.5 result): only public, bookable list prices that have been in effect for at least 30 days count; promotions, subsidies, credits and free tiers do not. The scoring price is never below the market reference price of the system's base model, found the same way as the estimates above. A later price change triggers a re-score with a visible note on the row.

  • Imajev-4B — ~$0.022 est. per 1,000 decisions: DeepInfra public model catalog for exact base Qwen/Qwen3.5-4B, retrieved 2026-09-27 00:52 UTC: USD 0.03/M input and USD 0.15/M output; see receipts/BASE-PRICE-REFERENCE.json; full-forward input tokens are counted once for the single pinned server pass; no generated output tokens; estimated, not charged
  • Plumb-4B — ~$0.030 est. per 1,000 decisions: EmpirioLabs qwen3-5-4b public pay-as-you-go list price $0.04/M input, $0.07/M output; exact Qwen3.5-4B base model, 25 Sep 2026 cutoff; no output generated by this system. https://empiriolabs.ai/models/qwen3-5-4b
  • decider-4b v2 — ~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B dense size class, as decider-2b), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • JevK5 v0.2.0 — ~$0.022 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Cygnet — ~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • Hopper — ~$0.024 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Winnow-12B Q8 — ~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • reflex 4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Jev-Omni — ~$0.037 est. per 1,000 decisions: OpenRouter Gemma 3 12B input rate list price $0.05/M in, $0.0/M out (a 12B one-pass model with no generated output; the same reference the author uses in his own model card for this model) x 384 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • metask-jev-4b — ~$0.033 est. per 1,000 decisions: hosted 4B reference rate USD 0.04/M input, USD 0/M output over the exact measured prompt-token counts of all 534 attempts; no generated answer tokens
  • SemIf — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
  • Jobe Qwen3.5-4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen3.5-4B hosted reference list price $0.03/M in, $0.0/M out (same underlying weights; one forward pass, no generated tokens) x 396 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • local-jev Qwen3.5-4B — ~$0.030 est. per 1,000 decisions: hosted 4B reference rate list price $0.04/M in, $0.0/M out (one forward pass over measured input tokens and no generated answer tokens) x 397 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • system-one-open — ~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
  • spark-s1-4b-v6 — ~$0.025 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 505 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Malkuth-4B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B one-pass size class, as kev-4b), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • jqv — ~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-35b-a3b — ~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 4B Instruct 2507 direct logits — ~$0.022 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenSourceJev — ~$0.016 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Malkuth-2B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (no hosted ~2B listed; the 4B price errs high, the decider-2b precedent), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • Raw Phi-4 mini direct logits — ~$0.048 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • JEV Qwen3.5-9B Base NVFP4 — ~$0.077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenJev — ~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
  • kev 4B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Decision 2B — ~$0.018 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 269 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Qwen3.5-9B Jev-like data-mix v2 — ~$0.083 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • SimpleJev Qwen3.8-27B — ~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-Flash-Next mixed — ~$0.109 est. per 1,000 decisions: OpenRouter qwen/qwen3.8-flash hosted list reference list price $0.15/M in, $0.0/M out (same underlying Flash-Next weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • swanOne — ~$0.111 est. per 1,000 decisions: OpenRouter qwen/qwen3.8-flash list price $0.15/M input (same underlying Flash-Next weights; the NInfer Flash-Next precedent), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • open-alternative-jev — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
  • jev-local — ~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Decision Fast — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-2b — ~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jeff — ~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Laya — ~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Standard One 8B — ~$0.104 est. per 1,000 decisions: OpenRouter mistralai/ministral-8b-2512 list price $0.15/M input (= the author's proposed Mistral API price for the exact base), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • lev-350m — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openjev-sglang — ~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
  • Von — ~$0.0055 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 8B — ~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • JevOne — ~$0.137 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • typecastlm — ~$0.021 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (the exact base weights; one pass, no output), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged
  • SimpleJev Qwen3.6-35B-A3B — ~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 0.6B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 8B direct logits — ~$0.087 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • system-one — ~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
  • OpenDecision — ~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • LitJev — ~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict 1.4 — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • kev 0.5B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Bespoke Nimble 9B — ~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Raw Qwen3 1.7B direct logits — ~$0.015 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • reflex-27b — ~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • JevAct — ~$0.015 est. per 1,000 decisions: DeepInfra Qwen3.5-2B size-class reference list price $0.02/M in, $0.0/M out (a 2B one-pass decision model with nothing generated - between the <=0.6B reference the kev 0.5B row uses and the Qwen3.5-4B reference of the reflex 4B, Jobe and SemIf rows) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • djev — ~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
  • GLiNER2 large — ~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • OpenJev — ~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decision
  • Qwen3.5-0.8B Decision Model — ~$0.0065 est. per 1,000 decisions: same-size DeepInfra Qwen3.5-0.8B reference tariff; measured JevLite tokenizer usage; estimated, not charged
  • open-jev-deberta-v3-large — ~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
  • smalljev semantic-v9 — ~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2 — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Instinct — ~$0.325 est. per 1,000 decisions: ESTIMATE from the BASE MODEL's public reference price (Florian's price rule, 24 Sep 2026): OpenRouter's model-level input price for qwen/qwen3.8-27b, $0.42 per 1M input tokens (catalog snapshot 2026-09-24T19:52Z), zero output-token price for this direct-logit system, x the board-standard input token counts (452 per standard/judge decision, 1,235 per hard decision; 774.6 per decision over the 534 frozen items, the same convention as the other run-11 rows). ZooWork publishes no bookable price; the author-announced $0.03/M is not used.
  • Open-Jev 9B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open-Jev 2B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2.5 multi — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • CLM-8B — ~$0.0052 est. per 1,000 decisions: Qwen3-Embedding-8B hosted list price (OpenRouter/DeepInfra $0.01/M input, read 2026-09-24); CLM usage.input_tokens (encoder tokens on cache misses, documented default action cache); estimated, not charged
  • SimpleJev — ~$0.011 est. per 1,000 decisions: Same nonzero hosted size-class reference and exact prompt-token accounting; see RESULT.md
  • GLiNER2.5 small — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Raw Qwen3 0.6B direct logits — ~$0.0074 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • verdict-small — ~$0.0013 est. per 1,000 decisions: deepinfra small multilingual encoders (multilingual-e5-small class, 118M) list price $0.004/M in, $0.0/M out (an encoder of the same size class, below the base-size encoders the open-jev-deberta and OpenDecision rows use; one forward pass, nothing generated) x 114 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Mirror — ~$0.0077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Certo v1 — ~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open Jev JSON Canvas — ~$0.065 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • classifier.dev — ~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.
  • Qwen3.8 27B — ~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
  • Needle 3, options as tools — ~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Needle 3 — ~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
  • dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
  • dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
  • dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
  • encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
  • generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
  • moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
  • moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9

Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).

Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted once

usd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.

  • The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.
  • Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.
  • A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.
  • No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.

No tariff, measurement, item, answer or rank changed. The prices before and after:

  • Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)
  • SemIf — $0.0230 → $0.0224 (-2.31 %)
  • system-one-open — $0.0157 → $0.0149 (-5.13 %)
  • OpenJev — $0.0672 → $0.0656 (-2.36 %)
  • GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)
  • openjev-sglang — $0.1346 → $0.1313 (-2.48 %)
  • Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)
  • Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)
  • DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)
  • system-one — $0.0915 → $0.0894 (-2.29 %)
  • openJev Verdict — $0.0039 → $0.0037 (-5.21 %)
  • GLiNER2 — $0.0039 → $0.0037 (-5.21 %)
  • open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)
  • Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)
  • Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)
  • Needle 3 — $0.0249 → $0.0238 (-4.36 %)

Who could not be measured, and why

An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.

Method and tiers

Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.

Capability and Jev-class (headline ranking). Capability = (Intelligence + Calibration) / 2. A system is Jev-class if its cost per decision is at most 2× Jev 1.13.0's and its median latency (the adjusted p50 — the median the speed chart plots, not the Speed-axis score) is at most 2× Jev 1.13.0's; rows without a recorded median latency use the Speed axis at the 2× equivalent (77.2). The headline ranks Jev-class systems by Capability; the others, including general-purpose LLMs, are listed below a divider. This is a way of presenting the same measurements; it changes no score and no official rank.

JevBench Score. Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost (power mean p=-1). If Intelligence <50 multiply by (I/50)^2. For Speed and Cost separately, if below 50 multiply by (axis/50)^2. The weight sliders re-score the same axes with other weights for exploration; only equal weights give the official score and rank.

Intelligence. 0.8 × v1.3 chance-corrected Intelligence on the frozen v1.2 items + 0.2 × 100 × max(0, (sealed accuracy − 0.293)/(1 − 0.293)); then multiply by 1 − max(0, public-minus-sealed accuracy gap in percentage points − 25)/100. Original tier weights: easy .14, standard .28, judge .28, hard .30.

Revision v1.4.2.2. v1.4.2.2 adds the verified Imajev-4B row to v1.4.2.1 using the exact live v1.4.2 scoring code. Earlier measurements and score fields are unchanged; only ranks and preset ranks move where the new row changes the ordering. A 28 Sep text-only amendment clarifies that full-forward input tokens are counted once for the single pinned server pass; no measurement, score, axis, rank, eligibility, or numeric value changed.

  • easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
  • standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
  • judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
  • hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
  • sealed: 308 fresh private decisions across ten families, run once per system. Only system-level aggregates — overall and per-family accuracy, calibration — are published; the item text, answers and per-item results stay private and rotate between versions.

Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.

A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.

v1.4 scores are not comparable with v1.3 or earlier (sealed blend, gap penalty and harmonic mean). The v1.3.0 board stays below as history; the v1.0 page keeps its own numbers, calibration plots and per-family tables.

Limits
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.

3D view: three.js r128 (MIT).

Loading capability views…

Loading context-length views…

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.