JevBench v1.5.0 — Jev alternatives ranking

JevBench measures Jev-class decision models on intelligence, calibration, speed and cost. v1.5 doubles the sample to 1,624 decisions per system, scores Choice, Noul and Score requests natively and gives the fresh sealed set half of Intelligence. The official headline uses equal axis weights and gives the three request types equal weight; option B remains a secondary view.

904 open + 720 sealed decisions · 89 ranked of 103 roster systems · only system-level sealed aggregates are published.

Published data: aggregate results JSON · SHA-256 6b2f6b058b36203c98ec5f585eb8376038bc905f11db944f4e0bcd29c278c643.

Making decisions from images? Explore Image JevBench v0.1.3 and compare its systems.

Share this version · View live board · Previous release: v1.4.2.2

JevBench v1.5.0 · headline option A

JevBench Score: 89 ranked systems

Official (A)weighted harmonic mean of four 0–100 axes, Intelligence · Calibration · Speed · Cost = 25 · 25 · 25 · 25, with the low-axis gates · Method ↓

Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).

Whiskers are 95% bootstrap intervals. 75 of the 88 adjacent pairs are statistical ties — read the order as a ranking, not the gaps as significant.

  1. 1Cygnet73.7I 71 · C 87 · S 91 · K 56 · B#2 · C#1 · $0.028
  2. 2Winnow-12B Q873.2I 74 · C 84 · S 86 · K 57 · B#1 · C#2 · $0.028
  3. 3Jev 1.13.0API72.1I 72 · C 88 · S 84 · K 55 · B#3 · C#3 · $0.032
  4. 4Jev-Omni71.5I 70 · C 83 · S 85 · K 56 · B#4 · C#4 · $0.029
  5. 5decider-4b v271.3I 56 · C 86 · S 91 · K 65 · B#5 · C#7 · $0.015
  6. 6SemIf68.7I 51 · C 84 · S 91 · K 63 · B#8 · C#13 · $0.017
  7. 7spark-s1-4b-v668.2I 62 · C 70 · S 86 · K 60 · B#6 · C#5 · $0.021
  8. 8metask-jev-4b67.5I 53 · C 83 · S 89 · K 58 · B#9 · C#10 · $0.026
  9. 9Hopper67.5I 50 · C 88 · S 87 · K 62 · B#12 · C#17 · $0.018
  10. 10Malkuth-4B66.8I 54 · C 83 · S 86 · K 56 · B#10 · C#9 · $0.030
  11. 11reflex 4B65.2I 51 · C 87 · S 69 · K 63 · B#13 · C#14 · $0.017
  12. 12jev-local65.2I 56 · C 78 · S 73 · K 59 · B#11 · C#8 · $0.024
  13. 13djev (Maisa, diffusion-gemma)64.2I 72 · C 80 · S 91 · K 48 · B#7 · C#6 · $0.053
  14. 14Raw Qwen3 4B Instruct 2507 direct logits62.1I 54 · C 53 · S 89 · K 63 · B#14 · C#12 · $0.017
  15. 15jqv60.8I 49 · C 87 · S 83 · K 51 · B#15 · C#19 · $0.042
  16. 16JevK5 v0.2.058.1I 47 · C 85 · S 91 · K 63 · B#16 · C#22 · $0.017
  17. 17Qwen3.5-9B Jev-like data-mix v253.0I 60 · C 81 · S 82 · K 46 · B#17 · C#11 · $0.065
  18. 18Standard One 8B47.8I 60 · C 83 · S 92 · K 43 · B#20 · C#16 · $0.078
  19. 19NInfer Qwen3.8-Flash-Next mixed47.5I 67 · C 89 · S 89 · K 43 · B#18 · C#15 · $0.082
  20. 20swanOne46.6I 71 · C 87 · S 85 · K 42 · B#19 · C#18 · $0.085
  21. 21Raw Qwen3 8B direct logits45.2I 51 · C 49 · S 87 · K 46 · B#21 · C#24 · $0.065
  22. 22decider-2b45.1I 42 · C 72 · S 94 · K 65 · B#26 · C#26 · $0.015
  23. 23system-one44.1I 50 · C 49 · S 91 · K 45 · B#22 · C#27 · $0.068
  24. 24system-one-openAPI42.4I 42 · C 72 · S 78 · K 68 · B#27 · C#29 · $0.011
  25. 25Autoloops – Gemma 4 31B ITAPI40.5I 77 · C 86 · S 84 · K 40 · B#24 · C#20 · tariff$0.103
  26. 26GPT-6 Luna (low reasoning effort)API40.5I 95 · C 95 · S 73 · K 39 · B#23 · C#21 · $0.108
  27. 27GPT-6 Luna (default medium reasoning effort)API38.8I 96 · C 96 · S 73 · K 38 · B#25 · C#23 · $0.114
  28. 28JevOne38.2I 54 · C 85 · S 90 · K 40 · B#28 · C#28 · $0.101
  29. 29kev 4B38.1I 40 · C 68 · S 85 · K 66 · B#29 · C#32 · $0.014
  30. 30kev 8B34.2I 48 · C 71 · S 84 · K 40 · B#30 · C#34 · $0.097
  31. 31open-alternative-jev33.6I 37 · C 77 · S 91 · K 63 · B#32 · C#35 · $0.017
  32. 32Bespoke Nimble 9B31.8I 64 · C 77 · S 83 · K 37 · B#31 · C#25 · $0.128
  33. 33Malkuth-2B29.9I 36 · C 75 · S 92 · K 66 · B#35 · C#37 · $0.014
  34. 34openjev-sglangAPI29.0I 59 · C 83 · S 78 · K 36 · B#33 · C#30 · $0.140
  35. 35decider-35b-a3b27.5I 60 · C 82 · S 91 · K 34 · B#34 · C#31 · $0.154
  36. 36local-jev Qwen3.5-4B25.8I 34 · C 82 · S 84 · K 59 · B#38 · C#43 · $0.023
  37. 37Open-Jev 9B24.4I 64 · C 82 · S 74 · K 33 · B#36 · C#33 · $0.170
  38. 38Decision 2B22.5I 31 · C 86 · S 90 · K 66 · B#41 · C#45 · $0.013
  39. 39GPT-5.6 LunaAPI22.4I 94 · C 95 · S 74 · K 31 · B#37 · C#36 · tariff$0.205
  40. 40typecastlm21.8I 31 · C 77 · S 92 · K 64 · B#45 · C#47 · $0.016
  41. 41JEV Qwen3.5-9B Base NVFP420.1I 32 · C 81 · S 94 · K 48 · B#47 · C#48 · $0.056
  42. 42Gemini 3.1 Flash-LiteAPI19.6I 78 · C 75 · S 80 · K 30 · B#39 · C#38 · tariff$0.219
  43. 43NInfer Qwen3.8-27B NVFP418.7I 65 · C 86 · S 90 · K 29 · B#40 · C#39 · $0.231
  44. 44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.5I 61 · C 86 · S 90 · K 29 · B#43 · C#40 · $0.231
  45. 45InstinctAPI18.3I 63 · C 85 · S 82 · K 29 · B#44 · C#41 · $0.230
  46. 46OpenJev (thinking, BF16)17.9I 84 · C 83 · S 74 · K 29 · B#42 · C#42 · $0.241
  47. 47djev (thinking)17.4I 77 · C 96 · S 72 · K 28 · B#46 · C#44 · $0.249
  48. 48LitJev16.3I 58 · C 84 · S 68 · K 28 · B#48 · C#46 · $0.244
  49. 49Raw Phi-4 mini direct logits15.2I 28 · C 71 · S 89 · K 53 · B#50 · C#50 · $0.036
  50. 50OpenSourceJev13.2I 26 · C 76 · S 73 · K 69 · B#51 · C#51 · $0.011
  51. 51reflex-27b13.2I 63 · C 86 · S 69 · K 26 · B#49 · C#49 · $0.297
  52. 52Open-Jev 2B9.1I 34 · C 74 · S 76 · K 33 · B#52 · C#53 · $0.170
  53. 53GLiNER2 large8.4I 22 · C 42 · S 65 · K 78 · B#54 · C#54 · $0.0056
  54. 54Qwen3-Reranker-4B7.1I 21 · C 76 · S 80 · K 48 · B#55 · C#55 · $0.052
  55. 55DeepSeek V4.1 FlashAPI6.6I 94 · C 97 · S 69 · K 19 · B#53 · C#52 · $0.498
  56. 56SimpleJev4.0I 17 · C 47 · S 59 · K 70 · B#57 · C#57 · $0.0098
  57. 57SimpleJev Qwen3.8-27BAPI3.4I 73 · C 87 · S 75 · K 15 · B#56 · C#56 · $0.687
  58. 58decision-machine-1API3.2I 15 · C 81 · S 93 · K 56 · B#58 · C#58 · $0.029
  59. 59GLiNER2.5 multi2.7I 14 · C 58 · S 67 · K 87 · B#59 · C#59 · $0.0028
  60. 60GLiNER22.3I 13 · C 36 · S 70 · K 87 · B#60 · C#60 · $0.0028
  61. 61JevActAPI1.5I 11 · C 63 · S 76 · K 69 · B#62 · C#61 · $0.011
  62. 62CLM-8B1.5I 11 · C 48 · S 93 · K 51 · B#61 · C#62 · $0.045
  63. 63kev 0.6B1.3I 10 · C 68 · S 87 · K 80 · B#63 · C#63 · $0.0046
  64. 64Raw Qwen3 0.6B direct logits1.1I 11 · C 21 · S 90 · K 78 · B#64 · C#64 · $0.0056
  65. 65GLiNER2.5 small0.9I 9 · C 56 · S 77 · K 87 · B#66 · C#65 · $0.0028
  66. 66Raw Qwen3 1.7B direct logits0.9I 10 · C 22 · S 90 · K 69 · B#65 · C#66 · $0.011
  67. 67Mirror0.2I 6 · C 43 · S 64 · K 89 · B#67 · C#67 · $0.0023
  68. 68ZeroEntropy zerank-20.1I 5 · C 82 · S 80 · K 48 · B#68 · C#68 · $0.052
  69. 69jeff0.1I 4 · C 80 · S 56 · K 81 · B#69 · C#69 · $0.0043
  70. 70smalljev semantic-v90.1I 4 · C 73 · S 86 · K 61 · B#70 · C#70 · $0.020
  71. 71OpenDecision0.0I 2 · C 73 · S 87 · K 79 · B#71 · C#71 · $0.0050
  72. 72BAAI bge-reranker-v2-m30.0I 0 · C 83 · S 91 · K 59 · B#72 · C#72 · $0.023
  73. 73Certo v10.0I 0 · C 88 · S 91 · K 97 · B#73 · C#73 · $0.0013
  74. 74Decision Fast0.0I 0 · C 76 · S 91 · K 80 · B#74 · C#74 · $0.0046
  75. 75Alibaba GTE Reranker ModernBERT-base0.0I 0 · C 75 · S 91 · K 69 · B#75 · C#75 · $0.011
  76. 76kev 0.5B0.0I 0 · C 65 · S 88 · K 80 · B#76 · C#76 · $0.0046
  77. 77Laya0.0I 0 · C 74 · S 74 · K 85 · B#77 · C#77 · $0.0032
  78. 78lev-350m0.0I 0 · C 78 · S 94 · K 80 · B#78 · C#78 · $0.0046
  79. 79Qwen3.5-0.8B Decision Model0.0I 0 · C 74 · S 72 · K 80 · B#79 · C#79 · $0.0048
  80. 80Mixedbread mxbai-rerank-base-v20.0I 0 · C 87 · S 89 · K 60 · B#80 · C#80 · $0.021
  81. 81Needle 30.0I 0 · C 0 · S 34 · K 62 · B#81 · C#81 · $0.019
  82. 82Needle 3, options as tools0.0I 0 · C 0 · S 41 · K 62 · B#82 · C#82 · $0.019
  83. 83open-jev-deberta-v3-large0.0I 0 · C 77 · S 68 · K 78 · B#83 · C#83 · $0.0056
  84. 84Open Jev JSON Canvas0.0I 77 · C 0 · S 86 · K 49 · B#84 · C#84 · $0.049
  85. 85openJev Verdict0.0I 0 · C 52 · S 84 · K 87 · B#85 · C#85 · $0.0028
  86. 86openJev Verdict 1.40.0I 0 · C 80 · S 81 · K 87 · B#86 · C#86 · $0.0028
  87. 87Qwen3.8 27BAPI0.0I 96 · C 98 · S 57 · K 0 · B#87 · C#87 · $2.18
  88. 88verdict-small0.0I 0 · C 59 · S 82 · K 100 · B#88 · C#88 · $0.0009
  89. 89Von0.0I 0 · C 83 · S 76 · K 83 · B#89 · C#89 · $0.0038

Costs are estimates (est.) unless marked tariff.

system-one-openJev rebuildJev (TypeSafe, closed)Raw-logit control (base model)Native-logit decision engineInstruction model, JSON schemaZero-shot classifierReranker (neutral adapter)Closed decision APISmall tool-calling modelService built on Jevunclassified

All three weight options

A is the official headline: equal 25/25/25/25 axis weights and an Intelligence floor of 50. B remains the secondary 40/20/20/20 axis-weight view; C retains equal axes with an Intelligence floor of 60. All three use equal Choice/Noul/Score weights. The CI column is the paired-bootstrap 95% interval of the A score.

#ASystemA · equal (headline)B · 40/20/20/20C · equal, I floor 60#B#CA 95% CI
1Cygnet73.773.273.72172.4–74.5
2Winnow-12B Q873.273.573.21272.0–74.0
3Jev 1.13.0API72.172.172.13371.0–72.6
4Jev-Omni71.571.371.54470.2–72.4
5decider-4b v271.367.561.65769.1–72.3
6SemIf68.764.350.281360.1–70.2
7spark-s1-4b-v668.266.968.26566.3–69.7
8metask-jev-4b67.564.153.691065.4–68.7
9Hopper67.562.946.9121756.3–69.2
10Malkuth-4B66.863.955.010964.9–67.9
11reflex 4B65.261.948.0131458.6–66.3
12jev-local65.263.257.411863.5–66.5
13djev (Maisa, diffusion-gemma)64.264.864.27663.1–65.0
14Raw Qwen3 4B Instruct 2507 direct logits62.160.350.4141259.4–64.2
15jqv60.857.542.2151951.3–64.3
16JevK5 v0.2.058.153.640.4162247.6–68.2
17Qwen3.5-9B Jev-like data-mix v253.052.553.0171151.9–53.9
18Standard One 8B47.847.147.2201646.7–48.5
19NInfer Qwen3.8-Flash-Next mixed47.547.747.5181546.7–47.9
20swanOne46.647.346.6191845.8–46.9
21Raw Qwen3 8B direct logits45.244.632.9212436.9–46.7
22decider-2b45.141.131.3262633.8–54.3
23system-one44.143.431.2222736.8–45.7
24system-one-openAPI42.438.829.5272933.8–51.9
25Autoloops – Gemma 4 31B ITAPI40.541.840.5242040.1–40.8
26GPT-6 Luna (low reasoning effort)API40.543.140.5232140.3–40.7
27GPT-6 Luna (default medium reasoning effort)API38.841.438.8252338.6–38.9
28JevOne38.237.431.1282837.4–38.8
29kev 4B38.134.626.4293227.5–46.8
30kev 8B34.233.123.7303426.8–37.5
31open-alternative-jev33.629.923.3323525.8–41.7
32Bespoke Nimble 9B31.832.331.8312531.2–32.3
33Malkuth-2B29.926.420.8353721.0–39.0
34openjev-sglangAPI29.029.227.7333028.4–29.4
35decider-35b-a3b27.527.727.5343126.9–27.9
36local-jev Qwen3.5-4B25.822.717.9384319.9–33.0
37Open-Jev 9B24.425.024.4363323.9–24.6
38Decision 2B22.519.315.6414516.7–28.8
39GPT-5.6 LunaAPI22.424.222.4373622.3–22.5
40typecastlm21.818.915.2454716.2–28.6
41JEV Qwen3.5-9B Base NVFP420.117.813.9474815.2–25.5
42Gemini 3.1 Flash-LiteAPI19.620.819.6393819.3–19.8
43NInfer Qwen3.8-27B NVFP418.719.318.7403918.5–18.9
44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.518.918.5434018.2–18.6
45InstinctAPI18.318.918.3444118.0–18.5
46OpenJev (thinking, BF16)17.919.317.9424217.8–18.1
47djev (thinking)17.418.517.4464417.3–17.5
48LitJev16.316.715.4484616.0–16.5
49Raw Phi-4 mini direct logits15.213.110.6505010.3–20.7
50OpenSourceJev13.211.29.251519.0–18.3
51reflex-27b13.213.813.2494913.0–13.3
52Open-Jev 2B9.18.46.352536.6–11.5
53GLiNER2 large8.47.25.854545.2–12.3
54Qwen3-Reranker-4B7.15.94.955554.3–10.7
55DeepSeek V4.1 FlashAPI6.67.46.653526.6–6.7
56SimpleJev4.03.32.857572.1–6.7
57SimpleJev Qwen3.8-27BAPI3.43.73.456563.3–3.4
58decision-machine-1API3.22.52.358581.7–5.4
59GLiNER2.5 multi2.72.11.959591.2–5.1
60GLiNER22.31.81.660600.9–4.5
61JevActAPI1.51.11.162610.5–3.2
62CLM-8B1.51.21.061620.5–3.2
63kev 0.6B1.30.90.963630.5–2.5
64Raw Qwen3 0.6B direct logits1.10.90.764640.3–2.6
65GLiNER2.5 small0.90.60.666650.2–2.2
66Raw Qwen3 1.7B direct logits0.90.70.665660.1–2.3
67Mirror0.20.20.267670.0–0.9
68ZeroEntropy zerank-20.10.10.168680.0–0.5
69jeff0.10.10.169690.0–0.5
70smalljev semantic-v90.10.00.070700.0–0.4
71OpenDecision0.00.00.071710.0–0.2
72BAAI bge-reranker-v2-m30.00.00.072720.0–0.0
73Certo v10.00.00.073730.0–0.0
74Decision Fast0.00.00.074740.0–0.0
75Alibaba GTE Reranker ModernBERT-base0.00.00.075750.0–0.0
76kev 0.5B0.00.00.076760.0–0.0
77Laya0.00.00.077770.0–0.0
78lev-350m0.00.00.078780.0–0.0
79Qwen3.5-0.8B Decision Model0.00.00.079790.0–0.0
80Mixedbread mxbai-rerank-base-v20.00.00.080800.0–0.0
81Needle 30.00.00.081810.0–0.0
82Needle 3, options as tools0.00.00.082820.0–0.0
83open-jev-deberta-v3-large0.00.00.083830.0–0.0
84Open Jev JSON Canvas0.00.00.084840.0–0.0
85openJev Verdict0.00.00.085850.0–0.0
86openJev Verdict 1.40.00.00.086860.0–0.0
87Qwen3.8 27BAPI0.00.00.087870.0–0.0
88verdict-small0.00.00.088880.0–0.0
89Von0.00.00.089890.0–0.0

Axes, request types, latency and cost

Every measured system. Intelligence is 50% open (904 decisions) and 50% sealed (720); per-type columns are chance-corrected competence (CC, 0 = chance) for Choice, Noul and Score, open / sealed. Gap = I_open − I_sealed; the penalty applies only above the field median gap (G_med 5.2) plus 8. Latency is adjusted p50 / p95; cost is per 1,000 decisions. On a phone the name column stays put while the table scrolls sideways.

#ASystemScoreIntel.Calib.SpeedCostI openI sealedGapPenaltyChoice o / sNoul o / sScore o / sp50 / p95$/1k decisionsEndpoint
1Cygnet73.771.187.091.056.472.569.7+2.7×1.00083 / 8058 / 5677 / 730.23 s / 0.35 s$0.028GPU pod (RTXPRO6000)
2Winnow-12B Q873.274.484.186.156.672.576.4-3.9×1.00078 / 8259 / 6681 / 810.34 s / 0.72 s$0.028GPU pod (RTX6000)
3Jev 1.13.0API72.172.088.083.854.771.572.5-0.9×1.00086 / 8848 / 4981 / 810.62 s / 0.67 s$0.032hosted API
4Jev-Omni71.570.582.684.756.170.570.5-0.0×1.00083 / 7955 / 5773 / 760.38 s / 0.91 s$0.029GPU pod (RTX6000)
5decider-4b v271.355.885.690.964.561.749.8+11.9×1.00080 / 7743 / 1962 / 530.20 s / 0.40 s$0.015GPU pod (RTX5090)
6SemIf68.751.384.090.963.154.048.6+5.4×1.00074 / 7130 / 1458 / 610.23 s / 0.36 s$0.017GPU pod (RTX5090)
7spark-s1-4b-v668.262.169.785.560.462.561.8+0.7×1.00073 / 7350 / 4664 / 660.41 s / 0.69 s$0.021GPU pod (RTX6000)
8metask-jev-4b67.553.582.789.557.756.750.2+6.5×1.00070 / 7243 / 2557 / 530.29 s / 0.39 s$0.026GPU pod (RTX5090)
9Hopper67.549.987.987.262.348.751.0-2.3×1.00070 / 7510 / 1366 / 650.39 s / 0.48 s$0.018GPU pod (A6000)
10Malkuth-4B66.854.583.286.255.858.450.5+7.8×1.00071 / 7039 / 2364 / 590.44 s / 0.55 s$0.030GPU pod (RTX5090)
11reflex 4B65.251.586.868.863.155.247.8+7.4×1.00074 / 6933 / 1959 / 552.9 s / 4.6 s$0.017GPU pod (H100)
12jev-local65.256.377.873.058.958.254.4+3.8×1.00067 / 6547 / 3761 / 610.94 s / 5.3 s$0.024GPU pod (H100)
13djev (Maisa, diffusion-gemma)64.272.380.491.048.272.771.9+0.8×1.00077 / 7366 / 7075 / 730.25 s / 0.32 s$0.053GPU pod (H100)
14Raw Qwen3 4B Instruct 2507 direct logits62.154.152.989.263.457.450.7+6.7×1.00065 / 6350 / 3857 / 510.26 s / 0.46 s$0.017GPU pod (A6000)
15jqv60.849.186.883.451.251.346.8+4.5×1.00073 / 7319 / 1062 / 570.31 s / 1.5 s$0.042GPU pod (H100)
16JevK5 v0.2.058.146.784.990.963.150.043.4+6.6×1.00078 / 7212 / -460 / 620.22 s / 0.38 s$0.017GPU pod (RTX5090)
17Qwen3.5-9B Jev-like data-mix v253.060.480.982.045.762.658.3+4.3×1.00073 / 7542 / 3273 / 680.58 s / 1.1 s$0.065GPU pod (A6000)
18Standard One 8B47.859.683.292.543.360.059.2+0.7×1.00074 / 7245 / 4361 / 630.20 s / 0.28 s$0.078GPU pod (RTX5090)
19NInfer Qwen3.8-Flash-Next mixed47.567.288.588.642.570.364.1+6.1×1.00085 / 7951 / 4374 / 700.31 s / 0.44 s$0.082GPU pod (RTXPRO6000)
20swanOne46.671.287.184.742.270.571.9-1.5×1.00088 / 8346 / 6078 / 730.57 s / 0.59 s$0.085GPU pod (RTXPRO6000)
21Raw Qwen3 8B direct logits45.251.149.286.845.556.845.5+11.3×1.00065 / 5748 / 3857 / 420.33 s / 0.62 s$0.065GPU pod (A6000)
22decider-2b45.142.371.594.464.948.835.9+12.9×1.00052 / 5335 / 659 / 480.18 s / 0.20 s$0.015GPU pod (H100)
23system-one44.150.549.490.945.053.547.4+6.1×1.00063 / 5750 / 4048 / 460.23 s / 0.36 s$0.068GPU pod (RTX5090)
24system-one-openAPI42.441.671.678.368.344.438.8+5.6×1.00060 / 5215 / 1558 / 501.2 s / 1.3 s$0.011author demo endpoint
25Autoloops – Gemma 4 31B ITAPI40.576.785.883.939.676.976.5+0.4×1.00086 / 8869 / 6776 / 750.61 s / 0.67 stariff$0.103hosted API
26GPT-6 Luna (low reasoning effort)API40.595.394.973.239.194.795.9-1.1×1.00099 / 9885 / 90100 / 1001.6 s / 3.0 s$0.108hosted API
27GPT-6 Luna (default medium reasoning effort)API38.896.295.673.238.396.396.2+0.2×1.00099 / 10091 / 9099 / 991.6 s / 3.1 s$0.114hosted API
28JevOne38.254.184.990.439.853.954.4-0.5×1.00082 / 779 / 1770 / 690.27 s / 0.34 s$0.101GPU pod (RTXPRO6000)
29kev 4B38.139.967.685.565.847.133.2+13.9×0.99354 / 4729 / 458 / 490.49 s / 0.57 s$0.014GPU pod (A6000)
30kev 8B34.248.371.084.140.454.342.3+12.0×1.00061 / 5938 / 1464 / 540.51 s / 0.75 s$0.097GPU pod (A6000)
31open-alternative-jev33.637.476.991.363.342.632.1+10.5×1.00063 / 655 / -1759 / 480.22 s / 0.33 s$0.017GPU pod (RTX5090)
32Bespoke Nimble 9B31.863.777.283.036.865.162.3+2.8×1.00072 / 7353 / 4571 / 700.55 s / 0.93 s$0.128GPU pod (A6000)
33Malkuth-2B29.935.575.291.865.644.228.6+15.6×0.97655 / 4826 / -652 / 430.23 s / 0.29 s$0.014GPU pod (RTX5090)
34openjev-sglangAPI29.058.682.978.135.661.455.9+5.4×1.00081 / 7536 / 2767 / 661.2 s / 1.3 s$0.140author demo endpoint
35decider-35b-a3b27.560.582.091.034.464.156.8+7.4×1.00074 / 7651 / 3167 / 630.24 s / 0.33 s$0.154GPU pod (H100)
36local-jev Qwen3.5-4B25.833.782.483.759.336.730.8+5.9×1.00074 / 69-23 / -3860 / 620.43 s / 0.99 s$0.023GPU pod (A6000)
37Open-Jev 9B24.463.881.573.633.161.865.9-4.1×1.00066 / 7149 / 5970 / 671.3 s / 3.4 s$0.170GPU pod (H100)
38Decision 2B22.531.386.290.466.434.628.0+6.7×1.00058 / 68-16 / -3562 / 510.30 s / 0.31 s$0.013GPU pod (RTX6000)
39GPT-5.6 LunaAPI22.494.394.774.130.793.295.5-2.3×1.00096 / 9487 / 9496 / 991.3 s / 3.0 stariff$0.205hosted API
40typecastlm21.831.276.792.064.023.738.8-15.1×1.00066 / 72-52 / -757 / 520.23 s / 0.28 s$0.016GPU pod (RTX5090)
41JEV Qwen3.5-9B Base NVFP420.132.280.993.547.632.931.6+1.3×1.00071 / 60-35 / -2262 / 570.19 s / 0.23 s$0.056GPU pod (RTX5090)
42Gemini 3.1 Flash-LiteAPI19.677.674.780.129.876.678.6-1.9×1.00084 / 8474 / 7572 / 770.86 s / 1.1 stariff$0.219hosted API
43NInfer Qwen3.8-27B NVFP418.765.585.989.929.167.064.0+3.0×1.00081 / 8048 / 4572 / 670.25 s / 0.42 s$0.231GPU pod (RTX5090)
44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.561.186.589.929.162.459.7+2.7×1.00081 / 8037 / 3670 / 630.25 s / 0.42 s$0.231GPU pod (RTX5090)
45InstinctAPI18.362.785.081.729.261.963.6-1.7×1.00084 / 8029 / 4073 / 720.52 s / 1.3 s$0.230author demo endpoint
46OpenJev (thinking, BF16)17.984.283.173.928.582.585.9-3.4×1.00082 / 7981 / 8985 / 911.6 s / 2.6 s$0.241GPU pod (H100)
47djev (thinking)17.477.395.772.328.176.678.1-1.6×1.00065 / 4875 / 9690 / 901.5 s / 4.0 s$0.249GPU pod (H100)
48LitJev16.358.384.568.228.458.158.6-0.5×1.00080 / 7923 / 3470 / 633.1 s / 4.9 s$0.244GPU pod (H100)
49Raw Phi-4 mini direct logits15.227.671.489.253.238.320.0+18.3×0.94954 / 576 / -4455 / 470.28 s / 0.43 s$0.036GPU pod (A6000)
50OpenSourceJev13.225.876.072.569.127.923.7+4.2×1.00064 / 51-36 / -3755 / 571.4 s / 4.0 s$0.011GPU pod (A6000)
51reflex-27b13.262.885.969.225.863.162.5+0.6×1.00085 / 8031 / 3674 / 712.7 s / 4.4 s$0.297GPU pod (H100)
52Open-Jev 2B9.133.673.675.833.138.828.3+10.4×1.00052 / 4517 / -848 / 481.0 s / 2.6 s$0.170GPU pod (H100)
53GLiNER2 large8.422.542.464.877.626.019.0+7.1×1.00027 / 2816 / -135 / 291.9 s / 17.7 s$0.0056CPU container
54Qwen3-Reranker-4B7.121.076.379.648.430.313.4+17.0×0.96257 / 46-10 / -4644 / 400.51 s / 2.1 s$0.052GPU pod (A6000)
55DeepSeek V4.1 FlashAPI6.693.796.969.419.191.895.5-3.7×1.00096 / 9886 / 9793 / 921.8 s / 6.5 s$0.498hosted API
56SimpleJev4.016.746.759.170.313.320.2-6.8×1.00028 / 28-13 / 425 / 287.4 s / 16.8 s$0.0098CPU container
57SimpleJev Qwen3.8-27BAPI3.472.887.274.914.971.973.6-1.7×1.00085 / 8151 / 6179 / 791.7 s / 1.9 s$0.687author demo endpoint
58decision-machine-1API3.214.881.292.856.327.55.1+22.4×0.90853 / 45-18 / -6847 / 380.18 s / 0.29 s$0.029hosted API
59GLiNER2.5 multi2.714.058.066.686.615.812.2+3.6×1.00019 / 170 / -328 / 221.4 s / 16.1 s$0.0028CPU container
60GLiNER22.313.535.670.386.617.29.7+7.5×1.00025 / 212 / -1025 / 181.00 s / 9.3 s$0.0028CPU container
61JevActAPI1.511.262.576.468.619.73.5+16.3×0.96931 / 35-11 / -5539 / 300.80 s / 2.8 s$0.011author demo endpoint
62CLM-8B1.511.448.593.350.517.15.8+11.3×1.00019 / 87 / -1325 / 230.18 s / 0.26 s$0.045GPU pod (RTXPRO6000)
63kev 0.6B1.310.367.787.280.125.5-1.5+27.0×0.86233 / 212 / -5841 / 320.42 s / 0.45 s$0.0046GPU pod (A6000)
64Raw Qwen3 0.6B direct logits1.110.621.590.477.610.011.1-1.1×1.00017 / 2413 / -4-1 / 140.28 s / 0.33 s$0.0056GPU pod (A6000)
65GLiNER2.5 small0.99.255.977.486.617.21.6+15.5×0.97716 / 124 / -3431 / 270.47 s / 3.9 s$0.0028CPU container
66Raw Qwen3 1.7B direct logits0.99.721.690.268.512.37.1+5.2×1.00041 / 2312 / -3-16 / 10.28 s / 0.34 s$0.011GPU pod (A6000)
67Mirror0.25.643.263.689.36.74.4+2.3×1.0002 / -9-1 / -519 / 274.6 s / 9.4 s$0.0023CPU container
68ZeroEntropy zerank-20.14.881.780.348.410.0-0.4+10.4×1.00057 / 47-68 / -8641 / 370.47 s / 2.0 s$0.052GPU pod (A6000)
69jeff0.14.480.355.881.112.7-3.6+16.3×0.96933 / 22-28 / -6133 / 287.1 s / 37.2 s$0.0043CPU container
70smalljev semantic-v90.13.773.186.560.810.9-3.5+14.5×0.98731 / 18-40 / -6442 / 350.45 s / 0.50 s$0.020GPU pod (A6000)
71OpenDecision0.02.172.586.679.19.5-5.3+14.7×0.98534 / 20-47 / -7141 / 350.30 s / 0.74 s$0.0050GPU pod (H100)
72BAAI bge-reranker-v2-m30.00.083.390.559.3-21.1-22.8+1.7×1.0004 / 4-100 / -10032 / 270.21 s / 0.42 s$0.023GPU pod (A6000)
73Certo v10.00.087.591.496.9-24.1-24.9+0.8×1.000-2 / -2-100 / -10029 / 270.26 s / 0.27 s$0.0013GPU pod (A6000)
74Decision Fast0.00.076.191.480.17.3-11.7+19.0×0.94238 / 18-59 / -8743 / 330.27 s / 0.27 s$0.0046GPU pod (RTX6000)
75Alibaba GTE Reranker ModernBERT-base0.00.075.291.569.3-18.6-23.2+4.5×1.00013 / 2-100 / -10031 / 280.21 s / 0.34 s$0.011GPU pod (A6000)
76kev 0.5B0.00.065.088.580.13.4-22.3+25.7×0.87525 / 11-48 / -9533 / 170.36 s / 0.40 s$0.0046GPU pod (A6000)
77Laya0.00.073.773.984.9-0.5-13.2+12.7×1.00037 / 13-76 / -8237 / 291.5 s / 2.7 s$0.0032CPU container
78lev-350m0.00.077.793.680.05.5-18.0+23.5×0.89633 / 15-52 / -10036 / 310.19 s / 0.22 s$0.0046GPU pod (RTX6000)
79Qwen3.5-0.8B Decision Model0.00.073.671.879.53.4-3.7+7.1×1.00035 / 40-70 / -8645 / 341.3 s / 5.0 s$0.0048CPU container
80Mixedbread mxbai-rerank-base-v20.00.086.689.360.3-21.0-24.0+3.0×1.0006 / 1-100 / -10031 / 270.23 s / 0.49 s$0.021GPU pod (A6000)
81Needle 30.00.00.034.161.6-11.1-21.1+10.0×1.000-4 / -8-15 / -24-15 / -32135.5 s / 285.4 s$0.019CPU container
82Needle 3, options as tools0.00.00.041.061.6-5.8-19.9+14.1×0.99111 / 2-12 / -30-16 / -3258.2 s / 136.7 s$0.019CPU container
83open-jev-deberta-v3-large0.00.077.168.377.6-9.3-14.6+5.3×1.00021 / 17-79 / -8730 / 272.9 s / 5.1 s$0.0056CPU container
84Open Jev JSON Canvas0.077.10.085.649.379.275.1+4.0×1.00081 / 7675 / 7481 / 750.44 s / 0.63 s$0.049GPU pod (H100)
85openJev Verdict0.00.052.283.986.612.6-14.1+26.7×0.86527 / 17-15 / -7926 / 200.40 s / 1.0 s$0.0028CPU container
86openJev Verdict 1.40.00.080.380.686.6-14.7-18.3+3.5×1.00024 / 18-100 / -10031 / 270.77 s / 1.1 s$0.0028CPU container
87Qwen3.8 27BAPI0.095.698.156.80.092.998.2-5.3×1.00097 / 9887 / 9995 / 986.5 s / 31.9 s$2.18hosted API
88verdict-small0.00.059.582.1100.0-0.9-2.9+2.0×1.00026 / 13-55 / -4426 / 220.22 s / 2.9 s$0.0009CPU container
89Von0.00.083.575.782.7-1.9-16.6+14.7×0.98534 / 18-79 / -9838 / 300.92 s / 2.9 s$0.0038CPU container
–classifier.devAPIhonorable mention—75.889.282.058.974.577.1-2.7×1.00084 / 8557 / 6482 / 820.52 s / 1.2 s$0.023hosted API
–JevK5 v0.3v1.5 roster addendum A1—56.388.393.663.161.650.9+10.7×1.00080 / 7442 / 1562 / 640.18 s / 0.24 s$0.017GPU pod (H100)
–Plumb-4Bv1.5 roster addendum A1—55.887.493.563.160.950.8+10.1×1.00080 / 7343 / 1759 / 630.18 s / 0.24 s$0.017GPU pod (H100)
–Decision 4B v1.2v1.5 roster addendum A1—53.788.693.563.156.351.1+5.2×1.00081 / 7420 / 1668 / 630.18 s / 0.24 s$0.017GPU pod (H100)
–Imajev-4Bv1.5 roster addendum A2—53.588.191.163.356.250.7+5.5×1.00080 / 8222 / 867 / 620.23 s / 0.33 s$0.017GPU pod (RTX 5090)
–Decision 4B v1.1v1.5 roster addendum A1—53.187.393.563.154.751.6+3.1×1.00080 / 7518 / 1666 / 640.18 s / 0.24 s$0.017GPU pod (H100)
–Surogate Rune 26B-A4B v3v1.5 roster addendum A2—69.788.386.049.069.070.5-1.6×1.00082 / 8246 / 5079 / 800.35 s / 0.72 s$0.050GPU pod (RTX PRO 6000)
–AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1—72.887.787.629.471.973.6-1.7×1.00085 / 8250 / 6280 / 770.35 s / 0.49 s$0.226GPU pod (H100)
–AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2—72.886.787.129.472.073.5-1.5×1.00085 / 8251 / 6280 / 770.33 s / 0.60 s$0.226GPU pod (RTX PRO 6000)
–Eikos-27Bv1.5 roster addendum A1—75.186.387.628.774.176.0-1.9×1.00088 / 8856 / 6478 / 760.35 s / 0.49 s$0.238GPU pod (H100)
–SimpleJev Qwen3.6-35B-A3BAPIpartial run—0.078.575.235.1-17.4-23.0+5.6×1.00016 / 9-36 / -43-32 / -351.7 s / 1.8 s$0.145author demo endpoint

Costs are estimates (est.) unless marked tariff.

API = the operator's endpoint received sealed item text during evaluation, without answers. Sealed item text, answers and item-level results stay private; only system-level aggregates appear here. Hover a cost for its price basis and a latency for its raw values and adjustment.

Listed, not ranked: honorable mention (1)

Services that run on Jev itself are measured and shown, but not ranked against Jev, as in v1.4.2. They do not enter the field median gap or the tie markers.

  • classifier.dev (fast tier)APIhonorable mention: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2). Official (A) score 74.7.

Roster addendum: newcomers scored on the same frozen protocol (9)

Added by separately hashed roster addenda before they ran. Same frozen sample, method, price rules and v1.5.0 median gap. These rows stay outside the v1.5.0 order and its tie markers. A and secondary B placement compare each row with the frozen base point estimates only; each row's interval is shown separately and does not establish a tie with a base row or another addendum.

SystemWould place (A)A score · 95% CIWould place (B)B score · 95% CI
JevK5 v0.3v1.5 roster addendum A1#471.9 69.4–72.9#568.1 64.9–69.8
Plumb-4Bv1.5 roster addendum A1#471.6 69.2–72.7#567.7 64.7–69.6
Decision 4B v1.2v1.5 roster addendum A1#670.8 68.5–72.0#766.6 63.7–68.4
Imajev-4Bv1.5 roster addendum A2#670.4 67.8–71.6#766.2 63.1–68.1
Decision 4B v1.1v1.5 roster addendum A1#670.4 66.9–71.6#766.1 62.1–67.9
Surogate Rune 26B-A4B v3v1.5 roster addendum A2#1166.5 65.4–67.0#766.6 65.1–67.5
AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1#4319.5 19.3–19.6#4020.4 20.1–20.6
AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2#4319.5 19.3–19.6#4020.4 20.1–20.6
Eikos-27Bv1.5 roster addendum A1#4418.5 18.3–18.6#4019.5 19.2–19.7

Score followed by its 95% interval. Placements compare point estimates with the frozen base only.

Not ranked: partial, unpriced and unmeasured systems

These systems are part of the 103-system v1.5 roster but have no rank. Their numbers are never shown as zero or free.

Partial runs (1)

  • SimpleJev Qwen3.6-35B-A3BAPIpartial run: Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.

Not measured in v1.5 (3)

Jobe Qwen3.5-4B · mica-v01-4bv1.5 roster addendum A1 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)

Not measured means no v1.5 run exists yet (for example no offline image, or a hosted API not cleared for sealed items). Their v1.4.2 results stay on the v1.4.2 page.

Method notes: what changed in v1.5

Frozen method METHOD-v1.5, SHA-256 c25d3d8b8512…; pricing addendum v1.5-M2, SHA-256 2fc44459ef80….

The method owner chose equal axis weights and equal weights for Choice, Noul and Score after reviewing the What-If Lab, preserving continuity with v1.4 and treating the three decision types equally. Disclosed headline amendment: equal-axis, equal-type A, SHA-256 752ddccc4e19…. B remains a secondary view.

  • 1,624 decisions per system: 904 open (601 published) and 720 sealed, drawn fresh from a private pool with the same tier mix as the open set. Sealed counts for 50% of Intelligence: base = 0.5 × I_open + 0.5 × I_sealed.
  • Three request types are scored natively and chance-corrected per item: Choice, Noul and Score each receive one third. Tier weights easy / standard / judge / hard = 10 / 20 / 30 / 40. A type a system does not support is excluded, never scored zero; only full-coverage systems are ranked.
  • Overfit penalty relative to the field: excess = gap − G_med, penalty = max(0, 1 − max(0, excess − 8) / 100). G_med for this batch is 5.2 CC points.
  • Calibration is typed (Choice ECE/TVD, Noul ECE with Brier, Score normalised RPS and top-level ECE), pooled over open and sealed. Speed and Cost formulas are unchanged from v1.4; self-hosted and demo endpoints carry the ×2 + 0.15 s adjustment. A manufacturer's standard, non-promotional launch list price counts from day one, but a newer price cut younger than 30 days does not. Rows without token counts use the measured proxy-token basis. A system without any eligible public, bookable price is listed as unpriced.
  • The frozen 25 Sep DeepInfra snapshot records Qwen3.5-4B as deprecated on 11 Jun 2026 and replaced by Qwen3.5-9B. Its frozen snapshot rates remain the v1.5 M2 reference; price basis tooltips and the correction note disclose this. Pricing disclosure correction SHA-256: 1b660648bd49….
  • Composite: weighted harmonic mean with the Intelligence, Speed and Cost gates below 50 (Intelligence below 60 in option C). The official headline A uses equal 25 / 25 / 25 / 25 axis weights and Intelligence floor 50. B remains the secondary 40 / 20 / 20 / 20 view; C keeps equal axes and Intelligence floor 60. Ties come from the paired bootstrap.
  • Rows marked with a v1.5 roster addendum label were added by separately hashed roster addenda: same frozen sample, method, pricing rules and G_med. They remain outside the base release order and its tie markers.
  • Before every release we review the leaderboard for anomalies and close loopholes with general, documented rules. The page and Git repository provide transparent data and method details; Benchmark Heaven owns its rules.

Data file SHA-256 6b2f6b058b36203c98ec5f585eb8376038bc905f11db944f4e0bcd29c278c643 · scorer output SHA-256 452885de2a84cd5b9ed393d541fd8d6c6a540f9d7f762f9d2df9349383f26173 · run kind official.

Previous release: JevBench v1.4.2.2 (frozen results).