JevBench v1.5.0 — Jev alternatives ranking
JevBench measures Jev-class decision models on intelligence, calibration, speed and cost. v1.5 doubles the sample to 1,624 decisions per system, scores Choice, Noul and Score requests natively and gives the fresh sealed set half of Intelligence. The official headline uses equal axis weights and gives the three request types equal weight; option B remains a secondary view.
904 open + 720 sealed decisions · 89 ranked of 103 roster systems · only system-level sealed aggregates are published.
Published data: aggregate results JSON · SHA-256 6b2f6b058b36203c98ec5f585eb8376038bc905f11db944f4e0bcd29c278c643.
Making decisions from images? Explore Image JevBench v0.1.3 and compare its systems.
Share this version · View live board · Previous release: v1.4.2.2
JevBench v1.5.0 · headline option A
JevBench Score: 89 ranked systems
Official (A)weighted harmonic mean of four 0–100 axes, Intelligence · Calibration · Speed · Cost = 25 · 25 · 25 · 25, with the low-axis gates · Method ↓
Cygnet and Winnow-12B Q8 are joint leaders (statistical tie).
Whiskers are 95% bootstrap intervals. 75 of the 88 adjacent pairs are statistical ties — read the order as a ranking, not the gaps as significant.
- 1Cygnet73.7I 71 · C 87 · S 91 · K 56 · B#2 · C#1 · $0.028
- 2Winnow-12B Q873.2I 74 · C 84 · S 86 · K 57 · B#1 · C#2 · $0.028
- 3Jev 1.13.0API72.1I 72 · C 88 · S 84 · K 55 · B#3 · C#3 · $0.032
- 4Jev-Omni71.5I 70 · C 83 · S 85 · K 56 · B#4 · C#4 · $0.029
- 5decider-4b v271.3I 56 · C 86 · S 91 · K 65 · B#5 · C#7 · $0.015
- 6SemIf68.7I 51 · C 84 · S 91 · K 63 · B#8 · C#13 · $0.017
- 7spark-s1-4b-v668.2I 62 · C 70 · S 86 · K 60 · B#6 · C#5 · $0.021
- 8metask-jev-4b67.5I 53 · C 83 · S 89 · K 58 · B#9 · C#10 · $0.026
- 9Hopper67.5I 50 · C 88 · S 87 · K 62 · B#12 · C#17 · $0.018
- 10Malkuth-4B66.8I 54 · C 83 · S 86 · K 56 · B#10 · C#9 · $0.030
- 11reflex 4B65.2I 51 · C 87 · S 69 · K 63 · B#13 · C#14 · $0.017
- 12jev-local65.2I 56 · C 78 · S 73 · K 59 · B#11 · C#8 · $0.024
- 13djev (Maisa, diffusion-gemma)64.2I 72 · C 80 · S 91 · K 48 · B#7 · C#6 · $0.053
- 14Raw Qwen3 4B Instruct 2507 direct logits62.1I 54 · C 53 · S 89 · K 63 · B#14 · C#12 · $0.017
- 15jqv60.8I 49 · C 87 · S 83 · K 51 · B#15 · C#19 · $0.042
- 16JevK5 v0.2.058.1I 47 · C 85 · S 91 · K 63 · B#16 · C#22 · $0.017
- 17Qwen3.5-9B Jev-like data-mix v253.0I 60 · C 81 · S 82 · K 46 · B#17 · C#11 · $0.065
- 18Standard One 8B47.8I 60 · C 83 · S 92 · K 43 · B#20 · C#16 · $0.078
- 19NInfer Qwen3.8-Flash-Next mixed47.5I 67 · C 89 · S 89 · K 43 · B#18 · C#15 · $0.082
- 20swanOne46.6I 71 · C 87 · S 85 · K 42 · B#19 · C#18 · $0.085
- 21Raw Qwen3 8B direct logits45.2I 51 · C 49 · S 87 · K 46 · B#21 · C#24 · $0.065
- 22decider-2b45.1I 42 · C 72 · S 94 · K 65 · B#26 · C#26 · $0.015
- 23system-one44.1I 50 · C 49 · S 91 · K 45 · B#22 · C#27 · $0.068
- 24system-one-openAPI42.4I 42 · C 72 · S 78 · K 68 · B#27 · C#29 · $0.011
- 25Autoloops – Gemma 4 31B ITAPI40.5I 77 · C 86 · S 84 · K 40 · B#24 · C#20 · tariff$0.103
- 26GPT-6 Luna (low reasoning effort)API40.5I 95 · C 95 · S 73 · K 39 · B#23 · C#21 · $0.108
- 27GPT-6 Luna (default medium reasoning effort)API38.8I 96 · C 96 · S 73 · K 38 · B#25 · C#23 · $0.114
- 28JevOne38.2I 54 · C 85 · S 90 · K 40 · B#28 · C#28 · $0.101
- 29kev 4B38.1I 40 · C 68 · S 85 · K 66 · B#29 · C#32 · $0.014
- 30kev 8B34.2I 48 · C 71 · S 84 · K 40 · B#30 · C#34 · $0.097
- 31open-alternative-jev33.6I 37 · C 77 · S 91 · K 63 · B#32 · C#35 · $0.017
- 32Bespoke Nimble 9B31.8I 64 · C 77 · S 83 · K 37 · B#31 · C#25 · $0.128
- 33Malkuth-2B29.9I 36 · C 75 · S 92 · K 66 · B#35 · C#37 · $0.014
- 34openjev-sglangAPI29.0I 59 · C 83 · S 78 · K 36 · B#33 · C#30 · $0.140
- 35decider-35b-a3b27.5I 60 · C 82 · S 91 · K 34 · B#34 · C#31 · $0.154
- 36local-jev Qwen3.5-4B25.8I 34 · C 82 · S 84 · K 59 · B#38 · C#43 · $0.023
- 37Open-Jev 9B24.4I 64 · C 82 · S 74 · K 33 · B#36 · C#33 · $0.170
- 38Decision 2B22.5I 31 · C 86 · S 90 · K 66 · B#41 · C#45 · $0.013
- 39GPT-5.6 LunaAPI22.4I 94 · C 95 · S 74 · K 31 · B#37 · C#36 · tariff$0.205
- 40typecastlm21.8I 31 · C 77 · S 92 · K 64 · B#45 · C#47 · $0.016
- 41JEV Qwen3.5-9B Base NVFP420.1I 32 · C 81 · S 94 · K 48 · B#47 · C#48 · $0.056
- 42Gemini 3.1 Flash-LiteAPI19.6I 78 · C 75 · S 80 · K 30 · B#39 · C#38 · tariff$0.219
- 43NInfer Qwen3.8-27B NVFP418.7I 65 · C 86 · S 90 · K 29 · B#40 · C#39 · $0.231
- 44NInfer Qwen3.8-27B NVFP4 (T=1.5)18.5I 61 · C 86 · S 90 · K 29 · B#43 · C#40 · $0.231
- 45InstinctAPI18.3I 63 · C 85 · S 82 · K 29 · B#44 · C#41 · $0.230
- 46OpenJev (thinking, BF16)17.9I 84 · C 83 · S 74 · K 29 · B#42 · C#42 · $0.241
- 47djev (thinking)17.4I 77 · C 96 · S 72 · K 28 · B#46 · C#44 · $0.249
- 48LitJev16.3I 58 · C 84 · S 68 · K 28 · B#48 · C#46 · $0.244
- 49Raw Phi-4 mini direct logits15.2I 28 · C 71 · S 89 · K 53 · B#50 · C#50 · $0.036
- 50OpenSourceJev13.2I 26 · C 76 · S 73 · K 69 · B#51 · C#51 · $0.011
- 51reflex-27b13.2I 63 · C 86 · S 69 · K 26 · B#49 · C#49 · $0.297
- 52Open-Jev 2B9.1I 34 · C 74 · S 76 · K 33 · B#52 · C#53 · $0.170
- 53GLiNER2 large8.4I 22 · C 42 · S 65 · K 78 · B#54 · C#54 · $0.0056
- 54Qwen3-Reranker-4B7.1I 21 · C 76 · S 80 · K 48 · B#55 · C#55 · $0.052
- 55DeepSeek V4.1 FlashAPI6.6I 94 · C 97 · S 69 · K 19 · B#53 · C#52 · $0.498
- 56SimpleJev4.0I 17 · C 47 · S 59 · K 70 · B#57 · C#57 · $0.0098
- 57SimpleJev Qwen3.8-27BAPI3.4I 73 · C 87 · S 75 · K 15 · B#56 · C#56 · $0.687
- 58decision-machine-1API3.2I 15 · C 81 · S 93 · K 56 · B#58 · C#58 · $0.029
- 59GLiNER2.5 multi2.7I 14 · C 58 · S 67 · K 87 · B#59 · C#59 · $0.0028
- 60GLiNER22.3I 13 · C 36 · S 70 · K 87 · B#60 · C#60 · $0.0028
- 61JevActAPI1.5I 11 · C 63 · S 76 · K 69 · B#62 · C#61 · $0.011
- 62CLM-8B1.5I 11 · C 48 · S 93 · K 51 · B#61 · C#62 · $0.045
- 63kev 0.6B1.3I 10 · C 68 · S 87 · K 80 · B#63 · C#63 · $0.0046
- 64Raw Qwen3 0.6B direct logits1.1I 11 · C 21 · S 90 · K 78 · B#64 · C#64 · $0.0056
- 65GLiNER2.5 small0.9I 9 · C 56 · S 77 · K 87 · B#66 · C#65 · $0.0028
- 66Raw Qwen3 1.7B direct logits0.9I 10 · C 22 · S 90 · K 69 · B#65 · C#66 · $0.011
- 67Mirror0.2I 6 · C 43 · S 64 · K 89 · B#67 · C#67 · $0.0023
- 68ZeroEntropy zerank-20.1I 5 · C 82 · S 80 · K 48 · B#68 · C#68 · $0.052
- 69jeff0.1I 4 · C 80 · S 56 · K 81 · B#69 · C#69 · $0.0043
- 70smalljev semantic-v90.1I 4 · C 73 · S 86 · K 61 · B#70 · C#70 · $0.020
- 71OpenDecision0.0I 2 · C 73 · S 87 · K 79 · B#71 · C#71 · $0.0050
- 72BAAI bge-reranker-v2-m30.0I 0 · C 83 · S 91 · K 59 · B#72 · C#72 · $0.023
- 73Certo v10.0I 0 · C 88 · S 91 · K 97 · B#73 · C#73 · $0.0013
- 74Decision Fast0.0I 0 · C 76 · S 91 · K 80 · B#74 · C#74 · $0.0046
- 75Alibaba GTE Reranker ModernBERT-base0.0I 0 · C 75 · S 91 · K 69 · B#75 · C#75 · $0.011
- 76kev 0.5B0.0I 0 · C 65 · S 88 · K 80 · B#76 · C#76 · $0.0046
- 77Laya0.0I 0 · C 74 · S 74 · K 85 · B#77 · C#77 · $0.0032
- 78lev-350m0.0I 0 · C 78 · S 94 · K 80 · B#78 · C#78 · $0.0046
- 79Qwen3.5-0.8B Decision Model0.0I 0 · C 74 · S 72 · K 80 · B#79 · C#79 · $0.0048
- 80Mixedbread mxbai-rerank-base-v20.0I 0 · C 87 · S 89 · K 60 · B#80 · C#80 · $0.021
- 81Needle 30.0I 0 · C 0 · S 34 · K 62 · B#81 · C#81 · $0.019
- 82Needle 3, options as tools0.0I 0 · C 0 · S 41 · K 62 · B#82 · C#82 · $0.019
- 83open-jev-deberta-v3-large0.0I 0 · C 77 · S 68 · K 78 · B#83 · C#83 · $0.0056
- 84Open Jev JSON Canvas0.0I 77 · C 0 · S 86 · K 49 · B#84 · C#84 · $0.049
- 85openJev Verdict0.0I 0 · C 52 · S 84 · K 87 · B#85 · C#85 · $0.0028
- 86openJev Verdict 1.40.0I 0 · C 80 · S 81 · K 87 · B#86 · C#86 · $0.0028
- 87Qwen3.8 27BAPI0.0I 96 · C 98 · S 57 · K 0 · B#87 · C#87 · $2.18
- 88verdict-small0.0I 0 · C 59 · S 82 · K 100 · B#88 · C#88 · $0.0009
- 89Von0.0I 0 · C 83 · S 76 · K 83 · B#89 · C#89 · $0.0038
Costs are estimates (est.) unless marked tariff.
system-one-openJev rebuildJev (TypeSafe, closed)Raw-logit control (base model)Native-logit decision engineInstruction model, JSON schemaZero-shot classifierReranker (neutral adapter)Closed decision APISmall tool-calling modelService built on Jevunclassified
All three weight options
A is the official headline: equal 25/25/25/25 axis weights and an Intelligence floor of 50. B remains the secondary 40/20/20/20 axis-weight view; C retains equal axes with an Intelligence floor of 60. All three use equal Choice/Noul/Score weights. The CI column is the paired-bootstrap 95% interval of the A score.
| #A | System | A · equal (headline) | B · 40/20/20/20 | C · equal, I floor 60 | #B | #C | A 95% CI |
|---|---|---|---|---|---|---|---|
| 1 | Cygnet | 73.7 | 73.2 | 73.7 | 2 | 1 | 72.4–74.5 |
| 2 | Winnow-12B Q8 | 73.2 | 73.5 | 73.2 | 1 | 2 | 72.0–74.0 |
| 3 | Jev 1.13.0API | 72.1 | 72.1 | 72.1 | 3 | 3 | 71.0–72.6 |
| 4 | Jev-Omni | 71.5 | 71.3 | 71.5 | 4 | 4 | 70.2–72.4 |
| 5 | decider-4b v2 | 71.3 | 67.5 | 61.6 | 5 | 7 | 69.1–72.3 |
| 6 | SemIf | 68.7 | 64.3 | 50.2 | 8 | 13 | 60.1–70.2 |
| 7 | spark-s1-4b-v6 | 68.2 | 66.9 | 68.2 | 6 | 5 | 66.3–69.7 |
| 8 | metask-jev-4b | 67.5 | 64.1 | 53.6 | 9 | 10 | 65.4–68.7 |
| 9 | Hopper | 67.5 | 62.9 | 46.9 | 12 | 17 | 56.3–69.2 |
| 10 | Malkuth-4B | 66.8 | 63.9 | 55.0 | 10 | 9 | 64.9–67.9 |
| 11 | reflex 4B | 65.2 | 61.9 | 48.0 | 13 | 14 | 58.6–66.3 |
| 12 | jev-local | 65.2 | 63.2 | 57.4 | 11 | 8 | 63.5–66.5 |
| 13 | djev (Maisa, diffusion-gemma) | 64.2 | 64.8 | 64.2 | 7 | 6 | 63.1–65.0 |
| 14 | Raw Qwen3 4B Instruct 2507 direct logits | 62.1 | 60.3 | 50.4 | 14 | 12 | 59.4–64.2 |
| 15 | jqv | 60.8 | 57.5 | 42.2 | 15 | 19 | 51.3–64.3 |
| 16 | JevK5 v0.2.0 | 58.1 | 53.6 | 40.4 | 16 | 22 | 47.6–68.2 |
| 17 | Qwen3.5-9B Jev-like data-mix v2 | 53.0 | 52.5 | 53.0 | 17 | 11 | 51.9–53.9 |
| 18 | Standard One 8B | 47.8 | 47.1 | 47.2 | 20 | 16 | 46.7–48.5 |
| 19 | NInfer Qwen3.8-Flash-Next mixed | 47.5 | 47.7 | 47.5 | 18 | 15 | 46.7–47.9 |
| 20 | swanOne | 46.6 | 47.3 | 46.6 | 19 | 18 | 45.8–46.9 |
| 21 | Raw Qwen3 8B direct logits | 45.2 | 44.6 | 32.9 | 21 | 24 | 36.9–46.7 |
| 22 | decider-2b | 45.1 | 41.1 | 31.3 | 26 | 26 | 33.8–54.3 |
| 23 | system-one | 44.1 | 43.4 | 31.2 | 22 | 27 | 36.8–45.7 |
| 24 | system-one-openAPI | 42.4 | 38.8 | 29.5 | 27 | 29 | 33.8–51.9 |
| 25 | Autoloops – Gemma 4 31B ITAPI | 40.5 | 41.8 | 40.5 | 24 | 20 | 40.1–40.8 |
| 26 | GPT-6 Luna (low reasoning effort)API | 40.5 | 43.1 | 40.5 | 23 | 21 | 40.3–40.7 |
| 27 | GPT-6 Luna (default medium reasoning effort)API | 38.8 | 41.4 | 38.8 | 25 | 23 | 38.6–38.9 |
| 28 | JevOne | 38.2 | 37.4 | 31.1 | 28 | 28 | 37.4–38.8 |
| 29 | kev 4B | 38.1 | 34.6 | 26.4 | 29 | 32 | 27.5–46.8 |
| 30 | kev 8B | 34.2 | 33.1 | 23.7 | 30 | 34 | 26.8–37.5 |
| 31 | open-alternative-jev | 33.6 | 29.9 | 23.3 | 32 | 35 | 25.8–41.7 |
| 32 | Bespoke Nimble 9B | 31.8 | 32.3 | 31.8 | 31 | 25 | 31.2–32.3 |
| 33 | Malkuth-2B | 29.9 | 26.4 | 20.8 | 35 | 37 | 21.0–39.0 |
| 34 | openjev-sglangAPI | 29.0 | 29.2 | 27.7 | 33 | 30 | 28.4–29.4 |
| 35 | decider-35b-a3b | 27.5 | 27.7 | 27.5 | 34 | 31 | 26.9–27.9 |
| 36 | local-jev Qwen3.5-4B | 25.8 | 22.7 | 17.9 | 38 | 43 | 19.9–33.0 |
| 37 | Open-Jev 9B | 24.4 | 25.0 | 24.4 | 36 | 33 | 23.9–24.6 |
| 38 | Decision 2B | 22.5 | 19.3 | 15.6 | 41 | 45 | 16.7–28.8 |
| 39 | GPT-5.6 LunaAPI | 22.4 | 24.2 | 22.4 | 37 | 36 | 22.3–22.5 |
| 40 | typecastlm | 21.8 | 18.9 | 15.2 | 45 | 47 | 16.2–28.6 |
| 41 | JEV Qwen3.5-9B Base NVFP4 | 20.1 | 17.8 | 13.9 | 47 | 48 | 15.2–25.5 |
| 42 | Gemini 3.1 Flash-LiteAPI | 19.6 | 20.8 | 19.6 | 39 | 38 | 19.3–19.8 |
| 43 | NInfer Qwen3.8-27B NVFP4 | 18.7 | 19.3 | 18.7 | 40 | 39 | 18.5–18.9 |
| 44 | NInfer Qwen3.8-27B NVFP4 (T=1.5) | 18.5 | 18.9 | 18.5 | 43 | 40 | 18.2–18.6 |
| 45 | InstinctAPI | 18.3 | 18.9 | 18.3 | 44 | 41 | 18.0–18.5 |
| 46 | OpenJev (thinking, BF16) | 17.9 | 19.3 | 17.9 | 42 | 42 | 17.8–18.1 |
| 47 | djev (thinking) | 17.4 | 18.5 | 17.4 | 46 | 44 | 17.3–17.5 |
| 48 | LitJev | 16.3 | 16.7 | 15.4 | 48 | 46 | 16.0–16.5 |
| 49 | Raw Phi-4 mini direct logits | 15.2 | 13.1 | 10.6 | 50 | 50 | 10.3–20.7 |
| 50 | OpenSourceJev | 13.2 | 11.2 | 9.2 | 51 | 51 | 9.0–18.3 |
| 51 | reflex-27b | 13.2 | 13.8 | 13.2 | 49 | 49 | 13.0–13.3 |
| 52 | Open-Jev 2B | 9.1 | 8.4 | 6.3 | 52 | 53 | 6.6–11.5 |
| 53 | GLiNER2 large | 8.4 | 7.2 | 5.8 | 54 | 54 | 5.2–12.3 |
| 54 | Qwen3-Reranker-4B | 7.1 | 5.9 | 4.9 | 55 | 55 | 4.3–10.7 |
| 55 | DeepSeek V4.1 FlashAPI | 6.6 | 7.4 | 6.6 | 53 | 52 | 6.6–6.7 |
| 56 | SimpleJev | 4.0 | 3.3 | 2.8 | 57 | 57 | 2.1–6.7 |
| 57 | SimpleJev Qwen3.8-27BAPI | 3.4 | 3.7 | 3.4 | 56 | 56 | 3.3–3.4 |
| 58 | decision-machine-1API | 3.2 | 2.5 | 2.3 | 58 | 58 | 1.7–5.4 |
| 59 | GLiNER2.5 multi | 2.7 | 2.1 | 1.9 | 59 | 59 | 1.2–5.1 |
| 60 | GLiNER2 | 2.3 | 1.8 | 1.6 | 60 | 60 | 0.9–4.5 |
| 61 | JevActAPI | 1.5 | 1.1 | 1.1 | 62 | 61 | 0.5–3.2 |
| 62 | CLM-8B | 1.5 | 1.2 | 1.0 | 61 | 62 | 0.5–3.2 |
| 63 | kev 0.6B | 1.3 | 0.9 | 0.9 | 63 | 63 | 0.5–2.5 |
| 64 | Raw Qwen3 0.6B direct logits | 1.1 | 0.9 | 0.7 | 64 | 64 | 0.3–2.6 |
| 65 | GLiNER2.5 small | 0.9 | 0.6 | 0.6 | 66 | 65 | 0.2–2.2 |
| 66 | Raw Qwen3 1.7B direct logits | 0.9 | 0.7 | 0.6 | 65 | 66 | 0.1–2.3 |
| 67 | Mirror | 0.2 | 0.2 | 0.2 | 67 | 67 | 0.0–0.9 |
| 68 | ZeroEntropy zerank-2 | 0.1 | 0.1 | 0.1 | 68 | 68 | 0.0–0.5 |
| 69 | jeff | 0.1 | 0.1 | 0.1 | 69 | 69 | 0.0–0.5 |
| 70 | smalljev semantic-v9 | 0.1 | 0.0 | 0.0 | 70 | 70 | 0.0–0.4 |
| 71 | OpenDecision | 0.0 | 0.0 | 0.0 | 71 | 71 | 0.0–0.2 |
| 72 | BAAI bge-reranker-v2-m3 | 0.0 | 0.0 | 0.0 | 72 | 72 | 0.0–0.0 |
| 73 | Certo v1 | 0.0 | 0.0 | 0.0 | 73 | 73 | 0.0–0.0 |
| 74 | Decision Fast | 0.0 | 0.0 | 0.0 | 74 | 74 | 0.0–0.0 |
| 75 | Alibaba GTE Reranker ModernBERT-base | 0.0 | 0.0 | 0.0 | 75 | 75 | 0.0–0.0 |
| 76 | kev 0.5B | 0.0 | 0.0 | 0.0 | 76 | 76 | 0.0–0.0 |
| 77 | Laya | 0.0 | 0.0 | 0.0 | 77 | 77 | 0.0–0.0 |
| 78 | lev-350m | 0.0 | 0.0 | 0.0 | 78 | 78 | 0.0–0.0 |
| 79 | Qwen3.5-0.8B Decision Model | 0.0 | 0.0 | 0.0 | 79 | 79 | 0.0–0.0 |
| 80 | Mixedbread mxbai-rerank-base-v2 | 0.0 | 0.0 | 0.0 | 80 | 80 | 0.0–0.0 |
| 81 | Needle 3 | 0.0 | 0.0 | 0.0 | 81 | 81 | 0.0–0.0 |
| 82 | Needle 3, options as tools | 0.0 | 0.0 | 0.0 | 82 | 82 | 0.0–0.0 |
| 83 | open-jev-deberta-v3-large | 0.0 | 0.0 | 0.0 | 83 | 83 | 0.0–0.0 |
| 84 | Open Jev JSON Canvas | 0.0 | 0.0 | 0.0 | 84 | 84 | 0.0–0.0 |
| 85 | openJev Verdict | 0.0 | 0.0 | 0.0 | 85 | 85 | 0.0–0.0 |
| 86 | openJev Verdict 1.4 | 0.0 | 0.0 | 0.0 | 86 | 86 | 0.0–0.0 |
| 87 | Qwen3.8 27BAPI | 0.0 | 0.0 | 0.0 | 87 | 87 | 0.0–0.0 |
| 88 | verdict-small | 0.0 | 0.0 | 0.0 | 88 | 88 | 0.0–0.0 |
| 89 | Von | 0.0 | 0.0 | 0.0 | 89 | 89 | 0.0–0.0 |
Axes, request types, latency and cost
Every measured system. Intelligence is 50% open (904 decisions) and 50% sealed (720); per-type columns are chance-corrected competence (CC, 0 = chance) for Choice, Noul and Score, open / sealed. Gap = I_open − I_sealed; the penalty applies only above the field median gap (G_med 5.2) plus 8. Latency is adjusted p50 / p95; cost is per 1,000 decisions. On a phone the name column stays put while the table scrolls sideways.
| #A | System | Score | Intel. | Calib. | Speed | Cost | I open | I sealed | Gap | Penalty | Choice o / s | Noul o / s | Score o / s | p50 / p95 | $/1k decisions | Endpoint |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Cygnet | 73.7 | 71.1 | 87.0 | 91.0 | 56.4 | 72.5 | 69.7 | +2.7 | ×1.000 | 83 / 80 | 58 / 56 | 77 / 73 | 0.23 s / 0.35 s | $0.028 | GPU pod (RTXPRO6000) |
| 2 | Winnow-12B Q8 | 73.2 | 74.4 | 84.1 | 86.1 | 56.6 | 72.5 | 76.4 | -3.9 | ×1.000 | 78 / 82 | 59 / 66 | 81 / 81 | 0.34 s / 0.72 s | $0.028 | GPU pod (RTX6000) |
| 3 | Jev 1.13.0API | 72.1 | 72.0 | 88.0 | 83.8 | 54.7 | 71.5 | 72.5 | -0.9 | ×1.000 | 86 / 88 | 48 / 49 | 81 / 81 | 0.62 s / 0.67 s | $0.032 | hosted API |
| 4 | Jev-Omni | 71.5 | 70.5 | 82.6 | 84.7 | 56.1 | 70.5 | 70.5 | -0.0 | ×1.000 | 83 / 79 | 55 / 57 | 73 / 76 | 0.38 s / 0.91 s | $0.029 | GPU pod (RTX6000) |
| 5 | decider-4b v2 | 71.3 | 55.8 | 85.6 | 90.9 | 64.5 | 61.7 | 49.8 | +11.9 | ×1.000 | 80 / 77 | 43 / 19 | 62 / 53 | 0.20 s / 0.40 s | $0.015 | GPU pod (RTX5090) |
| 6 | SemIf | 68.7 | 51.3 | 84.0 | 90.9 | 63.1 | 54.0 | 48.6 | +5.4 | ×1.000 | 74 / 71 | 30 / 14 | 58 / 61 | 0.23 s / 0.36 s | $0.017 | GPU pod (RTX5090) |
| 7 | spark-s1-4b-v6 | 68.2 | 62.1 | 69.7 | 85.5 | 60.4 | 62.5 | 61.8 | +0.7 | ×1.000 | 73 / 73 | 50 / 46 | 64 / 66 | 0.41 s / 0.69 s | $0.021 | GPU pod (RTX6000) |
| 8 | metask-jev-4b | 67.5 | 53.5 | 82.7 | 89.5 | 57.7 | 56.7 | 50.2 | +6.5 | ×1.000 | 70 / 72 | 43 / 25 | 57 / 53 | 0.29 s / 0.39 s | $0.026 | GPU pod (RTX5090) |
| 9 | Hopper | 67.5 | 49.9 | 87.9 | 87.2 | 62.3 | 48.7 | 51.0 | -2.3 | ×1.000 | 70 / 75 | 10 / 13 | 66 / 65 | 0.39 s / 0.48 s | $0.018 | GPU pod (A6000) |
| 10 | Malkuth-4B | 66.8 | 54.5 | 83.2 | 86.2 | 55.8 | 58.4 | 50.5 | +7.8 | ×1.000 | 71 / 70 | 39 / 23 | 64 / 59 | 0.44 s / 0.55 s | $0.030 | GPU pod (RTX5090) |
| 11 | reflex 4B | 65.2 | 51.5 | 86.8 | 68.8 | 63.1 | 55.2 | 47.8 | +7.4 | ×1.000 | 74 / 69 | 33 / 19 | 59 / 55 | 2.9 s / 4.6 s | $0.017 | GPU pod (H100) |
| 12 | jev-local | 65.2 | 56.3 | 77.8 | 73.0 | 58.9 | 58.2 | 54.4 | +3.8 | ×1.000 | 67 / 65 | 47 / 37 | 61 / 61 | 0.94 s / 5.3 s | $0.024 | GPU pod (H100) |
| 13 | djev (Maisa, diffusion-gemma) | 64.2 | 72.3 | 80.4 | 91.0 | 48.2 | 72.7 | 71.9 | +0.8 | ×1.000 | 77 / 73 | 66 / 70 | 75 / 73 | 0.25 s / 0.32 s | $0.053 | GPU pod (H100) |
| 14 | Raw Qwen3 4B Instruct 2507 direct logits | 62.1 | 54.1 | 52.9 | 89.2 | 63.4 | 57.4 | 50.7 | +6.7 | ×1.000 | 65 / 63 | 50 / 38 | 57 / 51 | 0.26 s / 0.46 s | $0.017 | GPU pod (A6000) |
| 15 | jqv | 60.8 | 49.1 | 86.8 | 83.4 | 51.2 | 51.3 | 46.8 | +4.5 | ×1.000 | 73 / 73 | 19 / 10 | 62 / 57 | 0.31 s / 1.5 s | $0.042 | GPU pod (H100) |
| 16 | JevK5 v0.2.0 | 58.1 | 46.7 | 84.9 | 90.9 | 63.1 | 50.0 | 43.4 | +6.6 | ×1.000 | 78 / 72 | 12 / -4 | 60 / 62 | 0.22 s / 0.38 s | $0.017 | GPU pod (RTX5090) |
| 17 | Qwen3.5-9B Jev-like data-mix v2 | 53.0 | 60.4 | 80.9 | 82.0 | 45.7 | 62.6 | 58.3 | +4.3 | ×1.000 | 73 / 75 | 42 / 32 | 73 / 68 | 0.58 s / 1.1 s | $0.065 | GPU pod (A6000) |
| 18 | Standard One 8B | 47.8 | 59.6 | 83.2 | 92.5 | 43.3 | 60.0 | 59.2 | +0.7 | ×1.000 | 74 / 72 | 45 / 43 | 61 / 63 | 0.20 s / 0.28 s | $0.078 | GPU pod (RTX5090) |
| 19 | NInfer Qwen3.8-Flash-Next mixed | 47.5 | 67.2 | 88.5 | 88.6 | 42.5 | 70.3 | 64.1 | +6.1 | ×1.000 | 85 / 79 | 51 / 43 | 74 / 70 | 0.31 s / 0.44 s | $0.082 | GPU pod (RTXPRO6000) |
| 20 | swanOne | 46.6 | 71.2 | 87.1 | 84.7 | 42.2 | 70.5 | 71.9 | -1.5 | ×1.000 | 88 / 83 | 46 / 60 | 78 / 73 | 0.57 s / 0.59 s | $0.085 | GPU pod (RTXPRO6000) |
| 21 | Raw Qwen3 8B direct logits | 45.2 | 51.1 | 49.2 | 86.8 | 45.5 | 56.8 | 45.5 | +11.3 | ×1.000 | 65 / 57 | 48 / 38 | 57 / 42 | 0.33 s / 0.62 s | $0.065 | GPU pod (A6000) |
| 22 | decider-2b | 45.1 | 42.3 | 71.5 | 94.4 | 64.9 | 48.8 | 35.9 | +12.9 | ×1.000 | 52 / 53 | 35 / 6 | 59 / 48 | 0.18 s / 0.20 s | $0.015 | GPU pod (H100) |
| 23 | system-one | 44.1 | 50.5 | 49.4 | 90.9 | 45.0 | 53.5 | 47.4 | +6.1 | ×1.000 | 63 / 57 | 50 / 40 | 48 / 46 | 0.23 s / 0.36 s | $0.068 | GPU pod (RTX5090) |
| 24 | system-one-openAPI | 42.4 | 41.6 | 71.6 | 78.3 | 68.3 | 44.4 | 38.8 | +5.6 | ×1.000 | 60 / 52 | 15 / 15 | 58 / 50 | 1.2 s / 1.3 s | $0.011 | author demo endpoint |
| 25 | Autoloops – Gemma 4 31B ITAPI | 40.5 | 76.7 | 85.8 | 83.9 | 39.6 | 76.9 | 76.5 | +0.4 | ×1.000 | 86 / 88 | 69 / 67 | 76 / 75 | 0.61 s / 0.67 s | tariff$0.103 | hosted API |
| 26 | GPT-6 Luna (low reasoning effort)API | 40.5 | 95.3 | 94.9 | 73.2 | 39.1 | 94.7 | 95.9 | -1.1 | ×1.000 | 99 / 98 | 85 / 90 | 100 / 100 | 1.6 s / 3.0 s | $0.108 | hosted API |
| 27 | GPT-6 Luna (default medium reasoning effort)API | 38.8 | 96.2 | 95.6 | 73.2 | 38.3 | 96.3 | 96.2 | +0.2 | ×1.000 | 99 / 100 | 91 / 90 | 99 / 99 | 1.6 s / 3.1 s | $0.114 | hosted API |
| 28 | JevOne | 38.2 | 54.1 | 84.9 | 90.4 | 39.8 | 53.9 | 54.4 | -0.5 | ×1.000 | 82 / 77 | 9 / 17 | 70 / 69 | 0.27 s / 0.34 s | $0.101 | GPU pod (RTXPRO6000) |
| 29 | kev 4B | 38.1 | 39.9 | 67.6 | 85.5 | 65.8 | 47.1 | 33.2 | +13.9 | ×0.993 | 54 / 47 | 29 / 4 | 58 / 49 | 0.49 s / 0.57 s | $0.014 | GPU pod (A6000) |
| 30 | kev 8B | 34.2 | 48.3 | 71.0 | 84.1 | 40.4 | 54.3 | 42.3 | +12.0 | ×1.000 | 61 / 59 | 38 / 14 | 64 / 54 | 0.51 s / 0.75 s | $0.097 | GPU pod (A6000) |
| 31 | open-alternative-jev | 33.6 | 37.4 | 76.9 | 91.3 | 63.3 | 42.6 | 32.1 | +10.5 | ×1.000 | 63 / 65 | 5 / -17 | 59 / 48 | 0.22 s / 0.33 s | $0.017 | GPU pod (RTX5090) |
| 32 | Bespoke Nimble 9B | 31.8 | 63.7 | 77.2 | 83.0 | 36.8 | 65.1 | 62.3 | +2.8 | ×1.000 | 72 / 73 | 53 / 45 | 71 / 70 | 0.55 s / 0.93 s | $0.128 | GPU pod (A6000) |
| 33 | Malkuth-2B | 29.9 | 35.5 | 75.2 | 91.8 | 65.6 | 44.2 | 28.6 | +15.6 | ×0.976 | 55 / 48 | 26 / -6 | 52 / 43 | 0.23 s / 0.29 s | $0.014 | GPU pod (RTX5090) |
| 34 | openjev-sglangAPI | 29.0 | 58.6 | 82.9 | 78.1 | 35.6 | 61.4 | 55.9 | +5.4 | ×1.000 | 81 / 75 | 36 / 27 | 67 / 66 | 1.2 s / 1.3 s | $0.140 | author demo endpoint |
| 35 | decider-35b-a3b | 27.5 | 60.5 | 82.0 | 91.0 | 34.4 | 64.1 | 56.8 | +7.4 | ×1.000 | 74 / 76 | 51 / 31 | 67 / 63 | 0.24 s / 0.33 s | $0.154 | GPU pod (H100) |
| 36 | local-jev Qwen3.5-4B | 25.8 | 33.7 | 82.4 | 83.7 | 59.3 | 36.7 | 30.8 | +5.9 | ×1.000 | 74 / 69 | -23 / -38 | 60 / 62 | 0.43 s / 0.99 s | $0.023 | GPU pod (A6000) |
| 37 | Open-Jev 9B | 24.4 | 63.8 | 81.5 | 73.6 | 33.1 | 61.8 | 65.9 | -4.1 | ×1.000 | 66 / 71 | 49 / 59 | 70 / 67 | 1.3 s / 3.4 s | $0.170 | GPU pod (H100) |
| 38 | Decision 2B | 22.5 | 31.3 | 86.2 | 90.4 | 66.4 | 34.6 | 28.0 | +6.7 | ×1.000 | 58 / 68 | -16 / -35 | 62 / 51 | 0.30 s / 0.31 s | $0.013 | GPU pod (RTX6000) |
| 39 | GPT-5.6 LunaAPI | 22.4 | 94.3 | 94.7 | 74.1 | 30.7 | 93.2 | 95.5 | -2.3 | ×1.000 | 96 / 94 | 87 / 94 | 96 / 99 | 1.3 s / 3.0 s | tariff$0.205 | hosted API |
| 40 | typecastlm | 21.8 | 31.2 | 76.7 | 92.0 | 64.0 | 23.7 | 38.8 | -15.1 | ×1.000 | 66 / 72 | -52 / -7 | 57 / 52 | 0.23 s / 0.28 s | $0.016 | GPU pod (RTX5090) |
| 41 | JEV Qwen3.5-9B Base NVFP4 | 20.1 | 32.2 | 80.9 | 93.5 | 47.6 | 32.9 | 31.6 | +1.3 | ×1.000 | 71 / 60 | -35 / -22 | 62 / 57 | 0.19 s / 0.23 s | $0.056 | GPU pod (RTX5090) |
| 42 | Gemini 3.1 Flash-LiteAPI | 19.6 | 77.6 | 74.7 | 80.1 | 29.8 | 76.6 | 78.6 | -1.9 | ×1.000 | 84 / 84 | 74 / 75 | 72 / 77 | 0.86 s / 1.1 s | tariff$0.219 | hosted API |
| 43 | NInfer Qwen3.8-27B NVFP4 | 18.7 | 65.5 | 85.9 | 89.9 | 29.1 | 67.0 | 64.0 | +3.0 | ×1.000 | 81 / 80 | 48 / 45 | 72 / 67 | 0.25 s / 0.42 s | $0.231 | GPU pod (RTX5090) |
| 44 | NInfer Qwen3.8-27B NVFP4 (T=1.5) | 18.5 | 61.1 | 86.5 | 89.9 | 29.1 | 62.4 | 59.7 | +2.7 | ×1.000 | 81 / 80 | 37 / 36 | 70 / 63 | 0.25 s / 0.42 s | $0.231 | GPU pod (RTX5090) |
| 45 | InstinctAPI | 18.3 | 62.7 | 85.0 | 81.7 | 29.2 | 61.9 | 63.6 | -1.7 | ×1.000 | 84 / 80 | 29 / 40 | 73 / 72 | 0.52 s / 1.3 s | $0.230 | author demo endpoint |
| 46 | OpenJev (thinking, BF16) | 17.9 | 84.2 | 83.1 | 73.9 | 28.5 | 82.5 | 85.9 | -3.4 | ×1.000 | 82 / 79 | 81 / 89 | 85 / 91 | 1.6 s / 2.6 s | $0.241 | GPU pod (H100) |
| 47 | djev (thinking) | 17.4 | 77.3 | 95.7 | 72.3 | 28.1 | 76.6 | 78.1 | -1.6 | ×1.000 | 65 / 48 | 75 / 96 | 90 / 90 | 1.5 s / 4.0 s | $0.249 | GPU pod (H100) |
| 48 | LitJev | 16.3 | 58.3 | 84.5 | 68.2 | 28.4 | 58.1 | 58.6 | -0.5 | ×1.000 | 80 / 79 | 23 / 34 | 70 / 63 | 3.1 s / 4.9 s | $0.244 | GPU pod (H100) |
| 49 | Raw Phi-4 mini direct logits | 15.2 | 27.6 | 71.4 | 89.2 | 53.2 | 38.3 | 20.0 | +18.3 | ×0.949 | 54 / 57 | 6 / -44 | 55 / 47 | 0.28 s / 0.43 s | $0.036 | GPU pod (A6000) |
| 50 | OpenSourceJev | 13.2 | 25.8 | 76.0 | 72.5 | 69.1 | 27.9 | 23.7 | +4.2 | ×1.000 | 64 / 51 | -36 / -37 | 55 / 57 | 1.4 s / 4.0 s | $0.011 | GPU pod (A6000) |
| 51 | reflex-27b | 13.2 | 62.8 | 85.9 | 69.2 | 25.8 | 63.1 | 62.5 | +0.6 | ×1.000 | 85 / 80 | 31 / 36 | 74 / 71 | 2.7 s / 4.4 s | $0.297 | GPU pod (H100) |
| 52 | Open-Jev 2B | 9.1 | 33.6 | 73.6 | 75.8 | 33.1 | 38.8 | 28.3 | +10.4 | ×1.000 | 52 / 45 | 17 / -8 | 48 / 48 | 1.0 s / 2.6 s | $0.170 | GPU pod (H100) |
| 53 | GLiNER2 large | 8.4 | 22.5 | 42.4 | 64.8 | 77.6 | 26.0 | 19.0 | +7.1 | ×1.000 | 27 / 28 | 16 / -1 | 35 / 29 | 1.9 s / 17.7 s | $0.0056 | CPU container |
| 54 | Qwen3-Reranker-4B | 7.1 | 21.0 | 76.3 | 79.6 | 48.4 | 30.3 | 13.4 | +17.0 | ×0.962 | 57 / 46 | -10 / -46 | 44 / 40 | 0.51 s / 2.1 s | $0.052 | GPU pod (A6000) |
| 55 | DeepSeek V4.1 FlashAPI | 6.6 | 93.7 | 96.9 | 69.4 | 19.1 | 91.8 | 95.5 | -3.7 | ×1.000 | 96 / 98 | 86 / 97 | 93 / 92 | 1.8 s / 6.5 s | $0.498 | hosted API |
| 56 | SimpleJev | 4.0 | 16.7 | 46.7 | 59.1 | 70.3 | 13.3 | 20.2 | -6.8 | ×1.000 | 28 / 28 | -13 / 4 | 25 / 28 | 7.4 s / 16.8 s | $0.0098 | CPU container |
| 57 | SimpleJev Qwen3.8-27BAPI | 3.4 | 72.8 | 87.2 | 74.9 | 14.9 | 71.9 | 73.6 | -1.7 | ×1.000 | 85 / 81 | 51 / 61 | 79 / 79 | 1.7 s / 1.9 s | $0.687 | author demo endpoint |
| 58 | decision-machine-1API | 3.2 | 14.8 | 81.2 | 92.8 | 56.3 | 27.5 | 5.1 | +22.4 | ×0.908 | 53 / 45 | -18 / -68 | 47 / 38 | 0.18 s / 0.29 s | $0.029 | hosted API |
| 59 | GLiNER2.5 multi | 2.7 | 14.0 | 58.0 | 66.6 | 86.6 | 15.8 | 12.2 | +3.6 | ×1.000 | 19 / 17 | 0 / -3 | 28 / 22 | 1.4 s / 16.1 s | $0.0028 | CPU container |
| 60 | GLiNER2 | 2.3 | 13.5 | 35.6 | 70.3 | 86.6 | 17.2 | 9.7 | +7.5 | ×1.000 | 25 / 21 | 2 / -10 | 25 / 18 | 1.00 s / 9.3 s | $0.0028 | CPU container |
| 61 | JevActAPI | 1.5 | 11.2 | 62.5 | 76.4 | 68.6 | 19.7 | 3.5 | +16.3 | ×0.969 | 31 / 35 | -11 / -55 | 39 / 30 | 0.80 s / 2.8 s | $0.011 | author demo endpoint |
| 62 | CLM-8B | 1.5 | 11.4 | 48.5 | 93.3 | 50.5 | 17.1 | 5.8 | +11.3 | ×1.000 | 19 / 8 | 7 / -13 | 25 / 23 | 0.18 s / 0.26 s | $0.045 | GPU pod (RTXPRO6000) |
| 63 | kev 0.6B | 1.3 | 10.3 | 67.7 | 87.2 | 80.1 | 25.5 | -1.5 | +27.0 | ×0.862 | 33 / 21 | 2 / -58 | 41 / 32 | 0.42 s / 0.45 s | $0.0046 | GPU pod (A6000) |
| 64 | Raw Qwen3 0.6B direct logits | 1.1 | 10.6 | 21.5 | 90.4 | 77.6 | 10.0 | 11.1 | -1.1 | ×1.000 | 17 / 24 | 13 / -4 | -1 / 14 | 0.28 s / 0.33 s | $0.0056 | GPU pod (A6000) |
| 65 | GLiNER2.5 small | 0.9 | 9.2 | 55.9 | 77.4 | 86.6 | 17.2 | 1.6 | +15.5 | ×0.977 | 16 / 12 | 4 / -34 | 31 / 27 | 0.47 s / 3.9 s | $0.0028 | CPU container |
| 66 | Raw Qwen3 1.7B direct logits | 0.9 | 9.7 | 21.6 | 90.2 | 68.5 | 12.3 | 7.1 | +5.2 | ×1.000 | 41 / 23 | 12 / -3 | -16 / 1 | 0.28 s / 0.34 s | $0.011 | GPU pod (A6000) |
| 67 | Mirror | 0.2 | 5.6 | 43.2 | 63.6 | 89.3 | 6.7 | 4.4 | +2.3 | ×1.000 | 2 / -9 | -1 / -5 | 19 / 27 | 4.6 s / 9.4 s | $0.0023 | CPU container |
| 68 | ZeroEntropy zerank-2 | 0.1 | 4.8 | 81.7 | 80.3 | 48.4 | 10.0 | -0.4 | +10.4 | ×1.000 | 57 / 47 | -68 / -86 | 41 / 37 | 0.47 s / 2.0 s | $0.052 | GPU pod (A6000) |
| 69 | jeff | 0.1 | 4.4 | 80.3 | 55.8 | 81.1 | 12.7 | -3.6 | +16.3 | ×0.969 | 33 / 22 | -28 / -61 | 33 / 28 | 7.1 s / 37.2 s | $0.0043 | CPU container |
| 70 | smalljev semantic-v9 | 0.1 | 3.7 | 73.1 | 86.5 | 60.8 | 10.9 | -3.5 | +14.5 | ×0.987 | 31 / 18 | -40 / -64 | 42 / 35 | 0.45 s / 0.50 s | $0.020 | GPU pod (A6000) |
| 71 | OpenDecision | 0.0 | 2.1 | 72.5 | 86.6 | 79.1 | 9.5 | -5.3 | +14.7 | ×0.985 | 34 / 20 | -47 / -71 | 41 / 35 | 0.30 s / 0.74 s | $0.0050 | GPU pod (H100) |
| 72 | BAAI bge-reranker-v2-m3 | 0.0 | 0.0 | 83.3 | 90.5 | 59.3 | -21.1 | -22.8 | +1.7 | ×1.000 | 4 / 4 | -100 / -100 | 32 / 27 | 0.21 s / 0.42 s | $0.023 | GPU pod (A6000) |
| 73 | Certo v1 | 0.0 | 0.0 | 87.5 | 91.4 | 96.9 | -24.1 | -24.9 | +0.8 | ×1.000 | -2 / -2 | -100 / -100 | 29 / 27 | 0.26 s / 0.27 s | $0.0013 | GPU pod (A6000) |
| 74 | Decision Fast | 0.0 | 0.0 | 76.1 | 91.4 | 80.1 | 7.3 | -11.7 | +19.0 | ×0.942 | 38 / 18 | -59 / -87 | 43 / 33 | 0.27 s / 0.27 s | $0.0046 | GPU pod (RTX6000) |
| 75 | Alibaba GTE Reranker ModernBERT-base | 0.0 | 0.0 | 75.2 | 91.5 | 69.3 | -18.6 | -23.2 | +4.5 | ×1.000 | 13 / 2 | -100 / -100 | 31 / 28 | 0.21 s / 0.34 s | $0.011 | GPU pod (A6000) |
| 76 | kev 0.5B | 0.0 | 0.0 | 65.0 | 88.5 | 80.1 | 3.4 | -22.3 | +25.7 | ×0.875 | 25 / 11 | -48 / -95 | 33 / 17 | 0.36 s / 0.40 s | $0.0046 | GPU pod (A6000) |
| 77 | Laya | 0.0 | 0.0 | 73.7 | 73.9 | 84.9 | -0.5 | -13.2 | +12.7 | ×1.000 | 37 / 13 | -76 / -82 | 37 / 29 | 1.5 s / 2.7 s | $0.0032 | CPU container |
| 78 | lev-350m | 0.0 | 0.0 | 77.7 | 93.6 | 80.0 | 5.5 | -18.0 | +23.5 | ×0.896 | 33 / 15 | -52 / -100 | 36 / 31 | 0.19 s / 0.22 s | $0.0046 | GPU pod (RTX6000) |
| 79 | Qwen3.5-0.8B Decision Model | 0.0 | 0.0 | 73.6 | 71.8 | 79.5 | 3.4 | -3.7 | +7.1 | ×1.000 | 35 / 40 | -70 / -86 | 45 / 34 | 1.3 s / 5.0 s | $0.0048 | CPU container |
| 80 | Mixedbread mxbai-rerank-base-v2 | 0.0 | 0.0 | 86.6 | 89.3 | 60.3 | -21.0 | -24.0 | +3.0 | ×1.000 | 6 / 1 | -100 / -100 | 31 / 27 | 0.23 s / 0.49 s | $0.021 | GPU pod (A6000) |
| 81 | Needle 3 | 0.0 | 0.0 | 0.0 | 34.1 | 61.6 | -11.1 | -21.1 | +10.0 | ×1.000 | -4 / -8 | -15 / -24 | -15 / -32 | 135.5 s / 285.4 s | $0.019 | CPU container |
| 82 | Needle 3, options as tools | 0.0 | 0.0 | 0.0 | 41.0 | 61.6 | -5.8 | -19.9 | +14.1 | ×0.991 | 11 / 2 | -12 / -30 | -16 / -32 | 58.2 s / 136.7 s | $0.019 | CPU container |
| 83 | open-jev-deberta-v3-large | 0.0 | 0.0 | 77.1 | 68.3 | 77.6 | -9.3 | -14.6 | +5.3 | ×1.000 | 21 / 17 | -79 / -87 | 30 / 27 | 2.9 s / 5.1 s | $0.0056 | CPU container |
| 84 | Open Jev JSON Canvas | 0.0 | 77.1 | 0.0 | 85.6 | 49.3 | 79.2 | 75.1 | +4.0 | ×1.000 | 81 / 76 | 75 / 74 | 81 / 75 | 0.44 s / 0.63 s | $0.049 | GPU pod (H100) |
| 85 | openJev Verdict | 0.0 | 0.0 | 52.2 | 83.9 | 86.6 | 12.6 | -14.1 | +26.7 | ×0.865 | 27 / 17 | -15 / -79 | 26 / 20 | 0.40 s / 1.0 s | $0.0028 | CPU container |
| 86 | openJev Verdict 1.4 | 0.0 | 0.0 | 80.3 | 80.6 | 86.6 | -14.7 | -18.3 | +3.5 | ×1.000 | 24 / 18 | -100 / -100 | 31 / 27 | 0.77 s / 1.1 s | $0.0028 | CPU container |
| 87 | Qwen3.8 27BAPI | 0.0 | 95.6 | 98.1 | 56.8 | 0.0 | 92.9 | 98.2 | -5.3 | ×1.000 | 97 / 98 | 87 / 99 | 95 / 98 | 6.5 s / 31.9 s | $2.18 | hosted API |
| 88 | verdict-small | 0.0 | 0.0 | 59.5 | 82.1 | 100.0 | -0.9 | -2.9 | +2.0 | ×1.000 | 26 / 13 | -55 / -44 | 26 / 22 | 0.22 s / 2.9 s | $0.0009 | CPU container |
| 89 | Von | 0.0 | 0.0 | 83.5 | 75.7 | 82.7 | -1.9 | -16.6 | +14.7 | ×0.985 | 34 / 18 | -79 / -98 | 38 / 30 | 0.92 s / 2.9 s | $0.0038 | CPU container |
| – | classifier.devAPIhonorable mention | — | 75.8 | 89.2 | 82.0 | 58.9 | 74.5 | 77.1 | -2.7 | ×1.000 | 84 / 85 | 57 / 64 | 82 / 82 | 0.52 s / 1.2 s | $0.023 | hosted API |
| – | JevK5 v0.3v1.5 roster addendum A1 | — | 56.3 | 88.3 | 93.6 | 63.1 | 61.6 | 50.9 | +10.7 | ×1.000 | 80 / 74 | 42 / 15 | 62 / 64 | 0.18 s / 0.24 s | $0.017 | GPU pod (H100) |
| – | Plumb-4Bv1.5 roster addendum A1 | — | 55.8 | 87.4 | 93.5 | 63.1 | 60.9 | 50.8 | +10.1 | ×1.000 | 80 / 73 | 43 / 17 | 59 / 63 | 0.18 s / 0.24 s | $0.017 | GPU pod (H100) |
| – | Decision 4B v1.2v1.5 roster addendum A1 | — | 53.7 | 88.6 | 93.5 | 63.1 | 56.3 | 51.1 | +5.2 | ×1.000 | 81 / 74 | 20 / 16 | 68 / 63 | 0.18 s / 0.24 s | $0.017 | GPU pod (H100) |
| – | Imajev-4Bv1.5 roster addendum A2 | — | 53.5 | 88.1 | 91.1 | 63.3 | 56.2 | 50.7 | +5.5 | ×1.000 | 80 / 82 | 22 / 8 | 67 / 62 | 0.23 s / 0.33 s | $0.017 | GPU pod (RTX 5090) |
| – | Decision 4B v1.1v1.5 roster addendum A1 | — | 53.1 | 87.3 | 93.5 | 63.1 | 54.7 | 51.6 | +3.1 | ×1.000 | 80 / 75 | 18 / 16 | 66 / 64 | 0.18 s / 0.24 s | $0.017 | GPU pod (H100) |
| – | Surogate Rune 26B-A4B v3v1.5 roster addendum A2 | — | 69.7 | 88.3 | 86.0 | 49.0 | 69.0 | 70.5 | -1.6 | ×1.000 | 82 / 82 | 46 / 50 | 79 / 80 | 0.35 s / 0.72 s | $0.050 | GPU pod (RTX PRO 6000) |
| – | AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1 | — | 72.8 | 87.7 | 87.6 | 29.4 | 71.9 | 73.6 | -1.7 | ×1.000 | 85 / 82 | 50 / 62 | 80 / 77 | 0.35 s / 0.49 s | $0.226 | GPU pod (H100) |
| – | AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2 | — | 72.8 | 86.7 | 87.1 | 29.4 | 72.0 | 73.5 | -1.5 | ×1.000 | 85 / 82 | 51 / 62 | 80 / 77 | 0.33 s / 0.60 s | $0.226 | GPU pod (RTX PRO 6000) |
| – | Eikos-27Bv1.5 roster addendum A1 | — | 75.1 | 86.3 | 87.6 | 28.7 | 74.1 | 76.0 | -1.9 | ×1.000 | 88 / 88 | 56 / 64 | 78 / 76 | 0.35 s / 0.49 s | $0.238 | GPU pod (H100) |
| – | SimpleJev Qwen3.6-35B-A3BAPIpartial run | — | 0.0 | 78.5 | 75.2 | 35.1 | -17.4 | -23.0 | +5.6 | ×1.000 | 16 / 9 | -36 / -43 | -32 / -35 | 1.7 s / 1.8 s | $0.145 | author demo endpoint |
Costs are estimates (est.) unless marked tariff.
API = the operator's endpoint received sealed item text during evaluation, without answers. Sealed item text, answers and item-level results stay private; only system-level aggregates appear here. Hover a cost for its price basis and a latency for its raw values and adjustment.
Listed, not ranked: honorable mention (1)
Services that run on Jev itself are measured and shown, but not ranked against Jev, as in v1.4.2. They do not enter the field median gap or the tie markers.
- classifier.dev (fast tier)APIhonorable mention: runs on Jev (TypeSafe) - listed, not ranked (honorable mention, as in v1.4.2). Official (A) score 74.7.
Roster addendum: newcomers scored on the same frozen protocol (9)
Added by separately hashed roster addenda before they ran. Same frozen sample, method, price rules and v1.5.0 median gap. These rows stay outside the v1.5.0 order and its tie markers. A and secondary B placement compare each row with the frozen base point estimates only; each row's interval is shown separately and does not establish a tie with a base row or another addendum.
| System | Would place (A) | A score · 95% CI | Would place (B) | B score · 95% CI |
|---|---|---|---|---|
| JevK5 v0.3v1.5 roster addendum A1 | #4 | 71.9 69.4–72.9 | #5 | 68.1 64.9–69.8 |
| Plumb-4Bv1.5 roster addendum A1 | #4 | 71.6 69.2–72.7 | #5 | 67.7 64.7–69.6 |
| Decision 4B v1.2v1.5 roster addendum A1 | #6 | 70.8 68.5–72.0 | #7 | 66.6 63.7–68.4 |
| Imajev-4Bv1.5 roster addendum A2 | #6 | 70.4 67.8–71.6 | #7 | 66.2 63.1–68.1 |
| Decision 4B v1.1v1.5 roster addendum A1 | #6 | 70.4 66.9–71.6 | #7 | 66.1 62.1–67.9 |
| Surogate Rune 26B-A4B v3v1.5 roster addendum A2 | #11 | 66.5 65.4–67.0 | #7 | 66.6 65.1–67.5 |
| AutoJev-27B (denis-pplx, Qwen3.8-27B)v1.5 roster addendum A1 | #43 | 19.5 19.3–19.6 | #40 | 20.4 20.1–20.6 |
| AutoJev-27B (RTX PRO 6000)v1.5 roster addendum A2 | #43 | 19.5 19.3–19.6 | #40 | 20.4 20.1–20.6 |
| Eikos-27Bv1.5 roster addendum A1 | #44 | 18.5 18.3–18.6 | #40 | 19.5 19.2–19.7 |
Score followed by its 95% interval. Placements compare point estimates with the frozen base only.
Not ranked: partial, unpriced and unmeasured systems
These systems are part of the 103-system v1.5 roster but have no rank. Their numbers are never shown as zero or free.
Partial runs (1)
- SimpleJev Qwen3.6-35B-A3BAPIpartial run: Partial run: 677 of 1,624 decisions answered; the missing ones count wrong and the row is not ranked.
Not measured in v1.5 (3)
Jobe Qwen3.5-4B · mica-v01-4bv1.5 roster addendum A1 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)
Not measured means no v1.5 run exists yet (for example no offline image, or a hosted API not cleared for sealed items). Their v1.4.2 results stay on the v1.4.2 page.
Method notes: what changed in v1.5
Frozen method METHOD-v1.5, SHA-256 c25d3d8b8512…; pricing addendum v1.5-M2, SHA-256 2fc44459ef80….
The method owner chose equal axis weights and equal weights for Choice, Noul and Score after reviewing the What-If Lab, preserving continuity with v1.4 and treating the three decision types equally. Disclosed headline amendment: equal-axis, equal-type A, SHA-256 752ddccc4e19…. B remains a secondary view.
- 1,624 decisions per system: 904 open (601 published) and 720 sealed, drawn fresh from a private pool with the same tier mix as the open set. Sealed counts for 50% of Intelligence:
base = 0.5 × I_open + 0.5 × I_sealed. - Three request types are scored natively and chance-corrected per item: Choice, Noul and Score each receive one third. Tier weights easy / standard / judge / hard = 10 / 20 / 30 / 40. A type a system does not support is excluded, never scored zero; only full-coverage systems are ranked.
- Overfit penalty relative to the field:
excess = gap − G_med,penalty = max(0, 1 − max(0, excess − 8) / 100). G_med for this batch is 5.2 CC points. - Calibration is typed (Choice ECE/TVD, Noul ECE with Brier, Score normalised RPS and top-level ECE), pooled over open and sealed. Speed and Cost formulas are unchanged from v1.4; self-hosted and demo endpoints carry the ×2 + 0.15 s adjustment. A manufacturer's standard, non-promotional launch list price counts from day one, but a newer price cut younger than 30 days does not. Rows without token counts use the measured proxy-token basis. A system without any eligible public, bookable price is listed as unpriced.
- The frozen 25 Sep DeepInfra snapshot records Qwen3.5-4B as deprecated on 11 Jun 2026 and replaced by Qwen3.5-9B. Its frozen snapshot rates remain the v1.5 M2 reference; price basis tooltips and the correction note disclose this. Pricing disclosure correction SHA-256:
1b660648bd49…. - Composite: weighted harmonic mean with the Intelligence, Speed and Cost gates below 50 (Intelligence below 60 in option C). The official headline A uses equal 25 / 25 / 25 / 25 axis weights and Intelligence floor 50. B remains the secondary 40 / 20 / 20 / 20 view; C keeps equal axes and Intelligence floor 60. Ties come from the paired bootstrap.
- Rows marked with a v1.5 roster addendum label were added by separately hashed roster addenda: same frozen sample, method, pricing rules and G_med. They remain outside the base release order and its tie markers.
- Before every release we review the leaderboard for anomalies and close loopholes with general, documented rules. The page and Git repository provide transparent data and method details; Benchmark Heaven owns its rules.
Data file SHA-256 6b2f6b058b36203c98ec5f585eb8376038bc905f11db944f4e0bcd29c278c643 · scorer output SHA-256 452885de2a84cd5b9ed393d541fd8d6c6a540f9d7f762f9d2df9349383f26173 · run kind official.
Previous release: JevBench v1.4.2.2 (frozen results).