Frozen JevBench release v1.4.1

JevBench v1.4.1 — frozen results

This page always uses the public, hash-checked v1.4.1 artifact. The live board may change when a later release is published.

Artifact SHA-256 e6754863056503fe2b010410fc7111df884ac1f9ce4449aa369aab61d98092cd.

Share this version · View live board

Frozen top five

  1. Jev 1.13.063.3
  2. JevK5 v0.2.062.0
  3. Hopper59.4
  4. Winnow-12B Q855.6
  5. reflex 4B54.0

These ranks and scores come from the frozen v1.4.1 release and do not follow changes to the live board.

JevBench v1.4.1 ranking

77 ranked systems and 5 unranked rows, measured on 534 public decisions plus 308 sealed decisions. The sealed text and answers remain private; only system-level aggregates appear here.

JevBench v1.4.1 · 534 public + 308 sealed decisions per system

JevBench Score: 77 ranked systems

OfficialIntelligence, Calibration, Speed and Cost, each 0–100 — equal-weight harmonic mean, with the generalization and Jev-class gates. What changed in v1.4 ↓

  1. 1Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
  2. 2JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
  3. 3Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
  4. 4Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
  6. 6djev52.2I 47 · C 55 · S 91 · K 58 · $0.026 ann.
  7. 7Jev-Omni51.3I 47 · C 64 · S 82 · K 53 · ~$0.037 est.
  8. 8metask-jev-4b47.8I 45 · C 67 · S 89 · K 55 · ~$0.033 est.
  9. 9SemIf47.7I 44 · C 67 · S 84 · K 59 · ~$0.022 est.
  10. 10Jobe Qwen3.5-4B46.9I 44 · C 66 · S 86 · K 60 · ~$0.022 est.
  11. 11local-jev Qwen3.5-4B46.8I 44 · C 73 · S 75 · K 56 · ~$0.030 est.
  12. 12system-one-openAPI45.1I 44 · C 55 · S 77 · K 65 · ~$0.015 est.
  13. 13spark-s1-4b-v644.6I 45 · C 48 · S 81 · K 58 · ~$0.025 est.
  14. 14jqv44.4I 46 · C 72 · S 75 · K 47 · ~$0.056 est.
  15. 15Qwen3-Reranker-4B43.5I 45 · C 65 · S 79 · K 49 · $0.050
  16. 16decider-35b-a3b41.2I 47 · C 65 · S 81 · K 45 · ~$0.067 est.
  17. 17Raw Qwen3 4B Instruct 2507 direct logits41.0I 46 · C 29 · S 88 · K 60 · ~$0.022 est.
  18. 18OpenSourceJev40.9I 42 · C 60 · S 82 · K 64 · ~$0.016 est.
  19. 19ZeroEntropy zerank-240.2I 42 · C 76 · S 79 · K 50 · $0.047
  20. 20decision-machine-1API39.9I 41 · C 68 · S 93 · K 54 · $0.035
Show all 82 systems (57 more ranked, 5 not ranked)
  1. 21Raw Phi-4 mini direct logits38.0I 42 · C 59 · S 89 · K 50 · ~$0.048 est.
  2. 22JEV Qwen3.5-9B Base NVFP437.7I 47 · C 68 · S 93 · K 43 · ~$0.077 est.
  3. 23OpenJev36.9I 45 · C 55 · S 83 · K 45 · ~$0.066 est.
  4. 24kev 4B36.1I 42 · C 40 · S 76 · K 62 · ~$0.019 est.
  5. 25Decision 2B35.8I 39 · C 74 · S 84 · K 63 · ~$0.018 est.
  6. 26Qwen3.5-9B Jev-like data-mix v235.2I 47 · C 61 · S 82 · K 42 · ~$0.083 est.
  7. 27GPT-6 LunaAPI35.1I 96 · C 92 · S 74 · K 37 · $0.127
  8. 28SimpleJev Qwen3.8-27BAPI34.6I 52 · C 74 · S 71 · K 39 · ~$0.104 est.
  9. 29NInfer Qwen3.8-Flash-Next mixed34.0I 50 · C 79 · S 88 · K 39 · ~$0.109 est.
  10. 30GPT-6 LunaAPI33.3I 97 · C 93 · S 73 · K 36 · $0.135
  11. 31open-alternative-jev33.2I 39 · C 59 · S 83 · K 60 · ~$0.022 est.
  12. 32jev-local32.5I 45 · C 64 · S 69 · K 43 · ~$0.077 est.
  13. 33Decision Fast32.5I 37 · C 65 · S 82 · K 76 · ~$0.0063 est.
  14. 34decider-2b30.7I 39 · C 43 · S 83 · K 61 · ~$0.020 est.
  15. 35jeff30.6I 37 · C 68 · S 63 · K 77 · ~$0.0060 est.
  16. 36Laya30.3I 36 · C 64 · S 71 · K 86 · ~$0.0029 est.
  17. 37lev-350m28.5I 35 · C 71 · S 85 · K 76 · ~$0.0063 est.
  18. 38openjev-sglangAPI27.7I 49 · C 69 · S 77 · K 36 · ~$0.131 est.
  19. 39Von27.5I 34 · C 76 · S 70 · K 78 · ~$0.0055 est.
  20. 40NInfer Qwen3.8-27B NVFP426.9I 51 · C 76 · S 80 · K 35 · ~$0.145 est.
  21. 41NInfer Qwen3.8-27B NVFP426.3I 51 · C 67 · S 80 · K 35 · ~$0.145 est.
  22. 42kev 8B25.6I 42 · C 40 · S 75 · K 44 · ~$0.073 est.
  23. 43JevOne25.5I 48 · C 78 · S 88 · K 36 · ~$0.137 est.
  24. 44SimpleJev Qwen3.6-35B-A3BAPI24.9I 46 · C 60 · S 75 · K 38 · ~$0.116 est.
  25. 45kev 0.6B24.8I 34 · C 50 · S 76 · K 76 · ~$0.0063 est.
  26. 46Raw Qwen3 8B direct logits23.7I 46 · C 24 · S 86 · K 42 · ~$0.087 est.
  27. 47system-one23.4I 44 · C 33 · S 84 · K 41 · ~$0.089 est.
  28. 48OpenDecision21.6I 32 · C 57 · S 80 · K 75 · ~$0.0066 est.
  29. 49LitJev19.5I 46 · C 77 · S 67 · K 34 · ~$0.163 est.
  30. 50openJev Verdict 1.419.0I 29 · C 72 · S 78 · K 82 · ~$0.0039 est.
  31. 51kev 0.5B18.9I 31 · C 50 · S 77 · K 76 · ~$0.0063 est.
  32. 52Bespoke Nimble 9B18.7I 46 · C 56 · S 79 · K 33 · ~$0.166 est.
  33. 53GPT-5.6 LunaAPI18.5I 93 · C 87 · S 78 · K 28 · $0.242
  34. 54openJev Verdict18.1I 30 · C 47 · S 77 · K 83 · ~$0.0037 est.
  35. 55Raw Qwen3 1.7B direct logits18.1I 33 · C 24 · S 90 · K 65 · ~$0.015 est.
  36. 56reflex-27b17.8I 46 · C 77 · S 67 · K 32 · ~$0.181 est.
  37. 57djev15.2I 72 · C 88 · S 75 · K 27 · ~$0.274 est.
  38. 58GLiNER2 large15.1I 31 · C 25 · S 62 · K 73 · ~$0.0077 est.
  39. 59OpenJev14.8I 58 · C 58 · S 76 · K 28 · ~$0.255 est.
  40. 60Qwen3.5-0.8B Decision Model14.5I 28 · C 68 · S 49 · K 76 · ~$0.0065 est.
  41. 61Gemini 3.1 Flash-LiteAPI14.3I 54 · C 59 · S 82 · K 27 · $0.264
  42. 62open-jev-deberta-v3-large12.6I 26 · C 67 · S 66 · K 74 · ~$0.0073 est.
  43. 63smalljev semantic-v912.3I 26 · C 59 · S 80 · K 58 · ~$0.025 est.
  44. 64GLiNER211.8I 27 · C 25 · S 72 · K 83 · ~$0.0037 est.
  45. 65Open-Jev 9B11.2I 44 · C 62 · S 72 · K 28 · ~$0.249 est.
  46. 66Open-Jev 2B10.0I 42 · C 55 · S 73 · K 28 · ~$0.249 est.
  47. 67GLiNER2.5 multi9.8I 23 · C 57 · S 68 · K 82 · ~$0.0039 est.
  48. 68SimpleJev7.5I 21 · C 49 · S 57 · K 68 · ~$0.011 est.
  49. 69GLiNER2.5 small7.2I 20 · C 51 · S 78 · K 82 · ~$0.0039 est.
  50. 70Raw Qwen3 0.6B direct logits7.1I 23 · C 21 · S 90 · K 74 · ~$0.0074 est.
  51. 71DeepSeek V4.1 FlashAPI4.8I 94 · C 96 · S 72 · K 17 · $0.594
  52. 72Mirror2.1I 14 · C 26 · S 71 · K 73 · ~$0.0077 est.
  53. 73Mixedbread mxbai-rerank-base-v20.4I 7 · C 84 · S 88 · K 68 · $0.012
  54. 74BAAI bge-reranker-v2-m30.2I 5 · C 84 · S 90 · K 73 · $0.0077
  55. 75Alibaba GTE Reranker ModernBERT-base0.2I 5 · C 79 · S 91 · K 70 · $0.010
  56. 76Certo v10.0I 0 · C 83 · S 94 · K 100 · ~$0.0010 est.
  57. 77Open Jev JSON Canvas0.0I 48 · C 0 · S 84 · K 46 · ~$0.065 est.
  58. classifier.dev (honorable mention)API70.8I 52 · C 72 · S 88 · K 84 · ~$0.0033 est.
  59. Qwen3.8 27B (partial run)API0.0I 40 · C 94 · S 61 · K 0 · ~$2.669 est.
  60. swanOne (partial run)I · C · S · K · ~$0.111 est.
  61. Needle 3, options as tools (partial run)I · C · S · K · ~$0.014 est.
  62. Needle 3 (partial run)I · C · S · K · ~$0.024 est.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; ~ est. = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers. Names link to each project.

Compare two systems

Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, and accuracy by family on the v1.2 hard tier and on the sealed set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 Jev (TypeSafe, closed) · Score 63.3 (#1)
  • B: JevK5 v0.2.0 Jev rebuild · Score 62.0 (#2)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs JevK5 v0.2.0. Intelligence: 53.1 vs 48.9; Calibration: 76.3 vs 74.5; Speed: 83.3 vs 91.1; Cost: 52.0 vs 59.5.50100Intelligence53.1 · 48.9Calibration76.3 · 74.5Speed83.3 · 91.1Cost52.0 · 59.5
0–100, the values in the table. A label-only system has no calibration (counted as 0).

Accuracy per tier, incl. sealed

Radar: accuracy per tier, incl. sealed, two systemsAccuracy per tier, incl. sealed, Jev 1.13.0 vs JevK5 v0.2.0. Easy: 100% vs 100%; Standard: 99% vs 96%; Judge: 95% vs 95%; Hard: 74% vs 70%; Sealed: 37% vs 33%.50100Easy100% · 100%Standard99% · 96%Judge95% · 95%Hard74% · 70%Sealed37% · 33%
Share correct per tier; Sealed = the 308 private decisions, aggregate only.

Hard tier by family (v1.2 topics)

Radar: hard tier by family (v1.2 topics), two systemsHard tier by family (v1.2 topics), Jev 1.13.0 vs JevK5 v0.2.0. Adversarial: 100% vs —; Ambiguous: 79% vs —; Judge: 79% vs —; Long policy: 61% vs —; Multi-hop: 86% vs —; Probability: 80% vs —; Routing: 100% vs —; Temporal / numeric: 27% vs —; Trade-off: 92% vs —; Trap: 100% vs —.50100Adversarial100% · Ambiguous79% · Judge79% · Long policy61% · Multi-hop86% · Probability80% · Routing100% · Temporal /numeric27% · Trade-off92% · Trap100% ·
Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out). “—” = not measured for that system, not plotted.

Sealed set by family

Radar: sealed set by family, two systemsSealed set by family, Jev 1.13.0 vs JevK5 v0.2.0. Ambiguous / abstain: 30% vs 43%; Judge: 34% vs 37%; Long policy: 28% vs 23%; Multi-hop: 45% vs 39%; Paraphrase: 64% vs 14%; Probability: 50% vs 39%; Safety judge: 38% vs 38%; Temporal / numeric: 29% vs 27%; Trade-off: 38% vs 23%; Trap / adversarial: 42% vs 58%.50100Ambiguous /abstain30% · 43%Judge34% · 37%Long policy28% · 23%Multi-hop45% · 39%Paraphrase64% · 14%Probability50% · 39%Safety judge38% · 38%Temporal /numeric29% · 27%Trade-off38% · 23%Trap /adversarial42% · 58%
Share correct within each sealed family — system-level aggregates; the items stay private.
All values as a table
SpokeA: Jev 1.13.0B: JevK5 v0.2.0
The four score axes
Intelligence53.148.9
Calibration76.374.5
Speed83.391.1
Cost52.059.5
Accuracy per tier, incl. sealed
Easy100%100%
Standard99%96%
Judge95%95%
Hard74%70%
Sealed37%33%
Hard tier by family (v1.2 topics)
Adversarial100%
Ambiguous79%
Judge79%
Long policy61%
Multi-hop86%
Probability80%
Routing100%
Temporal / numeric27%
Trade-off92%
Trap100%
Sealed set by family
Ambiguous / abstain30%43%
Judge34%37%
Long policy28%23%
Multi-hop45%39%
Paraphrase64%14%
Probability50%39%
Safety judge38%38%
Temporal / numeric29%27%
Trade-off38%23%
Trap / adversarial42%58%

What changed in v1.4

  • Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence: I = 0.8 × I_v1.3 + 0.2 × I_sealed, where I_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates.
  • Calibration blends toward the sealed-inclusive result at the approved weight: C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35).
  • The k = 1 generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points: I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half.
  • The four axes use an equal-weight harmonic mean (p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0.
  • The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.

Axes, accuracy, latency and cost

Every system with its four axes, public and sealed accuracy and the gap between them. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.

#SystemJevBench ScoreIntelligenceCalibrationSpeedCost axisPublic accuracy
534
Sealed accuracy
308
Public − sealed gapCost / 1,000p50 latencyEndpoint
1
Jev 1.13.0APIby TypeSafe AI
63.353.176.383.352.086.6%36.7%+49.9 pp$0.0400.65 sAPI
2
JevK5 v0.2.0
Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
by allebee
62.048.974.591.159.585.3%33.1%+52.2 pp~$0.022est.unknown
3
Hopper
Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
by HopitAI
59.448.079.186.858.782.3%34.1%+48.2 pp~$0.024est.0.13 sRunPod GPU
4
Winnow-12B Q8
The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
by Eldan Ring
55.648.364.882.352.985.7%33.1%+52.6 pp~$0.037est.0.23 sRunPod GPU
5
reflex 4B
The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
by kshetrajna12
54.047.570.468.059.779.2%28.2%+51.0 pp~$0.022est.1.80 sRunPod GPU
6
djev
The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
by Maisa (David Villalón) · Maisa, diffusion-gemma
52.247.055.491.457.684.0%29.9%+54.1 pp$0.026announced0.24 sAPI
7
Jev-Omni
Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by akhilaaa3 · akhilaaa3, Gemma-4-12B merged
51.346.864.181.553.088.7%32.1%+56.6 pp~$0.037est.0.22 sRunPod GPU
8
metask-jev-4b
Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
by Wayfind (metask-ai)
47.844.766.989.154.579.7%27.6%+52.1 pp~$0.033est.0.07 sRunPod GPU
9
SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ
47.744.466.883.759.581.0%26.3%+54.7 pp~$0.022est.0.20 sRunPod GPU
10
Jobe Qwen3.5-4B
No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
by MantisShrimpdev · frozen
46.944.166.185.659.581.0%25.6%+55.3 pp~$0.022est.0.13 sRunPod GPU
11
local-jev Qwen3.5-4B
Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
by Amith Chandrappa (amithgc)
46.844.473.375.055.880.5%26.0%+54.5 pp~$0.030est.0.71 sRunPod GPU
12
system-one-openAPIby mithalouni · Gemma 4 E2B LoRA on an L4
45.144.254.977.064.873.2%27.6%+45.6 pp~$0.015est.0.65 sauthor demo
13
spark-s1-4b-v6
Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085
44.645.147.681.057.979.2%26.6%+52.6 pp~$0.025est.0.31 sRunPod GPU
14
jqv
A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
by hjmurmur (Octalab) · Qwen3-32B zero-shot
44.446.471.674.647.580.1%28.2%+51.8 pp~$0.056est.0.75 sRunPod GPU
15
Qwen3-Reranker-4B
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Qwen
43.544.665.278.749.268.0%29.9%+38.1 pp$0.0500.13 sRunPod GPU
16
decider-35b-a3b
The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
by Mapika
41.247.265.380.845.383.1%31.5%+51.6 pp~$0.067est.0.29 sRunPod GPU
17
Raw Qwen3 4B Instruct 2507 direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
41.046.429.187.659.769.7%27.3%+42.4 pp~$0.022est.0.08 sRunPod GPU
18
OpenSourceJev
DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp
40.941.860.382.064.078.4%26.3%+52.1 pp~$0.016est.unknown
19
ZeroEntropy zerank-2
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by ZeroEntropy
40.242.175.879.049.870.1%28.6%+41.6 pp$0.0470.13 sRunPod GPU
20
decision-machine-1
A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
APIby milliseconds.ai (Baptiste Laget)
39.941.368.392.953.767.5%25.6%+41.9 pp$0.0350.17 sAPI
21
Raw Phi-4 mini direct logits
Neutral raw-logit control, not JevBench-directed.
by Microsoft / neutral reproduction
38.041.858.888.849.665.8%29.2%+36.6 pp~$0.048est.0.06 sRunPod GPU
22
JEV Qwen3.5-9B Base NVFP4
Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
by WilfLin
37.746.867.793.343.375.3%29.5%+45.8 pp~$0.077est.0.02 sRunPod GPU
23
OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16
36.945.455.083.245.581.8%28.6%+53.2 pp~$0.066est.0.24 sRunPod GPU
24
kev 4B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
36.142.139.675.761.866.2%22.4%+43.8 pp~$0.019est.0.55 sRunPod GPU
25
Decision 2B
Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v59
35.838.874.184.362.575.3%26.0%+49.4 pp~$0.018est.0.19 sRunPod GPU
26
Qwen3.5-9B Jev-like data-mix v2
The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
by jsaurabh
35.247.461.382.042.478.4%29.2%+49.1 pp~$0.083est.unknown
27
GPT-6 Luna
OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · low reasoning effort
35.195.892.073.736.999.1%92.9%+6.3 pp$0.1271.44 sAPI
28
SimpleJev Qwen3.8-27B
Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
34.651.674.571.239.586.6%35.7%+50.9 pp~$0.104est.1.01 sauthor demo
29
NInfer Qwen3.8-Flash-Next mixed
Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
by Igor L. / NInfer contributors
34.049.578.688.238.989.6%34.1%+55.5 pp~$0.109est.0.08 sRunPod GPU
30
GPT-6 Luna
OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · default medium reasoning effort
33.397.493.572.636.099.6%95.5%+4.1 pp$0.1351.48 sAPI
31
open-alternative-jev
With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
by IkerMoel · Qwen3.5-4B, IkerMoel
33.238.658.783.559.674.0%24.4%+49.7 pp~$0.022est.0.21 sRunPod GPU
32
jev-local
The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
by us (GitHub) · Qwen3.5-9B
32.545.264.269.243.374.9%29.5%+45.3 pp~$0.077est.1.05 sRunPod GPU
33
Decision Fast
Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v53a
32.537.165.381.676.163.2%25.6%+37.6 pp~$0.0063est.0.24 sRunPod GPU
34
decider-2b
The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
by Mapika
30.738.543.583.261.071.0%24.7%+46.3 pp~$0.020est.0.26 sRunPod GPU
35
jeff
Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
by Logan Markewich · Logan Markewich, GLiFormer 400M
30.636.867.963.576.662.8%33.1%+29.7 pp~$0.0060est.0.94 sCPU
36
Laya
The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
by Convai Innovations · Convai Innovations, ModernBERT-large 421M
30.336.163.771.186.258.4%30.8%+27.6 pp~$0.0029est.0.79 sCPU
37
lev-350m
Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M
28.534.870.685.376.158.4%25.0%+33.4 pp~$0.0063est.0.17 sRunPod GPU
38
openjev-sglangAPIby ekzhang · Qwen3.6-35B-A3B on SGLang
27.749.469.277.136.585.3%33.1%+52.2 pp~$0.131est.0.68 sauthor demo
39
Von
The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M
27.534.575.770.577.857.1%27.9%+29.2 pp~$0.0055est.unknown
40
NInfer Qwen3.8-27B NVFP4
T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
by Igor L. / NInfer contributors · T=1.5
26.951.576.080.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
41
NInfer Qwen3.8-27B NVFP4
Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
by Igor L. / NInfer contributors
26.351.567.280.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
42
kev 8B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
by Jared Palmer · research preview
25.641.840.274.944.071.4%21.8%+49.7 pp~$0.073est.0.59 sRunPod GPU
43
JevOne
Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
by Juspay
25.547.677.688.535.989.6%33.8%+55.8 pp~$0.137est.0.09 sRunPod GPU
44
SimpleJev Qwen3.6-35B-A3B
Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
24.945.759.875.038.181.4%28.2%+53.1 pp~$0.116est.0.85 sauthor demo
45
kev 0.6B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
24.834.250.075.676.166.7%24.0%+42.6 pp~$0.0063est.0.59 sRunPod GPU
46
Raw Qwen3 8B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
23.745.724.186.341.968.4%26.3%+42.1 pp~$0.087est.0.08 sRunPod GPU
47
system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke
23.443.632.884.441.571.9%24.4%+47.5 pp~$0.089est.0.17 sRunPod GPU
48
OpenDecision
A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
by Deepan Wadhwa · ModernBERT-large zero-shot
21.631.857.179.975.353.2%25.6%+27.6 pp~$0.0066est.0.34 sRunPod GPU
49
LitJev
The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
by Zhengxu Yu · Qwen3.8-27B
19.546.376.666.733.686.1%30.8%+55.3 pp~$0.163est.2.03 sRunPod GPU
50
openJev Verdict 1.4
Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
by Hemant (heman10x)
19.029.472.078.182.457.6%27.9%+29.7 pp~$0.0039est.0.31 sCPU
51
kev 0.5B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer
18.930.549.777.076.149.4%27.3%+22.1 pp~$0.0063est.0.43 sRunPod GPU
52
Bespoke Nimble 9B
Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
by Bespoke Labs
18.746.356.478.733.479.7%28.9%+50.8 pp~$0.166est.0.39 sRunPod GPU
53
GPT-5.6 LunaAPIby OpenAI · low reasoning effort
18.593.187.477.528.597.4%89.0%+8.4 pp$0.2420.97 sAPI
54
openJev Verdict
The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
by Hemant (heman10x) · heman10x, ModernBERT-base 151M
18.130.047.076.783.155.4%24.7%+30.7 pp~$0.0037est.0.28 sCPU
55
Raw Qwen3 1.7B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
18.133.224.389.764.954.1%26.0%+28.1 pp~$0.015est.0.07 sRunPod GPU
56
reflex-27b
The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
by kshetrajna12 · Qwen3.8-27B
17.846.477.267.532.387.0%29.5%+57.5 pp~$0.181est.1.89 sRunPod GPU
57
djev
Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
by David Villalon / Maisa · thinking
15.271.687.875.226.987.4%60.1%+27.4 pp~$0.274est.0.43 sRunPod GPU
58
GLiNER2 large
The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino
15.131.124.861.773.356.7%28.6%+28.1 pp~$0.0077est.1.10 sCPU
59
OpenJev
OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
by razorback16 · thinking, BF16
14.858.158.176.127.888.7%42.2%+46.5 pp~$0.255est.0.46 sRunPod GPU
60
Qwen3.5-0.8B Decision Model
JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
by Mourad Ghafiri
14.528.168.249.275.759.3%34.7%+24.6 pp~$0.0065est.7.15 sCPU
61
Gemini 3.1 Flash-LiteAPIby Google
14.354.559.381.827.487.0%38.6%+48.4 pp$0.2640.76 sAPI
62
open-jev-deberta-v3-large
297/308 sealed items answered validly (failures count as wrong)
by Kotoba Labs · local CPU
12.625.666.666.074.052.4%29.5%+22.8 pp~$0.0073est.1.77 sCPU
63
smalljev semantic-v9
The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
by Aditya (isHeSatoshi)
12.325.759.279.857.960.6%26.9%+33.7 pp~$0.025est.0.41 sRunPod GPU
64
GLiNER2
A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
by Fastino · Fastino, gliner2.5-base
11.827.425.271.883.158.0%29.2%+28.8 pp~$0.0037est.0.31 sCPU
65
Open-Jev 9B
The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
11.244.261.872.028.177.5%29.9%+47.6 pp~$0.249est.0.75 sRunPod GPU
66
Open-Jev 2B
The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
10.042.355.373.528.164.5%26.3%+38.2 pp~$0.249est.0.66 sRunPod GPU
67
GLiNER2.5 multi
The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
by Fastino · Fastino, 287M
9.823.157.267.882.448.9%32.8%+16.1 pp~$0.0039est.0.43 sCPU
68
SimpleJev
143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU
7.521.549.157.568.354.5%34.7%+19.8 pp~$0.011est.3.99 sCPU
69
GLiNER2.5 small
The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino · Fastino, 74M
7.220.550.777.882.445.9%28.6%+17.3 pp~$0.0039est.0.11 sCPU
70
Raw Qwen3 0.6B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
7.122.820.789.973.948.1%25.6%+22.4 pp~$0.0074est.0.07 sRunPod GPU
71
DeepSeek V4.1 FlashAPIby DeepSeek · thinking default
4.894.095.571.616.897.8%94.8%+3.0 pp$0.5941.42 sAPI
72
Mirror
171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
by Bluusun
2.113.626.070.873.341.1%9.1%+32.0 pp~$0.0077est.0.90 sauthor demo
73
Mixedbread mxbai-rerank-base-v2
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Mixedbread
0.46.884.187.567.937.2%34.4%+2.8 pp$0.0120.07 sRunPod GPU
74
BAAI bge-reranker-v2-m3
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by BAAI
0.25.084.289.573.439.4%27.9%+11.5 pp$0.00770.03 sRunPod GPU
75
Alibaba GTE Reranker ModernBERT-base
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Alibaba-NLP
0.24.878.990.669.633.8%33.4%+0.3 pp$0.0100.05 sRunPod GPU
76
Certo v1
The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
by AltSlate Labs
0.00.183.094.0100.031.6%29.5%+2.1 pp~$0.0010est.0.02 sRunPod GPU
77
Open Jev JSON Canvas
Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
by JoshuaSP
0.048.20.084.145.684.4%31.2%+53.2 pp~$0.065est.0.22 sRunPod GPU
classifier.dev
Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
APIby mrmps (@michael_chomsky) · fast tier
honorable mention · not ranked
70.851.672.487.684.385.3%34.4%+50.9 pp~$0.0033est.0.39 sAPI
Qwen3.8 27BAPIby Qwen / Chutes · Chutes TEE
partial · not ranked
0.040.493.661.30.071.9%21.8%+50.1 pp~$2.669est.5.75 sAPI
swanOne
Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
by swanOne submitter
partial · not ranked
88.7%~$0.111est.0.34 sRunPod GPU
Needle 3, options as tools
V1.4: Not re-measured: same as needle-3.
by Cactus Compute · post-hoc adapter mode
partial · not ranked
22.1%~$0.014est.3.78 sCPU
Needle 3
V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.
by Cactus Compute · Cactus, 2-bit, local CPU
partial · not ranked
22.5%~$0.024est.1.69 sCPU

API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.

All 72 system notes and disclosures
  • JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
  • Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
  • Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
  • local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
  • OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
  • JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
  • kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
  • Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
  • GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
  • GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
  • NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
  • NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
  • kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
  • Raw Qwen3 8B direct logits: Neutral raw-logit control.
  • OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
  • Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
  • reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
  • GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
  • open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
  • GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
  • Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
  • swanOne: Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
  • Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
  • Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.

Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.

Artifact: v1.4.1 results JSON · SHA-256 e67548630565 · JevBench v1.4.1 release and method

JevBench v1.4.1 · additional views

Capability, cost and speed

Capability is the arithmetic mean of Intelligence and Calibration: (Intelligence + Calibration) / 2, on a 0–100 scale. Cost is USD per 1,000 decisions; its axis is logarithmic, and lower is better. Speed uses the JevBench Speed axis, where higher is faster. Estimated costs are marked.

3 placeholder rows have no published Intelligence or Calibration values and are omitted.

Top 20 by Capability

Capability with cost alongside

Each system has a wide Capability bar and a narrower cost bar. The cost scale is logarithmic: longer bars mean higher cost, so shorter is cheaper.

  1. 1GPT-6 LunaAPI95.4Cost $0.14
  2. 2DeepSeek V4.1 FlashAPI94.7Cost $0.59
  3. 3GPT-6 LunaAPI93.9Cost $0.13
  4. 4GPT-5.6 LunaAPI90.3Cost $0.24
  5. 5djev79.7Cost $0.27 est.
  6. 6Qwen3.8 27BAPI67.0Cost $2.67 est.
  7. 7Jev 1.13.0API64.7Cost $0.040
  8. 8NInfer Qwen3.8-Flash-Next mixed64.1Cost $0.11 est.
  9. 9NInfer Qwen3.8-27B NVFP463.7Cost $0.14 est.
  10. 10Hopper63.5Cost $0.024 est.
  11. 11SimpleJev Qwen3.8-27BAPI63.0Cost $0.10 est.
  12. 12JevOne62.6Cost $0.14 est.
  13. 13classifier.devAPI62.0Cost $0.0033 est.
  14. 14reflex-27b61.8Cost $0.18 est.
  15. 15JevK5 v0.2.061.7Cost $0.022 est.
  16. 16LitJev61.4Cost $0.16 est.
  17. 17openjev-sglangAPI59.3Cost $0.13 est.
  18. 18NInfer Qwen3.8-27B NVFP459.3Cost $0.14 est.
  19. 19jqv59.0Cost $0.056 est.
  20. 20reflex 4B58.9Cost $0.022 est.
Show all 79 systems (59 more)
  1. 21ZeroEntropy zerank-258.9Cost $0.047
  2. 22local-jev Qwen3.5-4B58.9Cost $0.030 est.
  3. 23OpenJev58.1Cost $0.25 est.
  4. 24JEV Qwen3.5-9B Base NVFP457.3Cost $0.077 est.
  5. 25Gemini 3.1 Flash-LiteAPI56.9Cost $0.26
  6. 26Winnow-12B Q856.6Cost $0.037 est.
  7. 27Decision 2B56.4Cost $0.018 est.
  8. 28decider-35b-a3b56.2Cost $0.067 est.
  9. 29metask-jev-4b55.8Cost $0.033 est.
  10. 30SemIf55.6Cost $0.022 est.
  11. 31Jev-Omni55.4Cost $0.037 est.
  12. 32Von55.1Cost $0.0055 est.
  13. 33Jobe Qwen3.5-4B55.1Cost $0.022 est.
  14. 34Qwen3-Reranker-4B54.9Cost $0.050
  15. 35decision-machine-1API54.8Cost $0.035
  16. 36jev-local54.7Cost $0.077 est.
  17. 37Qwen3.5-9B Jev-like data-mix v254.4Cost $0.083 est.
  18. 38Open-Jev 9B53.0Cost $0.25 est.
  19. 39SimpleJev Qwen3.6-35B-A3BAPI52.8Cost $0.12 est.
  20. 40lev-350m52.7Cost $0.0063 est.
  21. 41jeff52.3Cost $0.0060 est.
  22. 42Bespoke Nimble 9B51.4Cost $0.17 est.
  23. 43Decision Fast51.2Cost $0.0063 est.
  24. 44djev51.2Cost $0.026
  25. 45OpenSourceJev51.1Cost $0.016 est.
  26. 46openJev Verdict 1.450.7Cost $0.0039 est.
  27. 47Raw Phi-4 mini direct logits50.3Cost $0.048 est.
  28. 48OpenJev50.2Cost $0.066 est.
  29. 49Laya49.9Cost $0.0029 est.
  30. 50system-one-openAPI49.5Cost $0.015 est.
  31. 51Open-Jev 2B48.8Cost $0.25 est.
  32. 52open-alternative-jev48.6Cost $0.022 est.
  33. 53Qwen3.5-0.8B Decision Model48.2Cost $0.0065 est.
  34. 54spark-s1-4b-v646.3Cost $0.025 est.
  35. 55open-jev-deberta-v3-large46.1Cost $0.0073 est.
  36. 56Mixedbread mxbai-rerank-base-v245.5Cost $0.012
  37. 57BAAI bge-reranker-v2-m344.6Cost $0.0077
  38. 58OpenDecision44.4Cost $0.0066 est.
  39. 59smalljev semantic-v942.4Cost $0.025 est.
  40. 60kev 0.6B42.1Cost $0.0063 est.
  41. 61Alibaba GTE Reranker ModernBERT-base41.9Cost $0.010
  42. 62Certo v141.5Cost $0.00097 est.
  43. 63decider-2b41.0Cost $0.020 est.
  44. 64kev 8B41.0Cost $0.073 est.
  45. 65kev 4B40.9Cost $0.019 est.
  46. 66GLiNER2.5 multi40.1Cost $0.0039 est.
  47. 67kev 0.5B40.1Cost $0.0063 est.
  48. 68openJev Verdict38.5Cost $0.0037 est.
  49. 69system-one38.2Cost $0.089 est.
  50. 70Raw Qwen3 4B Instruct 2507 direct logits37.7Cost $0.022 est.
  51. 71GLiNER2.5 small35.6Cost $0.0039 est.
  52. 72SimpleJev35.3Cost $0.011 est.
  53. 73Raw Qwen3 8B direct logits34.9Cost $0.087 est.
  54. 74Raw Qwen3 1.7B direct logits28.8Cost $0.015 est.
  55. 75GLiNER2 large28.0Cost $0.0077 est.
  56. 76GLiNER226.3Cost $0.0037 est.
  57. 77Open Jev JSON Canvas24.1Cost $0.065 est.
  58. 78Raw Qwen3 0.6B direct logits21.8Cost $0.0074 est.
  59. 79Mirror19.8Cost $0.0077 est.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Cost per 1,000 decisions · log scale
Cost bars use the right-hand scale, from $0.00097 to $2.67 per 1,000 decisions. Free cost is placed at the cheapest edge; missing cost is shown as —.

Capability vs cost

The upper-left is the more attractive area: higher Capability and lower cost.

Capability vs costThe upper-left is the more attractive area: higher Capability and lower cost. Each point has a tooltip with the system and its values.020406080100$0.0010$0.010$0.10$1.00USD per 1,000 decisions · log scale · cheaper ←Capability · higher ↑GPT-6 Luna · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Jev 1.13.0 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.JevK5 v0.2.0 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.metask-jev-4b · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Bespoke Nimble 9B · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.spark-s1-4b-v6 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.smalljev semantic-v9 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.kev 0.6B · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.Raw Qwen3 1.7B direct logits · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
79 systems plotted. Hover or focus a point to read its values.

Capability vs speed

The upper-right is the more attractive area: higher Capability and higher Speed.

Capability vs speedThe upper-right is the more attractive area: higher Capability and higher Speed. Each point has a tooltip with the system and its values.020406080100020406080100Speed axis · higher is faster →Capability · higher ↑GPT-6 Luna · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Jev 1.13.0 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.JevK5 v0.2.0 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.metask-jev-4b · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Bespoke Nimble 9B · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.spark-s1-4b-v6 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.smalljev semantic-v9 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.kev 0.6B · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.Raw Qwen3 1.7B direct logits · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
79 systems plotted. Hover or focus a point to read its values.

All three at once

The 3D view plots Capability vertically, lower cost to the right, and higher Speed toward you. Sphere size follows the JevBench score. Drag to rotate; pinch or scroll to zoom. The view loads when it scrolls into view.

Scroll here to load the interactive 3D view.

Vertical: Capability · Right: cheaper · Toward you: faster

The interactive 3D view loads when this panel scrolls into view.

79 systems plotted; systems missing cost or Speed are omitted. three.js r128 is included under its MIT license.