JevBench v1.4.1

JevBench by Benchmark Heaven

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Compare Jev alternativesChoose a Jev-class model by use case

Version v1.4.1 measures 82 systems on 534 public and 308 sealed decisions, with only system-level sealed aggregates published. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.

Scored 23 Sept 2026 · protocol jevbench::v1.4 · 534 public + 308 sealed aggregate decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 e67548630565 · v1.0 results

Share this version

JevBench v1.4.1 ranking

77 ranked systems and 5 unranked rows, measured on 534 public decisions plus 308 sealed decisions. The sealed text and answers remain private; only system-level aggregates appear here.

JevBench v1.4.1 · 534 public + 308 sealed decisions per system

JevBench Score: 77 ranked systems

OfficialIntelligence, Calibration, Speed and Cost, each 0–100 — equal-weight harmonic mean, with the generalization and Jev-class gates. What changed in v1.4 ↓

  1. 1Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
  2. 2JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
  3. 3Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
  4. 4Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
  6. 6djev52.2I 47 · C 55 · S 91 · K 58 · $0.026 ann.
  7. 7Jev-Omni51.3I 47 · C 64 · S 82 · K 53 · ~$0.037 est.
  8. 8metask-jev-4b47.8I 45 · C 67 · S 89 · K 55 · ~$0.033 est.
  9. 9SemIf47.7I 44 · C 67 · S 84 · K 59 · ~$0.022 est.
  10. 10Jobe Qwen3.5-4B46.9I 44 · C 66 · S 86 · K 60 · ~$0.022 est.
  11. 11local-jev Qwen3.5-4B46.8I 44 · C 73 · S 75 · K 56 · ~$0.030 est.
  12. 12system-one-openAPI45.1I 44 · C 55 · S 77 · K 65 · ~$0.015 est.
  13. 13spark-s1-4b-v644.6I 45 · C 48 · S 81 · K 58 · ~$0.025 est.
  14. 14jqv44.4I 46 · C 72 · S 75 · K 47 · ~$0.056 est.
  15. 15Qwen3-Reranker-4B43.5I 45 · C 65 · S 79 · K 49 · $0.050
  16. 16decider-35b-a3b41.2I 47 · C 65 · S 81 · K 45 · ~$0.067 est.
  17. 17Raw Qwen3 4B Instruct 2507 direct logits41.0I 46 · C 29 · S 88 · K 60 · ~$0.022 est.
  18. 18OpenSourceJev40.9I 42 · C 60 · S 82 · K 64 · ~$0.016 est.
  19. 19ZeroEntropy zerank-240.2I 42 · C 76 · S 79 · K 50 · $0.047
  20. 20decision-machine-1API39.9I 41 · C 68 · S 93 · K 54 · $0.035
Show all 82 systems (57 more ranked, 5 not ranked)
  1. 21Raw Phi-4 mini direct logits38.0I 42 · C 59 · S 89 · K 50 · ~$0.048 est.
  2. 22JEV Qwen3.5-9B Base NVFP437.7I 47 · C 68 · S 93 · K 43 · ~$0.077 est.
  3. 23OpenJev36.9I 45 · C 55 · S 83 · K 45 · ~$0.066 est.
  4. 24kev 4B36.1I 42 · C 40 · S 76 · K 62 · ~$0.019 est.
  5. 25Decision 2B35.8I 39 · C 74 · S 84 · K 63 · ~$0.018 est.
  6. 26Qwen3.5-9B Jev-like data-mix v235.2I 47 · C 61 · S 82 · K 42 · ~$0.083 est.
  7. 27GPT-6 LunaAPI35.1I 96 · C 92 · S 74 · K 37 · $0.127
  8. 28SimpleJev Qwen3.8-27BAPI34.6I 52 · C 74 · S 71 · K 39 · ~$0.104 est.
  9. 29NInfer Qwen3.8-Flash-Next mixed34.0I 50 · C 79 · S 88 · K 39 · ~$0.109 est.
  10. 30GPT-6 LunaAPI33.3I 97 · C 93 · S 73 · K 36 · $0.135
  11. 31open-alternative-jev33.2I 39 · C 59 · S 83 · K 60 · ~$0.022 est.
  12. 32jev-local32.5I 45 · C 64 · S 69 · K 43 · ~$0.077 est.
  13. 33Decision Fast32.5I 37 · C 65 · S 82 · K 76 · ~$0.0063 est.
  14. 34decider-2b30.7I 39 · C 43 · S 83 · K 61 · ~$0.020 est.
  15. 35jeff30.6I 37 · C 68 · S 63 · K 77 · ~$0.0060 est.
  16. 36Laya30.3I 36 · C 64 · S 71 · K 86 · ~$0.0029 est.
  17. 37lev-350m28.5I 35 · C 71 · S 85 · K 76 · ~$0.0063 est.
  18. 38openjev-sglangAPI27.7I 49 · C 69 · S 77 · K 36 · ~$0.131 est.
  19. 39Von27.5I 34 · C 76 · S 70 · K 78 · ~$0.0055 est.
  20. 40NInfer Qwen3.8-27B NVFP426.9I 51 · C 76 · S 80 · K 35 · ~$0.145 est.
  21. 41NInfer Qwen3.8-27B NVFP426.3I 51 · C 67 · S 80 · K 35 · ~$0.145 est.
  22. 42kev 8B25.6I 42 · C 40 · S 75 · K 44 · ~$0.073 est.
  23. 43JevOne25.5I 48 · C 78 · S 88 · K 36 · ~$0.137 est.
  24. 44SimpleJev Qwen3.6-35B-A3BAPI24.9I 46 · C 60 · S 75 · K 38 · ~$0.116 est.
  25. 45kev 0.6B24.8I 34 · C 50 · S 76 · K 76 · ~$0.0063 est.
  26. 46Raw Qwen3 8B direct logits23.7I 46 · C 24 · S 86 · K 42 · ~$0.087 est.
  27. 47system-one23.4I 44 · C 33 · S 84 · K 41 · ~$0.089 est.
  28. 48OpenDecision21.6I 32 · C 57 · S 80 · K 75 · ~$0.0066 est.
  29. 49LitJev19.5I 46 · C 77 · S 67 · K 34 · ~$0.163 est.
  30. 50openJev Verdict 1.419.0I 29 · C 72 · S 78 · K 82 · ~$0.0039 est.
  31. 51kev 0.5B18.9I 31 · C 50 · S 77 · K 76 · ~$0.0063 est.
  32. 52Bespoke Nimble 9B18.7I 46 · C 56 · S 79 · K 33 · ~$0.166 est.
  33. 53GPT-5.6 LunaAPI18.5I 93 · C 87 · S 78 · K 28 · $0.242
  34. 54openJev Verdict18.1I 30 · C 47 · S 77 · K 83 · ~$0.0037 est.
  35. 55Raw Qwen3 1.7B direct logits18.1I 33 · C 24 · S 90 · K 65 · ~$0.015 est.
  36. 56reflex-27b17.8I 46 · C 77 · S 67 · K 32 · ~$0.181 est.
  37. 57djev15.2I 72 · C 88 · S 75 · K 27 · ~$0.274 est.
  38. 58GLiNER2 large15.1I 31 · C 25 · S 62 · K 73 · ~$0.0077 est.
  39. 59OpenJev14.8I 58 · C 58 · S 76 · K 28 · ~$0.255 est.
  40. 60Qwen3.5-0.8B Decision Model14.5I 28 · C 68 · S 49 · K 76 · ~$0.0065 est.
  41. 61Gemini 3.1 Flash-LiteAPI14.3I 54 · C 59 · S 82 · K 27 · $0.264
  42. 62open-jev-deberta-v3-large12.6I 26 · C 67 · S 66 · K 74 · ~$0.0073 est.
  43. 63smalljev semantic-v912.3I 26 · C 59 · S 80 · K 58 · ~$0.025 est.
  44. 64GLiNER211.8I 27 · C 25 · S 72 · K 83 · ~$0.0037 est.
  45. 65Open-Jev 9B11.2I 44 · C 62 · S 72 · K 28 · ~$0.249 est.
  46. 66Open-Jev 2B10.0I 42 · C 55 · S 73 · K 28 · ~$0.249 est.
  47. 67GLiNER2.5 multi9.8I 23 · C 57 · S 68 · K 82 · ~$0.0039 est.
  48. 68SimpleJev7.5I 21 · C 49 · S 57 · K 68 · ~$0.011 est.
  49. 69GLiNER2.5 small7.2I 20 · C 51 · S 78 · K 82 · ~$0.0039 est.
  50. 70Raw Qwen3 0.6B direct logits7.1I 23 · C 21 · S 90 · K 74 · ~$0.0074 est.
  51. 71DeepSeek V4.1 FlashAPI4.8I 94 · C 96 · S 72 · K 17 · $0.594
  52. 72Mirror2.1I 14 · C 26 · S 71 · K 73 · ~$0.0077 est.
  53. 73Mixedbread mxbai-rerank-base-v20.4I 7 · C 84 · S 88 · K 68 · $0.012
  54. 74BAAI bge-reranker-v2-m30.2I 5 · C 84 · S 90 · K 73 · $0.0077
  55. 75Alibaba GTE Reranker ModernBERT-base0.2I 5 · C 79 · S 91 · K 70 · $0.010
  56. 76Certo v10.0I 0 · C 83 · S 94 · K 100 · ~$0.0010 est.
  57. 77Open Jev JSON Canvas0.0I 48 · C 0 · S 84 · K 46 · ~$0.065 est.
  58. classifier.dev (honorable mention)API70.8I 52 · C 72 · S 88 · K 84 · ~$0.0033 est.
  59. Qwen3.8 27B (partial run)API0.0I 40 · C 94 · S 61 · K 0 · ~$2.669 est.
  60. swanOne (partial run)I · C · S · K · ~$0.111 est.
  61. Needle 3, options as tools (partial run)I · C · S · K · ~$0.014 est.
  62. Needle 3 (partial run)I · C · S · K · ~$0.024 est.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; ~ est. = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers. Names link to each project.

Compare two systems

Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, and accuracy by family on the v1.2 hard tier and on the sealed set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 Jev (TypeSafe, closed) · Score 63.3 (#1)
  • B: JevK5 v0.2.0 Jev rebuild · Score 62.0 (#2)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs JevK5 v0.2.0. Intelligence: 53.1 vs 48.9; Calibration: 76.3 vs 74.5; Speed: 83.3 vs 91.1; Cost: 52.0 vs 59.5.50100Intelligence53.1 · 48.9Calibration76.3 · 74.5Speed83.3 · 91.1Cost52.0 · 59.5
0–100, the values in the table. A label-only system has no calibration (counted as 0).

Accuracy per tier, incl. sealed

Radar: accuracy per tier, incl. sealed, two systemsAccuracy per tier, incl. sealed, Jev 1.13.0 vs JevK5 v0.2.0. Easy: 100% vs 100%; Standard: 99% vs 96%; Judge: 95% vs 95%; Hard: 74% vs 70%; Sealed: 37% vs 33%.50100Easy100% · 100%Standard99% · 96%Judge95% · 95%Hard74% · 70%Sealed37% · 33%
Share correct per tier; Sealed = the 308 private decisions, aggregate only.

Hard tier by family (v1.2 topics)

Radar: hard tier by family (v1.2 topics), two systemsHard tier by family (v1.2 topics), Jev 1.13.0 vs JevK5 v0.2.0. Adversarial: 100% vs —; Ambiguous: 79% vs —; Judge: 79% vs —; Long policy: 61% vs —; Multi-hop: 86% vs —; Probability: 80% vs —; Routing: 100% vs —; Temporal / numeric: 27% vs —; Trade-off: 92% vs —; Trap: 100% vs —.50100Adversarial100% · Ambiguous79% · Judge79% · Long policy61% · Multi-hop86% · Probability80% · Routing100% · Temporal /numeric27% · Trade-off92% · Trap100% ·
Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out). “—” = not measured for that system, not plotted.

Sealed set by family

Radar: sealed set by family, two systemsSealed set by family, Jev 1.13.0 vs JevK5 v0.2.0. Ambiguous / abstain: 30% vs 43%; Judge: 34% vs 37%; Long policy: 28% vs 23%; Multi-hop: 45% vs 39%; Paraphrase: 64% vs 14%; Probability: 50% vs 39%; Safety judge: 38% vs 38%; Temporal / numeric: 29% vs 27%; Trade-off: 38% vs 23%; Trap / adversarial: 42% vs 58%.50100Ambiguous /abstain30% · 43%Judge34% · 37%Long policy28% · 23%Multi-hop45% · 39%Paraphrase64% · 14%Probability50% · 39%Safety judge38% · 38%Temporal /numeric29% · 27%Trade-off38% · 23%Trap /adversarial42% · 58%
Share correct within each sealed family — system-level aggregates; the items stay private.
All values as a table
SpokeA: Jev 1.13.0B: JevK5 v0.2.0
The four score axes
Intelligence53.148.9
Calibration76.374.5
Speed83.391.1
Cost52.059.5
Accuracy per tier, incl. sealed
Easy100%100%
Standard99%96%
Judge95%95%
Hard74%70%
Sealed37%33%
Hard tier by family (v1.2 topics)
Adversarial100%
Ambiguous79%
Judge79%
Long policy61%
Multi-hop86%
Probability80%
Routing100%
Temporal / numeric27%
Trade-off92%
Trap100%
Sealed set by family
Ambiguous / abstain30%43%
Judge34%37%
Long policy28%23%
Multi-hop45%39%
Paraphrase64%14%
Probability50%39%
Safety judge38%38%
Temporal / numeric29%27%
Trade-off38%23%
Trap / adversarial42%58%

What changed in v1.4

  • Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence: I = 0.8 × I_v1.3 + 0.2 × I_sealed, where I_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates.
  • Calibration blends toward the sealed-inclusive result at the approved weight: C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35).
  • The k = 1 generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points: I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half.
  • The four axes use an equal-weight harmonic mean (p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0.
  • The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.

Axes, accuracy, latency and cost

Every system with its four axes, public and sealed accuracy and the gap between them. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.

#SystemJevBench ScoreIntelligenceCalibrationSpeedCost axisPublic accuracy
534
Sealed accuracy
308
Public − sealed gapCost / 1,000p50 latencyEndpoint
1
Jev 1.13.0APIby TypeSafe AI
63.353.176.383.352.086.6%36.7%+49.9 pp$0.0400.65 sAPI
2
JevK5 v0.2.0
Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
by allebee
62.048.974.591.159.585.3%33.1%+52.2 pp~$0.022est.unknown
3
Hopper
Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
by HopitAI
59.448.079.186.858.782.3%34.1%+48.2 pp~$0.024est.0.13 sRunPod GPU
4
Winnow-12B Q8
The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
by Eldan Ring
55.648.364.882.352.985.7%33.1%+52.6 pp~$0.037est.0.23 sRunPod GPU
5
reflex 4B
The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
by kshetrajna12
54.047.570.468.059.779.2%28.2%+51.0 pp~$0.022est.1.80 sRunPod GPU
6
djev
The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
by Maisa (David Villalón) · Maisa, diffusion-gemma
52.247.055.491.457.684.0%29.9%+54.1 pp$0.026announced0.24 sAPI
7
Jev-Omni
Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by akhilaaa3 · akhilaaa3, Gemma-4-12B merged
51.346.864.181.553.088.7%32.1%+56.6 pp~$0.037est.0.22 sRunPod GPU
8
metask-jev-4b
Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
by Wayfind (metask-ai)
47.844.766.989.154.579.7%27.6%+52.1 pp~$0.033est.0.07 sRunPod GPU
9
SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ
47.744.466.883.759.581.0%26.3%+54.7 pp~$0.022est.0.20 sRunPod GPU
10
Jobe Qwen3.5-4B
No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
by MantisShrimpdev · frozen
46.944.166.185.659.581.0%25.6%+55.3 pp~$0.022est.0.13 sRunPod GPU
11
local-jev Qwen3.5-4B
Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
by Amith Chandrappa (amithgc)
46.844.473.375.055.880.5%26.0%+54.5 pp~$0.030est.0.71 sRunPod GPU
12
system-one-openAPIby mithalouni · Gemma 4 E2B LoRA on an L4
45.144.254.977.064.873.2%27.6%+45.6 pp~$0.015est.0.65 sauthor demo
13
spark-s1-4b-v6
Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085
44.645.147.681.057.979.2%26.6%+52.6 pp~$0.025est.0.31 sRunPod GPU
14
jqv
A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
by hjmurmur (Octalab) · Qwen3-32B zero-shot
44.446.471.674.647.580.1%28.2%+51.8 pp~$0.056est.0.75 sRunPod GPU
15
Qwen3-Reranker-4B
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Qwen
43.544.665.278.749.268.0%29.9%+38.1 pp$0.0500.13 sRunPod GPU
16
decider-35b-a3b
The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
by Mapika
41.247.265.380.845.383.1%31.5%+51.6 pp~$0.067est.0.29 sRunPod GPU
17
Raw Qwen3 4B Instruct 2507 direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
41.046.429.187.659.769.7%27.3%+42.4 pp~$0.022est.0.08 sRunPod GPU
18
OpenSourceJev
DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp
40.941.860.382.064.078.4%26.3%+52.1 pp~$0.016est.unknown
19
ZeroEntropy zerank-2
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by ZeroEntropy
40.242.175.879.049.870.1%28.6%+41.6 pp$0.0470.13 sRunPod GPU
20
decision-machine-1
A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
APIby milliseconds.ai (Baptiste Laget)
39.941.368.392.953.767.5%25.6%+41.9 pp$0.0350.17 sAPI
21
Raw Phi-4 mini direct logits
Neutral raw-logit control, not JevBench-directed.
by Microsoft / neutral reproduction
38.041.858.888.849.665.8%29.2%+36.6 pp~$0.048est.0.06 sRunPod GPU
22
JEV Qwen3.5-9B Base NVFP4
Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
by WilfLin
37.746.867.793.343.375.3%29.5%+45.8 pp~$0.077est.0.02 sRunPod GPU
23
OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16
36.945.455.083.245.581.8%28.6%+53.2 pp~$0.066est.0.24 sRunPod GPU
24
kev 4B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
36.142.139.675.761.866.2%22.4%+43.8 pp~$0.019est.0.55 sRunPod GPU
25
Decision 2B
Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v59
35.838.874.184.362.575.3%26.0%+49.4 pp~$0.018est.0.19 sRunPod GPU
26
Qwen3.5-9B Jev-like data-mix v2
The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
by jsaurabh
35.247.461.382.042.478.4%29.2%+49.1 pp~$0.083est.unknown
27
GPT-6 Luna
OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · low reasoning effort
35.195.892.073.736.999.1%92.9%+6.3 pp$0.1271.44 sAPI
28
SimpleJev Qwen3.8-27B
Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
34.651.674.571.239.586.6%35.7%+50.9 pp~$0.104est.1.01 sauthor demo
29
NInfer Qwen3.8-Flash-Next mixed
Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
by Igor L. / NInfer contributors
34.049.578.688.238.989.6%34.1%+55.5 pp~$0.109est.0.08 sRunPod GPU
30
GPT-6 Luna
OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · default medium reasoning effort
33.397.493.572.636.099.6%95.5%+4.1 pp$0.1351.48 sAPI
31
open-alternative-jev
With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
by IkerMoel · Qwen3.5-4B, IkerMoel
33.238.658.783.559.674.0%24.4%+49.7 pp~$0.022est.0.21 sRunPod GPU
32
jev-local
The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
by us (GitHub) · Qwen3.5-9B
32.545.264.269.243.374.9%29.5%+45.3 pp~$0.077est.1.05 sRunPod GPU
33
Decision Fast
Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v53a
32.537.165.381.676.163.2%25.6%+37.6 pp~$0.0063est.0.24 sRunPod GPU
34
decider-2b
The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
by Mapika
30.738.543.583.261.071.0%24.7%+46.3 pp~$0.020est.0.26 sRunPod GPU
35
jeff
Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
by Logan Markewich · Logan Markewich, GLiFormer 400M
30.636.867.963.576.662.8%33.1%+29.7 pp~$0.0060est.0.94 sCPU
36
Laya
The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
by Convai Innovations · Convai Innovations, ModernBERT-large 421M
30.336.163.771.186.258.4%30.8%+27.6 pp~$0.0029est.0.79 sCPU
37
lev-350m
Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M
28.534.870.685.376.158.4%25.0%+33.4 pp~$0.0063est.0.17 sRunPod GPU
38
openjev-sglangAPIby ekzhang · Qwen3.6-35B-A3B on SGLang
27.749.469.277.136.585.3%33.1%+52.2 pp~$0.131est.0.68 sauthor demo
39
Von
The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M
27.534.575.770.577.857.1%27.9%+29.2 pp~$0.0055est.unknown
40
NInfer Qwen3.8-27B NVFP4
T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
by Igor L. / NInfer contributors · T=1.5
26.951.576.080.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
41
NInfer Qwen3.8-27B NVFP4
Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
by Igor L. / NInfer contributors
26.351.567.280.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
42
kev 8B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
by Jared Palmer · research preview
25.641.840.274.944.071.4%21.8%+49.7 pp~$0.073est.0.59 sRunPod GPU
43
JevOne
Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
by Juspay
25.547.677.688.535.989.6%33.8%+55.8 pp~$0.137est.0.09 sRunPod GPU
44
SimpleJev Qwen3.6-35B-A3B
Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
24.945.759.875.038.181.4%28.2%+53.1 pp~$0.116est.0.85 sauthor demo
45
kev 0.6B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
24.834.250.075.676.166.7%24.0%+42.6 pp~$0.0063est.0.59 sRunPod GPU
46
Raw Qwen3 8B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
23.745.724.186.341.968.4%26.3%+42.1 pp~$0.087est.0.08 sRunPod GPU
47
system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke
23.443.632.884.441.571.9%24.4%+47.5 pp~$0.089est.0.17 sRunPod GPU
48
OpenDecision
A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
by Deepan Wadhwa · ModernBERT-large zero-shot
21.631.857.179.975.353.2%25.6%+27.6 pp~$0.0066est.0.34 sRunPod GPU
49
LitJev
The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
by Zhengxu Yu · Qwen3.8-27B
19.546.376.666.733.686.1%30.8%+55.3 pp~$0.163est.2.03 sRunPod GPU
50
openJev Verdict 1.4
Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
by Hemant (heman10x)
19.029.472.078.182.457.6%27.9%+29.7 pp~$0.0039est.0.31 sCPU
51
kev 0.5B
Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer
18.930.549.777.076.149.4%27.3%+22.1 pp~$0.0063est.0.43 sRunPod GPU
52
Bespoke Nimble 9B
Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
by Bespoke Labs
18.746.356.478.733.479.7%28.9%+50.8 pp~$0.166est.0.39 sRunPod GPU
53
GPT-5.6 LunaAPIby OpenAI · low reasoning effort
18.593.187.477.528.597.4%89.0%+8.4 pp$0.2420.97 sAPI
54
openJev Verdict
The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
by Hemant (heman10x) · heman10x, ModernBERT-base 151M
18.130.047.076.783.155.4%24.7%+30.7 pp~$0.0037est.0.28 sCPU
55
Raw Qwen3 1.7B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
18.133.224.389.764.954.1%26.0%+28.1 pp~$0.015est.0.07 sRunPod GPU
56
reflex-27b
The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
by kshetrajna12 · Qwen3.8-27B
17.846.477.267.532.387.0%29.5%+57.5 pp~$0.181est.1.89 sRunPod GPU
57
djev
Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
by David Villalon / Maisa · thinking
15.271.687.875.226.987.4%60.1%+27.4 pp~$0.274est.0.43 sRunPod GPU
58
GLiNER2 large
The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino
15.131.124.861.773.356.7%28.6%+28.1 pp~$0.0077est.1.10 sCPU
59
OpenJev
OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
by razorback16 · thinking, BF16
14.858.158.176.127.888.7%42.2%+46.5 pp~$0.255est.0.46 sRunPod GPU
60
Qwen3.5-0.8B Decision Model
JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
by Mourad Ghafiri
14.528.168.249.275.759.3%34.7%+24.6 pp~$0.0065est.7.15 sCPU
61
Gemini 3.1 Flash-LiteAPIby Google
14.354.559.381.827.487.0%38.6%+48.4 pp$0.2640.76 sAPI
62
open-jev-deberta-v3-large
297/308 sealed items answered validly (failures count as wrong)
by Kotoba Labs · local CPU
12.625.666.666.074.052.4%29.5%+22.8 pp~$0.0073est.1.77 sCPU
63
smalljev semantic-v9
The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
by Aditya (isHeSatoshi)
12.325.759.279.857.960.6%26.9%+33.7 pp~$0.025est.0.41 sRunPod GPU
64
GLiNER2
A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
by Fastino · Fastino, gliner2.5-base
11.827.425.271.883.158.0%29.2%+28.8 pp~$0.0037est.0.31 sCPU
65
Open-Jev 9B
The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
11.244.261.872.028.177.5%29.9%+47.6 pp~$0.249est.0.75 sRunPod GPU
66
Open-Jev 2B
The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
10.042.355.373.528.164.5%26.3%+38.2 pp~$0.249est.0.66 sRunPod GPU
67
GLiNER2.5 multi
The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
by Fastino · Fastino, 287M
9.823.157.267.882.448.9%32.8%+16.1 pp~$0.0039est.0.43 sCPU
68
SimpleJev
143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU
7.521.549.157.568.354.5%34.7%+19.8 pp~$0.011est.3.99 sCPU
69
GLiNER2.5 small
The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino · Fastino, 74M
7.220.550.777.882.445.9%28.6%+17.3 pp~$0.0039est.0.11 sCPU
70
Raw Qwen3 0.6B direct logits
Neutral raw-logit control.
by Alibaba Qwen / neutral reproduction
7.122.820.789.973.948.1%25.6%+22.4 pp~$0.0074est.0.07 sRunPod GPU
71
DeepSeek V4.1 FlashAPIby DeepSeek · thinking default
4.894.095.571.616.897.8%94.8%+3.0 pp$0.5941.42 sAPI
72
Mirror
171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
by Bluusun
2.113.626.070.873.341.1%9.1%+32.0 pp~$0.0077est.0.90 sauthor demo
73
Mixedbread mxbai-rerank-base-v2
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Mixedbread
0.46.884.187.567.937.2%34.4%+2.8 pp$0.0120.07 sRunPod GPU
74
BAAI bge-reranker-v2-m3
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by BAAI
0.25.084.289.573.439.4%27.9%+11.5 pp$0.00770.03 sRunPod GPU
75
Alibaba GTE Reranker ModernBERT-base
Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Alibaba-NLP
0.24.878.990.669.633.8%33.4%+0.3 pp$0.0100.05 sRunPod GPU
76
Certo v1
The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
by AltSlate Labs
0.00.183.094.0100.031.6%29.5%+2.1 pp~$0.0010est.0.02 sRunPod GPU
77
Open Jev JSON Canvas
Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
by JoshuaSP
0.048.20.084.145.684.4%31.2%+53.2 pp~$0.065est.0.22 sRunPod GPU
classifier.dev
Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
APIby mrmps (@michael_chomsky) · fast tier
honorable mention · not ranked
70.851.672.487.684.385.3%34.4%+50.9 pp~$0.0033est.0.39 sAPI
Qwen3.8 27BAPIby Qwen / Chutes · Chutes TEE
partial · not ranked
0.040.493.661.30.071.9%21.8%+50.1 pp~$2.669est.5.75 sAPI
swanOne
Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
by swanOne submitter
partial · not ranked
88.7%~$0.111est.0.34 sRunPod GPU
Needle 3, options as tools
V1.4: Not re-measured: same as needle-3.
by Cactus Compute · post-hoc adapter mode
partial · not ranked
22.1%~$0.014est.3.78 sCPU
Needle 3
V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.
by Cactus Compute · Cactus, 2-bit, local CPU
partial · not ranked
22.5%~$0.024est.1.69 sCPU

API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.

All 72 system notes and disclosures
  • JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
  • Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
  • Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
  • local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
  • OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
  • JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
  • kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
  • Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
  • GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
  • GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
  • NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
  • NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
  • kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
  • Raw Qwen3 8B direct logits: Neutral raw-logit control.
  • OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
  • Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
  • reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
  • GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
  • open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
  • GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
  • Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
  • swanOne: Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
  • Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
  • Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.

Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.

Artifact: v1.4.1 results JSON · SHA-256 e67548630565 · JevBench v1.4.1 release and method

What the run says

Jev alternatives, open source and self-hosting

The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are Hopper (#3, 59.4), Winnow-12B Q8 (#4, 55.6), reflex 4B (#5, 54.0), djev (#6, 52.2). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost. Version v1.4.1 measures 534 public and 308 sealed decisions: 20% of Intelligence comes from the sealed set, a public-minus-sealed gap above 25 points costs Intelligence, and Intelligence, Speed or Cost below 50 each pull the score down quadratically. What changed in v1.4 · method and tiers.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.

What a decision costs

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.

How costs are estimated

One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

  • JevK5 v0.2.0~$0.022 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Hopper~$0.024 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Winnow-12B Q8~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • reflex 4B~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Jev-Omni~$0.037 est. per 1,000 decisions: OpenRouter Gemma 3 12B input rate list price $0.05/M in, $0.0/M out (a 12B one-pass model with no generated output; the same reference the author uses in his own model card for this model) x 384 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • metask-jev-4b~$0.033 est. per 1,000 decisions: hosted 4B reference rate USD 0.04/M input, USD 0/M output over the exact measured prompt-token counts of all 534 attempts; no generated answer tokens
  • SemIf~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
  • Jobe Qwen3.5-4B~$0.022 est. per 1,000 decisions: DeepInfra Qwen3.5-4B hosted reference list price $0.03/M in, $0.0/M out (same underlying weights; one forward pass, no generated tokens) x 396 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • local-jev Qwen3.5-4B~$0.030 est. per 1,000 decisions: hosted 4B reference rate list price $0.04/M in, $0.0/M out (one forward pass over measured input tokens and no generated answer tokens) x 397 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • system-one-open~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
  • spark-s1-4b-v6~$0.025 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 505 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jqv~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-35b-a3b~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 4B Instruct 2507 direct logits~$0.022 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenSourceJev~$0.016 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Raw Phi-4 mini direct logits~$0.048 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • JEV Qwen3.5-9B Base NVFP4~$0.077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenJev~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
  • kev 4B~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Decision 2B~$0.018 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 269 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Qwen3.5-9B Jev-like data-mix v2~$0.083 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • SimpleJev Qwen3.8-27B~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-Flash-Next mixed~$0.109 est. per 1,000 decisions: OpenRouter qwen/qwen3.8-flash hosted list reference list price $0.15/M in, $0.0/M out (same underlying Flash-Next weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • open-alternative-jev~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
  • jev-local~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Decision Fast~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-2b~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jeff~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Laya~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • lev-350m~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openjev-sglang~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
  • Von~$0.0055 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • NInfer Qwen3.8-27B NVFP4~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-27B NVFP4~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 8B~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • JevOne~$0.137 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • SimpleJev Qwen3.6-35B-A3B~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 0.6B~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 8B direct logits~$0.087 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • system-one~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
  • OpenDecision~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • LitJev~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict 1.4~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • kev 0.5B~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Bespoke Nimble 9B~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Raw Qwen3 1.7B direct logits~$0.015 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • reflex-27b~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • djev~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
  • GLiNER2 large~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • OpenJev~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decision
  • Qwen3.5-0.8B Decision Model~$0.0065 est. per 1,000 decisions: same-size DeepInfra Qwen3.5-0.8B reference tariff; measured JevLite tokenizer usage; estimated, not charged
  • open-jev-deberta-v3-large~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
  • smalljev semantic-v9~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Open-Jev 9B~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open-Jev 2B~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2.5 multi~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • SimpleJev~$0.011 est. per 1,000 decisions: Same nonzero hosted size-class reference and exact prompt-token accounting; see RESULT.md
  • GLiNER2.5 small~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Raw Qwen3 0.6B direct logits~$0.0074 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Mirror~$0.0077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Certo v1~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open Jev JSON Canvas~$0.065 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • classifier.dev~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.
  • swanOne~$0.111 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Qwen3.8 27B~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
  • Needle 3, options as tools~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Needle 3~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
  • dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
  • dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
  • dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
  • encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
  • generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
  • moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
  • moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9

Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).

Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted once

usd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.

  • The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.
  • Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.
  • A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.
  • No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.

No tariff, measurement, item, answer or rank changed. The prices before and after:

  • Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)
  • SemIf — $0.0230 → $0.0224 (-2.31 %)
  • system-one-open — $0.0157 → $0.0149 (-5.13 %)
  • OpenJev — $0.0672 → $0.0656 (-2.36 %)
  • GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)
  • openjev-sglang — $0.1346 → $0.1313 (-2.48 %)
  • Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)
  • Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)
  • DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)
  • system-one — $0.0915 → $0.0894 (-2.29 %)
  • openJev Verdict — $0.0039 → $0.0037 (-5.21 %)
  • GLiNER2 — $0.0039 → $0.0037 (-5.21 %)
  • open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)
  • Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)
  • Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)
  • Needle 3 — $0.0249 → $0.0238 (-4.36 %)

Who could not be measured, and why

An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.

Method and tiers

JevBench Score. Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost (power mean p=-1). If Intelligence <50 multiply by (I/50)^2. For Speed and Cost separately, if below 50 multiply by (axis/50)^2.

Intelligence. 0.8 × v1.3 chance-corrected Intelligence on the frozen v1.2 items + 0.2 × 100 × max(0, (sealed accuracy − 0.293)/(1 − 0.293)); then multiply by 1 − max(0, public-minus-sealed accuracy gap in percentage points − 25)/100. Original tier weights: easy .14, standard .28, judge .28, hard .30.

Revision v1.4.1. v1.4.1 adds 6 systems omitted from the v1.4.0 freeze. The v1.4 scoring formulas and prior system measurements are unchanged.

  • easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
  • standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
  • judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
  • hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
  • sealed: 308 fresh private decisions across ten families, run once per system. Only system-level aggregates — overall and per-family accuracy, calibration — are published; the item text, answers and per-item results stay private and rotate between versions.

Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.

A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.

v1.4 scores are not comparable with v1.3 or earlier (sealed blend, gap penalty and harmonic mean). The v1.3.0 board stays below as history; the v1.0 page keeps its own numbers, calibration plots and per-family tables.

Limits
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.

JevBench v1.4.1 · additional views

Capability, cost and speed

Capability is the arithmetic mean of Intelligence and Calibration: (Intelligence + Calibration) / 2, on a 0–100 scale. Cost is USD per 1,000 decisions; its axis is logarithmic, and lower is better. Speed uses the JevBench Speed axis, where higher is faster. Estimated costs are marked.

3 placeholder rows have no published Intelligence or Calibration values and are omitted.

Top 20 by Capability

Capability with cost alongside

Each system has a wide Capability bar and a narrower cost bar. The cost scale is logarithmic: longer bars mean higher cost, so shorter is cheaper.

  1. 1GPT-6 LunaAPI95.4Cost $0.14
  2. 2DeepSeek V4.1 FlashAPI94.7Cost $0.59
  3. 3GPT-6 LunaAPI93.9Cost $0.13
  4. 4GPT-5.6 LunaAPI90.3Cost $0.24
  5. 5djev79.7Cost $0.27 est.
  6. 6Qwen3.8 27BAPI67.0Cost $2.67 est.
  7. 7Jev 1.13.0API64.7Cost $0.040
  8. 8NInfer Qwen3.8-Flash-Next mixed64.1Cost $0.11 est.
  9. 9NInfer Qwen3.8-27B NVFP463.7Cost $0.14 est.
  10. 10Hopper63.5Cost $0.024 est.
  11. 11SimpleJev Qwen3.8-27BAPI63.0Cost $0.10 est.
  12. 12JevOne62.6Cost $0.14 est.
  13. 13classifier.devAPI62.0Cost $0.0033 est.
  14. 14reflex-27b61.8Cost $0.18 est.
  15. 15JevK5 v0.2.061.7Cost $0.022 est.
  16. 16LitJev61.4Cost $0.16 est.
  17. 17openjev-sglangAPI59.3Cost $0.13 est.
  18. 18NInfer Qwen3.8-27B NVFP459.3Cost $0.14 est.
  19. 19jqv59.0Cost $0.056 est.
  20. 20reflex 4B58.9Cost $0.022 est.
Show all 79 systems (59 more)
  1. 21ZeroEntropy zerank-258.9Cost $0.047
  2. 22local-jev Qwen3.5-4B58.9Cost $0.030 est.
  3. 23OpenJev58.1Cost $0.25 est.
  4. 24JEV Qwen3.5-9B Base NVFP457.3Cost $0.077 est.
  5. 25Gemini 3.1 Flash-LiteAPI56.9Cost $0.26
  6. 26Winnow-12B Q856.6Cost $0.037 est.
  7. 27Decision 2B56.4Cost $0.018 est.
  8. 28decider-35b-a3b56.2Cost $0.067 est.
  9. 29metask-jev-4b55.8Cost $0.033 est.
  10. 30SemIf55.6Cost $0.022 est.
  11. 31Jev-Omni55.4Cost $0.037 est.
  12. 32Von55.1Cost $0.0055 est.
  13. 33Jobe Qwen3.5-4B55.1Cost $0.022 est.
  14. 34Qwen3-Reranker-4B54.9Cost $0.050
  15. 35decision-machine-1API54.8Cost $0.035
  16. 36jev-local54.7Cost $0.077 est.
  17. 37Qwen3.5-9B Jev-like data-mix v254.4Cost $0.083 est.
  18. 38Open-Jev 9B53.0Cost $0.25 est.
  19. 39SimpleJev Qwen3.6-35B-A3BAPI52.8Cost $0.12 est.
  20. 40lev-350m52.7Cost $0.0063 est.
  21. 41jeff52.3Cost $0.0060 est.
  22. 42Bespoke Nimble 9B51.4Cost $0.17 est.
  23. 43Decision Fast51.2Cost $0.0063 est.
  24. 44djev51.2Cost $0.026
  25. 45OpenSourceJev51.1Cost $0.016 est.
  26. 46openJev Verdict 1.450.7Cost $0.0039 est.
  27. 47Raw Phi-4 mini direct logits50.3Cost $0.048 est.
  28. 48OpenJev50.2Cost $0.066 est.
  29. 49Laya49.9Cost $0.0029 est.
  30. 50system-one-openAPI49.5Cost $0.015 est.
  31. 51Open-Jev 2B48.8Cost $0.25 est.
  32. 52open-alternative-jev48.6Cost $0.022 est.
  33. 53Qwen3.5-0.8B Decision Model48.2Cost $0.0065 est.
  34. 54spark-s1-4b-v646.3Cost $0.025 est.
  35. 55open-jev-deberta-v3-large46.1Cost $0.0073 est.
  36. 56Mixedbread mxbai-rerank-base-v245.5Cost $0.012
  37. 57BAAI bge-reranker-v2-m344.6Cost $0.0077
  38. 58OpenDecision44.4Cost $0.0066 est.
  39. 59smalljev semantic-v942.4Cost $0.025 est.
  40. 60kev 0.6B42.1Cost $0.0063 est.
  41. 61Alibaba GTE Reranker ModernBERT-base41.9Cost $0.010
  42. 62Certo v141.5Cost $0.00097 est.
  43. 63decider-2b41.0Cost $0.020 est.
  44. 64kev 8B41.0Cost $0.073 est.
  45. 65kev 4B40.9Cost $0.019 est.
  46. 66GLiNER2.5 multi40.1Cost $0.0039 est.
  47. 67kev 0.5B40.1Cost $0.0063 est.
  48. 68openJev Verdict38.5Cost $0.0037 est.
  49. 69system-one38.2Cost $0.089 est.
  50. 70Raw Qwen3 4B Instruct 2507 direct logits37.7Cost $0.022 est.
  51. 71GLiNER2.5 small35.6Cost $0.0039 est.
  52. 72SimpleJev35.3Cost $0.011 est.
  53. 73Raw Qwen3 8B direct logits34.9Cost $0.087 est.
  54. 74Raw Qwen3 1.7B direct logits28.8Cost $0.015 est.
  55. 75GLiNER2 large28.0Cost $0.0077 est.
  56. 76GLiNER226.3Cost $0.0037 est.
  57. 77Open Jev JSON Canvas24.1Cost $0.065 est.
  58. 78Raw Qwen3 0.6B direct logits21.8Cost $0.0074 est.
  59. 79Mirror19.8Cost $0.0077 est.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Cost per 1,000 decisions · log scale
Cost bars use the right-hand scale, from $0.00097 to $2.67 per 1,000 decisions. Free cost is placed at the cheapest edge; missing cost is shown as —.

Capability vs cost

The upper-left is the more attractive area: higher Capability and lower cost.

Capability vs costThe upper-left is the more attractive area: higher Capability and lower cost. Each point has a tooltip with the system and its values.020406080100$0.0010$0.010$0.10$1.00USD per 1,000 decisions · log scale · cheaper ←Capability · higher ↑GPT-6 Luna · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Jev 1.13.0 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.JevK5 v0.2.0 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.metask-jev-4b · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Bespoke Nimble 9B · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.spark-s1-4b-v6 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.smalljev semantic-v9 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.kev 0.6B · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.Raw Qwen3 1.7B direct logits · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
79 systems plotted. Hover or focus a point to read its values.

Capability vs speed

The upper-right is the more attractive area: higher Capability and higher Speed.

Capability vs speedThe upper-right is the more attractive area: higher Capability and higher Speed. Each point has a tooltip with the system and its values.020406080100020406080100Speed axis · higher is faster →Capability · higher ↑GPT-6 Luna · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Jev 1.13.0 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.JevK5 v0.2.0 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.metask-jev-4b · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Bespoke Nimble 9B · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.spark-s1-4b-v6 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.smalljev semantic-v9 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.kev 0.6B · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.Raw Qwen3 1.7B direct logits · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
79 systems plotted. Hover or focus a point to read its values.

All three at once

The 3D view plots Capability vertically, lower cost to the right, and higher Speed toward you. Sphere size follows the JevBench score. Drag to rotate; pinch or scroll to zoom. The view loads when it scrolls into view.

Scroll here to load the interactive 3D view.

Vertical: Capability · Right: cheaper · Toward you: faster

The interactive 3D view loads when this panel scrolls into view.

79 systems plotted; systems missing cost or Speed are omitted. three.js r128 is included under its MIT license.

JevBench v1.4.1 public split · input capacity and long inputs

Context length

Context length is the amount of input a model or service can accept in one request. It matters when an app sends a long conversation state, policy set, or document: a smaller window can force truncation or chunking. A larger window is a capacity ceiling, not a promise that the system will use every token well.

Across 82 v1.4.1 rows (77 ranked systems and five unranked additions), supported published values range from 512 tokens to 1,050,000 tokens; eight have no published maximum we could verify. For Jev-class rebuilds, the table records the base window and separately notes any published training truncation or missing runtime cap.

Public accuracy by actual input length

Each line is one of the 13 top-15 systems with reconciled public outcomes and input-token counts. A point's tooltip shows correct answers and the bucket size.

JevBench public accuracy across four input-length bucketsThirteen systems are plotted from fewer than 500 input tokens through 2,000 or more tokens. Bucket denominators differ by system and are available in the details table below.0%25%50%75%100%<500500–9991k–1,999≥2kActual input tokens per decisionJev 1.13.0 (TypeSafe AI) · <500 tokens: 94.0% (79/84)Jev 1.13.0 (TypeSafe AI) · 500–999 tokens: 83.3% (40/48)Jev 1.13.0 (TypeSafe AI) · 1,000–1,999 tokens: 42.9% (6/14)Jev 1.13.0 (TypeSafe AI) · >=2,000 tokens: 73.0% (27/37)Hopper · <500 tokens: 91.9% (148/161)Hopper · 500–999 tokens: 56.0% (14/25)Hopper · 1,000–1,999 tokens: 55.6% (5/9)Hopper · >=2,000 tokens: 63.9% (23/36)Winnow-12B Q8 · <500 tokens: 92.6% (151/163)Winnow-12B Q8 · 500–999 tokens: 66.7% (16/24)Winnow-12B Q8 · 1,000–1,999 tokens: 42.9% (3/7)Winnow-12B Q8 · >=2,000 tokens: 75.7% (28/37)reflex 4B (kshetrajna12) · <500 tokens: 89.5% (145/162)reflex 4B (kshetrajna12) · 500–999 tokens: 52.0% (13/25)reflex 4B (kshetrajna12) · 1,000–1,999 tokens: 50.0% (4/8)reflex 4B (kshetrajna12) · >=2,000 tokens: 58.3% (21/36)djev (Maisa, diffusion-gemma) · <500 tokens: 94.5% (154/163)djev (Maisa, diffusion-gemma) · 500–999 tokens: 50.0% (12/24)djev (Maisa, diffusion-gemma) · 1,000–1,999 tokens: 42.9% (3/7)djev (Maisa, diffusion-gemma) · >=2,000 tokens: 67.6% (25/37)Jev-Omni (akhilaaa3, Gemma-4-12B merged) · <500 tokens: 94.5% (154/163)Jev-Omni (akhilaaa3, Gemma-4-12B merged) · 500–999 tokens: 70.8% (17/24)Jev-Omni (akhilaaa3, Gemma-4-12B merged) · 1,000–1,999 tokens: 57.1% (4/7)Jev-Omni (akhilaaa3, Gemma-4-12B merged) · >=2,000 tokens: 81.1% (30/37)metask-jev-4b · <500 tokens: 92.5% (147/159)metask-jev-4b · 500–999 tokens: 52.0% (13/25)metask-jev-4b · 1,000–1,999 tokens: 50.0% (5/10)metask-jev-4b · >=2,000 tokens: 52.8% (19/36)SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · <500 tokens: 92.6% (151/163)SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · 500–999 tokens: 58.3% (14/24)SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · 1,000–1,999 tokens: 12.5% (1/8)SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · >=2,000 tokens: 58.3% (21/36)Jobe Qwen3.5-4B (frozen) · <500 tokens: 92.6% (151/163)Jobe Qwen3.5-4B (frozen) · 500–999 tokens: 58.3% (14/24)Jobe Qwen3.5-4B (frozen) · 1,000–1,999 tokens: 12.5% (1/8)Jobe Qwen3.5-4B (frozen) · >=2,000 tokens: 58.3% (21/36)local-jev Qwen3.5-4B · <500 tokens: 90.8% (148/163)local-jev Qwen3.5-4B · 500–999 tokens: 58.3% (14/24)local-jev Qwen3.5-4B · 1,000–1,999 tokens: 12.5% (1/8)local-jev Qwen3.5-4B · >=2,000 tokens: 63.9% (23/36)spark-s1-4b-v6 (Open Spark Jev, abhishek085) · <500 tokens: 92.1% (140/152)spark-s1-4b-v6 (Open Spark Jev, abhishek085) · 500–999 tokens: 56.3% (18/32)spark-s1-4b-v6 (Open Spark Jev, abhishek085) · 1,000–1,999 tokens: 30.0% (3/10)spark-s1-4b-v6 (Open Spark Jev, abhishek085) · >=2,000 tokens: 59.5% (22/37)jqv (Qwen3-32B zero-shot) · <500 tokens: 93.3% (152/163)jqv (Qwen3-32B zero-shot) · 500–999 tokens: 52.0% (13/25)jqv (Qwen3-32B zero-shot) · 1,000–1,999 tokens: 28.6% (2/7)jqv (Qwen3-32B zero-shot) · >=2,000 tokens: 50.0% (18/36)Qwen3-Reranker-4B · <500 tokens: 76.8% (53/69)Qwen3-Reranker-4B · 500–999 tokens: 80.2% (65/81)Qwen3-Reranker-4B · 1,000–1,999 tokens: 59.3% (16/27)Qwen3-Reranker-4B · >=2,000 tokens: 42.6% (23/54)
  • #1 Jev 1.13.0 (TypeSafe AI)
  • #3 Hopper
  • #4 Winnow-12B Q8
  • #5 reflex 4B (kshetrajna12)
  • #6 djev (Maisa, diffusion-gemma)
  • #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)
  • #8 metask-jev-4b
  • #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
  • #10 Jobe Qwen3.5-4B (frozen)
  • #11 local-jev Qwen3.5-4B
  • #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)
  • #14 jqv (Qwen3-32B zero-shot)
  • #15 Qwen3-Reranker-4B
Bucket counts vary because only stored usage telemetry is available. The chart describes these benchmark items; it does not show that context length alone caused a score change.
Exact correct counts and denominators by bucket
SystemMean input tokensLength n<500500–9991,000–1,999>=2,000
#1 Jev 1.13.0 (TypeSafe AI)1,057.8183/23179/84 · 94.0%40/48 · 83.3%6/14 · 42.9%27/37 · 73.0%
#3 Hopper739.2231/231148/161 · 91.9%14/25 · 56.0%5/9 · 55.6%23/36 · 63.9%
#4 Winnow-12B Q8692.3231/231151/163 · 92.6%16/24 · 66.7%3/7 · 42.9%28/37 · 75.7%
#5 reflex 4B (kshetrajna12)696.8231/231145/162 · 89.5%13/25 · 52.0%4/8 · 50.0%21/36 · 58.3%
#6 djev (Maisa, diffusion-gemma)692.3231/231154/163 · 94.5%12/24 · 50.0%3/7 · 42.9%25/37 · 67.6%
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)686.5231/231154/163 · 94.5%17/24 · 70.8%4/7 · 57.1%30/37 · 81.1%
#8 metask-jev-4b761.2230/231147/159 · 92.5%13/25 · 52.0%5/10 · 50.0%19/36 · 52.8%
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)698.6231/231151/163 · 92.6%14/24 · 58.3%1/8 · 12.5%21/36 · 58.3%
#10 Jobe Qwen3.5-4B (frozen)698.6231/231151/163 · 92.6%14/24 · 58.3%1/8 · 12.5%21/36 · 58.3%
#11 local-jev Qwen3.5-4B696.7231/231148/163 · 90.8%14/24 · 58.3%1/8 · 12.5%23/36 · 63.9%
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)794.6231/231140/152 · 92.1%18/32 · 56.3%3/10 · 30.0%22/37 · 59.5%
#14 jqv (Qwen3-32B zero-shot)662.7231/231152/163 · 93.3%13/25 · 52.0%2/7 · 28.6%18/36 · 50.0%
#15 Qwen3-Reranker-4B2,401.9231/23153/69 · 76.8%65/81 · 80.2%16/27 · 59.3%23/54 · 42.6%

Coverage: 13 of the top 15 systems are shown. JevK5 v0.2.0 is excluded because its retained per-item run is 199/231 while the published v1.4.1 public score is 197/231. SystemOne-open has matching outcomes but no per-item input-token telemetry. All 13 included systems exactly reproduce their published public accuracy; stored lengths cover 183–231 decisions per system.

Long-policy tasks show a separate stress point

Across 19 public items in the long_policy family, several systems scored well below their full public-set accuracy. The comparison uses the family label, not only the token buckets.

  • metask-jev-4b
    Long policy 6/19 (31.6%) vs 184/231 overall (79.7%): −48.1 pp.
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085)
    Long policy 7/19 (36.8%) vs 183/231 overall (79.2%): −42.4 pp.
  • system-one-open (Gemma 4 E2B LoRA on an L4)
    Long policy 6/19 (31.6%) vs 169/231 overall (73.2%): −41.6 pp.
  • Winnow-12B Q8
    Long policy 15/19 (78.9%) vs 198/231 overall (85.7%): −6.8 pp.
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged)
    Long policy 15/19 (78.9%) vs 205/231 overall (88.7%): −9.8 pp.
Show long_policy results for all 14 matched systems
SystemOverallLong policy (19 items)Change
#8 metask-jev-4b184/231 · 79.7%6/19 · 31.6%−48.1 pp
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)183/231 · 79.2%7/19 · 36.8%−42.4 pp
#12 system-one-open (Gemma 4 E2B LoRA on an L4)169/231 · 73.2%6/19 · 31.6%−41.6 pp
#14 jqv (Qwen3-32B zero-shot)185/231 · 80.1%9/19 · 47.4%−32.7 pp
#6 djev (Maisa, diffusion-gemma)194/231 · 84.0%10/19 · 52.6%−31.4 pp
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#10 Jobe Qwen3.5-4B (frozen)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#5 reflex 4B (kshetrajna12)183/231 · 79.2%10/19 · 52.6%−26.6 pp
#3 Hopper190/231 · 82.3%11/19 · 57.9%−24.4 pp
#1 Jev 1.13.0 (TypeSafe AI)200/231 · 86.6%12/19 · 63.2%−23.4 pp
#11 local-jev Qwen3.5-4B186/231 · 80.5%12/19 · 63.2%−17.4 pp
#15 Qwen3-Reranker-4B157/231 · 68.0%10/19 · 52.6%−15.3 pp
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)205/231 · 88.7%15/19 · 78.9%−9.8 pp
#4 Winnow-12B Q8198/231 · 85.7%15/19 · 78.9%−6.8 pp

For example, metask-jev-4b scored 31.6% on long_policy versus 79.7% overall (change −48.1 pp), while Winnow-12B Q8 scored 78.9% versus 85.7% (change −6.8 pp). These are descriptive public-set comparisons. Prompt wrappers and tokenizers differ by system, and the 19-item family is small, so the results do not isolate context length as the cause. Only public item results were used; sealed-set item rows were not used.

Context limits by system

82 systems · sources checked 23 Sept 2026

Sort by selecting a column heading. Source links open the primary model card, vendor documentation, or API documentation.

#1Jev 1.13.0 (TypeSafe AI)
Basis, training and serving notes

Evidence: TypeSafe's official Jev 1.13 docs specify a 64K request budget and a 32K state-plus-longest-question budget.

System repository

64,000 total request; 32,000 state + longest questionOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#2JevK5 v0.2.0
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 2,048 tokens.

Training utility defaults --max-len to 2,048 and skips longer rows. Runtime has no smaller total context cap documented; native Qwen3.5-4B window is 262,144.

System repository · 23 Sept 2026

Training configuration · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#3Hopper
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

JevBench registry identifies this as a LoRA on Qwen3.5-4B. Its public model card does not specify a shorter max_seq_len or serving truncation, so the base model's native 262,144 window is listed; adapter training length is unknown.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#4Winnow-12B Q8
Basis, training and serving notes

Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The Winnow model card documents a 65,536-position configured Q8 test profile; Gemma 4 12B base supports 262,144.

System repository · 21 Sept 2026

Base model source · 20 Jul 2026

65,536 configured/tested; base 262,144Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026API cap
#5reflex 4B (kshetrajna12)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Server refuses an input above the base model's context window rather than truncating it. The repo's 8,192 max-pack-tokens is a batching budget, not the per-request context cap.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#6djev (Maisa, diffusion-gemma)
Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Runtime setting is bounded to 1,024–32,768; DiffusionGemma base is 262,144.

System repository · 19 Sept 2026

Base model source · 15 Jul 2026

32,768 (prompt + reserved canvas)Open primary sourceSource date: 19 Sept 2026 · checked: 23 Sept 2026API cap
#7Jev-Omni (akhilaaa3, Gemma-4-12B merged)
Basis, training and serving notes

Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Gemma 4 12B card states a 256K context window.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026Trained length
#8metask-jev-4b
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Reported validation point: 4,096 tokens. This is not automatically the maximum accepted input.

Model card reports 4,096 as the validated evaluation point, not an architectural limit; the same card/config says native 262,144. Published validated evaluation point: 4,096 tokens.

System repository · 22 Sept 2026

Training configuration · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#9SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#10Jobe Qwen3.5-4B (frozen)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 4,096 tokens.

The optional adapter-training helper uses max_tokens=4,096 and rejects longer training examples. The ranked v1.4.1 entry is the frozen Qwen3.5-4B backbone, so this optional training helper does not set its inference window.

System repository · 23 Sept 2026

Training configuration · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#11local-jev Qwen3.5-4B
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

The README documents long-state shortening. Its 32,768-token context_tokens value is only an illustrative custom-model card; the effective Qwen3.5 deployment cap is not published. The listed 262,144 is the base model window.

System repository · 21 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#12system-one-open (Gemma 4 E2B LoRA on an L4)
Basis, training and serving notes

Base model: google/gemma-4-E2B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Gemma 4 E2B card/config state 131,072 (128K) positions.

Training state limit: 2,048 tokens.

Training batch token budget: 24,576 tokens.

Full training profile caps state at 2,048 tokens and total batch tokens at 24,576; this is a training profile, not an inference limit.

System repository · 17 Sept 2026

Training configuration · 17 Sept 2026

131,072Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026Trained length
#13spark-s1-4b-v6 (Open Spark Jev, abhishek085)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 2,048 tokens.

Published RLCD configs use max_len=2,048 for training; Qwen3.5-4B native inference window is 262,144.

System repository · 22 Sept 2026

Training configuration · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#14jqv (Qwen3-32B zero-shot)
Basis, training and serving notes

Base model: Qwen/Qwen3-32B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native. The config's 40,960 positions reserve output space; 131,072 requires YaRN.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#15Qwen3-Reranker-4B
Basis, training and serving notes

Base model: Qwen/Qwen3-Reranker-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official reranker card states 32K context; config has extra positions reserved for prompt/output.

Official card's stated context is 32K; do not substitute the larger config allocation because the card is explicit.

System repository · 16 Apr 2026

32,768Open primary sourceSource date: 16 Apr 2026 · checked: 23 Sept 2026Trained length
#16decider-35b-a3b (Mapika)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-35B-A3B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-35B-A3B-Base, whose native context is also 262,144.

System repository · 23 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#17Raw Qwen3 4B Instruct 2507 direct logits
Basis, training and serving notes

Base model: Qwen/Qwen3-4B-Instruct-2507. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official model card states 262,144 natively.

System repository · 17 Sept 2025

262,144Open primary sourceSource date: 17 Sept 2025 · checked: 23 Sept 2026Trained length
#18OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)
Basis, training and serving notes

Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-1.7B card specifies 32,768 context.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#19ZeroEntropy zerank-2
Basis, training and serving notes

Base model: zeroentropy/zerank-2-reranker. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official model card states 32,768 context.

System repository · 24 Jul 2026

32,768Open primary sourceSource date: 24 Jul 2026 · checked: 23 Sept 2026Trained length
#20decision-machine-1 (milliseconds.ai)
Basis, training and serving notes

Evidence: No primary public model-card or API context limit was found.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#21Raw Phi-4 mini direct logits
Basis, training and serving notes

Base model: microsoft/Phi-4-mini-instruct. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft card states 128K context; config uses long-RoPE scaling.

System repository · 10 Dec 2025

131,072Open primary sourceSource date: 10 Dec 2025 · checked: 23 Sept 2026Trained length
#22JEV Qwen3.5-9B Base NVFP4
Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#23OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)
Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: OPENJEV_MAX_MODEL_LEN defaults to 65,536 in the submitted runner; DiffusionGemma base supports 262,144.

System repository · 23 Sept 2026

Base model source · 15 Jul 2026

65,536; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#24kev 4B (research preview)
Basis, training and serving notes

Base model: Qwen/Qwen3-4B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-4B-Base is in the Qwen3 family with a 32,768-token native window.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#25Decision 2B (FlyMy.AI, v59)
Basis, training and serving notes

Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: MiniCPM5-2B card/config states 131,072 context.

System repository · 23 Sept 2026

131,072Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026Trained length
#26Qwen3.5-9B Jev-like data-mix v2
Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#27GPT-6 Luna (low reasoning effort)
Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026API cap
#28SimpleJev Qwen3.8-27B
Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The tested public Simple Jev demo API documents a 2,000-token context limit; the Qwen3.8-27B base window is 262,144.

System repository · 23 Sept 2026

Base model source · 14 Aug 2026

2,000 demo API context; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#29NInfer Qwen3.8-Flash-Next mixed
Basis, training and serving notes

Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026Trained length
#30GPT-6 Luna (default medium reasoning effort)
Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026API cap
#31open-alternative-jev (Qwen3.5-4B, IkerMoel)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#32jev-local (Qwen3.5-9B)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 18 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#33Decision Fast (FlyMy.AI, v53a)
Basis, training and serving notes

Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B-Base config specifies 32,768 positions.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#34decider-2b (Mapika)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-2B-Base, whose native context is also 262,144.

System repository · 23 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#35jeff (Logan Markewich, GLiFormer 400M)
Basis, training and serving notes

Base model: knowledgator/gliformer-large-v1. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: GLiFormer card states configured max_len=8,192.

System repository · 20 Sept 2026

8,192Open primary sourceSource date: 18 Sept 2026 · checked: 23 Sept 2026Hard limit
#36Laya (Convai Innovations, ModernBERT-large 421M)
Basis, training and serving notes

Base model: convaiinnovations/laya. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Laya model card documents a 512-token base context; head_max_len=192 is its answer-candidate budget.

System repository · 23 Sept 2026

512Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#37lev-350m (Franck Verrot, LFM2.5-350M)
Basis, training and serving notes

Base model: LiquidAI/LFM2.5-350M. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: LiquidAI model card states 32,768 context; config has a larger positional allocation.

LiquidAI card states a 32,768 context length. Config has a larger positional allocation; no larger trained/evaluated sequence is claimed.

System repository · 21 Sept 2026

32,768Open primary sourceSource date: 5 Aug 2026 · checked: 23 Sept 2026Trained length
#38openjev-sglang (Qwen3.6-35B-A3B on SGLang)
Basis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Runtime defaults max_input_tokens=32,768 and max_total_input_tokens=262,144.

System repository · 21 Sept 2026

Base model source · 24 Apr 2026

32,768 per question; 262,144 total across questionsOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#39Von (wfzyx, Option-Marker 395M)
Basis, training and serving notes

Evidence: The Von model card states an 8,192-token context for its ModernBERT-large scoring model and describes accurate premise reading to about 2,048 tokens.

System repository · 23 Sept 2026

8,192 model context; reads well to about 2,048Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Trained length
#40NInfer Qwen3.8-27B NVFP4 (T=1.5)
Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#41NInfer Qwen3.8-27B NVFP4
Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#42kev 8B (research preview)
Basis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#43JevOne
Basis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed.

JevOne is published as Qwen3.6-35B-A3B BF16 with a bidirectional option-logit mapping; no smaller serving or training sequence cap is documented.

System repository · 23 Sept 2026

Base model source · 24 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Trained length
#44SimpleJev Qwen3.6-35B-A3B
Basis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 24 Apr 2026 · checked: 23 Sept 2026Trained length
#45kev 0.6B (research preview)
Basis, training and serving notes

Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B-Base config specifies 32,768 positions.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#46Raw Qwen3 8B direct logits
Basis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#47system-one (Qwen3-8B, Sean Goedecke)
Basis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 18 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#48OpenDecision (ModernBERT-large zero-shot)
Basis, training and serving notes

Base model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Config has 8,192 positions. Card says v2.0 may not fully use the 8K window; exact trained length is not stated.

ModernBERT config allows 8,192 positions. The v2.0 card says the older zero-shot checkpoint may not fully use the long window; it does not state a smaller exact trained limit.

System repository · 21 Sept 2026

8,192Open primary sourceSource date: 16 Jan 2025 · checked: 23 Sept 2026Hard limit
#49LitJev (Qwen3.8-27B)
Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 21 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#50openJev Verdict 1.4
Basis, training and serving notes

Base model: heman10x/rlcd-modernbert-151m. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The fixed v1.4 inference path sets its context budget to 512; underlying model config is 8,192.

System repository · 20 Sept 2026

Base model source · 20 Sept 2026

512 service budget; model config 8,192Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026API cap
#51kev 0.5B
Basis, training and serving notes

Base model: Qwen/Qwen2.5-0.5B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Kev 0.5B model card specifies an 8,192-token serving allowance per branch and a 32K backbone window.

System repository · 23 Sept 2026

Base model source · 25 Sept 2024

8,192 per branch; backbone 32,768Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026API cap
#52Bespoke Nimble 9B (Bespoke Labs)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Serving code defaults NIMBLE_MAX_PROMPT_TOKENS to 8,192 and launches the backend with that request cap.

Training max_seq_len: 2,048 tokens.

The published adapter-training recipe uses --max-length=2,048; the ranked serving profile has an 8,192-token prompt cap.

System repository · 23 Sept 2026

Base model source · 2 Mar 2026

Training configuration · 23 Sept 2026

8,192 prompt cap; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#53GPT-5.6 Luna (low reasoning effort)
Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 9 Jul 2026 · checked: 23 Sept 2026API cap
#54openJev Verdict (heman10x, ModernBERT-base 151M)
Basis, training and serving notes

Base model: knowledgator/gliclass-modern-base-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Model tokenizer config sets an 8,192-token maximum.

Model config/tokenizer sets 8,192; its current separate v1.4 inference-engine cap is not documented in the available primary sources.

System repository · 20 Sept 2026

8,192Open primary sourceSource date: 12 Aug 2025 · checked: 23 Sept 2026Hard limit
#55Raw Qwen3 1.7B direct logits
Basis, training and serving notes

Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-1.7B card specifies 32,768 context.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#56reflex-27b (Qwen3.8-27B)
Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#57djev (thinking)
Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: DiffusionGemma card/config states 256K context.

The standard djev runtime caps requests at 32,768, but the separately measured full-generation thinking run has no matching runtime cap published. The listed 262,144 is the DiffusionGemma base window, not a verified cap for this run.

System repository · 19 Sept 2026

262,144Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026Trained length
#58GLiNER2 large (Fastino)
Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official docs provide chunked long-document helpers but no maximum total document size; the named DeBERTa-v3-large encoder has 512 positions.

System repository · 17 Sept 2026

Base model source · 19 Mar 2023

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#59OpenJev (thinking, BF16)
Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: OpenJev runtime defaults OPENJEV_MAX_MODEL_LEN to 65,536; the 512-token thinking allowance is generated output, not input context.

System repository · 23 Sept 2026

Base model source · 15 Jul 2026

65,536; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#60Qwen3.5-0.8B Decision Model (Mourad Ghafiri)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-0.8B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted decision model's config.json sets 262,144 positions; it is based on Qwen3.5-0.8B-Base, whose native context is 262,144.

System repository · 22 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026Hard limit
#61Gemini 3.1 Flash-Lite
Basis, training and serving notes

Evidence: Google model docs explicitly list a 1,048,576 input-token limit and 65,536 output-token limit.

1,048,576 input token limitOpen primary sourceSource date: 21 Jul 2026 · checked: 23 Sept 2026API cap
#62open-jev-deberta-v3-large (local CPU)
Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft config max_position_embeddings=512.

System repository · 22 Sept 2026

512Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026Hard limit
#63smalljev semantic-v9
Basis, training and serving notes

Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: MiniCPM5-2B card/config states 131,072 context.

System repository · 21 Sept 2026

131,072Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026Trained length
#64GLiNER2 (Fastino, gliner2.5-base)
Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official docs provide chunked long-document helpers but no maximum total document size; underlying DeBERTa-v3-base encoder has 512 positions.

System repository · 23 Sept 2026

Base model source · 19 Mar 2023

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#65Open-Jev 9B (Zefan Cai)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#66Open-Jev 2B (Zefan Cai)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Mapika model config and Qwen3.5 family card give 262,144 positions.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 23 Apr 2026 · checked: 23 Sept 2026Trained length
#67GLiNER2.5 multi (Fastino, 287M)
Basis, training and serving notes

Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published.

System repository · 20 Sept 2026

UnknownOpen primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026Unknown
#68SimpleJev (Qwen3.5-0.8B, CPU)
Basis, training and serving notes

Base model: Qwen/Qwen3.5-0.8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official Qwen3.5-0.8B card/config reports 262,144 natively; optional YaRN extends the window.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#69GLiNER2.5 small (Fastino, 74M)
Basis, training and serving notes

Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published.

System repository · 20 Sept 2026

UnknownOpen primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026Unknown
#70Raw Qwen3 0.6B direct logits
Basis, training and serving notes

Base model: Qwen/Qwen3-0.6B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B card specifies 32,768 context.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#71DeepSeek V4.1 Flash (thinking default)
Basis, training and serving notes

Evidence: DeepSeek API model/pricing docs list DeepSeek V4.1 Flash at 1M context and 384K maximum output.

1,000,000 context; 384,000 max outputOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#72Mirror
Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft config max_position_embeddings=512.

System repository

512Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026Hard limit
#73Mixedbread mxbai-rerank-base-v2
Basis, training and serving notes

Base model: mixedbread-ai/mxbai-rerank-base-v2. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official config max_position_embeddings=32,768.

System repository · 8 Apr 2026

32,768Open primary sourceSource date: 8 Apr 2026 · checked: 23 Sept 2026Hard limit
#74BAAI bge-reranker-v2-m3
Basis, training and serving notes

Base model: BAAI/bge-reranker-v2-m3. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official tokenizer limit is 8,192; config has 8,194 positions.

System repository · 24 Jun 2024

8,192Open primary sourceSource date: 24 Jun 2024 · checked: 23 Sept 2026Hard limit
#75Alibaba GTE Reranker ModernBERT-base
Basis, training and serving notes

Base model: Alibaba-NLP/gte-reranker-modernbert-base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official config max_position_embeddings=8,192.

System repository · 4 Jul 2025

8,192Open primary sourceSource date: 4 Jul 2025 · checked: 23 Sept 2026Hard limit
#76Certo v1 (AltSlate Labs)
Basis, training and serving notes

Base model: altslate/certo-decision-model. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Submitted model config/tokenizer sets 8,192.

System repository · 21 Sept 2026

8,192Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026Hard limit
#77Open Jev JSON Canvas (JoshuaSP)
Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: DiffusionGemma card/config states 256K context.

System repository · 16 Sept 2026

262,144Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026Trained length
Unrankedclassifier.dev (fast tier)
Basis, training and serving notes

Evidence: The Jev model behind the service has a documented limit, but no primary source documents a narrower or matching classifier.dev endpoint cap.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedNeedle 3 (Cactus, 2-bit, local CPU)
Basis, training and serving notes

Evidence: Official Needle 3 docs describe text input but publish no maximum context/token limit.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedNeedle 3, options as tools (post-hoc adapter mode)
Basis, training and serving notes

Evidence: Official Needle 3 docs describe tool inputs but publish no maximum context/token limit.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedQwen3.8 27B (Chutes TEE)
Basis, training and serving notes

Evidence: Chutes model catalog lists Qwen3.8-27B-TEE at 262K context.

262,144 contextOpen primary sourceSource date: 17 Aug 2026 · checked: 23 Sept 2026API cap
UnrankedswanOne
Basis, training and serving notes

Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN.

The submitted NVFP4 runner is based on Qwen3.8-Flash-Next; 262,144 is its native window. No larger runtime setting is documented for the submitted patch.

System repository

262,144Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026Trained length

“Hard limit” is an explicit model or tokenizer ceiling; “Trained length” is a published base-model or training length; “API cap” is a published service limit. These are different kinds of evidence. A base-model window does not prove that a particular hosted endpoint accepts the same length; row notes identify cases where its serving cap is unpublished. Some API docs publish a combined context window and a separate output ceiling, so the usable input can be lower when output tokens share that window. “Unknown” means no supported maximum was found.

The input-length chart uses each system's existing usage.input_tokens telemetry and public item outcomes; no new model runs were made. Source dates are listed beside each primary-source link, and every source was checked 23 Sept 2026. Training limits and inference or API caps are shown separately where published.

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

The following public-only tables and diagnostics preserve the earlier JevBench v1.3.0 view. The ranking above is the current v1.4.1 result.

JevBench v1.3.0 · 534 decisions per system

JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)

OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓

  1. 1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.040
  2. 2SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · ~$0.022 est.
  3. 3djev (Maisa, diffusion-gemma)73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.
  4. 4Winnow-12B Q871.2I 82 · C 72 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B70.3I 80 · C 75 · S 68 · K 60 · ~$0.022 est.
  6. 6jqv68.6I 79 · C 79 · S 75 · K 47 · ~$0.056 est.
  7. 7decision-machine-168.3I 62 · C 70 · S 93 · K 54 · $0.035
  8. 8decider-35b-a3b67.6I 80 · C 72 · S 81 · K 45 · ~$0.067 est.
  9. 9open-alternative-jev (Qwen3.5-4B)67.0I 64 · C 63 · S 83 · K 60 · ~$0.022 est.
  10. 10system-one-open66.6I 70 · C 57 · S 77 · K 65 · ~$0.015 est.
  11. 11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · ~$0.066 est.
  12. 12SimpleJev Qwen3.8-27B66.3I 85 · C 81 · S 71 · K 39 · ~$0.104 est.
  13. 13ZeroEntropy zerank-266.0I 63 · C 76 · S 79 · K 50 · $0.047
  14. 14GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.242
  15. 15openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · ~$0.131 est.
  16. 16Qwen3-Reranker-4B63.8I 64 · C 67 · S 79 · K 49 · $0.050
  17. 17reflex-27b63.3I 86 · C 86 · S 67 · K 32 · ~$0.181 est.
  18. 18LitJev62.7I 82 · C 84 · S 67 · K 34 · ~$0.163 est.
  19. 19kev 0.6B62.5I 52 · C 51 · S 76 · K 76 · ~$0.0063 est.
  20. 20SimpleJev Qwen3.6-35B-A3B62.5I 80 · C 67 · S 75 · K 38 · ~$0.116 est.
  21. 21djev62.4I 81 · C 93 · S 75 · K 27 · ~$0.274 est.
  22. 22jev-local61.8I 71 · C 69 · S 69 · K 43 · ~$0.077 est.
  23. 23decider-2b61.7I 61 · C 47 · S 83 · K 61 · ~$0.020 est.
  24. 24Bespoke Nimble 9B60.5I 78 · C 65 · S 79 · K 33 · ~$0.166 est.
  25. 25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.264
  26. 26OpenJev60.0I 88 · C 70 · S 76 · K 28 · ~$0.255 est.
  27. 27kev 4B59.7I 65 · C 42 · S 76 · K 62 · ~$0.019 est.
  28. 28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.594
  29. 29kev 8B56.4I 69 · C 44 · S 75 · K 44 · ~$0.073 est.
  30. 30Open-Jev 9B55.0I 71 · C 63 · S 72 · K 28 · ~$0.249 est.
  31. 31system-one54.8I 70 · C 37 · S 84 · K 41 · ~$0.089 est.
  32. 32jeff54.4I 47 · C 65 · S 63 · K 77 · ~$0.0060 est.
  33. 33Laya54.4I 46 · C 62 · S 71 · K 86 · ~$0.0029 est.
  34. 34Open-Jev 2B51.3I 61 · C 55 · S 73 · K 28 · ~$0.249 est.
  35. 35OpenDecision40.6I 41 · C 56 · S 80 · K 75 · ~$0.0066 est.
  36. 36openJev Verdict 1.438.9I 39 · C 74 · S 78 · K 82 · ~$0.0039 est.
  37. 37openJev Verdict38.1I 40 · C 51 · S 77 · K 83 · ~$0.0037 est.
  38. 38kev 0.5B33.2I 38 · C 47 · S 77 · K 76 · ~$0.0063 est.
  39. 39GLiNER2 large29.6I 40 · C 24 · S 62 · K 73 · ~$0.0077 est.
  40. 40smalljev semantic-v927.4I 35 · C 59 · S 80 · K 58 · ~$0.025 est.
  41. 41GLiNER224.0I 36 · C 24 · S 72 · K 83 · ~$0.0037 est.
  42. 42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · ~$0.0073 est.
  43. 43GLiNER2.5 multi16.6I 28 · C 56 · S 68 · K 82 · ~$0.0039 est.
  44. 44GLiNER2.5 small13.8I 26 · C 47 · S 78 · K 82 · ~$0.0039 est.
  45. 45Mixedbread mxbai-rerank-base-v20.8I 7 · C 83 · S 88 · K 68 · $0.012
  46. 46BAAI bge-reranker-v2-m30.7I 6 · C 84 · S 90 · K 73 · $0.0077
  47. 47Alibaba GTE Reranker ModernBERT-base0.3I 5 · C 77 · S 91 · K 70 · $0.010
  48. 48Certo v10.0I 0 · C 82 · S 94 · K 100 · ~$0.0010 est.
  49. classifier.dev (fast tier) (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · ~$0.0033 est.
  50. Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · ~$2.669 est.
  51. Needle 3, options as tools (partial run)1.1I 14 · C · S 53 · K 65 · ~$0.014 est.
  52. Needle 3 (partial run)0.1I 5 · C · S 60 · K 59 · ~$0.024 est.

Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)

  • Jev (TypeSafe, closed)
  • Jev rebuild (open, or open source planned)
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier (not a Jev rebuild)
  • Closed decision model (API only, not Jev)
  • Shown, not ranked — honorable mention (runs another entrant's model) · partial run
Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.
Legend and notes
  • ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
  • ann. = the provider’s announced price, not yet charged.
  • Names link to each project.
  • A label-only system has no calibration (–, counted as 0).
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Weighting: Intelligence : Calibration : Speed : Cost

Official default
Custom

The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.

Explore by task difficulty

All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.

Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.

What the run says (JevBench Score)

  • Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
  • classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
  • Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.11.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
  • GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
  • Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.

Axes, tiers, latency and cost

Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.💲 $ per 1,000 decisions, not per 1,000 tokens — one decision ≈ 950 input tokens.
Rank#SystemEndpoint
1by TypeSafe AIJev 1.13.074.485.782.783.352.0$0.040100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API
2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ73.179.072.683.759.5~$0.022 est.100.0%97.9%95.2%59.5%0.20 s raw0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU
3by Maisa (David Villalón)djevMaisa, diffusion-gemma73.082.765.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API
4by Eldan RingWinnow-12B Q871.282.072.082.352.9~$0.037 est.100.0%96.9%91.1%70.9%0.23 s raw0.60 s adjustedp95 0.41 s raw → 0.98 sour RunPod GPU
5by kshetrajna12reflex 4B70.380.175.268.059.7~$0.022 est.100.0%94.8%97.3%63.2%1.80 s raw3.75 s adjustedp95 2.05 s raw → 4.26 sour RunPod GPU
6by hjmurmur (Octalab)jqvQwen3-32B zero-shot68.679.379.074.647.5~$0.056 est.100.0%95.8%92.5%64.5%0.75 s raw1.64 s adjustedp95 0.97 s raw → 2.10 sour RunPod GPU
7by milliseconds.ai (Baptiste Laget)decision-machine-1milliseconds.ai68.362.170.492.953.7$0.035100.0%76.0%89.7%46.8%0.17 s rawp95 0.30 s rawproduction API
8by Mapikadecider-35b-a3b67.679.671.580.845.3~$0.067 est.100.0%96.9%91.1%65.5%0.29 s raw0.73 s adjustedp95 0.49 s raw → 1.14 sour RunPod GPU
9by IkerMoelopen-alternative-jevQwen3.5-4B, IkerMoel67.064.063.283.559.6~$0.022 est.100.0%84.4%74.7%56.8%0.21 s raw0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU
10by mithalounisystem-one-openGemma 4 E2B LoRA on an L466.669.556.777.064.8~$0.015 est.100.0%93.8%87.7%49.1%0.65 s raw1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server
11by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1666.479.264.883.245.5~$0.066 est.100.0%95.8%91.1%65.5%0.24 s raw0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU
12by Featherless AISimpleJev Qwen3.8-27B66.384.781.171.239.5~$0.104 est.100.0%96.9%93.2%75.0%1.01 s raw2.03 s adjustedp95 1.88 s raw → 3.76 sauthor's demo server
13by ZeroEntropyZeroEntropy zerank-266.063.076.579.049.8$0.047100.0%79.2%88.4%47.3%0.13 s raw0.40 s adjustedp95 1.50 s raw → 3.15 sour RunPod GPU
14by OpenAIGPT-5.6 Lunalow reasoning effort65.995.389.877.528.5$0.242100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API
15by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang65.383.477.477.136.5~$0.131 est.100.0%95.8%95.2%71.4%0.68 s raw1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server
16by QwenQwen3-Reranker-4B63.864.067.078.749.2$0.050100.0%79.2%87.7%50.0%0.13 s raw0.41 s adjustedp95 1.56 s raw → 3.27 sour RunPod GPU
17by kshetrajna12reflex-27bQwen3.8-27B63.385.886.267.532.3~$0.181 est.100.0%95.8%95.9%75.9%1.89 s raw3.93 s adjustedp95 2.21 s raw → 4.57 sour RunPod GPU
18by Zhengxu YuLitJevQwen3.8-27B62.782.483.566.733.6~$0.163 est.100.0%97.9%88.4%73.2%2.03 s raw4.20 s adjustedp95 2.46 s raw → 5.06 sour RunPod GPU
19by Jared Palmerkev 0.6Bresearch preview62.551.951.175.676.1~$0.0063 est.100.0%81.3%66.4%40.0%0.59 s raw1.33 s adjustedp95 0.97 s raw → 2.09 sour RunPod GPU
20by Featherless AISimpleJev Qwen3.6-35B-A3B62.579.567.175.038.1~$0.116 est.100.0%93.8%93.2%66.4%0.85 s raw1.70 s adjustedp95 0.93 s raw → 1.86 sauthor's demo server
21by David Villalon / Maisadjevthinking62.480.892.775.226.9~$0.274 est.95.8%99.0%80.1%77.7%0.43 s raw1.00 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
22by us (GitHub)jev-localQwen3.5-9B61.870.868.769.243.3~$0.077 est.100.0%84.4%89.0%59.1%1.05 s raw2.24 s adjustedp95 2.62 s raw → 5.38 sour RunPod GPU
23by Mapikadecider-2b61.761.246.683.261.0~$0.020 est.100.0%85.4%77.4%47.3%0.26 s raw0.67 s adjustedp95 0.28 s raw → 0.72 sour RunPod GPU
24by Bespoke LabsBespoke Nimble 9B60.577.965.378.733.4~$0.166 est.100.0%94.8%89.0%65.5%0.39 s raw0.93 s adjustedp95 0.65 s raw → 1.46 sour RunPod GPU
25by GoogleGemini 3.1 Flash-Lite60.185.668.181.827.4$0.264100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API
26by razorback16OpenJevthinking, BF1660.088.069.676.127.8~$0.255 est.100.0%100.0%94.5%78.2%0.46 s raw1.08 s adjustedp95 1.08 s raw → 2.31 sour RunPod GPU
27by Jared Palmerkev 4Bresearch preview59.764.842.075.761.8~$0.019 est.100.0%91.7%85.6%42.3%0.55 s raw1.25 s adjustedp95 0.99 s raw → 2.13 sour RunPod GPU
28by DeepSeekDeepSeek V4.1 Flashthinking default57.594.396.771.616.8$0.59498.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API
29by Jared Palmerkev 8Bresearch preview56.469.444.274.944.0~$0.073 est.100.0%92.7%90.4%47.3%0.59 s raw1.33 s adjustedp95 1.15 s raw → 2.45 sour RunPod GPU
30by Zefan Cai (@Zefan_Cai)Open-Jev 9BZefan Cai55.071.263.372.028.1~$0.249 est.100.0%90.6%81.5%60.9%0.75 s raw1.66 s adjustedp95 1.81 s raw → 3.77 sour RunPod GPU
31by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke54.870.336.884.441.5~$0.089 est.100.0%90.6%91.8%50.0%0.17 s raw0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU
32by Logan MarkewichjeffLogan Markewich, GLiFormer 400M54.446.964.663.576.6~$0.0060 est.100.0%76.0%61.6%37.7%0.94 s raw2.03 s adjustedp95 10.97 s raw → 22.09 sour CPU
33by Convai InnovationsLayaConvai Innovations, ModernBERT-large 421M54.445.862.571.186.2~$0.0029 est.94.4%72.9%69.2%34.1%0.79 s raw1.72 s adjustedp95 2.20 s raw → 4.54 sour CPU
34by Zefan Cai (@Zefan_Cai)Open-Jev 2BZefan Cai51.361.055.173.528.1~$0.249 est.100.0%79.2%88.4%42.7%0.66 s raw1.48 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
35by Deepan WadhwaOpenDecisionModernBERT-large zero-shot40.640.856.179.975.3~$0.0066 est.87.5%62.5%71.2%33.2%0.34 s raw0.83 s adjustedp95 0.54 s raw → 1.24 sour RunPod GPU
36by Hemant (heman10x)openJev Verdict 1.438.938.674.178.182.4~$0.0039 est.86.1%67.7%56.2%37.7%0.31 s raw0.78 s adjustedp95 0.92 s raw → 2.00 sour CPU
37by Hemant (heman10x)openJev Verdictheman10x, ModernBERT-base 151M38.139.851.376.783.1~$0.0037 est.86.1%65.6%61.0%38.2%0.28 s raw0.71 s adjustedp95 1.45 s raw → 3.04 sour CPU
38by Jared Palmerkev 0.5B33.238.247.477.076.1~$0.0063 est.95.8%52.1%71.2%30.9%0.43 s raw1.01 s adjustedp95 0.92 s raw → 1.99 sour RunPod GPU
39by FastinoGLiNER2 large29.640.124.361.773.3~$0.0077 est.98.6%62.5%61.0%36.4%1.10 s raw2.34 s adjustedp95 14.49 s raw → 29.13 sour CPU
40by Aditya (isHeSatoshi)smalljev semantic-v927.435.158.979.857.9~$0.025 est.97.2%68.8%40.4%38.2%0.41 s raw0.98 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU
41by FastinoGLiNER2Fastino, gliner2.5-base24.035.623.771.883.1~$0.0037 est.97.2%66.7%45.9%36.4%0.31 s raw0.78 s adjustedp95 4.15 s raw → 8.46 sour CPU
42by Kotoba Labsopen-jev-deberta-v3-largelocal CPU23.131.966.466.074.0~$0.0073 est.100.0%49.0%53.4%36.4%1.77 s raw3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU
43by FastinoGLiNER2.5 multiFastino, 287M16.627.756.167.882.4~$0.0039 est.90.3%51.0%43.8%37.7%0.43 s raw1.01 s adjustedp95 8.18 s raw → 16.50 sour CPU
44by FastinoGLiNER2.5 smallFastino, 74M13.825.647.277.882.4~$0.0039 est.83.3%47.9%50.0%33.2%0.11 s raw0.38 s adjustedp95 2.10 s raw → 4.35 sour CPU
45by MixedbreadMixedbread mxbai-rerank-base-v20.86.783.187.567.9$0.01244.4%33.3%26.7%40.0%0.07 s raw0.29 s adjustedp95 0.23 s raw → 0.62 sour RunPod GPU
46by BAAIBAAI bge-reranker-v2-m30.76.383.889.573.4$0.007743.1%36.5%8.9%36.8%0.03 s raw0.22 s adjustedp95 0.18 s raw → 0.51 sour RunPod GPU
47by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base0.34.676.890.669.6$0.01033.3%39.6%30.1%33.6%0.05 s raw0.25 s adjustedp95 0.10 s raw → 0.35 sour RunPod GPU
48by AltSlate LabsCerto v10.00.082.094.0100.0~$0.0010 est.27.8%30.2%21.9%31.8%0.02 s raw0.19 s adjustedp95 0.03 s raw → 0.21 sour RunPod GPU
Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
by mrmps (@michael_chomsky)classifier.devfast tierhonorable mention · not ranked83.685.177.987.684.3~$0.0033 est.100.0%99.0%97.3%70.5%0.39 s rawp95 0.45 s rawproduction API
Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.
by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked24.867.492.161.30.0~$2.669 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction API
by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked1.113.5none (label only)52.865.3~$0.014 est.66.7%31.3%34.2%3.78 s raw7.71 s adjustedp95 33.64 s raw → 67.42 sour CPU
by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked0.14.6none (label only)59.958.7~$0.024 est.47.2%16.7%31.5%7.7%1.69 s raw3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU
† Notes on 39 marked systems — how each was run
  • djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev: Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Honorable mentions — services built on another entrant's model

A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.

classifier.devno rank

83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4

Runs on Jev (TypeSafe).

  • Intelligence85.1
  • Calibration77.9
  • Speed87.6
  • Cost84.3
  • $ per 1,000 decisions~$0.0033 est.

Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Why it is not ranked, what its price assumes, and what we found

classifier.dev is not its own model. Its own pages say so: "The fast tier is Jev, TypeSafe's decision model" (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with "model": "jev-1.13.0" — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Only the fast tier was measured. The smart tier's escalation was never run, so nothing here scores it.

Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: "The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key" (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench's whole questions.

Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev's 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev's own explanation for differences of this kind is batching ("The fast tier is Jev, packed a thousand to a request"); on their own two test sets they measured the same difference as noise.

A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/about

Which public tasks did each system get right?

This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped.

Show 231 public task outcomes across 52 systems
TaskJev 1.13.0SemIfdjevWinnow-12B Q8reflex 4Bjqvdecision-machine-1decider-35b-a3bopen-alternative-jevsystem-one-openOpenJevSimpleJev Qwen3.8-27BZeroEntropy zerank-2GPT-5.6 Lunaopenjev-sglangQwen3-Reranker-4Breflex-27bLitJevkev 0.6BSimpleJev Qwen3.6-35B-A3Bdjevjev-localdecider-2bBespoke Nimble 9BGemini 3.1 Flash-LiteOpenJevkev 4BDeepSeek V4.1 Flashkev 8BOpen-Jev 9Bsystem-onejeffLayaOpen-Jev 2BOpenDecisionopenJev Verdict 1.4openJev Verdictkev 0.5BGLiNER2 largesmalljev semantic-v9GLiNER2open-jev-deberta-v3-largeGLiNER2.5 multiGLiNER2.5 smallMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-baseCerto v1classifier.devQwen3.8 27BNeedle 3, options as toolsNeedle 3
Easy · 48 of 72 public48/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4842/4842/4841/4846/4848/4847/4847/4848/4844/4841/4822/4823/4817/4812/4848/4848/4832/4823/48
××××××
××××
××××××
×××
××××
×××××
×××××
×××
××××××
×××××
×××
××
××
×××××××××
×××
××××××××
××××
××××××××
×××
×××××××
××
××××××××
××××
×××××××
×××××
××××
×××
×××
×××××
×
××××
××
××
××
××××
××
××
××××
×××
!××××
××××
××××
××××
!×××
××××
××××
××
×××
Medium (standard) · 72 of 96 public71/7271/7271/7269/7268/7269/7254/7270/7260/7267/7270/7270/7257/7270/7268/7254/7269/7271/7258/7267/7271/7260/7261/7267/7271/7272/7264/7271/7267/7265/7264/7254/7250/7255/7243/7250/7245/7235/7242/7249/7246/7231/7232/7230/7224/7226/7226/7224/7271/7271/7219/7212/72
××××××
!××××××××
××××××××××
××××××××××××××××××××
×××××××××××××××××
×××××××××××××××××××
×××××××××××
×××××××××
×××××××××
××××××
××××××××××
××××××××××××××××
××××××××××
×××××××××××
×××××××××××××××××××××××
×××××××××××
×××××××××××××××
××××
××××
××××××××××
×××××××××××××××××
×××××××××××××××××
××××××××××××
××××××××××××
××××××××××
×××××××
××××××××
××××××××
×××××××
×××××××××
××××××
××××××
×××!×××××××××××
××××××××××××
×××××
××××
×××××××××××××××××
××××××××××××××
×××××××
×××××××
××××××××××××××
××××××××××
×××××××××××
×××××××××
×××××××
×××××××
××××××
××××××××××
××××××××××
×××××××
×××××××××××××××××
××××××
××××××××××××××××××××××××××××××××
××××××××××××××××××××××××××××××××
××××××××××
××××××××××××
×××××××××××××××××××××××××××
×××××××××××××××××××××××××
×××××××
×××××××××××
××××××
×××××××××××××××
××××××××
×××××××××
×××××
×××××
×××××××××××××××
××××××××××××××××××××××
×××××××××××××××××
××××××××××××××××××××××××
×××××××××××!×
×××××××××××××
Hard · 111 of 220 public81/11168/11175/11181/11167/11168/11154/11174/11163/11154/11171/11182/11157/111107/11181/11155/11184/11180/11148/11173/11185/11165/11155/11169/11182/11185/11141/111107/11150/11166/11154/11143/11139/11146/11138/11141/11142/11133/11141/11144/11141/11142/11137/11135/11140/11142/11135/11137/11178/11147/510/017/44
×!×××××××××××!××××××·×
×××××××××××××××××××!×××××××××××××××××××××××××××·
××××××××××××××××××××××·×
××××××××××××××××××××××××·×
××××××××××××××××××·×
×××××××××××××××××××××××××××××!·
××××××××!×××××××××××××××××××·×
××××××××××××××××××××××××××××××·×
××××××××××××××·
××××××××××××××××××××××××××××××××××××××××××·×
××××××××××××××××××××××××××××××·×
××××××××××××××××××××·×
×××××××××××××××××××××××××××××××××××·
××××××××××××××××××××·×
×××××××××××××××××××!××××××××××××××××·
×××××××××××××××××××××××××·×
××××××××××××××××!××××××××××××××××××××××××·
××××××××××××××××·×
×××××××××××××!×××××××××××××·
×××××××××××××!××××××××××××××××××·
××××××××××××××××××××××××××·
×××××××××××××××××××××·×
×××××××××××××××××·×
××!×××××××××××××××××·×
××××××××××!×××××××××××××××××××××·×
××××××××××××××××××××××××××××××××××××××××××·×
×××××××××××××××××××××××××××××××××××××××××·×
××××××××××!×××××××××××××××××××××·×
×××××××××××××××××·×
×××××××××××××××××××××·×
××××××××××××××××××××××××××××××××××××××××××·×
××××××××××××××·
××××××××××××××·
××××××××××××××!××××××××××××·
××××××××××××××××××××××××××××××××××××××·×
××××××××××××××××××××·×
××××××××××××××××××××·
×××××××××××××××××××××××××·×
××××××××××××××××××××××××××××××·
××!×××××××××·×
××!××××××××××××·×
×××××××××××××××××××××××××××××·
××××××××××·
×××××××××××××××××××××××××××××××××××××××××!·
××××××××××××××××××××××××××××××××××××××××××··
×××××××××××··
×××××××××××××××××××××××××××××××××××××!··
×××××××××××××××××··
×××××××××××××××××××××××××××××··
×××××××××××××××!××××××××××××××××××××××××··
×××××××××××××××××××!××××××××××××××××××××××··
×××××××××××××××××××××××××××××××××××××××××××···
×××××××××××××××××××××××××××××××···
×××××××××××××××××××××××××···
×××××!××××××××××××···
×××···
×××××××××××···
××××××××××···
×××××××××···
×××××××××××××××××···
××××××××···
×××××××××!××××××××××××××××···
××××××××××××××××××××××···
××××××××××××××××××××××···
×××××××××××××××××××××××××××××××××××···
××××××××××××××···
××××××××××!×××××××××××××××××···
×××××××××××××××××××××××××···
×××××××××···
×××××···
××××××××××××××××××××···
!×××××××××××···
×××××××××××××××××××···
×××××××××××××···
×××××××××××···
××××××××××···
××××···
××××××××××××××××××××!×××××××××××××××××××××××···
×××××××···
××××××××××××××××!×××××××××××××××××××××···
××××××××××××××××××××××××××××××···
×××××××××××××××××···
×××···
××××××××××××××××××××××××···
×××××!××××××××××××···
××××××××××××××××××××××××···
××!××××××××××××××···
×××××××××××××××!×××××××××××××××···
×××××××···
×××××××××××···
×××××××××××××××···
××××××××××××××···
×××××××××××××···
×××××××××××××××××××××××××××××××××××××××××××××···
×××××××××××××××××××××××××××···
××××××××××××××!×××××××××××××××××···
×××××××××××××××××××××××××××××···
×××××××××××××××××××××××···
×××××××××××××××···
××××××××××××××××××××××××××···
××××××××××××××××××××××××···
×××××××···
××××···
××××××××××××××××···
××××××××××××××××××××××××···
××××××××···
×××××××××××××···
××!××××××××××××···
×××!×××××××××××××××···
×××××××××××××××××××××××××···
×××××××××···

✓ correct · × wrong · ! failed (scored wrong) · · not attempted · — no public outcome in the pinned artifact. A group row counts the public tasks it lists; the tier columns of the table above use all decisions of the tier. Every task id carries its topic (hard-opus-a-long_policy-01 is a long_policy task), and the tag after the id is the published question type: choice — pick one of a defined set of options · noul — whether a stated condition holds · score — a degree along a described dimension. Task ids are shown without their tier prefix; the tier is the group row. Task descriptions are intentionally not included; the task id, tier, topic and type are the published public metadata.

How the JevBench Score works

JevBench Score = (Intelligence × Calibration × Speed × Cost)1/4, each axis on 0–100 — the geometric mean. A weak axis pulls the score down hard: a strong axis cannot buy it back. Below 50 Intelligence the score is also multiplied by (Intelligence ÷ 50)², so a system barely better than guessing cannot rank on speed and price.

  • Intelligence — accuracy above chance: per tier, how much of the gap between guessing and all-correct a system closes (0 = guessing, 100 = all correct), weighted: hard 30 %, easy 14 %, standard 28 %, judge 28 % (220 / 72 / 96 / 146 decisions).
  • Calibration — on the hard tier: does “80 % sure” come true 80 % of the time, and does the returned distribution match the exact gold distribution on the probability items.
  • Speed — median and 95th-percentile latency, one request at a time: 0.1 s scores 100, each 10× slower costs 20 points (1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.
  • Cost — dollars per 1,000 decisions, never per 1,000 tokens: $0.001 scores 100, each 10× more expensive costs 30 points ($0.01 = 70, $0.10 = 40, $1 = 10). Models without a tariff are priced at hosted-provider prices, marked “est.” (how). 💲 $ per 1,000 decisions, not $ per 1,000 tokens. One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.
Full scoring rules
JevBench Score.
Geometric mean of Intelligence, Calibration, Speed and Cost, 25 % each. If chance-corrected Intelligence is below 50, multiply by (Intelligence / 50)^2; at or above 50 there is no penalty.
Intelligence.
Per tier: 100 x (accuracy - chance) / (1 - chance), clipped at 0. Chance is 1 / options for each item (1 / levels for score items), then averaged within the tier. Tier weights: hard 30 %, easy 14 %, standard 28 %, judge 28 %. Failed, timed-out or unparseable answers count as wrong.
Calibration.
Hard tier only, systems that return a probability distribution: mean of (a) 100 x (1 - ECE/0.5), ECE = top-label expected calibration error in 10 bins, and (b) probability fidelity = 100 x (1 - mean total-variation distance) between the returned distribution and the exact gold distribution on the 20 probability items. Label-only systems have none; it counts as 0 in the JevBench Score.
Speed.
Mean of score(p50) and score(p95) of the serial 242-decision standard+judge run; score(s) = 100 - 20 log10(s / 0.1 s), clipped to 0..100 (0.1 s = 100, 1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo. Production APIs (Jev, djev, classifier.dev, OpenAI, Google, DeepSeek, Chutes) are not adjusted.
Cost.
US dollars per 1,000 DECISIONS — not per 1,000 tokens. One decision is one whole question: its state, its rubric and its options, which is hundreds to thousands of input tokens. Pooled over all 534 v1.2 decisions; score = 100 - 30 log10(usd / 0.001), clipped to 0..100 ($0.001 = 100, $0.01 = 70, $0.10 = 40, $1 = 10). Measured = public tariff x measured tokens. est. = hosted-provider list price of the same weights or size class x tokens (for a flat-rate service, its published plan price at full use). announced = the provider's published price, not yet charged (free preview), x measured tokens.
Ranked.
Ranked: a system's own model, with every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank. A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number.
Presets.
Other views reweight the same four axes and combine them the same way (geometric mean). They are not the JevBench Score.

How the ranking moves with other weights

Rank and score under the JevBench Score and the earlier views, all combined as a geometric mean (Intelligence : Calibration : Speed : Cost). Highlighted = a different rank than the JevBench Score. Ranked systems only.

SystemJevBench Score25:25:25:25 · officialBalanced, no calibration33:0:33:33 · not the defaultEmphasis on Accuracy60:0:20:20 · not the defaultEmphasis on Speed20:0:60:20 · not the defaultEmphasis on Cost20:0:20:60 · not the default
Jev 1.13.0#1 74.4#3 71.8 (rank differs from the JevBench Score)#2 77.1 (rank differs from the JevBench Score)#4 76.2 (rank differs from the JevBench Score)#9 63.1 (rank differs from the JevBench Score)
SemIf#2 73.1#2 73.3#3 75.5 (rank differs from the JevBench Score)#2 77.3#4 67.4 (rank differs from the JevBench Score)
djev#3 73.0#1 75.8 (rank differs from the JevBench Score)#1 78.5 (rank differs from the JevBench Score)#1 81.7 (rank differs from the JevBench Score)#3 67.9
Winnow-12B Q8#4 71.2#4 71.0#4 75.2#5 75.3 (rank differs from the JevBench Score)#10 63.1 (rank differs from the JevBench Score)
reflex 4B#5 70.3#6 68.8 (rank differs from the JevBench Score)#5 73.1#17 68.4 (rank differs from the JevBench Score)#5 65.0
jqv#6 68.6#14 65.5 (rank differs from the JevBench Score)#9 70.7 (rank differs from the JevBench Score)#14 69.0 (rank differs from the JevBench Score)#14 57.6 (rank differs from the JevBench Score)
decision-machine-1#7 68.3#9 67.6 (rank differs from the JevBench Score)#22 65.4 (rank differs from the JevBench Score)#3 76.8 (rank differs from the JevBench Score)#11 61.7 (rank differs from the JevBench Score)
decider-35b-a3b#8 67.6#13 66.3 (rank differs from the JevBench Score)#8 71.3#10 71.7 (rank differs from the JevBench Score)#18 56.9 (rank differs from the JevBench Score)
open-alternative-jev#9 67.0#7 68.3 (rank differs from the JevBench Score)#17 66.6 (rank differs from the JevBench Score)#6 74.0 (rank differs from the JevBench Score)#8 64.7 (rank differs from the JevBench Score)
system-one-open#10 66.6#5 70.3 (rank differs from the JevBench Score)#11 70.0 (rank differs from the JevBench Score)#9 72.9 (rank differs from the JevBench Score)#2 68.0 (rank differs from the JevBench Score)
OpenJev#11 66.4#11 66.9#7 71.6 (rank differs from the JevBench Score)#8 73.0 (rank differs from the JevBench Score)#15 57.3 (rank differs from the JevBench Score)
SimpleJev Qwen3.8-27B#12 66.3#18 62.0 (rank differs from the JevBench Score)#10 70.2 (rank differs from the JevBench Score)#24 65.5 (rank differs from the JevBench Score)#22 51.8 (rank differs from the JevBench Score)
ZeroEntropy zerank-2#13 66.0#15 62.8 (rank differs from the JevBench Score)#29 62.9 (rank differs from the JevBench Score)#15 68.8 (rank differs from the JevBench Score)#16 57.2 (rank differs from the JevBench Score)
GPT-5.6 Luna#14 65.9#23 59.5 (rank differs from the JevBench Score)#6 71.8 (rank differs from the JevBench Score)#23 66.1 (rank differs from the JevBench Score)#30 44.3 (rank differs from the JevBench Score)
openjev-sglang#15 65.3#19 61.7 (rank differs from the JevBench Score)#12 69.6 (rank differs from the JevBench Score)#18 67.4 (rank differs from the JevBench Score)#24 50.0 (rank differs from the JevBench Score)
Qwen3-Reranker-4B#16 63.8#16 62.8#27 63.3 (rank differs from the JevBench Score)#16 68.7#17 56.9 (rank differs from the JevBench Score)
reflex-27b#17 63.3#26 57.2 (rank differs from the JevBench Score)#16 67.2 (rank differs from the JevBench Score)#28 61.1 (rank differs from the JevBench Score)#27 45.5 (rank differs from the JevBench Score)
LitJev#18 62.7#28 57.0 (rank differs from the JevBench Score)#19 66.0 (rank differs from the JevBench Score)#29 60.7 (rank differs from the JevBench Score)#26 46.1 (rank differs from the JevBench Score)
kev 0.6B#19 62.5#12 66.8 (rank differs from the JevBench Score)#30 60.4 (rank differs from the JevBench Score)#13 70.2 (rank differs from the JevBench Score)#1 70.4 (rank differs from the JevBench Score)
SimpleJev Qwen3.6-35B-A3B#20 62.5#21 61.0 (rank differs from the JevBench Score)#14 67.8 (rank differs from the JevBench Score)#21 66.3 (rank differs from the JevBench Score)#23 50.6 (rank differs from the JevBench Score)
djev#21 62.4#30 54.6 (rank differs from the JevBench Score)#25 63.9 (rank differs from the JevBench Score)#27 62.1 (rank differs from the JevBench Score)#34 41.1 (rank differs from the JevBench Score)
jev-local#22 61.8#22 59.6#26 63.9 (rank differs from the JevBench Score)#26 63.3 (rank differs from the JevBench Score)#21 52.5 (rank differs from the JevBench Score)
decider-2b#23 61.7#8 67.7 (rank differs from the JevBench Score)#23 65.1#7 73.5 (rank differs from the JevBench Score)#7 64.9 (rank differs from the JevBench Score)
Bespoke Nimble 9B#24 60.5#24 58.9#20 65.9 (rank differs from the JevBench Score)#22 66.2 (rank differs from the JevBench Score)#25 47.0 (rank differs from the JevBench Score)
Gemini 3.1 Flash-Lite#25 60.1#25 57.6#15 67.5 (rank differs from the JevBench Score)#20 66.3 (rank differs from the JevBench Score)#32 42.8 (rank differs from the JevBench Score)
OpenJev#26 60.0#27 57.1 (rank differs from the JevBench Score)#13 67.9 (rank differs from the JevBench Score)#25 64.0 (rank differs from the JevBench Score)#31 42.8 (rank differs from the JevBench Score)
kev 4B#27 59.7#10 67.2 (rank differs from the JevBench Score)#18 66.2 (rank differs from the JevBench Score)#12 70.5 (rank differs from the JevBench Score)#6 65.0 (rank differs from the JevBench Score)
DeepSeek V4.1 Flash#28 57.5#34 48.4 (rank differs from the JevBench Score)#28 63.2#33 56.6 (rank differs from the JevBench Score)#40 31.7 (rank differs from the JevBench Score)
kev 8B#29 56.4#20 61.2 (rank differs from the JevBench Score)#24 64.3 (rank differs from the JevBench Score)#19 66.3 (rank differs from the JevBench Score)#19 53.6 (rank differs from the JevBench Score)
Open-Jev 9B#30 55.0#32 52.4 (rank differs from the JevBench Score)#31 59.3 (rank differs from the JevBench Score)#30 59.5#35 40.9 (rank differs from the JevBench Score)
system-one#31 54.8#17 62.6 (rank differs from the JevBench Score)#21 65.6 (rank differs from the JevBench Score)#11 70.6 (rank differs from the JevBench Score)#20 53.1 (rank differs from the JevBench Score)
jeff#32 54.4#31 53.6 (rank differs from the JevBench Score)#33 48.2 (rank differs from the JevBench Score)#34 54.5 (rank differs from the JevBench Score)#13 58.7 (rank differs from the JevBench Score)
Laya#33 54.4#29 55.0 (rank differs from the JevBench Score)#34 47.7 (rank differs from the JevBench Score)#32 56.8 (rank differs from the JevBench Score)#12 61.4 (rank differs from the JevBench Score)
Open-Jev 2B#34 51.3#33 50.1 (rank differs from the JevBench Score)#32 54.2 (rank differs from the JevBench Score)#31 58.4 (rank differs from the JevBench Score)#37 39.8 (rank differs from the JevBench Score)
OpenDecision#35 40.6#35 41.7#35 35.2#35 46.0#28 44.9 (rank differs from the JevBench Score)
openJev Verdict 1.4#36 38.9#37 37.4 (rank differs from the JevBench Score)#38 30.7 (rank differs from the JevBench Score)#37 40.8 (rank differs from the JevBench Score)#33 41.6 (rank differs from the JevBench Score)
openJev Verdict#37 38.1#36 40.1 (rank differs from the JevBench Score)#36 33.3 (rank differs from the JevBench Score)#36 43.3 (rank differs from the JevBench Score)#29 44.7 (rank differs from the JevBench Score)
kev 0.5B#38 33.2#39 35.4 (rank differs from the JevBench Score)#39 29.4 (rank differs from the JevBench Score)#38 38.9#38 38.7
GLiNER2 large#39 29.6#38 36.5 (rank differs from the JevBench Score)#37 31.8 (rank differs from the JevBench Score)#39 37.8#36 40.5 (rank differs from the JevBench Score)
smalljev semantic-v9#40 27.4#41 26.9 (rank differs from the JevBench Score)#41 22.6 (rank differs from the JevBench Score)#41 31.4 (rank differs from the JevBench Score)#41 27.6 (rank differs from the JevBench Score)
GLiNER2#41 24.0#40 30.3 (rank differs from the JevBench Score)#40 24.6 (rank differs from the JevBench Score)#40 32.6 (rank differs from the JevBench Score)#39 34.6 (rank differs from the JevBench Score)
open-jev-deberta-v3-large#42 23.1#42 21.9#42 17.8#42 23.7#42 24.9
GLiNER2.5 multi#43 16.6#43 16.4#43 12.6#43 18.0#43 19.5
GLiNER2.5 small#44 13.8#44 14.4#44 10.6#44 16.5#44 16.9
Mixedbread mxbai-rerank-base-v2#45 0.8#45 0.6#45 0.3#45 0.9#45 0.8
BAAI bge-reranker-v2-m3#46 0.7#46 0.5#46 0.3#46 0.8#46 0.7
Alibaba GTE Reranker ModernBERT-base#47 0.3#47 0.3#47 0.1#47 0.4#47 0.4
Certo v1#48 0.0#48 0.0#48 0.0#48 0.0#48 0.0

Compare two systems

Pick any two. The first radar shows the four axes of the JevBench Score (0–100, the values in the table above); the second shows accuracy by subject topic, over all tiers. Further out is better on every spoke.

  • A: Jev 1.13.0 Jev · JevBench Score 74.4 (#1)
  • B: SemIf Jev rebuild · JevBench Score 73.1 (#2)

The four score axes

Radar: the four JevBench Score axes, two systemsJev 1.13.0 vs SemIf. Intelligence: 85.7 vs 79.0; Calibration: 82.7 vs 72.6; Speed: 83.3 vs 83.7; Cost: 52.0 vs 59.5.50100Intelligence85.7 · 79.0Calibration82.7 · 72.6Speed83.3 · 83.7Cost52.0 · 59.5
Speed includes the latency adjustment for self-hosted and demo endpoints — an assumption, see Limits. A label-only system has no calibration (counted as 0).
Values as a table
AxisA: Jev 1.13.0B: SemIf
Intelligence85.779.0
Calibration82.772.6
Speed83.383.7
Cost52.059.5
JevBench Score74.473.1

Accuracy by subject topic — not part of the score

Radar: accuracy by subject topic, two systemsAccuracy by subject topic, Jev 1.13.0 vs SemIf. Math & numbers (129 items): 87.6% vs 79.1%; Coding & software (56 items): 83.9% vs 96.4%; Rules, policy & law (67 items): 83.6% vs 64.2%; Finance & commerce (64 items): 73.4% vs 60.9%; Support & operations (119 items): 89.1% vs 87.4%; Everyday language (79 items): 100.0% vs 100.0%; Safety & security (20 items): 100.0% vs 75.0%.50100Math87.6% · 79.1%Coding83.9% · 96.4%Rules & law83.6% · 64.2%Finance73.4% · 60.9%Support & ops89.1% · 87.4%Everydaylanguage100.0% · 100.0%Safety &security100.0% · 75.0%
Share of each topic's decisions answered correctly, all tiers together — compare the two systems within a topic, not topics with each other.
Values and notes
  • Topics mix tiers differently — Everyday language is mostly easy items, Rules & law and Finance mostly hard ones — which is why topics are not compared with each other.
  • Topics: one per item, drafted by a model and checked by hand — method. Held-out items count in the totals; their texts stay private.
Topic (items)A: Jev 1.13.0B: SemIf
Math & numbers (129)a calculation decides the answer: arithmetic, word problems, probability, dates, units87.6% 113 of 12979.1% 102 of 129
Coding & software (56)code, SQL, repositories, developer tools and IT systems83.9% 47 of 5696.4% 54 of 56
Rules, policy & law (67)applying written rules: company policies, contracts, regulations, eligibility83.6% 56 of 6764.2% 43 of 67
Finance & commerce (64)money: payments, refunds, invoices, orders, expenses, insurance payouts73.4% 47 of 6460.9% 39 of 64
Support & operations (119)support tickets, incidents, logistics, scheduling desks and routing work to a team89.1% 106 of 11987.4% 104 of 119
Everyday language (79)short everyday messages: intents, assistant requests, reading a detail out of a text100.0% 79 of 79100.0% 79 of 79
Safety & security (20)untrusted or injected instructions, fraud, moderation, access and security triage100.0% 20 of 2075.0% 15 of 20
Held-out hard-tier detail

With about 110 items on each side, ordinary noise is roughly ±9 percentage points. Read a system's public-minus-held-out gap against the field mean (-0.7 points across 49 complete systems): only an outlier against that field is meaningful. “Not public” does not mean “not seen”, because held-out items were sent to hosted APIs.

SystemHard publicHard held-outPublic − held-out gap (95% interval)Field mean gap
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)61.3% (68/111)57.8% (63/109)+3.5 points [-9.5, +16.4]-0.7 points
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)64.0% (71/111)67.0% (73/109)-3.0 points [-15.6, +9.6]-0.7 points
system-one (Qwen3-8B, Sean Goedecke)48.6% (54/111)51.4% (56/109)-2.7 points [-15.9, +10.5]-0.7 points
Jev 1.13.0 (TypeSafe AI)73.0% (81/111)75.2% (82/109)-2.3 points [-13.8, +9.3]-0.7 points
system-one-open (Gemma 4 E2B LoRA on an L4)48.6% (54/111)49.5% (54/109)-0.9 points [-14.1, +12.3]-0.7 points
Bespoke Nimble 9B (Bespoke Labs)62.2% (69/111)68.8% (75/109)-6.6 points [-19.2, +5.9]-0.7 points
openjev-sglang (Qwen3.6-35B-A3B on SGLang)73.0% (81/111)69.7% (76/109)+3.2 points [-8.7, +15.2]-0.7 points
GPT-5.6 Luna (low reasoning effort)96.4% (107/111)92.7% (101/109)+3.7 points [-2.3, +9.7]-0.7 points
Gemini 3.1 Flash-Lite73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 points
open-jev-deberta-v3-large (local CPU)37.8% (42/111)34.9% (38/109)+3.0 points [-9.7, +15.7]-0.7 points
DeepSeek V4.1 Flash (thinking default)96.4% (107/111)93.6% (102/109)+2.8 points [-2.9, +8.6]-0.7 points
Needle 3 (Cactus, 2-bit, local CPU) (partial)38.6% (17/44) (0/0) points [, ]-0.7 points
Qwen3.8 27B (Chutes TEE) (partial)92.2% (47/51) (0/0) points [, ]-0.7 points
Needle 3, options as tools (post-hoc adapter mode) (partial) (0/0) (0/0) points [, ]-0.7 points
open-alternative-jev (Qwen3.5-4B, IkerMoel)56.8% (63/111)56.9% (62/109)-0.1 points [-13.2, +13.0]-0.7 points
BAAI bge-reranker-v2-m337.8% (42/111)35.8% (39/109)+2.1 points [-10.7, +14.8]-0.7 points
Certo v1 (AltSlate Labs)33.3% (37/111)30.3% (33/109)+3.1 points [-9.2, +15.4]-0.7 points
classifier.dev (fast tier)70.3% (78/111)70.6% (77/109)-0.4 points [-12.4, +11.7]-0.7 points
decider-2b (Mapika)49.5% (55/111)45.0% (49/109)+4.6 points [-8.6, +17.8]-0.7 points
decider-35b-a3b (Mapika)66.7% (74/111)64.2% (70/109)+2.4 points [-10.1, +15.0]-0.7 points
decision-machine-1 (milliseconds.ai)48.6% (54/111)45.0% (49/109)+3.7 points [-9.5, +16.9]-0.7 points
djev (thinking)76.6% (85/111)78.9% (86/109)-2.3 points [-13.3, +8.7]-0.7 points
djev (Maisa, diffusion-gemma)67.6% (75/111)71.6% (78/109)-4.0 points [-16.1, +8.2]-0.7 points
GLiNER2 large (Fastino)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 points
GLiNER2.5 multi (Fastino, 287M)33.3% (37/111)42.2% (46/109)-8.9 points [-21.6, +3.9]-0.7 points
GLiNER2.5 small (Fastino, 74M)31.5% (35/111)34.9% (38/109)-3.3 points [-15.8, +9.1]-0.7 points
GLiNER2 (Fastino, gliner2.5-base)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 points
Alibaba GTE Reranker ModernBERT-base31.5% (35/111)35.8% (39/109)-4.2 points [-16.7, +8.2]-0.7 points
jeff (Logan Markewich, GLiFormer 400M)38.7% (43/111)36.7% (40/109)+2.0 points [-10.8, +14.8]-0.7 points
jev-local (Qwen3.5-9B)58.6% (65/111)59.6% (65/109)-1.1 points [-14.1, +11.9]-0.7 points
jqv (Qwen3-32B zero-shot)61.3% (68/111)67.9% (74/109)-6.6 points [-19.2, +6.0]-0.7 points
kev 0.5B29.7% (33/111)32.1% (35/109)-2.4 points [-14.6, +9.8]-0.7 points
kev 0.6B (research preview)43.2% (48/111)36.7% (40/109)+6.5 points [-6.4, +19.5]-0.7 points
kev 4B (research preview)36.9% (41/111)47.7% (52/109)-10.8 points [-23.8, +2.2]-0.7 points
kev 8B (research preview)45.0% (50/111)49.5% (54/109)-4.5 points [-17.7, +8.7]-0.7 points
Laya (Convai Innovations, ModernBERT-large 421M)35.1% (39/111)33.0% (36/109)+2.1 points [-10.4, +14.6]-0.7 points
LitJev (Qwen3.8-27B)72.1% (80/111)74.3% (81/109)-2.2 points [-13.9, +9.5]-0.7 points
Mixedbread mxbai-rerank-base-v236.0% (40/111)44.0% (48/109)-8.0 points [-20.9, +4.9]-0.7 points
Open-Jev 2B (Zefan Cai)41.4% (46/111)44.0% (48/109)-2.6 points [-15.7, +10.5]-0.7 points
Open-Jev 9B (Zefan Cai)59.5% (66/111)62.4% (68/109)-2.9 points [-15.8, +10.0]-0.7 points
OpenDecision (ModernBERT-large zero-shot)34.2% (38/111)32.1% (35/109)+2.1 points [-10.3, +14.6]-0.7 points
OpenJev (thinking, BF16)76.6% (85/111)79.8% (87/109)-3.2 points [-14.1, +7.7]-0.7 points
openJev Verdict 1.436.9% (41/111)38.5% (42/109)-1.6 points [-14.4, +11.2]-0.7 points
openJev Verdict (heman10x, ModernBERT-base 151M)37.8% (42/111)38.5% (42/109)-0.7 points [-13.5, +12.1]-0.7 points
Qwen3-Reranker-4B49.5% (55/111)50.5% (55/109)-0.9 points [-14.1, +12.3]-0.7 points
reflex-27b (Qwen3.8-27B)75.7% (84/111)76.1% (83/109)-0.5 points [-11.8, +10.8]-0.7 points
reflex 4B (kshetrajna12)60.4% (67/111)66.1% (72/109)-5.7 points [-18.4, +7.0]-0.7 points
SimpleJev Qwen3.6-35B-A3B65.8% (73/111)67.0% (73/109)-1.2 points [-13.7, +11.3]-0.7 points
SimpleJev Qwen3.8-27B73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 points
smalljev semantic-v939.6% (44/111)36.7% (40/109)+2.9 points [-9.9, +15.8]-0.7 points
Winnow-12B Q873.0% (81/111)68.8% (75/109)+4.2 points [-7.8, +16.2]-0.7 points
ZeroEntropy zerank-251.4% (57/111)43.1% (47/109)+8.2 points [-4.9, +21.4]-0.7 points

Accuracy is correct / attempted; invalid responses count as incorrect. The interval is the unpooled two-sample normal 95% interval for a difference in proportions. Partial systems are shown but excluded from the field mean.

Public-split policy. Training on JevBench's public split is allowed and should be declared with each submission. Rankings continue to use all benchmark items. We report held-out results separately so that specialisation on public tasks is visible. Held-out means not publicly released, not guaranteed unseen: hosted systems receive these tasks during evaluation. We periodically issue fresh tasks to reduce the value of prior exposure.