Frozen JevBench release v1.4.2

JevBench v1.4.2 — frozen results

This page always uses the public, hash-checked v1.4.2 artifact. The live board may change when a later release is published.

Artifact SHA-256 ac14e206dde51ae28e40dc1ea2ff1fecc4a449b941d098e9ecb5618bd533e5be.

534 public + 308 sealed decisions · only system-level sealed aggregates are published.

Share this version · View live board

Frozen top five

  1. decider-4b v2 — 64.1
  2. Jev 1.13.0 — 63.3
  3. JevK5 v0.2.0 — 62.0
  4. Cygnet — 61.8
  5. Hopper — 59.4

These ranks and scores come from the frozen v1.4.2 release and do not follow changes to the live board.

JevBench v1.4.2 ranking

JevBench v1.4.2

JevBench Score: 89 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · What changed in v1.4 ↓

Jev 1.13.0 out-reasons decider-4b v2 (Intelligence 53.1 vs 49.4) and is better calibrated; decider-4b v2 leads on speed and cost. JevBench weighs the four axes equally; sort by Intelligence for raw reasoning.

Sort by Intelligence ↓
By IntelligenceSystemIntelligenceCalibrationJevBench ScoreOverall rank
1GPT-6 LunaAPI97.493.533.3#35
2GPT-6 LunaAPI95.892.035.1#31
3DeepSeek V4.1 FlashAPI94.095.54.8#83
4GPT-5.6 LunaAPI93.187.418.5#61
5djev71.687.815.2#66
6Kushal Patil — Gemma 4 31B ITAPI59.879.230.0#42
7OpenJev58.158.114.8#68
8Gemini 3.1 Flash-LiteAPI54.559.314.3#70
9Jev 1.13.0API53.176.363.3#2
10swanOne52.671.033.6#34
11SimpleJev Qwen3.8-27BAPI51.674.534.6#32
12NInfer Qwen3.8-27B NVFP451.576.026.9#47
13NInfer Qwen3.8-27B NVFP451.567.226.3#48
14InstinctAPI51.178.711.4#74
15NInfer Qwen3.8-Flash-Next mixed49.578.634.0#33
16Cygnet49.574.961.8#4
17openjev-sglangAPI49.469.227.7#45
18decider-4b v249.475.064.1#1
19JevK5 v0.2.048.974.562.0#3
20Winnow-12B Q848.364.855.6#6
21Open Jev JSON Canvas48.20.00.0#89
22Hopper48.079.159.4#5
23JevOne47.677.625.5#50
24reflex 4B47.570.454.0#7
25Qwen3.5-9B Jev-like data-mix v247.461.335.2#30
26decider-35b-a3b47.265.341.2#19
27djev47.055.452.2#8
28JEV Qwen3.5-9B Base NVFP446.867.737.7#26
29Jev-Omni46.864.151.3#9
30jqv46.471.644.4#17
31Raw Qwen3 4B Instruct 2507 direct logits46.429.141.0#20
32reflex-27b46.477.217.8#64
33Bespoke Nimble 9B46.356.418.7#60
34LitJev46.376.619.5#57
35Standard One 8B46.265.929.1#43
36SimpleJev Qwen3.6-35B-A3BAPI45.759.824.9#52
37Raw Qwen3 8B direct logits45.724.123.7#54
38OpenJev45.455.036.9#27
39jev-local45.264.232.5#37
40spark-s1-4b-v645.147.644.6#15
41metask-jev-4b44.766.947.8#10
42Qwen3-Reranker-4B44.665.243.5#18
43SemIf44.466.847.7#11
44local-jev Qwen3.5-4B44.473.346.8#13
45system-one-openAPI44.254.945.1#14
46Open-Jev 9B44.261.811.2#75
47Jobe Qwen3.5-4B44.166.146.9#12
48system-one43.632.823.4#55
49Malkuth-4B43.161.544.5#16
50Open-Jev 2B42.355.310.0#76
51kev 4B42.139.636.1#28
52ZeroEntropy zerank-242.175.840.2#22
53kev 8B41.840.225.6#49
54Raw Phi-4 mini direct logits41.858.838.0#25
55OpenSourceJev41.860.340.9#21
56Malkuth-2B41.353.538.9#24
57decision-machine-1API41.368.339.9#23
58Decision 2B38.874.135.8#29
59open-alternative-jev38.658.733.2#36
60decider-2b38.543.530.7#39
61Decision Fast37.165.332.5#38
62jeff36.867.930.6#40
63Laya36.163.730.3#41
64lev-350m34.870.628.5#44
65Von34.575.727.5#46
66kev 0.6B34.250.024.8#53
67typecastlm34.062.125.3#51
68Raw Qwen3 1.7B direct logits33.224.318.1#63
69OpenDecision31.857.121.6#56
70GLiNER2 large31.124.815.1#67
71kev 0.5B30.549.718.9#59
72openJev Verdict30.047.018.1#62
73openJev Verdict 1.429.472.019.0#58
74JevActAPI29.355.116.9#65
75Qwen3.5-0.8B Decision Model28.168.214.5#69
76GLiNER227.425.211.8#73
77smalljev semantic-v925.759.212.3#72
78open-jev-deberta-v3-large25.666.612.6#71
79GLiNER2.5 multi23.157.29.8#77
80Raw Qwen3 0.6B direct logits22.820.77.1#81
81CLM-8B22.439.88.6#78
82SimpleJev21.549.17.5#79
83GLiNER2.5 small20.550.77.2#80
84verdict-small18.166.95.7#82
85Mirror13.626.02.1#84
86Mixedbread mxbai-rerank-base-v26.884.10.4#85
87BAAI bge-reranker-v2-m35.084.20.2#86
88Alibaba GTE Reranker ModernBERT-base4.878.90.2#87
89Certo v10.183.00.0#88
  1. 1decider-4b v264.1I 49 · C 75 · S 93 · K 61 · ~$0.020 est.
  2. 2Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
  3. 3JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
  4. 4Cygnet61.8I 50 · C 75 · S 91 · K 53 · ~$0.037 est.
  5. 5Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
  6. 6Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
  7. 7reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
  8. 8djev52.2I 47 · C 55 · S 91 · K 58 · $0.026 ann.
  9. 9Jev-Omni51.3I 47 · C 64 · S 82 · K 53 · ~$0.037 est.
  10. 10metask-jev-4b47.8I 45 · C 67 · S 89 · K 55 · ~$0.033 est.
  11. 11SemIf47.7I 44 · C 67 · S 84 · K 59 · ~$0.022 est.
  12. 12Jobe Qwen3.5-4B46.9I 44 · C 66 · S 86 · K 60 · ~$0.022 est.
  13. 13local-jev Qwen3.5-4B46.8I 44 · C 73 · S 75 · K 56 · ~$0.030 est.
  14. 14system-one-openAPI45.1I 44 · C 55 · S 77 · K 65 · ~$0.015 est.
  15. 15spark-s1-4b-v644.6I 45 · C 48 · S 81 · K 58 · ~$0.025 est.
  16. 16Malkuth-4B44.5I 43 · C 61 · S 88 · K 62 · ~$0.019 est.
  17. 17jqv44.4I 46 · C 72 · S 75 · K 47 · ~$0.056 est.
  18. 18Qwen3-Reranker-4B43.5I 45 · C 65 · S 79 · K 49 · $0.050
  19. 19decider-35b-a3b41.2I 47 · C 65 · S 81 · K 45 · ~$0.067 est.
  20. 20Raw Qwen3 4B Instruct 2507 direct logits41.0I 46 · C 29 · S 88 · K 60 · ~$0.022 est.
Show all 93 systems (69 more ranked, 4 not ranked)
  1. 21OpenSourceJev40.9I 42 · C 60 · S 82 · K 64 · ~$0.016 est.
  2. 22ZeroEntropy zerank-240.2I 42 · C 76 · S 79 · K 50 · $0.047
  3. 23decision-machine-1API39.9I 41 · C 68 · S 93 · K 54 · $0.035
  4. 24Malkuth-2B38.9I 41 · C 53 · S 91 · K 62 · ~$0.019 est.
  5. 25Raw Phi-4 mini direct logits38.0I 42 · C 59 · S 89 · K 50 · ~$0.048 est.
  6. 26JEV Qwen3.5-9B Base NVFP437.7I 47 · C 68 · S 93 · K 43 · ~$0.077 est.
  7. 27OpenJev36.9I 45 · C 55 · S 83 · K 45 · ~$0.066 est.
  8. 28kev 4B36.1I 42 · C 40 · S 76 · K 62 · ~$0.019 est.
  9. 29Decision 2B35.8I 39 · C 74 · S 84 · K 63 · ~$0.018 est.
  10. 30Qwen3.5-9B Jev-like data-mix v235.2I 47 · C 61 · S 82 · K 42 · ~$0.083 est.
  11. 31GPT-6 LunaAPI35.1I 96 · C 92 · S 74 · K 37 · $0.127
  12. 32SimpleJev Qwen3.8-27BAPI34.6I 52 · C 74 · S 71 · K 39 · ~$0.104 est.
  13. 33NInfer Qwen3.8-Flash-Next mixed34.0I 50 · C 79 · S 88 · K 39 · ~$0.109 est.
  14. 34swanOne33.6I 53 · C 71 · S 83 · K 39 · ~$0.111 est.
  15. 35GPT-6 LunaAPI33.3I 97 · C 93 · S 73 · K 36 · $0.135
  16. 36open-alternative-jev33.2I 39 · C 59 · S 83 · K 60 · ~$0.022 est.
  17. 37jev-local32.5I 45 · C 64 · S 69 · K 43 · ~$0.077 est.
  18. 38Decision Fast32.5I 37 · C 65 · S 82 · K 76 · ~$0.0063 est.
  19. 39decider-2b30.7I 39 · C 43 · S 83 · K 61 · ~$0.020 est.
  20. 40jeff30.6I 37 · C 68 · S 63 · K 77 · ~$0.0060 est.
  21. 41Laya30.3I 36 · C 64 · S 71 · K 86 · ~$0.0029 est.
  22. 42Kushal Patil — Gemma 4 31B ITAPI30.0I 60 · C 79 · S 84 · K 36 · $0.136
  23. 43Standard One 8B29.1I 46 · C 66 · S 92 · K 39 · ~$0.104 est.
  24. 44lev-350m28.5I 35 · C 71 · S 85 · K 76 · ~$0.0063 est.
  25. 45openjev-sglangAPI27.7I 49 · C 69 · S 77 · K 36 · ~$0.131 est.
  26. 46Von27.5I 34 · C 76 · S 70 · K 78 · ~$0.0055 est.
  27. 47NInfer Qwen3.8-27B NVFP426.9I 51 · C 76 · S 80 · K 35 · ~$0.145 est.
  28. 48NInfer Qwen3.8-27B NVFP426.3I 51 · C 67 · S 80 · K 35 · ~$0.145 est.
  29. 49kev 8B25.6I 42 · C 40 · S 75 · K 44 · ~$0.073 est.
  30. 50JevOne25.5I 48 · C 78 · S 88 · K 36 · ~$0.137 est.
  31. 51typecastlm25.3I 34 · C 62 · S 93 · K 60 · ~$0.021 est.
  32. 52SimpleJev Qwen3.6-35B-A3BAPI24.9I 46 · C 60 · S 75 · K 38 · ~$0.116 est.
  33. 53kev 0.6B24.8I 34 · C 50 · S 76 · K 76 · ~$0.0063 est.
  34. 54Raw Qwen3 8B direct logits23.7I 46 · C 24 · S 86 · K 42 · ~$0.087 est.
  35. 55system-one23.4I 44 · C 33 · S 84 · K 41 · ~$0.089 est.
  36. 56OpenDecision21.6I 32 · C 57 · S 80 · K 75 · ~$0.0066 est.
  37. 57LitJev19.5I 46 · C 77 · S 67 · K 34 · ~$0.163 est.
  38. 58openJev Verdict 1.419.0I 29 · C 72 · S 78 · K 82 · ~$0.0039 est.
  39. 59kev 0.5B18.9I 31 · C 50 · S 77 · K 76 · ~$0.0063 est.
  40. 60Bespoke Nimble 9B18.7I 46 · C 56 · S 79 · K 33 · ~$0.166 est.
  41. 61GPT-5.6 LunaAPI18.5I 93 · C 87 · S 78 · K 28 · $0.242
  42. 62openJev Verdict18.1I 30 · C 47 · S 77 · K 83 · ~$0.0037 est.
  43. 63Raw Qwen3 1.7B direct logits18.1I 33 · C 24 · S 90 · K 65 · ~$0.015 est.
  44. 64reflex-27b17.8I 46 · C 77 · S 67 · K 32 · ~$0.181 est.
  45. 65JevActAPI16.9I 29 · C 55 · S 76 · K 64 · ~$0.015 est.
  46. 66djev15.2I 72 · C 88 · S 75 · K 27 · ~$0.274 est.
  47. 67GLiNER2 large15.1I 31 · C 25 · S 62 · K 73 · ~$0.0077 est.
  48. 68OpenJev14.8I 58 · C 58 · S 76 · K 28 · ~$0.255 est.
  49. 69Qwen3.5-0.8B Decision Model14.5I 28 · C 68 · S 49 · K 76 · ~$0.0065 est.
  50. 70Gemini 3.1 Flash-LiteAPI14.3I 54 · C 59 · S 82 · K 27 · $0.264
  51. 71open-jev-deberta-v3-large12.6I 26 · C 67 · S 66 · K 74 · ~$0.0073 est.
  52. 72smalljev semantic-v912.3I 26 · C 59 · S 80 · K 58 · ~$0.025 est.
  53. 73GLiNER211.8I 27 · C 25 · S 72 · K 83 · ~$0.0037 est.
  54. 74InstinctAPI11.4I 51 · C 79 · S 84 · K 25 · ~$0.325 est.
  55. 75Open-Jev 9B11.2I 44 · C 62 · S 72 · K 28 · ~$0.249 est.
  56. 76Open-Jev 2B10.0I 42 · C 55 · S 73 · K 28 · ~$0.249 est.
  57. 77GLiNER2.5 multi9.8I 23 · C 57 · S 68 · K 82 · ~$0.0039 est.
  58. 78CLM-8B8.6I 22 · C 40 · S 94 · K 78 · ~$0.0052 est.
  59. 79SimpleJev7.5I 21 · C 49 · S 57 · K 68 · ~$0.011 est.
  60. 80GLiNER2.5 small7.2I 20 · C 51 · S 78 · K 82 · ~$0.0039 est.
  61. 81Raw Qwen3 0.6B direct logits7.1I 23 · C 21 · S 90 · K 74 · ~$0.0074 est.
  62. 82verdict-small5.7I 18 · C 67 · S 85 · K 97 · ~$0.0013 est.
  63. 83DeepSeek V4.1 FlashAPI4.8I 94 · C 96 · S 72 · K 17 · $0.594
  64. 84Mirror2.1I 14 · C 26 · S 71 · K 73 · ~$0.0077 est.
  65. 85Mixedbread mxbai-rerank-base-v20.4I 7 · C 84 · S 88 · K 68 · $0.012
  66. 86BAAI bge-reranker-v2-m30.2I 5 · C 84 · S 90 · K 73 · $0.0077
  67. 87Alibaba GTE Reranker ModernBERT-base0.2I 5 · C 79 · S 91 · K 70 · $0.010
  68. 88Certo v10.0I 0 · C 83 · S 94 · K 100 · ~$0.0010 est.
  69. 89Open Jev JSON Canvas0.0I 48 · C 0 · S 84 · K 46 · ~$0.065 est.
  70. classifier.dev (honorable mention)API70.8I 52 · C 72 · S 88 · K 84 · ~$0.0033 est.
  71. Qwen3.8 27B (partial run)API0.0I 40 · C 94 · S 61 · K 0 · ~$2.669 est.
  72. Needle 3, options as tools (partial run)—I – · C – · S – · K – · ~$0.014 est.
  73. Needle 3 (partial run)—I – · C – · S – · K – · ~$0.024 est.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; ~ est. = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers. Names link to each project.

Axes, accuracy, latency and cost

Every system with its four axes, public and sealed accuracy and the gap between them. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.

#SystemJevBench ScoreIntelligenceCalibrationSpeedCost axisPublic accuracy
534
Sealed accuracy
308
Public − sealed gapCost / 1,000p50 latencyEndpoint
1
decider-4b v2
†Decider-ai 1.2.2 (PyPI wheel identical to tag v1.2.2 abadc94), weights Mapika/decider-4b rev 7ab294cbdf6be6ac17fc818c10cdead744393d92 (decider_config version 4b-v2, T=1.935), uvicorn decider.serve:app. Disclosed by the author: 8,000 of the v2 LoRA rows come from generators written from the published names of the ten sealed families (no item read). Independent #1 gate (24 Sep): LEGIT. The author's private stage-2 training rows could not be audited for overlap with public items. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B dense size class, as decider-2b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Mapika
64.149.475.092.960.983.5%34.7%+48.8 pp~$0.020est.0.02 sRunPod GPU
263.353.176.383.352.086.6%36.7%+49.9 pp$0.0400.65 sAPI
3
JevK5 v0.2.0
†Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
by allebee
62.048.974.591.159.585.3%33.1%+52.2 pp~$0.022est.—unknown
4
Cygnet
†Blockbrain-ai/cygnet-recipe 81974de: frozen google/gemma-4-12B-it rev 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 on unmodified vLLM 0.30.0 (image vllm/vllm-openai:v0.30.0@sha256:8a69ffad…), the author's shim: options as letters in the benchmark's label order, logits masked to the option letters, one calibration temperature T=3.4 fitted on the author's own items. The shim sends its own system prompt and renders structured state with json indent=1. Over-context/over-26-option inputs are HTTP 422 (a wrong answer). Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by blockbrain · blockbrain, frozen Gemma-4-12B-it
61.849.574.990.752.887.9%33.8%+54.1 pp~$0.037est.0.04 sRunPod GPU
5
Hopper
†Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
by HopitAI
59.448.079.186.858.782.3%34.1%+48.2 pp~$0.024est.0.13 sRunPod GPU
6
Winnow-12B Q8
†The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
by Eldan Ring
55.648.364.882.352.985.7%33.1%+52.6 pp~$0.037est.0.23 sRunPod GPU
7
reflex 4B
†The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
by kshetrajna12
54.047.570.468.059.779.2%28.2%+51.0 pp~$0.022est.1.80 sRunPod GPU
8
djev
†The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
by Maisa (David Villalón) · Maisa, diffusion-gemma
52.247.055.491.457.684.0%29.9%+54.1 pp$0.026announced0.24 sAPI
9
Jev-Omni
†Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by akhilaaa3 · akhilaaa3, Gemma-4-12B merged
51.346.864.181.553.088.7%32.1%+56.6 pp~$0.037est.0.22 sRunPod GPU
10
metask-jev-4b
†Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
by Wayfind (metask-ai)
47.844.766.989.154.579.7%27.6%+52.1 pp~$0.033est.0.07 sRunPod GPU
11
SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ
47.744.466.883.759.581.0%26.3%+54.7 pp~$0.022est.0.20 sRunPod GPU
12
Jobe Qwen3.5-4B
†No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
by MantisShrimpdev · frozen
46.944.166.185.659.581.0%25.6%+55.3 pp~$0.022est.0.13 sRunPod GPU
13
local-jev Qwen3.5-4B
†Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
by Amith Chandrappa (amithgc)
46.844.473.375.055.880.5%26.0%+54.5 pp~$0.030est.0.71 sRunPod GPU
14
system-one-openAPIby mithalouni · Gemma 4 E2B LoRA on an L4
45.144.254.977.064.873.2%27.6%+45.6 pp~$0.015est.0.65 sauthor demo
15
spark-s1-4b-v6
†Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085
44.645.147.681.057.979.2%26.6%+52.6 pp~$0.025est.0.31 sRunPod GPU
16
Malkuth-4B
†Dhtocks/malkuth-4b rev 11dc416995dab324803cb6c533f1d5c69e19d630 (rank-16 LoRA + pointer head over Qwen/Qwen3.5-4B-Base rev 1001bb4d826a52d1f399e183466143f4da7b741b), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb (the author's evaluation revision). The released head carries temperature 1.0 (no fitted temperature), although the card says calibrated. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B one-pass size class, as kev-4b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by newfull5 (dhtocks) · newfull5, Kev post-train
44.543.161.588.361.574.9%23.4%+51.5 pp~$0.019est.0.09 sRunPod GPU
17
jqv
†A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
by hjmurmur (Octalab) · Qwen3-32B zero-shot
44.446.471.674.647.580.1%28.2%+51.8 pp~$0.056est.0.75 sRunPod GPU
18
Qwen3-Reranker-4B
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Qwen
43.544.665.278.749.268.0%29.9%+38.1 pp$0.0500.13 sRunPod GPU
19
decider-35b-a3b
†The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
by Mapika
41.247.265.380.845.383.1%31.5%+51.6 pp~$0.067est.0.29 sRunPod GPU
2041.046.429.187.659.769.7%27.3%+42.4 pp~$0.022est.0.08 sRunPod GPU
21
OpenSourceJev
†DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp
40.941.860.382.064.078.4%26.3%+52.1 pp~$0.016est.—unknown
22
ZeroEntropy zerank-2
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by ZeroEntropy
40.242.175.879.049.870.1%28.6%+41.6 pp$0.0470.13 sRunPod GPU
23
decision-machine-1
†A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
APIby milliseconds.ai (Baptiste Laget)
39.941.368.392.953.767.5%25.6%+41.9 pp$0.0350.17 sAPI
24
Malkuth-2B
†Dhtocks/malkuth-2b rev 401304b989451876070d83c488271b4f927d03ab (rank-16 LoRA + pointer head over empero-ai/Qwen3.8-2B-Distill rev e37a2dc4acc68ad75a91e07e63168cb04cc06345, fitted T=1.1755), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (no hosted ~2B listed; the 4B price errs high, the decider-2b precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by newfull5 (dhtocks) · newfull5, Kev post-train
38.941.353.591.461.569.7%24.4%+45.3 pp~$0.019est.0.04 sRunPod GPU
25
Raw Phi-4 mini direct logits
†Neutral raw-logit control, not JevBench-directed.
by Microsoft / neutral reproduction
38.041.858.888.849.665.8%29.2%+36.6 pp~$0.048est.0.06 sRunPod GPU
26
JEV Qwen3.5-9B Base NVFP4
†Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
by WilfLin
37.746.867.793.343.375.3%29.5%+45.8 pp~$0.077est.0.02 sRunPod GPU
27
OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16
36.945.455.083.245.581.8%28.6%+53.2 pp~$0.066est.0.24 sRunPod GPU
28
kev 4B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
36.142.139.675.761.866.2%22.4%+43.8 pp~$0.019est.0.55 sRunPod GPU
29
Decision 2B
†Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v59
35.838.874.184.362.575.3%26.0%+49.4 pp~$0.018est.0.19 sRunPod GPU
30
Qwen3.5-9B Jev-like data-mix v2
†The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
by jsaurabh
35.247.461.382.042.478.4%29.2%+49.1 pp~$0.083est.—unknown
31
GPT-6 Luna
†OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · low reasoning effort
35.195.892.073.736.999.1%92.9%+6.3 pp$0.1271.44 sAPI
32
SimpleJev Qwen3.8-27B
†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
34.651.674.571.239.586.6%35.7%+50.9 pp~$0.104est.1.01 sauthor demo
33
NInfer Qwen3.8-Flash-Next mixed
†Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
by Igor L. / NInfer contributors
34.049.578.688.238.989.6%34.1%+55.5 pp~$0.109est.0.08 sRunPod GPU
34
swanOne
†Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. blockbrain-ai/swanone-recipe abcca2e789316472e58d4e824b61c04f419c3ba5: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4 rev 925d7be6c14c6c9442ef83e8f05b5a3c39304f69 on vllm/vllm-openai:qwen38-flash-next@sha256:0aea3024… with MiaAI Lab's nine patched vLLM files (hash-verified) and the author's shim (option letters in the benchmark's order, the model's own renormalised letter mass, no temperature), H100.md serve command, here on an RTX PRO 6000 Blackwell. The shim sends its own system prompt; structured state is rendered with json indent=1. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter qwen/qwen3.8-flash list price $0.15/M input (same underlying Flash-Next weights; the NInfer Flash-Next precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by blockbrain · blockbrain, Qwen3.8-Flash-Next NVFP4
33.652.671.082.538.688.7%38.6%+50.1 pp~$0.111est.0.28 sRunPod GPU
35
GPT-6 Luna
†OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
APIby OpenAI · default medium reasoning effort
33.397.493.572.636.099.6%95.5%+4.1 pp$0.1351.48 sAPI
36
open-alternative-jev
†With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
by IkerMoel · Qwen3.5-4B, IkerMoel
33.238.658.783.559.674.0%24.4%+49.7 pp~$0.022est.0.21 sRunPod GPU
37
jev-local
†The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
by us (GitHub) · Qwen3.5-9B
32.545.264.269.243.374.9%29.5%+45.3 pp~$0.077est.1.05 sRunPod GPU
38
Decision Fast
†Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by FlyMy.AI (@denti) · FlyMy.AI, v53a
32.537.165.381.676.163.2%25.6%+37.6 pp~$0.0063est.0.24 sRunPod GPU
39
decider-2b
†The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
by Mapika
30.738.543.583.261.071.0%24.7%+46.3 pp~$0.020est.0.26 sRunPod GPU
40
jeff
†Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
by Logan Markewich · Logan Markewich, GLiFormer 400M
30.636.867.963.576.662.8%33.1%+29.7 pp~$0.0060est.0.94 sCPU
41
Laya
†The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
by Convai Innovations · Convai Innovations, ModernBERT-large 421M
30.336.163.771.186.258.4%30.8%+27.6 pp~$0.0029est.0.79 sCPU
42
Kushal Patil — Gemma 4 31B IT
†Kushal Patil's Gemma 4 31B IT served through the Autoloops systemone API. Cost = the API's own token usage x Autoloops' published rates ($0.20/M input, $0.45/M output; no output tokens were used). API measurement: public and sealed item content reached the operator endpoint; no gold labels or answers were sent.
APIby Kushal Patil · Autoloops
30.059.879.284.036.092.8%45.8%+47.1 pp$0.1360.60 sAPI
43
Standard One 8B
†StandardThinking/StandardOne-8B rev 0f14d009a9400e55ea5a00a89b4d859882db704e (Ministral-3-8B-Instruct-2512 + LoRA, merged) on stock SGLang 0.5.20 behind the author's jev-adapter from server/ at the same revision, nominated configuration: --prompt-wording native --native-system-prompt none --default-temperature 1.35, context 8192. Disclosed by the author: benchmark-directed development; the native wording was selected by an ablation on the public hard tier; 181 of 359,497 training rows reuse one generic 58-character public-hard instruction. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter mistralai/ministral-8b-2512 list price $0.15/M input (= the author's proposed Mistral API price for the exact base), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Standard Thinking (myeongho12)
29.146.265.992.039.476.6%26.6%+50.0 pp~$0.104est.0.02 sRunPod GPU
44
lev-350m
†Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M
28.534.870.685.376.158.4%25.0%+33.4 pp~$0.0063est.0.17 sRunPod GPU
45
openjev-sglangAPIby ekzhang · Qwen3.6-35B-A3B on SGLang
27.749.469.277.136.585.3%33.1%+52.2 pp~$0.131est.0.68 sauthor demo
46
Von
†The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M
27.534.575.770.577.857.1%27.9%+29.2 pp~$0.0055est.—unknown
47
NInfer Qwen3.8-27B NVFP4
†T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
by Igor L. / NInfer contributors · T=1.5
26.951.576.080.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
48
NInfer Qwen3.8-27B NVFP4
†Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
by Igor L. / NInfer contributors
26.351.567.280.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
49
kev 8B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
by Jared Palmer · research preview
25.641.840.274.944.071.4%21.8%+49.7 pp~$0.073est.0.59 sRunPod GPU
50
JevOne
†Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
by Juspay
25.547.677.688.535.989.6%33.8%+55.8 pp~$0.137est.0.09 sRunPod GPU
51
typecastlm
†Typecastlm[server] 1.1.2 (wheel identical to git e95911e), checkpoint mihailgribov/typecastlm-qwen3.5-3.8b rev dcfecfdc28e44ef82631901628535b71574e27af: frozen Qwen/Qwen3.5-4B blocks 0-27 plus a computed 39-row head, one forward pass, four calibration temperatures by question type; transformers 5.3.0, flash-linear-attention 0.5.2. States over 32,768 tokens are folded in the middle, not refused. The server reads a noul question's meaning from the order of its two criteria (first = true), not from their keys. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (the exact base weights; one pass, no output), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
by Mikhail Gribov · Mikhail Gribov, Qwen3.5-4B computed head
25.334.062.192.660.277.9%29.9%+48.1 pp~$0.021est.0.03 sRunPod GPU
52
SimpleJev Qwen3.6-35B-A3B
†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
APIby Featherless AI
24.945.759.875.038.181.4%28.2%+53.1 pp~$0.116est.0.85 sauthor demo
53
kev 0.6B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer · research preview
24.834.250.075.676.166.7%24.0%+42.6 pp~$0.0063est.0.59 sRunPod GPU
5423.745.724.186.341.968.4%26.3%+42.1 pp~$0.087est.0.08 sRunPod GPU
55
system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke
23.443.632.884.441.571.9%24.4%+47.5 pp~$0.089est.0.17 sRunPod GPU
56
OpenDecision
†A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
by Deepan Wadhwa · ModernBERT-large zero-shot
21.631.857.179.975.353.2%25.6%+27.6 pp~$0.0066est.0.34 sRunPod GPU
57
LitJev
†The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
by Zhengxu Yu · Qwen3.8-27B
19.546.376.666.733.686.1%30.8%+55.3 pp~$0.163est.2.03 sRunPod GPU
58
openJev Verdict 1.4
†Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
by Hemant (heman10x)
19.029.472.078.182.457.6%27.9%+29.7 pp~$0.0039est.0.31 sCPU
59
kev 0.5B
†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
by Jared Palmer
18.930.549.777.076.149.4%27.3%+22.1 pp~$0.0063est.0.43 sRunPod GPU
60
Bespoke Nimble 9B
†Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
by Bespoke Labs
18.746.356.478.733.479.7%28.9%+50.8 pp~$0.166est.0.39 sRunPod GPU
61
GPT-5.6 LunaAPIby OpenAI · low reasoning effort
18.593.187.477.528.597.4%89.0%+8.4 pp$0.2420.97 sAPI
62
openJev Verdict
†The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
by Hemant (heman10x) · heman10x, ModernBERT-base 151M
18.130.047.076.783.155.4%24.7%+30.7 pp~$0.0037est.0.28 sCPU
6318.133.224.389.764.954.1%26.0%+28.1 pp~$0.015est.0.07 sRunPod GPU
64
reflex-27b
†The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
by kshetrajna12 · Qwen3.8-27B
17.846.477.267.532.387.0%29.5%+57.5 pp~$0.181est.1.89 sRunPod GPU
65
JevAct
†Requested in GitHub issue #66. The author published an endpoint and an evaluation client, not weights, and his API is not the TypeSafe wire format, so JevBench's own `jevact_api` adapter implements exactly the mapping table in his eval_code.zip: state to state, instructions to the question, the labels with their criteria text as the options in label order, noul as false/true, and results[0].options[i].probability read back as the probability of labels[i] by index. His server answers HTTP 400 with `all inference items exceeded max_context_tokens or contained reserved markers` for an input it cannot take, and his own client and test script record exactly that as one failed item and continue; the adapter therefore surfaces that one documented refusal as the harness's 422 refusal path, which scores it as a wrong answer and does not count toward the stop rule - the same treatment the swanOne and Cygnet packages get for their own 422. Any other 400 or an outage still stops the run. The endpoint reports no token accounting, so cost is the labelled 2B size-class estimate. The endpoint is one machine in China and the latency includes that distance; it is not a production API and gets the standard x2 self-host adjustment, without the +0.15 s that only our own servers carry.
APIby einptein · einptein, jev1-2b-v2
16.929.355.176.564.361.9%23.4%+38.5 pp~$0.015est.0.40 sauthor demo
66
djev
†Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
by David Villalon / Maisa · thinking
15.271.687.875.226.987.4%60.1%+27.4 pp~$0.274est.0.43 sRunPod GPU
67
GLiNER2 large
†The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino
15.131.124.861.773.356.7%28.6%+28.1 pp~$0.0077est.1.10 sCPU
68
OpenJev
†OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
by razorback16 · thinking, BF16
14.858.158.176.127.888.7%42.2%+46.5 pp~$0.255est.0.46 sRunPod GPU
69
Qwen3.5-0.8B Decision Model
†JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
by Mourad Ghafiri
14.528.168.249.275.759.3%34.7%+24.6 pp~$0.0065est.7.15 sCPU
7014.354.559.381.827.487.0%38.6%+48.4 pp$0.2640.76 sAPI
71
open-jev-deberta-v3-large
†297/308 sealed items answered validly (failures count as wrong)
by Kotoba Labs · local CPU
12.625.666.666.074.052.4%29.5%+22.8 pp~$0.0073est.1.77 sCPU
72
smalljev semantic-v9
†The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
by Aditya (isHeSatoshi)
12.325.759.279.857.960.6%26.9%+33.7 pp~$0.025est.0.41 sRunPod GPU
73
GLiNER2
†A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
by Fastino · Fastino, gliner2.5-base
11.827.425.271.883.158.0%29.2%+28.8 pp~$0.0037est.0.31 sCPU
74
Instinct
†Requested in GitHub issue #69. A free evaluation/demo endpoint that speaks TypeSafe's /v1/systemone wire format, so the unchanged typesafe adapter ran it and no mapping of ours was involved; we registered the account and created the key ourselves in their console. The author states frozen Qwen3.8-27B base weights with no fine-tuning, read in one forward pass per question with no autoregressive decoding; the serving stack is not public, so nothing about it could be reviewed and the row rests on the API's own behaviour. Cost is an ESTIMATE: ZooWork publishes no bookable price, so the base model's public reference price is used (OpenRouter model-level qwen/qwen3.8-27b, $0.42/M input, zero output for a direct-logit readout); the author-announced tariff is not used. Classified as a free evaluation/demo endpoint (no SLA, no status page, no terms, pricing 'to be announced'), so the x2 latency adjustment applies. Re-scored when a bookable price is published.
APIby rayrain-srp (ZooWork) · ZooWork, Qwen3.8-27B
11.451.178.783.924.686.6%35.1%+51.5 pp~$0.325est.0.26 sauthor demo
75
Open-Jev 9B
†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
11.244.261.872.028.177.5%29.9%+47.6 pp~$0.249est.0.75 sRunPod GPU
76
Open-Jev 2B
†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
by Zefan Cai (@Zefan_Cai)
10.042.355.373.528.164.5%26.3%+38.2 pp~$0.249est.0.66 sRunPod GPU
77
GLiNER2.5 multi
†The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
by Fastino · Fastino, 287M
9.823.157.267.882.448.9%32.8%+16.1 pp~$0.0039est.0.43 sCPU
78
CLM-8B
†Contrastive-LM/CLM commit cca045ffdb07b3ebcfe6938537cdeac5e14899c9, head Contrastive-LM/CLM-v0.1-8B rev 87655cb835bd76fd66c2da78e1e3709f7fa11a94 (clm-latest), Qwen3-8B rev b968826d9c46dd6066d109eabc6255188de91218 last-token pooling via vLLM. Authors' documented system_one path: state + instructions as state text, each option description as a candidate action, softmax over contrastive scores at temperature 1.0. Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod; no operator endpoint; no golds were exposed. Latency is in-process Engine.answer time on the serial standard+judge items with the self-hosted adjustment. Cost is estimated at the Qwen3-Embedding-8B hosted list price ($0.01/M input, same-size 8B pooling encoder) over CLM's measured encoder tokens; it is not a GPU bill.
by Contrastive-LM (Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini) · Contrastive-LM, clm-latest
8.622.439.893.678.440.7%24.0%+16.7 pp~$0.0052est.0.02 sRunPod GPU
79
SimpleJev
†143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU
7.521.549.157.568.354.5%34.7%+19.8 pp~$0.011est.3.99 sCPU
80
GLiNER2.5 small
†The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
by Fastino · Fastino, 74M
7.220.550.777.882.445.9%28.6%+17.3 pp~$0.0039est.0.11 sCPU
817.122.820.789.973.948.1%25.6%+22.4 pp~$0.0074est.0.07 sRunPod GPU
82
verdict-small
†Requested in GitHub issue #73. Run through the author's own `verdict serve` on our CPU, which speaks TypeSafe's /v1/systemone wire format, so JevBench's unchanged typesafe adapter ran it and no mapping of ours was involved. A 118M multilingual bi-encoder: every option is scored against the rendered state by cosine similarity, at the model's own scale (temperature 1.0, no calibrator fitted on JevBench items, as the author states). Structured state is rendered as `key: value` lines by his own code. Code review before the run: the only network call is the Hugging Face download of his own checkpoint, no telemetry, no key, no rule written against public items. The `usage.input_tokens` his server reports is a word count and not a tokeniser count, so the cost is the labelled size-class estimate rather than a measured token price. Self-host latency gets the standard x2 + 0.15 s adjustment. The author's own public-set figures were easy 0.938, standard 0.486, hard 0.396 on an Apple M5 CPU. Offline measurement of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) on our own CPU through the author's server; no operator endpoint.
by Manavarya09 (Manav) · Manavarya09, multilingual-e5-small 118M
5.718.166.985.096.653.7%28.9%+24.8 pp~$0.0013est.0.05 sCPU
83
DeepSeek V4.1 FlashAPIby DeepSeek · thinking default
4.894.095.571.616.897.8%94.8%+3.0 pp$0.5941.42 sAPI
84
Mirror
†171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
by Bluusun
2.113.626.070.873.341.1%9.1%+32.0 pp~$0.0077est.0.90 sauthor demo
85
Mixedbread mxbai-rerank-base-v2
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Mixedbread
0.46.884.187.567.937.2%34.4%+2.8 pp$0.0120.07 sRunPod GPU
86
BAAI bge-reranker-v2-m3
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by BAAI
0.25.084.289.573.439.4%27.9%+11.5 pp$0.00770.03 sRunPod GPU
87
Alibaba GTE Reranker ModernBERT-base
†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
by Alibaba-NLP
0.24.878.990.669.633.8%33.4%+0.3 pp$0.0100.05 sRunPod GPU
88
Certo v1
†The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
by AltSlate Labs
0.00.183.094.0100.031.6%29.5%+2.1 pp~$0.0010est.0.02 sRunPod GPU
89
Open Jev JSON Canvas
†Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
by JoshuaSP
0.048.20.084.145.684.4%31.2%+53.2 pp~$0.065est.0.22 sRunPod GPU
—
classifier.dev
†Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
APIby mrmps (@michael_chomsky) · fast tier
honorable mention · not ranked
70.851.672.487.684.385.3%34.4%+50.9 pp~$0.0033est.0.39 sAPI
—
Qwen3.8 27BAPIby Qwen / Chutes · Chutes TEE
partial · not ranked
0.040.493.661.30.071.9%21.8%+50.1 pp~$2.669est.5.75 sAPI
—
Needle 3, options as tools
†V1.4: Not re-measured: same as needle-3.
by Cactus Compute · post-hoc adapter mode
partial · not ranked
—————22.1%——~$0.014est.3.78 sCPU
—
Needle 3
†V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.
by Cactus Compute · Cactus, 2-bit, local CPU
partial · not ranked
—————22.5%——~$0.024est.1.69 sCPU

API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. The sealed text and answers remain private; only system-level aggregates appear here. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.

All 83 system notes and disclosures
  • † decider-4b v2 (Mapika): Decider-ai 1.2.2 (PyPI wheel identical to tag v1.2.2 abadc94), weights Mapika/decider-4b rev 7ab294cbdf6be6ac17fc818c10cdead744393d92 (decider_config version 4b-v2, T=1.935), uvicorn decider.serve:app. Disclosed by the author: 8,000 of the v2 LoRA rows come from generators written from the published names of the ten sealed families (no item read). Independent #1 gate (24 Sep): LEGIT. The author's private stage-2 training rows could not be audited for overlap with public items. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B dense size class, as decider-2b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
  • † Cygnet (blockbrain, frozen Gemma-4-12B-it): Blockbrain-ai/cygnet-recipe 81974de: frozen google/gemma-4-12B-it rev 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 on unmodified vLLM 0.30.0 (image vllm/vllm-openai:v0.30.0@sha256:8a69ffad…), the author's shim: options as letters in the benchmark's label order, logits masked to the option letters, one calibration temperature T=3.4 fitted on the author's own items. The shim sends its own system prompt and renders structured state with json indent=1. Over-context/over-26-option inputs are HTTP 422 (a wrong answer). Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
  • † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • † reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • † djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
  • † Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
  • † Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
  • † local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
  • † spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Malkuth-4B (newfull5, Kev post-train): Dhtocks/malkuth-4b rev 11dc416995dab324803cb6c533f1d5c69e19d630 (rank-16 LoRA + pointer head over Qwen/Qwen3.5-4B-Base rev 1001bb4d826a52d1f399e183466143f4da7b741b), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb (the author's evaluation revision). The released head carries temperature 1.0 (no fitted temperature), although the card says calibrated. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (4B one-pass size class, as kev-4b), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • † Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
  • † OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
  • † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • † Malkuth-2B (newfull5, Kev post-train): Dhtocks/malkuth-2b rev 401304b989451876070d83c488271b4f927d03ab (rank-16 LoRA + pointer head over empero-ai/Qwen3.8-2B-Distill rev e37a2dc4acc68ad75a91e07e63168cb04cc06345, fitted T=1.1755), served by kev.serve from jaredpalmer/kev 557598fced1dada75dfbf36ed144dce309ac6ceb. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (no hosted ~2B listed; the 4B price errs high, the decider-2b precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
  • † JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
  • † kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
  • † Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
  • † GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
  • † swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4): Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. blockbrain-ai/swanone-recipe abcca2e789316472e58d4e824b61c04f419c3ba5: Mia-AiLab/Qwen3.8-Flash-Next-NVFP4 rev 925d7be6c14c6c9442ef83e8f05b5a3c39304f69 on vllm/vllm-openai:qwen38-flash-next@sha256:0aea3024… with MiaAI Lab's nine patched vLLM files (hash-verified) and the author's shim (option letters in the benchmark's order, the model's own renormalised letter mass, no temperature), H100.md serve command, here on an RTX PRO 6000 Blackwell. The shim sends its own system prompt; structured state is rendered with json indent=1. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter qwen/qwen3.8-flash list price $0.15/M input (same underlying Flash-Next weights; the NInfer Flash-Next precedent), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • † jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • † Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • † jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • † Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • † Kushal Patil — Gemma 4 31B IT (Autoloops): Kushal Patil's Gemma 4 31B IT served through the Autoloops systemone API. Cost = the API's own token usage x Autoloops' published rates ($0.20/M input, $0.45/M output; no output tokens were used). API measurement: public and sealed item content reached the operator endpoint; no gold labels or answers were sent.
  • † Standard One 8B (Standard Thinking): StandardThinking/StandardOne-8B rev 0f14d009a9400e55ea5a00a89b4d859882db704e (Ministral-3-8B-Instruct-2512 + LoRA, merged) on stock SGLang 0.5.20 behind the author's jev-adapter from server/ at the same revision, nominated configuration: --prompt-wording native --native-system-prompt none --default-temperature 1.35, context 8192. Disclosed by the author: benchmark-directed development; the native wording was selected by an ablation on the public hard tier; 181 of 359,497 training rows reuse one generic 58-character public-hard instruction. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: OpenRouter mistralai/ministral-8b-2512 list price $0.15/M input (= the author's proposed Mistral API price for the exact base), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
  • † NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
  • † NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
  • † kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
  • † typecastlm (Mikhail Gribov, Qwen3.5-4B computed head): Typecastlm[server] 1.1.2 (wheel identical to git e95911e), checkpoint mihailgribov/typecastlm-qwen3.5-3.8b rev dcfecfdc28e44ef82631901628535b71574e27af: frozen Qwen/Qwen3.5-4B blocks 0-27 plus a computed 39-row head, one forward pass, four calibration temperatures by question type; transformers 5.3.0, flash-linear-attention 0.5.2. States over 32,768 tokens are folded in the middle, not refused. The server reads a noul question's meaning from the order of its two criteria (first = true), not from their keys. Offline self-hosted inference of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 Blackwell 96 GB pod, through JevBench's unchanged typesafe adapter against the author's own server on loopback; no operator endpoint; no golds were exposed. Latency is the serial request wall time on the standard+judge items with the self-hosted adjustment. Cost is a labelled estimate: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M input (the exact base weights; one pass, no output), read 2026-09-24, over the server's own usage.input_tokens; it is not a GPU bill.
  • † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
  • † Raw Qwen3 8B direct logits: Neutral raw-logit control.
  • † OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • † LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
  • † Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • † openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • † Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
  • † reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • † JevAct (einptein, jev1-2b-v2): Requested in GitHub issue #66. The author published an endpoint and an evaluation client, not weights, and his API is not the TypeSafe wire format, so JevBench's own `jevact_api` adapter implements exactly the mapping table in his eval_code.zip: state to state, instructions to the question, the labels with their criteria text as the options in label order, noul as false/true, and results[0].options[i].probability read back as the probability of labels[i] by index. His server answers HTTP 400 with `all inference items exceeded max_context_tokens or contained reserved markers` for an input it cannot take, and his own client and test script record exactly that as one failed item and continue; the adapter therefore surfaces that one documented refusal as the harness's 422 refusal path, which scores it as a wrong answer and does not count toward the stop rule - the same treatment the swanOne and Cygnet packages get for their own 422. Any other 400 or an outage still stops the run. The endpoint reports no token accounting, so cost is the labelled 2B size-class estimate. The endpoint is one machine in China and the latency includes that distance; it is not a production API and gets the standard x2 self-host adjustment, without the +0.15 s that only our own servers carry.
  • † djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
  • † GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • † Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
  • † open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
  • † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • † GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • † Instinct (ZooWork, Qwen3.8-27B): Requested in GitHub issue #69. A free evaluation/demo endpoint that speaks TypeSafe's /v1/systemone wire format, so the unchanged typesafe adapter ran it and no mapping of ours was involved; we registered the account and created the key ourselves in their console. The author states frozen Qwen3.8-27B base weights with no fine-tuning, read in one forward pass per question with no autoregressive decoding; the serving stack is not public, so nothing about it could be reviewed and the row rests on the API's own behaviour. Cost is an ESTIMATE: ZooWork publishes no bookable price, so the base model's public reference price is used (OpenRouter model-level qwen/qwen3.8-27b, $0.42/M input, zero output for a direct-logit readout); the author-announced tariff is not used. Classified as a free evaluation/demo endpoint (no SLA, no status page, no terms, pricing 'to be announced'), so the x2 latency adjustment applies. Re-scored when a bookable price is published.
  • † Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • † CLM-8B (Contrastive-LM, clm-latest): Contrastive-LM/CLM commit cca045ffdb07b3ebcfe6938537cdeac5e14899c9, head Contrastive-LM/CLM-v0.1-8B rev 87655cb835bd76fd66c2da78e1e3709f7fa11a94 (clm-latest), Qwen3-8B rev b968826d9c46dd6066d109eabc6255188de91218 last-token pooling via vLLM. Authors' documented system_one path: state + instructions as state text, each option description as a candidate action, softmax over contrastive scores at temperature 1.0. Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container on an evaluator-owned Lium RTX PRO 6000 pod; no operator endpoint; no golds were exposed. Latency is in-process Engine.answer time on the serial standard+judge items with the self-hosted adjustment. Cost is estimated at the Qwen3-Embedding-8B hosted list price ($0.01/M input, same-size 8B pooling encoder) over CLM's measured encoder tokens; it is not a GPU bill.
  • † SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
  • † GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
  • † verdict-small (Manavarya09, multilingual-e5-small 118M): Requested in GitHub issue #73. Run through the author's own `verdict serve` on our CPU, which speaks TypeSafe's /v1/systemone wire format, so JevBench's unchanged typesafe adapter ran it and no mapping of ours was involved. A 118M multilingual bi-encoder: every option is scored against the rendered state by cosine similarity, at the model's own scale (temperature 1.0, no calibrator fitted on JevBench items, as the author states). Structured state is rendered as `key: value` lines by his own code. Code review before the run: the only network call is the Hugging Face download of his own checkpoint, no telemetry, no key, no rule written against public items. The `usage.input_tokens` his server reports is a word count and not a tokeniser count, so the cost is the labelled size-class estimate rather than a measured token price. Self-host latency gets the standard x2 + 0.15 s adjustment. The author's own public-set figures were easy 0.938, standard 0.486, hard 0.396 on an Apple M5 CPU. Offline measurement of all 842 decisions (534 frozen v1.2 + 308 sealed v1.4) on our own CPU through the author's server; no operator endpoint.
  • † Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
  • † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • † Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
  • † classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
  • † Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
  • † Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.

Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.

Artifact: v1.4.2 results JSON · SHA-256 ac14e206dde5… · JevBench v1.4.2 release and method

Compare two systems

Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, and accuracy by family on the v1.2 hard tier and on the sealed set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#2)
  • B: decider-4b v2 — system-one-open · Score 64.1 (#1)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs decider-4b v2. Intelligence: 53.1 vs 49.4; Calibration: 76.3 vs 75.0; Speed: 83.3 vs 92.9; Cost: 52.0 vs 60.9.50100Intelligence53.1 · 49.4Calibration76.3 · 75.0Speed83.3 · 92.9Cost52.0 · 60.9
0–100, the values in the table. A label-only system has no calibration (counted as 0).

Accuracy per tier, incl. sealed

Radar: accuracy per tier, incl. sealed, two systemsAccuracy per tier, incl. sealed, Jev 1.13.0 vs decider-4b v2. Easy: 100% vs 100%; Standard: 99% vs 97%; Judge: 95% vs 88%; Hard: 74% vs 67%; Sealed: 37% vs 35%.50100Easy100% · 100%Standard99% · 97%Judge95% · 88%Hard74% · 67%Sealed37% · 35%
Share correct per tier; Sealed = the 308 private decisions, aggregate only.

Hard tier by family (v1.2 topics)

decider-4b v2 has no published v1.2 hard-tier family breakdown.

Radar: hard tier by family (v1.2 topics), two systemsHard tier by family (v1.2 topics), Jev 1.13.0 vs decider-4b v2. Adversarial: 100% vs —; Ambiguous: 79% vs —; Judge: 79% vs —; Long policy: 61% vs —; Multi-hop: 86% vs —; Probability: 80% vs —; Routing: 100% vs —; Temporal / numeric: 27% vs —; Trade-off: 92% vs —; Trap: 100% vs —.50100Adversarial100%Ambiguous79%Judge79%Long policy61%Multi-hop86%Probability80%Routing100%Temporal /numeric27%Trade-off92%Trap100%
Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out).

Sealed set by family

Radar: sealed set by family, two systemsSealed set by family, Jev 1.13.0 vs decider-4b v2. Ambiguous / abstain: 30% vs 43%; Judge: 34% vs 32%; Long policy: 28% vs 30%; Multi-hop: 45% vs 34%; Paraphrase: 64% vs 50%; Probability: 50% vs 36%; Safety judge: 38% vs 38%; Temporal / numeric: 29% vs 30%; Trade-off: 38% vs 35%; Trap / adversarial: 42% vs 33%.50100Ambiguous /abstain30% · 43%Judge34% · 32%Long policy28% · 30%Multi-hop45% · 34%Paraphrase64% · 50%Probability50% · 36%Safety judge38% · 38%Temporal /numeric29% · 30%Trade-off38% · 35%Trap /adversarial42% · 33%
Share correct within each sealed family — system-level aggregates; the items stay private.
All values as a table
SpokeA: Jev 1.13.0B: decider-4b v2
The four score axes
Intelligence53.149.4
Calibration76.375.0
Speed83.392.9
Cost52.060.9
Accuracy per tier, incl. sealed
Easy100%100%
Standard99%97%
Judge95%88%
Hard74%67%
Sealed37%35%
Hard tier by family (v1.2 topics)
Adversarial100%—
Ambiguous79%—
Judge79%—
Long policy61%—
Multi-hop86%—
Probability80%—
Routing100%—
Temporal / numeric27%—
Trade-off92%—
Trap100%—
Sealed set by family
Ambiguous / abstain30%43%
Judge34%32%
Long policy28%30%
Multi-hop45%34%
Paraphrase64%50%
Probability50%36%
Safety judge38%38%
Temporal / numeric29%30%
Trade-off38%35%
Trap / adversarial42%33%

What changed in v1.4

  • Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence: I = 0.8 × I_v1.3 + 0.2 × I_sealed, where I_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates.
  • Calibration blends toward the sealed-inclusive result at the approved weight: C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35).
  • The k = 1 generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points: I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half.
  • The four axes use an equal-weight harmonic mean (p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0.
  • The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.

JevBench v1.4.2 · additional views

Capability, cost and speed

Capability is the arithmetic mean of Intelligence and Calibration: (Intelligence + Calibration) / 2, on a 0–100 scale. Cost is USD per 1,000 decisions; its axis is logarithmic, and lower is better. Speed uses the JevBench Speed axis, where higher is faster. Estimated costs are marked.

2 placeholder rows have no published Intelligence or Calibration values and are omitted.

Top 20 by Capability

Capability with cost alongside

Each system has a wide Capability bar and a narrower cost bar. The cost scale is logarithmic: longer bars mean higher cost, so shorter is cheaper.

  1. 1GPT-6 Luna (medium)API95.4Cost $0.14
  2. 2DeepSeek V4.1 FlashAPI94.7Cost $0.59
  3. 3GPT-6 Luna (low)API93.9Cost $0.13
  4. 4GPT-5.6 LunaAPI90.3Cost $0.24
  5. 5djev79.7Cost $0.27 est.
  6. 6Kushal Patil — Gemma 4 31B ITAPI69.5Cost $0.14
  7. 7Qwen3.8 27BAPI67.0Cost $2.67 est.
  8. 8InstinctAPI64.9Cost $0.33 est.
  9. 9Jev 1.13.0API64.7Cost $0.040
  10. 10NInfer Qwen3.8-Flash-Next mixed64.1Cost $0.11 est.
  11. 11NInfer Qwen3.8-27B NVFP463.7Cost $0.14 est.
  12. 12Hopper63.5Cost $0.024 est.
  13. 13SimpleJev Qwen3.8-27BAPI63.0Cost $0.10 est.
  14. 14JevOne62.6Cost $0.14 est.
  15. 15Cygnet62.2Cost $0.037 est.
  16. 16decider-4b v262.2Cost $0.020 est.
  17. 17classifier.devAPI62.0Cost $0.0033 est.
  18. 18reflex-27b61.8Cost $0.18 est.
  19. 19swanOne61.8Cost $0.11 est.
  20. 20JevK5 v0.2.061.7Cost $0.022 est.
Show all 91 systems (71 more)
  1. 21LitJev61.4Cost $0.16 est.
  2. 22openjev-sglangAPI59.3Cost $0.13 est.
  3. 23NInfer Qwen3.8-27B NVFP459.3Cost $0.14 est.
  4. 24jqv59.0Cost $0.056 est.
  5. 25reflex 4B58.9Cost $0.022 est.
  6. 26ZeroEntropy zerank-258.9Cost $0.047
  7. 27local-jev Qwen3.5-4B58.9Cost $0.030 est.
  8. 28OpenJev58.1Cost $0.25 est.
  9. 29JEV Qwen3.5-9B Base NVFP457.3Cost $0.077 est.
  10. 30Gemini 3.1 Flash-LiteAPI56.9Cost $0.26
  11. 31Winnow-12B Q856.6Cost $0.037 est.
  12. 32Decision 2B56.4Cost $0.018 est.
  13. 33decider-35b-a3b56.2Cost $0.067 est.
  14. 34Standard One 8B56.0Cost $0.10 est.
  15. 35metask-jev-4b55.8Cost $0.033 est.
  16. 36SemIf55.6Cost $0.022 est.
  17. 37Jev-Omni55.4Cost $0.037 est.
  18. 38Von55.1Cost $0.0055 est.
  19. 39Jobe Qwen3.5-4B55.1Cost $0.022 est.
  20. 40Qwen3-Reranker-4B54.9Cost $0.050
  21. 41decision-machine-1API54.8Cost $0.035
  22. 42jev-local54.7Cost $0.077 est.
  23. 43Qwen3.5-9B Jev-like data-mix v254.4Cost $0.083 est.
  24. 44Open-Jev 9B53.0Cost $0.25 est.
  25. 45SimpleJev Qwen3.6-35B-A3BAPI52.8Cost $0.12 est.
  26. 46lev-350m52.7Cost $0.0063 est.
  27. 47jeff52.3Cost $0.0060 est.
  28. 48Malkuth-4B52.3Cost $0.019 est.
  29. 49Bespoke Nimble 9B51.4Cost $0.17 est.
  30. 50Decision Fast51.2Cost $0.0063 est.
  31. 51djev51.2Cost $0.026
  32. 52OpenSourceJev51.1Cost $0.016 est.
  33. 53openJev Verdict 1.450.7Cost $0.0039 est.
  34. 54Raw Phi-4 mini direct logits50.3Cost $0.048 est.
  35. 55OpenJev50.2Cost $0.066 est.
  36. 56Laya49.9Cost $0.0029 est.
  37. 57system-one-openAPI49.5Cost $0.015 est.
  38. 58Open-Jev 2B48.8Cost $0.25 est.
  39. 59open-alternative-jev48.6Cost $0.022 est.
  40. 60Qwen3.5-0.8B Decision Model48.2Cost $0.0065 est.
  41. 61typecastlm48.0Cost $0.021 est.
  42. 62Malkuth-2B47.4Cost $0.019 est.
  43. 63spark-s1-4b-v646.3Cost $0.025 est.
  44. 64open-jev-deberta-v3-large46.1Cost $0.0073 est.
  45. 65Mixedbread mxbai-rerank-base-v245.5Cost $0.012
  46. 66BAAI bge-reranker-v2-m344.6Cost $0.0077
  47. 67OpenDecision44.4Cost $0.0066 est.
  48. 68verdict-small42.5Cost $0.0013 est.
  49. 69smalljev semantic-v942.4Cost $0.025 est.
  50. 70JevActAPI42.2Cost $0.015 est.
  51. 71kev 0.6B42.1Cost $0.0063 est.
  52. 72Alibaba GTE Reranker ModernBERT-base41.9Cost $0.010
  53. 73Certo v141.5Cost $0.00097 est.
  54. 74decider-2b41.0Cost $0.020 est.
  55. 75kev 8B41.0Cost $0.073 est.
  56. 76kev 4B40.9Cost $0.019 est.
  57. 77GLiNER2.5 multi40.1Cost $0.0039 est.
  58. 78kev 0.5B40.1Cost $0.0063 est.
  59. 79openJev Verdict38.5Cost $0.0037 est.
  60. 80system-one38.2Cost $0.089 est.
  61. 81Raw Qwen3 4B Instruct 2507 direct logits37.7Cost $0.022 est.
  62. 82GLiNER2.5 small35.6Cost $0.0039 est.
  63. 83SimpleJev35.3Cost $0.011 est.
  64. 84Raw Qwen3 8B direct logits34.9Cost $0.087 est.
  65. 85CLM-8B31.1Cost $0.0052 est.
  66. 86Raw Qwen3 1.7B direct logits28.8Cost $0.015 est.
  67. 87GLiNER2 large28.0Cost $0.0077 est.
  68. 88GLiNER226.3Cost $0.0037 est.
  69. 89Open Jev JSON Canvas24.1Cost $0.065 est.
  70. 90Raw Qwen3 0.6B direct logits21.8Cost $0.0074 est.
  71. 91Mirror19.8Cost $0.0077 est.
Cost per 1,000 decisions · logarithmic · lower is better; free is at the left edge.
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Cost per 1,000 decisions · log scale
Cost bars use the right-hand scale, from $0.00097 to $2.67 per 1,000 decisions. Free cost is placed at the cheapest edge; missing cost is shown as —.

Capability vs cost

The upper-left is the more attractive area: higher Capability and lower cost.

Capability vs costThe upper-left is the more attractive area: higher Capability and lower cost. Each point has a tooltip with the system and its values.020406080100$0.0010$0.010$0.10$1.00USD per 1,000 decisions · log scale · cheaper ←Capability · higher ↑GPT-6 Luna (medium) · JevBench rank 35 · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · JevBench rank 83 · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna (low) · JevBench rank 31 · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · JevBench rank 61 · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · JevBench rank 66 · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Kushal Patil — Gemma 4 31B IT · JevBench rank 42 · Capability 69.5 · Cost $0.14 per 1,000 decisions · Speed 84.0.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Instinct · JevBench rank 74 · Capability 64.9 · Cost $0.33 estimated per 1,000 decisions · Speed 83.9.Jev 1.13.0 · JevBench rank 2 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · JevBench rank 33 · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · JevBench rank 47 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · JevBench rank 5 · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · JevBench rank 32 · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · JevBench rank 50 · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.Cygnet · JevBench rank 4 · Capability 62.2 · Cost $0.037 estimated per 1,000 decisions · Speed 90.7.decider-4b v2 · JevBench rank 1 · Capability 62.2 · Cost $0.020 estimated per 1,000 decisions · Speed 92.9.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · JevBench rank 64 · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.swanOne · JevBench rank 34 · Capability 61.8 · Cost $0.11 estimated per 1,000 decisions · Speed 82.5.JevK5 v0.2.0 · JevBench rank 3 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · JevBench rank 57 · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · JevBench rank 45 · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · JevBench rank 48 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · JevBench rank 17 · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · JevBench rank 7 · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · JevBench rank 22 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · JevBench rank 13 · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · JevBench rank 68 · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · JevBench rank 26 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · JevBench rank 70 · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · JevBench rank 6 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · JevBench rank 29 · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · JevBench rank 19 · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.Standard One 8B · JevBench rank 43 · Capability 56.0 · Cost $0.10 estimated per 1,000 decisions · Speed 92.0.metask-jev-4b · JevBench rank 10 · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · JevBench rank 11 · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · JevBench rank 9 · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · JevBench rank 46 · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · JevBench rank 12 · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · JevBench rank 18 · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · JevBench rank 23 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · JevBench rank 37 · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · JevBench rank 30 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · JevBench rank 75 · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · JevBench rank 52 · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · JevBench rank 44 · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · JevBench rank 40 · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Malkuth-4B · JevBench rank 16 · Capability 52.3 · Cost $0.019 estimated per 1,000 decisions · Speed 88.3.Bespoke Nimble 9B · JevBench rank 60 · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · JevBench rank 38 · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · JevBench rank 8 · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · JevBench rank 21 · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · JevBench rank 58 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · JevBench rank 25 · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · JevBench rank 27 · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · JevBench rank 41 · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · JevBench rank 14 · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · JevBench rank 76 · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · JevBench rank 36 · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · JevBench rank 69 · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.typecastlm · JevBench rank 51 · Capability 48.0 · Cost $0.021 estimated per 1,000 decisions · Speed 92.6.Malkuth-2B · JevBench rank 24 · Capability 47.4 · Cost $0.019 estimated per 1,000 decisions · Speed 91.4.spark-s1-4b-v6 · JevBench rank 15 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · JevBench rank 71 · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · JevBench rank 85 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · JevBench rank 86 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · JevBench rank 56 · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.verdict-small · JevBench rank 82 · Capability 42.5 · Cost $0.0013 estimated per 1,000 decisions · Speed 85.0.smalljev semantic-v9 · JevBench rank 72 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.JevAct · JevBench rank 65 · Capability 42.2 · Cost $0.015 estimated per 1,000 decisions · Speed 76.5.kev 0.6B · JevBench rank 53 · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · JevBench rank 87 · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · JevBench rank 88 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · JevBench rank 39 · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · JevBench rank 49 · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · JevBench rank 28 · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · JevBench rank 77 · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · JevBench rank 59 · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · JevBench rank 62 · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · JevBench rank 55 · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · JevBench rank 20 · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · JevBench rank 80 · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · JevBench rank 79 · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · JevBench rank 54 · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.CLM-8B · JevBench rank 78 · Capability 31.1 · Cost $0.0052 estimated per 1,000 decisions · Speed 93.6.Raw Qwen3 1.7B direct logits · JevBench rank 63 · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · JevBench rank 67 · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · JevBench rank 73 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · JevBench rank 89 · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · JevBench rank 81 · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · JevBench rank 84 · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #35
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #83
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #31
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #61
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #66
91 systems plotted. Hover or focus a point to read its values.

Capability vs speed

The upper-right is the more attractive area: higher Capability and higher Speed.

Capability vs speedThe upper-right is the more attractive area: higher Capability and higher Speed. Each point has a tooltip with the system and its values.020406080100020406080100Speed axis · higher is faster →Capability · higher ↑GPT-6 Luna (medium) · JevBench rank 35 · Capability 95.4 · Cost $0.14 per 1,000 decisions · Speed 72.6.DeepSeek V4.1 Flash · JevBench rank 83 · Capability 94.7 · Cost $0.59 per 1,000 decisions · Speed 71.6.GPT-6 Luna (low) · JevBench rank 31 · Capability 93.9 · Cost $0.13 per 1,000 decisions · Speed 73.7.GPT-5.6 Luna · JevBench rank 61 · Capability 90.3 · Cost $0.24 per 1,000 decisions · Speed 77.5.djev · JevBench rank 66 · Capability 79.7 · Cost $0.27 estimated per 1,000 decisions · Speed 75.2.Kushal Patil — Gemma 4 31B IT · JevBench rank 42 · Capability 69.5 · Cost $0.14 per 1,000 decisions · Speed 84.0.Qwen3.8 27B · Capability 67.0 · Cost $2.67 estimated per 1,000 decisions · Speed 61.3.Instinct · JevBench rank 74 · Capability 64.9 · Cost $0.33 estimated per 1,000 decisions · Speed 83.9.Jev 1.13.0 · JevBench rank 2 · Capability 64.7 · Cost $0.040 per 1,000 decisions · Speed 83.3.NInfer Qwen3.8-Flash-Next mixed · JevBench rank 33 · Capability 64.1 · Cost $0.11 estimated per 1,000 decisions · Speed 88.2.NInfer Qwen3.8-27B NVFP4 · JevBench rank 47 · Capability 63.7 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.Hopper · JevBench rank 5 · Capability 63.5 · Cost $0.024 estimated per 1,000 decisions · Speed 86.8.SimpleJev Qwen3.8-27B · JevBench rank 32 · Capability 63.0 · Cost $0.10 estimated per 1,000 decisions · Speed 71.2.JevOne · JevBench rank 50 · Capability 62.6 · Cost $0.14 estimated per 1,000 decisions · Speed 88.5.Cygnet · JevBench rank 4 · Capability 62.2 · Cost $0.037 estimated per 1,000 decisions · Speed 90.7.decider-4b v2 · JevBench rank 1 · Capability 62.2 · Cost $0.020 estimated per 1,000 decisions · Speed 92.9.classifier.dev · Capability 62.0 · Cost $0.0033 estimated per 1,000 decisions · Speed 87.6.reflex-27b · JevBench rank 64 · Capability 61.8 · Cost $0.18 estimated per 1,000 decisions · Speed 67.5.swanOne · JevBench rank 34 · Capability 61.8 · Cost $0.11 estimated per 1,000 decisions · Speed 82.5.JevK5 v0.2.0 · JevBench rank 3 · Capability 61.7 · Cost $0.022 estimated per 1,000 decisions · Speed 91.1.LitJev · JevBench rank 57 · Capability 61.4 · Cost $0.16 estimated per 1,000 decisions · Speed 66.7.openjev-sglang · JevBench rank 45 · Capability 59.3 · Cost $0.13 estimated per 1,000 decisions · Speed 77.1.NInfer Qwen3.8-27B NVFP4 · JevBench rank 48 · Capability 59.3 · Cost $0.14 estimated per 1,000 decisions · Speed 80.1.jqv · JevBench rank 17 · Capability 59.0 · Cost $0.056 estimated per 1,000 decisions · Speed 74.6.reflex 4B · JevBench rank 7 · Capability 58.9 · Cost $0.022 estimated per 1,000 decisions · Speed 68.0.ZeroEntropy zerank-2 · JevBench rank 22 · Capability 58.9 · Cost $0.047 per 1,000 decisions · Speed 79.0.local-jev Qwen3.5-4B · JevBench rank 13 · Capability 58.9 · Cost $0.030 estimated per 1,000 decisions · Speed 75.0.OpenJev · JevBench rank 68 · Capability 58.1 · Cost $0.25 estimated per 1,000 decisions · Speed 76.1.JEV Qwen3.5-9B Base NVFP4 · JevBench rank 26 · Capability 57.3 · Cost $0.077 estimated per 1,000 decisions · Speed 93.3.Gemini 3.1 Flash-Lite · JevBench rank 70 · Capability 56.9 · Cost $0.26 per 1,000 decisions · Speed 81.8.Winnow-12B Q8 · JevBench rank 6 · Capability 56.6 · Cost $0.037 estimated per 1,000 decisions · Speed 82.3.Decision 2B · JevBench rank 29 · Capability 56.4 · Cost $0.018 estimated per 1,000 decisions · Speed 84.3.decider-35b-a3b · JevBench rank 19 · Capability 56.2 · Cost $0.067 estimated per 1,000 decisions · Speed 80.8.Standard One 8B · JevBench rank 43 · Capability 56.0 · Cost $0.10 estimated per 1,000 decisions · Speed 92.0.metask-jev-4b · JevBench rank 10 · Capability 55.8 · Cost $0.033 estimated per 1,000 decisions · Speed 89.1.SemIf · JevBench rank 11 · Capability 55.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.7.Jev-Omni · JevBench rank 9 · Capability 55.4 · Cost $0.037 estimated per 1,000 decisions · Speed 81.5.Von · JevBench rank 46 · Capability 55.1 · Cost $0.0055 estimated per 1,000 decisions · Speed 70.5.Jobe Qwen3.5-4B · JevBench rank 12 · Capability 55.1 · Cost $0.022 estimated per 1,000 decisions · Speed 85.6.Qwen3-Reranker-4B · JevBench rank 18 · Capability 54.9 · Cost $0.050 per 1,000 decisions · Speed 78.7.decision-machine-1 · JevBench rank 23 · Capability 54.8 · Cost $0.035 per 1,000 decisions · Speed 92.9.jev-local · JevBench rank 37 · Capability 54.7 · Cost $0.077 estimated per 1,000 decisions · Speed 69.2.Qwen3.5-9B Jev-like data-mix v2 · JevBench rank 30 · Capability 54.4 · Cost $0.083 estimated per 1,000 decisions · Speed 82.0.Open-Jev 9B · JevBench rank 75 · Capability 53.0 · Cost $0.25 estimated per 1,000 decisions · Speed 72.0.SimpleJev Qwen3.6-35B-A3B · JevBench rank 52 · Capability 52.8 · Cost $0.12 estimated per 1,000 decisions · Speed 75.0.lev-350m · JevBench rank 44 · Capability 52.7 · Cost $0.0063 estimated per 1,000 decisions · Speed 85.3.jeff · JevBench rank 40 · Capability 52.3 · Cost $0.0060 estimated per 1,000 decisions · Speed 63.5.Malkuth-4B · JevBench rank 16 · Capability 52.3 · Cost $0.019 estimated per 1,000 decisions · Speed 88.3.Bespoke Nimble 9B · JevBench rank 60 · Capability 51.4 · Cost $0.17 estimated per 1,000 decisions · Speed 78.7.Decision Fast · JevBench rank 38 · Capability 51.2 · Cost $0.0063 estimated per 1,000 decisions · Speed 81.6.djev · JevBench rank 8 · Capability 51.2 · Cost $0.026 per 1,000 decisions · Speed 91.4.OpenSourceJev · JevBench rank 21 · Capability 51.1 · Cost $0.016 estimated per 1,000 decisions · Speed 82.0.openJev Verdict 1.4 · JevBench rank 58 · Capability 50.7 · Cost $0.0039 estimated per 1,000 decisions · Speed 78.1.Raw Phi-4 mini direct logits · JevBench rank 25 · Capability 50.3 · Cost $0.048 estimated per 1,000 decisions · Speed 88.8.OpenJev · JevBench rank 27 · Capability 50.2 · Cost $0.066 estimated per 1,000 decisions · Speed 83.2.Laya · JevBench rank 41 · Capability 49.9 · Cost $0.0029 estimated per 1,000 decisions · Speed 71.1.system-one-open · JevBench rank 14 · Capability 49.5 · Cost $0.015 estimated per 1,000 decisions · Speed 77.0.Open-Jev 2B · JevBench rank 76 · Capability 48.8 · Cost $0.25 estimated per 1,000 decisions · Speed 73.5.open-alternative-jev · JevBench rank 36 · Capability 48.6 · Cost $0.022 estimated per 1,000 decisions · Speed 83.5.Qwen3.5-0.8B Decision Model · JevBench rank 69 · Capability 48.2 · Cost $0.0065 estimated per 1,000 decisions · Speed 49.2.typecastlm · JevBench rank 51 · Capability 48.0 · Cost $0.021 estimated per 1,000 decisions · Speed 92.6.Malkuth-2B · JevBench rank 24 · Capability 47.4 · Cost $0.019 estimated per 1,000 decisions · Speed 91.4.spark-s1-4b-v6 · JevBench rank 15 · Capability 46.3 · Cost $0.025 estimated per 1,000 decisions · Speed 81.0.open-jev-deberta-v3-large · JevBench rank 71 · Capability 46.1 · Cost $0.0073 estimated per 1,000 decisions · Speed 66.0.Mixedbread mxbai-rerank-base-v2 · JevBench rank 85 · Capability 45.5 · Cost $0.012 per 1,000 decisions · Speed 87.5.BAAI bge-reranker-v2-m3 · JevBench rank 86 · Capability 44.6 · Cost $0.0077 per 1,000 decisions · Speed 89.5.OpenDecision · JevBench rank 56 · Capability 44.4 · Cost $0.0066 estimated per 1,000 decisions · Speed 79.9.verdict-small · JevBench rank 82 · Capability 42.5 · Cost $0.0013 estimated per 1,000 decisions · Speed 85.0.smalljev semantic-v9 · JevBench rank 72 · Capability 42.4 · Cost $0.025 estimated per 1,000 decisions · Speed 79.8.JevAct · JevBench rank 65 · Capability 42.2 · Cost $0.015 estimated per 1,000 decisions · Speed 76.5.kev 0.6B · JevBench rank 53 · Capability 42.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 75.6.Alibaba GTE Reranker ModernBERT-base · JevBench rank 87 · Capability 41.9 · Cost $0.010 per 1,000 decisions · Speed 90.6.Certo v1 · JevBench rank 88 · Capability 41.5 · Cost $0.00097 estimated per 1,000 decisions · Speed 94.0.decider-2b · JevBench rank 39 · Capability 41.0 · Cost $0.020 estimated per 1,000 decisions · Speed 83.2.kev 8B · JevBench rank 49 · Capability 41.0 · Cost $0.073 estimated per 1,000 decisions · Speed 74.9.kev 4B · JevBench rank 28 · Capability 40.9 · Cost $0.019 estimated per 1,000 decisions · Speed 75.7.GLiNER2.5 multi · JevBench rank 77 · Capability 40.1 · Cost $0.0039 estimated per 1,000 decisions · Speed 67.8.kev 0.5B · JevBench rank 59 · Capability 40.1 · Cost $0.0063 estimated per 1,000 decisions · Speed 77.0.openJev Verdict · JevBench rank 62 · Capability 38.5 · Cost $0.0037 estimated per 1,000 decisions · Speed 76.7.system-one · JevBench rank 55 · Capability 38.2 · Cost $0.089 estimated per 1,000 decisions · Speed 84.4.Raw Qwen3 4B Instruct 2507 direct logits · JevBench rank 20 · Capability 37.7 · Cost $0.022 estimated per 1,000 decisions · Speed 87.6.GLiNER2.5 small · JevBench rank 80 · Capability 35.6 · Cost $0.0039 estimated per 1,000 decisions · Speed 77.8.SimpleJev · JevBench rank 79 · Capability 35.3 · Cost $0.011 estimated per 1,000 decisions · Speed 57.5.Raw Qwen3 8B direct logits · JevBench rank 54 · Capability 34.9 · Cost $0.087 estimated per 1,000 decisions · Speed 86.3.CLM-8B · JevBench rank 78 · Capability 31.1 · Cost $0.0052 estimated per 1,000 decisions · Speed 93.6.Raw Qwen3 1.7B direct logits · JevBench rank 63 · Capability 28.8 · Cost $0.015 estimated per 1,000 decisions · Speed 89.7.GLiNER2 large · JevBench rank 67 · Capability 28.0 · Cost $0.0077 estimated per 1,000 decisions · Speed 61.7.GLiNER2 · JevBench rank 73 · Capability 26.3 · Cost $0.0037 estimated per 1,000 decisions · Speed 71.8.Open Jev JSON Canvas · JevBench rank 89 · Capability 24.1 · Cost $0.065 estimated per 1,000 decisions · Speed 84.1.Raw Qwen3 0.6B direct logits · JevBench rank 81 · Capability 21.8 · Cost $0.0074 estimated per 1,000 decisions · Speed 89.9.Mirror · JevBench rank 84 · Capability 19.8 · Cost $0.0077 estimated per 1,000 decisions · Speed 70.8.
  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #35
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #83
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #31
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #61
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #66
91 systems plotted. Hover or focus a point to read its values.

All three at once

The 3D view plots Capability vertically, lower cost to the right, and higher Speed toward you. Sphere size follows the JevBench score. Drag to rotate; pinch or scroll to zoom. The view loads when it scrolls into view.

Scroll here to load the interactive 3D view.

  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 decisions · Speed 72.6 · JevBench #35
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 decisions · Speed 71.6 · JevBench #83
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 decisions · Speed 73.7 · JevBench #31
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 decisions · Speed 77.5 · JevBench #61
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 / 1,000 decisions · Speed 75.2 · JevBench #66

Vertical: Capability · Right: cheaper · Toward you: faster

The interactive 3D view loads when this panel scrolls into view.

91 systems plotted; systems missing cost or Speed are omitted.

JevBench v1.4.1 public split · input capacity and long inputs

Context length

Context length is how much input a system accepts in one request; a smaller window forces truncation or chunking.

82 rows · published limits 512 to 1,050,000 tokens · 8 without a published maximum · sources checked 24 Sept 2026

Public accuracy by actual input length

13 of the top 15 systems are plotted; JevK5 v0.2.0 (per-item accuracy 199/231 does not match published v1.4.1) and system-one-open (no per-item token counts) are not.

JevBench public accuracy across seven input-length buckets13 systems are plotted across <2k, 2–8k, 8–16k, 16–64k, 64–256k, 256k–1M, ≥1M input-token buckets. Bucket denominators differ by system and are available in the details table below.0%25%50%75%100%<2k2–8k8–16k16–64kActual input tokens per decision#1 Jev 1.13.0 (TypeSafe AI) · <2k tokens: 85.6% (125/146)#1 Jev 1.13.0 (TypeSafe AI) · 2–8k tokens: 73.0% (27/37)#3 Hopper · <2k tokens: 85.6% (167/195)#3 Hopper · 2–8k tokens: 63.9% (23/36)#4 Winnow-12B Q8 · <2k tokens: 87.6% (170/194)#4 Winnow-12B Q8 · 2–8k tokens: 75.7% (28/37)#5 reflex 4B (kshetrajna12) · <2k tokens: 83.1% (162/195)#5 reflex 4B (kshetrajna12) · 2–8k tokens: 58.3% (21/36)#6 djev (Maisa, diffusion-gemma) · <2k tokens: 87.1% (169/194)#6 djev (Maisa, diffusion-gemma) · 2–8k tokens: 67.6% (25/37)#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged) · <2k tokens: 90.2% (175/194)#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged) · 2–8k tokens: 81.1% (30/37)#8 metask-jev-4b · <2k tokens: 85.1% (165/194)#8 metask-jev-4b · 2–8k tokens: 52.8% (19/36)#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · <2k tokens: 85.1% (166/195)#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · 2–8k tokens: 58.3% (21/36)#10 Jobe Qwen3.5-4B (frozen) · <2k tokens: 85.1% (166/195)#10 Jobe Qwen3.5-4B (frozen) · 2–8k tokens: 58.3% (21/36)#11 local-jev Qwen3.5-4B · <2k tokens: 83.6% (163/195)#11 local-jev Qwen3.5-4B · 2–8k tokens: 63.9% (23/36)#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085) · <2k tokens: 83.0% (161/194)#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085) · 2–8k tokens: 59.5% (22/37)#14 jqv (Qwen3-32B zero-shot) · <2k tokens: 85.6% (167/195)#14 jqv (Qwen3-32B zero-shot) · 2–8k tokens: 50.0% (18/36)#15 Qwen3-Reranker-4B · <2k tokens: 75.7% (134/177)#15 Qwen3-Reranker-4B · 2–8k tokens: 29.2% (7/24)#15 Qwen3-Reranker-4B · 8–16k tokens: 48.1% (13/27)#15 Qwen3-Reranker-4B · 16–64k tokens: 100.0% (3/3) · n = 3
  • Top five · #1 Jev 1.13.0 (TypeSafe AI)
  • Top five · #3 Hopper
  • Top five · #4 Winnow-12B Q8
  • Top five · #5 reflex 4B (kshetrajna12)
  • Top five · #6 djev (Maisa, diffusion-gemma)
  • #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)
  • #8 metask-jev-4b
  • #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
  • #10 Jobe Qwen3.5-4B (frozen)
  • #11 local-jev Qwen3.5-4B
  • #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)
  • #14 jqv (Qwen3-32B zero-shot)
  • #15 Qwen3-Reranker-4B

No public item exceeds 64k input tokens; the 64–256k, 256k–1M and ≥1M buckets are empty. A hollow point comes from fewer than 20 decisions and is not joined to the line.

From each run's recorded input-token counts; no new runs. The chart describes these benchmark items; it does not show that context length alone caused a score change.
Exact correct counts and denominators by bucket

All 13 plotted systems exactly reproduce their published public accuracy; stored lengths cover 183–231 decisions per system.

SystemMean input tokensLength n<2k2–8k8–16k16–64k64–256k256k–1M≥1M
#1 Jev 1.13.0 (TypeSafe AI)1,057.8183/231125/146 · 85.6%27/37 · 73.0%No dataNo dataNo dataNo dataNo data
#3 Hopper739.2231/231167/195 · 85.6%23/36 · 63.9%No dataNo dataNo dataNo dataNo data
#4 Winnow-12B Q8692.3231/231170/194 · 87.6%28/37 · 75.7%No dataNo dataNo dataNo dataNo data
#5 reflex 4B (kshetrajna12)696.8231/231162/195 · 83.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#6 djev (Maisa, diffusion-gemma)692.3231/231169/194 · 87.1%25/37 · 67.6%No dataNo dataNo dataNo dataNo data
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)686.5231/231175/194 · 90.2%30/37 · 81.1%No dataNo dataNo dataNo dataNo data
#8 metask-jev-4b761.2230/231165/194 · 85.1%19/36 · 52.8%No dataNo dataNo dataNo dataNo data
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)698.6231/231166/195 · 85.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#10 Jobe Qwen3.5-4B (frozen)698.6231/231166/195 · 85.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#11 local-jev Qwen3.5-4B696.7231/231163/195 · 83.6%23/36 · 63.9%No dataNo dataNo dataNo dataNo data
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)794.6231/231161/194 · 83.0%22/37 · 59.5%No dataNo dataNo dataNo dataNo data
#14 jqv (Qwen3-32B zero-shot)662.7231/231167/195 · 85.6%18/36 · 50.0%No dataNo dataNo dataNo dataNo data
#15 Qwen3-Reranker-4B2,401.9231/231134/177 · 75.7%7/24 · 29.2%13/27 · 48.1%3/3 · 100.0%No dataNo dataNo data

Long-policy tasks show a separate stress point

Across 19 public items in the long-policy family, several systems scored well below their full public-set accuracy. The comparison uses the family label, not only the token buckets.

  • metask-jev-4b
    Long policy 6/19 (31.6%) vs 184/231 overall (79.7%): −48.1 pp.
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085)
    Long policy 7/19 (36.8%) vs 183/231 overall (79.2%): −42.4 pp.
  • system-one-open (Gemma 4 E2B LoRA on an L4)
    Long policy 6/19 (31.6%) vs 169/231 overall (73.2%): −41.6 pp.
  • Winnow-12B Q8
    Long policy 15/19 (78.9%) vs 198/231 overall (85.7%): −6.8 pp.
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged)
    Long policy 15/19 (78.9%) vs 205/231 overall (88.7%): −9.8 pp.
All 14 matched systems on long-policy tasks
SystemOverallLong policy (19 items)Change
#8 metask-jev-4b184/231 · 79.7%6/19 · 31.6%−48.1 pp
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)183/231 · 79.2%7/19 · 36.8%−42.4 pp
#12 system-one-open (Gemma 4 E2B LoRA on an L4)169/231 · 73.2%6/19 · 31.6%−41.6 pp
#14 jqv (Qwen3-32B zero-shot)185/231 · 80.1%9/19 · 47.4%−32.7 pp
#6 djev (Maisa, diffusion-gemma)194/231 · 84.0%10/19 · 52.6%−31.4 pp
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#10 Jobe Qwen3.5-4B (frozen)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#5 reflex 4B (kshetrajna12)183/231 · 79.2%10/19 · 52.6%−26.6 pp
#3 Hopper190/231 · 82.3%11/19 · 57.9%−24.4 pp
#1 Jev 1.13.0 (TypeSafe AI)200/231 · 86.6%12/19 · 63.2%−23.4 pp
#11 local-jev Qwen3.5-4B186/231 · 80.5%12/19 · 63.2%−17.4 pp
#15 Qwen3-Reranker-4B157/231 · 68.0%10/19 · 52.6%−15.3 pp
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)205/231 · 88.7%15/19 · 78.9%−9.8 pp
#4 Winnow-12B Q8198/231 · 85.7%15/19 · 78.9%−6.8 pp

For example, metask-jev-4b scored 31.6% on long-policy tasks versus 79.7% overall (change −48.1 pp), while Winnow-12B Q8 scored 78.9% versus 85.7% (change −6.8 pp). These are descriptive public-set comparisons. Prompt wrappers and tokenizers differ by system, and the 19-item family is small, so the results do not isolate context length as the cause. Only public item results were used; sealed-set item rows were not used.

Published context limits · logarithmic scale

77 ranked systems and 5 unranked additions, each with its exact published maximum input context. API/serving caps, hard limits and trained lengths use different bar colors; training configuration limits appear as separate markers and values where published.

  1. #1 Jev 1.13.0 (TypeSafe AI)64,000 · API cap
  2. #2 JevK5 v0.2.0262,144 · Trained lengthTraining configuration: trained sequence length 2,048 tokens.
  3. #3 Hopper262,144 · Trained length
  4. #4 Winnow-12B Q865,536 · API cap
  5. #5 reflex 4B (kshetrajna12)262,144 · Trained length
  6. #6 djev (Maisa, diffusion-gemma)32,768 · API cap
  7. #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)262,144 · Trained length
  8. #8 metask-jev-4b262,144 · Trained length
  9. #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)262,144 · Trained length
  10. #10 Jobe Qwen3.5-4B (frozen)262,144 · Trained lengthTraining configuration: trained sequence length 4,096 tokens.
  11. #11 local-jev Qwen3.5-4B262,144 · Trained length
  12. #12 system-one-open (Gemma 4 E2B LoRA on an L4)131,072 · Trained lengthTraining configuration: training state limit 2,048 tokens.
  13. #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)262,144 · Trained lengthTraining configuration: trained sequence length 2,048 tokens.
  14. #14 jqv (Qwen3-32B zero-shot)32,768 · Trained length
  15. #15 Qwen3-Reranker-4B32,768 · Trained length
  16. #16 decider-35b-a3b (Mapika)262,144 · Hard limit
  17. #17 Raw Qwen3 4B Instruct 2507 direct logits262,144 · Trained length
  18. #18 OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)32,768 · Trained length
  19. #19 ZeroEntropy zerank-232,768 · Trained length
  20. #20 decision-machine-1 (milliseconds.ai)Unknown
  21. #21 Raw Phi-4 mini direct logits131,072 · Trained length
  22. #22 JEV Qwen3.5-9B Base NVFP4262,144 · Trained length
  23. #23 OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)65,536 · API cap
  24. #24 kev 4B (research preview)32,768 · Trained length
  25. #25 Decision 2B (FlyMy.AI, v59)131,072 · Trained length
Show all 82 systems (57 more)
  1. #26 Qwen3.5-9B Jev-like data-mix v2262,144 · Trained length
  2. #27 GPT-6 Luna (low reasoning effort)1,050,000 · API cap
  3. #28 SimpleJev Qwen3.8-27B2,000 · API cap
  4. #29 NInfer Qwen3.8-Flash-Next mixed262,144 · Trained length
  5. #30 GPT-6 Luna (default medium reasoning effort)1,050,000 · API cap
  6. #31 open-alternative-jev (Qwen3.5-4B, IkerMoel)262,144 · Trained length
  7. #32 jev-local (Qwen3.5-9B)262,144 · Trained length
  8. #33 Decision Fast (FlyMy.AI, v53a)32,768 · Trained length
  9. #34 decider-2b (Mapika)262,144 · Hard limit
  10. #35 jeff (Logan Markewich, GLiFormer 400M)8,192 · Hard limit
  11. #36 Laya (Convai Innovations, ModernBERT-large 421M)512 · Hard limit
  12. #37 lev-350m (Franck Verrot, LFM2.5-350M)32,768 · Trained length
  13. #38 openjev-sglang (Qwen3.6-35B-A3B on SGLang)32,768 · API cap
  14. #39 Von (wfzyx, Option-Marker 395M)8,192 · Trained length
  15. #40 NInfer Qwen3.8-27B NVFP4 (T=1.5)262,144 · Trained length
  16. #41 NInfer Qwen3.8-27B NVFP4262,144 · Trained length
  17. #42 kev 8B (research preview)32,768 · Trained length
  18. #43 JevOne262,144 · Trained length
  19. #44 SimpleJev Qwen3.6-35B-A3B262,144 · Trained length
  20. #45 kev 0.6B (research preview)32,768 · Trained length
  21. #46 Raw Qwen3 8B direct logits32,768 · Trained length
  22. #47 system-one (Qwen3-8B, Sean Goedecke)32,768 · Trained length
  23. #48 OpenDecision (ModernBERT-large zero-shot)8,192 · Hard limit
  24. #49 LitJev (Qwen3.8-27B)262,144 · Trained length
  25. #50 openJev Verdict 1.4512 · API cap
  26. #51 kev 0.5B8,192 · API cap
  27. #52 Bespoke Nimble 9B (Bespoke Labs)8,192 · API capTraining configuration: trained sequence length 2,048 tokens.
  28. #53 GPT-5.6 Luna (low reasoning effort)1,050,000 · API cap
  29. #54 openJev Verdict (heman10x, ModernBERT-base 151M)8,192 · Hard limit
  30. #55 Raw Qwen3 1.7B direct logits32,768 · Trained length
  31. #56 reflex-27b (Qwen3.8-27B)262,144 · Trained length
  32. #57 djev (thinking)262,144 · Trained length
  33. #58 GLiNER2 large (Fastino)Unknown
  34. #59 OpenJev (thinking, BF16)65,536 · API cap
  35. #60 Qwen3.5-0.8B Decision Model (Mourad Ghafiri)262,144 · Hard limit
  36. #61 Gemini 3.1 Flash-Lite1,048,576 · API cap
  37. #62 open-jev-deberta-v3-large (local CPU)512 · Hard limit
  38. #63 smalljev semantic-v9131,072 · Trained length
  39. #64 GLiNER2 (Fastino, gliner2.5-base)Unknown
  40. #65 Open-Jev 9B (Zefan Cai)262,144 · Trained length
  41. #66 Open-Jev 2B (Zefan Cai)262,144 · Trained length
  42. #67 GLiNER2.5 multi (Fastino, 287M)Unknown
  43. #68 SimpleJev (Qwen3.5-0.8B, CPU)262,144 · Trained length
  44. #69 GLiNER2.5 small (Fastino, 74M)Unknown
  45. #70 Raw Qwen3 0.6B direct logits32,768 · Trained length
  46. #71 DeepSeek V4.1 Flash (thinking default)1,000,000 · API cap
  47. #72 Mirror512 · Hard limit
  48. #73 Mixedbread mxbai-rerank-base-v232,768 · Hard limit
  49. #74 BAAI bge-reranker-v2-m38,192 · Hard limit
  50. #75 Alibaba GTE Reranker ModernBERT-base8,192 · Hard limit
  51. #76 Certo v1 (AltSlate Labs)8,192 · Hard limit
  52. #77 Open Jev JSON Canvas (JoshuaSP)262,144 · Trained length
  53. Unranked classifier.dev (fast tier)Unknown
  54. Unranked Needle 3 (Cactus, 2-bit, local CPU)Unknown
  55. Unranked Needle 3, options as tools (post-hoc adapter mode)Unknown
  56. Unranked Qwen3.8 27B (Chutes TEE)262,144 · API cap
  57. Unranked swanOne262,144 · Trained length
  • API / serving cap
  • Hard limit
  • Trained length
  • Trained sequence length
  • Training state limit
The scale runs from 512 tokens to 1,050,000 tokens. Source links, dates, evidence notes and exact training details remain in the table below.

Context limits by system

82 systems · sources checked 24 Sept 2026
All 82 limits as a table
#1Jev 1.13.0 (TypeSafe AI)†64,000Open primary sourceAPI cap
#2JevK5 v0.2.0†262,144Open primary sourceTrained length
#3Hopper†262,144Open primary sourceTrained length
#4Winnow-12B Q8†65,536Open primary sourceAPI cap
#5reflex 4B (kshetrajna12)†262,144Open primary sourceTrained length
#6djev (Maisa, diffusion-gemma)†32,768Open primary sourceAPI cap
#7Jev-Omni (akhilaaa3, Gemma-4-12B merged)262,144Open primary sourceTrained length
#8metask-jev-4b†262,144Open primary sourceTrained length
#9SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)262,144Open primary sourceTrained length
#10Jobe Qwen3.5-4B (frozen)†262,144Open primary sourceTrained length
#11local-jev Qwen3.5-4B†262,144Open primary sourceTrained length
#12system-one-open (Gemma 4 E2B LoRA on an L4)†131,072Open primary sourceTrained length
#13spark-s1-4b-v6 (Open Spark Jev, abhishek085)†262,144Open primary sourceTrained length
#14jqv (Qwen3-32B zero-shot)32,768Open primary sourceTrained length
#15Qwen3-Reranker-4B†32,768Open primary sourceTrained length
#16decider-35b-a3b (Mapika)262,144Open primary sourceHard limit
#17Raw Qwen3 4B Instruct 2507 direct logits262,144Open primary sourceTrained length
#18OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)32,768Open primary sourceTrained length
#19ZeroEntropy zerank-232,768Open primary sourceTrained length
#20decision-machine-1 (milliseconds.ai)UnknownOpen primary sourceUnknown
#21Raw Phi-4 mini direct logits131,072Open primary sourceTrained length
#22JEV Qwen3.5-9B Base NVFP4262,144Open primary sourceTrained length
#23OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)†65,536Open primary sourceAPI cap
#24kev 4B (research preview)32,768Open primary sourceTrained length
#25Decision 2B (FlyMy.AI, v59)131,072Open primary sourceTrained length
#26Qwen3.5-9B Jev-like data-mix v2262,144Open primary sourceTrained length
#27GPT-6 Luna (low reasoning effort)†1,050,000Open primary sourceAPI cap
#28SimpleJev Qwen3.8-27B†2,000Open primary sourceAPI cap
#29NInfer Qwen3.8-Flash-Next mixed262,144Open primary sourceTrained length
#30GPT-6 Luna (default medium reasoning effort)†1,050,000Open primary sourceAPI cap
#31open-alternative-jev (Qwen3.5-4B, IkerMoel)262,144Open primary sourceTrained length
#32jev-local (Qwen3.5-9B)262,144Open primary sourceTrained length
#33Decision Fast (FlyMy.AI, v53a)32,768Open primary sourceTrained length
#34decider-2b (Mapika)262,144Open primary sourceHard limit
#35jeff (Logan Markewich, GLiFormer 400M)8,192Open primary sourceHard limit
#36Laya (Convai Innovations, ModernBERT-large 421M)512Open primary sourceHard limit
#37lev-350m (Franck Verrot, LFM2.5-350M)†32,768Open primary sourceTrained length
#38openjev-sglang (Qwen3.6-35B-A3B on SGLang)†32,768Open primary sourceAPI cap
#39Von (wfzyx, Option-Marker 395M)†8,192Open primary sourceTrained length
#40NInfer Qwen3.8-27B NVFP4 (T=1.5)262,144Open primary sourceTrained length
#41NInfer Qwen3.8-27B NVFP4262,144Open primary sourceTrained length
#42kev 8B (research preview)32,768Open primary sourceTrained length
#43JevOne†262,144Open primary sourceTrained length
#44SimpleJev Qwen3.6-35B-A3B262,144Open primary sourceTrained length
#45kev 0.6B (research preview)32,768Open primary sourceTrained length
#46Raw Qwen3 8B direct logits32,768Open primary sourceTrained length
#47system-one (Qwen3-8B, Sean Goedecke)32,768Open primary sourceTrained length
#48OpenDecision (ModernBERT-large zero-shot)†8,192Open primary sourceHard limit
#49LitJev (Qwen3.8-27B)262,144Open primary sourceTrained length
#50openJev Verdict 1.4†512Open primary sourceAPI cap
#51kev 0.5B†8,192Open primary sourceAPI cap
#52Bespoke Nimble 9B (Bespoke Labs)†8,192Open primary sourceAPI cap
#53GPT-5.6 Luna (low reasoning effort)†1,050,000Open primary sourceAPI cap
#54openJev Verdict (heman10x, ModernBERT-base 151M)†8,192Open primary sourceHard limit
#55Raw Qwen3 1.7B direct logits32,768Open primary sourceTrained length
#56reflex-27b (Qwen3.8-27B)262,144Open primary sourceTrained length
#57djev (thinking)†262,144Open primary sourceTrained length
#58GLiNER2 large (Fastino)†UnknownOpen primary sourceUnknown
#59OpenJev (thinking, BF16)†65,536Open primary sourceAPI cap
#60Qwen3.5-0.8B Decision Model (Mourad Ghafiri)262,144Open primary sourceHard limit
#61Gemini 3.1 Flash-Lite†1,048,576Open primary sourceAPI cap
#62open-jev-deberta-v3-large (local CPU)512Open primary sourceHard limit
#63smalljev semantic-v9131,072Open primary sourceTrained length
#64GLiNER2 (Fastino, gliner2.5-base)†UnknownOpen primary sourceUnknown
#65Open-Jev 9B (Zefan Cai)262,144Open primary sourceTrained length
#66Open-Jev 2B (Zefan Cai)262,144Open primary sourceTrained length
#67GLiNER2.5 multi (Fastino, 287M)†UnknownOpen primary sourceUnknown
#68SimpleJev (Qwen3.5-0.8B, CPU)262,144Open primary sourceTrained length
#69GLiNER2.5 small (Fastino, 74M)†UnknownOpen primary sourceUnknown
#70Raw Qwen3 0.6B direct logits32,768Open primary sourceTrained length
#71DeepSeek V4.1 Flash (thinking default)†1,000,000Open primary sourceAPI cap
#72Mirror512Open primary sourceHard limit
#73Mixedbread mxbai-rerank-base-v232,768Open primary sourceHard limit
#74BAAI bge-reranker-v2-m38,192Open primary sourceHard limit
#75Alibaba GTE Reranker ModernBERT-base8,192Open primary sourceHard limit
#76Certo v1 (AltSlate Labs)8,192Open primary sourceHard limit
#77Open Jev JSON Canvas (JoshuaSP)262,144Open primary sourceTrained length
Unrankedclassifier.dev (fast tier)UnknownOpen primary sourceUnknown
UnrankedNeedle 3 (Cactus, 2-bit, local CPU)UnknownOpen primary sourceUnknown
UnrankedNeedle 3, options as tools (post-hoc adapter mode)UnknownOpen primary sourceUnknown
UnrankedQwen3.8 27B (Chutes TEE)†262,144Open primary sourceAPI cap
UnrankedswanOne†262,144Open primary sourceTrained length
Notes for 82 systems
  • Jev 1.13.0 (TypeSafe AI) — TypeSafe's official Jev 1.13 docs specify a 64K request budget and a 32K state-plus-longest-question budget. Published as “64,000 total request; 32,000 state + longest question”. System repository
  • JevK5 v0.2.0 — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Trained sequence length: 2,048 tokens. Training utility defaults --max-len to 2,048 and skips longer rows. Runtime has no smaller total context cap documented; native Qwen3.5-4B window is 262,144. System repository · 23 Sept 2026 Training configuration · 23 Sept 2026
  • Hopper — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. JevBench registry identifies this as a LoRA on Qwen3.5-4B. Its public model card does not specify a shorter max_seq_len or serving truncation, so the base model's native 262,144 window is listed; adapter training length is unknown. System repository · 22 Sept 2026
  • Winnow-12B Q8 — Base model google/gemma-4-12B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The Winnow model card documents a 65,536-position configured Q8 test profile; Gemma 4 12B base supports 262,144. Published as “65,536 configured/tested; base 262,144”. System repository · 21 Sept 2026 Base model source · 20 Jul 2026
  • reflex 4B (kshetrajna12) — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Server refuses an input above the base model's context window rather than truncating it. The repo's 8,192 max-pack-tokens is a batching budget, not the per-request context cap. System repository · 23 Sept 2026
  • djev (Maisa, diffusion-gemma) — Base model google/diffusiongemma-26B-A4B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Runtime setting is bounded to 1,024–32,768; DiffusionGemma base is 262,144. Published as “32,768 (prompt + reserved canvas)”. System repository · 19 Sept 2026 Base model source · 15 Jul 2026
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged) — Base model google/gemma-4-12B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Gemma 4 12B card states a 256K context window. System repository · 22 Sept 2026
  • metask-jev-4b — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Reported validation point: 4,096 tokens; that is not automatically the maximum accepted input. Model card reports 4,096 as the validated evaluation point, not an architectural limit; the same card/config says native 262,144. Published validated evaluation point: 4,096 tokens. System repository · 22 Sept 2026 Training configuration · 22 Sept 2026
  • SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 23 Sept 2026
  • Jobe Qwen3.5-4B (frozen) — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Trained sequence length: 4,096 tokens. The optional adapter-training helper uses max_tokens=4,096 and rejects longer training examples. The ranked v1.4.1 entry is the frozen Qwen3.5-4B backbone, so this optional training helper does not set its inference window. System repository · 23 Sept 2026 Training configuration · 23 Sept 2026
  • local-jev Qwen3.5-4B — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. The README documents long-state shortening. Its 32,768-token context_tokens value is only an illustrative custom-model card; the effective Qwen3.5 deployment cap is not published. The listed 262,144 is the base model window. System repository · 21 Sept 2026
  • system-one-open (Gemma 4 E2B LoRA on an L4) — Base model google/gemma-4-E2B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Gemma 4 E2B card/config state 131,072 (128K) positions. Training state limit: 2,048 tokens. Training batch token budget: 24,576 tokens. Full training profile caps state at 2,048 tokens and total batch tokens at 24,576; this is a training profile, not an inference limit. System repository · 17 Sept 2026 Training configuration · 17 Sept 2026
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085) — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Trained sequence length: 2,048 tokens. Published RLCD configs use max_len=2,048 for training; Qwen3.5-4B native inference window is 262,144. System repository · 22 Sept 2026 Training configuration · 22 Sept 2026
  • jqv (Qwen3-32B zero-shot) — Base model Qwen/Qwen3-32B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3 card specifies 32,768 native. The config's 40,960 positions reserve output space; 131,072 requires YaRN. System repository · 23 Sept 2026
  • Qwen3-Reranker-4B — Base model Qwen/Qwen3-Reranker-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official reranker card states 32K context; config has extra positions reserved for prompt/output. Official card's stated context is 32K; do not substitute the larger config allocation because the card is explicit. System repository · 16 Apr 2026
  • decider-35b-a3b (Mapika) — Base model Qwen/Qwen3.5-35B-A3B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-35B-A3B-Base, whose native context is also 262,144. System repository · 23 Sept 2026 Base model source · 23 Apr 2026
  • Raw Qwen3 4B Instruct 2507 direct logits — Base model Qwen/Qwen3-4B-Instruct-2507; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official model card states 262,144 natively. System repository · 17 Sept 2025
  • OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp) — Base model Qwen/Qwen3-1.7B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-1.7B card specifies 32,768 context. System repository · 23 Sept 2026
  • ZeroEntropy zerank-2 — Base model zeroentropy/zerank-2-reranker; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official model card states 32,768 context. System repository · 24 Jul 2026
  • decision-machine-1 (milliseconds.ai) — No primary public model-card or API context limit was found. System repository
  • Raw Phi-4 mini direct logits — Base model microsoft/Phi-4-mini-instruct; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Microsoft card states 128K context; config uses long-RoPE scaling. System repository · 10 Dec 2025
  • JEV Qwen3.5-9B Base NVFP4 — Base model Qwen/Qwen3.5-9B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 22 Sept 2026
  • OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) — Base model google/diffusiongemma-26B-A4B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. OPENJEV_MAX_MODEL_LEN defaults to 65,536 in the submitted runner; DiffusionGemma base supports 262,144. Published as “65,536; base 262,144”. System repository · 23 Sept 2026 Base model source · 15 Jul 2026
  • kev 4B (research preview) — Base model Qwen/Qwen3-4B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-4B-Base is in the Qwen3 family with a 32,768-token native window. System repository · 23 Sept 2026
  • Decision 2B (FlyMy.AI, v59) — Base model openbmb/MiniCPM5-2B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. MiniCPM5-2B card/config states 131,072 context. System repository · 23 Sept 2026
  • Qwen3.5-9B Jev-like data-mix v2 — Base model Qwen/Qwen3.5-9B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 22 Sept 2026
  • GPT-6 Luna (low reasoning effort) — OpenAI model docs: 1,050,000 context window and 128,000 maximum output. Published as “1,050,000 context window; 128,000 max output”.
  • SimpleJev Qwen3.8-27B — Base model Qwen/Qwen3.8-27B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The tested public Simple Jev demo API documents a 2,000-token context limit; the Qwen3.8-27B base window is 262,144. Published as “2,000 demo API context; base 262,144”. System repository · 23 Sept 2026 Base model source · 14 Aug 2026
  • NInfer Qwen3.8-Flash-Next mixed — Base model Qwen/Qwen3.8-Flash-Next; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN. System repository · 22 Sept 2026
  • GPT-6 Luna (default medium reasoning effort) — OpenAI model docs: 1,050,000 context window and 128,000 maximum output. Published as “1,050,000 context window; 128,000 max output”.
  • open-alternative-jev (Qwen3.5-4B, IkerMoel) — Base model Qwen/Qwen3.5-4B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 22 Sept 2026
  • jev-local (Qwen3.5-9B) — Base model Qwen/Qwen3.5-9B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 18 Sept 2026
  • Decision Fast (FlyMy.AI, v53a) — Base model Qwen/Qwen3-0.6B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-0.6B-Base config specifies 32,768 positions. System repository · 23 Sept 2026
  • decider-2b (Mapika) — Base model Qwen/Qwen3.5-2B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-2B-Base, whose native context is also 262,144. System repository · 23 Sept 2026 Base model source · 23 Apr 2026
  • jeff (Logan Markewich, GLiFormer 400M) — Base model knowledgator/gliformer-large-v1; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. GLiFormer card states configured max_len=8,192. System repository · 20 Sept 2026
  • Laya (Convai Innovations, ModernBERT-large 421M) — Base model convaiinnovations/laya; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Laya model card documents a 512-token base context; head_max_len=192 is its answer-candidate budget. System repository · 23 Sept 2026
  • lev-350m (Franck Verrot, LFM2.5-350M) — Base model LiquidAI/LFM2.5-350M; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. LiquidAI model card states 32,768 context; config has a larger positional allocation. LiquidAI card states a 32,768 context length. Config has a larger positional allocation; no larger trained/evaluated sequence is claimed. System repository · 21 Sept 2026
  • openjev-sglang (Qwen3.6-35B-A3B on SGLang) — Base model Qwen/Qwen3.6-35B-A3B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Runtime defaults max_input_tokens=32,768 and max_total_input_tokens=262,144. Published as “32,768 per question; 262,144 total across questions”. System repository · 21 Sept 2026 Base model source · 24 Apr 2026
  • Von (wfzyx, Option-Marker 395M) — The Von model card states an 8,192-token context for its ModernBERT-large scoring model and describes accurate premise reading to about 2,048 tokens. Published as “8,192 model context; reads well to about 2,048”. System repository · 23 Sept 2026
  • NInfer Qwen3.8-27B NVFP4 (T=1.5) — Base model Qwen/Qwen3.8-27B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. System repository · 22 Sept 2026
  • NInfer Qwen3.8-27B NVFP4 — Base model Qwen/Qwen3.8-27B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. System repository · 22 Sept 2026
  • kev 8B (research preview) — Base model Qwen/Qwen3-8B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3 card specifies 32,768 native; 131,072 requires YaRN. System repository · 23 Sept 2026
  • JevOne — Base model Qwen/Qwen3.6-35B-A3B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed. JevOne is published as Qwen3.6-35B-A3B BF16 with a bidirectional option-logit mapping; no smaller serving or training sequence cap is documented. System repository · 23 Sept 2026 Base model source · 24 Apr 2026
  • SimpleJev Qwen3.6-35B-A3B — Base model Qwen/Qwen3.6-35B-A3B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed. System repository · 23 Sept 2026
  • kev 0.6B (research preview) — Base model Qwen/Qwen3-0.6B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-0.6B-Base config specifies 32,768 positions. System repository · 23 Sept 2026
  • Raw Qwen3 8B direct logits — Base model Qwen/Qwen3-8B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3 card specifies 32,768 native; 131,072 requires YaRN. System repository · 26 Jul 2025
  • system-one (Qwen3-8B, Sean Goedecke) — Base model Qwen/Qwen3-8B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3 card specifies 32,768 native; 131,072 requires YaRN. System repository · 18 Sept 2026
  • OpenDecision (ModernBERT-large zero-shot) — Base model MoritzLaurer/ModernBERT-large-zeroshot-v2.0; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Config has 8,192 positions. Card says v2.0 may not fully use the 8K window; exact trained length is not stated. ModernBERT config allows 8,192 positions. The v2.0 card says the older zero-shot checkpoint may not fully use the long window; it does not state a smaller exact trained limit. System repository · 21 Sept 2026
  • LitJev (Qwen3.8-27B) — Base model Qwen/Qwen3.8-27B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. System repository · 21 Sept 2026
  • openJev Verdict 1.4 — Base model heman10x/rlcd-modernbert-151m; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The fixed v1.4 inference path sets its context budget to 512; underlying model config is 8,192. Published as “512 service budget; model config 8,192”. System repository · 20 Sept 2026 Base model source · 20 Sept 2026
  • kev 0.5B — Base model Qwen/Qwen2.5-0.5B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Kev 0.5B model card specifies an 8,192-token serving allowance per branch and a 32K backbone window. Published as “8,192 per branch; backbone 32,768”. System repository · 23 Sept 2026 Base model source · 25 Sept 2024
  • Bespoke Nimble 9B (Bespoke Labs) — Base model Qwen/Qwen3.5-9B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Serving code defaults NIMBLE_MAX_PROMPT_TOKENS to 8,192 and launches the backend with that request cap. Published as “8,192 prompt cap; base 262,144”. Trained sequence length: 2,048 tokens. The published adapter-training recipe uses --max-length=2,048; the ranked serving profile has an 8,192-token prompt cap. System repository · 23 Sept 2026 Base model source · 2 Mar 2026 Training configuration · 23 Sept 2026
  • GPT-5.6 Luna (low reasoning effort) — OpenAI model docs: 1,050,000 context window and 128,000 maximum output. Published as “1,050,000 context window; 128,000 max output”.
  • openJev Verdict (heman10x, ModernBERT-base 151M) — Base model knowledgator/gliclass-modern-base-v2.0; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Model tokenizer config sets an 8,192-token maximum. Model config/tokenizer sets 8,192; its current separate v1.4 inference-engine cap is not documented in the available primary sources. System repository · 20 Sept 2026
  • Raw Qwen3 1.7B direct logits — Base model Qwen/Qwen3-1.7B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-1.7B card specifies 32,768 context. System repository · 26 Jul 2025
  • reflex-27b (Qwen3.8-27B) — Base model Qwen/Qwen3.8-27B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. System repository · 23 Sept 2026
  • djev (thinking) — Base model google/diffusiongemma-26B-A4B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. DiffusionGemma card/config states 256K context. The standard djev runtime caps requests at 32,768, but the separately measured full-generation thinking run has no matching runtime cap published. The listed 262,144 is the DiffusionGemma base window, not a verified cap for this run. System repository · 19 Sept 2026
  • GLiNER2 large (Fastino) — Base model microsoft/deberta-v3-large; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official docs provide chunked long-document helpers but no maximum total document size; the named DeBERTa-v3-large encoder has 512 positions. Published as “Unknown total document limit; 512-token DeBERTa chunks”. System repository · 17 Sept 2026 Base model source · 19 Mar 2023
  • OpenJev (thinking, BF16) — Base model google/diffusiongemma-26B-A4B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. OpenJev runtime defaults OPENJEV_MAX_MODEL_LEN to 65,536; the 512-token thinking allowance is generated output, not input context. Published as “65,536; base 262,144”. System repository · 23 Sept 2026 Base model source · 15 Jul 2026
  • Qwen3.5-0.8B Decision Model (Mourad Ghafiri) — Base model Qwen/Qwen3.5-0.8B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. The submitted decision model's config.json sets 262,144 positions; it is based on Qwen3.5-0.8B-Base, whose native context is 262,144. System repository · 22 Sept 2026 Base model source · 23 Apr 2026
  • Gemini 3.1 Flash-Lite — Google model docs explicitly list a 1,048,576 input-token limit and 65,536 output-token limit. Published as “1,048,576 input token limit”.
  • open-jev-deberta-v3-large (local CPU) — Base model microsoft/deberta-v3-large; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Microsoft config max_position_embeddings=512. System repository · 22 Sept 2026
  • smalljev semantic-v9 — Base model openbmb/MiniCPM5-2B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. MiniCPM5-2B card/config states 131,072 context. System repository · 21 Sept 2026
  • GLiNER2 (Fastino, gliner2.5-base) — Base model microsoft/deberta-v3-large; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official docs provide chunked long-document helpers but no maximum total document size; underlying DeBERTa-v3-base encoder has 512 positions. Published as “Unknown total document limit; 512-token DeBERTa chunks”. System repository · 23 Sept 2026 Base model source · 19 Mar 2023
  • Open-Jev 9B (Zefan Cai) — Base model Qwen/Qwen3.5-9B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. System repository · 23 Sept 2026
  • Open-Jev 2B (Zefan Cai) — Base model Qwen/Qwen3.5-2B-Base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Mapika model config and Qwen3.5 family card give 262,144 positions. System repository · 23 Sept 2026
  • GLiNER2.5 multi (Fastino, 287M) — Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published. Published as “Unknown; long-document chunking documented”. System repository · 20 Sept 2026
  • SimpleJev (Qwen3.5-0.8B, CPU) — Base model Qwen/Qwen3.5-0.8B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official Qwen3.5-0.8B card/config reports 262,144 natively; optional YaRN extends the window. System repository · 23 Sept 2026
  • GLiNER2.5 small (Fastino, 74M) — Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published. Published as “Unknown; long-document chunking documented”. System repository · 20 Sept 2026
  • Raw Qwen3 0.6B direct logits — Base model Qwen/Qwen3-0.6B; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3-0.6B card specifies 32,768 context. System repository · 26 Jul 2025
  • DeepSeek V4.1 Flash (thinking default) — DeepSeek API model/pricing docs list DeepSeek V4.1 Flash at 1M context and 384K maximum output. Published as “1,000,000 context; 384,000 max output”.
  • Mirror — Base model microsoft/deberta-v3-large; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Microsoft config max_position_embeddings=512. System repository
  • Mixedbread mxbai-rerank-base-v2 — Base model mixedbread-ai/mxbai-rerank-base-v2; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official config max_position_embeddings=32,768. System repository · 8 Apr 2026
  • BAAI bge-reranker-v2-m3 — Base model BAAI/bge-reranker-v2-m3; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official tokenizer limit is 8,192; config has 8,194 positions. System repository · 24 Jun 2024
  • Alibaba GTE Reranker ModernBERT-base — Base model Alibaba-NLP/gte-reranker-modernbert-base; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Official config max_position_embeddings=8,192. System repository · 4 Jul 2025
  • Certo v1 (AltSlate Labs) — Base model altslate/certo-decision-model; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Submitted model config/tokenizer sets 8,192. System repository · 21 Sept 2026
  • Open Jev JSON Canvas (JoshuaSP) — Base model google/diffusiongemma-26B-A4B-it; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. DiffusionGemma card/config states 256K context. System repository · 16 Sept 2026
  • classifier.dev (fast tier) — The Jev model behind the service has a documented limit, but no primary source documents a narrower or matching classifier.dev endpoint cap. System repository
  • Needle 3 (Cactus, 2-bit, local CPU) — Official Needle 3 docs describe text input but publish no maximum context/token limit. System repository
  • Needle 3, options as tools (post-hoc adapter mode) — Official Needle 3 docs describe tool inputs but publish no maximum context/token limit. System repository
  • Qwen3.8 27B (Chutes TEE) — Chutes model catalog lists Qwen3.8-27B-TEE at 262K context. Published as “262,144 context”.
  • swanOne — Base model Qwen/Qwen3.8-Flash-Next; a Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN. The submitted NVFP4 runner is based on Qwen3.8-Flash-Next; 262,144 is its native window. No larger runtime setting is documented for the submitted patch. System repository

“Hard limit” is an explicit model or tokenizer ceiling; “Trained length” is a published base-model or training length; “API cap” is a published service limit. These are different kinds of evidence. A base-model window does not prove that a particular hosted endpoint accepts the same length; row notes identify cases where its serving cap is unpublished. Some API docs publish a combined context window and a separate output ceiling, so the usable input can be lower when output tokens share that window. “Unknown” means no supported maximum was found.