JevBench v1.2 · our own benchmark

Jev-class models

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Version 1.2 measures 16 systems on 534 decisions, including 220 hard ones, and ranks them by the JevBench Score. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.

Scored 19 Sept 2026 · protocol jevbench::v1.2 · 72 easy + 96 standard + 146 judge + 220 hard decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 88ab9abde5cc · v1.0 results

JevBench v1.2.1 · 534 decisions per system

JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)

OfficialIntelligence, Calibration, Speed, Cost — 25 % each, geometric mean: a weak axis pulls the score down hard. Change the weighting ↓

  1. 1Jev 1.13.075.3I 90 · C 83 · S 83 · K 52 · $0.041
  2. 2SemIf (Qwen3.5-4B)74.6I 86 · C 73 · S 84 · K 59 · ~$0.023 est.
  3. 3djev (Maisa, diffusion-gemma)74.3I 88 · C 65 · S 91 · K 58 · $0.026 ann.
  4. 4open-alternative-jev (Qwen3.5-4B)69.8I 76 · C 63 · S 83 · K 60 · ~$0.022 est.
  5. 5system-one-open68.7I 80 · C 57 · S 77 · K 64 · ~$0.016 est.
  6. 6OpenJev (razorback16)67.6I 86 · C 65 · S 83 · K 45 · ~$0.067 est.
  7. 7openjev-sglang66.2I 89 · C 77 · S 77 · K 36 · ~$0.135 est.
  8. 8GPT-5.6 Luna (low)66.0I 97 · C 90 · S 78 · K 28 · $0.247
  9. 9open-jev-deberta-v3-large64.4I 54 · C 66 · S 66 · K 73 · ~$0.0077 est.
  10. 10Bespoke Nimble 9B63.5I 79 · C 65 · S 83 · K 39 · ~$0.109 est.
  11. 11Gemini 3.1 Flash-Lite60.8I 90 · C 68 · S 82 · K 27 · $0.268
  12. 12DeepSeek V4.1 Flash58.1I 96 · C 97 · S 72 · K 17 · $0.579
  13. 13system-one56.5I 80 · C 37 · S 84 · K 41 · ~$0.092 est.
  14. Qwen3.8 27B (partial run)25.5I 75 · C 92 · S 61 · K 0 · ~$2.711 est.
  15. Needle 3, options as tools (partial run)19.1I 40 · C · S 53 · K 64 · ~$0.016 est.
  16. Needle 3 (partial run)16.7I 22 · C · S 60 · K 58 · ~$0.025 est.

Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean)

  • Jev (TypeSafe, closed)
  • Jev rebuild (open, or open source planned)
  • Instruction model, JSON schema
  • Small tool-calling model
  • Partial run — shown, not ranked
Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.
Legend and notes
  • ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
  • ann. = the provider’s announced price, not yet charged.
  • Names link to each project.
  • A label-only system has no calibration (–, counted as 0).
  • djev (Maisa, diffusion-gemma): Hosted API in free preview: the cost uses djev's announced price ($0.035 per million input tokens, output free); nothing is charged yet. Open-sourcing is planned, not yet released. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.

Weighting: Intelligence : Calibration : Speed : Cost

Official default
Custom

The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.

Explore by task difficulty

All tasks is the published default. Choose an easier scope to see how the ranking changes when your work is mostly straightforward; Intelligence and the JevBench Score are recomputed from the published tier aggregates.

Tier mapping: Easy = easy; Medium = standard; Judge and Hard remain in All tasks. The score still includes the published Calibration, Speed and Cost axes.

What the run says (JevBench Score)

Axes, tiers, latency and cost

Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.
Rank#SystemEndpoint
1by TypeSafe AIJev 1.13.075.390.482.783.351.7$0.041100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API
2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ74.685.972.683.759.2~$0.023 est.100.0%97.9%95.2%59.5%0.20 s raw0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU
3by Maisa (David Villalón)djevMaisa, diffusion-gemma74.388.465.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API
4by IkerMoelopen-alternative-jevQwen3.5-4B, IkerMoel69.875.663.283.559.6~$0.022 est.100.0%84.4%74.7%56.8%0.21 s raw0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU
5by mithalounisystem-one-openGemma 4 E2B LoRA on an L468.779.556.777.064.1~$0.016 est.100.0%93.8%87.7%49.1%0.65 s raw1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server
6by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1667.686.064.883.245.2~$0.067 est.100.0%95.8%91.1%65.5%0.24 s raw0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU
7by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang66.288.977.477.136.1~$0.135 est.100.0%95.8%95.2%71.4%0.68 s raw1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server
8by OpenAIGPT-5.6 Lunalow reasoning effort66.096.889.877.528.2$0.247100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API
9by Kotoba Labsopen-jev-deberta-v3-largelocal CPU64.453.666.466.073.3~$0.0077 est.100.0%49.0%53.4%36.4%1.77 s raw3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU
10by Bespoke LabsBespoke Nimble 9B63.578.664.582.538.9~$0.109 est.100.0%94.8%89.0%43.6%0.19 s raw0.52 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU
11by GoogleGemini 3.1 Flash-Lite60.890.368.181.827.1$0.268100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API
12by DeepSeekDeepSeek V4.1 Flashthinking default58.196.196.771.617.1$0.57998.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API
13by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke56.580.136.884.441.2~$0.092 est.100.0%90.6%91.8%50.0%0.17 s raw0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU
Partial runs — shown, not ranked: ranked = every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank.
by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked25.574.692.161.30.0~$2.711 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction API
by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked19.139.5none (label only)52.863.7~$0.016 est.66.7%31.3%34.2%3.78 s raw7.71 s adjustedp95 33.64 s raw → 67.42 sour CPU
by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked16.722.4none (label only)59.958.1~$0.025 est.47.2%16.7%31.5%7.7%1.69 s raw3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU

Which public tasks did each system get right?

This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped. The pinned artifact has no public-task outcomes for djev yet, so its cells remain unavailable.

Show 231 public task outcomes across 16 systems
TaskJev 1.13.0SemIfdjevopen-alternative-jevsystem-one-openOpenJevopenjev-sglangGPT-5.6 Lunaopen-jev-deberta-v3-largeBespoke Nimble 9BGemini 3.1 Flash-LiteDeepSeek V4.1 Flashsystem-oneQwen3.8 27BNeedle 3, options as toolsNeedle 3
Easy · 48 public tasks
easy-intent-00intent · choice×
easy-intent-01intent · choice
easy-intent-02intent · choice××
easy-intent-03intent · choice×
easy-intent-04intent · choice××
easy-intent-05intent · choice×
easy-intent-06intent · choice×
easy-intent-07intent · choice
easy-intent-08intent · choice×
easy-intent-09intent · choice×
easy-intent-10intent · choice×
easy-intent-11intent · choice
easy-fact-00fact · noul
easy-fact-01fact · noul
easy-fact-02fact · noul
easy-fact-03fact · noul×
easy-fact-04fact · noul×
easy-fact-05fact · noul××
easy-fact-06fact · noul
easy-fact-07fact · noul××
easy-fact-08fact · noul
easy-fact-09fact · noul××
easy-fact-10fact · noul××
easy-fact-11fact · noul××
easy-extraction-00extraction · choice×
easy-extraction-01extraction · choice×
easy-extraction-02extraction · choice×
easy-extraction-03extraction · choice
easy-extraction-04extraction · choice×
easy-extraction-05extraction · choice
easy-extraction-06extraction · choice
easy-extraction-07extraction · choice×
easy-extraction-08extraction · choice
easy-extraction-09extraction · choice×
easy-extraction-10extraction · choice×
easy-extraction-11extraction · choice
easy-tool_selection-00tool_selection · choice
easy-tool_selection-01tool_selection · choice×
easy-tool_selection-02tool_selection · choice×
easy-tool_selection-03tool_selection · choice×
easy-tool_selection-04tool_selection · choice×
easy-tool_selection-05tool_selection · choice×
easy-tool_selection-06tool_selection · choice×
easy-tool_selection-07tool_selection · choice
easy-tool_selection-08tool_selection · choice××
easy-tool_selection-09tool_selection · choice×
easy-tool_selection-10tool_selection · choice×
easy-tool_selection-11tool_selection · choice×
Medium (standard) · 72 public tasks
original-policy-01-0policy · noul××
original-policy-01-1policy · noul!××
original-policy-02-0policy · noul××
original-policy-02-1policy · noul××××××
original-policy-03-0policy · noul×××××
original-policy-03-1policy · noul×××××
original-policy-04-0policy · noul×××
original-policy-04-1policy · noul××
original-policy-05-0policy · noul××
original-policy-05-1policy · noul××
original-policy-06-0policy · noul×××
original-policy-06-1policy · noul××××××
original-intent-01-0intent · choice×××
original-intent-01-1intent · choice×
original-intent-02-0intent · choice××××
original-intent-02-1intent · choice×
original-intent-03-0intent · choice××
original-intent-03-1intent · choice×
original-intent-04-0intent · choice×
original-intent-04-1intent · choice×
original-intent-05-0intent · choice×××××
original-intent-05-1intent · choice×××××
original-intent-06-0intent · choice××
original-intent-06-1intent · choice××
original-ordinal-01-0ordinal · score×××
original-ordinal-01-1ordinal · score×××
original-ordinal-02-0ordinal · score×××
original-ordinal-02-1ordinal · score×××
original-ordinal-03-0ordinal · score××
original-ordinal-03-1ordinal · score××
original-ordinal-04-0ordinal · score×××
original-ordinal-04-1ordinal · score××
original-ordinal-05-0ordinal · score×××
original-ordinal-05-1ordinal · score×××
original-ordinal-06-0ordinal · score××
original-ordinal-06-1ordinal · score××
original-extraction-01-0extraction · choice×××
original-extraction-01-1extraction · choice×
original-extraction-02-0extraction · choice×××
original-extraction-02-1extraction · choice×××
original-extraction-03-0extraction · choice××
original-extraction-03-1extraction · choice××
original-extraction-04-0extraction · choice××
original-extraction-04-1extraction · choice
original-extraction-05-0extraction · choice×××
original-extraction-05-1extraction · choice×××
original-extraction-06-0extraction · choice
original-extraction-06-1extraction · choice×
original-adequacy-01-0adequacy · noul×
original-adequacy-01-1adequacy · noul×
original-adequacy-02-0adequacy · noul×××
original-adequacy-02-1adequacy · noul
original-adequacy-03-0adequacy · noul×××××××
original-adequacy-03-1adequacy · noul××××××××
original-adequacy-04-0adequacy · noul×××
original-adequacy-04-1adequacy · noul×××
original-adequacy-05-0adequacy · noul××××××
original-adequacy-05-1adequacy · noul×××××
original-adequacy-06-0adequacy · noul×
original-adequacy-06-1adequacy · noul××
original-routing-01-0routing · choice××
original-routing-01-1routing · choice×××
original-routing-02-0routing · choice××
original-routing-02-1routing · choice××
original-routing-03-0routing · choice××
original-routing-03-1routing · choice××
original-routing-04-0routing · choice××
original-routing-04-1routing · choice×××××
original-routing-05-0routing · choice×××
original-routing-05-1routing · choice×××××
original-routing-06-0routing · choice×!×
original-routing-06-1routing · choice×××
Hard · 111 public tasks
hard-opus-a-long_policy-01long_policy · choice!!·×
hard-opus-a-long_policy-04long_policy · choice×××××××!××·
hard-opus-a-long_policy-08long_policy · choice××!·×
hard-opus-a-long_policy-09long_policy · choice××!×·×
hard-opus-a-long_policy-11long_policy · noul××!·×
hard-opus-a-long_policy-13long_policy · noul××××××!××!·
hard-opus-a-long_policy-17long_policy · choice×××××!×·×
hard-opus-a-long_policy-19long_policy · noul×××!×·×
hard-opus-a-probability-03probability · noul×××·
hard-opus-a-probability-04probability · choice×××××××××·×
hard-opus-a-probability-07probability · choice××××××××·×
hard-opus-a-probability-08probability · noul×××·×
hard-opus-a-temporal_numeric-03temporal_numeric · noul×××××××××·
hard-opus-a-temporal_numeric-06temporal_numeric · noul×××××·×
hard-opus-a-temporal_numeric-07temporal_numeric · choice××××××!×·
hard-opus-a-temporal_numeric-09temporal_numeric · noul×××××·×
hard-opus-a-temporal_numeric-12temporal_numeric · score××××××××××·
hard-opus-b-ambiguous-02ambiguous · choice××·×
hard-opus-b-ambiguous-03ambiguous · choice×××××××·
hard-opus-b-ambiguous-07ambiguous · choice×××××·
hard-opus-b-ambiguous-09ambiguous · choice××××·
hard-opus-b-ambiguous-10ambiguous · choice××·×
hard-opus-b-ambiguous-11ambiguous · choice·×
hard-opus-b-ambiguous-13ambiguous · choice××·×
hard-opus-b-multi_hop-03multi_hop · choice×××××!××·×
hard-opus-b-multi_hop-04multi_hop · choice××××××!××·×
hard-opus-b-multi_hop-05multi_hop · score×××××××!××·×
hard-opus-b-multi_hop-07multi_hop · choice×××!××·×
hard-opus-b-multi_hop-08multi_hop · noul××!×·×
hard-opus-b-probability-01probability · choice××××!×·×
hard-opus-b-probability-02probability · choice×××××××××·×
hard-opus-b-probability-03probability · noul××××·
hard-opus-b-probability-04probability · choice·
hard-opus-b-probability-06probability · choice××××××!×·
hard-opus-b-tradeoff-01tradeoff · choice×××××××××·×
hard-opus-b-tradeoff-03tradeoff · choice×××·×
hard-opus-b-tradeoff-06tradeoff · noul××××××××·
hard-opus-b-tradeoff-07tradeoff · choice××××××·×
hard-opus-b-tradeoff-08tradeoff · noul××××××·
hard-opus-b-tradeoff-12tradeoff · choice×·×
hard-opus-c-long_policy-02long_policy · choice××!·×
hard-opus-c-long_policy-03long_policy · choice×××××××!×·
hard-opus-c-long_policy-04long_policy · noul×!·
hard-opus-c-long_policy-05long_policy · score××××××!××!·
hard-opus-c-long_policy-08long_policy · score×××××××!××··
hard-opus-c-long_policy-10long_policy · choice×××!×··
hard-opus-c-long_policy-11long_policy · noul××××××!×!··
hard-opus-c-probability-03probability · choice××··
hard-opus-c-temporal_numeric-02temporal_numeric · noul×××××××··
hard-opus-c-temporal_numeric-03temporal_numeric · choice×××××××··
hard-opus-c-temporal_numeric-04temporal_numeric · choice××××××××××··
hard-opus-c-temporal_numeric-06temporal_numeric · choice×××××××××···
hard-opus-c-temporal_numeric-08temporal_numeric · choice××××××××···
hard-opus-c-temporal_numeric-12temporal_numeric · choice××××···
hard-sol-a-adversarial-01adversarial · choice×××···
hard-sol-a-adversarial-06adversarial · noul···
hard-sol-a-adversarial-07adversarial · choice×××···
hard-sol-a-adversarial-08adversarial · noul××···
hard-sol-a-adversarial-09adversarial · score×···
hard-sol-a-adversarial-11adversarial · choice×××××···
hard-sol-a-multi_hop-01multi_hop · choice!···
hard-sol-a-multi_hop-02multi_hop · choice××!···
hard-sol-a-multi_hop-05multi_hop · choice××!···
hard-sol-a-multi_hop-07multi_hop · choice×!×···
hard-sol-a-multi_hop-08multi_hop · choice×××××!××···
hard-sol-a-multi_hop-09multi_hop · choice!×···
hard-sol-a-multi_hop-10multi_hop · choice×××!×···
hard-sol-a-multi_hop-12multi_hop · choice×××!···
hard-sol-a-trap-02trap · noul×···
hard-sol-a-trap-04trap · noul···
hard-sol-a-trap-05trap · score×××···
hard-sol-a-trap-06trap · choice···
hard-sol-a-trap-08trap · choice×××···
hard-sol-a-trap-10trap · choice×···
hard-sol-a-trap-13trap · noul···
hard-sol-a-trap-15trap · noul×···
hard-sol-b-judge_hard-01judge_hard · noul···
hard-sol-b-judge_hard-02judge_hard · noul××××××××××××···
hard-sol-b-judge_hard-05judge_hard · noul···
hard-sol-b-judge_hard-08judge_hard · noul××××××××···
hard-sol-b-judge_hard-10judge_hard · noul×××××···
hard-sol-b-judge_hard-14judge_hard · noul×···
hard-sol-b-judge_hard-15judge_hard · noul···
hard-sol-b-judge_hard-18judge_hard · noul××××···
hard-sol-b-long_policy-01long_policy · choice××!···
hard-sol-b-long_policy-02long_policy · choice×××!×···
hard-sol-b-long_policy-05long_policy · choice××!···
hard-sol-b-long_policy-06long_policy · choice×××!!×···
hard-sol-b-routing_hard-01routing_hard · choice···
hard-sol-b-routing_hard-02routing_hard · choice××···
hard-sol-b-routing_hard-03routing_hard · choice×···
hard-sol-b-routing_hard-07routing_hard · choice×···
hard-sol-b-routing_hard-09routing_hard · choice×···
hard-sol-b-temporal_numeric-01temporal_numeric · choice××××××××××···
hard-sol-b-temporal_numeric-02temporal_numeric · choice×××××···
hard-sol-b-temporal_numeric-03temporal_numeric · choice××××××××···
hard-sol-b-temporal_numeric-04temporal_numeric · choice×××××××···
hard-sol-c-judge_hard-03judge_hard · noul×××···
hard-sol-c-judge_hard-05judge_hard · noul××···
hard-sol-c-judge_hard-07judge_hard · noul××××···
hard-sol-c-judge_hard-08judge_hard · noul×××···
hard-sol-c-judge_hard-09judge_hard · noul×···
hard-sol-c-judge_hard-10judge_hard · noul···
hard-sol-c-judge_hard-11judge_hard · noul×···
hard-sol-c-judge_hard-13judge_hard · noul×××××···
hard-sol-c-judge_hard-15judge_hard · noul×···
hard-sol-c-multi_hop-07multi_hop · choice××!×···
hard-sol-c-multi_hop-09multi_hop · choice!···
hard-sol-c-multi_hop-11multi_hop · choice××!···
hard-sol-c-multi_hop-12multi_hop · choice×××···
hard-sol-c-multi_hop-13multi_hop · choice×!···

✓ correct · × wrong · ! failed (scored wrong) · · not attempted · — no public outcome in the pinned artifact. Task descriptions are intentionally not included; the task id, tier and topic are the published public metadata.

How the JevBench Score works

JevBench Score = (Intelligence × Calibration × Speed × Cost)1/4, each axis on 0–100 — the geometric mean. A weak axis pulls the score down hard: a strong axis cannot buy it back.

Full scoring rules
JevBench Score.
exp(sum over the four axes of 0.25 x ln(max(axis, 1))) — the geometric mean of Intelligence, Calibration, Speed and Cost. A weak axis pulls the score down hard; a strong axis cannot buy it back.
Intelligence.
100 x weighted accuracy: hard 30 %, easy 14 %, standard 28 %, judge 28 %. Accuracy = correct / all items; failed, timed-out or unparseable answers count as wrong.
Calibration.
Hard tier only, systems that return a probability distribution: mean of (a) 100 x (1 - ECE/0.5), ECE = top-label expected calibration error in 10 bins, and (b) probability fidelity = 100 x (1 - mean total-variation distance) between the returned distribution and the exact gold distribution on the 20 probability items. Label-only systems have none; it counts as 0 in the JevBench Score.
Speed.
Mean of score(p50) and score(p95) of the serial 242-decision standard+judge run; score(s) = 100 - 20 log10(s / 0.1 s), clipped to 0..100 (0.1 s = 100, 1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo. Production APIs (Jev, djev, OpenAI, Google, DeepSeek, Chutes) are not adjusted.
Cost.
Dollars per 1,000 decisions pooled over all 534 v1.2 decisions; score = 100 - 30 log10(usd / 0.001), clipped to 0..100 ($0.001 = 100, $0.01 = 70, $0.10 = 40, $1 = 10). Measured = public tariff x measured tokens. est. = hosted-provider list price of the same weights or size class x tokens. announced = the provider's published price, not yet charged (free preview), x measured tokens.
Ranked.
Ranked: every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank.
Presets.
Other views reweight the same four axes and combine them the same way (geometric mean). They are not the JevBench Score.

How the ranking moves with other weights

Rank and score under the JevBench Score and the earlier views, all combined as a geometric mean (Intelligence : Calibration : Speed : Cost). Highlighted = a different rank than the JevBench Score. Ranked systems only.

SystemJevBench Score25:25:25:25 · officialBalanced, no calibration33:0:33:33 · not the defaultEmphasis on Accuracy60:0:20:20 · not the defaultEmphasis on Speed20:0:60:20 · not the defaultEmphasis on Cost20:0:20:60 · not the default
Jev 1.13.0#1 75.3#4 73.0 (rank differs from the JevBench Score)#2 79.5 (rank differs from the JevBench Score)#3 77.0 (rank differs from the JevBench Score)#6 63.6 (rank differs from the JevBench Score)
SemIf#2 74.6#2 75.2#3 79.3 (rank differs from the JevBench Score)#2 78.5#3 68.3 (rank differs from the JevBench Score)
djev#3 74.3#1 77.5 (rank differs from the JevBench Score)#1 81.6 (rank differs from the JevBench Score)#1 82.7 (rank differs from the JevBench Score)#2 68.8 (rank differs from the JevBench Score)
open-alternative-jev#4 69.8#5 72.2 (rank differs from the JevBench Score)#6 73.5 (rank differs from the JevBench Score)#4 76.5#5 66.9 (rank differs from the JevBench Score)
system-one-open#5 68.7#3 73.2 (rank differs from the JevBench Score)#4 75.7 (rank differs from the JevBench Score)#5 74.7#1 69.4 (rank differs from the JevBench Score)
OpenJev#6 67.6#6 68.6#5 75.1 (rank differs from the JevBench Score)#6 74.1#7 58.1 (rank differs from the JevBench Score)
openjev-sglang#7 66.2#10 62.8 (rank differs from the JevBench Score)#8 72.2 (rank differs from the JevBench Score)#9 68.1 (rank differs from the JevBench Score)#10 50.3 (rank differs from the JevBench Score)
GPT-5.6 Luna#8 66.0#11 59.6 (rank differs from the JevBench Score)#7 72.4 (rank differs from the JevBench Score)#11 66.2 (rank differs from the JevBench Score)#11 44.2 (rank differs from the JevBench Score)
open-jev-deberta-v3-large#9 64.4#8 63.8 (rank differs from the JevBench Score)#13 59.5 (rank differs from the JevBench Score)#12 64.6 (rank differs from the JevBench Score)#4 67.4 (rank differs from the JevBench Score)
Bespoke Nimble 9B#10 63.5#9 63.2 (rank differs from the JevBench Score)#11 69.0 (rank differs from the JevBench Score)#8 70.3 (rank differs from the JevBench Score)#9 52.1 (rank differs from the JevBench Score)
Gemini 3.1 Flash-Lite#11 60.8#12 58.5 (rank differs from the JevBench Score)#10 69.6 (rank differs from the JevBench Score)#10 66.9 (rank differs from the JevBench Score)#12 43.0 (rank differs from the JevBench Score)
DeepSeek V4.1 Flash#12 58.1#13 49.0 (rank differs from the JevBench Score)#12 64.2#13 57.0 (rank differs from the JevBench Score)#13 32.2 (rank differs from the JevBench Score)
system-one#13 56.5#7 65.3 (rank differs from the JevBench Score)#9 70.8 (rank differs from the JevBench Score)#7 72.3 (rank differs from the JevBench Score)#8 54.3 (rank differs from the JevBench Score)

Compare two systems

Pick any two. The first radar shows the four axes of the JevBench Score (0–100, the values in the table above); the second shows accuracy by subject topic, over all tiers. Further out is better on every spoke.

  • A: Jev 1.13.0 Jev · JevBench Score 75.3 (#1)
  • B: SemIf Jev rebuild · JevBench Score 74.6 (#2)

The four score axes

Radar: the four JevBench Score axes, two systemsJev 1.13.0 vs SemIf. Intelligence: 90.4 vs 85.9; Calibration: 82.7 vs 72.6; Speed: 83.3 vs 83.7; Cost: 51.7 vs 59.2.50100Intelligence90.4 · 85.9Calibration82.7 · 72.6Speed83.3 · 83.7Cost51.7 · 59.2
Speed includes the latency adjustment for self-hosted and demo endpoints — an assumption, see Limits. A label-only system has no calibration (counted as 0).
Values as a table
AxisA: Jev 1.13.0B: SemIf
Intelligence90.485.9
Calibration82.772.6
Speed83.383.7
Cost51.759.2
JevBench Score75.374.6

Accuracy by subject topic

Radar: accuracy by subject topic, two systemsAccuracy by subject topic, Jev 1.13.0 vs SemIf. Math & numbers (129 items): 87.6% vs 79.1%; Coding & software (56 items): 83.9% vs 96.4%; Rules, policy & law (67 items): 83.6% vs 64.2%; Finance & commerce (64 items): 73.4% vs 60.9%; Support & operations (119 items): 89.1% vs 87.4%; Everyday language (79 items): 100.0% vs 100.0%; Safety & security (20 items): 100.0% vs 75.0%.50100Math87.6% · 79.1%Coding83.9% · 96.4%Rules & law83.6% · 64.2%Finance73.4% · 60.9%Support & ops89.1% · 87.4%Everydaylanguage100.0% · 100.0%Safety &security100.0% · 75.0%
Share of a topic's decisions answered correctly, easy to hard together (items per topic in the table). Topics mix tiers differently — Everyday language is mostly easy items, Rules & law and Finance mostly hard ones — so compare the two systems within a topic, not topics with each other. Not part of the JevBench Score.Topics: one per item, drafted by a model and checked by hand — method. Held-out items count in the totals; their texts stay private.
Values as a table
Topic (items)A: Jev 1.13.0B: SemIf
Math & numbers (129)a calculation decides the answer: arithmetic, word problems, probability, dates, units87.6% 113 of 12979.1% 102 of 129
Coding & software (56)code, SQL, repositories, developer tools and IT systems83.9% 47 of 5696.4% 54 of 56
Rules, policy & law (67)applying written rules: company policies, contracts, regulations, eligibility83.6% 56 of 6764.2% 43 of 67
Finance & commerce (64)money: payments, refunds, invoices, orders, expenses, insurance payouts73.4% 47 of 6460.9% 39 of 64
Support & operations (119)support tickets, incidents, logistics, scheduling desks and routing work to a team89.1% 106 of 11987.4% 104 of 119
Everyday language (79)short everyday messages: intents, assistant requests, reading a detail out of a text100.0% 79 of 79100.0% 79 of 79
Safety & security (20)untrusted or injected instructions, fraud, moderation, access and security triage100.0% 20 of 2075.0% 15 of 20

Systems with a public tariff use that tariff and measured tokens. For systems without one, we use a clearly marked estimate based on a large inference provider's list price for the same weights or size class.

How costs are estimated

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

  • SemIf~$0.023 est. per 1,000: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 426 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
  • open-alternative-jev~$0.022 est. per 1,000: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
  • system-one-open~$0.016 est. per 1,000: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 452 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
  • OpenJev~$0.067 est. per 1,000: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 410 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
  • openjev-sglang~$0.135 est. per 1,000: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 667 input and 2 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
  • open-jev-deberta-v3-large~$0.0077 est. per 1,000: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
  • Bespoke Nimble 9B~$0.109 est. per 1,000: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B)) x 990 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3.5-9b $0.1/M in, $0.15/M out x 1215 in / 2 out tokens per hard decision
  • system-one~$0.092 est. per 1,000: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 443 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
  • Qwen3.8 27B~$2.711 est. per 1,000: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 445 input and 393 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
  • Needle 3, options as tools~$0.016 est. per 1,000: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 452 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed.
  • Needle 3~$0.025 est. per 1,000: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 452 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
  • dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
  • dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
  • dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
  • encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
  • generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
  • moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
  • moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9

Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).

Who could not be measured, and why

An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank.

Method and tiers

exp(sum over the four axes of 0.25 x ln(max(axis, 1))) — the geometric mean of Intelligence, Calibration, Speed and Cost. A weak axis pulls the score down hard; a strong axis cannot buy it back.

Revision v1.2.1. v1.2 final: 4 axes (Intelligence, Calibration, Speed, Cost), 25 % each, geometric mean; hard tier 30 % of Intelligence; latency of non-production endpoints adjusted (assumption); one open-alternative-jev row (author's option order); Needle 3 options-as-tools priced. v1.2.1: added djev (Maisa, diffusion-gemma).

  • easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
  • standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
  • judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
  • hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.

Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.

v1.2 numbers are not comparable with v1.1 or v1.0 (different tiers and scoring). The v1.0 page keeps its own numbers, calibration plots and per-family tables.

Limits
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.