JevBench v1.2 · our own benchmark
Jev-class models
JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
Version 1.2 measures 16 systems on 534 decisions, including 220 hard ones, and ranks them by the JevBench Score. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.
Scored 19 Sept 2026 · protocol jevbench::v1.2 · 72 easy + 96 standard + 146 judge + 220 hard decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 88ab9abde5cc… · v1.0 results
JevBench v1.2.1 · 534 decisions per system
JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)
OfficialIntelligence, Calibration, Speed, Cost — 25 % each, geometric mean: a weak axis pulls the score down hard. Change the weighting ↓
- 1Jev 1.13.075.3I 90 · C 83 · S 83 · K 52 · $0.041
- 2SemIf (Qwen3.5-4B)74.6I 86 · C 73 · S 84 · K 59 · ~$0.023 est.
- 3djev (Maisa, diffusion-gemma)†74.3I 88 · C 65 · S 91 · K 58 · $0.026 ann.
- 4open-alternative-jev (Qwen3.5-4B)†69.8I 76 · C 63 · S 83 · K 60 · ~$0.022 est.
- 5system-one-open68.7I 80 · C 57 · S 77 · K 64 · ~$0.016 est.
- 6OpenJev (razorback16)67.6I 86 · C 65 · S 83 · K 45 · ~$0.067 est.
- 7openjev-sglang66.2I 89 · C 77 · S 77 · K 36 · ~$0.135 est.
- 8GPT-5.6 Luna (low)66.0I 97 · C 90 · S 78 · K 28 · $0.247
- 9open-jev-deberta-v3-large64.4I 54 · C 66 · S 66 · K 73 · ~$0.0077 est.
- 10Bespoke Nimble 9B63.5I 79 · C 65 · S 83 · K 39 · ~$0.109 est.
- 11Gemini 3.1 Flash-Lite60.8I 90 · C 68 · S 82 · K 27 · $0.268
- 12DeepSeek V4.1 Flash58.1I 96 · C 97 · S 72 · K 17 · $0.579
- 13system-one56.5I 80 · C 37 · S 84 · K 41 · ~$0.092 est.
- Qwen3.8 27B (partial run)25.5I 75 · C 92 · S 61 · K 0 · ~$2.711 est.
- Needle 3, options as tools (partial run)19.1I 40 · C – · S 53 · K 64 · ~$0.016 est.
- Needle 3 (partial run)16.7I 22 · C – · S 60 · K 58 · ~$0.025 est.
Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean)
- Jev (TypeSafe, closed)
- Jev rebuild (open, or open source planned)
- Instruction model, JSON schema
- Small tool-calling model
- Partial run — shown, not ranked
Legend and notes
- ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
- ann. = the provider’s announced price, not yet charged.
- Names link to each project.
- A label-only system has no calibration (–, counted as 0).
- djev (Maisa, diffusion-gemma): Hosted API in free preview: the cost uses djev's announced price ($0.035 per million input tokens, output free); nothing is charged yet. Open-sourcing is planned, not yet released. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
Weighting: Intelligence : Calibration : Speed : Cost
Official defaultCustom
The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.
Explore by task difficulty
All tasks is the published default. Choose an easier scope to see how the ranking changes when your work is mostly straightforward; Intelligence and the JevBench Score are recomputed from the published tier aggregates.
Tier mapping: Easy = easy; Medium = standard; Judge and Hard remain in All tasks. The score still includes the published Calibration, Speed and Cost axes.
What the run says (JevBench Score)
- Jev 1.13.0 (TypeSafe AI) leads with 75.3: Intelligence 90.4, Calibration 82.7, Speed 83.3, Cost 51.7 ($0.041 per 1,000 decisions).
- Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 74.6 — 0.7 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
- GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (96.8) but places #8: its cost score is 28.2 ($0.247 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
- Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not finish every tier in time; they are shown below the ranking as partial runs, without a rank.
Axes, tiers, latency and cost
Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.
⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.| Rank# | System | Endpoint | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | by TypeSafe AIJev 1.13.0 | 75.3 | 90.4 | 82.7 | 83.3 | 51.7 | $0.041 | 100.0% | 99.0% | 94.5% | 74.1% | 0.65 s rawp95 0.72 s raw | production API | |
| 2 | by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ | 74.6 | 85.9 | 72.6 | 83.7 | 59.2 | ~$0.023 est. | 100.0% | 97.9% | 95.2% | 59.5% | 0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 s | our RunPod GPU | |
| 3 | by Maisa (David Villalón)djev†Maisa, diffusion-gemma | 74.3 | 88.4 | 65.4 | 91.4 | 57.6 | $0.026 announced | 100.0% | 97.9% | 93.2% | 69.5% | 0.24 s rawp95 0.31 s raw | production API | |
| 4 | by IkerMoelopen-alternative-jev†Qwen3.5-4B, IkerMoel | 69.8 | 75.6 | 63.2 | 83.5 | 59.6 | ~$0.022 est. | 100.0% | 84.4% | 74.7% | 56.8% | 0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 s | our RunPod GPU | |
| 5 | by mithalounisystem-one-openGemma 4 E2B LoRA on an L4 | 68.7 | 79.5 | 56.7 | 77.0 | 64.1 | ~$0.016 est. | 100.0% | 93.8% | 87.7% | 49.1% | 0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 s | author's demo server | |
| 6 | by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback16 | 67.6 | 86.0 | 64.8 | 83.2 | 45.2 | ~$0.067 est. | 100.0% | 95.8% | 91.1% | 65.5% | 0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 s | our RunPod GPU | |
| 7 | by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang | 66.2 | 88.9 | 77.4 | 77.1 | 36.1 | ~$0.135 est. | 100.0% | 95.8% | 95.2% | 71.4% | 0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 s | author's demo server | |
| 8 | by OpenAIGPT-5.6 Lunalow reasoning effort | 66.0 | 96.8 | 89.8 | 77.5 | 28.2 | $0.247 | 100.0% | 97.9% | 96.6% | 94.5% | 0.97 s rawp95 1.82 s raw | production API | |
| 9 | by Kotoba Labsopen-jev-deberta-v3-largelocal CPU | 64.4 | 53.6 | 66.4 | 66.0 | 73.3 | ~$0.0077 est. | 100.0% | 49.0% | 53.4% | 36.4% | 1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 s | our CPU | |
| 10 | by Bespoke LabsBespoke Nimble 9B | 63.5 | 78.6 | 64.5 | 82.5 | 38.9 | ~$0.109 est. | 100.0% | 94.8% | 89.0% | 43.6% | 0.19 s raw→ 0.52 s adjustedp95 0.46 s raw → 1.07 s | our RunPod GPU | |
| 11 | by GoogleGemini 3.1 Flash-Lite | 60.8 | 90.3 | 68.1 | 81.8 | 27.1 | $0.268 | 100.0% | 99.0% | 93.2% | 75.0% | 0.76 s rawp95 0.88 s raw | production API | |
| 12 | by DeepSeekDeepSeek V4.1 Flashthinking default | 58.1 | 96.1 | 96.7 | 71.6 | 17.1 | $0.579 | 98.6% | 99.0% | 93.2% | 95.0% | 1.42 s rawp95 4.89 s raw | production API | |
| 13 | by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke | 56.5 | 80.1 | 36.8 | 84.4 | 41.2 | ~$0.092 est. | 100.0% | 90.6% | 91.8% | 50.0% | 0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 s | our RunPod GPU | |
| Partial runs — shown, not ranked: ranked = every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank. | ||||||||||||||
| by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked | 25.5 | 74.6 | 92.1 | 61.3 | 0.0 | ~$2.711 est. | 98.6% | 99.0% | 95.3% | 21.4% | 5.75 s rawp95 12.97 s raw | production API | ||
| by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked | 19.1 | 39.5 | none (label only) | 52.8 | 63.7 | ~$0.016 est. | 66.7% | 31.3% | 34.2% | — | 3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 s | our CPU | ||
| by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked | 16.7 | 22.4 | none (label only) | 59.9 | 58.1 | ~$0.025 est. | 47.2% | 16.7% | 31.5% | 7.7% | 1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 s | our CPU | ||
- † djev: Hosted API in free preview: the cost uses djev's announced price ($0.035 per million input tokens, output free); nothing is charged yet. Open-sourcing is planned, not yet released. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- † open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
Which public tasks did each system get right?
This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped. The pinned artifact has no public-task outcomes for djev yet, so its cells remain unavailable.
Show 231 public task outcomes across 16 systems
| Task | Jev 1.13.0 | SemIf | djev | open-alternative-jev | system-one-open | OpenJev | openjev-sglang | GPT-5.6 Luna | open-jev-deberta-v3-large | Bespoke Nimble 9B | Gemini 3.1 Flash-Lite | DeepSeek V4.1 Flash | system-one | Qwen3.8 27B | Needle 3, options as tools | Needle 3 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Easy · 48 public tasks | ||||||||||||||||
| easy-intent-00intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-01intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-intent-02intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-intent-03intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-04intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-intent-05intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-06intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-07intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-intent-08intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-09intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-10intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-11intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-00fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-01fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-02fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-03fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-fact-04fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-fact-05fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-fact-06fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-07fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-fact-08fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-09fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-fact-10fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-fact-11fact · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-extraction-00extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-01extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-02extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-03extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-04extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-05extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-06extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-07extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-08extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-09extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-10extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-11extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-00tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-01tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-02tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-03tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-04tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-05tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-06tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-07tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-08tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| easy-tool_selection-09tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-10tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-11tool_selection · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| Medium (standard) · 72 public tasks | ||||||||||||||||
| original-policy-01-0policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-01-1policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | × |
| original-policy-02-0policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| original-policy-02-1policy · noul | ✓ | × | — | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | × |
| original-policy-03-0policy · noul | ✓ | ✓ | — | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-03-1policy · noul | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | × |
| original-policy-04-0policy · noul | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-04-1policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-05-0policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-05-1policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-06-0policy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-06-1policy · noul | × | ✓ | — | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-intent-01-0intent · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-01-1intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| original-intent-02-0intent · choice | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × |
| original-intent-02-1intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-03-0intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-intent-03-1intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-04-0intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-04-1intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-05-0intent · choice | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-intent-05-1intent · choice | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-intent-06-0intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-06-1intent · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-ordinal-01-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-01-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-02-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-02-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-03-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-03-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-04-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-04-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-05-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-05-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-06-0ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-ordinal-06-1ordinal · score | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-01-0extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-01-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| original-extraction-02-0extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-02-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-03-0extraction · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| original-extraction-03-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-04-0extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-04-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-extraction-05-0extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-05-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-extraction-06-0extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-extraction-06-1extraction · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| original-adequacy-01-0adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-01-1adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-02-0adequacy · noul | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-adequacy-02-1adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-03-0adequacy · noul | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | × | ✓ | × | × |
| original-adequacy-03-1adequacy · noul | ✓ | ✓ | — | × | ✓ | × | × | ✓ | × | × | ✓ | ✓ | × | ✓ | × | × |
| original-adequacy-04-0adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-adequacy-04-1adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-adequacy-05-0adequacy · noul | ✓ | ✓ | — | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × |
| original-adequacy-05-1adequacy · noul | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | × |
| original-adequacy-06-0adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-06-1adequacy · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-routing-01-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-routing-01-1routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-02-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-02-1routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-03-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-03-1routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-04-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| original-routing-04-1routing · choice | ✓ | ✓ | — | × | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × |
| original-routing-05-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-05-1routing · choice | ✓ | ✓ | — | × | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| original-routing-06-0routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ! | ✓ | × |
| original-routing-06-1routing · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| Hard · 111 public tasks | ||||||||||||||||
| hard-opus-a-long_policy-01long_policy · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ! | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-a-long_policy-04long_policy · choice | × | × | — | × | × | × | × | ✓ | × | ! | × | ✓ | × | ✓ | · | ✓ |
| hard-opus-a-long_policy-08long_policy · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-a-long_policy-09long_policy · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-long_policy-11long_policy · noul | ✓ | ✓ | — | × | ✓ | × | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-a-long_policy-13long_policy · noul | × | × | — | × | × | × | × | ✓ | ✓ | ! | × | ✓ | × | ! | · | ✓ |
| hard-opus-a-long_policy-17long_policy · choice | ✓ | × | — | × | × | ✓ | × | ✓ | × | ! | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-long_policy-19long_policy · noul | × | ✓ | — | ✓ | × | × | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-probability-03probability · noul | ✓ | × | — | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-a-probability-04probability · choice | × | × | — | × | × | × | × | ✓ | × | × | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-probability-07probability · choice | × | × | — | × | × | ✓ | × | ✓ | × | × | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-probability-08probability · noul | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-a-temporal_numeric-03temporal_numeric · noul | × | ✓ | — | × | × | × | × | × | ✓ | × | × | ✓ | × | ✓ | · | ✓ |
| hard-opus-a-temporal_numeric-06temporal_numeric · noul | ✓ | × | — | ✓ | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-temporal_numeric-07temporal_numeric · choice | × | × | — | ✓ | × | ✓ | ✓ | ✓ | × | × | × | ! | × | ✓ | · | ✓ |
| hard-opus-a-temporal_numeric-09temporal_numeric · noul | × | ✓ | — | ✓ | × | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-a-temporal_numeric-12temporal_numeric · score | × | × | — | × | × | × | × | ✓ | × | × | × | ✓ | × | ✓ | · | ✓ |
| hard-opus-b-ambiguous-02ambiguous · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-b-ambiguous-03ambiguous · choice | × | × | — | ✓ | × | × | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-b-ambiguous-07ambiguous · choice | ✓ | × | — | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | · | ✓ |
| hard-opus-b-ambiguous-09ambiguous · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ | · | ✓ |
| hard-opus-b-ambiguous-10ambiguous · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-ambiguous-11ambiguous · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-b-ambiguous-13ambiguous · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-multi_hop-03multi_hop · choice | × | × | — | × | × | ✓ | ✓ | ✓ | × | ! | × | ✓ | × | ✓ | · | × |
| hard-opus-b-multi_hop-04multi_hop · choice | × | × | — | × | × | ✓ | × | ✓ | × | ! | × | ✓ | × | ✓ | · | × |
| hard-opus-b-multi_hop-05multi_hop · score | × | × | — | × | × | × | × | ✓ | × | ! | × | ✓ | × | ✓ | · | × |
| hard-opus-b-multi_hop-07multi_hop · choice | ✓ | × | — | × | × | ✓ | ✓ | ✓ | ✓ | ! | × | ✓ | × | ✓ | · | × |
| hard-opus-b-multi_hop-08multi_hop · noul | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-probability-01probability · choice | ✓ | ✓ | — | × | × | ✓ | × | ✓ | × | ! | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-probability-02probability · choice | × | × | — | × | × | × | ✓ | × | × | × | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-probability-03probability · noul | ✓ | × | — | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-b-probability-04probability · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-b-probability-06probability · choice | ✓ | × | — | × | × | × | × | ✓ | × | ✓ | ✓ | ! | × | ✓ | · | ✓ |
| hard-opus-b-tradeoff-01tradeoff · choice | × | ✓ | — | × | × | × | × | ✓ | × | × | × | ✓ | × | ✓ | · | × |
| hard-opus-b-tradeoff-03tradeoff · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-tradeoff-06tradeoff · noul | ✓ | × | — | × | × | × | ✓ | ✓ | × | × | × | ✓ | × | ✓ | · | ✓ |
| hard-opus-b-tradeoff-07tradeoff · choice | ✓ | × | — | × | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | · | × |
| hard-opus-b-tradeoff-08tradeoff · noul | ✓ | ✓ | — | × | × | × | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-b-tradeoff-12tradeoff · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-c-long_policy-02long_policy · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | ✓ | · | × |
| hard-opus-c-long_policy-03long_policy · choice | × | × | — | × | × | × | × | ✓ | × | ! | × | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-c-long_policy-04long_policy · noul | ✓ | ✓ | — | ✓ | ✓ | × | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | ✓ | · | ✓ |
| hard-opus-c-long_policy-05long_policy · score | × | × | — | × | × | ✓ | × | ✓ | × | ! | × | ✓ | × | ! | · | ✓ |
| hard-opus-c-long_policy-08long_policy · score | × | × | — | × | × | × | × | ✓ | × | ! | × | ✓ | × | ✓ | · | · |
| hard-opus-c-long_policy-10long_policy · choice | ✓ | × | — | × | × | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | ✓ | · | · |
| hard-opus-c-long_policy-11long_policy · noul | × | × | — | × | × | ✓ | × | ✓ | × | ! | ✓ | ✓ | × | ! | · | · |
| hard-opus-c-probability-03probability · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | · | · |
| hard-opus-c-temporal_numeric-02temporal_numeric · noul | × | × | — | × | ✓ | × | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | · | · |
| hard-opus-c-temporal_numeric-03temporal_numeric · choice | × | × | — | × | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | · | · |
| hard-opus-c-temporal_numeric-04temporal_numeric · choice | × | × | — | × | × | × | × | ✓ | × | × | × | ✓ | × | ✓ | · | · |
| hard-opus-c-temporal_numeric-06temporal_numeric · choice | × | × | — | × | × | × | × | ✓ | × | × | × | ✓ | ✓ | · | · | · |
| hard-opus-c-temporal_numeric-08temporal_numeric · choice | ✓ | × | — | × | × | × | ✓ | ✓ | × | × | × | ✓ | × | · | · | · |
| hard-opus-c-temporal_numeric-12temporal_numeric · choice | ✓ | ✓ | — | × | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | · | · | · |
| hard-sol-a-adversarial-01adversarial · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | · | · | · |
| hard-sol-a-adversarial-06adversarial · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-adversarial-07adversarial · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-a-adversarial-08adversarial · noul | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-adversarial-09adversarial · score | ✓ | × | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-adversarial-11adversarial · choice | ✓ | × | — | × | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | · | · | · |
| hard-sol-a-multi_hop-01multi_hop · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-multi_hop-02multi_hop · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-multi_hop-05multi_hop · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-multi_hop-07multi_hop · choice | ✓ | ✓ | — | × | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | · | · | · |
| hard-sol-a-multi_hop-08multi_hop · choice | ✓ | ✓ | — | × | × | ✓ | × | × | × | ! | × | ✓ | × | · | · | · |
| hard-sol-a-multi_hop-09multi_hop · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | · | · | · |
| hard-sol-a-multi_hop-10multi_hop · choice | ✓ | × | — | × | × | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | × | · | · | · |
| hard-sol-a-multi_hop-12multi_hop · choice | ✓ | × | — | × | ✓ | × | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-02trap · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-04trap · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-05trap · score | ✓ | ✓ | — | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-06trap · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-08trap · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | · | · | · |
| hard-sol-a-trap-10trap · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-13trap · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-a-trap-15trap · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-judge_hard-01judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-judge_hard-02judge_hard · noul | × | × | — | × | × | × | × | × | × | × | × | × | × | · | · | · |
| hard-sol-b-judge_hard-05judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-judge_hard-08judge_hard · noul | × | × | — | × | ✓ | × | ✓ | ✓ | × | × | × | ✓ | × | · | · | · |
| hard-sol-b-judge_hard-10judge_hard · noul | × | ✓ | — | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-b-judge_hard-14judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-judge_hard-15judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-judge_hard-18judge_hard · noul | × | ✓ | — | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | · | · | · |
| hard-sol-b-long_policy-01long_policy · choice | ✓ | × | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-long_policy-02long_policy · choice | ✓ | ✓ | — | × | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | × | · | · | · |
| hard-sol-b-long_policy-05long_policy · choice | ✓ | ✓ | — | ✓ | ✓ | × | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-long_policy-06long_policy · choice | ✓ | ✓ | — | × | ✓ | × | × | ✓ | ✓ | ! | ✓ | ! | × | · | · | · |
| hard-sol-b-routing_hard-01routing_hard · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-routing_hard-02routing_hard · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-routing_hard-03routing_hard · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-routing_hard-07routing_hard · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-routing_hard-09routing_hard · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-b-temporal_numeric-01temporal_numeric · choice | × | × | — | × | × | × | × | ✓ | × | × | × | ✓ | × | · | · | · |
| hard-sol-b-temporal_numeric-02temporal_numeric · choice | × | × | — | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | · | · | · |
| hard-sol-b-temporal_numeric-03temporal_numeric · choice | ✓ | × | — | × | ✓ | × | × | ✓ | × | × | × | ✓ | × | · | · | · |
| hard-sol-b-temporal_numeric-04temporal_numeric · choice | × | × | — | ✓ | × | × | ✓ | ✓ | ✓ | × | × | ✓ | × | · | · | · |
| hard-sol-c-judge_hard-03judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-c-judge_hard-05judge_hard · noul | ✓ | × | — | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-judge_hard-07judge_hard · noul | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-c-judge_hard-08judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-c-judge_hard-09judge_hard · noul | ✓ | × | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-judge_hard-10judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-judge_hard-11judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-judge_hard-13judge_hard · noul | ✓ | ✓ | — | × | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | · | · | · |
| hard-sol-c-judge_hard-15judge_hard · noul | ✓ | ✓ | — | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-multi_hop-07multi_hop · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | × | · | · | · |
| hard-sol-c-multi_hop-09multi_hop · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-multi_hop-11multi_hop · choice | ✓ | ✓ | — | ✓ | × | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
| hard-sol-c-multi_hop-12multi_hop · choice | ✓ | ✓ | — | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | · | · | · |
| hard-sol-c-multi_hop-13multi_hop · choice | ✓ | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | × | ! | ✓ | ✓ | ✓ | · | · | · |
✓ correct · × wrong · ! failed (scored wrong) · · not attempted · — no public outcome in the pinned artifact. Task descriptions are intentionally not included; the task id, tier and topic are the published public metadata.
How the JevBench Score works
JevBench Score = (Intelligence × Calibration × Speed × Cost)1/4, each axis on 0–100 — the geometric mean. A weak axis pulls the score down hard: a strong axis cannot buy it back.
- Intelligence — weighted accuracy: hard 30 %, easy 14 %, standard 28 %, judge 28 % (220 / 72 / 96 / 146 decisions).
- Calibration — on the hard tier: does “80 % sure” come true 80 % of the time, and does the returned distribution match the exact gold distribution on the probability items.
- Speed — median and 95th-percentile latency, one request at a time: 0.1 s scores 100, each 10× slower costs 20 points (1 s = 80, 10 s = 60). ⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.
- Cost — dollars per 1,000 decisions: $0.001 scores 100, each 10× more expensive costs 30 points ($0.01 = 70, $0.10 = 40, $1 = 10). Models without a tariff are priced at hosted-provider prices, marked “est.” (how).
Full scoring rules
- JevBench Score.
- exp(sum over the four axes of 0.25 x ln(max(axis, 1))) — the geometric mean of Intelligence, Calibration, Speed and Cost. A weak axis pulls the score down hard; a strong axis cannot buy it back.
- Intelligence.
- 100 x weighted accuracy: hard 30 %, easy 14 %, standard 28 %, judge 28 %. Accuracy = correct / all items; failed, timed-out or unparseable answers count as wrong.
- Calibration.
- Hard tier only, systems that return a probability distribution: mean of (a) 100 x (1 - ECE/0.5), ECE = top-label expected calibration error in 10 bins, and (b) probability fidelity = 100 x (1 - mean total-variation distance) between the returned distribution and the exact gold distribution on the 20 probability items. Label-only systems have none; it counts as 0 in the JevBench Score.
- Speed.
- Mean of score(p50) and score(p95) of the serial 242-decision standard+judge run; score(s) = 100 - 20 log10(s / 0.1 s), clipped to 0..100 (0.1 s = 100, 1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo. Production APIs (Jev, djev, OpenAI, Google, DeepSeek, Chutes) are not adjusted.
- Cost.
- Dollars per 1,000 decisions pooled over all 534 v1.2 decisions; score = 100 - 30 log10(usd / 0.001), clipped to 0..100 ($0.001 = 100, $0.01 = 70, $0.10 = 40, $1 = 10). Measured = public tariff x measured tokens. est. = hosted-provider list price of the same weights or size class x tokens. announced = the provider's published price, not yet charged (free preview), x measured tokens.
- Ranked.
- Ranked: every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank.
- Presets.
- Other views reweight the same four axes and combine them the same way (geometric mean). They are not the JevBench Score.
How the ranking moves with other weights
Rank and score under the JevBench Score and the earlier views, all combined as a geometric mean (Intelligence : Calibration : Speed : Cost). Highlighted = a different rank than the JevBench Score. Ranked systems only.
| System | JevBench Score25:25:25:25 · official | Balanced, no calibration33:0:33:33 · not the default | Emphasis on Accuracy60:0:20:20 · not the default | Emphasis on Speed20:0:60:20 · not the default | Emphasis on Cost20:0:20:60 · not the default |
|---|---|---|---|---|---|
| Jev 1.13.0 | #1 75.3 | #4 73.0 (rank differs from the JevBench Score) | #2 79.5 (rank differs from the JevBench Score) | #3 77.0 (rank differs from the JevBench Score) | #6 63.6 (rank differs from the JevBench Score) |
| SemIf | #2 74.6 | #2 75.2 | #3 79.3 (rank differs from the JevBench Score) | #2 78.5 | #3 68.3 (rank differs from the JevBench Score) |
| djev | #3 74.3 | #1 77.5 (rank differs from the JevBench Score) | #1 81.6 (rank differs from the JevBench Score) | #1 82.7 (rank differs from the JevBench Score) | #2 68.8 (rank differs from the JevBench Score) |
| open-alternative-jev | #4 69.8 | #5 72.2 (rank differs from the JevBench Score) | #6 73.5 (rank differs from the JevBench Score) | #4 76.5 | #5 66.9 (rank differs from the JevBench Score) |
| system-one-open | #5 68.7 | #3 73.2 (rank differs from the JevBench Score) | #4 75.7 (rank differs from the JevBench Score) | #5 74.7 | #1 69.4 (rank differs from the JevBench Score) |
| OpenJev | #6 67.6 | #6 68.6 | #5 75.1 (rank differs from the JevBench Score) | #6 74.1 | #7 58.1 (rank differs from the JevBench Score) |
| openjev-sglang | #7 66.2 | #10 62.8 (rank differs from the JevBench Score) | #8 72.2 (rank differs from the JevBench Score) | #9 68.1 (rank differs from the JevBench Score) | #10 50.3 (rank differs from the JevBench Score) |
| GPT-5.6 Luna | #8 66.0 | #11 59.6 (rank differs from the JevBench Score) | #7 72.4 (rank differs from the JevBench Score) | #11 66.2 (rank differs from the JevBench Score) | #11 44.2 (rank differs from the JevBench Score) |
| open-jev-deberta-v3-large | #9 64.4 | #8 63.8 (rank differs from the JevBench Score) | #13 59.5 (rank differs from the JevBench Score) | #12 64.6 (rank differs from the JevBench Score) | #4 67.4 (rank differs from the JevBench Score) |
| Bespoke Nimble 9B | #10 63.5 | #9 63.2 (rank differs from the JevBench Score) | #11 69.0 (rank differs from the JevBench Score) | #8 70.3 (rank differs from the JevBench Score) | #9 52.1 (rank differs from the JevBench Score) |
| Gemini 3.1 Flash-Lite | #11 60.8 | #12 58.5 (rank differs from the JevBench Score) | #10 69.6 (rank differs from the JevBench Score) | #10 66.9 (rank differs from the JevBench Score) | #12 43.0 (rank differs from the JevBench Score) |
| DeepSeek V4.1 Flash | #12 58.1 | #13 49.0 (rank differs from the JevBench Score) | #12 64.2 | #13 57.0 (rank differs from the JevBench Score) | #13 32.2 (rank differs from the JevBench Score) |
| system-one | #13 56.5 | #7 65.3 (rank differs from the JevBench Score) | #9 70.8 (rank differs from the JevBench Score) | #7 72.3 (rank differs from the JevBench Score) | #8 54.3 (rank differs from the JevBench Score) |
Compare two systems
Pick any two. The first radar shows the four axes of the JevBench Score (0–100, the values in the table above); the second shows accuracy by subject topic, over all tiers. Further out is better on every spoke.
- A: Jev 1.13.0 — Jev · JevBench Score 75.3 (#1)
- B: SemIf — Jev rebuild · JevBench Score 74.6 (#2)
The four score axes
Values as a table
| Axis | A: Jev 1.13.0 | B: SemIf |
|---|---|---|
| Intelligence | 90.4 | 85.9 |
| Calibration | 82.7 | 72.6 |
| Speed | 83.3 | 83.7 |
| Cost | 51.7 | 59.2 |
| JevBench Score | 75.3 | 74.6 |
Accuracy by subject topic
Values as a table
| Topic (items) | A: Jev 1.13.0 | B: SemIf |
|---|---|---|
| Math & numbers (129)a calculation decides the answer: arithmetic, word problems, probability, dates, units | 87.6% 113 of 129 | 79.1% 102 of 129 |
| Coding & software (56)code, SQL, repositories, developer tools and IT systems | 83.9% 47 of 56 | 96.4% 54 of 56 |
| Rules, policy & law (67)applying written rules: company policies, contracts, regulations, eligibility | 83.6% 56 of 67 | 64.2% 43 of 67 |
| Finance & commerce (64)money: payments, refunds, invoices, orders, expenses, insurance payouts | 73.4% 47 of 64 | 60.9% 39 of 64 |
| Support & operations (119)support tickets, incidents, logistics, scheduling desks and routing work to a team | 89.1% 106 of 119 | 87.4% 104 of 119 |
| Everyday language (79)short everyday messages: intents, assistant requests, reading a detail out of a text | 100.0% 79 of 79 | 100.0% 79 of 79 |
| Safety & security (20)untrusted or injected instructions, fraud, moderation, access and security triage | 100.0% 20 of 20 | 75.0% 15 of 20 |
Systems with a public tariff use that tariff and measured tokens. For systems without one, we use a clearly marked estimate based on a large inference provider's list price for the same weights or size class.
How costs are estimated
Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.
- SemIf — ~$0.023 est. per 1,000: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 426 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
- open-alternative-jev — ~$0.022 est. per 1,000: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
- system-one-open — ~$0.016 est. per 1,000: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 452 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
- OpenJev — ~$0.067 est. per 1,000: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 410 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
- openjev-sglang — ~$0.135 est. per 1,000: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 667 input and 2 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
- open-jev-deberta-v3-large — ~$0.0077 est. per 1,000: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
- Bespoke Nimble 9B — ~$0.109 est. per 1,000: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B)) x 990 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3.5-9b $0.1/M in, $0.15/M out x 1215 in / 2 out tokens per hard decision
- system-one — ~$0.092 est. per 1,000: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 443 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
- Qwen3.8 27B — ~$2.711 est. per 1,000: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 445 input and 393 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
- Needle 3, options as tools — ~$0.016 est. per 1,000: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 452 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed.
- Needle 3 — ~$0.025 est. per 1,000: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 452 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
- dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
- dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
- dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
- encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
- generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
- moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
- moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9
Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).
Who could not be measured, and why
An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank.
- SemIf / openjev (Theodore Lee (TheoLeeCJ)) — README: "Python 3.10+, CUDA, and a GPU that can hold a 4B BF16 model". The alternative is a browser-only WebGPU demo, which is not a programmable endpoint. No GPU available (see above).
- open-jev (Dasein Labs) — MLX on Apple Silicon. The README itself says "there is no container path because Linux containers cannot reach the Apple GPU". We have no Apple hardware.
- open-jev (JoshuaSP) — DiffusionGemma 26B-A4B measured on an H100 through the author's own Modal account. No public endpoint, no GPU on our side.
- mini-jev (Mikhail Rakutko (r-ms)) — README: "~9 GB of disk for the weights … the 4B model needs about 8.5 GB of memory" plus Apple Silicon or CUDA. Over our 4 GB memory bound and our disk headroom.
- system-one-gemma (Akash Kamat) — The base model is gated: the README requires accepting Google's Gemma licence on Hugging Face first. We do not accept binding terms on Florian's behalf.
- jevlike (Vincent Wang-Maścianica) — The released checkpoints are the Doom and chess vision scorers; there is no released general text-decision checkpoint to run against this suite.
- AlexWortega/openjev (Alex Wortega) — A three-way NLI head (entailment / contradiction / neutral) plus task-specific heads. Turning that into a distribution over our arbitrary label sets needs an assumption we would then be measuring instead of the model.
- GLiNER2 (Fastino) — A multi-label classification head. Its per-label scores are not a categorical posterior over our label set without choosing a normalization, and that choice would drive the calibration number. Kept as a candidate for a later version with a documented mapping.
- Succinct Router 14M (Pedro Marques) — A trained router over three fixed GPT settings, not a general typed-decision interface. In MARKET.md, not in the suite.
- jev-model-router, Director, Loki (various) — Applications built on a decision model, not decision models. In MARKET.md.
Method and tiers
exp(sum over the four axes of 0.25 x ln(max(axis, 1))) — the geometric mean of Intelligence, Calibration, Speed and Cost. A weak axis pulls the score down hard; a strong axis cannot buy it back.
Revision v1.2.1. v1.2 final: 4 axes (Intelligence, Calibration, Speed, Cost), 25 % each, geometric mean; hard tier 30 % of Intelligence; latency of non-production endpoints adjusted (assumption); one open-alternative-jev row (author's option order); Needle 3 options-as-tools priced. v1.2.1: added djev (Maisa, diffusion-gemma).
- easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
- standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
- judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
- hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.
v1.2 numbers are not comparable with v1.1 or v1.0 (different tiers and scoring). The v1.0 page keeps its own numbers, calibration plots and per-family tables.
Limits
- 534 decisions is a pilot, not a census, and it is English-only.
- The weights are a choice. The JevBench Score weights the four axes equally and multiplies rather than adds them; if a wrong decision costs you more than a slow or expensive one, pick “Emphasis on Accuracy” above — the table of views shows what other weightings would do.
- The latency adjustment (×2, +0.15 s on our own servers) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. The official Jev API is presumably under high load, given the public interest. Serving under load trades per-user speed for throughput: in the NVIDIA chart shown by SemiAnalysis, moving to the throughput-maximising setting cuts per-user tokens per second by far more than 2×. That chart is a 1.8T mixture-of-experts model on GPU clusters, not a 4B model on one GPU, so it supports the direction and size of the effect, not our exact factor. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo, and a measurement under load is planned.
- Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
- Latency is one origin at one time of day; hosted endpoints, public demos and a local CPU are different kinds of latency. Public demo endpoints are shared with everyone else using them.
- Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit
Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.
- Bespoke Nimble 9B (Bespoke Labs) — Bespoke Labs, Apache-2.0 (weights); repository without a licence file as of 19 Sep — github.com/bespokelabsai/nimble
- DeepSeek V4.1 Flash (thinking default) — DeepSeek, open weights, proprietary API route — api-docs.deepseek.com
- djev (Maisa, diffusion-gemma) — Maisa (David Villalón), hosted API; open-sourcing announced, not yet released — djev.dev
- Gemini 3.1 Flash-Lite — Google, proprietary API — ai.google.dev
- GPT-5.6 Luna (low reasoning effort) — OpenAI, proprietary API — platform.openai.com
- Jev 1.13.0 (TypeSafe AI) — TypeSafe AI, proprietary API — docs.typesafe.ai
- Needle 3 (Cactus, 2-bit, local CPU) — Cactus Compute, Apache-2.0 (model and package) — github.com/cactus-compute/needle
- open-alternative-jev (Qwen3.5-4B, IkerMoel) — IkerMoel, Apache-2.0 (code and weights) — github.com/ikermoel/open-alternative-jev
- open-jev-deberta-v3-large (local CPU) — Kotoba Labs, Apache-2.0 (model card); DeBERTa-v3 keeps its own terms — github.com/kotoba-lang/typed-decisions
- OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) — razorback16 / Codiv, Apache-2.0 (repo and weights) — github.com/razorback16/openjev
- openjev-sglang (Qwen3.6-35B-A3B on SGLang) — ekzhang, no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms — github.com/ekzhang/openjev-sglang
- Qwen3.8 27B (Chutes TEE) — Qwen / Chutes, open weights — chutes.ai
- SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) — Theodore Lee (TheoLeeCJ), MIT (code); Qwen3.5 weights Apache-2.0 — github.com/TheoLeeCJ/openjev
- system-one (Qwen3-8B, Sean Goedecke) — Sean Goedecke, no licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0 — github.com/sgoedecke/system-one
- system-one-open (Gemma 4 E2B LoRA on an L4) — mithalouni, MIT (repository LICENSE; Gemma weights keep Google’s terms) — github.com/mithalouni/system-one-open
Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.