This is the archived v1.0 run (242 decisions, five axes, calibration plots). Current results: JevBench v1.2 with the JevBench Score →
JevBench v1.0 · our own benchmark · pilot
Jev-class models
Typed decisions compared on accuracy, cost, latency, reliability and openness.
JevBench v1 is our own benchmark — built and run by Benchmark Heaven, not collected from someone else's leaderboard. These results describe the tested configurations and tasks, not every application or a vendor-wide ranking. Five axes, deliberately no combined score: the trade-off between them is the point.
Measured 19 Sept 2026 · protocol jevbench::v1 · 242 decisions per system (72 published · 24 held out · 146 imported) · one request at a time from a Hetzner server in Germany · harness & published decisions (MIT) · results JSON sha256 38fc5f1d6fd8…
What the run says
- GPT-5.6 Luna (low reasoning effort) leads at 97.1%, Jev 1.13.0 (TypeSafe AI) follows at 96.3% — their 95% intervals overlap, so read the top rows as a group, not a podium.
- Best open rebuild: openjev-sglang (Qwen3.6-35B-A3B on SGLang) at 95.5%.
- Cheapest metered route: Jev 1.13.0 (TypeSafe AI) at $0.027 per 1,000 decisions. 3 of 7 complete runs have no per-token tariff to us, which is not the same as free.
- Fastest median: system-one-open (Gemma 4 E2B LoRA on an L4), Jev 1.13.0 (TypeSafe AI), openjev-sglang (Qwen3.6-35B-A3B on SGLang) (0.65 s to 0.68 s — indistinguishable), measured serially from one server in Germany.
- Best calibrated: DeepSeek V4.1 Flash (thinking default) (ECE 0.009, verbalized probabilities).
- Only models that write their probabilities needed renormalizing: GPT-5.6 Luna (low reasoning effort) (20), DeepSeek V4.1 Flash (thinking default) (4), Qwen3.8 27B (Chutes TEE) (1). A native distribution sums to 1 by construction.
- open-jev-deberta-v3-large (local CPU) did not see the whole question 80 times out of 242: its input window is shorter than several requests, so its 51.7% is a context limit as much as a judgement one.
Up is more accurate, left is cheaper. ● native probabilities · ● verbalized. Complete runs only.
Not plotted: openjev-sglang, system-one-open, open-jev-deberta-v3-large — no per-token tariff on the measured route, so there is no cost to place — never drawn as free. Runs that stopped early (Qwen3.8 27B) answered a different subset and are left out of both charts.
Stated confidence against how often the top answer was right, ten fixed bins; on the diagonal is perfectly calibrated. Dot size = decisions in the bin; empty bins are left out, never interpolated. ECE 0.016 · verbalized probabilities.
Bin table
| Bin | n | Mean confidence | Accuracy |
|---|---|---|---|
| 0.0–0.1 | 0 | — | — |
| 0.1–0.2 | 0 | — | — |
| 0.2–0.3 | 0 | — | — |
| 0.3–0.4 | 0 | — | — |
| 0.4–0.5 | 0 | — | — |
| 0.5–0.6 | 0 | — | — |
| 0.6–0.7 | 1 | 69.7% | 100.0% |
| 0.7–0.8 | 3 | 71.5% | 100.0% |
| 0.8–0.9 | 5 | 85.7% | 60.0% |
| 0.9–1.0 | 233 | 98.5% | 97.9% |
All five axes
Sort by any axis; nothing is combined. Unknown values always sort last. Open a row for per-family accuracy, settings, price receipts and licence.
| System | Validanswers | Same answeron a rephrasing | |||||
|---|---|---|---|---|---|---|---|
| 97.1%94.7%–98.8% | $0.176 | 0.97 sp95 1.82 s | 0.016verbalized | 100.0% | 94.4%36 pairs | Closed | |
| 96.3%93.7%–98.4% | $0.027 | 0.65 sp95 0.72 s | 0.027native | 100.0% | 97.2%36 pairs | Closed | |
| 95.5%92.5%–97.9% | $0.245over 234/242 metered | 1.42 sp95 4.89 s | 0.009verbalized | 96.7% | 100.0%36 pairs | Open weights | |
| 95.5%92.6%–97.9% | $0.195 | 0.76 sp95 0.88 s | 0.041verbalized | 100.0% | 97.2%36 pairs | Closed | |
| 95.5%92.8%–97.8% | no tariff | 0.68 sp95 0.73 s | 0.042native | 100.0% | 88.9%36 pairs | Open | |
| 90.1%86.3%–93.7% | no tariff | 0.65 sp95 0.77 s | 0.068native | 100.0% | 86.1%36 pairs | Open | |
| 51.7%45.0%–58.4% | no tariff | 1.77 sp95 3.35 s | 0.147native | 100.0% | 83.3%36 pairs | Open | |
| Stopped early — shown with its own denominator, not ranked against complete runs: | |||||||
| 96.9%94.4%–99.1% | no tariff | 5.75 sp95 12.97 s | 0.014verbalized | 97.8% | 100.0%36 pairs | Open weights | |
Who could not be measured, and why
A benchmark that quietly drops what it could not run is a benchmark you cannot check. An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict.
- open-alternative-jev (Qwen3.5-4B, HF Space) (IkerMoel) — reached, but the run stopped after 6 of 242 decisions: AppError: You have exceeded your ZeroGPU runs limit. Subscribe to Hugging Face PRO to get 40 min of ZeroGPU quota a day - https://huggingface.co/subscribe/pro?from=ZeroGPU. A handful of answers is evidence that we tried, not a measurement, so it is not ranked.
- Bespoke Nimble 9B (Bespoke Labs) — Needs an NVIDIA GPU (~18 GB for the 9B weights, per the README) and there is no public endpoint. Our RunPod API returns HTTP 403 with Cloudflare error code: 1010, a bot rejection; we do not change client identity to get past one, so no GPU could be rented.
- SemIf / openjev (Theodore Lee (TheoLeeCJ)) — README: "Python 3.10+, CUDA, and a GPU that can hold a 4B BF16 model". The alternative is a browser-only WebGPU demo, which is not a programmable endpoint. No GPU available (see above).
- open-jev (Dasein Labs) — MLX on Apple Silicon. The README itself says "there is no container path because Linux containers cannot reach the Apple GPU". We have no Apple hardware.
- open-jev (JoshuaSP) — DiffusionGemma 26B-A4B measured on an H100 through the author's own Modal account. No public endpoint, no GPU on our side.
- OpenJev (razorback16 / Codiv) — README: "at least 24 GB of memory for the NVFP4 checkpoint". The hosted route, codiv.ai, returns HTTP 403 to us and we hold no account there.
- mini-jev (Mikhail Rakutko (r-ms)) — README: "~9 GB of disk for the weights … the 4B model needs about 8.5 GB of memory" plus Apple Silicon or CUDA. Over our 4 GB memory bound and our disk headroom.
- system-one (Sean Goedecke) — Frozen Qwen3-8B scored locally; same memory and GPU wall, no public endpoint.
- system-one-gemma (Akash Kamat) — The base model is gated: the README requires accepting Google's Gemma licence on Hugging Face first. We do not accept binding terms on Florian's behalf.
- jevlike (Vincent Wang-Maścianica) — The released checkpoints are the Doom and chess vision scorers; there is no released general text-decision checkpoint to run against this suite.
- AlexWortega/openjev (Alex Wortega) — A three-way NLI head (entailment / contradiction / neutral) plus task-specific heads. Turning that into a distribution over our arbitrary label sets needs an assumption we would then be measuring instead of the model.
- Needle 3 (Cactus Compute) — Its interface returns a chosen label plus one accept/refuse confidence, not a distribution over the exact label set, and we do not synthesize a distribution from a confidence scalar. Measured separately on the same 78 routing tasks in our earlier head-to-head (~/jobs/needle3-vs-jev-20260918): 0/78 category accuracy, ~4.2 s median on the same CPU.
- GLiNER2 (Fastino) — A multi-label classification head. Its per-label scores are not a categorical posterior over our label set without choosing a normalization, and that choice would drive the calibration number. Kept as a candidate for a later version with a documented mapping.
- Succinct Router 14M (Pedro Marques) — A trained router over three fixed GPT settings, not a general typed-decision interface. In MARKET.md, not in the suite.
- jev-model-router, Director, Loki (various) — Applications built on a decision model, not decision models. In MARKET.md.
What would change this list: a public endpoint, a CPU-runnable checkpoint under ~3 GB, or a GPU budget. Authors who want their project measured can tell us what to run; a later entry becomes v1.1 rather than silently changing v1's cohort.
Method
Every system sees the same state, the same instructions, the same rubric and the same exact label set. Only the transport differs.
The native adapters read the model's own probability distribution — a single forward pass, nothing generated. The openai_compat adapter asks an ordinary instruction model to write a distribution under a JSON schema. Those are different objects. They are labelled native and verbalized everywhere on this page and in the artifact, and they are never pooled into one calibration claim. Token-level logprobs are not used for anyone.
Requests go out one at a time from a Hetzner server in Germany, with no retries and no concurrency, so the latency you see includes the network. The first request to each system is reported separately, because a scale-to-zero endpoint bills its cold start to whoever knocks first.
Accuracy is argmax over the exact label set. Confidence intervals resample whole scenarios rather than individual decisions, because a paraphrase pair is one scenario asked twice. Brier is the multi-class sum over the label set. ECE is top-label confidence in ten equal-width bins, and empty bins are absent rather than zero.
Price is the provider's own published tariff, read on the run day, multiplied by the token usage that provider reported. It is marked derived_usage_times_tariff in the artifact, not presented as an invoice. A route with no billable account — a public demo, a flat-rate subscription, open weights on our own CPU — has no per-token tariff; the compute is still real, so it is shown as “no tariff”, never as $0.
The suite: 72 decisions we wrote and publish in the repo · 24 decisions we wrote and keep private, so the suite cannot be trained on · 146 decisions imported from our auto-router job's labelled tasks (source text not redistributed). Private items, labels and raw replies never reach this page or its JSON; only aggregate counts and whole-split hashes are published (original 72: 5c2414edb3… · heldout 24: 8766a0cccb… · router 78: d845fe3672… · judge 68: e0b9805689…). The held-out decisions are sent to the evaluated services to get predictions, so the split is not contamination-proof.
How much of a gap is noise? Jev answered the same 242 decisions twice, about 16 minutes apart. 3 answers changed (1.2% of the suite), all in routing, and accuracy moved from 96.7% to 96.3%. These endpoints are not deterministic: read a gap of about a point between two rows as noise and use the confidence intervals.
One protocol change, stated openly. v1 froze a 0.001 tolerance on “do the probabilities sum to 1” before the run. The run showed that this mostly measures rounding: models that write probabilities to three decimals land on 0.999 for a nine-option question. So the headline renormalizes any distribution that sums to within 2% of 1, uniformly for every system, and the page reports both — “valid answers” under the headline rule and “exact-sum answers” under the original one (in each row's details). Distributions outside the 2% band are still invalid and still count as wrong.
Limits
- 242 decisions is a pilot, not a census, and it is English-only.
- The answer-adequacy family has a majority-class floor of 82.1%; read that family against its floor, which the per-family table shows.
- Two of the instruction-model baselines judge some of their own earlier answers in that family, because the saved answers came from four models and two of them are also measured here.
- The held-out decisions are sent to the services being evaluated in order to get predictions. Not public is not the same as not seen. This is not a contamination proof.
- Latency is one origin at one time of day. A hosted endpoint and a local CPU are not the same kind of latency and should not be read as one ranking.
- The public demo endpoints are shared with everyone else using them. Their numbers describe that deployment on that day, not the model's ceiling on your hardware.
- An earlier head-to-head on 78 routing tasks (Needle 3 vs Jev, 18 Sep) used a different protocol; it is kept in the artifact's legacy appendix and never ranked against v1.
Credit
Every open rebuild here is someone's weekend project published for free, and several of them run on their author's own money. Links go to their repositories.
- Jev 1.13.0 (TypeSafe AI) — TypeSafe AI, proprietary API — docs.typesafe.ai
- open-alternative-jev (Qwen3.5-4B, HF Space) — IkerMoel, Apache-2.0 (repository); Qwen3.5 weights keep their own terms — github.com/ikermoel/open-alternative-jev
- open-jev-deberta-v3-large (local CPU) — Kotoba Labs, Apache-2.0 (model card); DeBERTa-v3 keeps its own terms — github.com/kotoba-lang/typed-decisions
- openjev-sglang (Qwen3.6-35B-A3B on SGLang) — ekzhang, no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms — github.com/ekzhang/openjev-sglang
- system-one-open (Gemma 4 E2B LoRA on an L4) — mithalouni, MIT (repository LICENSE; Gemma weights keep Google’s terms) — github.com/mithalouni/system-one-open
Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become v1.1 rather than silently changing v1's cohort.