Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.
Capability and Jev-class (headline ranking). Capability = (Intelligence + Calibration) / 2. A system is Jev-class if its cost per decision is at most 2× Jev 1.13.0's and its median latency (the adjusted p50 — the median the speed chart plots, not the Speed-axis score) is at most 2× Jev 1.13.0's; rows without a recorded median latency use the Speed axis at the 2× equivalent (77.2). The headline ranks Jev-class systems by Capability; the others, including general-purpose LLMs, are listed below a divider. This is a way of presenting the same measurements; it changes no score and no official rank.
JevBench Score. Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost (power mean p=-1). If Intelligence <50 multiply by (I/50)^2. For Speed and Cost separately, if below 50 multiply by (axis/50)^2. The weight sliders re-score the same axes with other weights for exploration; only equal weights give the official score and rank.
Intelligence. 0.8 × v1.3 chance-corrected Intelligence on the frozen v1.2 items + 0.2 × 100 × max(0, (sealed accuracy − 0.293)/(1 − 0.293)); then multiply by 1 − max(0, public-minus-sealed accuracy gap in percentage points − 25)/100. Original tier weights: easy .14, standard .28, judge .28, hard .30.
Revision v1.4.2.2. v1.4.2.2 adds the verified Imajev-4B row to v1.4.2.1 using the exact live v1.4.2 scoring code. Earlier measurements and score fields are unchanged; only ranks and preset ranks move where the new row changes the ordering. A 28 Sep text-only amendment clarifies that full-forward input tokens are counted once for the single pinned server pass; no measurement, score, axis, rank, eligibility, or numeric value changed.
- easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
- standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
- judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
- hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
- sealed: 308 fresh private decisions across ten families, run once per system. Only system-level aggregates — overall and per-family accuracy, calibration — are published; the item text, answers and per-item results stay private and rotate between versions.
Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.
A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.
v1.4 scores are not comparable with v1.3 or earlier (sealed blend, gap penalty and harmonic mean). The v1.3.0 board stays below as history; the v1.0 page keeps its own numbers, calibration plots and per-family tables.