JevBench · fresh regular draw
The first regular-queue release on a fresh seeded draw: 8 self-hosted open systems each answered the same 1,500 decisions of draw v1.6-regular-20261010-a2 (1,200 sealed plus the 300-item public set). Every drawn item had no prior scored use, and the set does not overlap v1.6.0 or the paid fast-lane draw. All 8 rows are complete and ranked. Historical scores keep their original dates and are listed separately below.
Aggregate results JSON · SHA-256 1f97a539745a95cec0bb2feac754788cffa4ef35f689ad92b1e75441d8acfbfc · Previous v1.6.4 paid cohort · Full historical board
JevBench v1.6.5 · headline
JevBench Capability Score
Capability ranking of Jev-class systems
Capability Score averages Intelligence and Calibration. Wald 4B v2.1 leads the Jev-class systems with 69.3.
Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘
Adjust cost / latency caps · 2× official
ScoreCCost
ScoreCap.Capability
Score$/1k$/1k decisions
- 1Wald 4B v2.152.046.369.3$0.062*
- 2Liquid AI d1-3B22.150.649.5$0.044*
- 3Vega 4B16.747.237.7$0.058*
- 4Vega 0.8B3.161.533.0$0.019*
- 5Liquid AI d1-omni-600M3.373.132.8$0.0079*
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = -0.05, n = 8).
- Open weights · LLM decoder (7)
- Open weights · encoder / classifier (1)
- green: ≤ reference
- amber: ≤ cap (2× reference)
- red: > cap
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).
- –RSI-Jev v6.1-VL 27B61.631.173.5$0.20*Outside: cost 6.1× Jev (v1.5 reference)
- –Standard One 8B SH34.543.260.0$0.078*Outside: cost 2.4× Jev (v1.5 reference)
- –Gutsy 0.8B v0.38.678.443.1$0.0053*Outside: latency 2.7× Jev (v1.5 reference)
Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 5 of 8 systems qualify; the other 3, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed
Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
- Open weights · LLM decoder (7)
- Open weights · encoder / classifier (1)
- faint = outside Jev-class
Model kind
Jev-class
Release
No family reported in this release.
No system in this release reports an exact parameter count.
No system in this release reports a developer/API price.
No system in this release reports a base-model reference price.
No system in this release reports an alternative pricing scenario.
JevBench v1.6.5
JevBench Composite Score: 8 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Adjust weights ↓
The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis, and the column headings sort too.
Greener = stronger within its column.
8 of 8 systems, sorted by official rank, #1 first.
- 1Wald 4B v2.1newfine-tune · 4BBase model: undisclosed54.3I 52C 87S 93K 46est.$0.062
- 2RSI-Jev v6.1-VL 27Bnewmerge · 27BBase model: undisclosed21.6I 62C 85S 85K 31est.$0.197
- 3Standard One 8B SHnewfine-tune · 8BBase model: undisclosed18.7I 34C 85S 84K 43est.$0.078
- 4Liquid AI d1-3Bnewfine-tune · 3BBase model: undisclosed8.9I 22C 77S 94K 51est.$0.044
- 5Vega 4Bnewadapter · 4BBase model: undisclosed3.6I 17C 59S 84K 47est.$0.058
- 6Gutsy 0.8B v0.3newfine-tune · 0.8B · GGUF Q8_0Base model: undisclosed0.8I 9C 78S 73K 78est.$0.0053
- 7Liquid AI d1-omni-600Mnewfine-tune · 600MBase model: undisclosed0.0I 3C 62S 94K 73est.$0.0079
- 8Vega 0.8Bnewadapter · 0.8BBase model: undisclosed0.0I 3C 63S 89K 61est.$0.019
Adjust weights ↓
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Open weights · LLM decoder (7)
- Open weights · encoder / classifier (1)
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
- A: Wald 4B v2.1 — Open weights · LLM decoder · Score 54.3 (#1)
- B: RSI-Jev v6.1-VL 27B — Open weights · LLM decoder · Score 21.6 (#2)
The four score axes
Capability by subject topic
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
What each category means · items per category
- Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 711 items (151 open / 560 sealed)
- Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 334 items (37 open / 297 sealed)
- Coding & software — code, SQL, repositories, developer tools and IT systems. 160 items (62 open / 98 sealed)
- Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 125 items (21 open / 104 sealed)
- Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 71 items (12 open / 59 sealed)
- Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 54 items (11 open / 43 sealed)
- Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 45 items (6 open / 39 sealed)
Use cases (TypeSafe categories)
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Low sample, n < 30 — indicative only
| Category (items) | A: Wald 4B v2.1 | B: RSI-Jev v6.1-VL 27B |
|---|---|---|
| LLM guardrails (29) | 70.7 n=29 | 89.9 n=29 |
| Gaming (29) | 29.1 n=29 | 25.4 n=29 |
| Lead generation (28) | 35.3 n=28 | 61.4 n=28 |
| Demand forecasting (26) | 7.6 n=26 | 0.0 n=26 |
| Advertising (24) | 47.4 n=24 | 37.0 n=24 |
| Moderation and trust and safety (24) | 45.2 n=24 | 66.6 n=24 |
| Semantic code linting (24) | 41.5 n=24 | 69.0 n=24 |
| Graphs and knowledge graphs (23) | 34.6 n=23 | 42.7 n=23 |
| Insurance claims (22) | 48.3 n=22 | 47.2 n=22 |
| Search and retrieval (22) | 76.7 n=22 | 87.0 n=22 |
| Recruiting (21) | 38.8 n=21 | 54.0 n=21 |
Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.
What each category means · items per category
- Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 319 items (74 open / 245 sealed)
- Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 273 items (86 open / 187 sealed)
- Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 255 items (24 open / 231 sealed)
- Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 155 items (37 open / 118 sealed)
- Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 60 items (9 open / 51 sealed)
- E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 58 items (9 open / 49 sealed)
- Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 46 items (8 open / 38 sealed)
- Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 32 items (2 open / 30 sealed)
- Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 30 items (5 open / 25 sealed)
Competence per request type, open / sealed
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — open set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — sealed set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
How the categories were made
Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).
Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 items of this draw was labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use case by Winnow-12B Q8 on our own GPU pod (same model, questions, taxonomy and T1-T3/U1/U2 item-group rules as v1.5 and v1.6.4); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items were checked by hand afterwards and no rule was derived from them: topic agreement 71/75 = 94.7 %. 17 of 27 router_policy items are labelled Coding rather than Rules, policy & law. All non-English uc1 items are machine-authored and not native-reviewed.
- Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
- Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
| Spoke | A: Wald 4B v2.1 | B: RSI-Jev v6.1-VL 27B |
|---|---|---|
| The four score axes | ||
| Intelligence | 52.0 | 61.6 |
| Calibration | 86.6 | 85.3 |
| Speed | 93.1 | 84.7 |
| Cost | 46.3 | 31.1 |
| Capability by subject topic | ||
| Rules, policy & law | 54.8 | 71.2 |
| Math & numbers | 32.1 | 26.0 |
| Coding & software | 70.2 | 79.8 |
| Support & operations | 38.0 | 30.9 |
| Finance & commerce | 45.3 | 50.2 |
| Everyday language | 21.7 | 33.1 |
| Safety & security | 53.2 | 59.4 |
| Use cases (TypeSafe categories) | ||
| Legal & compliance | 59.2 | 76.5 |
| Model routing | 66.1 | 78.7 |
| Other | 28.1 | 20.0 |
| Customer support | 47.5 | 47.8 |
| Financial crime | 34.2 | 52.9 |
| E-commerce | 60.8 | 35.1 |
| Risk assessment | 63.5 | 65.8 |
| Science | 6.8 | 25.8 |
| Feature extraction | 31.8 | 64.3 |
| Competence per request type, open / sealed | ||
| Choice · open | 71.0 | 75.8 |
| Choice · sealed | 55.1 | 70.3 |
| Noul · open | 66.2 | 66.2 |
| Noul · sealed | 53.1 | 41.7 |
| Score · open | 33.0 | 63.5 |
| Score · sealed | 33.4 | 51.9 |
| Competence per tier — open set | ||
| Easy | 81.6 | 90.3 |
| Standard | 60.8 | 63.9 |
| Judge | 59.1 | 60.2 |
| Hard | 54.8 | 73.6 |
| Competence per tier — sealed set | ||
| Easy | 53.4 | 75.9 |
| Standard | 51.1 | 65.7 |
| Judge | 52.0 | 64.3 |
| Hard | 44.9 | 46.4 |
All eight rows are self-hosted on evaluator-owned pods; there is no API lane in this release, so no equating was applied to any published cell. Sealed counts refer to self-hosted S (1,200); API rows use their original measured sealed basis.
Languages
Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1200 + P 300; API rows use their measured sealed subset plus P. Sage (A3) language cells and A2/A3 topic/use-case cells cover public P300 only. Raw and outside the Composite. Cells under 15 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.5. English (1,119 items) is listed first; the other 21 languages and the mixed-language group share 381 items. In 12 further groups no system reaches the 15-item reporting minimum, so they get no column (items in the pool shown): Dutch (13), Korean (11), Turkish (11), Ukrainian (11), Arabic (11), Mixed-language (11), Finnish (9), Chinese (8), Czech (8), Swedish (8), Norwegian (7), Indonesian (6) — 114 items, scored like every other item.
| System | en1119 | de45 | pl40 | ja32 | hi30 | es†28 | da†24 | pt†18 | el†18 | it†17 | fr†15 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Wald 4B v2.1self-hosted · S+P | 50 | 55 | 34 | 49 | 25 | 72† | 43† | 75† | 12† | 54† | 54† |
| RSI-Jev v6.1-VL 27Bself-hosted · S+P | 57 | 72 | 51 | 32 | 53 | 83† | 44† | 83† | 0† | 90† | 42† |
| Standard One 8B SHself-hosted · S+P | 30 | 45 | 41 | 9 | 4 | 44† | 23† | 77† | 0† | 43† | 0† |
| Liquid AI d1-3Bself-hosted · S+P | 14 | 6 | 30 | 0 | 1 | 34† | 0† | 55† | 0† | 41† | 29† |
| Vega 4Bself-hosted · S+P | 11 | 4 | 43 | 6 | 0 | 24† | 22† | 15† | 0† | 42† | 10† |
| Gutsy 0.8B v0.3self-hosted · S+P | 0 | 2 | 12 | 0 | 0 | 14† | 0† | 27† | 0† | 28† | 0† |
| Liquid AI d1-omni-600Mself-hosted · S+P | 0 | 0 | 0 | 0 | 0 | 0† | 0† | 0† | 0† | 0† | 0† |
| Vega 0.8Bself-hosted · S+P | 0 | 0 | 0 | 0 | 0 | 0† | 0† | 22† | 0† | 5† | 0† |
Intelligence gate and Noul decisiveness
The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.
| System | Intelligence | Choice | Noul | Score | Noul decisive rate | Accuracy among decisive |
|---|---|---|---|---|---|---|
| Wald 4B v2.1 | 52.0 | native | native | native | 100 % | 78 % |
| RSI-Jev v6.1-VL 27B | 61.6 | native | native | native | 83 % | 91 % |
| Standard One 8B SH | 34.5below gate | native | native | native | 61 % | 88 % |
| Liquid AI d1-3B | 22.1below gate | native | native | native | 49 % | 83 % |
| Vega 4B | 16.7below gate | native | native | native | 70 % | 68 % |
| Gutsy 0.8B v0.3 | 8.6below gate | native | native | native | 29 % | 74 % |
| Liquid AI d1-omni-600M | 3.3below gate | native | native | native | 27 % | 56 % |
| Vega 0.8B | 3.1below gate | native | native | native | 32 % | 63 % |
All data (8 systems)
Method notes · fresh native cohort · v1.6.5
Each system answered the same 1,200 freshly drawn sealed decisions and 300 public decisions on an offline GPU pod. Scores use methodology v1.6, O1S, with 1,000 bootstrap samples and seed 16.
Capability is the mean of Intelligence and Calibration within the frozen Jev-class cost and median-latency caps. Composite is secondary. Failed and refused answers remain in the full 1,500-decision denominator.
The public-versus-sealed gap reference is the median of this completed native cohort: 10.1 points. Historical scores are shown separately with their original dates and do not enter this cohort’s median or ranking. No API equating is applied.
Costs use each row’s documented price reference and measured usage on this draw. Subject-topic and use-case radars come from stored per-item results; only aggregates are published.
Sliders, presets and What-If change your view; they do not change the official result. Wrappers and subsidised systems are listed below the ranking.
Results SHA-256 1f97a539745a95cec0bb2feac754788cffa4ef35f689ad92b1e75441d8acfbfc · categories SHA-256 a96d1571b16419b1c6457115b8b34d0d79fbdfe33ba2d70707b0bf3e57126cb6 · scoring source SHA-256 64404997986ab427500d6fabc569089e2e0efea98e25c47046e54382e9f1c13f.
Read these scores against this release only
Each JevBench release since v1.6 scores on its own fresh seeded draw, and the gap penalty is measured against the field median of that draw. This release’s median is 10.073244581339713, so its penalty threshold sits at 18.073; the live v1.6.1 board’s median is 2.573387642438244. That makes these numbers not directly comparable with the headline board or with earlier draws, in a direction that flatters this field rather than penalising it.
The field median here is above 10, which has not happened in an earlier published JevBench release. No row in this release loses anything to the gap penalty; on the live board’s median, four of the eight would. The measured difference is at most two Intelligence points, and the measurement notes below give it per row.
Scoring keeps the frozen method: O1S, 1,000 bootstrap samples with seed 16, a fixed field median, all 1,500 decisions including recorded failures, and no API lane or equating. Every price on this board is a clearly labelled estimate from a base-model market reference or a documented per-1,000 figure, never a bill one of these systems issued.
Refusals, errors and context-bound failures stay in the full 1,500 denominator. Every row kept its own package, model, runtime and context settings; nothing was re-run for presentation.
NOT DIRECTLY COMPARABLE WITH THE LIVE v1.6.1 BOARD. This release is a different draw with its own field median: G_med = 10.073244581339713 over eight members, against 2.573387642438244 on the live field. G_med_flag_gt10 is TRUE, which has not been true in an earlier published JevBench artifact. Concretely: the gap penalty is max(0, 1 - max(0, (gap - G_med) - 8)/100), so this field's threshold is G_med + 8 = 18.073 and no row here is penalised for its public-vs-sealed gap, whereas at the live field's median four of the eight would be. The difference was measured, not estimated, and is small: rsi-jev-v6.1-vl-27b Intelligence 61.564 would be 59.551, standard_one_8b_sh 34.465 would be 32.460, liquid-d1-3b 22.139 would be 21.949, vegaml-4b 16.689 would be 16.673, and wald-4b-v2-1 (gap 9.478, below the live threshold 10.573) and the other three are unchanged. The direction is one-sided: this field flatters its own members.
Cost: no item is excluded from the cost basis on this draw. Every row is priced from its own measured usage block against the frozen 25 Sep 2026 snapshot, or from a documented per-1,000 override, so the paid draw's public-counter-overflow exclusions have no analogue here. The cost normalizer C0 is USD 0.001 per 1,000 decisions with 30 cost points per decade. Every price on these rows is a clearly-labelled ESTIMATE; none is a bill any of these systems issued.
gutsy-0.8b-v0.3: 29 of 1,500 items returned HTTP 500 "llama_decode returned 1" - the longest items, which llama.cpp cannot decode into the author's 8,192-token context. They are scored wrong, as the method requires. This is a crash class rather than a declared refusal. Its cost is input-only and marked *est.: the author's server emits no output-token field, and under the 9 Oct 2026 rule a missing output_tokens on a model with no generation path at all is treated as 0 output tokens.
standard_one_8b_sh and rsi-jev-v6.1-vl-27b: the same 29 longest items exceed each system's declared 32,768-token context and were refused with HTTP 422 rather than truncated; scored wrong per method. For RSI-Jev the refusal boundary was proven before the scored run: 13 of 13 boundary checks passed, including that a 32,768-token question is accepted and seen whole and a 32,769-token one returns 422 with nothing silently cut.
liquid-d1-omni-600m truncates states on the right at 16,384 tokens by design, and its published run is the second of two: a first run on a 24 GB card hit 29 CUDA out-of-memory errors, which was our under-provisioning and not model behaviour; it is retained and not scored. vegaml-0.8b and vegaml-4b truncate 29 states on the right at 73,728 tokens by design. The Vega checkpoint repository declares no licence field; the vegaml package is Apache-2.0.
wald-4b-v2-1: the shipped v2.1 server applies a documented Noul decisive floor (serving.json noul_decisive_floor 0.81): every calibrated P(yes) inside (0.19, 0.81) is moved to the nearer edge, on every Noul answer and not keyed to any item, so it answers all 375 Noul items decisively. Measured as shipped, server default effort none (one pass, 0 output tokens). It is the successor of the published wald-4b-v2 row on the paid draw and carries the same base-model cost basis.
Subject topics and TypeSafe use cases were labelled for this draw's 1,500 items in their own run (Winnow-12B Q8 on our own pod, egress cut before the sealed upload, the unchanged T1-T3/U1/U2 rules; no new rule was added). On 75 held-out public items checked by hand, topic agreement after the rules is 71/75 = 94.7 %. Known systematic miss: 17 of 27 router_policy items are filed under Coding by the subject of the request rather than under Rules, policy & law by the routing task; left as labelled because adding a rule after seeing the sample would change the method.
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
Open this section to load the earlier public-only board and diagnostics.