Official JevBench v1.6.8

JevBench · fresh regular draw

modernbert-jev v2 joins the ten systems of v1.6.7 on the same fresh seeded draw: 11 self-hosted open systems each answered the same 1,500 decisions of draw v1.6-regular-20261010-a2 (1,200 sealed plus the 300-item public set). The field median stays frozen at its v1.6.5 value, so the ten earlier rows keep their scores; only ranks and confidence intervals are recomputed for the larger field. Every drawn item had no prior scored use, and the set does not overlap v1.6.0 or the paid fast-lane draw. All 11 rows are complete and ranked. Historical scores keep their original dates and are listed separately below.

Aggregate results JSON · SHA-256 ed739acee1bbcee401e4e756ecb92efb41e654750317801977853f6a1b33c068 · Previous v1.6.7 (ten rows) · Full historical board

JevBench v1.6.8 · headline

JevBench Capability Score

Capability ranking of Jev-class systems

Capability Score averages Intelligence and Calibration. Blink v0.3 26B-A4B leads the Jev-class systems with 69.9.

Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘

Adjust cost / latency caps · 2× official
Official
#SystemICCap.$/1k
  1. 1Blink v0.3 26B-A4B49.249.669.9$0.048*
  2. 2Wald 4B v2.152.046.369.3$0.062*
  3. 3Liquid AI d1-3B22.150.649.5$0.044*
  4. 4Vega 4B16.747.237.7$0.058*
  5. 5Vega 0.8B3.161.533.0$0.019*
  6. 6Liquid AI d1-omni-600M3.373.132.8$0.0079*

Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.

Cost and latency are shown separately because they are nearly independent across systems (Spearman ρ = -0.05, n = 11).

  • Open weights · LLM decoder (9)
  • Open weights · encoder / classifier (2)
  • green: ≤ reference
  • amber: ≤ cap (2× reference)
  • red: > cap
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).

  1. –RSI-Jev v6.1-VL 27B61.631.173.5$0.20*Outside: cost 6.1× Jev (v1.5 reference)
  2. –Standard One 8B SH34.543.260.0$0.078*Outside: cost 2.4× Jev (v1.5 reference)
  3. –Clef-omni34.526.858.3$0.28*Outside: cost 8.5× Jev (v1.5 reference)
  4. –Gutsy 0.8B v0.38.678.443.1$0.0053*Outside: latency 2.7× Jev (v1.5 reference)
  5. –modernbert-jev v22.670.434.6$0.0097*Outside: latency 8.7× Jev (v1.5 reference)

Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 6 of 11 systems qualify; the other 5, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.

Capability against cost and speed

Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.

Capability vs cost

Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

Capability vs cost: 6 systems. Upper right is best: more capable and cheaper.30405060708090100$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev (v1.5 reference)← priciercheaper →1. Blink v0.3 26B-A4B2. Wald 4B v2.13. Liquid AI d1-3B4. Vega 4B5. Vega 0.8B
6 systems. Tap a bubble for its values.

Capability vs speed

Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

Capability vs speed: 6 systems. Upper right is best: more capable and faster.3040506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev (v1.5 reference)← slowerfaster →1. Blink v0.3 26B-A4B2. Wald 4B v2.13. Liquid AI d1-3B4. Vega 4B5. Vega 0.8B
6 systems. Tap a bubble for its values.
  • Open weights · LLM decoder (9)
  • Open weights · encoder / classifier (2)
  • faint = outside Jev-class
11 of 11 systems

Model kind

Jev-class

Release

No family reported in this release.

Parameters

No system in this release reports an exact parameter count.

Developer/API price $/1k

No system in this release reports a developer/API price.

Base-model price $/1k

No system in this release reports a base-model reference price.

Scored cost $/1k
Alternative price $/1k

No system in this release reports an alternative pricing scenario.

p50 latency (s)
p95 latency (s)

JevBench v1.6.8

JevBench Composite Score: 11 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

Weights:
Adjust weights ↓

View by:Capability ↑

The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis.

Greener = stronger within its column.

11 of 11 systems, sorted by official rank, #1 first.

  1. 1Blink v0.3 26B-A4Bnewfine-tune · 26B-A4B · NVFP4Base model: undisclosed61.2I 49C 91S 91K 50est.$0.048
  2. 2Wald 4B v2.1newfine-tune · 4BBase model: undisclosed54.3I 52C 87S 93K 46est.$0.062
  3. 3RSI-Jev v6.1-VL 27Bnewmerge · 27BBase model: undisclosed21.6I 62C 85S 85K 31est.$0.197
  4. 4Standard One 8B SHnewfine-tune · 8BBase model: undisclosed18.7I 34C 85S 84K 43est.$0.078
  5. 5Liquid AI d1-3Bnewfine-tune · 3BBase model: undisclosed8.9I 22C 77S 94K 51est.$0.044
  6. 6Clef-omninewfine-tune · 30B-A3BBase model: undisclosed6.1I 34C 82S 87K 27est.$0.276
  7. 7Vega 4Bnewadapter · 4BBase model: undisclosed3.6I 17C 59S 84K 47est.$0.058
  8. 8Gutsy 0.8B v0.3newfine-tune · 0.8B · GGUF Q8_0Base model: undisclosed0.8I 9C 78S 73K 78est.$0.0053
  9. 9Liquid AI d1-omni-600Mnewfine-tune · 600MBase model: undisclosed0.0I 3C 62S 94K 73est.$0.0079
  10. 10Vega 0.8Bnewadapter · 0.8BBase model: undisclosed0.0I 3C 63S 89K 61est.$0.019
  11. 11modernbert-jev v2newfine-tune · 150M · ONNX INT8Base model: undisclosed0.0I 3C 67S 58K 70est.$0.0097
Weights:
Adjust weights ↓

Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Open weights · LLM decoder (9)
  • Open weights · encoder / classifier (2)
I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

Compare two systems

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

The four score axes

Radar: the four score axes, two systemsThe four score axes, Blink v0.3 26B-A4B vs Wald 4B v2.1. Intelligence: 49.2 vs 52.0; Calibration: 90.7 vs 86.6; Speed: 91.5 vs 93.1; Cost: 49.6 vs 46.3.50100Intelligence: Blink v0.3 26B-A4B: 49.2; Wald 4B v2.1: 52.0Intelligence49.2 · 52.0Calibration: Blink v0.3 26B-A4B: 90.7; Wald 4B v2.1: 86.6Calibration90.7 · 86.6Speed: Blink v0.3 26B-A4B: 91.5; Wald 4B v2.1: 93.1Speed91.5 · 93.1Cost: Blink v0.3 26B-A4B: 49.6; Wald 4B v2.1: 46.3Cost49.6 · 46.3
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Blink v0.3 26B-A4B vs Wald 4B v2.1. Rules, policy & law: 58.8 vs 54.8; Math & numbers: 16.6 vs 32.1; Coding & software: 70.1 vs 70.2; Support & operations: 36.6 vs 38.0; Finance & commerce: 33.7 vs 45.3; Everyday language: 27.7 vs 21.7; Safety & security: 52.6 vs 53.2. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Rules, policy & law: Blink v0.3 26B-A4B: 58.8; Wald 4B v2.1: 54.8 — Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 711 items (151 open, 560 sealed).Rules, policy& law58.8 · 54.8Math & numbers: Blink v0.3 26B-A4B: 16.6; Wald 4B v2.1: 32.1 — Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 334 items (37 open, 297 sealed).Math &numbers16.6 · 32.1Coding & software: Blink v0.3 26B-A4B: 70.1; Wald 4B v2.1: 70.2 — Coding & software: code, SQL, repositories, developer tools and IT systems. 160 items (62 open, 98 sealed).Coding &software70.1 · 70.2Support & operations: Blink v0.3 26B-A4B: 36.6; Wald 4B v2.1: 38.0 — Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 125 items (21 open, 104 sealed).Support &operations36.6 · 38.0Finance & commerce: Blink v0.3 26B-A4B: 33.7; Wald 4B v2.1: 45.3 — Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 71 items (12 open, 59 sealed).Finance &commerce33.7 · 45.3Everyday language: Blink v0.3 26B-A4B: 27.7; Wald 4B v2.1: 21.7 — Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 54 items (11 open, 43 sealed).Everydaylanguage27.7 · 21.7Safety & security: Blink v0.3 26B-A4B: 52.6; Wald 4B v2.1: 53.2 — Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 45 items (6 open, 39 sealed).Safety &security52.6 · 53.2

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Measured subject-topic competence on this release’s item set. Chance-corrected competence per category (0 = chance, below chance negative, 100 = perfect), Blink v0.3 26B-A4B: open+sealed; Wald 4B v2.1: open+sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 711 items (151 open / 560 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 334 items (37 open / 297 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 160 items (62 open / 98 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 125 items (21 open / 104 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 71 items (12 open / 59 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 54 items (11 open / 43 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 45 items (6 open / 39 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Blink v0.3 26B-A4B vs Wald 4B v2.1. Legal & compliance: 65.5 vs 59.2; Model routing: 71.0 vs 66.1; Other: 14.6 vs 28.1; Customer support: 40.3 vs 47.5; Financial crime: 29.2 vs 34.2; E-commerce: 31.4 vs 60.8; Risk assessment: 42.1 vs 63.5; Science: 27.3 vs 6.8; Feature extraction: 40.7 vs 31.8. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Legal & compliance: Blink v0.3 26B-A4B: 65.5; Wald 4B v2.1: 59.2 — Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 319 items (74 open, 245 sealed).Legal &compliance65.5 · 59.2Model routing: Blink v0.3 26B-A4B: 71.0; Wald 4B v2.1: 66.1 — Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 273 items (86 open, 187 sealed).Modelrouting71.0 · 66.1Other: Blink v0.3 26B-A4B: 14.6; Wald 4B v2.1: 28.1 — Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 255 items (24 open, 231 sealed).Other14.6 · 28.1Customer support: Blink v0.3 26B-A4B: 40.3; Wald 4B v2.1: 47.5 — Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 155 items (37 open, 118 sealed).Customersupport40.3 · 47.5Financial crime: Blink v0.3 26B-A4B: 29.2; Wald 4B v2.1: 34.2 — Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 60 items (9 open, 51 sealed).Financialcrime29.2 · 34.2E-commerce: Blink v0.3 26B-A4B: 31.4; Wald 4B v2.1: 60.8 — E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 58 items (9 open, 49 sealed).E-commerce31.4 · 60.8Risk assessment: Blink v0.3 26B-A4B: 42.1; Wald 4B v2.1: 63.5 — Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 46 items (8 open, 38 sealed).Riskassessment42.1 · 63.5Science: Blink v0.3 26B-A4B: 27.3; Wald 4B v2.1: 6.8 — Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 32 items (2 open, 30 sealed).Science27.3 · 6.8Feature extraction: Blink v0.3 26B-A4B: 40.7; Wald 4B v2.1: 31.8 — Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 30 items (5 open, 25 sealed).Featureextraction40.7 · 31.8

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Measured application-category competence on this release’s item set. Chance-corrected competence per category (0 = chance, below chance negative, 100 = perfect), Blink v0.3 26B-A4B: open+sealed; Wald 4B v2.1: open+sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.

Low sample, n < 30 — indicative only

Category (items)A: Blink v0.3 26B-A4BB: Wald 4B v2.1
LLM guardrails (29)60.6 n=2970.7 n=29
Gaming (29)13.4 n=2929.1 n=29
Lead generation (28)38.6 n=2835.3 n=28
Demand forecasting (26)0.0 n=267.6 n=26
Advertising (24)43.8 n=2447.4 n=24
Moderation and trust and safety (24)84.6 n=2445.2 n=24
Semantic code linting (24)30.5 n=2441.5 n=24
Graphs and knowledge graphs (23)30.3 n=2334.6 n=23
Insurance claims (22)30.7 n=2248.3 n=22
Search and retrieval (22)84.2 n=2276.7 n=22
Recruiting (21)38.8 n=2138.8 n=21

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 319 items (74 open / 245 sealed)
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 273 items (86 open / 187 sealed)
  • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 255 items (24 open / 231 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 155 items (37 open / 118 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 60 items (9 open / 51 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 58 items (9 open / 49 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 46 items (8 open / 38 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 32 items (2 open / 30 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 30 items (5 open / 25 sealed)

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Blink v0.3 26B-A4B vs Wald 4B v2.1. Choice · open: 74.0 vs 71.0; Choice · sealed: 68.4 vs 55.1; Noul · open: 28.2 vs 66.2; Noul · sealed: 20.8 vs 53.1; Score · open: 57.7 vs 33.0; Score · sealed: 46.1 vs 33.4. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Choice · open: Blink v0.3 26B-A4B: 74.0; Wald 4B v2.1: 71.0Choice · open74.0 · 71.0Choice · sealed: Blink v0.3 26B-A4B: 68.4; Wald 4B v2.1: 55.1Choice ·sealed68.4 · 55.1Noul · open: Blink v0.3 26B-A4B: 28.2; Wald 4B v2.1: 66.2Noul · open28.2 · 66.2Noul · sealed: Blink v0.3 26B-A4B: 20.8; Wald 4B v2.1: 53.1Noul · sealed20.8 · 53.1Score · open: Blink v0.3 26B-A4B: 57.7; Wald 4B v2.1: 33.0Score · open57.7 · 33.0Score · sealed: Blink v0.3 26B-A4B: 46.1; Wald 4B v2.1: 33.4Score ·sealed46.1 · 33.4

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Chance-corrected competence for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Blink v0.3 26B-A4B vs Wald 4B v2.1. Easy: 87.5 vs 81.6; Standard: 56.4 vs 60.8; Judge: 49.0 vs 59.1; Hard: 58.6 vs 54.8. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Blink v0.3 26B-A4B: 87.5; Wald 4B v2.1: 81.6Easy87.5 · 81.6Standard: Blink v0.3 26B-A4B: 56.4; Wald 4B v2.1: 60.8Standard56.4 · 60.8Judge: Blink v0.3 26B-A4B: 49.0; Wald 4B v2.1: 59.1Judge49.0 · 59.1Hard: Blink v0.3 26B-A4B: 58.6; Wald 4B v2.1: 54.8Hard58.6 · 54.8

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Blink v0.3 26B-A4B vs Wald 4B v2.1. Easy: 72.5 vs 53.4; Standard: 58.8 vs 51.1; Judge: 56.4 vs 52.0; Hard: 37.8 vs 44.9. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Blink v0.3 26B-A4B: 72.5; Wald 4B v2.1: 53.4Easy72.5 · 53.4Standard: Blink v0.3 26B-A4B: 58.8; Wald 4B v2.1: 51.1Standard58.8 · 51.1Judge: Blink v0.3 26B-A4B: 56.4; Wald 4B v2.1: 52.0Judge56.4 · 52.0Hard: Blink v0.3 26B-A4B: 37.8; Wald 4B v2.1: 44.9Hard37.8 · 44.9

Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).

Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 items of this draw was labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use case by Winnow-12B Q8 on our own GPU pod (same model, questions, taxonomy and T1-T3/U1/U2 item-group rules as v1.5 and v1.6.4); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items were checked by hand afterwards and no rule was derived from them: topic agreement 71/75 = 94.7 %. 17 of 27 router_policy items are labelled Coding rather than Rules, policy & law. All non-English uc1 items are machine-authored and not native-reviewed.

  • Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
  • Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
SpokeA: Blink v0.3 26B-A4BB: Wald 4B v2.1
The four score axes
Intelligence49.252.0
Calibration90.786.6
Speed91.593.1
Cost49.646.3
Capability by subject topic
Rules, policy & law58.854.8
Math & numbers16.632.1
Coding & software70.170.2
Support & operations36.638.0
Finance & commerce33.745.3
Everyday language27.721.7
Safety & security52.653.2
Use cases (TypeSafe categories)
Legal & compliance65.559.2
Model routing71.066.1
Other14.628.1
Customer support40.347.5
Financial crime29.234.2
E-commerce31.460.8
Risk assessment42.163.5
Science27.36.8
Feature extraction40.731.8
Competence per request type, open / sealed
Choice · open74.071.0
Choice · sealed68.455.1
Noul · open28.266.2
Noul · sealed20.853.1
Score · open57.733.0
Score · sealed46.133.4
Competence per tier — open set
Easy87.581.6
Standard56.460.8
Judge49.059.1
Hard58.654.8
Competence per tier — sealed set
Easy72.553.4
Standard58.851.1
Judge56.452.0
Hard37.844.9

All eleven rows are self-hosted on evaluator-owned pods; there is no API lane in this release, so no equating was applied to any published cell. Sealed counts refer to self-hosted S (1,200); API rows use their original measured sealed basis.

Languages

Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1200 + P 300; API rows use their measured sealed subset plus P. Sage (A3) language cells and A2/A3 topic/use-case cells cover public P300 only. Raw and outside the Composite. Cells under 15 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.8. English (1,119 items) is listed first; the other 21 languages and the mixed-language group share 381 items. In 12 further groups no system reaches the 15-item reporting minimum, so they get no column (items in the pool shown): Dutch (13), Korean (11), Turkish (11), Ukrainian (11), Arabic (11), Mixed-language (11), Finnish (9), Chinese (8), Czech (8), Swedish (8), Norwegian (7), Indonesian (6) — 114 items, scored like every other item.

11 systems

JevBench v1.6.8 competence by item language and system
111945403230282418181715
Blink v0.3 26B-A4Bself-hosted · S+P467045253954†54†90†0†77†39†
Wald 4B v2.1self-hosted · S+P505534492572†43†75†12†54†54†
RSI-Jev v6.1-VL 27Bself-hosted · S+P577251325383†44†83†0†90†42†
Standard One 8B SHself-hosted · S+P3045419444†23†77†0†43†0†
Liquid AI d1-3Bself-hosted · S+P146300134†0†55†0†41†29†
Clef-omniself-hosted · S+P355236184554†24†74†0†51†17†
Vega 4Bself-hosted · S+P114436024†22†15†0†42†10†
Gutsy 0.8B v0.3self-hosted · S+P02120014†0†27†0†28†0†
Liquid AI d1-omni-600Mself-hosted · S+P000000†0†0†0†0†0†
Vega 0.8Bself-hosted · S+P000000†0†22†0†5†0†
modernbert-jev v2self-hosted · S+P000000†0†0†0†0†0†

Intelligence gate and Noul decisiveness

The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.

SystemIntelligenceChoiceNoulScoreNoul decisive rateAccuracy among decisive
Blink v0.3 26B-A4B49.2below gatenativenativenative67 %95 %
Wald 4B v2.152.0nativenativenative100 %78 %
RSI-Jev v6.1-VL 27B61.6nativenativenative83 %91 %
Standard One 8B SH34.5below gatenativenativenative61 %88 %
Liquid AI d1-3B22.1below gatenativenativenative49 %83 %
Clef-omni34.5below gatenativenativenative65 %86 %
Vega 4B16.7below gatenativenativenative70 %68 %
Gutsy 0.8B v0.38.6below gatenativenativenative29 %74 %
Liquid AI d1-omni-600M3.3below gatenativenativenative27 %56 %
Vega 0.8B3.1below gatenativenativenative32 %63 %
modernbert-jev v22.6below gatenativenativenative0 %—
All data (11 systems)

Method notes · fresh native cohort · v1.6.8

Each system answered the same 1,200 freshly drawn sealed decisions and 300 public decisions on an offline GPU pod. Scores use methodology v1.6, O1S, with 1,000 bootstrap samples and seed 16.

Capability is the mean of Intelligence and Calibration within the frozen Jev-class cost and median-latency caps. Composite is secondary. Failed and refused answers remain in the full 1,500-decision denominator.

The public-versus-sealed gap reference is the median of this completed native cohort: 10.1 points. Historical scores are shown separately with their original dates and do not enter this cohort’s median or ranking. No API equating is applied.

Costs use each row’s documented price reference and measured usage on this draw. Subject-topic and use-case radars come from stored per-item results; only aggregates are published.

Sliders, presets and What-If change your view; they do not change the official result. Wrappers and subsidised systems are listed below the ranking.

Results SHA-256 ed739acee1bbcee401e4e756ecb92efb41e654750317801977853f6a1b33c068 · categories SHA-256 0dd07bc493f63793fb17c809c1c1ed1c914aeca61c5881acf14cadbae6f11c26 · scoring source SHA-256 b61ad7cd7ea5a1d5d0a45dbaf98669ab480bc86430d64c89181aac7769f99d24.

Read these scores against this release only

Each JevBench release since v1.6 scores on its own fresh seeded draw, and the gap penalty is measured against the field median of that draw. This release’s median is 10.073244581339713, so its penalty threshold sits at 18.073; the live v1.6.1 board’s median is 2.573387642438244. That makes these numbers not directly comparable with the headline board or with earlier draws, in a direction that flatters this field rather than penalising it.

The field median here is above 10, which has not happened in an earlier published JevBench release. No row in this release loses anything to the gap penalty; on the live board’s median, four of the eleven would. The measured difference is at most 2.013 Intelligence points, and the measurement notes below give it per row.

Scoring keeps the frozen method: O1S, 1,000 bootstrap samples with seed 16, a fixed field median, all 1,500 decisions including recorded failures, and no API lane or equating. Every price on this board is a clearly labelled estimate from a base-model market reference or a documented per-1,000 figure, never a bill one of these systems issued.

Refusals, errors and context-bound failures stay in the full 1,500 denominator. Every row kept its own package, model, runtime and context settings; nothing was re-run for presentation.

NOT DIRECTLY COMPARABLE WITH THE LIVE v1.6.1 BOARD. This release is a different draw with its own field median: G_med = 10.073244581339713, frozen from the eight v1.6.5 members (Clef-omni, Blink and modernbert-jev v2 do not move it), against 2.573387642438244 on the live field. G_med_flag_gt10 is TRUE, which has not been true in an earlier published JevBench artifact. Concretely: the gap penalty is max(0, 1 - max(0, (gap - G_med) - 8)/100), so this field's threshold is G_med + 8 = 18.073 and no row here is penalised for its public-vs-sealed gap, whereas at the live field's median four of the eleven would be. The difference was measured, not estimated, and is small: rsi-jev-v6.1-vl-27b Intelligence 61.564 would be 59.551, standard_one_8b_sh 34.465 would be 32.460, liquid-d1-3b 22.139 would be 21.949, vegaml-4b 16.689 would be 16.673, and wald-4b-v2-1 (gap 9.478, below the live threshold 10.573), blink-v0-3-26b-a4b-nvfp4 (gap 8.161), modernbert-jev-v2 (gap 2.433) and the other four (clef-omni, gutsy-0.8b-v0.3, liquid-d1-omni-600m, vegaml-0.8b) are unchanged. The direction is one-sided: this field flatters its own members.

Cost: no item is excluded from the cost basis on this draw. Every row is priced from its own measured usage block against the frozen 25 Sep 2026 snapshot, or from a documented per-1,000 override, so the paid draw's public-counter-overflow exclusions have no analogue here. The cost normalizer C0 is USD 0.001 per 1,000 decisions with 30 cost points per decade. Every price on these rows is a clearly-labelled ESTIMATE; none is a bill any of these systems issued.

gutsy-0.8b-v0.3: 29 of 1,500 items returned HTTP 500 "llama_decode returned 1" - the longest items, which llama.cpp cannot decode into the author's 8,192-token context. They are scored wrong, as the method requires. This is a crash class rather than a declared refusal. Its cost is input-only and marked *est.: the author's server emits no output-token field, and under the 9 Oct 2026 rule a missing output_tokens on a model with no generation path at all is treated as 0 output tokens.

standard_one_8b_sh and rsi-jev-v6.1-vl-27b: the same 29 longest items exceed each system's declared 32,768-token context and were refused with HTTP 422 rather than truncated; scored wrong per method. For RSI-Jev the refusal boundary was proven before the scored run: 13 of 13 boundary checks passed, including that a 32,768-token question is accepted and seen whole and a 32,769-token one returns 422 with nothing silently cut.

liquid-d1-omni-600m truncates states on the right at 16,384 tokens by design, and its published run is the second of two: a first run on a 24 GB card hit 29 CUDA out-of-memory errors, which was our under-provisioning and not model behaviour; it is retained and not scored. vegaml-0.8b and vegaml-4b truncate 29 states on the right at 73,728 tokens by design. The Vega checkpoint repository declares no licence field; the vegaml package is Apache-2.0.

v1.6.8 ADDS ONE ROW, modernbert-jev v2, to the ten of v1.6.7 on the same draw, under the same rule: the field median stays the frozen v1.6.5 value, so all ten earlier rows keep identical point scores, axes and capability; ranks and bootstrap confidence intervals are recomputed for the eleven-row roster, and raw_on_A of standard_one_8b_sh and the field-level diagnostics (G_med_api_basis_P_vs_A, the API-basis offsets) move again for the reason given below; none of them enters a score.

modernbert-jev-v2: a 150M ModernBERT-base NLI encoder, int8 (unsigned) ONNX, served by the author's own CLI at its own defaults on ONE CPU thread of the evaluator's shared Sandy host (AMD Ryzen 5 3600, Zen 2), not a dedicated machine; the 1-minute load average was logged every 60 s and ranged about 5.5 to 27.9 (median 10.2) during the run, so Speed is a lower bound for this model on an idle core. The run was interrupted once by the evaluator's session ending after 41 answered items and resumed with the unchanged runner's --resume, which keeps answered rows and sends only the rest; the server is deterministic (byte-identical answers on repeated requests). By the author's documented design (README: "windowing retrieval is future work"), a state is cut at 20,000 characters and read through at most six 300-word windows spread across it: the 29 longest states of this draw are answered from that sample rather than refused, so words between the windows are never seen. Probabilities are native (per-head temperatures from the shipped calibration.json). With the shipped Noul temperature of 2.5, every one of its 375 Noul answers lies between P(yes) 0.23 and 0.80, so none counts as decisive under the O1S Noul method, which puts every Noul tier of this row at its floor; this is the model as shipped, not a serving fault. Output tokens are a literal 0; cost is a labelled estimate at the base-size encoder rate.

v1.6.7 ADDED ONE ROW, Blink v0.3 26B-A4B, to the nine of v1.6.6 on the same draw, under the same rule: the field median stays the frozen v1.6.5 value, so all nine earlier rows keep identical point scores, axes and capability; ranks and bootstrap confidence intervals are recomputed for the ten-row roster, and raw_on_A of standard_one_8b_sh moves again for the reason given below. Blink's gap (8.161) is below the threshold.

blink-v0-3-26b-a4b-nvfp4: measured from the model card's own plain-vLLM 0.30.0 serve command (NVFP4, fp8 KV cache, 32,768-token context, T=1.3 option-letter readout through a loopback shim reproducing the card's decide() client) with ONE serving-only deviation: VLLM_BATCH_INVARIANT=1. With the command as published, four identical repeats of 50 public items differed by up to 0.126 in probability, failing the pre-registered repeatability gate (0.005); with batch-invariant kernels the repeats were byte-identical. This makes the run repeatable, not more correct: between serving settings the same item's probabilities differ by up to 0.39 and the argmax by 2 of 49 public items, so a different stack or card can score differently. Latency is of the batch-invariant path (about twice the default per request). Run on an RTX 5090 32 GB (no RTX PRO 6000 was free). 29 of 1,500 states (77,706 to 82,401 tokens) exceed the card's 32,768-token context and were refused with HTTP 422; scored wrong per method. Cost is a labelled estimate from the base model google/gemma-4-26b-a4b-it (USD 0.09 in / 0.30 out per 1M, frozen 25 Sep 2026 OpenRouter snapshot).

v1.6.6 ADDED ONE ROW, Clef-omni, to the eight of v1.6.5 on the same draw. The field median is held at the frozen v1.6.5 value (g_med_fixed, as for earlier same-draw additions), so all eight earlier rows keep identical point scores, axes and capability. Ranks and bootstrap confidence intervals are recomputed for the nine-row roster, as the v1.6.3 and v1.6.4 addenda did. One diagnostic also moves: raw_on_A (scores on the 300-item A subset) of standard_one_8b_sh, because the A-basis median is not part of the frozen value. Clef-omni's gap (5.555) is below the threshold.

clef-omni: measured as shipped on its own in-process entry point at the pinned revision, defaults unchanged. 29 of 1,500 items reach the model's own 64,000-token default and are answered on the truncated state, as the authors' encoder does by design; none was refused. The Omni video processor needs torchvision to load, which the model card's tested stack does not list; torchvision 0.26.0+cu130, the build matching its torch 2.11.0+cu130, was added and nothing else changed. Output tokens are a literal 0, so the cost is input-only.

wald-4b-v2-1: the shipped v2.1 server applies a documented Noul decisive floor (serving.json noul_decisive_floor 0.81): every calibrated P(yes) inside (0.19, 0.81) is moved to the nearer edge, on every Noul answer and not keyed to any item, so it answers all 375 Noul items decisively. Measured as shipped, server default effort none (one pass, 0 output tokens). It is the successor of the published wald-4b-v2 row on the paid draw and carries the same base-model cost basis.

Subject topics and TypeSafe use cases were labelled for this draw's 1,500 items in their own run (Winnow-12B Q8 on our own pod, egress cut before the sealed upload, the unchanged T1-T3/U1/U2 rules; no new rule was added). On 75 held-out public items checked by hand, topic agreement after the rules is 71/75 = 94.7 %. Known systematic miss: 17 of 27 router_policy items are filed under Coding by the subject of the request rather than under Rules, policy & law by the routing task; left as labelled because adding a rule after seeing the sample would change the method.

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

Open this section to load the earlier public-only board and diagnostics.