JevBench · completed fresh native cohort
Same-draw addendum to the published five-model v1.6.3: 6 completed preregistered paid submissions answered the same fresh 1,500-decision draw with methodology v1.6. Four native systems are ranked; 2 measured wrappers are listed below the rankings. Historical scores retain their original dates and are listed separately below.
Aggregate results JSON · SHA-256 f3d57e0705da2c565bc0925a03a0fe6aa351d3cff37e71816a92e6d27cc3e7d7 · Previous v1.6.3 five-system release · Full historical board
JevBench v1.6.4 · headline
JevBench Capability Score
Capability ranking of Jev-class systems
Capability Score averages Intelligence and Calibration. Jeff-1.0-Large leads the Jev-class systems with 81.5.
Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘
Adjust cost / latency caps · 2× official
ScoreCCost
ScoreCap.Capability
Score$/1k$/1k decisions
- 1Jeff-1.0-Large71.949.281.5$0.049*
- 2Metask rain 4B51.151.866.0$0.040*
- 3Wald 4B41.154.359.8$0.033*
- 4decisor-4b39.651.445.7$0.042*
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency correlate here (Spearman ρ = 0.80, n = 4).
- Open weights · LLM decoder (4)
- green: ≤ reference
- amber: ≤ cap (2× reference)
- red: > cap
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).
Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 4 of 4 systems qualify; the other 0, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed
Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
- Open weights · LLM decoder (4)
- faint = outside Jev-class
Model kind
Jev-class
Release
No provider reported in this release.
No family reported in this release.
No licence reported in this release.
No system in this release reports an exact parameter count.
No system in this release reports a developer/API price.
No system in this release reports a base-model reference price.
No system in this release reports an alternative pricing scenario.
JevBench v1.6.4
JevBench Composite Score: 4 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Adjust weights ↓
The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis, and the column headings sort too.
Greener = stronger within its column.
4 of 4 systems, sorted by official rank, #1 first.
- 1Jeff-1.0-Largenewfine-tune · 31BBase model: undisclosed68.6I 72C 91S 89K 49est.$0.049
- 2Metask rain 4Bnewfine-tune · 4BBase model: undisclosed62.7I 51C 81S 80K 52est.$0.040
- 3Wald 4Bnewfine-tune · 4BBase model: undisclosed40.0I 41C 78S 82K 54est.$0.033
- 4decisor-4bnewfine-tune · 4B · FP8Base model: undisclosed33.5I 40C 52S 92K 51est.$0.042
Adjust weights ↓
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Open weights · LLM decoder (4)
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
- A: Jeff-1.0-Large — Open weights · LLM decoder · Score 68.6 (#1)
- B: Metask rain 4B — Open weights · LLM decoder · Score 62.7 (#2)
The four score axes
Capability by subject topic
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
What each category means · items per category
- Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 747 items (153 open / 594 sealed)
- Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 282 items (37 open / 245 sealed)
- Coding & software — code, SQL, repositories, developer tools and IT systems. 197 items (61 open / 136 sealed)
- Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 100 items (22 open / 78 sealed)
- Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 79 items (11 open / 68 sealed)
- Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 53 items (10 open / 43 sealed)
- Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 42 items (6 open / 36 sealed)
Use cases (TypeSafe categories)
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Low sample, n < 30 — indicative only
| Category (items) | A: Jeff-1.0-Large | B: Metask rain 4B |
|---|---|---|
| Insurance claims (29) | 65.5 n=29 | 50.9 n=29 |
| LLM guardrails (25) | 74.8 n=25 | 47.7 n=25 |
| Search and retrieval (24) | 82.2 n=24 | 70.4 n=24 |
| Semantic code linting (22) | 84.2 n=22 | 28.3 n=22 |
| Graphs and knowledge graphs (22) | 64.3 n=22 | 13.4 n=22 |
| Demand forecasting (21) | 19.3 n=21 | 21.9 n=21 |
| Scientific discovery (21) | 59.9 n=21 | 17.5 n=21 |
| Gaming (18) | 45.3 n=18 | 25.9 n=18 |
| Recruiting (17) | 66.1 n=17 | 67.4 n=17 |
| Feature extraction for predictive modeling (16) | 71.1 n=16 | 5.9 n=16 |
Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.
What each category means · items per category
- Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 343 items (77 open / 266 sealed)
- Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 318 items (86 open / 232 sealed)
- Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 213 items (23 open / 190 sealed)
- Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 158 items (37 open / 121 sealed)
- Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 54 items (9 open / 45 sealed)
- Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 53 items (7 open / 46 sealed)
- E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 43 items (9 open / 34 sealed)
- Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 37 items (7 open / 30 sealed)
- Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 36 items (4 open / 32 sealed)
- Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open / 27 sealed)
Competence per request type, open / sealed
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — open set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
Competence per tier — sealed set
Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.
How the categories were made
O1S chance-corrected competence: equal mean over request types present, clipped to 0–100, using the pinned official group_stats function. Recorded failures remain in the denominator. Category scores are descriptive and do not alter headline Intelligence, Calibration, composite score or G_med.
Subject topics and TypeSafe use cases use the actual shared local reference labels, with frozen T1/T2/T3/U1/U2 authoring rules. The original probabilities and model choice remain preserved in protected custody. The genuine 75-public-item handcheck is required. Family, type and language are frozen authoring metadata; non-English machine-authored items are not native-reviewed.
- Every nonempty category has a measured cell; zero-count taxonomy categories have no cell.
- Categories below 30 items are indicative only; radar display retains its 30-item minimum.
- No historic overlays, estimated cells, outcome exclusions or private item-level data are included.
All values as a table
| Spoke | A: Jeff-1.0-Large | B: Metask rain 4B |
|---|---|---|
| The four score axes | ||
| Intelligence | 71.9 | 51.1 |
| Calibration | 91.1 | 80.9 |
| Speed | 88.8 | 79.6 |
| Cost | 49.2 | 51.8 |
| Capability by subject topic | ||
| Rules, policy & law | 79.7 | 51.2 |
| Math & numbers | 48.3 | 29.4 |
| Coding & software | 83.9 | 61.0 |
| Support & operations | 56.7 | 50.9 |
| Finance & commerce | 74.7 | 55.1 |
| Everyday language | 52.4 | 48.0 |
| Safety & security | 90.2 | 53.7 |
| Use cases (TypeSafe categories) | ||
| Legal & compliance | 84.5 | 51.5 |
| Model routing | 84.4 | 68.1 |
| Other | 36.7 | 23.6 |
| Customer support | 66.5 | 49.6 |
| Financial crime | 62.7 | 53.4 |
| Risk assessment | 76.2 | 40.7 |
| E-commerce marketplaces | 67.8 | 40.4 |
| Moderation and trust and safety | 87.0 | 39.9 |
| Lead generation | 78.7 | 60.6 |
| Advertising | 49.6 | 28.1 |
| Competence per request type, open / sealed | ||
| Choice · open | 76.9 | 73.6 |
| Choice · sealed | 81.3 | 58.7 |
| Noul · open | 85.3 | 54.9 |
| Noul · sealed | 77.2 | 56.7 |
| Score · open | 59.6 | 39.8 |
| Score · sealed | 51.0 | 22.8 |
| Competence per tier — open set | ||
| Easy | 95.4 | 83.7 |
| Standard | 68.3 | 57.5 |
| Judge | 67.4 | 60.6 |
| Hard | 76.1 | 54.7 |
| Competence per tier — sealed set | ||
| Easy | 86.4 | 58.7 |
| Standard | 81.0 | 53.5 |
| Judge | 68.1 | 46.5 |
| Hard | 68.7 | 46.4 |
All completed measurements cover the same fresh 1,500-item draw (1,200 sealed and 300 public); pending candidates are outside this completed category artifact. Sealed counts refer to self-hosted S (1,200); API rows use their original measured sealed basis.
Languages
Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1200 + P 300; API rows use their measured sealed subset plus P. Sage (A3) language cells and A2/A3 topic/use-case cells cover public P300 only. Raw and outside the Composite. Cells under 1 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.4. The mixed-language group is listed first, then English (1,135 items); the other 21 languages and the mixed-language group share 365 items.
| System | mixed†15 | en1135 | de40 | pl39 | es30 | hi†26 | pt†24 | ja†22 | ar†17 | el†17 | nl†17 | da†16 | fr†16 | it†13 | zh†12 | tr†11 | cs†9 | ko†9 | uk†9 | sv†8 | fi†7 | id†4 | no†4 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Jeff-1.0-Largeself-hosted · S+P | 99† | 71 | 91 | 64 | 89 | 84† | 67† | 49† | 62† | 59† | 77† | 56† | 98† | 79† | 77† | 77† | 46† | 91† | 70† | 34† | 94† | 100† | 68† |
| Metask rain 4Bself-hosted · S+P | 100† | 48 | 62 | 60 | 65 | 41† | 74† | 8† | 0† | 73† | 0† | 61† | 30† | 93† | 52† | 6† | 36† | 65† | 81† | 30† | 0† | 34† | 68† |
| Wald 4Bself-hosted · S+P | 95† | 44 | 54 | 35 | 51 | 26† | 52† | 0† | 14† | 15† | 8† | 51† | 47† | 76† | 0† | 77† | 69† | 74† | 13† | 8† | 2† | 34† | 0† |
| decisor-4bself-hosted · S+P | 100† | 36 | 58 | 44 | 70 | 12† | 53† | 0† | 11† | 13† | 50† | 24† | 45† | 95† | 20† | 42† | 56† | 1† | 0† | 12† | 34† | 34† | 68† |
| Wrappers and subsidized systems (listed, never ranked) | |||||||||||||||||||||||
| metask-jev-rain-12Bself-hosted · S+P | 99† | 64 | 83 | 55 | 79 | 65† | 78† | 42† | 47† | 71† | 67† | 67† | 71† | 85† | 14† | 41† | 79† | 91† | 65† | 31† | 70† | 100† | 18† |
| ryotide_qwen9self-hosted · S+P | 92† | 20 | 60 | 29 | 70 | 1† | 45† | 0† | 15† | 28† | 0† | 15† | 13† | 56† | 0† | 9† | 56† | 0† | 51† | 0† | 0† | 34† | 0† |
Intelligence gate and Noul decisiveness
The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.
| System | Intelligence | Choice | Noul | Score | Noul decisive rate | Accuracy among decisive |
|---|---|---|---|---|---|---|
| Jeff-1.0-Large | 71.9 | native | native | native | not supported | — |
| Metask rain 4B | 51.1 | native | native | native | not supported | — |
| Wald 4B | 41.1below gate | native | native | native | not supported | — |
| decisor-4b | 39.6below gate | native | native | native | not supported | — |
All data (6 systems)
Method notes · fresh native cohort · v1.6.4
Each system answered the same 1,200 freshly drawn sealed decisions and 300 public decisions on an offline GPU pod. Scores use methodology v1.6, O1S, with 1,000 bootstrap samples and seed 16.
Capability is the mean of Intelligence and Calibration within the frozen Jev-class cost and median-latency caps. Composite is secondary. Failed and refused answers remain in the full 1,500-decision denominator.
The public-versus-sealed gap reference is the median of this completed native cohort: 7.1 points. Historical scores are shown separately with their original dates and do not enter this cohort’s median or ranking. No API equating is applied.
This same-draw six-model addendum adds the measured RYOTIDE wrapper to v1.6.3. The five prior measurements, point scores and four-native field median are retained. The standard scorer recalculates confidence intervals for the expanded roster. Wrappers remain listed below ranking and excluded from the median.
Costs use each row’s documented price reference and measured usage on this draw. Subject-topic and use-case radars come from stored per-item results; only aggregates are published.
Sliders, presets and What-If change your view; they do not change the official result. Wrappers and subsidised systems are listed below the ranking.
Results SHA-256 f3d57e0705da2c565bc0925a03a0fe6aa351d3cff37e71816a92e6d27cc3e7d7 · categories SHA-256 884a83dd21aef50895e1eec391d07f490b15c0f5c0a8cba7abb05955b6cfbac0 · scoring source SHA-256 6326428cdee639ae758d71ea2674111c7889943f0ee8af1e26315feb84cf7ffb.
Addendum method · whole-field rescoring
v1.6.4 computes the six-system completed field together, adding the measured RYOTIDE wrapper to the five systems published in v1.6.3. Their original measurements (raw answers, completion status, cost basis, latency and category cells) are retained unchanged; no inference was rerun.
The four-native field median G_med remains unchanged from v1.6.3. RYOTIDE is a wrapper and does not enter that median or the ranked cohort. This is a new immutable publication; the published v1.6.3 artifacts remain available.
The field median is the median Intelligence gap of exactly the four completed native systems (7.088097744249485). Wrappers and historical rows never enter it. Scoring keeps O1S, 1,000 bootstrap samples with seed 16 and a fixed median, all 1,500 decisions including recorded failures, and the published 1,479-item common cost basis. No API cohort or equating is applied.
Refusals, errors and context-bound failures remain in the full 1,500 denominator. Native package/model/runtime/context settings remain those of each independently reviewed measurement; no inference is rerun for this metadata enrichment.
Reference token costs use a common 1,479-item set: 21 items are excluded from cost using an input-only rule based on state plus the longest question and the current public reference limit of 32,736 tokens. The empirical historical accepted/refused gap does not identify the exact historical cap, and token-length reconstruction extrapolates beyond the historical approximately 2,300-token region. This is a dated convention with uncertainty, not an exact or stationary historical tariff/input claim. The 21 exclusions do not reduce the full 1,500 capability denominator. The official cost normalizer C0 is USD 0.001 per 1,000 decisions, with 30 cost points per decade.
All six fixed roster members are now complete. The four eligible native nonwrapper field-median members remain Jeff-1.0-Large, Wald 4B, Metask rain 4B and Decisor 4B. Metask rain 12B and RYOTIDE are measured wrappers listed below the ranking and excluded from G_med. Bootstrap B=1000, seed=16 holds G_med fixed; published confidence intervals therefore omit field-median uncertainty. This small-field release is not directly comparable with larger-field or earlier-draw scores. In the canonical five-to-six-system restatement, Decisor occupies canonical bootstrap job-list index 5 rather than 4; its per-system CI seed consequently changes from 20 to 21 under the unchanged seed-16-plus-job-index rule. Its confidence intervals are recomputed by the official scorer; all five existing point estimates, measurements, prices and native G_med remain unchanged. This is a canonical full-field uncertainty restatement, not a new API request, model inference, input draw or changed bootstrap algorithm. The canonical API-equating reference pool grows from five to six complete full-coverage self-hosted systems, including wrappers, and its offsets are recomputed; this separate pool does not change the four native G_med members. No API system is measured in this addendum.
Base-reference prices are labelled estimates with mixed historical cut-off and current public-provider dates; they are neither invoices nor bookable package hosting tariffs. Per-system cost basis preserves its exact documented source/date/precision convention.
Capability is mean(Intelligence, Calibration). The eligibility envelope remains the historical fixed Jev 1.13.0 v1.5.7 reference, not a newly measured API: USD 0.06459465517241379 per 1,000 and p50 1.2329566404223442 seconds caps. Composite is secondary.
Historical catalogue and scores/dates remain unchanged. The fresh version is added separately and retains all required page sections; pending candidates are not disguised as completed results.
Wrappers · listed, not ranked
These measured serving wrappers are listed below the rankings and are excluded from the native field median. Values are their stored v1.6.4 aggregates.
| System | Listing | Capability | Composite | Cost / 1,000 | Median latency |
|---|---|---|---|---|---|
| metask-jev-rain-12B | wrapper · not ranked · excluded from G_med | 78.3 | 68.7 | $0.0378 | 0.33 s |
| ryotide_qwen9 | wrapper · not ranked · excluded from G_med | 53.5 | 16.1 | $0.0488 | 0.41 s |
Historical model catalogue
All 173 previously listed systems retain their original published values. These scores were measured on earlier item sets and do not enter the fresh cohort’s ranking or field median.
| System | Measured | Capability | Composite | Cost / 1,000 | Median latency |
|---|---|---|---|---|---|
| Sage 1.3.0 (Levanto Labs) | 2026-10-05 | 78.6 | 74.0 | $0.0247 | 0.15 s |
| H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim) | 2026-10-04 | 75.0 | 72.5 | $0.0211 | 0.21 s |
| Mercury Decide (Inception; System One decisions API, served free on OpenRouter as inception/mercury-decide:free) | 2026-10-07 | 73.6 | 72.4 | $0.0184 | 0.33 s |
| decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted) | 2026-10-06 | 79.6 | 71.7 | $0.0408 | 0.24 s |
| Jev 1.13.0 (TypeSafe AI) | 2026-10-05 | 77.1 | 71.5 | $0.0323 | 0.24 s |
| Quyet-1.0-Large (Chinh Nguyen, Gemma-4-31B decoder, option-letter logits) | 2026-10-04 | 81.7 | 71.4 | $0.0445 | 0.38 s |
| decider-12b v2 (Mapika) | 2026-10-04 | 72.5 | 70.9 | $0.0256 | 0.26 s |
| wity-1 (Wity, reasoning auto) | 2026-10-06 | 79.1 | 70.8 | $0.0236 | 1.57 s |
| decider-12b v1, stock Gemma-4-12B-it (Mapika) | 2026-10-04 | 72.2 | 70.4 | $0.0256 | 0.26 s |
| torchcast-decision-12b (Torchcast AI, Gemma-4-12B fine-tune, option-letter logprob readout) | 2026-10-04 | 71.7 | 69.9 | $0.0285 | 0.22 s |
| Winnow-12B Q8 | 2026-10-02 | 71.2 | 68.9 | $0.0281 | 0.38 s |
| deck-31B (krishna765, frozen Gemma-4-31B-it, TorchAO FP8 dynamic) | 2026-10-05 | 77.6 | 68.7 | $0.0475 | 0.40 s |
| Cygnet (blockbrain, frozen Gemma-4-12B-it) | 2026-10-02 | 70.9 | 68.6 | $0.0283 | 0.22 s |
| decisio v0.8.0 on gemma-4-12B-it (frozen, one forward pass, self-hosted) | 2026-10-06 | 70.7 | 68.2 | $0.0227 | 0.35 s |
| decisio v0.9.0 on gemma-4-12B-it (frozen, one prefill per question, self-hosted) | 2026-10-07 | 70.0 | 67.8 | $0.0227 | 0.41 s |
| Jev-Omni (akhilaaa3, Gemma-4-12B merged) | 2026-10-02 | 71.3 | 67.7 | $0.0292 | 0.46 s |
| Xor 26B-A4B (Juspay, Gemma-4-26B-A4B, bf16) | 2026-10-06 | 72.4 | 67.4 | $0.0440 | 0.27 s |
| Bobcat Flash 1.2 (Gemma-4-26B-A4B-it + two merged rank-64 LoRA adapters, typed-decision readout at the first answer position, FP8, self-hosted) | 2026-10-07 | 73.5 | 67.3 | $0.0475 | 0.27 s |
| SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint) | 2026-10-06 | 74.3 | 66.7 | $0.0395 | 0.98 s |
| decider chat on Gemma-4-31B-it (Mapika, frozen base, inference technique) | 2026-10-06 | 70.6 | 66.1 | $0.0437 | 0.40 s |
| Hopper 12B trained (gemma-4-12B-it frozen + LoRA r32 unmerged, one forward pass, self-hosted) | 2026-10-06 | 68.1 | 65.8 | $0.0289 | 0.48 s |
| Surogate Rune 26B-A4B v3 (v1.6 pool) | 2026-10-06 | 73.5 | 65.0 | $0.0479 | 0.38 s |
| GEV-26B-Decide (AutoTrust, Gemma-4-26B-A4B + LoRA + head) | 2026-10-06 | 70.1 | 64.2 | $0.0426 | 0.25 s |
| SPX-CD-Omni (SurdAI, google/gemma-4-12B-it + LoRA r32, one forward pass, self-hosted) | 2026-10-06 | 68.1 | 63.2 | $0.0246 | 0.30 s |
| Diffusion Jev (DiffusionGemma 26B-A4B on patched SGLang, 48-step diffusion readout, self-hosted) | 2026-10-06 | 58.1 | 59.3 | $0.0480 | 0.41 s |
| René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout) | 2026-10-07 | 76.0 | 55.8 | $0.0633 | 0.37 s |
| Aplomb 1 (5.3B decision model on Qwen3.5-4B, trained readout head, self-hosted) | 2026-10-06 | 67.5 | 54.6 | $0.0254 | 0.34 s |
| Bespoke Nimble 9B v3 (Qwen3.5-9B LoRA) | 2026-10-06 | 66.8 | 53.5 | $0.0626 | 0.34 s |
| APUS-OpenJev-v1-9B (merged bf16 checkpoint-3000 on Qwen3.5-9B, their own native runtime, full depth, self-hosted) | 2026-10-07 | 56.9 | 49.2 | $0.0511 | 0.71 s |
| Plumb-4B (crh225, JevK5 v0.2 + LoRA) | 2026-10-02 | 65.4 | 48.3 | $0.0170 | 0.19 s |
| SPX-CD Pro (SurdAI, hosted /v1/systemone, Oct-4 checkpoint) | 2026-10-06 | 78.6 | 47.7 | $0.0790 | 1.13 s |
| Messier One v0.2 (Qwen3.5-4B fine-tune, one prefill, self-hosted) | 2026-10-07 | 58.1 | 45.3 | $0.0149 | 0.20 s |
| spark-s1-4b-v6 (Open Spark Jev, abhishek085) | 2026-10-05 | 54.0 | 44.9 | $0.0209 | 0.41 s |
| swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4) | 2026-10-02 | 72.4 | 43.8 | $0.0848 | 0.71 s |
| Quyet-1.0-Medium (Chinh Nguyen, Qwen3.5-4B decoder, option-letter logits) | 2026-10-04 | 59.2 | 43.6 | $0.0166 | 0.30 s |
| lev (Interfaze AI, Qwen3.5-4B + LoRA r32, label-token readout + candidate-path head, per-bucket temperatures) | 2026-10-04 | 61.7 | 42.2 | $0.0188 | 0.35 s |
| NInfer Qwen3.8-Flash-Next mixed | 2026-10-02 | 69.7 | 42.0 | $0.0823 | 0.31 s |
| jev-local (Qwen3.5-9B) | 2026-10-02 | 57.5 | 41.3 | $0.0235 | 0.88 s |
| decider-4b v2 (Mapika) | 2026-10-02 | 64.9 | 41.2 | $0.0152 | 0.20 s |
| ClassOne Qwen 3.5 9B (Qwen3.5-9B backbone with ClassOne decision heads, schema-constrained single forward pass, self-hosted) | 2026-10-07 | 50.5 | 38.5 | $0.0488 | 0.26 s |
| JevK5 v0.3 (4B) | 2026-10-02 | 64.2 | 37.4 | $0.0170 | 0.19 s |
| metask-jev-4b | 2026-10-02 | 61.8 | 37.3 | $0.0256 | 0.26 s |
| deck-4B v1.0 (krishna765, JevK5 + LoRA, FP8 weight-only) | 2026-10-05 | 63.6 | 37.3 | $0.0165 | 0.30 s |
| Clef-Flash (Cloudflare, Qwen3.5-9B post-train with a joint schema head, multimodal, measured on text) | 2026-10-04 | 63.8 | 36.6 | $0.0593 | 0.36 s |
| Decision-4B (Eval Engine / Chromia) | 2026-10-04 | 59.1 | 35.4 | $0.0137 | 0.24 s |
| janus 4B (Icarus AI / cmxu, Qwen3.5-4B + LoRA r64 and pointer decision head) | 2026-10-04 | 62.9 | 35.2 | $0.0154 | 0.21 s |
| Qwen3.5-9B Jev-like data-mix v2 | 2026-10-02 | 60.9 | 33.6 | $0.0646 | 0.49 s |
| JevK5 v0.2.0 | 2026-10-02 | 62.8 | 32.1 | $0.0170 | 0.23 s |
| JevOne | 2026-10-02 | 68.9 | 31.1 | $0.1012 | 0.27 s |
| Kahn1 4B (Okura66, Qwen3.5-4B LoRA merge, author's sysone engine) | 2026-10-07 | 61.7 | 31.1 | $0.0613 | 0.28 s |
| TypeCastLM 1.4.0 (Mikhail Gribov, Qwen3.5-4B computed head) | 2026-10-07 | 56.4 | 29.8 | $0.0149 | 0.21 s |
| Hopper | 2026-10-02 | 62.4 | 28.9 | $0.0181 | 0.26 s |
| Imajev-4B (RTX 5090) | 2026-10-02 | 61.9 | 28.7 | $0.0168 | 0.29 s |
| Malkuth-4B (newfull5, Kev post-train) | 2026-10-02 | 61.1 | 26.2 | $0.0297 | 0.30 s |
| jqv (Qwen3-32B zero-shot) | 2026-10-02 | 61.9 | 25.4 | $0.0424 | 0.34 s |
| JPT-35B-A3B (Qwen3.5-35B-A3B fine-tune) | 2026-10-06 | 70.7 | 25.4 | $0.1631 | 0.29 s |
| decider-35b-a3b (Mapika) | 2026-10-02 | 66.2 | 24.8 | $0.1539 | 0.25 s |
| CoCo-Decision-4B (corners-ai, LoRA on Qwen3.5-4B, option-label logits of one pass, served by oh-my-jev) | 2026-10-07 | 56.9 | 24.6 | $0.0145 | 0.24 s |
| Manchego v2.1 | 2026-10-04 | 58.3 | 24.5 | $0.0155 | 0.26 s |
| Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA) | 2026-10-02 | 60.4 | 24.5 | $0.0170 | 0.19 s |
| Raw Qwen3 4B Instruct 2507 direct logits | 2026-10-02 | 34.6 | 21.8 | $0.0166 | 0.24 s |
| JEV-27B (AutoTrust, Qwen3.8-27B) | 2026-10-06 | 73.9 | 21.3 | $0.1991 | 0.36 s |
| torchcast-decision-27b (Torchcast AI, Qwen3.8-27B fine-tune) | 2026-10-06 | 81.6 | 21.3 | $0.2099 | 0.34 s |
| Bespoke Nimble 9B (Bespoke Labs) | 2026-10-02 | 56.5 | 21.1 | $0.1283 | 0.44 s |
| reflex 4B (kshetrajna12) | 2026-10-02 | 59.0 | 20.6 | $0.0169 | 2.87 s |
| Perplexity Decider v1.1 27B (Qwen3.8-27B, noncausal decision readout) | 2026-10-06 | 82.8 | 20.6 | $0.2173 | 0.39 s |
| JADE (Qwen3.8-27B LoRA) | 2026-10-06 | 49.2 | 20.2 | $0.1637 | 0.40 s |
| typecastlm (Mikhail Gribov, Qwen3.5-4B computed head) | 2026-10-02 | 56.5 | 19.7 | $0.0159 | 0.21 s |
| Jebadiah 27B (Frontier Infra, Qwen3.8-27B LoRA) | 2026-10-06 | 70.3 | 19.3 | $0.2165 | 0.43 s |
| Decision 2.0 Vega 27B (vLLM-SR, Qwen3.8-27B) | 2026-10-06 | 71.3 | 18.9 | $0.2216 | 0.53 s |
| Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA) | 2026-10-02 | 57.2 | 18.7 | $0.0170 | 0.19 s |
| AutoJev-27B (denis-pplx, Qwen3.8-27B) | 2026-10-02 | 71.6 | 18.5 | $0.2262 | 0.34 s |
| Eikos-27B (caiovicentino1, Qwen3.8-27B) | 2026-10-02 | 77.2 | 18.1 | $0.2382 | 0.38 s |
| NInfer Qwen3.8-27B NVFP4 | 2026-10-02 | 68.0 | 17.7 | $0.2305 | 0.26 s |
| OpenJev (thinking, BF16) | 2026-10-02 | 75.9 | 17.3 | $0.2415 | 1.86 s |
| Open-Jev 9B (Zefan Cai) | 2026-10-02 | 61.8 | 17.0 | $0.1704 | 1.39 s |
| Standard One 8B (Standard Thinking) | 2026-10-02 | 58.4 | 16.9 | $0.0777 | 0.21 s |
| Clef (Cloudflare, Qwen3.8-27B post-train with a joint schema head, multimodal, measured on text) | 2026-10-04 | 75.3 | 16.8 | $0.2492 | 0.58 s |
| JEV Qwen3.5-9B Base NVFP4 | 2026-10-02 | 57.1 | 16.0 | $0.0558 | 0.20 s |
| SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) | 2026-10-02 | 55.4 | 15.5 | $0.0170 | 0.25 s |
| Bev / Bonsai 27B | 2026-10-04 | 66.1 | 11.7 | $0.2468 | 2.12 s |
| LitJev (Qwen3.8-27B) | 2026-10-02 | 66.6 | 11.5 | $0.2444 | 3.10 s |
| local-jev Qwen3.5-4B | 2026-10-02 | 55.0 | 11.4 | $0.0228 | 0.40 s |
| system-one (Qwen3-8B, Sean Goedecke) | 2026-10-02 | 29.7 | 11.0 | $0.0683 | 0.25 s |
| kev 8B (research preview) | 2026-10-02 | 50.0 | 10.6 | $0.0968 | 0.30 s |
| Decision 2B (FlyMy.AI, v59) | 2026-10-05 | 54.8 | 10.3 | $0.0132 | 0.30 s |
| OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp) | 2026-10-02 | 48.6 | 10.3 | $0.0107 | 0.74 s |
| open-alternative-jev (Qwen3.5-4B, IkerMoel) | 2026-10-02 | 49.1 | 9.4 | $0.0168 | 0.24 s |
| Raw Qwen3 8B direct logits | 2026-10-02 | 30.1 | 9.3 | $0.0654 | 0.31 s |
| decider-2b (Mapika) | 2026-10-02 | 44.6 | 7.7 | $0.0148 | 0.18 s |
| Nemotron Diffusion 8B (pst2154, optimized vLLM) | 2026-10-04 | 47.5 | 7.3 | $0.0305 | 0.19 s |
| kev 4B (research preview) | 2026-10-02 | 44.0 | 6.6 | $0.0138 | 0.29 s |
| Kev 27B (Kev 1.0, Qwen3.8-27B) | 2026-10-06 | 78.9 | 5.8 | $0.5288 | 0.51 s |
| Malkuth-2B (newfull5, Kev post-train) | 2026-10-02 | 47.6 | 4.0 | $0.0141 | 0.24 s |
| Qwen3-Reranker-4B | 2026-10-02 | 44.7 | 3.7 | $0.0525 | 0.56 s |
| SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool) | 2026-10-06 | 74.9 | 3.4 | $0.6760 | 1.39 s |
| ZeroEntropy zerank-2 | 2026-10-02 | 49.4 | 3.3 | $0.0525 | 0.50 s |
| EXAONE-4.0-1.2B-JEV v0.3 (unofficial fine-tune of LGAI-EXAONE/EXAONE-4.0-1.2B, native option-label softmax, two reads averaged, self-hosted) | 2026-10-07 | 47.7 | 3.3 | $0.0253 | 0.19 s |
| Open-Jev 2B (Zefan Cai) | 2026-10-02 | 47.7 | 3.2 | $0.1704 | 1.02 s |
| Open-Jev 27B v1.1 (Zefan Cai, LoRA + decision head on Qwen3.8-27B) | 2026-10-04 | 68.0 | 2.9 | $0.7157 | 1.73 s |
| Raw Phi-4 mini direct logits | 2026-10-02 | 35.3 | 2.5 | $0.0363 | 0.23 s |
| SimpleJev (Qwen3.5-0.8B, CPU) | 2026-10-02 | 31.6 | 2.1 | $0.0098 | 8.80 s |
| smalljev semantic-v9 | 2026-10-02 | 43.0 | 0.8 | $0.0203 | 0.28 s |
| kev 0.6B (research preview) | 2026-10-02 | 36.9 | 0.5 | $0.0046 | 0.26 s |
| jul fast (usejul/jul-decision-minicpm5-2b: MiniCPM5-2B fine-tune with a pointer head, option probabilities from one pass, self-hosted via `jul serve`) | 2026-10-07 | 42.8 | 0.5 | $0.0108 | 0.24 s |
| OpenDecision (ModernBERT-large zero-shot) | 2026-10-02 | 38.0 | 0.5 | $0.0050 | 0.29 s |
| ClassOne Gemma 4 E2B (Gemma-4-E2B backbone with ClassOne decision heads, schema-constrained single forward pass, self-hosted) | 2026-10-07 | 15.2 | 0.5 | $0.0098 | 0.21 s |
| Qwen3.5-0.8B Decision Model (Mourad Ghafiri) | 2026-10-02 | 42.6 | 0.5 | $0.0048 | 11.25 s |
| Fastino GLiNER-2.5-Decide (hosted API) | 2026-10-05 | 34.6 | 0.3 | $0.0640 | 0.44 s |
| Deem 0.8B v1 | 2026-10-04 | 18.1 | 0.3 | $0.0045 | 0.46 s |
| Decision Fast (FlyMy.AI, v53a) | 2026-10-05 | 40.6 | 0.3 | $0.0046 | 0.27 s |
| Quyet-1.0-Small-EN (Chinh Nguyen, ModernBERT-base encoder with fixed heads) | 2026-10-04 | 39.5 | 0.3 | $0.0023 | 0.18 s |
| Laya typed-decisions | 2026-10-04 | 44.2 | 0.2 | $0.0038 | 1.88 s |
| WaterSheep (Samrat Dutta, ModernBERT-base encoder with calibrated heads, served by the author's own `watersheep --serve`) | 2026-10-07 | 35.6 | 0.2 | $0.0111 | 0.95 s |
| openJev Verdict 1.4 | 2026-10-02 | 40.4 | 0.2 | $0.0028 | 0.74 s |
| Raw Qwen3 0.6B direct logits | 2026-10-02 | 8.2 | 0.2 | $0.0056 | 0.21 s |
| lev-350m (Franck Verrot, LFM2.5-350M) | 2026-10-05 | 43.3 | 0.2 | $0.0046 | 0.19 s |
| Raw Qwen3 1.7B direct logits | 2026-10-02 | 8.3 | 0.1 | $0.0112 | 0.21 s |
| kev 0.5B | 2026-10-02 | 32.9 | 0.1 | $0.0046 | 0.24 s |
| openJev Verdict (heman10x, ModernBERT-base 151M) | 2026-10-02 | 26.6 | 0.1 | $0.0028 | 0.69 s |
| Quyet-1.0-Small (Chinh Nguyen, SEA-LION-ModernBERT-300M encoder with fixed heads) | 2026-10-04 | 32.5 | 0.1 | $0.0048 | 0.18 s |
| watt-flash-0.1 (Zaitgeist Labs, 140M mmBERT-small encoder with an option-marker scorer, one forward pass per request, self-hosted) | 2026-10-07 | 41.6 | 0.1 | $0.0023 | 1.01 s |
| Quyet-1.0-Tiny (Chinh Nguyen, mmBERT-small encoder (16 layers kept) with fixed heads) | 2026-10-04 | 31.1 | 0.1 | $0.0024 | 0.18 s |
| open-jev-deberta-v3-large (local CPU) | 2026-10-02 | 35.4 | 0.0 | $0.0056 | 3.53 s |
| verdict-small (Manavarya09, multilingual-e5-small 118M) | 2026-10-02 | 37.6 | 0.0 | $0.0009 | 0.25 s |
| Tacet Sonata (CodePawl, 144M packed encoder on mmBERT-small, option-marker softmax in one pass, in-process via the author's tacet package) | 2026-10-07 | 24.3 | 0.0 | $0.0023 | 0.61 s |
| Laya (Convai Innovations, ModernBERT-large 421M) | 2026-10-02 | 32.5 | 0.0 | $0.0032 | 1.80 s |
| Mixedbread mxbai-rerank-base-v2 | 2026-10-02 | 43.9 | 0.0 | $0.0210 | 0.21 s |
| Laya multilingual | 2026-10-04 | 17.0 | 0.0 | $0.0039 | 0.87 s |
| Certo v1 (AltSlate Labs) | 2026-10-02 | 43.2 | 0.0 | $0.0013 | 0.23 s |
| CLM-8B (Contrastive-LM, clm-latest) | 2026-10-02 | 20.1 | 0.0 | $0.0447 | 0.19 s |
| BAAI bge-reranker-v2-m3 | 2026-10-02 | 44.1 | 0.0 | $0.0227 | 0.18 s |
| Alibaba GTE Reranker ModernBERT-base | 2026-10-02 | 41.7 | 0.0 | $0.0106 | 0.20 s |
| Mirror | 2026-10-02 | 25.8 | 0.0 | $0.0023 | 1.45 s |
| Open Jev JSON Canvas (JoshuaSP) | 2026-10-02 | 30.6 | 0.0 | $0.0490 | 0.45 s |
| APUS-OpenJev-v1-35B-A3B (merged bf16 checkpoint-5949 MoE on Qwen3.5-35B-A3B, their own native runtime, full depth, self-hosted) | 2026-10-07 | 60.4 | — | — | 0.86 s |
| BB-Qwen3.5-4B-LoRA (Babak Barazandeh, LoRA on Qwen3.5-4B-Base, hosted /v1/systemone) | 2026-10-07 | 60.4 | — | — | 0.70 s |
| Seb-9B (Qwen3.5-9B-shaped multimodal decision model, one forward pass, stock vLLM, self-hosted) | 2026-10-07 | 66.8 | — | — | 0.27 s |
| wity-1 (Wity, reasoning always) | 2026-10-06 | 79.3 | 70.6 | $0.0236 | 1.69 s |
| wity-1 (Wity, reasoning off) | 2026-10-06 | 63.0 | 49.4 | $0.0236 | 0.28 s |
| AutoJev-27B (RTX PRO 6000) | measured on v1.5.0 (2026-09-28) | 79.7 | 19.5 | $0.2262 | 0.33 s |
| Bosun v3.1 0.6B | measured on v1.5.3 (2026-09-29) | 39.4 | 2.5 | $0.0056 | 4.04 s |
| classifier.dev (fast tier) | Not yet measured | — | — | — | — |
| decision-machine-1 (milliseconds.ai) | measured on v1.5.0 (2026-09-28) | 48.0 | 3.2 | $0.0286 | 0.18 s |
| DeepSeek V4.1 Flash (thinking default) | measured on v1.5.0 (2026-09-28) | 95.3 | 6.6 | $0.4976 | 1.78 s |
| Gemini 3.1 Flash-Lite | measured on v1.5.0 (2026-09-28) | 76.1 | 19.6 | $0.2194 | 0.86 s |
| GLiNER2 (Fastino, gliner2.5-base) | measured on v1.5.0 (2026-09-28) | 24.5 | 2.3 | $0.0028 | 1.00 s |
| GLiNER2 large (Fastino) | measured on v1.5.0 (2026-09-28) | 32.4 | 8.4 | $0.0056 | 1.88 s |
| GLiNER2.5 multi (Fastino, 287M) | measured on v1.5.0 (2026-09-28) | 36.0 | 2.7 | $0.0028 | 1.37 s |
| GLiNER2.5 small (Fastino, 74M) | measured on v1.5.0 (2026-09-28) | 32.5 | 0.9 | $0.0028 | 0.47 s |
| GPT-5.6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 94.5 | 22.4 | $0.2047 | 1.32 s |
| GPT-6 Luna (default medium reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.9 | 38.8 | $0.1138 | 1.56 s |
| GPT-6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.1 | 40.5 | $0.1075 | 1.58 s |
| Instinct (ZooWork, Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 73.9 | 18.3 | $0.2298 | 0.52 s |
| Instinct Dual 4B | measured on v1.5.3 (2026-09-29) | 65.7 | 47.0 | $0.0212 | 0.51 s |
| jeff (Logan Markewich, GLiFormer 400M) | measured on v1.5.0 (2026-09-28) | 42.4 | 0.1 | $0.0043 | 7.14 s |
| JevAct (einptein, jev1-2b-v2) | measured on v1.5.0 (2026-09-28) | 36.9 | 1.5 | $0.0112 | 0.80 s |
| Jobe Qwen3.5-4B (frozen) | Not yet measured | — | — | — | — |
| Autoloops – Gemma 4 31B IT | measured on v1.5.0 (2026-09-28) | 81.3 | 40.5 | $0.1032 | 0.61 s |
| mica-v01-4b | Not yet measured | — | — | — | — |
| Needle 3 (Cactus, 2-bit, local CPU) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 | $0.0191 | 135.51 s |
| Needle 3, options as tools (post-hoc adapter mode) | measured on v1.5.0 (2026-09-28) | 0.0 | 0.0 | $0.0191 | 58.16 s |
| NInfer Qwen3.8-27B NVFP4 (T=1.5) | measured on v1.5.0 (2026-09-28) | 73.8 | 18.5 | $0.2305 | 0.25 s |
| OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) | Not yet measured | — | — | — | — |
| openjev-sglang (Qwen3.6-35B-A3B on SGLang) | measured on v1.5.0 (2026-09-28) | 70.8 | 29.0 | $0.1398 | 1.20 s |
| Qwen3.8 27B (Chutes TEE) | measured on v1.5.0 (2026-09-28) | 96.8 | 0.0 | $2.1784 | 6.49 s |
| reflex-27b (Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 74.4 | 13.2 | $0.2973 | 2.75 s |
| SimpleJev Qwen3.6-35B-A3B | Not yet measured | — | — | — | — |
| SimpleJev Qwen3.8-27B | measured on v1.5.0 (2026-09-28) | 80.0 | 3.4 | $0.6868 | 1.68 s |
| Surogate Rune 26B-A4B v3 (RTX PRO 6000) | measured on v1.5.0 (2026-09-28) | 79.0 | 66.5 | $0.0502 | 0.35 s |
| system-one-open (Gemma 4 E2B LoRA on an L4) | measured on v1.5.0 (2026-09-28) | 56.6 | 42.4 | $0.0114 | 1.15 s |
| Vansa-3.4 (Vansa, hosted System One API) | measured on v1.5.6 (2026-10-03) | 72.8 | 71.6 | $0.0193 | 0.14 s |
| Von (wfzyx, Option-Marker 395M) | measured on v1.5.0 (2026-09-28) | 41.7 | 0.0 | $0.0038 | 0.92 s |
Historical results SHA-256 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1. Previous boards: v1.6.2 · v1.6.1.
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
Open this section to load the earlier public-only board and diagnostics.