Official JevBench v1.6.2

JevBench · fresh native cohort

Paid submissions measured on the same fresh 1,500-decision draw with methodology v1.6. Historical scores retain their original dates and are listed separately below.

Aggregate results JSON · SHA-256 4464184cb5a867f2694ae709d1daf36185f615a08807bb2e5f11f7a6d852eead · Full historical board

JevBench v1.6.2 · headline

JevBench Capability Score

Capability ranking of Jev-class systems

Capability Score averages Intelligence and Calibration. Jeff-1.0-Large leads the Jev-class systems with 81.5.

Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘

Adjust cost / latency caps · 2× official
Official
#SystemICCap.$/1k
  1. 1Jeff-1.0-Large71.949.281.5$0.049*
  2. 2Metask rain 4B51.151.866.0$0.040*
  3. 3Wald 4B41.154.359.8$0.033*

Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.

Cost and latency correlate here (Spearman ρ = 1.00, n = 3).

  • Open weights · LLM decoder (3)
  • green: ≤ reference
  • amber: ≤ cap (2× reference)
  • red: > cap
Show general-purpose LLMs and other systems outside the limits

Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).

    Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 3 of 3 systems qualify; the other 0, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.

    Capability against cost and speed

    Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.

    Capability vs cost

    Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.

    Capability vs cost: 3 systems. Upper right is best: more capable and cheaper.5060708090100$0.010$0.10$ per 1,000 decisions (log)Capability↑2× Jev (v1.5 reference)← priciercheaper →1. Jeff-1.0-Large2. Metask rain 4B3. Wald 4B
    3 systems. Tap a bubble for its values.

    Capability vs speed

    Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.

    Capability vs speed: 3 systems. Upper right is best: more capable and faster.506070809010070≈3.2 s80≈1.0 s90≈316 ms100≈100 msMedian-latency speedCapability↑2× Jev (v1.5 reference)← slowerfaster →1. Jeff-1.0-Large2. Metask rain 4B3. Wald 4B
    3 systems. Tap a bubble for its values.
    • Open weights · LLM decoder (3)
    • faint = outside Jev-class
    4 of 4 systems

    Model kind

    Jev-class

    Release

    No provider reported in this release.

    No family reported in this release.

    No licence reported in this release.

    Parameters

    No system in this release reports an exact parameter count.

    Developer/API price $/1k

    No system in this release reports a developer/API price.

    Base-model price $/1k

    No system in this release reports a base-model reference price.

    Scored cost $/1k
    Alternative price $/1k

    No system in this release reports an alternative pricing scenario.

    p50 latency (s)
    p95 latency (s)

    JevBench v1.6.2

    JevBench Composite Score: 3 ranked systems

    Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓

    Weights:
    Adjust weights ↓

    View by:Capability ↑

    The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis.

    Greener = stronger within its column.

    3 of 3 systems, sorted by official rank, #1 first.

    1. 1Jeff-1.0-Largenewfine-tune · 31BBase model: undisclosed68.6I 72C 91S 89K 49est.$0.049
    2. 2Metask rain 4Bnewfine-tune · 4BBase model: undisclosed62.7I 51C 81S 80K 52est.$0.040
    3. 3Wald 4Bnewfine-tune · 4BBase model: undisclosed40.0I 41C 78S 82K 54est.$0.033
    Weights:
    Adjust weights ↓

    Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.

    Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

    • Open weights · LLM decoder (3)
    I, C, S, K = Intelligence, Calibration, Speed, Cost; the est. pill = estimated cost; ann. = announced price; API = the operator's endpoint saw held-out benchmark inputs, without answers; $/1k decisions = US dollars per 1,000 decisions (not heat-shaded). Names link to each project.

    Compare two systems

    Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

    The four score axes

    Radar: the four score axes, two systemsThe four score axes, Jeff-1.0-Large vs Metask rain 4B. Intelligence: 71.9 vs 51.1; Calibration: 91.1 vs 80.9; Speed: 88.8 vs 79.6; Cost: 49.2 vs 51.8.50100Intelligence: Jeff-1.0-Large: 71.9; Metask rain 4B: 51.1Intelligence71.9 · 51.1Calibration: Jeff-1.0-Large: 91.1; Metask rain 4B: 80.9Calibration91.1 · 80.9Speed: Jeff-1.0-Large: 88.8; Metask rain 4B: 79.6Speed88.8 · 79.6Cost: Jeff-1.0-Large: 49.2; Metask rain 4B: 51.8Cost49.2 · 51.8
    0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

    Capability by subject topic

    Radar: capability by subject topic, two systemsCapability by subject topic, Jeff-1.0-Large vs Metask rain 4B. Rules, policy & law: 79.7 vs 51.2; Math & numbers: 48.3 vs 29.4; Coding & software: 83.9 vs 61.0; Support & operations: 56.7 vs 50.9; Finance & commerce: 74.7 vs 55.1; Everyday language: 52.4 vs 48.0; Safety & security: 90.2 vs 53.7. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Rules, policy & law: Jeff-1.0-Large: 79.7; Metask rain 4B: 51.2 — Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 747 items (153 open, 594 sealed).Rules, policy& law79.7 · 51.2Math & numbers: Jeff-1.0-Large: 48.3; Metask rain 4B: 29.4 — Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 282 items (37 open, 245 sealed).Math &numbers48.3 · 29.4Coding & software: Jeff-1.0-Large: 83.9; Metask rain 4B: 61.0 — Coding & software: code, SQL, repositories, developer tools and IT systems. 197 items (61 open, 136 sealed).Coding &software83.9 · 61.0Support & operations: Jeff-1.0-Large: 56.7; Metask rain 4B: 50.9 — Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 100 items (22 open, 78 sealed).Support &operations56.7 · 50.9Finance & commerce: Jeff-1.0-Large: 74.7; Metask rain 4B: 55.1 — Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 79 items (11 open, 68 sealed).Finance &commerce74.7 · 55.1Everyday language: Jeff-1.0-Large: 52.4; Metask rain 4B: 48.0 — Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 53 items (10 open, 43 sealed).Everydaylanguage52.4 · 48.0Safety & security: Jeff-1.0-Large: 90.2; Metask rain 4B: 53.7 — Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 42 items (6 open, 36 sealed).Safety &security90.2 · 53.7

    Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

    Measured subject-topic competence on this release’s item set. Chance-corrected competence per category (0 = at or below chance; negative averages are reported as 0, 100 = perfect), Jeff-1.0-Large: open+sealed; Metask rain 4B: open+sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
    What each category means · items per category
    • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 747 items (153 open / 594 sealed)
    • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 282 items (37 open / 245 sealed)
    • Coding & software — code, SQL, repositories, developer tools and IT systems. 197 items (61 open / 136 sealed)
    • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 100 items (22 open / 78 sealed)
    • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 79 items (11 open / 68 sealed)
    • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 53 items (10 open / 43 sealed)
    • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 42 items (6 open / 36 sealed)

    Use cases (TypeSafe categories)

    Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jeff-1.0-Large vs Metask rain 4B. Legal & compliance: 84.5 vs 51.5; Model routing: 84.4 vs 68.1; Other: 36.7 vs 23.6; Customer support: 66.5 vs 49.6; Financial crime: 62.7 vs 53.4; Risk assessment: 76.2 vs 40.7; E-commerce marketplaces: 67.8 vs 40.4; Moderation and trust and safety: 87.0 vs 39.9; Lead generation: 78.7 vs 60.6; Advertising: 49.6 vs 28.1. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Legal & compliance: Jeff-1.0-Large: 84.5; Metask rain 4B: 51.5 — Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 343 items (77 open, 266 sealed).Legal &compliance84.5 · 51.5Model routing: Jeff-1.0-Large: 84.4; Metask rain 4B: 68.1 — Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 318 items (86 open, 232 sealed).Modelrouting84.4 · 68.1Other: Jeff-1.0-Large: 36.7; Metask rain 4B: 23.6 — Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 213 items (23 open, 190 sealed).Other36.7 · 23.6Customer support: Jeff-1.0-Large: 66.5; Metask rain 4B: 49.6 — Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 158 items (37 open, 121 sealed).Customersupport66.5 · 49.6Financial crime: Jeff-1.0-Large: 62.7; Metask rain 4B: 53.4 — Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 54 items (9 open, 45 sealed).Financialcrime62.7 · 53.4Risk assessment: Jeff-1.0-Large: 76.2; Metask rain 4B: 40.7 — Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 53 items (7 open, 46 sealed).Riskassessment76.2 · 40.7E-commerce marketplaces: Jeff-1.0-Large: 67.8; Metask rain 4B: 40.4 — E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 43 items (9 open, 34 sealed).E-commercemarketplaces67.8 · 40.4Moderation and trust and safety: Jeff-1.0-Large: 87.0; Metask rain 4B: 39.9 — Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 37 items (7 open, 30 sealed).Moderationand trustand safety87.0 · 39.9Lead generation: Jeff-1.0-Large: 78.7; Metask rain 4B: 60.6 — Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 36 items (4 open, 32 sealed).Leadgeneration78.7 · 60.6Advertising: Jeff-1.0-Large: 49.6; Metask rain 4B: 28.1 — Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open, 27 sealed).Advertising49.6 · 28.1

    Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

    Measured application-category competence on this release’s item set. Chance-corrected competence per category (0 = at or below chance; negative averages are reported as 0, 100 = perfect), Jeff-1.0-Large: open+sealed; Metask rain 4B: open+sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.

    Low sample, n < 30 — indicative only

    Category (items)A: Jeff-1.0-LargeB: Metask rain 4B
    Insurance claims (29)65.5 n=2950.9 n=29
    LLM guardrails (25)74.8 n=2547.7 n=25
    Search and retrieval (24)82.2 n=2470.4 n=24
    Semantic code linting (22)84.2 n=2228.3 n=22
    Graphs and knowledge graphs (22)64.3 n=2213.4 n=22
    Demand forecasting (21)19.3 n=2121.9 n=21
    Scientific discovery (21)59.9 n=2117.5 n=21
    Gaming (18)45.3 n=1825.9 n=18
    Recruiting (17)66.1 n=1767.4 n=17
    Feature extraction for predictive modeling (16)71.1 n=165.9 n=16

    Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

    What each category means · items per category
    • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 343 items (77 open / 266 sealed)
    • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 318 items (86 open / 232 sealed)
    • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 213 items (23 open / 190 sealed)
    • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 158 items (37 open / 121 sealed)
    • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 54 items (9 open / 45 sealed)
    • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 53 items (7 open / 46 sealed)
    • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 43 items (9 open / 34 sealed)
    • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 37 items (7 open / 30 sealed)
    • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 36 items (4 open / 32 sealed)
    • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open / 27 sealed)

    Competence per request type, open / sealed

    Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jeff-1.0-Large vs Metask rain 4B. Choice · open: 76.9 vs 73.6; Choice · sealed: 81.3 vs 58.7; Noul · open: 85.3 vs 54.9; Noul · sealed: 77.2 vs 56.7; Score · open: 59.6 vs 39.8; Score · sealed: 51.0 vs 22.8. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Choice · open: Jeff-1.0-Large: 76.9; Metask rain 4B: 73.6Choice · open76.9 · 73.6Choice · sealed: Jeff-1.0-Large: 81.3; Metask rain 4B: 58.7Choice ·sealed81.3 · 58.7Noul · open: Jeff-1.0-Large: 85.3; Metask rain 4B: 54.9Noul · open85.3 · 54.9Noul · sealed: Jeff-1.0-Large: 77.2; Metask rain 4B: 56.7Noul · sealed77.2 · 56.7Score · open: Jeff-1.0-Large: 59.6; Metask rain 4B: 39.8Score · open59.6 · 39.8Score · sealed: Jeff-1.0-Large: 51.0; Metask rain 4B: 22.8Score ·sealed51.0 · 22.8

    Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

    Chance-corrected competence for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

    Competence per tier — open set

    Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jeff-1.0-Large vs Metask rain 4B. Easy: 95.4 vs 83.7; Standard: 68.3 vs 57.5; Judge: 67.4 vs 60.6; Hard: 76.1 vs 54.7. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Jeff-1.0-Large: 95.4; Metask rain 4B: 83.7Easy95.4 · 83.7Standard: Jeff-1.0-Large: 68.3; Metask rain 4B: 57.5Standard68.3 · 57.5Judge: Jeff-1.0-Large: 67.4; Metask rain 4B: 60.6Judge67.4 · 60.6Hard: Jeff-1.0-Large: 76.1; Metask rain 4B: 54.7Hard76.1 · 54.7

    Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

    Per-tier competence, the three request types pooled by their published decision counts.

    Competence per tier — sealed set

    Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jeff-1.0-Large vs Metask rain 4B. Easy: 86.4 vs 58.7; Standard: 81.0 vs 53.5; Judge: 68.1 vs 46.5; Hard: 68.7 vs 46.4. Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.−1000100Easy: Jeff-1.0-Large: 86.4; Metask rain 4B: 58.7Easy86.4 · 58.7Standard: Jeff-1.0-Large: 81.0; Metask rain 4B: 53.5Standard81.0 · 53.5Judge: Jeff-1.0-Large: 68.1; Metask rain 4B: 46.5Judge68.1 · 46.5Hard: Jeff-1.0-Large: 68.7; Metask rain 4B: 46.4Hard68.7 · 46.4

    Linear signed scale: −100 at the centre, 0 on the bold middle ring, 100 at the rim.

    Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
    How the categories were made

    O1S chance-corrected competence: equal mean over request types present, clipped to 0–100, using the pinned official group_stats function. Recorded failures remain in the denominator. Category scores are descriptive and do not alter headline Intelligence, Calibration, composite score or G_med.

    Subject topics and TypeSafe use cases use the actual shared local reference labels, with frozen T1/T2/T3/U1/U2 authoring rules. The original probabilities and model choice remain preserved in protected custody. The genuine 75-public-item handcheck is required. Family, type and language are frozen authoring metadata; non-English machine-authored items are not native-reviewed.

    • Every nonempty category has a measured cell; zero-count taxonomy categories have no cell.
    • Categories below 30 items are indicative only; radar display retains its 30-item minimum.
    • No historic overlays, estimated cells, outcome exclusions or private item-level data are included.
    All values as a table
    SpokeA: Jeff-1.0-LargeB: Metask rain 4B
    The four score axes
    Intelligence71.951.1
    Calibration91.180.9
    Speed88.879.6
    Cost49.251.8
    Capability by subject topic
    Rules, policy & law79.751.2
    Math & numbers48.329.4
    Coding & software83.961.0
    Support & operations56.750.9
    Finance & commerce74.755.1
    Everyday language52.448.0
    Safety & security90.253.7
    Use cases (TypeSafe categories)
    Legal & compliance84.551.5
    Model routing84.468.1
    Other36.723.6
    Customer support66.549.6
    Financial crime62.753.4
    Risk assessment76.240.7
    E-commerce marketplaces67.840.4
    Moderation and trust and safety87.039.9
    Lead generation78.760.6
    Advertising49.628.1
    Competence per request type, open / sealed
    Choice · open76.973.6
    Choice · sealed81.358.7
    Noul · open85.354.9
    Noul · sealed77.256.7
    Score · open59.639.8
    Score · sealed51.022.8
    Competence per tier — open set
    Easy95.483.7
    Standard68.357.5
    Judge67.460.6
    Hard76.154.7
    Competence per tier — sealed set
    Easy86.458.7
    Standard81.053.5
    Judge68.146.5
    Hard68.746.4

    All four completed native measurements cover the same fresh 1,500-item draw (1,200 sealed and 300 public); pending candidates are outside this completed category artifact. Sealed counts refer to self-hosted S (1,200); API rows use their original measured sealed basis.

    Languages

    Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: Self-hosted rows use S 1200 + P 300; API rows use their measured sealed subset plus P. Sage (A3) language cells and A2/A3 topic/use-case cells cover public P300 only. Raw and outside the Composite. Cells under 1 completed supported responses are left empty; recorded input refusals count toward coverage. A dagger (†) marks fewer than 30 completed responses. Competence retains all supported scored observations, including failures and refusals. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.2. The mixed-language group is listed first, then English (1,135 items); the other 21 languages and the mixed-language group share 365 items.

    JevBench v1.6.2 competence by item language and system
    Systemmixed†15en1135de40pl39es30hi†26pt†24ja†22ar†17el†17nl†17da†16fr†16it†13zh†12tr†11cs†9ko†9uk†9sv†8fi†7id†4no†4
    Jeff-1.0-Largeself-hosted · S+P99†7191648984†67†49†62†59†77†56†98†79†77†77†46†91†70†34†94†100†68†
    Metask rain 4Bself-hosted · S+P100†4862606541†74†8†0†73†0†61†30†93†52†6†36†65†81†30†0†34†68†
    Wald 4Bself-hosted · S+P95†4454355126†52†0†14†15†8†51†47†76†0†77†69†74†13†8†2†34†0†

    Intelligence gate and Noul decisiveness

    The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.

    SystemIntelligenceChoiceNoulScoreNoul decisive rateAccuracy among decisive
    Jeff-1.0-Large71.9nativenativenativenot supported—
    Metask rain 4B51.1nativenativenativenot supported—
    Wald 4B41.1below gatenativenativenativenot supported—
    All data (4 systems)

    Method notes · fresh native cohort

    Each system answered the same 1,200 freshly drawn sealed decisions and 300 public decisions on an offline GPU pod. Scores use methodology v1.6, O1S, with 1,000 bootstrap samples and seed 16.

    Capability is the mean of Intelligence and Calibration within the frozen Jev-class cost and median-latency caps. Composite is secondary. Failed and refused answers remain in the full 1,500-decision denominator.

    The public-versus-sealed gap reference is the median of this completed native cohort: 4.1 points. Historical scores are shown separately with their original dates and do not enter this cohort’s median or ranking. No API equating is applied.

    Costs use each row’s documented price reference and measured usage on this draw. Subject-topic and use-case radars come from stored per-item results; only aggregates are published.

    Sliders, presets and What-If change your view; they do not change the official result. Wrappers and subsidised systems are listed below the ranking.

    Results SHA-256 4464184cb5a867f2694ae709d1daf36185f615a08807bb2e5f11f7a6d852eead · categories SHA-256 21c7b174f693f307bf53bcaef6c795701388acdf73b576ea0fddfd7c9c6b1280 · scoring source SHA-256 db4b1244b4660b364f5cb9f1d90e4ca3e2ec7cd60fd9b0a918991d1f512fcb1d.

    Wrappers and listed systems · not ranked

    These measured systems are listed below the rankings and do not enter the native field median. Their published scores and category values remain available in All data above.

    Measured unranked systems in the fresh JevBench v1.6.2 cohort
    SystemListingCapabilityCompositeCost / 1,000Median latency
    metask-jev-rain-12Bwrapper · not ranked78.368.7$0.03780.33 s

    Historical model catalogue

    All 173 previously listed systems retain their original published values. These scores were measured on earlier item sets and do not enter the fresh cohort’s ranking or field median.

    Full historical JevBench catalogue, outside the fresh cohort ranking
    SystemMeasuredCapabilityCompositeCost / 1,000Median latency
    Sage 1.3.0 (Levanto Labs)2026-10-0578.674.0$0.02470.15 s
    H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim)2026-10-0475.072.5$0.02110.21 s
    Mercury Decide (Inception; System One decisions API, served free on OpenRouter as inception/mercury-decide:free)2026-10-0773.672.4$0.01840.33 s
    decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted)2026-10-0679.671.7$0.04080.24 s
    Jev 1.13.0 (TypeSafe AI)2026-10-0577.171.5$0.03230.24 s
    Quyet-1.0-Large (Chinh Nguyen, Gemma-4-31B decoder, option-letter logits)2026-10-0481.771.4$0.04450.38 s
    decider-12b v2 (Mapika)2026-10-0472.570.9$0.02560.26 s
    wity-1 (Wity, reasoning auto)2026-10-0679.170.8$0.02361.57 s
    decider-12b v1, stock Gemma-4-12B-it (Mapika)2026-10-0472.270.4$0.02560.26 s
    torchcast-decision-12b (Torchcast AI, Gemma-4-12B fine-tune, option-letter logprob readout)2026-10-0471.769.9$0.02850.22 s
    Winnow-12B Q82026-10-0271.268.9$0.02810.38 s
    deck-31B (krishna765, frozen Gemma-4-31B-it, TorchAO FP8 dynamic)2026-10-0577.668.7$0.04750.40 s
    Cygnet (blockbrain, frozen Gemma-4-12B-it)2026-10-0270.968.6$0.02830.22 s
    decisio v0.8.0 on gemma-4-12B-it (frozen, one forward pass, self-hosted)2026-10-0670.768.2$0.02270.35 s
    decisio v0.9.0 on gemma-4-12B-it (frozen, one prefill per question, self-hosted)2026-10-0770.067.8$0.02270.41 s
    Jev-Omni (akhilaaa3, Gemma-4-12B merged)2026-10-0271.367.7$0.02920.46 s
    Xor 26B-A4B (Juspay, Gemma-4-26B-A4B, bf16)2026-10-0672.467.4$0.04400.27 s
    Bobcat Flash 1.2 (Gemma-4-26B-A4B-it + two merged rank-64 LoRA adapters, typed-decision readout at the first answer position, FP8, self-hosted)2026-10-0773.567.3$0.04750.27 s
    SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)2026-10-0674.366.7$0.03950.98 s
    decider chat on Gemma-4-31B-it (Mapika, frozen base, inference technique)2026-10-0670.666.1$0.04370.40 s
    Hopper 12B trained (gemma-4-12B-it frozen + LoRA r32 unmerged, one forward pass, self-hosted)2026-10-0668.165.8$0.02890.48 s
    Surogate Rune 26B-A4B v3 (v1.6 pool)2026-10-0673.565.0$0.04790.38 s
    GEV-26B-Decide (AutoTrust, Gemma-4-26B-A4B + LoRA + head)2026-10-0670.164.2$0.04260.25 s
    SPX-CD-Omni (SurdAI, google/gemma-4-12B-it + LoRA r32, one forward pass, self-hosted)2026-10-0668.163.2$0.02460.30 s
    Diffusion Jev (DiffusionGemma 26B-A4B on patched SGLang, 48-step diffusion readout, self-hosted)2026-10-0658.159.3$0.04800.41 s
    René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout)2026-10-0776.055.8$0.06330.37 s
    Aplomb 1 (5.3B decision model on Qwen3.5-4B, trained readout head, self-hosted)2026-10-0667.554.6$0.02540.34 s
    Bespoke Nimble 9B v3 (Qwen3.5-9B LoRA)2026-10-0666.853.5$0.06260.34 s
    APUS-OpenJev-v1-9B (merged bf16 checkpoint-3000 on Qwen3.5-9B, their own native runtime, full depth, self-hosted)2026-10-0756.949.2$0.05110.71 s
    Plumb-4B (crh225, JevK5 v0.2 + LoRA)2026-10-0265.448.3$0.01700.19 s
    SPX-CD Pro (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)2026-10-0678.647.7$0.07901.13 s
    Messier One v0.2 (Qwen3.5-4B fine-tune, one prefill, self-hosted)2026-10-0758.145.3$0.01490.20 s
    spark-s1-4b-v6 (Open Spark Jev, abhishek085)2026-10-0554.044.9$0.02090.41 s
    swanOne (blockbrain, Qwen3.8-Flash-Next NVFP4)2026-10-0272.443.8$0.08480.71 s
    Quyet-1.0-Medium (Chinh Nguyen, Qwen3.5-4B decoder, option-letter logits)2026-10-0459.243.6$0.01660.30 s
    lev (Interfaze AI, Qwen3.5-4B + LoRA r32, label-token readout + candidate-path head, per-bucket temperatures)2026-10-0461.742.2$0.01880.35 s
    NInfer Qwen3.8-Flash-Next mixed2026-10-0269.742.0$0.08230.31 s
    jev-local (Qwen3.5-9B)2026-10-0257.541.3$0.02350.88 s
    decider-4b v2 (Mapika)2026-10-0264.941.2$0.01520.20 s
    ClassOne Qwen 3.5 9B (Qwen3.5-9B backbone with ClassOne decision heads, schema-constrained single forward pass, self-hosted)2026-10-0750.538.5$0.04880.26 s
    JevK5 v0.3 (4B)2026-10-0264.237.4$0.01700.19 s
    metask-jev-4b2026-10-0261.837.3$0.02560.26 s
    deck-4B v1.0 (krishna765, JevK5 + LoRA, FP8 weight-only)2026-10-0563.637.3$0.01650.30 s
    Clef-Flash (Cloudflare, Qwen3.5-9B post-train with a joint schema head, multimodal, measured on text)2026-10-0463.836.6$0.05930.36 s
    Decision-4B (Eval Engine / Chromia)2026-10-0459.135.4$0.01370.24 s
    janus 4B (Icarus AI / cmxu, Qwen3.5-4B + LoRA r64 and pointer decision head)2026-10-0462.935.2$0.01540.21 s
    Qwen3.5-9B Jev-like data-mix v22026-10-0260.933.6$0.06460.49 s
    JevK5 v0.2.02026-10-0262.832.1$0.01700.23 s
    JevOne2026-10-0268.931.1$0.10120.27 s
    Kahn1 4B (Okura66, Qwen3.5-4B LoRA merge, author's sysone engine)2026-10-0761.731.1$0.06130.28 s
    TypeCastLM 1.4.0 (Mikhail Gribov, Qwen3.5-4B computed head)2026-10-0756.429.8$0.01490.21 s
    Hopper2026-10-0262.428.9$0.01810.26 s
    Imajev-4B (RTX 5090)2026-10-0261.928.7$0.01680.29 s
    Malkuth-4B (newfull5, Kev post-train)2026-10-0261.126.2$0.02970.30 s
    jqv (Qwen3-32B zero-shot)2026-10-0261.925.4$0.04240.34 s
    JPT-35B-A3B (Qwen3.5-35B-A3B fine-tune)2026-10-0670.725.4$0.16310.29 s
    decider-35b-a3b (Mapika)2026-10-0266.224.8$0.15390.25 s
    CoCo-Decision-4B (corners-ai, LoRA on Qwen3.5-4B, option-label logits of one pass, served by oh-my-jev)2026-10-0756.924.6$0.01450.24 s
    Manchego v2.12026-10-0458.324.5$0.01550.26 s
    Decision 4B v1.2 (FlyMyJev, Qwen3.5-4B + LoRA)2026-10-0260.424.5$0.01700.19 s
    Raw Qwen3 4B Instruct 2507 direct logits2026-10-0234.621.8$0.01660.24 s
    JEV-27B (AutoTrust, Qwen3.8-27B)2026-10-0673.921.3$0.19910.36 s
    torchcast-decision-27b (Torchcast AI, Qwen3.8-27B fine-tune)2026-10-0681.621.3$0.20990.34 s
    Bespoke Nimble 9B (Bespoke Labs)2026-10-0256.521.1$0.12830.44 s
    reflex 4B (kshetrajna12)2026-10-0259.020.6$0.01692.87 s
    Perplexity Decider v1.1 27B (Qwen3.8-27B, noncausal decision readout)2026-10-0682.820.6$0.21730.39 s
    JADE (Qwen3.8-27B LoRA)2026-10-0649.220.2$0.16370.40 s
    typecastlm (Mikhail Gribov, Qwen3.5-4B computed head)2026-10-0256.519.7$0.01590.21 s
    Jebadiah 27B (Frontier Infra, Qwen3.8-27B LoRA)2026-10-0670.319.3$0.21650.43 s
    Decision 2.0 Vega 27B (vLLM-SR, Qwen3.8-27B)2026-10-0671.318.9$0.22160.53 s
    Decision 4B v1.1 (FlyMyJev, Qwen3.5-4B + LoRA)2026-10-0257.218.7$0.01700.19 s
    AutoJev-27B (denis-pplx, Qwen3.8-27B)2026-10-0271.618.5$0.22620.34 s
    Eikos-27B (caiovicentino1, Qwen3.8-27B)2026-10-0277.218.1$0.23820.38 s
    NInfer Qwen3.8-27B NVFP42026-10-0268.017.7$0.23050.26 s
    OpenJev (thinking, BF16)2026-10-0275.917.3$0.24151.86 s
    Open-Jev 9B (Zefan Cai)2026-10-0261.817.0$0.17041.39 s
    Standard One 8B (Standard Thinking)2026-10-0258.416.9$0.07770.21 s
    Clef (Cloudflare, Qwen3.8-27B post-train with a joint schema head, multimodal, measured on text)2026-10-0475.316.8$0.24920.58 s
    JEV Qwen3.5-9B Base NVFP42026-10-0257.116.0$0.05580.20 s
    SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)2026-10-0255.415.5$0.01700.25 s
    Bev / Bonsai 27B2026-10-0466.111.7$0.24682.12 s
    LitJev (Qwen3.8-27B)2026-10-0266.611.5$0.24443.10 s
    local-jev Qwen3.5-4B2026-10-0255.011.4$0.02280.40 s
    system-one (Qwen3-8B, Sean Goedecke)2026-10-0229.711.0$0.06830.25 s
    kev 8B (research preview)2026-10-0250.010.6$0.09680.30 s
    Decision 2B (FlyMy.AI, v59)2026-10-0554.810.3$0.01320.30 s
    OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)2026-10-0248.610.3$0.01070.74 s
    open-alternative-jev (Qwen3.5-4B, IkerMoel)2026-10-0249.19.4$0.01680.24 s
    Raw Qwen3 8B direct logits2026-10-0230.19.3$0.06540.31 s
    decider-2b (Mapika)2026-10-0244.67.7$0.01480.18 s
    Nemotron Diffusion 8B (pst2154, optimized vLLM)2026-10-0447.57.3$0.03050.19 s
    kev 4B (research preview)2026-10-0244.06.6$0.01380.29 s
    Kev 27B (Kev 1.0, Qwen3.8-27B)2026-10-0678.95.8$0.52880.51 s
    Malkuth-2B (newfull5, Kev post-train)2026-10-0247.64.0$0.01410.24 s
    Qwen3-Reranker-4B2026-10-0244.73.7$0.05250.56 s
    SimpleJev Qwen3.8-27B (self-hosted, v1.6 pool)2026-10-0674.93.4$0.67601.39 s
    ZeroEntropy zerank-22026-10-0249.43.3$0.05250.50 s
    EXAONE-4.0-1.2B-JEV v0.3 (unofficial fine-tune of LGAI-EXAONE/EXAONE-4.0-1.2B, native option-label softmax, two reads averaged, self-hosted)2026-10-0747.73.3$0.02530.19 s
    Open-Jev 2B (Zefan Cai)2026-10-0247.73.2$0.17041.02 s
    Open-Jev 27B v1.1 (Zefan Cai, LoRA + decision head on Qwen3.8-27B)2026-10-0468.02.9$0.71571.73 s
    Raw Phi-4 mini direct logits2026-10-0235.32.5$0.03630.23 s
    SimpleJev (Qwen3.5-0.8B, CPU)2026-10-0231.62.1$0.00988.80 s
    smalljev semantic-v92026-10-0243.00.8$0.02030.28 s
    kev 0.6B (research preview)2026-10-0236.90.5$0.00460.26 s
    jul fast (usejul/jul-decision-minicpm5-2b: MiniCPM5-2B fine-tune with a pointer head, option probabilities from one pass, self-hosted via `jul serve`)2026-10-0742.80.5$0.01080.24 s
    OpenDecision (ModernBERT-large zero-shot)2026-10-0238.00.5$0.00500.29 s
    ClassOne Gemma 4 E2B (Gemma-4-E2B backbone with ClassOne decision heads, schema-constrained single forward pass, self-hosted)2026-10-0715.20.5$0.00980.21 s
    Qwen3.5-0.8B Decision Model (Mourad Ghafiri)2026-10-0242.60.5$0.004811.25 s
    Fastino GLiNER-2.5-Decide (hosted API)2026-10-0534.60.3$0.06400.44 s
    Deem 0.8B v12026-10-0418.10.3$0.00450.46 s
    Decision Fast (FlyMy.AI, v53a)2026-10-0540.60.3$0.00460.27 s
    Quyet-1.0-Small-EN (Chinh Nguyen, ModernBERT-base encoder with fixed heads)2026-10-0439.50.3$0.00230.18 s
    Laya typed-decisions2026-10-0444.20.2$0.00381.88 s
    WaterSheep (Samrat Dutta, ModernBERT-base encoder with calibrated heads, served by the author's own `watersheep --serve`)2026-10-0735.60.2$0.01110.95 s
    openJev Verdict 1.42026-10-0240.40.2$0.00280.74 s
    Raw Qwen3 0.6B direct logits2026-10-028.20.2$0.00560.21 s
    lev-350m (Franck Verrot, LFM2.5-350M)2026-10-0543.30.2$0.00460.19 s
    Raw Qwen3 1.7B direct logits2026-10-028.30.1$0.01120.21 s
    kev 0.5B2026-10-0232.90.1$0.00460.24 s
    openJev Verdict (heman10x, ModernBERT-base 151M)2026-10-0226.60.1$0.00280.69 s
    Quyet-1.0-Small (Chinh Nguyen, SEA-LION-ModernBERT-300M encoder with fixed heads)2026-10-0432.50.1$0.00480.18 s
    watt-flash-0.1 (Zaitgeist Labs, 140M mmBERT-small encoder with an option-marker scorer, one forward pass per request, self-hosted)2026-10-0741.60.1$0.00231.01 s
    Quyet-1.0-Tiny (Chinh Nguyen, mmBERT-small encoder (16 layers kept) with fixed heads)2026-10-0431.10.1$0.00240.18 s
    open-jev-deberta-v3-large (local CPU)2026-10-0235.40.0$0.00563.53 s
    verdict-small (Manavarya09, multilingual-e5-small 118M)2026-10-0237.60.0$0.00090.25 s
    Tacet Sonata (CodePawl, 144M packed encoder on mmBERT-small, option-marker softmax in one pass, in-process via the author's tacet package)2026-10-0724.30.0$0.00230.61 s
    Laya (Convai Innovations, ModernBERT-large 421M)2026-10-0232.50.0$0.00321.80 s
    Mixedbread mxbai-rerank-base-v22026-10-0243.90.0$0.02100.21 s
    Laya multilingual2026-10-0417.00.0$0.00390.87 s
    Certo v1 (AltSlate Labs)2026-10-0243.20.0$0.00130.23 s
    CLM-8B (Contrastive-LM, clm-latest)2026-10-0220.10.0$0.04470.19 s
    BAAI bge-reranker-v2-m32026-10-0244.10.0$0.02270.18 s
    Alibaba GTE Reranker ModernBERT-base2026-10-0241.70.0$0.01060.20 s
    Mirror2026-10-0225.80.0$0.00231.45 s
    Open Jev JSON Canvas (JoshuaSP)2026-10-0230.60.0$0.04900.45 s
    APUS-OpenJev-v1-35B-A3B (merged bf16 checkpoint-5949 MoE on Qwen3.5-35B-A3B, their own native runtime, full depth, self-hosted)2026-10-0760.4——0.86 s
    BB-Qwen3.5-4B-LoRA (Babak Barazandeh, LoRA on Qwen3.5-4B-Base, hosted /v1/systemone)2026-10-0760.4——0.70 s
    Seb-9B (Qwen3.5-9B-shaped multimodal decision model, one forward pass, stock vLLM, self-hosted)2026-10-0766.8——0.27 s
    wity-1 (Wity, reasoning always)2026-10-0679.370.6$0.02361.69 s
    wity-1 (Wity, reasoning off)2026-10-0663.049.4$0.02360.28 s
    AutoJev-27B (RTX PRO 6000)measured on v1.5.0 (2026-09-28)79.719.5$0.22620.33 s
    Bosun v3.1 0.6Bmeasured on v1.5.3 (2026-09-29)39.42.5$0.00564.04 s
    classifier.dev (fast tier)Not yet measured————
    decision-machine-1 (milliseconds.ai)measured on v1.5.0 (2026-09-28)48.03.2$0.02860.18 s
    DeepSeek V4.1 Flash (thinking default)measured on v1.5.0 (2026-09-28)95.36.6$0.49761.78 s
    Gemini 3.1 Flash-Litemeasured on v1.5.0 (2026-09-28)76.119.6$0.21940.86 s
    GLiNER2 (Fastino, gliner2.5-base)measured on v1.5.0 (2026-09-28)24.52.3$0.00281.00 s
    GLiNER2 large (Fastino)measured on v1.5.0 (2026-09-28)32.48.4$0.00561.88 s
    GLiNER2.5 multi (Fastino, 287M)measured on v1.5.0 (2026-09-28)36.02.7$0.00281.37 s
    GLiNER2.5 small (Fastino, 74M)measured on v1.5.0 (2026-09-28)32.50.9$0.00280.47 s
    GPT-5.6 Luna (low reasoning effort)measured on v1.5.0 (2026-09-28)94.522.4$0.20471.32 s
    GPT-6 Luna (default medium reasoning effort)measured on v1.5.0 (2026-09-28)95.938.8$0.11381.56 s
    GPT-6 Luna (low reasoning effort)measured on v1.5.0 (2026-09-28)95.140.5$0.10751.58 s
    Instinct (ZooWork, Qwen3.8-27B)measured on v1.5.0 (2026-09-28)73.918.3$0.22980.52 s
    Instinct Dual 4Bmeasured on v1.5.3 (2026-09-29)65.747.0$0.02120.51 s
    jeff (Logan Markewich, GLiFormer 400M)measured on v1.5.0 (2026-09-28)42.40.1$0.00437.14 s
    JevAct (einptein, jev1-2b-v2)measured on v1.5.0 (2026-09-28)36.91.5$0.01120.80 s
    Jobe Qwen3.5-4B (frozen)Not yet measured————
    Autoloops – Gemma 4 31B ITmeasured on v1.5.0 (2026-09-28)81.340.5$0.10320.61 s
    mica-v01-4bNot yet measured————
    Needle 3 (Cactus, 2-bit, local CPU)measured on v1.5.0 (2026-09-28)0.00.0$0.0191135.51 s
    Needle 3, options as tools (post-hoc adapter mode)measured on v1.5.0 (2026-09-28)0.00.0$0.019158.16 s
    NInfer Qwen3.8-27B NVFP4 (T=1.5)measured on v1.5.0 (2026-09-28)73.818.5$0.23050.25 s
    OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)Not yet measured————
    openjev-sglang (Qwen3.6-35B-A3B on SGLang)measured on v1.5.0 (2026-09-28)70.829.0$0.13981.20 s
    Qwen3.8 27B (Chutes TEE)measured on v1.5.0 (2026-09-28)96.80.0$2.17846.49 s
    reflex-27b (Qwen3.8-27B)measured on v1.5.0 (2026-09-28)74.413.2$0.29732.75 s
    SimpleJev Qwen3.6-35B-A3BNot yet measured————
    SimpleJev Qwen3.8-27Bmeasured on v1.5.0 (2026-09-28)80.03.4$0.68681.68 s
    Surogate Rune 26B-A4B v3 (RTX PRO 6000)measured on v1.5.0 (2026-09-28)79.066.5$0.05020.35 s
    system-one-open (Gemma 4 E2B LoRA on an L4)measured on v1.5.0 (2026-09-28)56.642.4$0.01141.15 s
    Vansa-3.4 (Vansa, hosted System One API)measured on v1.5.6 (2026-10-03)72.871.6$0.01930.14 s
    Von (wfzyx, Option-Marker 395M)measured on v1.5.0 (2026-09-28)41.70.0$0.00380.92 s

    Historical results SHA-256 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1. Previous board: v1.6.1.

    Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

    Open this section to load the earlier public-only board and diagnostics.