Jev vs Surogate Rune 26B-A4B v3: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

Rank #3 on the open-weights board in v1.6.1.

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and Surogate Rune 26B-A4B v3: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

The four score axes

Radar: the four score axes, two systemsThe four score axes, Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Intelligence: 56.1 vs 63.6; Calibration: 90.9 vs 90.6; Speed: 87.0 vs 91.5; Cost: 49.6 vs 54.7.50100Intelligence56.1 · 63.6Calibration90.9 · 90.6Speed87.0 · 91.5Cost49.6 · 54.7
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Rules, policy & law: 60.3 vs 63.3; Coding & software: 71.3 vs 70.5; Math & numbers: 13.9 vs 22.6; Support & operations: 38.4 vs 44.4; Finance & commerce: 66.5 vs 50.2; Everyday language: 37.0 vs 57.4; Safety & security: 60.0 vs 32.4.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1011 items (151 open, 860 sealed).Rules, policy& law60.3 · 63.3Coding & software: code, SQL, repositories, developer tools and IT systems. 330 items (62 open, 268 sealed).Coding &software71.3 · 70.5Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 234 items (37 open, 197 sealed).Math &numbers13.9 · 22.6Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 113 items (21 open, 92 sealed).Support &operations38.4 · 44.4Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 99 items (12 open, 87 sealed).Finance &commerce66.5 · 50.2Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 52 items (11 open, 41 sealed).Everydaylanguage37.0 · 57.4Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 48 items (6 open, 42 sealed).Safety &security60.0 · 32.4
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1011 items (151 open / 860 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 330 items (62 open / 268 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 234 items (37 open / 197 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 113 items (21 open / 92 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 99 items (12 open / 87 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 52 items (11 open / 41 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 48 items (6 open / 42 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Model routing: 70.7 vs 77.4; Legal & compliance: 65.2 vs 75.3; Customer support: 39.5 vs 53.0; E-commerce: 48.2 vs 40.3; Insurance claims: 41.8 vs 58.0; Risk assessment: 40.1 vs 34.4; Financial crime: 33.7 vs 36.7; Moderation: 61.7 vs 54.7; Code linting: n=16 vs 15.4; Lead generation: n=17 vs 56.5; Feature extraction: n=19 vs 19.9; LLM guardrails: n=17 vs 68.6; Search & retrieval: n=17 vs 76.2; Advertising: — vs 28.2; Demand forecasting: n=16 vs 0.0; Recruiting: — vs 3.1; Knowledge graphs: n=18 vs 22.4; Gaming: — vs 14.5; Science: — vs 25.0.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 471 items (86 open, 385 sealed).Model routing70.7 · 77.4Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 466 items (74 open, 392 sealed).Legal &compliance65.2 · 75.3Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 179 items (37 open, 142 sealed).Customersupport39.5 · 53.0E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 65 items (9 open, 56 sealed).E-commerce48.2 · 40.3Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 63 items (9 open, 54 sealed).Insuranceclaims41.8 · 58.0Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 62 items (8 open, 54 sealed).Riskassessment40.1 · 34.4Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 51 items (9 open, 42 sealed).Financialcrime33.7 · 36.7Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 47 items (7 open, 40 sealed).Moderation61.7 · 54.7Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 33 items (3 open, 30 sealed).Code lintingn=16 · 15.4Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 31 items (4 open, 27 sealed).Leadgenerationn=17 · 56.5Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 31 items (5 open, 26 sealed).Featureextractionn=19 · 19.9LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 30 items (5 open, 25 sealed).LLMguardrailsn=17 · 68.6Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 30 items (5 open, 25 sealed).Search &retrievaln=17 · 76.2Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open, 27 sealed).Advertisingn/a · 28.2Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 30 items (3 open, 27 sealed).Demandforecastingn=16 · 0.0Recruiting: resumes, applications, interview feedback, matching candidates to roles. 30 items (1 open, 29 sealed).Recruitingn/a · 3.1Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 30 items (3 open, 27 sealed).Knowledgegraphsn=18 · 22.4Gaming: player reports, in-game chat, game support. 30 items (3 open, 27 sealed).Gamingn/a · 14.5Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 30 items (2 open, 28 sealed).Sciencen/a · 25.0
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.A (Surogate Rune 26B-A4B v3) is measured on 8 of 19 spokes only, so it is drawn as points, not a shape — measured on fewer items; full radar after the next re-measure.Gaps, not zeros: Open spokes (n/a or n=…) have fewer than 30 answered items for that system or no published value.

Low sample, n < 30 — indicative only

Category (items)A: Surogate Rune 26B-A4B v3B: Jev 1.13.0
Semantic code linting (33; a system answered fewer than 30)38.2 n=1615.4 n=33
Lead generation (31; a system answered fewer than 30)75.6 n=1756.5 n=31
Feature extraction for predictive modeling (31; a system answered fewer than 30)32.4 n=1919.9 n=31
LLM guardrails (30; a system answered fewer than 30)72.2 n=1768.6 n=30
Search and retrieval (30; a system answered fewer than 30)90.0 n=1776.2 n=30
Demand forecasting (30; a system answered fewer than 30)0.0 n=160.0 n=30
Graphs and knowledge graphs (30; a system answered fewer than 30)42.4 n=1822.4 n=30

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 471 items (86 open / 385 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 466 items (74 open / 392 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 179 items (37 open / 142 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 65 items (9 open / 56 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 63 items (9 open / 54 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 62 items (8 open / 54 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 51 items (9 open / 42 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 47 items (7 open / 40 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 33 items (3 open / 30 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 31 items (4 open / 27 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 31 items (5 open / 26 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 30 items (5 open / 25 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 30 items (5 open / 25 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open / 27 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 30 items (3 open / 27 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 30 items (1 open / 29 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 30 items (3 open / 27 sealed)
  • Gaming — player reports, in-game chat, game support. 30 items (3 open / 27 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 30 items (2 open / 28 sealed)

Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 148 items (24 open / 124 sealed) — not a use case of its own, so it is counted but not drawn.

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Choice · open: 76.0 vs 79.0; Choice · sealed: 73.1 vs 75.8; Noul · open: 42.6 vs 54.5; Noul · sealed: 34.3 vs 46.1; Score · open: 56.0 vs 63.3; Score · sealed: 54.4 vs 63.0.50100Choice · open76.0 · 79.0Choice ·sealed73.1 · 75.8Noul · open42.6 · 54.5Noul · sealed34.3 · 46.1Score · open56.0 · 63.3Score ·sealed54.4 · 63.0
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Easy: 90.4 vs 93.9; Standard: 53.4 vs 63.2; Judge: 58.2 vs 67.7; Hard: 61.0 vs 65.6.50100Easy90.4 · 93.9Standard53.4 · 63.2Judge58.2 · 67.7Hard61.0 · 65.6
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Surogate Rune 26B-A4B v3 vs Jev 1.13.0. Easy: 77.0 vs 78.4; Standard: 57.2 vs 68.7; Judge: 62.1 vs 66.9; Hard: 52.4 vs 58.9.50100Easy77.0 · 78.4Standard57.2 · 68.7Judge62.1 · 66.9Hard52.4 · 58.9
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).

Family and language are authoring metadata of every item. Each of the 1,500 v1.6 items and each of the 387 supplement items (L1 language supplement 354, L2 use-case supplement 33) was labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod, with the same model, questions, taxonomy and item-group rules (labelling rounds r13 and r14); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) of the main pool were checked by hand (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems and full-set hosted APIs answered S u P (1,500 items) plus the supplements where their row is tagged so; API rows re-run on A4/A5 have no cells yet.

  • Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
  • Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items of S u P and the L2 supplement.
All values as a table
SpokeA: Surogate Rune 26B-A4B v3B: Jev 1.13.0
The four score axes
Intelligence56.163.6
Calibration90.990.6
Speed87.091.5
Cost49.654.7
Capability by subject topic
Rules, policy & law60.363.3
Coding & software71.370.5
Math & numbers13.922.6
Support & operations38.444.4
Finance & commerce66.550.2
Everyday language37.057.4
Safety & security60.032.4
Use cases (TypeSafe categories)
Model routing70.777.4
Legal & compliance65.275.3
Customer support39.553.0
E-commerce48.240.3
Insurance claims41.858.0
Risk assessment40.134.4
Financial crime33.736.7
Moderation61.754.7
Code lintingn=1615.4
Lead generationn=1756.5
Feature extractionn=1919.9
LLM guardrailsn=1768.6
Search & retrievaln=1776.2
Advertising—28.2
Demand forecastingn=160.0
Recruiting—3.1
Knowledge graphsn=1822.4
Gaming—14.5
Science—25.0
Competence per request type, open / sealed
Choice · open76.079.0
Choice · sealed73.175.8
Noul · open42.654.5
Noul · sealed34.346.1
Score · open56.063.3
Score · sealed54.463.0
Competence per tier — open set
Easy90.493.9
Standard53.463.2
Judge58.267.7
Hard61.065.6
Competence per tier — sealed set
Easy77.078.4
Standard57.268.7
Judge62.166.9
Hard52.458.9

Published values

MeasureJev 1.13.0 (TypeSafe AI)Surogate Rune 26B-A4B v3 (v1.6 pool)
Capability Score77.173.5
Rank (open-weights board)Reference, not ranked#3
Composite (secondary)71.565.0
intelligence63.656.1
calibration90.690.9
speed91.587.0
cost54.749.6
USD per 1,000 decisions$0.032 per 1,000 decisions$0.048 per 1,000 decisions
p50 latency (seconds)0.20.4
Last measured2026-10-052026-10-06

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

Surogate Rune 26B-A4B v3 (v1.6 pool)

apache-2.0

Published source
Measured on
evaluator-owned Lium GPU pod, 1x NVIDIA RTX PRO 6000 Blackwell 96GB, a thin transport shim around the author's engine in its own pinned environment, offline (Hub offline mode, all credentials unset), driven over loopback only by the JevBench client with its loopback guard; label-free inputs; pod destroyed afterwards; HF Jev Decision Index intake, 6-7 Oct 2026
Cost basis (estimated)
ESTIMATE (base-model reference; self-hosted open weights, no public tariff): OpenRouter google/gemma-4-26b-a4b-it in the frozen 25 Sep snapshot, the M2 reference of the published surogate-rune-26b-a4b-v3 row (the frozen module has no explicit listing-decision line for this base: release-lane line), USD 0.09 per 1M input and USD 0.3 per 1M output x this row's own measured tokens. Estimated, not charged
Latency adjustment
x2 + 0.15 s (assumption, not measured)

Frequently asked questions

What does JevBench show for Jev and Surogate Rune 26B-A4B v3?
Jev 1.13.0 has Capability 77.1 and is the unranked reference (the open-weights board ranks 66 open-weights systems; hosted API offerings are ranked on the API leaderboard). Surogate Rune 26B-A4B v3 (v1.6 pool) has Capability 73.5. Rank #3 on the open-weights board in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. Surogate Rune 26B-A4B v3 (v1.6 pool): estimated, $0.048 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is Surogate Rune 26B-A4B v3 open source?
The published licence note for Surogate Rune 26B-A4B v3 (v1.6 pool) is apache-2.0. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.