Jev vs decisio v0.8.0 on gemma-4-31B-it: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

Rank #2 on the open-weights board in v1.6.1.

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and decisio v0.8.0 on gemma-4-31B-it: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: decisio v0.8.0 on gemma-4-31B-it — Open weights · LLM decoder · Score 71.7 (#2)Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.
  • B: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (ranked, not ranked)

The four score axes

Radar: the four score axes, two systemsThe four score axes, decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Intelligence: 70.4 vs 63.6; Calibration: 88.7 vs 90.6; Speed: 91.2 vs 91.5; Cost: 51.7 vs 54.7.50100Intelligence70.4 · 63.6Calibration88.7 · 90.6Speed91.2 · 91.5Cost51.7 · 54.7
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

decisio v0.8.0 on gemma-4-31B-it: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: capability by subject topic, two systemsCapability by subject topic, decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Rules, policy & law: 79.0 vs 49.4; Coding & software: 81.7 vs 64.3; Math & numbers: 43.5 vs 10.4; Finance & commerce: 82.7 vs 40.9; Support & operations: 42.8 vs 48.8; Everyday language: 64.1 vs 77.5; Safety & security: 75.5 vs 52.2.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open, 1465 sealed).Rules, policy& law79.0 · 49.4Coding & software: code, SQL, repositories, developer tools and IT systems. 382 items (62 open, 320 sealed).Coding &software81.7 · 64.3Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open, 333 sealed).Math &numbers43.5 · 10.4Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open, 211 sealed).Finance &commerce82.7 · 40.9Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open, 139 sealed).Support &operations42.8 · 48.8Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open, 134 sealed).Everydaylanguage64.1 · 77.5Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open, 103 sealed).Safety &security75.5 · 52.2
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), decisio v0.8.0 on gemma-4-31B-it: S+P; Jev 1.13.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open / 1465 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 382 items (62 open / 320 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open / 333 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open / 211 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open / 139 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open / 134 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open / 103 sealed)

Use cases (TypeSafe categories)

decisio v0.8.0 on gemma-4-31B-it: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Model routing: 84.4 vs 73.0; Legal & compliance: 81.3 vs 68.6; Customer support: 61.5 vs 49.6; Other: 37.0 vs 40.1; E-commerce: 64.7 vs 40.5; Insurance claims: 76.9 vs 46.1; Risk assessment: 67.8 vs 32.8; Financial crime: 50.1 vs 35.8; Feature extraction: n=19 vs 28.3; Lead generation: n=17 vs 49.2; Recruiting: — vs 0.0; Knowledge graphs: n=18 vs 28.8; LLM guardrails: n=17 vs 68.0; Moderation: 62.0 vs 51.6; Code linting: n=16 vs 29.1; Search & retrieval: n=17 vs 80.9; Science: — vs 21.2; Advertising: — vs 16.4; Gaming: — vs 5.1; Demand forecasting: n=16 vs 0.0.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open, 438 sealed).Model routing84.4 · 73.0Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open, 441 sealed).Legal &compliance81.3 · 68.6Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open, 203 sealed).Customersupport61.5 · 49.6Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open, 150 sealed).Other37.0 · 40.1E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open, 110 sealed).E-commerce64.7 · 40.5Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open, 110 sealed).Insuranceclaims76.9 · 46.1Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open, 110 sealed).Riskassessment67.8 · 32.8Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open, 92 sealed).Financialcrime50.1 · 35.8Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open, 94 sealed).Featureextractionn=19 · 28.3Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open, 91 sealed).Leadgenerationn=17 · 49.2Recruiting: resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open, 93 sealed).Recruitingn/a · 0.0Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open, 91 sealed).Knowledgegraphsn=18 · 28.8LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open, 88 sealed).LLMguardrailsn=17 · 68.0Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open, 86 sealed).Moderation62.0 · 51.6Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open, 87 sealed).Code lintingn=16 · 29.1Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open, 85 sealed).Search &retrievaln=17 · 80.9Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open, 88 sealed).Sciencen/a · 21.2Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open, 85 sealed).Advertisingn/a · 16.4Gaming: player reports, in-game chat, game support. 86 items (3 open, 83 sealed).Gamingn/a · 5.1Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open, 80 sealed).Demandforecastingn=16 · 0.0
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), decisio v0.8.0 on gemma-4-31B-it: S+P; Jev 1.13.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.A (decisio v0.8.0 on gemma-4-31B-it) is measured on 9 of 20 spokes only, so it is drawn as points, not a shape — measured on fewer items; full radar after the next re-measure.Gaps, not zeros: Open spokes (n/a or n=…) have fewer than 30 answered items for that system or no published value.

Low sample, n < 30 — indicative only

Category (items)A: decisio v0.8.0 on gemma-4-31B-itB: Jev 1.13.0
Feature extraction for predictive modeling (99; a system answered fewer than 30)72.0 n=1928.3 n=99
Lead generation (95; a system answered fewer than 30)98.6 n=1749.2 n=95
Graphs and knowledge graphs (94; a system answered fewer than 30)78.6 n=1828.8 n=94
LLM guardrails (93; a system answered fewer than 30)76.9 n=1768.0 n=93
Semantic code linting (90; a system answered fewer than 30)45.6 n=1629.1 n=90
Search and retrieval (90; a system answered fewer than 30)93.9 n=1780.9 n=90
Demand forecasting (83; a system answered fewer than 30)0.0 n=160.0 n=83

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open / 438 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open / 441 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open / 203 sealed)
  • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open / 150 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open / 110 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open / 110 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open / 110 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open / 92 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open / 94 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open / 91 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open / 93 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open / 91 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open / 88 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open / 86 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open / 87 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open / 85 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open / 88 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open / 85 sealed)
  • Gaming — player reports, in-game chat, game support. 86 items (3 open / 83 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open / 80 sealed)

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Choice · open: 70.8 vs 79.0; Choice · sealed: 83.1 vs 75.8; Noul · open: 73.1 vs 54.5; Noul · sealed: 76.0 vs 46.1; Score · open: 58.8 vs 63.3; Score · sealed: 60.6 vs 63.0.50100Choice · open70.8 · 79.0Choice ·sealed83.1 · 75.8Noul · open73.1 · 54.5Noul · sealed76.0 · 46.1Score · open58.8 · 63.3Score ·sealed60.6 · 63.0
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Easy: 93.0 vs 93.9; Standard: 65.2 vs 63.2; Judge: 61.2 vs 67.7; Hard: 69.0 vs 65.6.50100Easy93.0 · 93.9Standard65.2 · 63.2Judge61.2 · 67.7Hard69.0 · 65.6
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, decisio v0.8.0 on gemma-4-31B-it vs Jev 1.13.0. Easy: 85.9 vs 78.4; Standard: 81.2 vs 68.7; Judge: 80.0 vs 66.9; Hard: 67.7 vs 58.9.50100Easy85.9 · 78.4Standard81.2 · 68.7Judge80.0 · 66.9Hard67.7 · 58.9
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

O1S chance-corrected competence: equal mean across the request types present, clipped to 0–100. Zero means the equal-weight mean across present request types was at or below baseline; negative means are clipped. It does not mean every type scored zero. Choice uses a random-option baseline; Noul uses a fixed 50% accuracy baseline (abstentions count as wrong); Score uses the midpoint-guess error (random-guess fallback when every gold is at the midpoint). Compare models within a category; row pools and supported types can differ.

Categories use ruled subject-topic labels and authoring use-case precedence across the answered pool union, with stable item identities counted once.

  • Official O1S; recorded failures count where supported; no missing response is imputed; stable identity dedupe parent first.
  • S-based category cells pool recorded supported responses from S+P and the answered L1/L2/L3 supplements. API overlay category cells keep their separately recorded A4/A5+P union. Headline and per-type/tier scores are unchanged.
All values as a table
SpokeA: decisio v0.8.0 on gemma-4-31B-itB: Jev 1.13.0
The four score axes
Intelligence70.463.6
Calibration88.790.6
Speed91.291.5
Cost51.754.7
Capability by subject topic
Rules, policy & law79.049.4
Coding & software81.764.3
Math & numbers43.510.4
Finance & commerce82.740.9
Support & operations42.848.8
Everyday language64.177.5
Safety & security75.552.2
Use cases (TypeSafe categories)
Model routing84.473.0
Legal & compliance81.368.6
Customer support61.549.6
Other37.040.1
E-commerce64.740.5
Insurance claims76.946.1
Risk assessment67.832.8
Financial crime50.135.8
Feature extractionn=1928.3
Lead generationn=1749.2
Recruiting—0.0
Knowledge graphsn=1828.8
LLM guardrailsn=1768.0
Moderation62.051.6
Code lintingn=1629.1
Search & retrievaln=1780.9
Science—21.2
Advertising—16.4
Gaming—5.1
Demand forecastingn=160.0
Competence per request type, open / sealed
Choice · open70.879.0
Choice · sealed83.175.8
Noul · open73.154.5
Noul · sealed76.046.1
Score · open58.863.3
Score · sealed60.663.0
Competence per tier — open set
Easy93.093.9
Standard65.263.2
Judge61.267.7
Hard69.065.6
Competence per tier — sealed set
Easy85.978.4
Standard81.268.7
Judge80.066.9
Hard67.758.9

Published values

MeasureJev 1.13.0 (TypeSafe AI)decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted)
Capability Score77.179.6
Rank (open-weights board)Reference, not ranked#2
Composite (secondary)71.571.7
intelligence63.670.4
calibration90.688.7
speed91.591.2
cost54.751.7
USD per 1,000 decisions$0.032 per 1,000 decisions$0.041 per 1,000 decisions
p50 latency (seconds)0.20.2
Last measured2026-10-052026-10-06

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted)

Apache-2.0 software; Gemma model terms

Published source
Measured on
measured 2026-10-06 on the live v1.6.0/v1.6.1 pool. github.com/aminry/decisio tag v0.8.0 (5b42101a191d222062a795194b0aedb5003bf6ea) - the same tag whose other two entries run 54 measured - installed with their own `uv sync --extra serve --frozen` so their lockfile fixes every version, and served with their own `python -m decisio.serve.vllm_engine --base gemma-4-31b`. --base carries google/gemma-4-31B-it@842da379, the official weights, FROZEN: no fine-tuning, no adapter, no task registered; vLLM 0.30.0 quantizes it to FP8 when it loads (30.6 GiB). That profile's measured settings: temperature 4.672 for choice questions and 5.252 for the rest, the 12B entry's prompt unchanged, a system turn, the chat template's own answer position, every single-token form of each letter summed, no tokens generated. New in 0.8.0 on this base and the reason it is worth its own row: several questions of a state are scored in ONE pass after a single prefill. The author's own limit is a 32,768-token prompt, so the pool's 23 items at ~80,000 tokens are a documented capacity limit, not a defect. THE AUTHOR STATES ONE 96 GB CARD; this is one H100 80 GB (Lium) and no flag of theirs was changed to make it fit. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 55, GitHub #169
Cost basis (estimated)
ESTIMATE (base-model reference; self-hosted open weights, no public tariff) and CLEAN: unlike run 54's FP8 sibling, the frozen pricing snapshot resolves this EXACT base - google/gemma-4-31B-it - to USD 0.09 per 1M input and USD 0.34 per 1M output. Read before the pod and asserted in make_registry_r55.py. x this row's own measured tokens, the deck31b method. Nothing is provisional and no release-lane line is open on it. Estimated, not charged
Latency adjustment
x2 + 0.15 s (assumption, not measured)

Frequently asked questions

What does JevBench show for Jev and decisio v0.8.0 on gemma-4-31B-it?
Jev 1.13.0 has Capability 77.1 and is the unranked reference (the open-weights board ranks 92 open-weights systems; hosted API offerings are ranked on the API leaderboard). decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted) has Capability 79.6. Rank #2 on the open-weights board in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted): estimated, $0.041 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is decisio v0.8.0 on gemma-4-31B-it open source?
The published licence note for decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted) is Apache-2.0 software; Gemma model terms. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.