Jev vs René-1 31B FP8: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

Rank #4 on the open-weights board in v1.6.1.

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and René-1 31B FP8: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: René-1 31B FP8 — Open weights · LLM decoder · Score 55.8 (#4)Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.
  • B: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (ranked, not ranked)

The four score axes

Radar: the four score axes, two systemsThe four score axes, René-1 31B FP8 vs Jev 1.13.0. Intelligence: 61.7 vs 63.6; Calibration: 90.2 vs 90.6; Speed: 87.1 vs 91.5; Cost: 46.0 vs 54.7.50100Intelligence61.7 · 63.6Calibration90.2 · 90.6Speed87.1 · 91.5Cost46.0 · 54.7
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

René-1 31B FP8: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: capability by subject topic, two systemsCapability by subject topic, René-1 31B FP8 vs Jev 1.13.0. Rules, policy & law: 71.2 vs 49.4; Coding & software: 69.0 vs 64.3; Math & numbers: 25.5 vs 10.4; Finance & commerce: 65.8 vs 40.9; Support & operations: 35.9 vs 48.8; Everyday language: 36.0 vs 77.5; Safety & security: 24.7 vs 52.2.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open, 1465 sealed).Rules, policy& law71.2 · 49.4Coding & software: code, SQL, repositories, developer tools and IT systems. 382 items (62 open, 320 sealed).Coding &software69.0 · 64.3Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open, 333 sealed).Math &numbers25.5 · 10.4Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open, 211 sealed).Finance &commerce65.8 · 40.9Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open, 139 sealed).Support &operations35.9 · 48.8Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open, 134 sealed).Everydaylanguage36.0 · 77.5Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open, 103 sealed).Safety &security24.7 · 52.2
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), René-1 31B FP8: S+P; Jev 1.13.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open / 1465 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 382 items (62 open / 320 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open / 333 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open / 211 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open / 139 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open / 134 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open / 103 sealed)

Use cases (TypeSafe categories)

René-1 31B FP8: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), René-1 31B FP8 vs Jev 1.13.0. Model routing: 72.1 vs 73.0; Legal & compliance: 76.8 vs 68.6; Customer support: 49.1 vs 49.6; Other: 8.9 vs 40.1; E-commerce: 53.8 vs 40.5; Insurance claims: 73.0 vs 46.1; Risk assessment: 39.8 vs 32.8; Financial crime: 33.3 vs 35.8; Feature extraction: n=19 vs 28.3; Lead generation: n=17 vs 49.2; Recruiting: — vs 0.0; Knowledge graphs: n=18 vs 28.8; LLM guardrails: n=17 vs 68.0; Moderation: 69.6 vs 51.6; Code linting: n=16 vs 29.1; Search & retrieval: n=17 vs 80.9; Science: — vs 21.2; Advertising: — vs 16.4; Gaming: — vs 5.1; Demand forecasting: n=16 vs 0.0.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open, 438 sealed).Model routing72.1 · 73.0Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open, 441 sealed).Legal &compliance76.8 · 68.6Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open, 203 sealed).Customersupport49.1 · 49.6Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open, 150 sealed).Other8.9 · 40.1E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open, 110 sealed).E-commerce53.8 · 40.5Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open, 110 sealed).Insuranceclaims73.0 · 46.1Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open, 110 sealed).Riskassessment39.8 · 32.8Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open, 92 sealed).Financialcrime33.3 · 35.8Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open, 94 sealed).Featureextractionn=19 · 28.3Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open, 91 sealed).Leadgenerationn=17 · 49.2Recruiting: resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open, 93 sealed).Recruitingn/a · 0.0Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open, 91 sealed).Knowledgegraphsn=18 · 28.8LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open, 88 sealed).LLMguardrailsn=17 · 68.0Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open, 86 sealed).Moderation69.6 · 51.6Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open, 87 sealed).Code lintingn=16 · 29.1Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open, 85 sealed).Search &retrievaln=17 · 80.9Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open, 88 sealed).Sciencen/a · 21.2Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open, 85 sealed).Advertisingn/a · 16.4Gaming: player reports, in-game chat, game support. 86 items (3 open, 83 sealed).Gamingn/a · 5.1Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open, 80 sealed).Demandforecastingn=16 · 0.0
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), René-1 31B FP8: S+P; Jev 1.13.0: S+P+L1+L2+L3 items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.A (René-1 31B FP8) is measured on 9 of 20 spokes only, so it is drawn as points, not a shape — measured on fewer items; full radar after the next re-measure.Gaps, not zeros: Open spokes (n/a or n=…) have fewer than 30 answered items for that system or no published value.

Low sample, n < 30 — indicative only

Category (items)A: René-1 31B FP8B: Jev 1.13.0
Feature extraction for predictive modeling (99; a system answered fewer than 30)19.5 n=1928.3 n=99
Lead generation (95; a system answered fewer than 30)93.8 n=1749.2 n=95
Graphs and knowledge graphs (94; a system answered fewer than 30)69.6 n=1828.8 n=94
LLM guardrails (93; a system answered fewer than 30)55.3 n=1768.0 n=93
Semantic code linting (90; a system answered fewer than 30)37.1 n=1629.1 n=90
Search and retrieval (90; a system answered fewer than 30)77.0 n=1780.9 n=90
Demand forecasting (83; a system answered fewer than 30)15.6 n=160.0 n=83

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open / 438 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open / 441 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open / 203 sealed)
  • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open / 150 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open / 110 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open / 110 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open / 110 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open / 92 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open / 94 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open / 91 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open / 93 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open / 91 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open / 88 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open / 86 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open / 87 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open / 85 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open / 88 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open / 85 sealed)
  • Gaming — player reports, in-game chat, game support. 86 items (3 open / 83 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open / 80 sealed)

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, René-1 31B FP8 vs Jev 1.13.0. Choice · open: 78.7 vs 79.0; Choice · sealed: 78.9 vs 75.8; Noul · open: 53.3 vs 54.5; Noul · sealed: 47.4 vs 46.1; Score · open: 56.8 vs 63.3; Score · sealed: 55.3 vs 63.0.50100Choice · open78.7 · 79.0Choice ·sealed78.9 · 75.8Noul · open53.3 · 54.5Noul · sealed47.4 · 46.1Score · open56.8 · 63.3Score ·sealed55.3 · 63.0
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, René-1 31B FP8 vs Jev 1.13.0. Easy: 90.2 vs 93.9; Standard: 48.2 vs 63.2; Judge: 63.4 vs 67.7; Hard: 69.9 vs 65.6.50100Easy90.2 · 93.9Standard48.2 · 63.2Judge63.4 · 67.7Hard69.9 · 65.6
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, René-1 31B FP8 vs Jev 1.13.0. Easy: 77.5 vs 78.4; Standard: 64.0 vs 68.7; Judge: 71.7 vs 66.9; Hard: 57.7 vs 58.9.50100Easy77.5 · 78.4Standard64.0 · 68.7Judge71.7 · 66.9Hard57.7 · 58.9
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

O1S chance-corrected competence: equal mean across the request types present, clipped to 0–100. Zero means the equal-weight mean across present request types was at or below baseline; negative means are clipped. It does not mean every type scored zero. Choice uses a random-option baseline; Noul uses a fixed 50% accuracy baseline (abstentions count as wrong); Score uses the midpoint-guess error (random-guess fallback when every gold is at the midpoint). Compare models within a category; row pools and supported types can differ.

Categories use ruled subject-topic labels and authoring use-case precedence across the answered pool union, with stable item identities counted once.

  • Official O1S; recorded failures count where supported; no missing response is imputed; stable identity dedupe parent first.
  • S-based category cells pool recorded supported responses from S+P and the answered L1/L2/L3 supplements. API overlay category cells keep their separately recorded A4/A5+P union. Headline and per-type/tier scores are unchanged.
All values as a table
SpokeA: René-1 31B FP8B: Jev 1.13.0
The four score axes
Intelligence61.763.6
Calibration90.290.6
Speed87.191.5
Cost46.054.7
Capability by subject topic
Rules, policy & law71.249.4
Coding & software69.064.3
Math & numbers25.510.4
Finance & commerce65.840.9
Support & operations35.948.8
Everyday language36.077.5
Safety & security24.752.2
Use cases (TypeSafe categories)
Model routing72.173.0
Legal & compliance76.868.6
Customer support49.149.6
Other8.940.1
E-commerce53.840.5
Insurance claims73.046.1
Risk assessment39.832.8
Financial crime33.335.8
Feature extractionn=1928.3
Lead generationn=1749.2
Recruiting—0.0
Knowledge graphsn=1828.8
LLM guardrailsn=1768.0
Moderation69.651.6
Code lintingn=1629.1
Search & retrievaln=1780.9
Science—21.2
Advertising—16.4
Gaming—5.1
Demand forecastingn=160.0
Competence per request type, open / sealed
Choice · open78.779.0
Choice · sealed78.975.8
Noul · open53.354.5
Noul · sealed47.446.1
Score · open56.863.3
Score · sealed55.363.0
Competence per tier — open set
Easy90.293.9
Standard48.263.2
Judge63.467.7
Hard69.965.6
Competence per tier — sealed set
Easy77.578.4
Standard64.068.7
Judge71.766.9
Hard57.758.9

Published values

MeasureJev 1.13.0 (TypeSafe AI)René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout)
Capability Score77.176.0
Rank (open-weights board)Reference, not ranked#4
Composite (secondary)71.555.8
intelligence63.661.7
calibration90.690.2
speed91.587.1
cost54.746.0
USD per 1,000 decisions$0.032 per 1,000 decisions$0.063 per 1,000 decisions
p50 latency (seconds)0.20.4
Last measured2026-10-052026-10-07

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout)

apache-2.0

Published source
Measured on
Candidate re-measured on 2026-10-07, current v1.6.0/v1.6.1 1500-item pool (sha256901983ae...). Immutable salfatigroup/rene-1-31b-fp8@bd634489957f8da63ccce858ee33fd6c2929d286, Apache-2.0. Same unmodified reviewed contributor source and run-39 transport shim. Author noul/yesno type mapping only; two-decimal probability rounding retained. One evaluator-owned RTX PRO6000 Blackwell 96GB pod, loopback-only network namespace, offline credential-free inference; no external hosted API, no retry or batching. Unmodified run_v16/typesafe adapter, serial one request at a time, selfhosted x2+0.15s latency adjustment. Missing ATTRIBUTIONS.md and NOTICE disclosed: training-data overlap cannot be checked. Florian GO in DECISIONS.md (go gor it, Apache2.0 confirmed) resolves prior attribution hold; publication remains release lane responsibility. COMPLETE: 1500/1500 distinct input IDs, 1477 HTTP200 native answers,23 HTTP422 refusals under the author's own32768-token state+question /65536-token request budget; every refusal counted wrong. Smoke passed all8 checks. All42 release-manifest files and6 prior source files hash-verified. Raw sha256 a28a8f0068f38b548f064a6e678bf63b13304f9add0a67b448dcc8003c9f38f1. Stack torch2.14.0, transformers5.17.0, compressed-tensors0.19.0, accelerate1.15.0. Provider pod91916684-9980-419c-98b7-c7a714a3912c removed and verified absent; approxUSD0.17. Fresh latest feed places this candidate inside both official caps; material open-board top-five change requires release-lane preview disposition before publication.
Cost basis (estimated)
ESTIMATE (base-model reference, not charged). Unmodified pricing_v15.reference resolves google/gemma-4-31B-it to google/gemma-4-31b-it at USD0.09/M input and USD0.34/M output in frozen snapshot. Measured output tokens are 0 by construction (decision_contract.py and release.json); one forward-pass option readout. Common v1.6.1 cost input range applied unchanged.
Latency adjustment
x2 + 0.15 s (assumption, not measured)

Frequently asked questions

What does JevBench show for Jev and René-1 31B FP8?
Jev 1.13.0 has Capability 77.1 and is the unranked reference (the open-weights board ranks 92 open-weights systems; hosted API offerings are ranked on the API leaderboard). René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout) has Capability 76.0. Rank #4 on the open-weights board in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout): estimated, $0.063 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is René-1 31B FP8 open source?
The published licence note for René-1 31B FP8 (salfatigroup, Gemma 4 31B full fine-tune, one-pass option readout) is apache-2.0. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.