Jev vs Jev-Omni: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

Capability rank #6 in v1.6.1.

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and Jev-Omni: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (#4)
  • B: Jev-Omni — Open weights · LLM decoder · Score 67.7 (#6)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Jev-Omni. Intelligence: 63.6 vs 55.5; Calibration: 90.6 vs 87.0; Speed: 91.5 vs 85.4; Cost: 54.7 vs 56.1.50100Intelligence63.6 · 55.5Calibration90.6 · 87.0Speed91.5 · 85.4Cost54.7 · 56.1
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, Jev 1.13.0 vs Jev-Omni. Rules, policy & law: 70.6 vs 58.7; Coding & software: 74.5 vs 63.0; Math & numbers: 30.5 vs 23.3; Support & operations: 34.9 vs 39.9; Finance & commerce: 60.3 vs 76.6; Everyday language: 51.5 vs 32.4; Safety & security: 35.2 vs 63.3.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open, 612 sealed).Rules, policy& law70.6 · 58.7Coding & software: code, SQL, repositories, developer tools and IT systems. 302 items (62 open, 240 sealed).Coding &software74.5 · 63.0Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open, 156 sealed).Math &numbers30.5 · 23.3Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open, 72 sealed).Support &operations34.9 · 39.9Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open, 54 sealed).Finance &commerce60.3 · 76.6Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open, 35 sealed).Everydaylanguage51.5 · 32.4Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open, 31 sealed).Safety &security35.2 · 63.3
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open / 612 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 302 items (62 open / 240 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open / 156 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open / 72 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open / 54 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open / 35 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open / 31 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jev 1.13.0 vs Jev-Omni. Model routing: 78.2 vs 69.1; Legal & compliance: 81.3 vs 60.5; Customer support: 47.6 vs 34.5; E-commerce: 44.6 vs 52.9; Risk assessment: 42.1 vs 43.2; Insurance claims: 76.5 vs 78.6; Financial crime: 36.0 vs 34.5; Moderation: 53.4 vs 80.3.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open, 356 sealed).Model routing78.2 · 69.1Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open, 313 sealed).Legal &compliance81.3 · 60.5Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open, 112 sealed).Customersupport47.6 · 34.5E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open, 39 sealed).E-commerce44.6 · 52.9Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open, 37 sealed).Riskassessment42.1 · 43.2Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open, 32 sealed).Insuranceclaims76.5 · 78.6Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open, 30 sealed).Financialcrime36.0 · 34.5Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open, 25 sealed).Moderation53.4 · 80.3
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.

Low sample, n < 30 — indicative only

Category (items)A: Jev 1.13.0B: Jev-Omni
Feature extraction for predictive modeling (19)22.5 n=1941.4 n=19
Graphs and knowledge graphs (18)49.8 n=1842.2 n=18
LLM guardrails (17)44.1 n=1765.6 n=17
Search and retrieval (17)80.5 n=1780.9 n=17
Lead generation (17)67.5 n=1780.0 n=17
Semantic code linting (16)12.2 n=169.8 n=16
Demand forecasting (16)0.0 n=160.0 n=16

No published value for either system (under 15 answered items, or no per-category values): Gaming (14), Scientific discovery (14), Recruiting (13), Advertising (11).

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open / 356 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open / 313 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open / 112 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open / 39 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open / 37 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open / 32 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open / 30 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open / 25 sealed)

Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 145 items (24 open / 121 sealed) — not a use case of its own, so it is counted but not drawn.

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs Jev-Omni. Choice · open: 79.0 vs 75.0; Choice · sealed: 75.8 vs 68.9; Noul · open: 54.5 vs 45.1; Noul · sealed: 46.1 vs 32.5; Score · open: 63.3 vs 57.3; Score · sealed: 63.0 vs 54.4.50100Choice · open79.0 · 75.0Choice ·sealed75.8 · 68.9Noul · open54.5 · 45.1Noul · sealed46.1 · 32.5Score · open63.3 · 57.3Score ·sealed63.0 · 54.4
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs Jev-Omni. Easy: 93.9 vs 83.9; Standard: 63.2 vs 58.9; Judge: 67.7 vs 51.3; Hard: 65.6 vs 66.8.50100Easy93.9 · 83.9Standard63.2 · 58.9Judge67.7 · 51.3Hard65.6 · 66.8
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs Jev-Omni. Easy: 78.4 vs 67.6; Standard: 68.7 vs 61.2; Judge: 66.9 vs 60.6; Hard: 58.9 vs 48.0.50100Easy78.4 · 67.6Standard68.7 · 61.2Judge66.9 · 60.6Hard58.9 · 48.0
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).

Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 v1.6 items was also labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod (same model, questions and taxonomy as v1.5); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) were checked by hand; item-group rules fix the systematic misses (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems saw S u P (1,500 items); hosted API systems saw only A u P (600 items), so their cells cover fewer items.

  • Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
  • Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
SpokeA: Jev 1.13.0B: Jev-Omni
The four score axes
Intelligence63.655.5
Calibration90.687.0
Speed91.585.4
Cost54.756.1
Capability by subject topic
Rules, policy & law70.658.7
Coding & software74.563.0
Math & numbers30.523.3
Support & operations34.939.9
Finance & commerce60.376.6
Everyday language51.532.4
Safety & security35.263.3
Use cases (TypeSafe categories)
Model routing78.269.1
Legal & compliance81.360.5
Customer support47.634.5
E-commerce44.652.9
Risk assessment42.143.2
Insurance claims76.578.6
Financial crime36.034.5
Moderation53.480.3
Competence per request type, open / sealed
Choice · open79.075.0
Choice · sealed75.868.9
Noul · open54.545.1
Noul · sealed46.132.5
Score · open63.357.3
Score · sealed63.054.4
Competence per tier — open set
Easy93.983.9
Standard63.258.9
Judge67.751.3
Hard65.666.8
Competence per tier — sealed set
Easy78.467.6
Standard68.761.2
Judge66.960.6
Hard58.948.0

Published values

MeasureJev 1.13.0 (TypeSafe AI)Jev-Omni (akhilaaa3, Gemma-4-12B merged)
Capability Score77.171.3
Capability rank#4#6
Composite (secondary)71.567.7
intelligence63.655.5
calibration90.687.0
speed91.585.4
cost54.756.1
USD per 1,000 decisions$0.032 per 1,000 decisions$0.029 per 1,000 decisions
p50 latency (seconds)0.20.5
Last measured2026-10-052026-10-02

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

Jev-Omni (akhilaaa3, Gemma-4-12B merged)

Apache-2.0, following Gemma 4; dataset rights stated separately by the author

Published source
Measured on
evaluator-owned Lium GPU pod (RTX6000), offline read-only container
Cost basis (estimated)
v1.5.4 published cost per 1,000 decisions carried (pricing rules unchanged; v1.6 item lengths differ): documented hosted-model estimate; no exact base-model floor applies
Latency adjustment
x2 + 0.15 s (assumption, not measured)

Frequently asked questions

What does JevBench show for Jev and Jev-Omni?
Jev 1.13.0 has Capability 77.1, rank #4 of 64. Jev-Omni (akhilaaa3, Gemma-4-12B merged) has Capability 71.3. Capability rank #6 in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. Jev-Omni (akhilaaa3, Gemma-4-12B merged): estimated, $0.029 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is Jev-Omni open source?
The published licence note for Jev-Omni (akhilaaa3, Gemma-4-12B merged) is Apache-2.0, following Gemma 4; dataset rights stated separately by the author. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.