Jev vs H2O-Lightning-4B v1.1: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

Rank #3 on the open-weights board in v1.6.1.

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and H2O-Lightning-4B v1.1: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: H2O-Lightning-4B v1.1 — Open weights · LLM decoder · Score 72.5 (#3)
  • B: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (ranked, not ranked)

The four score axes

Radar: the four score axes, two systemsThe four score axes, H2O-Lightning-4B v1.1 vs Jev 1.13.0. Intelligence: 60.0 vs 63.6; Calibration: 90.0 vs 90.6; Speed: 92.6 vs 91.5; Cost: 60.3 vs 54.7.50100Intelligence60.0 · 63.6Calibration90.0 · 90.6Speed92.6 · 91.5Cost60.3 · 54.7
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

Radar: capability by subject topic, two systemsCapability by subject topic, H2O-Lightning-4B v1.1 vs Jev 1.13.0. Rules, policy & law: 61.7 vs 63.3; Coding & software: 68.4 vs 70.5; Math & numbers: 37.2 vs 22.6; Support & operations: 51.0 vs 44.4; Finance & commerce: 65.9 vs 50.2; Everyday language: 61.9 vs 57.4; Safety & security: 70.5 vs 32.4.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1011 items (151 open, 860 sealed).Rules, policy& law61.7 · 63.3Coding & software: code, SQL, repositories, developer tools and IT systems. 330 items (62 open, 268 sealed).Coding &software68.4 · 70.5Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 234 items (37 open, 197 sealed).Math &numbers37.2 · 22.6Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 113 items (21 open, 92 sealed).Support &operations51.0 · 44.4Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 99 items (12 open, 87 sealed).Finance &commerce65.9 · 50.2Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 52 items (11 open, 41 sealed).Everydaylanguage61.9 · 57.4Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 48 items (6 open, 42 sealed).Safety &security70.5 · 32.4
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1011 items (151 open / 860 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 330 items (62 open / 268 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 234 items (37 open / 197 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 113 items (21 open / 92 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 99 items (12 open / 87 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 52 items (11 open / 41 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 48 items (6 open / 42 sealed)

Use cases (TypeSafe categories)

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), H2O-Lightning-4B v1.1 vs Jev 1.13.0. Model routing: 70.7 vs 77.4; Legal & compliance: 67.2 vs 75.3; Customer support: 61.5 vs 53.0; E-commerce: 59.6 vs 40.3; Insurance claims: 83.1 vs 58.0; Risk assessment: 62.0 vs 34.4; Financial crime: 37.7 vs 36.7; Moderation: 56.2 vs 54.7; Code linting: n=16 vs 15.4; Lead generation: n=17 vs 56.5; Feature extraction: n=19 vs 19.9; LLM guardrails: n=17 vs 68.6; Search & retrieval: n=17 vs 76.2; Advertising: — vs 28.2; Demand forecasting: n=16 vs 0.0; Recruiting: — vs 3.1; Knowledge graphs: n=18 vs 22.4; Gaming: — vs 14.5; Science: — vs 25.0.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 471 items (86 open, 385 sealed).Model routing70.7 · 77.4Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 466 items (74 open, 392 sealed).Legal &compliance67.2 · 75.3Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 179 items (37 open, 142 sealed).Customersupport61.5 · 53.0E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 65 items (9 open, 56 sealed).E-commerce59.6 · 40.3Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 63 items (9 open, 54 sealed).Insuranceclaims83.1 · 58.0Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 62 items (8 open, 54 sealed).Riskassessment62.0 · 34.4Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 51 items (9 open, 42 sealed).Financialcrime37.7 · 36.7Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 47 items (7 open, 40 sealed).Moderation56.2 · 54.7Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 33 items (3 open, 30 sealed).Code lintingn=16 · 15.4Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 31 items (4 open, 27 sealed).Leadgenerationn=17 · 56.5Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 31 items (5 open, 26 sealed).Featureextractionn=19 · 19.9LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 30 items (5 open, 25 sealed).LLMguardrailsn=17 · 68.6Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 30 items (5 open, 25 sealed).Search &retrievaln=17 · 76.2Advertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open, 27 sealed).Advertisingn/a · 28.2Demand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 30 items (3 open, 27 sealed).Demandforecastingn=16 · 0.0Recruiting: resumes, applications, interview feedback, matching candidates to roles. 30 items (1 open, 29 sealed).Recruitingn/a · 3.1Graphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 30 items (3 open, 27 sealed).Knowledgegraphsn=18 · 22.4Gaming: player reports, in-game chat, game support. 30 items (3 open, 27 sealed).Gamingn/a · 14.5Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 30 items (2 open, 28 sealed).Sciencen/a · 25.0
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), open and sealed items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.A (H2O-Lightning-4B v1.1) is measured on 8 of 19 spokes only, so it is drawn as points, not a shape — measured on fewer items; full radar after the next re-measure.Gaps, not zeros: Open spokes (n/a or n=…) have fewer than 30 answered items for that system or no published value.

Low sample, n < 30 — indicative only

Category (items)A: H2O-Lightning-4B v1.1B: Jev 1.13.0
Semantic code linting (33; a system answered fewer than 30)20.9 n=1615.4 n=33
Lead generation (31; a system answered fewer than 30)54.4 n=1756.5 n=31
Feature extraction for predictive modeling (31; a system answered fewer than 30)50.6 n=1919.9 n=31
LLM guardrails (30; a system answered fewer than 30)48.5 n=1768.6 n=30
Search and retrieval (30; a system answered fewer than 30)54.7 n=1776.2 n=30
Demand forecasting (30; a system answered fewer than 30)22.2 n=160.0 n=30
Graphs and knowledge graphs (30; a system answered fewer than 30)42.3 n=1822.4 n=30

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 471 items (86 open / 385 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 466 items (74 open / 392 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 179 items (37 open / 142 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 65 items (9 open / 56 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 63 items (9 open / 54 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 62 items (8 open / 54 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 51 items (9 open / 42 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 47 items (7 open / 40 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 33 items (3 open / 30 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 31 items (4 open / 27 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 31 items (5 open / 26 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 30 items (5 open / 25 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 30 items (5 open / 25 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 30 items (3 open / 27 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 30 items (3 open / 27 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 30 items (1 open / 29 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 30 items (3 open / 27 sealed)
  • Gaming — player reports, in-game chat, game support. 30 items (3 open / 27 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 30 items (2 open / 28 sealed)

Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 148 items (24 open / 124 sealed) — not a use case of its own, so it is counted but not drawn.

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, H2O-Lightning-4B v1.1 vs Jev 1.13.0. Choice · open: 75.2 vs 79.0; Choice · sealed: 64.7 vs 75.8; Noul · open: 60.9 vs 54.5; Noul · sealed: 68.1 vs 46.1; Score · open: 47.0 vs 63.3; Score · sealed: 44.3 vs 63.0.50100Choice · open75.2 · 79.0Choice ·sealed64.7 · 75.8Noul · open60.9 · 54.5Noul · sealed68.1 · 46.1Score · open47.0 · 63.3Score ·sealed44.3 · 63.0
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, H2O-Lightning-4B v1.1 vs Jev 1.13.0. Easy: 88.2 vs 93.9; Standard: 57.6 vs 63.2; Judge: 61.0 vs 67.7; Hard: 63.7 vs 65.6.50100Easy88.2 · 93.9Standard57.6 · 63.2Judge61.0 · 67.7Hard63.7 · 65.6
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, H2O-Lightning-4B v1.1 vs Jev 1.13.0. Easy: 74.4 vs 78.4; Standard: 62.6 vs 68.7; Judge: 63.2 vs 66.9; Hard: 54.3 vs 58.9.50100Easy74.4 · 78.4Standard62.6 · 68.7Judge63.2 · 66.9Hard54.3 · 58.9
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).

Family and language are authoring metadata of every item. Each of the 1,500 v1.6 items and each of the 387 supplement items (L1 language supplement 354, L2 use-case supplement 33) was labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod, with the same model, questions, taxonomy and item-group rules (labelling rounds r13 and r14); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) of the main pool were checked by hand (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems and full-set hosted APIs answered S u P (1,500 items) plus the supplements where their row is tagged so; API rows re-run on A4/A5 have no cells yet.

  • Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
  • Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items of S u P and the L2 supplement.
All values as a table
SpokeA: H2O-Lightning-4B v1.1B: Jev 1.13.0
The four score axes
Intelligence60.063.6
Calibration90.090.6
Speed92.691.5
Cost60.354.7
Capability by subject topic
Rules, policy & law61.763.3
Coding & software68.470.5
Math & numbers37.222.6
Support & operations51.044.4
Finance & commerce65.950.2
Everyday language61.957.4
Safety & security70.532.4
Use cases (TypeSafe categories)
Model routing70.777.4
Legal & compliance67.275.3
Customer support61.553.0
E-commerce59.640.3
Insurance claims83.158.0
Risk assessment62.034.4
Financial crime37.736.7
Moderation56.254.7
Code lintingn=1615.4
Lead generationn=1756.5
Feature extractionn=1919.9
LLM guardrailsn=1768.6
Search & retrievaln=1776.2
Advertising—28.2
Demand forecastingn=160.0
Recruiting—3.1
Knowledge graphsn=1822.4
Gaming—14.5
Science—25.0
Competence per request type, open / sealed
Choice · open75.279.0
Choice · sealed64.775.8
Noul · open60.954.5
Noul · sealed68.146.1
Score · open47.063.3
Score · sealed44.363.0
Competence per tier — open set
Easy88.293.9
Standard57.663.2
Judge61.067.7
Hard63.765.6
Competence per tier — sealed set
Easy74.478.4
Standard62.668.7
Judge63.266.9
Hard54.358.9

Published values

MeasureJev 1.13.0 (TypeSafe AI)H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim)
Capability Score77.175.0
Rank (open-weights board)Reference, not ranked#3
Composite (secondary)71.572.5
intelligence63.660.0
calibration90.690.0
speed91.592.6
cost54.760.3
USD per 1,000 decisions$0.032 per 1,000 decisions$0.021 per 1,000 decisions
p50 latency (seconds)0.20.2
Last measured2026-10-052026-10-04

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim)

apache-2.0

Published source
Measured on
evaluator-owned Lium GPU pod, 1x NVIDIA RTX 5090, official vLLM 0.30.0 image + the author's open standard-library shim (h2o_lightning_shim.py) as published on the model card, offline (all credentials unset), driven over loopback only by the JevBench client; label-free inputs; pod destroyed afterwards; JevBench add-request (GitHub #181), measured 4 Oct 2026
Cost basis (estimated)
ESTIMATE (self-hosted open weights, no public tariff): cost per 1,000 decisions measured for H2O-Lightning-4B v1.0 on the v1.5 pool (same base, shim and token path) carried for v1.1, the practice of every Qwen3.5-4B row on this board; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off. Cross-check: v1.1's own v1.6 input tokens on the common cost basis give USD 0.0204 per 1,000 (within 3 %)
Latency adjustment
x2 + 0.15 s (assumption, not measured)

Frequently asked questions

What does JevBench show for Jev and H2O-Lightning-4B v1.1?
Jev 1.13.0 has Capability 77.1 and is the unranked reference (the open-weights board ranks 68 open-weights systems; hosted API offerings are ranked on the API leaderboard). H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim) has Capability 75.0. Rank #3 on the open-weights board in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim): estimated, $0.021 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is H2O-Lightning-4B v1.1 open source?
The published licence note for H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim) is apache-2.0. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.