Jev vs SPX-CD Flash: published benchmark comparison

Data: JevBench v1.6.1, published 2026-10-06

API offering, ranked on the API leaderboard in v1.6.1. JevBench API leaderboard

Compare Jev-compatible models by separate measures and their measurement conditions. A historical measurement is not a current rank.

Jev and SPX-CD Flash: published axes and category radars

Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (ranked, not ranked)
  • B: SPX-CD Flash — System (router / cascade / ensemble) · Score 66.7 (ranked, not ranked)Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs SPX-CD Flash. Intelligence: 63.6 vs 58.4; Calibration: 90.6 vs 90.2; Speed: 91.5 vs 79.4; Cost: 54.7 vs 52.1.50100Intelligence63.6 · 58.4Calibration90.6 · 90.2Speed91.5 · 79.4Cost54.7 · 52.1
0–100, the values in the table. An axis a system has no published value for is left as a gap (it counts as 0 in the composite).

Capability by subject topic

SPX-CD Flash: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: capability by subject topic, two systemsCapability by subject topic, Jev 1.13.0 vs SPX-CD Flash. Rules, policy & law: 49.4 vs 64.7; Coding & software: 64.3 vs 71.8; Math & numbers: 10.4 vs 28.0; Finance & commerce: 40.9 vs 58.5; Support & operations: 48.8 vs 27.6; Everyday language: 77.5 vs 43.2; Safety & security: 52.2 vs 46.8.50100Rules, policy & law: applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open, 1465 sealed).Rules, policy& law49.4 · 64.7Coding & software: code, SQL, repositories, developer tools and IT systems. 382 items (62 open, 320 sealed).Coding &software64.3 · 71.8Math & numbers: a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open, 333 sealed).Math &numbers10.4 · 28.0Finance & commerce: money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open, 211 sealed).Finance &commerce40.9 · 58.5Support & operations: support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open, 139 sealed).Support &operations48.8 · 27.6Everyday language: short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open, 134 sealed).Everydaylanguage77.5 · 43.2Safety & security: untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open, 103 sealed).Safety &security52.2 · 46.8
What the decision is about — math and dates, coding, rules, money, support work, everyday messages, safety. Chance-corrected competence per category (0 = chance, 100 = perfect), Jev 1.13.0: S+P+L1+L2+L3; SPX-CD Flash: S+P items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.
What each category means · items per category
  • Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 1616 items (151 open / 1465 sealed)
  • Coding & software — code, SQL, repositories, developer tools and IT systems. 382 items (62 open / 320 sealed)
  • Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 370 items (37 open / 333 sealed)
  • Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 223 items (12 open / 211 sealed)
  • Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 160 items (21 open / 139 sealed)
  • Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 145 items (11 open / 134 sealed)
  • Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 109 items (6 open / 103 sealed)

Use cases (TypeSafe categories)

SPX-CD Flash: Language and category supplements have not been measured for this newly released row; the original current-pool category counts are shown.

Radar: use cases (typesafe categories), two systemsUse cases (TypeSafe categories), Jev 1.13.0 vs SPX-CD Flash. Model routing: 73.0 vs 71.3; Legal & compliance: 68.6 vs 75.1; Customer support: 49.6 vs 38.5; Other: 40.1 vs 19.3; E-commerce: 40.5 vs 67.6; Insurance claims: 46.1 vs 79.4; Risk assessment: 32.8 vs 62.9; Financial crime: 35.8 vs 27.0; Feature extraction: 28.3 vs n=19; Lead generation: 49.2 vs n=17; Recruiting: 0.0 vs —; Knowledge graphs: 28.8 vs n=18; LLM guardrails: 68.0 vs n=17; Moderation: 51.6 vs 56.6; Code linting: 29.1 vs n=16; Search & retrieval: 80.9 vs n=17; Science: 21.2 vs —; Advertising: 16.4 vs —; Gaming: 5.1 vs —; Demand forecasting: 0.0 vs n=16.50100Model routing: deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open, 438 sealed).Model routing73.0 · 71.3Legal and compliance: contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open, 441 sealed).Legal &compliance68.6 · 75.1Customer support: classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open, 203 sealed).Customersupport49.6 · 38.5Other: none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open, 150 sealed).Other40.1 · 19.3E-commerce marketplaces: product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open, 110 sealed).E-commerce40.5 · 67.6Insurance claims: first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open, 110 sealed).Insuranceclaims46.1 · 79.4Risk assessment: estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open, 110 sealed).Riskassessment32.8 · 62.9Financial crime: suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open, 92 sealed).Financialcrime35.8 · 27.0Feature extraction for predictive modeling: turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open, 94 sealed).Featureextraction28.3 · n=19Lead generation: matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open, 91 sealed).Leadgeneration49.2 · n=17Recruiting: resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open, 93 sealed).Recruiting0.0 · n/aGraphs and knowledge graphs: entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open, 91 sealed).Knowledgegraphs28.8 · n=18LLM guardrails: checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open, 88 sealed).LLMguardrails68.0 · n=17Moderation and trust and safety: user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open, 86 sealed).Moderation51.6 · 56.6Semantic code linting: checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open, 87 sealed).Code linting29.1 · n=16Search and retrieval: scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open, 85 sealed).Search &retrieval80.9 · n=17Scientific discovery: screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open, 88 sealed).Science21.2 · n/aAdvertising: ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open, 85 sealed).Advertising16.4 · n/aGaming: player reports, in-game chat, game support. 86 items (3 open, 83 sealed).Gaming5.1 · n/aDemand forecasting: purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open, 80 sealed).Demandforecasting0.0 · n=16
The real-world application area of each decision, after the TypeSafe use-case map; an item can count in two. Chance-corrected competence per category (0 = chance, 100 = perfect), Jev 1.13.0: S+P+L1+L2+L3; SPX-CD Flash: S+P items pooled; only categories with at least 30 items are spokes, smaller ones are listed below. Hover a category for its definition and item count.B (SPX-CD Flash) is measured on 9 of 20 spokes only, so it is drawn as points, not a shape — measured on fewer items; full radar after the next re-measure.Gaps, not zeros: Open spokes (n/a or n=…) have fewer than 30 answered items for that system or no published value; hosted APIs answer a smaller item set, so more of their category cells stay under that.

Low sample, n < 30 — indicative only

Category (items)A: Jev 1.13.0B: SPX-CD Flash
Feature extraction for predictive modeling (99; a system answered fewer than 30)28.3 n=9938.9 n=19
Lead generation (95; a system answered fewer than 30)49.2 n=9553.4 n=17
Graphs and knowledge graphs (94; a system answered fewer than 30)28.8 n=9442.4 n=18
LLM guardrails (93; a system answered fewer than 30)68.0 n=9315.3 n=17
Semantic code linting (90; a system answered fewer than 30)29.1 n=9021.0 n=16
Search and retrieval (90; a system answered fewer than 30)80.9 n=9080.5 n=17
Demand forecasting (83; a system answered fewer than 30)0.0 n=830.0 n=16

Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.

What each category means · items per category
  • Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 524 items (86 open / 438 sealed)
  • Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 515 items (74 open / 441 sealed)
  • Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 240 items (37 open / 203 sealed)
  • Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 174 items (24 open / 150 sealed)
  • E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 119 items (9 open / 110 sealed)
  • Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 119 items (9 open / 110 sealed)
  • Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 118 items (8 open / 110 sealed)
  • Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 101 items (9 open / 92 sealed)
  • Feature extraction for predictive modeling — turning natural-language data into probabilistic features or estimates for a downstream prediction. 99 items (5 open / 94 sealed)
  • Lead generation — matching companies or inbound messages to an ideal customer profile, buyer intent, prioritising and routing sales leads. 95 items (4 open / 91 sealed)
  • Recruiting — resumes, applications, interview feedback, matching candidates to roles. 94 items (1 open / 93 sealed)
  • Graphs and knowledge graphs — entity types, relationships between records, contradictions between facts, multi-step lookups across linked facts. 94 items (3 open / 91 sealed)
  • LLM guardrails — checking LLM inputs, outputs and tool calls: prompt injection, jailbreaks, policy violations, tool-call errors, answer quality. 93 items (5 open / 88 sealed)
  • Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 93 items (7 open / 86 sealed)
  • Semantic code linting — checking code or writing against conventions and guidelines, as in CI review of code. 90 items (3 open / 87 sealed)
  • Search and retrieval — scoring query-to-candidate relevance, reranking results, selecting useful context for RAG. 90 items (5 open / 85 sealed)
  • Scientific discovery — screening papers, labelling research passages or survey answers, checking that citations support claims, methodology checks. 90 items (2 open / 88 sealed)
  • Advertising — ad creatives, campaign copy, brand safety, prohibited claims in ads. 88 items (3 open / 85 sealed)
  • Gaming — player reports, in-game chat, game support. 86 items (3 open / 83 sealed)
  • Demand forecasting — purchase intent, product interest, demand and supply signals for forecasting. 83 items (3 open / 80 sealed)

Competence per request type, open / sealed

Radar: competence per request type, open / sealed, two systemsCompetence per request type, open / sealed, Jev 1.13.0 vs SPX-CD Flash. Choice · open: 79.0 vs 75.4; Choice · sealed: 75.8 vs 71.0; Noul · open: 54.5 vs 52.9; Noul · sealed: 46.1 vs 45.3; Score · open: 63.3 vs 51.7; Score · sealed: 63.0 vs 54.2.50100Choice · open79.0 · 75.4Choice ·sealed75.8 · 71.0Noul · open54.5 · 52.9Noul · sealed46.1 · 45.3Score · open63.3 · 51.7Score ·sealed63.0 · 54.2
Chance-corrected competence (0 = chance) for Choice, Noul and Score on the 300 open and 1200 sealed decisions.

Competence per tier — open set

Radar: competence per tier — open set, two systemsCompetence per tier — open set, Jev 1.13.0 vs SPX-CD Flash. Easy: 93.9 vs 92.7; Standard: 63.2 vs 62.8; Judge: 67.7 vs 45.7; Hard: 65.6 vs 68.7.50100Easy93.9 · 92.7Standard63.2 · 62.8Judge67.7 · 45.7Hard65.6 · 68.7
Per-tier competence, the three request types pooled by their published decision counts.

Competence per tier — sealed set

Radar: competence per tier — sealed set, two systemsCompetence per tier — sealed set, Jev 1.13.0 vs SPX-CD Flash. Easy: 78.4 vs 79.1; Standard: 68.7 vs 67.2; Judge: 66.9 vs 61.2; Hard: 58.9 vs 52.6.50100Easy78.4 · 79.1Standard68.7 · 67.2Judge66.9 · 61.2Hard58.9 · 52.6
Per-tier competence on the sealed decisions, types pooled the same way; item text stays private.
How the categories were made

O1S chance-corrected competence: equal mean across the request types present, clipped to 0–100. Zero means the equal-weight mean across present request types was at or below baseline; negative means are clipped. It does not mean every type scored zero. Choice uses a random-option baseline; Noul uses a fixed 50% accuracy baseline (abstentions count as wrong); Score uses the midpoint-guess error (random-guess fallback when every gold is at the midpoint). Compare models within a category; row pools and supported types can differ.

Categories use ruled subject-topic labels and authoring use-case precedence across the answered pool union, with stable item identities counted once.

  • Official O1S; recorded failures count where supported; no missing response is imputed; stable identity dedupe parent first.
  • S-based category cells pool recorded supported responses from S+P and the answered L1/L2/L3 supplements. API overlay category cells keep their separately recorded A4/A5+P union. Headline and per-type/tier scores are unchanged.
All values as a table
SpokeA: Jev 1.13.0B: SPX-CD Flash
The four score axes
Intelligence63.658.4
Calibration90.690.2
Speed91.579.4
Cost54.752.1
Capability by subject topic
Rules, policy & law49.464.7
Coding & software64.371.8
Math & numbers10.428.0
Finance & commerce40.958.5
Support & operations48.827.6
Everyday language77.543.2
Safety & security52.246.8
Use cases (TypeSafe categories)
Model routing73.071.3
Legal & compliance68.675.1
Customer support49.638.5
Other40.119.3
E-commerce40.567.6
Insurance claims46.179.4
Risk assessment32.862.9
Financial crime35.827.0
Feature extraction28.3n=19
Lead generation49.2n=17
Recruiting0.0—
Knowledge graphs28.8n=18
LLM guardrails68.0n=17
Moderation51.656.6
Code linting29.1n=16
Search & retrieval80.9n=17
Science21.2—
Advertising16.4—
Gaming5.1—
Demand forecasting0.0n=16
Competence per request type, open / sealed
Choice · open79.075.4
Choice · sealed75.871.0
Noul · open54.552.9
Noul · sealed46.145.3
Score · open63.351.7
Score · sealed63.054.2
Competence per tier — open set
Easy93.992.7
Standard63.262.8
Judge67.745.7
Hard65.668.7
Competence per tier — sealed set
Easy78.479.1
Standard68.767.2
Judge66.961.2
Hard58.952.6

Published values

MeasureJev 1.13.0 (TypeSafe AI)SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)
Capability Score77.174.3
Rank (open-weights board)Reference, not rankedAPI offering, ranked on the API leaderboard
Composite (secondary)71.566.7
intelligence63.658.4
calibration90.690.2
speed91.579.4
cost54.752.1
USD per 1,000 decisions$0.032 per 1,000 decisions$0.039 per 1,000 decisions
p50 latency (seconds)0.21.0
Last measured2026-10-052026-10-06

Jev 1.13.0 (TypeSafe AI)

proprietary API

Published source
Measured on
the operator's hosted API
Cost basis (not published)
TypeSafe list tariff USD 0.042 per 1M input tokens, output free; measured on the v1.6.1 common cost basis (all items except those outside the Jev reference input range) from this run's token usage
Latency adjustment
none (API)

SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)

Licence not stated

Published source
Measured on
measured 2026-10-06 on the live v1.6.0 S u P (1,500 rows) under Florian's 5 Oct 2026 full-set amendment, through the release lane's own full-set API lane (run_api_full_v161.py, rotation.py expose --full-set logged 1,200 sealed items to provider 'surd' before the first request); stock jevbench.adapters.typesafe with an endpoint override only, reasoning at the provider default, one request at a time, no retries; add-requests run 53, benchmarkheaven.com/submit ref 3FF00545
Cost basis (estimated)
Operator's own public-beta overage tariff, conservative reading USD 0.04 per 1M input tokens, output USD 0 (the provider's /v1/models payload reports output_price "0"): three dated primary readings of the operator's own published overage tariff: 24 Sep 2026 provider console (flash 0.025, pro 0.08 per 1M input), 1 Oct 2026 docs page (flash 0.04, pro 0.02), 6 Oct 2026 docs page AND /v1/models payload agreeing (flash 0.02, pro 0.04). Sources conflict, so the per-model MAXIMUM is used, which is the published fastino-gliner-2-5-decide precedent ("published sources conflict ... conservative higher"). Today's agreeing pair would be USD 0.02/M, i.e. half this row's cost. No base-model floor applies: the submitter states the base model stays sealed until a later open release, so no base model exists to reference. Priced on usage.input_tokens as the provider reports them. The provider also reports billable_input_tokens = input_tokens / 2 on every single item (297/297 of the public set, exactly 2.0): its default effort 2 renders the prompt twice and its docs bundle the second pass free during public beta ("default effort 2 is included at the single-pass input price"). That bundle is a public-beta promotion, and TASK.md forbids a free or promotional tier setting the price, so the two passes actually consumed are charged. The operator's own billed figure is exactly half of this row's cost
Latency adjustment
none (API)

Frequently asked questions

What does JevBench show for Jev and SPX-CD Flash?
Jev 1.13.0 has Capability 77.1 and is the unranked reference (the open-weights board ranks 92 open-weights systems; hosted API offerings are ranked on the API leaderboard). SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint) has Capability 74.3. API offering, ranked on the API leaderboard in v1.6.1. Historical scores use their original method and are not ranked against the current release.
How does Capability differ from Composite?
Capability is the mean of Intelligence and Calibration; headline ranks require both official cost and median-latency caps. Composite additionally scores Speed and Cost and is secondary.
Can I compare the prices as actual bills?
Jev 1.13.0 (TypeSafe AI): not published, $0.032 per 1,000 decisions. SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint): estimated, $0.039 per 1,000 decisions. Estimates and self-reported vendor prices are not measured charges.
Is SPX-CD Flash open source?
The published licence note for SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint) is not stated. Jev 1.13.0 is listed as proprietary API. Check the linked sources for terms.