JevBench by Benchmark Heaven · released v1.4.2

Jev vs Cygnet: published benchmark comparison

A side-by-side view of Jev 1.13.0 and Cygnet (blockbrain, frozen Gemma-4-12B-it) from the same hash-checked v1.4.2 aggregate. The overall score is a composite; compare the separate measures against your use case.

Four-radar comparison: Jev and Cygnet

Four radars compare this fixed pair across the score axes, accuracy per tier, and accuracy by family on the hard tier and sealed set. Further out is better on every spoke.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#2)
  • B: Cygnet — system-one-open · Score 61.8 (#4)

The four score axes

Radar: the four score axes, two systemsThe four score axes, Jev 1.13.0 vs Cygnet. Intelligence: 53.1 vs 49.5; Calibration: 76.3 vs 74.9; Speed: 83.3 vs 90.7; Cost: 52.0 vs 52.8.50100Intelligence53.1 · 49.5Calibration76.3 · 74.9Speed83.3 · 90.7Cost52.0 · 52.8
0–100, the values in the table. A label-only system has no calibration (counted as 0).

Accuracy per tier, incl. sealed

Radar: accuracy per tier, incl. sealed, two systemsAccuracy per tier, incl. sealed, Jev 1.13.0 vs Cygnet. Easy: 100% vs 100%; Standard: 99% vs 97%; Judge: 95% vs 95%; Hard: 74% vs 75%; Sealed: 37% vs 34%.50100Easy100% · 100%Standard99% · 97%Judge95% · 95%Hard74% · 75%Sealed37% · 34%
Share correct per tier; Sealed = the 308 private decisions, aggregate only.

Hard tier by family (v1.2 topics)

Cygnet has no published v1.2 hard-tier family breakdown.

Radar: hard tier by family (v1.2 topics), two systemsHard tier by family (v1.2 topics), Jev 1.13.0 vs Cygnet. Adversarial: 100% vs —; Ambiguous: 79% vs —; Judge: 79% vs —; Long policy: 61% vs —; Multi-hop: 86% vs —; Probability: 80% vs —; Routing: 100% vs —; Temporal / numeric: 27% vs —; Trade-off: 92% vs —; Trap: 100% vs —.50100Adversarial100%Ambiguous79%Judge79%Long policy61%Multi-hop86%Probability80%Routing100%Temporal /numeric27%Trade-off92%Trap100%
Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out).

Sealed set by family

Radar: sealed set by family, two systemsSealed set by family, Jev 1.13.0 vs Cygnet. Ambiguous / abstain: 30% vs 38%; Judge: 34% vs 54%; Long policy: 28% vs 23%; Multi-hop: 45% vs 32%; Paraphrase: 64% vs 64%; Probability: 50% vs 36%; Safety judge: 38% vs 44%; Temporal / numeric: 29% vs 18%; Trade-off: 38% vs 19%; Trap / adversarial: 42% vs 50%.50100Ambiguous /abstain30% · 38%Judge34% · 54%Long policy28% · 23%Multi-hop45% · 32%Paraphrase64% · 64%Probability50% · 36%Safety judge38% · 44%Temporal /numeric29% · 18%Trade-off38% · 19%Trap /adversarial42% · 50%
Share correct within each sealed family — system-level aggregates; the items stay private.

Published values

Jev 1.13.0 (TypeSafe AI) has the higher published JevBench Score. The displayed score uses one decimal place; the comparison above uses the same published release.

All values as a table
MeasureJev 1.13.0 (TypeSafe AI)Cygnet (blockbrain, frozen Gemma-4-12B-it)
Published rank#2#4
JevBench Score63.361.8
Sealed-set accuracy36.7%33.8%
Intelligence axis53.149.5
Calibration axis76.374.9
Speed axis83.390.7
Cost axis52.052.8
Cost evidencemeasuredestimated
Cost per 1,000 decisions$0.040 per 1,000 decisions$0.037 per 1,000 decisions

Intelligence, Calibration, Speed and Cost — equal-weight harmonic mean, with generalization and Jev-class gates. Cost evidence is labeled per row. Speed is a benchmark axis; deployment latency depends on the endpoint and conditions.

Jev 1.13.0 (TypeSafe AI)

Marked closed in the published row. License note: proprietary API.

Published source

Measured on
production API (api.typesafe.ai)
Speed adjustment
none (production API)
Cost basis
public tariff x measured tokens (https://docs.typesafe.ai/models (output tokens not billed)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)

Cygnet (blockbrain, frozen Gemma-4-12B-it)

Code and weights marked open in the published row. License note: shim MIT; weights Apache-2.0 with Google's Gemma Prohibited Use Policy.

Published source

Measured on
RTX PRO 6000 (evaluator-owned Lium pod) · our evaluator-owned Lium GPU pod (RTX PRO 6000), offline read-only container, author's server on loopback
Speed adjustment
x2 + 0.15 s (assumption, not measured)
Cost basis
ESTIMATE: OpenRouter google/gemma-3-12b-it list price $0.05/M input (the nearest hosted 12B Gemma; gemma-4-12b-it is not listed; the Winnow-12B / Jev-Omni precedent), read 2026-09-24; the server's own usage.input_tokens; nothing generated; estimated, not charged

Compare more JevBench pairs: Jev vs decider-4b v2 · Jev vs JevK5 · Jev vs Hopper · Jev vs Winnow-12B Q8 · Jev vs reflex 4B · Jev vs Laya

See Jev alternatives · Choose by use case

Frequently asked questions

What does JevBench show for Jev and Cygnet?
In v1.4.2, Jev is rank 2 with a JevBench Score of 63.3; Cygnet (blockbrain, frozen Gemma-4-12B-it) is rank 4 with a score of 61.8. The table shows their published axes and sealed-set accuracy separately.
Which system has higher sealed-set accuracy?
Jev 1.13.0 (TypeSafe AI) has the higher published sealed-set accuracy (36.7% versus 33.8%). This is separate from the composite JevBench Score.
Can I compare the cost values as actual bills?
Jev 1.13.0 (TypeSafe AI): measured at $0.040 per 1,000 decisions. Cygnet (blockbrain, frozen Gemma-4-12B-it): estimated at $0.037 per 1,000 decisions. Estimated and announced bases are not measured charges; inspect the board’s full row disclosure.
Is Cygnet open source, and is Jev?
The published row lists Cygnet (blockbrain, frozen Gemma-4-12B-it) with license note “shim MIT; weights Apache-2.0 with Google's Gemma Prohibited Use Policy”. Jev 1.13.0 is listed as “proprietary API”: TypeSafe AI serves it through its own API and has not published its weights. Check each linked source for the exact terms.