JevBench by Benchmark Heaven · released v1.4.2

Is Jev open source?

No. Jev 1.13.0 is TypeSafe AI’s proprietary API; its weights are not published. But 60 open-weight Jev-class models were measured on the same benchmark, and one of them ranks above Jev on the composite score.

Jev 1.13.0 out-reasons decider-4b v2 (Intelligence 53.1 vs 49.4) and is better calibrated; decider-4b v2 leads on speed and cost. JevBench weighs the four axes equally; sort by Intelligence for raw reasoning.

The 10 highest-ranked open-weight Jev-class models

Ranks are overall ranks in v1.4.2, with Jev 1.13.0 at #2 (JevBench Score 63.3) for reference.

RankSystemScoreIntelligenceCost per 1,000LicenseMeasured on
1decider-4b v2 (Mapika)64.149.4$0.020 (estimated)Apache-2.0 (package and weights)RTX PRO 6000 (evaluator-owned Lium pod)
3JevK5 v0.2.062.048.9$0.022 (estimated)Apache-2.0 (code/adapter); Apache-2.0 (Qwen3.5-4B base)re-run on a throwaway RunPod pod with the original recipe; deviations in its manifest
4Cygnet (blockbrain, frozen Gemma-4-12B-it)61.849.5$0.037 (estimated)shim MIT; weights Apache-2.0 with Google's Gemma Prohibited Use PolicyRTX PRO 6000 (evaluator-owned Lium pod)
5Hopper59.448.0$0.024 (estimated)Component-specific terms recorded in RESULT.md; submitted adapter release and Qwen base retain their respective termsour GPU (lium.io RTX A6000 48 GB), local loopback HTTP; serial
6Winnow-12B Q855.648.3$0.037 (estimated)Apache-2.0, including the applicable Gemma 4 base/derivative licence termsour GPU (lium.io RTX 4090 24 GB), reached over the internet from Germany; serial, one request at a time
7reflex 4B (kshetrajna12)54.047.5$0.022 (estimated)MIT (code, adapter); Apache-2.0 (base)our RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany
8djev (Maisa, diffusion-gemma)52.247.0$0.026 (announced price)Apache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weightsproduction API (api.djev.dev, free preview)
9Jev-Omni (akhilaaa3, Gemma-4-12B merged)51.346.8$0.037 (estimated)Apache-2.0, following Gemma 4; dataset rights stated separately by the authorour RunPod GPU (L40 48 GB, Czechia), reached over the internet from Germany
10metask-jev-4b47.844.7$0.033 (estimated)Apache-2.0our GPU (lium.io RTX 5090 32 GB); serial, in-process candidate-logit inference
11SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)47.744.4$0.022 (estimated)MIT (code); Qwen3.5 weights Apache-2.0RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1)

All 60 open-weight rows, searchable and sortable

What you give up and gain by self-hosting

Jev was measured through its production API from a server in Germany, network included, at a measured $0.040 per 1,000 decisions. Most open models were measured on GPUs we rented or on their authors’ endpoints; rows on our own hardware carry an assumed speed adjustment, and their cost is usually an estimate from a comparable hosted price. Self-hosting keeps decisions on your own hardware, which the API cannot.

Read the license, not just the label

“Open” here means the published row lists public code or weights. Terms differ: some rows are Apache-2.0 or MIT, others inherit a base model’s use policy, and a few are non-commercial research licenses. Each row keeps its own license text.

Head-to-head: Jev vs decider-4b v2 · Jev vs Cygnet · Jev vs Hopper · Jev vs Laya

Frequently asked questions

Is Jev open source?
No. Jev 1.13.0 by TypeSafe AI is listed in JevBench v1.4.2 as “proprietary API”: it is served through TypeSafe AI’s production API and its weights are not published. JevBench’s harness and public tasks are open source (MIT), but that does not make Jev itself open.
What is the best open-source Jev alternative?
On the v1.4.2 composite, decider-4b v2 (Mapika) ranks #1 with a JevBench Score of 64.1 (Apache-2.0 (package and weights)). Jev 1.13.0 out-reasons decider-4b v2 (Intelligence 53.1 vs 49.4) and is better calibrated; decider-4b v2 leads on speed and cost. JevBench weighs the four axes equally; sort by Intelligence for raw reasoning. The right choice depends on whether you weight reasoning, calibration, speed or cost most.
How many open-weight Jev-class models does JevBench measure?
60 Jev-style systems with public code or weights are ranked in v1.4.2. One of them ranks above Jev 1.13.0 (#2) on the composite score.
Can I run an open Jev alternative on a CPU?
Yes, some were measured on four CPU threads of an AMD Ryzen 5 3600. The highest-ranked of them is jeff (Logan Markewich, GLiFormer 400M) at #40 (JevBench Score 30.6). Most higher-ranked open models were measured on a rented GPU or the author’s own endpoint.
Are the costs of open models measured?
Mostly not. A self-hosted model has no bill of its own, so its cost is usually estimated from a comparable hosted model’s price and labeled as an estimate. Rows served on our own hardware also carry a published speed adjustment (such as ×2 + 0.15 s), which is an assumption rather than a measurement.