JevBench by Benchmark Heaven · released v1.4.1
Jev alternatives, compared on the published board
The best Jev alternative depends on your use case. Compare Jev-class decision models using the same released benchmark: Intelligence, Calibration, Speed and Cost, with each row’s evidence and openness notes.
JevBench Scores in the current top five
The Jev row is the reference; the four rows below it are current alternatives. The overall score is a composite, so check the separate axes for your use case.
- 1Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
- 2JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
- 3Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
- 4Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
- 5reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
All top-five values as a table
| Rank | System | Score | Intelligence | Calibration | Speed | Cost | Cost basis |
|---|---|---|---|---|---|---|---|
| 1 | Jev 1.13.0 (TypeSafe AI) Reference system | 63.3 | 53.1 | 76.3 | 83.3 | 52.0 | measured |
| 2 | JevK5 v0.2.0 | 62.0 | 48.9 | 74.5 | 91.1 | 59.5 | estimated |
| 3 | Hopper | 59.4 | 48.0 | 79.1 | 86.8 | 58.7 | estimated |
| 4 | Winnow-12B Q8 | 55.6 | 48.3 | 64.8 | 82.3 | 52.9 | estimated |
| 5 | reflex 4B (kshetrajna12) | 54.0 | 47.5 | 70.4 | 68.0 | 59.7 | estimated |
Cost evidence is labeled by basis so estimated and announced values are not presented as measured charges.
Looking for an open source Jev alternative?
These rows have explicit openness, license and repository fields in the published artifact. “Open” is the board’s status; read the linked source and exact terms before deploying a system.
Hopper · rank 3
Component-specific terms recorded in RESULT.md; submitted adapter release and Qwen base retain their respective terms
Published repository or weights
Sealed-set accuracy: 34.1% · 86.8 Speed axis
Winnow-12B Q8 · rank 4
Apache-2.0, including the applicable Gemma 4 base/derivative licence terms
Published repository or weights
Sealed-set accuracy: 33.1% · 82.3 Speed axis
reflex 4B (kshetrajna12) · rank 5
MIT (code, adapter); Apache-2.0 (base)
Published repository or weights
Sealed-set accuracy: 28.2% · 68.0 Speed axis
djev (Maisa, diffusion-gemma) · rank 6
Apache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weights
Published repository or weights
Sealed-set accuracy: 29.9% · 91.4 Speed axis
Jev-Omni (akhilaaa3, Gemma-4-12B merged) · rank 7
Apache-2.0, following Gemma 4; dataset rights stated separately by the author
Published repository or weights
Sealed-set accuracy: 32.1% · 81.5 Speed axis
metask-jev-4b · rank 8
Apache-2.0
Published repository or weights
Sealed-set accuracy: 27.6% · 89.1 Speed axis
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · rank 9
MIT (code); Qwen3.5 weights Apache-2.0
Published repository or weights
Sealed-set accuracy: 26.3% · 83.7 Speed axis
Jobe Qwen3.5-4B (frozen) · rank 10
MIT code; Apache-2.0 weights
Published repository or weights
Sealed-set accuracy: 25.6% · 85.6 Speed axis
For a broader list, use the board’s row disclosures. Rows without explicit openness, license and repository evidence are not classified here.
See the use-case chooser, including accuracy, speed, cost and self-hosting evidence.
Frequently asked questions
- What does JevBench compare?
- JevBench compares published Jev-class decision systems across Intelligence, Calibration, Speed and Cost. This page uses the released v1.4.1 aggregate.
- How should I choose between Jev alternatives?
- There is no single best fit for every use. Compare the published axes and their evidence, then choose for your use case. Intelligence, Calibration, Speed and Cost — equal-weight harmonic mean, with generalization and Jev-class gates.
- Which Jev alternatives have public self-hosting evidence?
- The list below includes only rows marked as open or open weights in the published artifact that also include a license note and public repository link. Check the linked source and its terms before deployment; missing evidence remains unknown.
- Are cost and speed values directly measured?
- The board labels cost evidence as measured, estimated or announced and shows the conditions behind speed measurements where available. The Cost and Speed axes are benchmark scores, not a universal bill or wall-clock guarantee.