JevBench by Benchmark Heaven · released v1.4.1

How to choose a Jev-class model

Start with the use case, then read the metric behind each recommendation. The guide uses the latest published JevBench artifact; it does not treat the overall composite as raw accuracy.

Most accurate (sealed-set accuracy)

GPT-6 Luna (default medium reasoning effort)

95.5%

Published sealed-set accuracy in v1.4.1. The JevBench Score is a separate composite.

Fastest (Speed axis)

Certo v1 (AltSlate Labs)

94.0/ 100 benchmark score

This is the published Speed axis. Use the row’s latency conditions when estimating real performance; it is not a universal wall-clock guarantee.

Cheapest per decision (Cost axis)

Certo v1 (AltSlate Labs)

100.0/ 100 benchmark score

Published cost: $0.0010 per 1,000 decisions (estimated)

Cost axis is a comparison score. The row’s cost basis distinguishes measured, estimated and announced values.

Self-hosting evidence

Use rows with public openness, license and repository evidence

The list below is ordered by the published JevBench rank. The openness status and license terms are copied from each public aggregate row.

Systems with explicit self-hosting evidence

Rows without explicit openness, license or repository evidence are left unclassified. Confirm each model’s terms and requirements before using it in production.

Keep the measures separate

Intelligence, Calibration, Speed and Cost — equal-weight harmonic mean, with generalization and Jev-class gates. Accuracy, speed, cost and the composite answer different questions. Review the full live board and its row disclosures before making a deployment choice.

Frequently asked questions

Which Jev-class model has the highest published accuracy?
GPT-6 Luna (default medium reasoning effort) has the highest sealed-set accuracy in the published v1.4.1 ranked rows (95.5%). This is a separate measure from the composite JevBench Score.
Which Jev-class model is fastest?
Certo v1 (AltSlate Labs) has the highest published Speed axis score (94.0). The axis is a benchmark score; it does not promise the same wall-clock latency on every host or workload.
Which Jev-class model is cheapest?
Certo v1 (AltSlate Labs) has the highest Cost axis score (100.0). Its published cost is estimated at $0.0010 per 1,000 decisions; check the board’s row disclosure for the basis before comparing bills.
How do I find a self-hostable Jev alternative?
Use only systems marked open or open weights in the published row, with a license note and public repository link. This page leaves absent openness, license or repository evidence unclassified; review the linked terms before deployment.