JevBench by Benchmark Heaven · Decision model evaluation

JevBench data card & downloadable results

Download only Benchmark Heaven’s own published JevBench and ImageJevBench system-level measurements. JevBench JSON and CSV identify the named release, and JSON includes the exact source hash; they contain no Artificial Analysis numbers, sealed questions or per-item predictions.

Reviewed · Maintained by Benchmark Heaven, independently of TypeSafe AI.

Versioned downloads

JevBench v1.6.1

Initially published 2026-10-06 · JSON · CSV

Source SHA-256: 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1

ImageJevBench v0.3.0

Artifact built 2026-10-07 · JSON · CSV

Source SHA-256: 1f5014f77a9945c2098f50460872a3d5c193b3751a3fb44543bcaf0931365ffc

The live board may also include later dated API reruns and presentation revisions. Image exports contain the core track, with unrankable measurements labeled. Dated carries and computer-use measurements are separate and never silently blended.

Schema, missing values and units

One row per measured system. JSON also contains schema_version, version, source_sha256 and method definitions; JevBench includes published_at and ImageJevBench includes built_at. CSV uses the same row values, repeats the release version, quotes text fields and leaves missing values blank.

FieldsMeaning
key, name, source_urlStable system key, published name and model/source URL.
model_pin, last_measured_onPublished model pin (commit) or the API model identifier we requested, the provider reported or the author declared, and the measurement date where available; API identifiers may change. Missing values are null. ImageJevBench exports have no model_pin field.
capability, composite, intelligence, calibration, speed, cost_axisNormalized 0–100 scores. The method explains the distinction between Capability and Composite.
ranked, capability_eligible, open_capability_rankAdmission, cap eligibility and scoped open-board rank; raw Capability alone is not a rank. Image admission uses ranked and not_ranked_because.
usd_per_1000_decisions, price_kind, price_basisJevBench modeled USD per 1,000 complete decisions and its evidence basis; not an invoice.
p50_raw_s, p50_adjusted_s, latency_adjustmentJevBench median latency in seconds and the published adjustment. Image exports retain normalized axes, not invented raw seconds or USD.

Null means unavailable, never zero. Reuse of these aggregate exports is under CC BY 4.0; cite Benchmark Heaven, the exact release, access date and source hash. Original model weights, task sources and third-party data keep their own licenses.

Scope and limitations

JevBench evaluates typed text decisions on public and sealed splits. ImageJevBench uses an independent image pool. These samples do not guarantee coverage of your domain; low-count slices, runtime differences, dated carries and admission holds must be read with the result. Full sealed reruns cannot be independently reconstructed from public aggregates.

Hugging Face results dataset · Static Hugging Face leaderboard · Release manifest.