JevBench by Benchmark Heaven · Decision model evaluation

Decision model benchmark methodology — JevBench v1.6.1

JevBench measures typed decisions: application state and a bounded rubric go in; a structured answer and, where supported, probabilities come out. Compare accuracy and calibration alongside measured latency and modeled cost. The open-weights board leads with Capability; the hosted API board leads with Composite.

Reviewed · Maintained by Benchmark Heaven, independently of TypeSafe AI.

Release and measurement dates

Published release v1.6.1, initially published . The versioned leaderboard retains its release context. Additions stored in the results file: a6 (2026-10-07), a7 (2026-10-07), a8 (2026-10-07), a11 (2026-10-08), a12 (2026-10-08); amended through 2026-10-08; latest row measurement 2026-10-07. Cite source_sha256 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1 for these exact result bytes. The live board may also include later dated API reruns and presentation revisions.

The release draw contains 1,200 sealed and 300 public decisions. Self-hosted systems answer 1,500 decisions. The v1.6.1 full-set API rows answer the same set. Older subsets and later reruns are explicitly labeled and may be equated. Missing or refused answers are counted by the scorer, rather than silently removed.

Capability and Composite are different scores

Capability: Arithmetic mean of Intelligence and Calibration; ranked models within both official caps. Headline on the open-weights board. The cost cap is USD 0.0646 per 1,000 decisions, and the adjusted median-latency cap is 1.23 seconds. Eligibility is separate from the raw score.

Composite: Weighted harmonic mean of Intelligence, Calibration, Speed and Cost, multiplied by squared penalties when Intelligence is below its configured floor, or Speed or Cost is below 50. Equal 25% weights do not mean an arithmetic average.

Official option A uses weights Calibration 25%, Cost 25%, Intelligence 25%, Speed 25%; Intelligence floor 50. A zero axis gives zero Composite; a missing axis gives no score.

H = Σw / Σ(w / axis). Composite = H × min(1, I / 50)² × min(1, Speed / 50)² × min(1, Cost / 50)².

Intelligence is chance-corrected decision accuracy, with tier weighting and a public/sealed gap adjustment. Calibration evaluates probability distributions; it is not simply accuracy among confident answers. Noul handling, per-type normalization and equating are versioned in the scorer. The four normalized axes are on a 0–100 scale; actual seconds and USD remain separate values.

Reproducibility and model versions

Scorer and public benchmark repository · Board source · Release manifest and scorer hashes · Published aggregate artifact.

Artifact SHA-256: 5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1. The data card documents model pins, measurement dates, price bases and latency adjustments. Unknown pins and dates remain null; hosted service names do not guarantee immutable weights.

Sealed questions, identifiers and predictions are not downloadable. Public-split experiments can be reproduced, and aggregate Composite scores can be recalculated from the published axes. Complete independent reruns of the sealed evaluation require the official evaluation process.

Limitations and historical methods

Coverage is finite and uneven across languages and domains. Cost is modeled from a stated hardware or list-price basis; latency depends on the measured runtime and location. Confidence, coherence, robustness and real-world error costs need additional workload-specific checks. Training on the public split must be disclosed, and a sealed split is not a guarantee of no prior exposure.

Historical releases retain their original scoring definitions and measurement sets. Do not compare scores across method versions as if they were the same experiment. ImageJevBench v0.3.0 has a separate image pool and frozen eligibility envelope; core and computer-use tracks remain separate. Its reserve size is not the number of decisions answered by each model.

Another project named JevBench tests metamorphic coherence. It is independent of Benchmark Heaven and measures consistency rather than this board's accuracy/cost/latency score.