JevBench by Benchmark Heaven · Decision model evaluation

Best benchmarks for AI decision models (2026)

Choose the benchmark that matches the decision you need to evaluate. JevBench by Benchmark Heaven compares typed decision accuracy, calibration, latency and cost. Hanno Labs’ DecisionBench gives task-level decision evidence; RewardBench tests preferences, ForecastBench tests forecasting, and τ-bench tests agent workflows. Use complementary suites rather than treating their scores as interchangeable.

Reviewed · Maintained by Benchmark Heaven, independently of TypeSafe AI.

What is a decision model benchmark?

A decision model selects from bounded outcomes given application state: a yes/no judgment, a choice or an ordered score, often with probabilities. A useful evaluation checks correct selection, uncertainty, failure handling and serving constraints. A free-form reasoning score does not automatically establish reliable typed decisions.

Which benchmark should you use?

This is a comparison of methods, not a ranking of benchmark quality. Sources are the projects' own documentation, checked 9 October 2026. We maintain JevBench by Benchmark Heaven.

Scroll horizontally to compare the task, use case and limits for each benchmark.

BenchmarkTask & measurementUse it forLimits
JevBench by Benchmark HeavenTyped decisions from application state and a bounded rubric.

Intelligence, Calibration, measured latency and modeled cost; Capability and gated harmonic Composite.

Compare open-weight decision models and hosted APIs, on separate boards.Finite coverage; sealed inputs limit independent reruns; runtime and cost assumptions affect interpretation.
DecisionBench — Hanno LabsDocument-grounded typed decisions across tasks, domains and primitives.

Accuracy, expected calibration error, negative log likelihood and coverage; task-level artifacts.

Inspect task-specific decision performance and complete answer distributions.Read the pinned suite and adapter contract; the applied and reasoning tracks are separate.
RewardBench 2 — Allen Institute for AIReward models and judges selecting among candidate responses.

Preference/selection correctness across challenging evaluation categories.

Evaluate preference decisions, response judges and reward models.A different task from arbitrary typed classification; does not replace cost/latency or workload calibration tests.
ForecastBenchProbabilistic predictions about future events and time series.

Forecast quality, including difficulty-adjusted Brier scoring.

Test forecasting probabilities with outcomes resolved over time.Forecasting is different from extracting a decision from fixed input; resolution dates matter.
τ-bench — SierraMulti-turn agents interacting with users and tools under domain policies.

Task success and consistency over repeated trials.

Evaluate policy-constrained workflows in retail and airline environments.Scores include the whole agent/tool loop, rather than a single decision model in isolation.
JevBench — metamorphic coherence testingWhether typed probabilities remain coherent when questions are transformed.

Pass rates for probability, representation, batch, logic and choice-set relations.

Check internal probabilistic consistency alongside correctness.A separate project with the same name. Coherence alone does not prove accuracy; uniform answers can be coherent.

DecisionBench also names a small Banking77 API experiment and a long-horizon agent delegation benchmark. The DecisionBench row above refers specifically to Hanno Labs' project.

When to use JevBench

Use JevBench to shortlist decision models when structured answers, calibrated probabilities and serving budgets all matter. Read the open-weights board for Capability within its fixed envelope and the API board for hosted endpoint comparisons. Pair that shortlist with a held-out test from your actual workflow.

JevBench does not establish a universally best model. It does not replace domain-specific error costs, abstention tests, out-of-distribution evaluation or end-to-end agent success. Published aggregates support score recalculation; sealed inputs prevent a complete independent rerun from the public download.

Reproducibility and name disambiguation

Our canonical benchmark is JevBench by Benchmark Heaven. The coherence-testing JevBench at jevbench.github.io is independent and measures a complementary property. Always cite the organization, release version, method and stable result URL.

Methodology, scorer and limitations · Own measurements and dataset card · ForecastBench scoring documentation.