JevBench by Benchmark Heaven · Decision model evaluation
How to measure calibration and reliability in decision models
Measure correctness and probability quality on held-out decisions, then test the confidence thresholds you will use in production. Calibration means probabilities agree with observed outcomes; reliability also requires failure handling, robustness and acceptable error costs. JevBench’s Calibration axis is one part of this evaluation.
Reviewed · Maintained by Benchmark Heaven, independently of TypeSafe AI.
Keep the measurements separate
Report accuracy or task-appropriate class metrics alongside a proper probability score such as Brier score or negative log likelihood. Inspect reliability plots and calibration error by domain and confidence range; report sample counts because small bins are noisy. A high-confidence correct answer is useful evidence, but accuracy alone does not establish calibration.
JevBench publishes its versioned distribution-quality Calibration axis. That normalized score is not interchangeable with another benchmark's ECE, NLL or Brier score. Read the scorer and method before interpreting it.
Evaluate abstention and failure cost
At each proposed confidence threshold, measure the fraction of decisions accepted (coverage) and their error rate (selective risk). Count refusals, invalid distributions, timeouts and unsupported classes. Test routing to a stronger model or a human with your own costs for false positives, false negatives and escalations. Select thresholds on validation data, then evaluate once on a separate test set.
Add coherence and robustness checks
Check that probabilities follow option reordering, negation and equivalent wording. Internal coherence does not imply correct answers: a model can be consistently wrong. The independent metamorphic coherence JevBench complements our accuracy and serving measurements. ForecastBench is relevant when the decisions are predictions about future outcomes.
Retest under realistic class imbalance, new domains and input changes. Record model pins, dataset versions, measurement dates, latency percentiles and cost assumptions. A benchmark shortlist helps choose candidates; it cannot certify your workload.