Audio benchmark · v0.1 · 29 Sep 2026
AudioJevBench v0.1
Which audio models can make the decisions a voice agent needs — straight from the caller's audio, with honest probabilities, fast and cheaply?
17 systems · 988 scored items (112 public + 876 sealed) · intent, escalation, command safety, sentiment, urgency, speaker verification and sound events · only system-level aggregates are published; sealed items stay private.
Text decisions: JevBench · Image decisions: Image JevBench · How it is measured
Method notes
What it measures. Each item is a 16 kHz mono clip plus a short operator context and a typed question. The decisive evidence is only in the audio. The system answers with a probability for every allowed label: a choice among named options, a yes/no, or a 0–4 score — the same contract as JevBench. Decision families: speech intent routing, escalation to a human, command safety, caller sentiment and urgency, overlapping and background talkers, speaker verification (enrolment sample vs caller) and sound events (alarms, horns, animals, household sounds), under clean and noisy conditions (15 to 0 dB SNR) and in 13 languages and 12 English accents.
Public and sealed. 112 public items (authored separately, commercially usable voices) and 876 sealed items are scored; a pre-registered audio-quality gate removed a further 6 public and 18 sealed items for every system alike. Sealed items are never published and go only to systems we run ourselves, on our own rented GPUs (offline container, no gold, pod removed afterwards) or on this server. Intelligence blends 25% public / 75% sealed (the public set is about an eighth of the items); a public-minus-sealed gap beyond 15 points above the field median costs Intelligence.
Hosted APIs are public-only. Sealed items are never sent to third-party APIs: no enforced zero-retention route was available for these providers. Gemini and OpenAI models are therefore scored on the public items only, shown as provisional with a 95% bootstrap interval, and never ranked with the full-coverage systems. Their numbers can move once a sealed-safe route exists.
Partial coverage. CLAP and AST are sound classifiers: they answer only the sound-event family (choice), so their Intelligence covers that family alone. Their label softmax is not treated as calibrated, so Calibration is 0.
Axes and scores. Four axes, 0–100: Intelligence (tier- and type-weighted accuracy, chance-corrected; missing or invalid answers count wrong), Calibration (are the probabilities honest), Speed (100 − 20·log10(latency / 0.1 s), mean of p50 and p95; self-hosted latency gets the JevBench ×2 + 0.15 s adjustment) and Cost (100 − 30·log10(USD per 1,000 / 0.001)). Capability = mean(Intelligence, Calibration). The composite (option B) is a weighted harmonic mean 40:20:20:20 with the JevBench v1.5 gates.
Jev-class gate. Adjusted median latency ≤ 1.0 s and ≤ $0.25 per 1,000 decisions, fixed before any result because there is no audio Jev reference model. It decides who gets a Capability rank number; the composite ranks every full-coverage system.
Pricing basis. Hosted rows use the provider's standard list price in force for at least 30 days, applied to the measured audio and text tokens; promotional prices with an end date do not count, so Gemini 3.8 Flash is priced at its announced regular rate. Self-hosted rows (marked *) are estimates: the exact model's public catalogue price if one exists, otherwise its base model's or the nearest same-family size (erring high), applied to the tokens counted by the model's own processor. Classifiers use the JevBench size-class rule for small encoders.
Limits. Speech clips are synthetic (text-to-speech voices with added real and synthetic noise); sound events are real recordings. One run per system at temperature 0. A system at Intelligence 0 answered at or below chance after correction, often by refusing or returning invalid output. Results describe these exact configurations.
AudioJevBench is a separate benchmark from the text-only JevBench and from Image JevBench. Sealed item content stays private; this page shows system-level aggregates only.