Audio benchmark · v0.1 · 29 Sep 2026

AudioJevBench v0.1

Which audio models can make the decisions a voice agent needs — straight from the caller's audio, with honest probabilities, fast and cheaply?

17 systems · 988 scored items (112 public + 876 sealed) · intent, escalation, command safety, sentiment, urgency, speaker verification and sound events · only system-level aggregates are published; sealed items stay private.

Text decisions: JevBench · Image decisions: Image JevBench · How it is measured

Capability ranking

Capability is the mean of Intelligence and Calibration. Qwen3-Omni-30B-A3B-Instruct leads the Jev-class audio systems with 83.2.

Jev-class: median latency ≤ 1.0 s and ≤ $0.25 per 1,000 decisions. 7 of 9 full-coverage systems meet the ranking gate. 1 additional provisional or partial row meets the speed/cost limits but is not ranked.

Ranked · full coverage (sealed + public items)

  1. 1Qwen3-Omni-30B-A3B-Instructself-hosted GPU83.2
  2. 2Qwen2.5-Omni-7Bself-hosted GPU68.1
  3. 3Qwen2.5-Omni-3Bself-hosted GPU53.8
  4. 4Voxtral Mini 3B (2507)self-hosted GPU47.8
  5. 5Ultravox v0.5 (Llama 3.2 1B)self-hosted GPU31.4
  6. 6MOSS-Audio-4B-Instructself-hosted GPU27.9
  7. 7Qwen2-Audio-7B-Instructself-hosted GPU12.8

Not ranked · provisional and partial rows

Hosted APIs are scored on the 112 public items only, because sealed items never go to third-party APIs; the whisker is the 95% interval of their Capability. Sound classifiers answer only the sound-event family. Neither is comparable with the ranking above. Shown regardless of the Jev-class filter.

  1. –Gemini 3.8 Flashpublic-onlyGemini API · public items only94.4Outside Jev-class: median latency 1.65 s > 1.0 s; $1.5 > $0.25 per 1,000
  2. –Gemini 3.5 Flashpublic-onlyGemini API · public items only83.1Outside Jev-class: median latency 1.27 s > 1.0 s; $1.1 > $0.25 per 1,000
  3. –Gemini 3.1 Flash-Litepublic-onlyGemini API · public items only72.1Outside Jev-class: median latency 1.27 s > 1.0 s
  4. –GPT Audiopublic-onlyOpenAI API · public items only69.3Outside Jev-class: $4.9 > $0.25 per 1,000
  5. –GPT Audio Minipublic-onlyOpenAI API · public items only66.3Outside Jev-class: $1.2 > $0.25 per 1,000
  6. –Gemini 3.5 Flash-Litepublic-onlyGemini API · public items only65.6Outside Jev-class: median latency 1.32 s > 1.0 s
  7. –CLAP (laion/clap-htsat-fused)partiallocal CPU · sound events only47.5
  8. –AST (MIT AudioSet fine-tune)partiallocal CPU · sound events only29.9Outside Jev-class: median latency 3.94 s > 1.0 s
  • Full coverage: sealed + public items
  • Hosted API, public items only (provisional)
  • Partial coverage (sound events only)

# numbers only Jev-class systems with full coverage. Right column: USD per 1,000 decisions (* = estimated from a catalogue price; see method) and adjusted median latency.

Capability against speed and cost

The two charts are linked: pointing at a system highlights it in both. The five most capable ranked Jev-class systems are labelled. Jev-class filter on; untick it above to add every system.

Capability vs speed

Median latency (adjusted p50, the Jev-class measure), log scale, faster to the right. The dashed line is the Jev-class limit (1.0 s). Upper right is better.

020406080100100 ms1.0 s300 ms3.0 sJev-class limit →← slower · median latency · faster →Capability ↑Qwen3-Omni-30B-A3B-Instr…Qwen2.5-Omni-7BQwen2.5-Omni-3BVoxtral Mini 3B (2507)Ultravox v0.5 (Llama 3.2…
8 systems. Hover, focus or tap a bubble; both charts highlight the same system.
Capability vs cost

USD per 1,000 decisions, log scale, cheaper to the right. The dashed line is the Jev-class limit ($0.25). Upper right is better.

020406080100$0.0010$0.010$0.10Jev-class limit →← pricier · USD per 1,000 decisions · cheaper →Capability ↑Qwen3-Omni-30B-A3B-Instr…Qwen2.5-Omni-7BQwen2.5-Omni-3BVoxtral Mini 3B (2507)Ultravox v0.5 (Llama 3.2…
8 systems. Hover, focus or tap a bubble; both charts highlight the same system.
  • Full coverage: sealed + public items
  • Hosted API, public items only (provisional)
  • Partial coverage (sound events only)

Composite score

The official AudioJevBench score (option B): a weighted harmonic mean of Intelligence 40, Calibration 20, Speed 20 and Cost 20, with the JevBench v1.5 low-axis gates. Every full-coverage system is ranked here, including those outside Jev-class; a weak axis pulls the score down hard, and the row says which one.

  1. 1Qwen3-Omni-30B-A3B-Instruct70.4
  2. 2Qwen2.5-Omni-7B52.1
  3. 3Qwen2.5-Omni-3B48.2
  4. 4Voxtral Mini 3B (2507)36.6
  5. 5Whisper large-v3-turbo → Qwen3-4B-Instruct-250719.3
  6. 6MOSS-Audio-4B-Instruct12.7
  7. 7AudioJev (mocomoco edge ONNX)0.0Intelligence 0: at or below chance after chance correction
  8. 8Qwen2-Audio-7B-Instruct0.0Intelligence 0: at or below chance after chance correction
  9. 9Ultravox v0.5 (Llama 3.2 1B)0.0Intelligence 0: at or below chance after chance correction

Not ranked · provisional and partial rows

Same formula on what these systems could be measured on; not comparable with the ranking above.

  1. –Gemini 3.5 Flash-Lite19.4
  2. –Gemini 3.1 Flash-Lite19.1
  3. –Gemini 3.5 Flash1.1gated by Cost ($1.07 per 1,000 decisions)
  4. –GPT Audio Mini0.5gated by Cost ($1.24 per 1,000 decisions)
  5. –Gemini 3.8 Flash0.1gated by Cost ($1.53 per 1,000 decisions)
  6. –AST (MIT AudioSet fine-tune)0.0no calibrated probabilities (Calibration 0) and only one family covered
  7. –CLAP (laion/clap-htsat-fused)0.0no calibrated probabilities (Calibration 0) and only one family covered
  8. –GPT Audio0.0gated by Cost ($4.94 per 1,000 decisions)

Direct comparison

Pick two systems and compare every published aggregate side by side. The better value in each row is bold.

Qwen3-Omni-30B-A3B-Instruct

self-hosted GPU · full coverage · composite #1

Qwen2.5-Omni-7B

self-hosted GPU · full coverage · composite #2

Head-to-head radar

Up to four 0–100 axes; farther from the centre is better. Public-only rows use public Intelligence, and unmeasured axes are omitted.

Head-to-head radar: Qwen3-Omni-30B-A3B-Instruct and Qwen2.5-Omni-7BIntelligenceCalibrationSpeedCost
Qwen3-Omni-30B-A3B-InstructQwen2.5-Omni-7B
Composite score (B)
70.452.1
Capability
83.268.1
Intelligence
81.057.8
Public Intelligence
86.958.2
Sealed Intelligence
79.157.6
Calibration
85.478.5
Speed axis
82.480.7
Cost axis
49.246.1
Median latency (adj.)
660 ms676 ms
USD per 1,000 decisions
$0.049$0.063

Full table

All 17 measured systems in three labelled groups. Only the full-coverage group is ranked. Public-only rows show their public Intelligence with a stratified 95% bootstrap interval. Scroll sideways on small screens.

#SystemComposite (B)CapabilityIntelligencePublic I (95% CI)Sealed IGap · penaltyCalibrationSpeedCostp50 / p95 adj.$/1kItemsJev-class
Full coverage (sealed + public) · 9 · ranked by composite
1Qwen3-Omni-30B-A3B-Instructself-hosted GPU70.483.281.086.979.17.7 · ×1.0085.482.449.2660 ms / 880 ms$0.049*988yes
2Qwen2.5-Omni-7Bself-hosted GPU52.168.157.858.257.60.6 · ×1.0078.580.746.1676 ms / 1.26 s$0.063*988yes
3Qwen2.5-Omni-3Bself-hosted GPU48.253.850.150.849.90.9 · ×1.0057.483.846.7522 ms / 794 ms$0.060*988yes
4Voxtral Mini 3B (2507)self-hosted GPU36.647.847.241.749.0-7.3 · ×1.0048.485.444.7383 ms / 753 ms$0.070*988yes
5Whisper large-v3-turbo → Qwen3-4B-Instruct-2507cascade: transcript, then text model · self-hosted GPU19.357.648.244.649.4-4.7 · ×1.0067.161.432.88.04 s / 8.97 s$0.17*988no
6MOSS-Audio-4B-Instructself-hosted GPU12.727.931.829.832.4-2.7 · ×1.0023.985.446.5396 ms / 731 ms$0.061*988yes
7AudioJev (mocomoco edge ONNX)local CPU0.00.00.0-55.2-55.50.3 · ×1.000.073.877.11.65 s / 2.52 s$0.0058*988no
8Qwen2-Audio-7B-Instructself-hosted GPU0.012.80.0-42.4-49.57.1 · ×1.0025.681.838.6653 ms / 1.00 s$0.11*988yes
9Ultravox v0.5 (Llama 3.2 1B)self-hosted GPU0.031.40.0-44.0-29.3-14.7 · ×1.0062.991.161.4224 ms / 347 ms$0.019*988yes
Public-only, provisional · 6 · not ranked
–Gemini 3.8 FlashAPIGemini API · public items only0.194.4—97.0 [91.9, 99.9]——91.974.14.51.65 s / 2.37 s$1.53112no
–Gemini 3.5 FlashAPIGemini API · public items only1.183.1—87.2 [78.9, 94.5]——79.076.79.11.27 s / 1.69 s$1.07112no
–Gemini 3.1 Flash-LiteAPIGemini API · public items only19.172.1—80.1 [70.9, 88.2]——64.077.229.01.27 s / 1.51 s$0.23112no
–Gemini 3.5 Flash-LiteAPIGemini API · public items only19.465.6—73.5 [62.6, 83.9]——57.876.529.71.32 s / 1.69 s$0.22112no
–GPT Audio MiniAPIOpenAI API · public items only0.566.3—68.6 [57.9, 79.6]——64.079.87.2844 ms / 1.24 s$1.24112no
–GPT AudioAPIOpenAI API · public items only0.069.3—67.3 [56.1, 77.9]——71.378.30.0947 ms / 1.55 s$4.94112no
Partial coverage · 2 · not ranked
–CLAP (laion/clap-htsat-fused)local CPU · sound events onlycovers: sound event0.047.595.1100.093.46.6 · ×1.000.081.796.3613 ms / 1.11 s$0.0013*95yes
–AST (MIT AudioSet fine-tune)local CPU · sound events onlycovers: sound event0.029.959.837.567.2-29.7 · ×1.000.068.198.53.94 s / 3.96 s$0.0011*95no
Robustness diagnostics (not scored)

Accuracy per condition slice (Score items count 1 − normalised error). Diagnostics only — not part of any axis. Small n means wide uncertainty; public-only rows cover public items only.

SystemClean audioNoisy (≤ 5 dB SNR)15 dB10 dB5 dB0 dBEnglishNon-EnglishBackground talkerOverlapping talkersSpeaker verificationSound eventsNoise drop (clean − ≤5 dB)Language gap (EN − non-EN)
Qwen3-Omni-30B-A3B-Instructself-hosted GPU87% n=69189% n=12592% n=8293% n=8990% n=8983% n=2389% n=63388% n=35584% n=11274% n=5864% n=11690% n=179-1 pts+1 pts
Qwen2.5-Omni-7Bself-hosted GPU77% n=69176% n=12583% n=8285% n=8974% n=8979% n=2380% n=63375% n=35574% n=11276% n=5851% n=11690% n=179+1 pts+6 pts
Qwen2.5-Omni-3Bself-hosted GPU74% n=69178% n=12571% n=8274% n=8975% n=8986% n=2375% n=63373% n=35571% n=11266% n=5849% n=11682% n=179-3 pts+2 pts
Voxtral Mini 3B (2507)self-hosted GPU70% n=69172% n=12584% n=8277% n=8973% n=8982% n=2369% n=63377% n=35577% n=11269% n=5849% n=11650% n=179-2 pts-8 pts
Whisper large-v3-turbo → Qwen3-4B-Instruct-2507cascade: transcript, then text model · self-hosted GPU65% n=69175% n=12582% n=8282% n=8978% n=8982% n=2365% n=63376% n=35578% n=11274% n=5822% n=11637% n=179-10 pts-11 pts
MOSS-Audio-4B-Instructself-hosted GPU63% n=69159% n=12566% n=8261% n=8954% n=8975% n=2365% n=63358% n=35558% n=11259% n=5850% n=11667% n=179+5 pts+7 pts
AudioJev (mocomoco edge ONNX)local CPU9% n=6915% n=1258% n=8212% n=894% n=895% n=239% n=6338% n=3555% n=1125% n=580% n=11613% n=179+4 pts+1 pts
Qwen2-Audio-7B-Instructself-hosted GPU10% n=6916% n=12513% n=8214% n=895% n=895% n=2313% n=6335% n=3557% n=1129% n=580% n=11610% n=179+4 pts+7 pts
Ultravox v0.5 (Llama 3.2 1B)self-hosted GPU16% n=69115% n=12518% n=8228% n=8916% n=8912% n=2317% n=63318% n=35510% n=1127% n=582% n=11619% n=179+1 pts-1 pts
Gemini 3.8 FlashAPIGemini API · public items only97% n=80100% n=10100% n=10100% n=12100% n=6100% n=299% n=7297% n=40100% n=14100% n=1295% n=2096% n=24-3 pts+1 pts
Gemini 3.5 FlashAPIGemini API · public items only89% n=80100% n=1090% n=10100% n=12100% n=6100% n=292% n=7290% n=4093% n=1492% n=1265% n=2096% n=24-11 pts+2 pts
Gemini 3.1 Flash-LiteAPIGemini API · public items only82% n=80100% n=10100% n=10100% n=12100% n=6100% n=290% n=7282% n=4071% n=1467% n=1260% n=2095% n=24-18 pts+8 pts
Gemini 3.5 Flash-LiteAPIGemini API · public items only79% n=80100% n=1087% n=1092% n=12100% n=6100% n=289% n=7272% n=4079% n=1475% n=1255% n=2088% n=24-21 pts+17 pts
GPT Audio MiniAPIOpenAI API · public items only77% n=8099% n=1088% n=1080% n=1299% n=6100% n=282% n=7277% n=4064% n=1467% n=1250% n=2084% n=24-22 pts+6 pts
GPT AudioAPIOpenAI API · public items only72% n=8099% n=1089% n=1083% n=1298% n=6100% n=277% n=7277% n=4071% n=1467% n=1215% n=2095% n=24-27 pts-0 pts
CLAP (laion/clap-htsat-fused)local CPU · sound events only95% n=8091% n=11—100% n=388% n=8—95% n=95————95% n=95+4 pts—
AST (MIT AudioSet fine-tune)local CPU · sound events only85% n=8082% n=11—100% n=388% n=8—85% n=95————85% n=95+3 pts—

Public example items

Public items as the system receives them: an operator context, a typed question and the allowed labels, plus the audio clip (not embedded here). Gold is marked. Sealed items are never shown.

sound event · choice · 9.5 s clip

Context: You are the sound-event classifier inside the yard monitor of Redfern Smallholding, used at dawn. Label the main acoustic event in the clip.

Which event best describes the main sound in the clip?

  • car_alarm — a car alarm going off
  • rooster — a rooster crowinggold
  • car_horn — a car horn honking
  • siren — a siren wailing (emergency vehicle or warning siren)

intent routing · choice · 9.0 s clip

Context: You are the phone assistant of Pharmacie Exemple Demo. The caller is the first voice you hear; another person near the caller may talk at the same time. Route only the caller's request.

Which service does the caller need?

  • prescription_refill — Renew or prepare a repeat prescription.gold
  • opening_hours — When the pharmacy is open.
  • speak_to_pharmacist — Medical advice about a medicine or symptom.
  • delivery_status — Where a home delivery is.

urgency · 0–4 score · 12.5 s clip

Context: You are the dispatcher assistant of Soccorso Stradale Demo, a roadside assistance club. Assign the dispatch priority defined in the question from the member's description.

Which dispatch priority applies?

  • 0 — 0 - Information only (membership, coverage questions); no dispatch.
  • 1 — 1 - Book a workshop appointment; the car is still drivable.
  • 2 — 2 - Dispatch within 4 hours: the car will not drive but is at home or in a car park or other safe place.gold
  • 3 — 3 - Dispatch within 1 hour: stranded at the side of an ordinary road, people are safe.
  • 4 — 4 - Immediate priority: stopped on a motorway or in a traffic lane, smoke or fire, or vulnerable people (children, elderly, ill) exposed to danger.

caller escalation · yes/no · 11.2 s clip

Context: You are the repairs line assistant of Willow Court Housing Demo. Policy: if there are signs of immediate danger in the caller's home (fire or smoke, a smoke alarm going off because of a real hazard, gas, sparking electrics, water reaching electrics), transfer to the emergency repairs line. Ordinary repairs are booked by you. Sounds from the caller's surroundings may be audible.

Should this call be transferred to the emergency repairs line?

  • yes — There are signs of immediate danger in the home.
  • no — It is an ordinary repair with no sign of immediate danger.gold

speaker verification · yes/no · 19.4 s clip

Context: Voice ID check for Hilltop Federal Savings. The first part of the clip, before the beep, is the enrolment recording of the account holder made at sign-up. After the beep is the current caller.

Is the current caller the same person as the enrolled account holder?

  • yes — The probe voice (after the beep) is the same speaker as the enrolment voice (before the beep).gold
  • no — The probe voice (after the beep) is a different speaker from the enrolment voice (before the beep).

sentiment · choice · 7.8 s clip

Context: You are the post-delivery survey assistant of Tiendita Sample, an online shop. Tag the customer's spoken answer with one label.

Which label best describes the customer's answer?

  • satisfied — Positive about the order or delivery.
  • neutral — Neither positive nor negative; the delivery was simply as expected.gold
  • frustrated — Negative about the order or delivery, but no statement about leaving.
  • churn_risk — Says they will stop buying here, cancel, or switch to another shop.

command safety · yes/no · 6.0 s clip

Context: You are the in-car assistant of Northwind Motors Example Co. The car is currently moving at 80 km/h. While the car is moving, you may change climate, navigation, music and phone calls, but you must not play video or open the web browser on the centre screen.

Should the assistant execute the driver's request now?

  • yes — The request is allowed while the car is moving.
  • no — The request is blocked while the car is moving.gold

Method notes

What it measures. Each item is a 16 kHz mono clip plus a short operator context and a typed question. The decisive evidence is only in the audio. The system answers with a probability for every allowed label: a choice among named options, a yes/no, or a 0–4 score — the same contract as JevBench. Decision families: speech intent routing, escalation to a human, command safety, caller sentiment and urgency, overlapping and background talkers, speaker verification (enrolment sample vs caller) and sound events (alarms, horns, animals, household sounds), under clean and noisy conditions (15 to 0 dB SNR) and in 13 languages and 12 English accents.

Public and sealed. 112 public items (authored separately, commercially usable voices) and 876 sealed items are scored; a pre-registered audio-quality gate removed a further 6 public and 18 sealed items for every system alike. Sealed items are never published and go only to systems we run ourselves, on our own rented GPUs (offline container, no gold, pod removed afterwards) or on this server. Intelligence blends 25% public / 75% sealed (the public set is about an eighth of the items); a public-minus-sealed gap beyond 15 points above the field median costs Intelligence.

Hosted APIs are public-only. Sealed items are never sent to third-party APIs: no enforced zero-retention route was available for these providers. Gemini and OpenAI models are therefore scored on the public items only, shown as provisional with a 95% bootstrap interval, and never ranked with the full-coverage systems. Their numbers can move once a sealed-safe route exists.

Partial coverage. CLAP and AST are sound classifiers: they answer only the sound-event family (choice), so their Intelligence covers that family alone. Their label softmax is not treated as calibrated, so Calibration is 0.

Axes and scores. Four axes, 0–100: Intelligence (tier- and type-weighted accuracy, chance-corrected; missing or invalid answers count wrong), Calibration (are the probabilities honest), Speed (100 − 20·log10(latency / 0.1 s), mean of p50 and p95; self-hosted latency gets the JevBench ×2 + 0.15 s adjustment) and Cost (100 − 30·log10(USD per 1,000 / 0.001)). Capability = mean(Intelligence, Calibration). The composite (option B) is a weighted harmonic mean 40:20:20:20 with the JevBench v1.5 gates.

Jev-class gate. Adjusted median latency ≤ 1.0 s and ≤ $0.25 per 1,000 decisions, fixed before any result because there is no audio Jev reference model. It decides who gets a Capability rank number; the composite ranks every full-coverage system.

Pricing basis. Hosted rows use the provider's standard list price in force for at least 30 days, applied to the measured audio and text tokens; promotional prices with an end date do not count, so Gemini 3.8 Flash is priced at its announced regular rate. Self-hosted rows (marked *) are estimates: the exact model's public catalogue price if one exists, otherwise its base model's or the nearest same-family size (erring high), applied to the tokens counted by the model's own processor. Classifiers use the JevBench size-class rule for small encoders.

Limits. Speech clips are synthetic (text-to-speech voices with added real and synthetic noise); sound events are real recordings. One run per system at temperature 0. A system at Intelligence 0 answered at or below chance after correction, often by refusing or returning invalid output. Results describe these exact configurations.

Revision history

  • v0.1 · 29 Sep 2026 — first public release. Method frozen on 26 Sep 2026 before any sealed response (freeze sha256 7f78d3ee…d50b), inheriting the JevBench v1.5 method and pricing rules; two scorer implementation errata logged before scoring (gate exclusions, public-only scoring), no formula change. 17 systems: 9 full coverage, 6 public-only, 2 partial.

AudioJevBench is a separate benchmark from the text-only JevBench and from Image JevBench. Sealed item content stays private; this page shows system-level aggregates only.