Diffusion Jev (DiffusionGemma 26B-A4B on patched SGLang, 48-step diffusion readout, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 58.1 · rank #37 of 87 on the open-weights board

Last measured: 2026-10-06 · measurement release: v1.6.1

apache-2.0 (model card metadata); linked Gemma 4 model terms

Base model: google/diffusiongemma-26B-A4B-itsource

intelligence

54.9

calibration

61.3

speed

86.8

cost

49.6

Composite (secondary)

59.3 · Composite rank #16

Cost and measurement conditions

$0.048 per 1,000 decisions (estimated)

ESTIMATE, and the caveat travels with the number rather than sitting in a footnote. The frozen pricing snapshot returns None for google/diffusiongemma-26B-A4B-it (it is unlisted: no exact rate and no floor), and a scan of all 129 registry entries finds NO published row on any DiffusionGemma base at all - the page's openjev-razorback16 is the same model but a different submitter and carries no v1.6 score - so there was no precedent to copy. The reference used is google/gemma-4-26b-a4b-it, USD 0.09 per 1M input and USD 0.30 per 1M output: the identically sized and identically shaped publicly hosted Gemma 4 26B-A4B sibling of the unlisted diffusion checkpoint, a closer sibling than the one run 54 had to use for gemma-4-12B-it. x this row's own measured tokens. THE CAVEAT: a token-priced cost axis charges this row as though it ran one forward pass, when the author's own documentation denoises a 256-position canvas for up to 48 adaptive steps per question. The Cost axis therefore UNDERSTATES this configuration's compute, by construction and not by accident; the measured per-question denoising_steps are kept in the raw so the understatement is quantifiable. Estimated, not charged

p50 latency: 0.4 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-06 on the live v1.6.0/v1.6.1 pool. github.com/Hangzhi/diffusion-jev-sglang@d61f0a9c0c9b6d78631aa250125dc4b20e9ea0a2 serving google/diffusiongemma-26B-A4B-it@f7f5b7f5fa82ffc52addd066915886d497f5517b on SGLang@ef90bf677e4e6b280bf796f1c5f4ef0f05dda27f (their open PR 34061) with their own version-checked readout and context-admission patch. Built from THEIR OWN tested container recipe (deployment/modal_app.py + deployment/gemma-requirements.txt), which their doc names as the recipe to use; their `engine` extra and top-level Dockerfile target the older LLaDA deployment and their doc says plainly not to use them as a DiffusionGemma installer, so they were not used. SGLang is a pure-python install imported through PYTHONPATH, no source compile. Served with their own scripts/launch_diffusiongemma.py: native Gemma4Renoise, a 256-position canvas, up to 48 adaptive denoising steps, seed 42, thinking disabled, 8,192-token context, Triton attention, no CUDA graphs, no prefix cache, T=1 softmax over the options offered - their configuration, unchanged, on one H100 80 GB (Lium) against the A100 80 GB they tested. THE AUTHOR'S OWN STATEMENT, CARRIED UNSOFTENED: these are self-conditioned denoiser scores, NOT independent masked-token likelihoods and NOT calibrated probabilities of correctness; this is not a one-forward-pass or zero-generation method. There are no new trained weights - it is an inference configuration over Google's unchanged checkpoint. A format mismatch returns their own HTTP 503 and is recorded invalid, the same way for everyone. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 55, GitHub #30

Published source