decisio v0.8.0 on gemma-4-31B-it (frozen, FP8 on load, one prefill per state, self-hosted)
Data: JevBench v1.6.1, published 2026-10-06
Capability 79.6 · rank #2 of 92 on the open-weights board
Last measured: 2026-10-06 · measurement release: v1.6.1
Apache-2.0 software; Gemma model terms
intelligence
70.4
calibration
88.7
speed
91.2
cost
51.7
Composite (secondary)
71.7 · Composite rank #2
Cost and measurement conditions
$0.041 per 1,000 decisions (estimated)
ESTIMATE (base-model reference; self-hosted open weights, no public tariff) and CLEAN: unlike run 54's FP8 sibling, the frozen pricing snapshot resolves this EXACT base - google/gemma-4-31B-it - to USD 0.09 per 1M input and USD 0.34 per 1M output. Read before the pod and asserted in make_registry_r55.py. x this row's own measured tokens, the deck31b method. Nothing is provisional and no release-lane line is open on it. Estimated, not charged
p50 latency: 0.2 s
x2 + 0.15 s (assumption, not measured)
measured 2026-10-06 on the live v1.6.0/v1.6.1 pool. github.com/aminry/decisio tag v0.8.0 (5b42101a191d222062a795194b0aedb5003bf6ea) - the same tag whose other two entries run 54 measured - installed with their own `uv sync --extra serve --frozen` so their lockfile fixes every version, and served with their own `python -m decisio.serve.vllm_engine --base gemma-4-31b`. --base carries google/gemma-4-31B-it@842da379, the official weights, FROZEN: no fine-tuning, no adapter, no task registered; vLLM 0.30.0 quantizes it to FP8 when it loads (30.6 GiB). That profile's measured settings: temperature 4.672 for choice questions and 5.252 for the rest, the 12B entry's prompt unchanged, a system turn, the chat template's own answer position, every single-token form of each letter summed, no tokens generated. New in 0.8.0 on this base and the reason it is worth its own row: several questions of a state are scored in ONE pass after a single prefill. The author's own limit is a 32,768-token prompt, so the pool's 23 items at ~80,000 tokens are a documented capacity limit, not a defect. THE AUTHOR STATES ONE 96 GB CARD; this is one H100 80 GB (Lium) and no flag of theirs was changed to make it fit. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 55, GitHub #169