decisio v0.9.0 on gemma-4-12B-it (frozen, one prefill per question, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 70.0 · rank #13 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

Apache-2.0 software; Gemma model terms

Base model: google/gemma-4-12B-itsource

intelligence

52.3

calibration

87.6

speed

85.9

cost

59.3

Composite (secondary)

67.8 · Composite rank #8

Cost and measurement conditions

$0.023 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights, no public tariff), CARRIED AND NOT NEW: the exact base google/gemma-4-12B-it is in the frozen pricing module's UNLISTED_BASE_MODELS set, so no exact base-model rate and no floor exist, and (0.05, 0.15) is the reference the rows PUBLISHED on this same base already use - cygnet, jev-omni, winnow-12b and torchcast-decision-12b - OpenRouter google/gemma-3-12b-it, the nearest publicly hosted 12B Gemma sibling, at its FULL catalog pair USD 0.05 per 1M input and USD 0.15 per 1M output. It is also the rate THIS ROW'S OWN v0.8.0 SIBLING carries, so the two versions are priced identically and the comparison between them is clean. PREREG-R54-AMENDMENT-1's full pair rather than output=0, x this row's own measured input/output tokens, the deck31b method. The server reports its own usage.input_tokens/output_tokens on every answer (664,759 in and 1,477 out over the answered rows: one prefill, one option-label readout step per decision), so no token-accounting amendment was needed. Estimated, not charged.

p50 latency: 0.4 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 21:11:36-21:15:32Z on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...), 1,477 of 1,500 answered. github.com/aminry/decisio tag v0.9.0 = 333a0a849de7658969a8f8933b34a8097a8c3489 (verified against the GitHub refs API before the run and re-read as the pod's checked-out HEAD after it), Apache-2.0, installed with their own `uv sync --extra serve --frozen` so their lockfile fixes every version (decisio 0.9.0, torch 2.13.0+cu130, vllm 0.30.0, transformers 5.18.0) and served with their own command `python -m decisio.serve.vllm_engine --base gemma-4-12b`. --base carries google/gemma-4-12B-it@707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 (bf16, 22.8 GiB) and that profile's own settings - temperature 3.592 for every question type, a system turn, the spaced layout read at the chat template's own answer position, every single-token form of each letter summed. Frozen base: no fine-tuning, no adapter, no task registered. UNCHANGED typesafe adapter and UNCHANGED run_v16.py; all eight checks of the unchanged smoke gate passed on synthetic requests before any sealed item was sent. A FOLLOW-UP OF THIS JOB'S OWN decisio-gemma-4-12b-v080 ROW (run 54, same pool, same basis, same rate), measured as a CANDIDATE: which version is listed is the release lane's decision. WHAT v0.9.0 CHANGES is prefix-cache state-boundary registration, request ticketing and engine-health probing; the whole readout package, serve/systemone.py, serve/temperature.py, serve/abstention.py and the entire vllm_plugin are BYTE-IDENTICAL between the two tags, and no temperature or prompt line moved (r61/receipts/DECISIO-V090-SOURCE-DIFF.json). THE 23 LONG ITEMS FAIL WITH A BARE HTTP 500 from the author's own server - vLLM raises VLLMValidationError because the rendered prompt exceeds the 32,768-token context the profile declares, and the server does not translate that into the 422 the harness scores as a refusal. They are counted wrong either way and they are exactly the 23 items the v1.6.1 common cost basis excludes (the ones the Jev reference itself refuses), but it is the author's own code and A FINDING TO REPORT, NOT A THING TO PATCH AROUND; no setting of theirs was changed and no retry was made. One NVIDIA RTX 6000 Ada Generation 48 GB (Lium, Toronto, CANADA) - the same card model and city on which run 54 measured the v0.8.0 row, chosen over the 96 GB class the author names because the 12B profile is proven to fit it. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, offline after setup, credential-free. add-requests run 61, PREREG-R61-AMENDMENT-2, benchmarkheaven.com/submit ref 65B71BD3 (bh-submit f05c3298-702f-4314-bc8e-f823d9bce858); previously GitHub #169 for v0.8.0

Published source