decisio v0.8.0 on gemma-4-12B-it (frozen, one forward pass, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 70.7 · rank #10 of 87 on the open-weights board

Last measured: 2026-10-06 · measurement release: v1.6.1

Apache-2.0 software; Gemma model terms

Base model: google/gemma-4-12B-itsource

intelligence

52.3

calibration

89.2

speed

87.0

cost

59.3

Composite (secondary)

68.2 · Composite rank #7

Cost and measurement conditions

$0.023 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights, no public tariff): the exact base google/gemma-4-12B-it is in the frozen pricing module's UNLISTED_BASE_MODELS set, so no exact base-model rate and no floor exist. The reference is the one the rows PUBLISHED on this same base use - cygnet, jev-omni and winnow-12b - OpenRouter google/gemma-3-12b-it, the nearest publicly hosted 12B Gemma sibling, at its FULL catalog pair USD 0.05 per 1M input and USD 0.15 per 1M output. PREREG-R54-AMENDMENT-1 uses the full pair rather than run 53's output=0, so a row that emits tokens pays for them; where a row generates nothing the output term is zero anyway. x this row's own measured input/output tokens, the deck31b method. Estimated, not charged

p50 latency: 0.4 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-06 on the live v1.6.0/v1.6.1 pool. github.com/aminry/decisio tag v0.8.0 (5b42101a191d222062a795194b0aedb5003bf6ea, verified against the GitHub refs API), installed with their own `uv sync --extra serve --frozen` so their lockfile fixes every version, and served with their own `python -m decisio.serve.vllm_engine --base gemma-4-12b`. --base carries google/gemma-4-12B-it@707f0a3b (bf16, 22.8 GiB) and that profile's measured settings - temperature 3.592 for every question type, a system turn, the chat template's own answer position, every single-token form of each letter summed. Frozen base: no fine-tuning, no adapter, no task registered. One RTX 6000 Ada 48 GB (Lium, Toronto). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 54, GitHub #169

Published source