decider chat on Gemma-4-31B-it (Mapika, frozen base, inference technique)
Data: JevBench v1.6.1, published 2026-10-06
Capability 70.6 · rank #7 of 64 on the open-weights board
Last measured: 2026-10-06 · measurement release: v1.6.1
apache-2.0
intelligence
58.5
calibration
82.6
speed
85.8
cost
50.8
Composite (secondary)
66.1 · Composite rank #7
Cost and measurement conditions
$0.044 per 1,000 decisions (estimated)
ESTIMATE (base-model reference; self-hosted open weights, no public tariff): OpenRouter google/gemma-4-31b-it, the reference of the published deck31b and quyet-1-0-large rows, USD 0.09 per 1M input and USD 0.34 per 1M output x this row's own measured tokens. Estimated, not charged
p50 latency: 0.4 s
x2 + 0.15 s (assumption, not measured)
evaluator-owned Lium GPU pod, 1x NVIDIA RTX PRO 6000 Blackwell 96GB, the author's own server in its own pinned environment, offline (Hub offline mode, all credentials unset), driven over loopback only by the JevBench client with its loopback guard; label-free inputs; pod destroyed afterwards; HF Jev Decision Index intake, 6-7 Oct 2026