Hopper 12B trained (gemma-4-12B-it frozen + LoRA r32 unmerged, one forward pass, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 68.1 · rank #14 of 87 on the open-weights board

Last measured: 2026-10-06 · measurement release: v1.6.1

apache-2.0

Base model: google/gemma-4-12B-itsource

intelligence

51.7

calibration

84.6

speed

84.6

cost

56.2

Composite (secondary)

65.8 · Composite rank #12

Cost and measurement conditions

$0.029 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights, no public tariff): the exact base google/gemma-4-12B-it is in the frozen pricing module's UNLISTED_BASE_MODELS set, so no exact base-model rate and no floor exist. The reference is the one the rows PUBLISHED on this same base use - cygnet, jev-omni and winnow-12b - OpenRouter google/gemma-3-12b-it, the nearest publicly hosted 12B Gemma sibling, at its FULL catalog pair USD 0.05 per 1M input and USD 0.15 per 1M output. PREREG-R54-AMENDMENT-1 uses the full pair rather than run 53's output=0, so a row that emits tokens pays for them; where a row generates nothing the output term is zero anyway. x this row's own measured input/output tokens, the deck31b method. Estimated, not charged

p50 latency: 0.5 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-06 on the live v1.6.0/v1.6.1 pool. Serving code github.com/hopit-ai/hopper tag 12b-1.1.1, installed with the author's own one-line pip install; adapter HopitAI/hopper@4b605f25eb903c2cf4d121040ffb097240b3bc37 (branch hopper-12b), LoRA r32 alpha 64, loaded UNMERGED because the author states a bf16 merge changes answers; base google/gemma-4-12B-it@707f0a3b, frozen; readout map maps/hopper-12b-trained-readout.json as shipped in their package. 1.1.1 is their own fix for corrupted option scores at prompt lengths one above a multiple of 32. Their flags left at their defaults: --shortlist tournament (their measured default) for menus over 26 options, --cuda-graphs OFF. The author asked on 6 Oct 2026 that the TRAINED version replace the untrained Hopper 12B request in GitHub #164; this is that row. NOT the published `hopper` row (HopitAI, rank 25), a different system. One RTX 6000 Ada 48 GB (Lium, Toronto); the author ran theirs on an L40S. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 54, GitHub #164

Published source