SPX-CD-Omni (SurdAI, google/gemma-4-12B-it + LoRA r32, one forward pass, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 68.1 · rank #15 of 87 on the open-weights board

Last measured: 2026-10-06 · measurement release: v1.6.1

apache-2.0

Base model: google/gemma-4-12B-itsource

intelligence

48.8

calibration

87.5

speed

89.4

cost

58.3

Composite (secondary)

63.2 · Composite rank #15

Cost and measurement conditions

$0.025 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights, no public tariff): the exact base google/gemma-4-12B-it is in the frozen pricing module's UNLISTED_BASE_MODELS set, so no exact base-model rate and no floor exist. The reference is the one the three rows PUBLISHED on this same base all use - cygnet, jev-omni and winnow-12b - OpenRouter google/gemma-3-12b-it list price USD 0.05 per 1M input, USD 0.00 per 1M output (the nearest publicly hosted 12B Gemma sibling; output is 0 because the runner reads candidate logits in one forward pass and generates nothing), x this row's own measured usage.input_tokens, the deck31b method. Estimated, not charged

p50 latency: 0.3 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-06 on the live v1.6.0 S u P (1,500 rows, 1,477 ok; the 23 refusals are exactly the 23 long-context items, through the author's own "Prompt exceeds context limit; no automatic truncation" guard at their shipped --max-context 16384, although the base model itself declares 262,144 positions). Author's own code unmodified (infer.py / decision.py, sha256 1c5db5de.../43389f44...) with their shipped --backend transformers on the unquantized base and the flags their own EVALUATION.md prescribes FOR JEVBENCH: --prompt-format open-format --effort 1 --temperature 1.0. Their requirements-transformers.txt is missing torchvision, which transformers 5.16.1 needs for the Gemma 4 unified processor; adding it was the only deviation and it changes no numerics. One RTX PRO 6000 Blackwell (Lium, Brasov, Romania), torch 2.12.0+cu130, transformers 5.16.1, peft 0.20.0. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment; raw p50/p95 are in the run 53 receipts; add-requests run 53, benchmarkheaven.com/submit ref 75D5EF0C

Published source