H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim)
Data: JevBench v1.6.1, published 2026-10-06
Capability 75.0 · rank #3 of 68 on the open-weights board
Last measured: 2026-10-04 · measurement release: v1.6.1
apache-2.0
intelligence
60.0
calibration
90.0
speed
92.6
cost
60.3
Composite (secondary)
72.5 · Composite rank #1
Cost and measurement conditions
$0.021 per 1,000 decisions (estimated)
ESTIMATE (self-hosted open weights, no public tariff): cost per 1,000 decisions measured for H2O-Lightning-4B v1.0 on the v1.5 pool (same base, shim and token path) carried for v1.1, the practice of every Qwen3.5-4B row on this board; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off. Cross-check: v1.1's own v1.6 input tokens on the common cost basis give USD 0.0204 per 1,000 (within 3 %)
p50 latency: 0.2 s
x2 + 0.15 s (assumption, not measured)
evaluator-owned Lium GPU pod, 1x NVIDIA RTX 5090, official vLLM 0.30.0 image + the author's open standard-library shim (h2o_lightning_shim.py) as published on the model card, offline (all credentials unset), driven over loopback only by the JevBench client; label-free inputs; pod destroyed afterwards; JevBench add-request (GitHub #181), measured 4 Oct 2026