H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B fine-tune, stock vLLM + open shim)

Data: JevBench v1.6.1, published 2026-10-06

Capability 75.0 · rank #3 of 68 on the open-weights board

Last measured: 2026-10-04 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

60.0

calibration

90.0

speed

92.6

cost

60.3

Composite (secondary)

72.5 · Composite rank #1

Cost and measurement conditions

$0.021 per 1,000 decisions (estimated)

ESTIMATE (self-hosted open weights, no public tariff): cost per 1,000 decisions measured for H2O-Lightning-4B v1.0 on the v1.5 pool (same base, shim and token path) carried for v1.1, the practice of every Qwen3.5-4B row on this board; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off. Cross-check: v1.1's own v1.6 input tokens on the common cost basis give USD 0.0204 per 1,000 (within 3 %)

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

evaluator-owned Lium GPU pod, 1x NVIDIA RTX 5090, official vLLM 0.30.0 image + the author's open standard-library shim (h2o_lightning_shim.py) as published on the model card, offline (all credentials unset), driven over loopback only by the JevBench client; label-free inputs; pod destroyed afterwards; JevBench add-request (GitHub #181), measured 4 Oct 2026

Published source