Kahn1 4B (Okura66, Qwen3.5-4B LoRA merge, author's sysone engine)

Data: JevBench v1.6.1, published 2026-10-06

Capability 61.7 · rank #23 of 68 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

39.7

calibration

83.8

speed

90.0

cost

46.4

Composite (secondary)

31.1 · Composite rank #28

Cost and measurement conditions

$0.061 per 1,000 decisions (estimated)

ESTIMATE, ours (the author's server returns no token usage; Florian 7 Oct 2026: use our own estimate): 829,523 input tokens counted with the author's tokenizer and prompt builder over the 1,624 v1.5 items x n_permutations 3 (the author's server default) = 2,488,569 tokens / 1,624 decisions x USD 0.04 per 1M input (deepinfra:Qwen/Qwen3.5-4B, frozen 25 Sep snapshot), 0 output tokens = USD 0.061295 per 1,000 decisions; carried like every Qwen3.5-4B row on this board. Cross-check: the same count on the v1.6 common cost basis (1,477 items, x3) gives USD 0.0597 per 1,000 (within 3 %)

p50 latency: 0.3 s

x2 + 0.15 s (assumption, not measured)

evaluator-owned Lium GPU pod, 1x NVIDIA RTX 5090, the author's own sysone engine (github.com/Okura66/kahn1) on vLLM 0.30.0 with the author's documented max-model-len 4096 and calibration.json, offline (all credentials unset), driven over loopback only by the JevBench client through a thin pre-registered adapter for the author's /v1/evaluate/jev endpoint; label-free inputs; pod destroyed afterwards; JevBench add-request, measured 7 Oct 2026

Published source