TypeCastLM 1.4.0 (Mikhail Gribov, Qwen3.5-4B computed head)

Data: JevBench v1.6.1, published 2026-10-06

Capability 56.4 · rank #44 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

35.4

calibration

77.4

speed

93.0

cost

64.8

Composite (secondary)

29.8 · Composite rank #40

Cost and measurement conditions

$0.015 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights). NO JUDGEMENT OF OURS: pricing_v15.reference() resolves Qwen/Qwen3.5-4B straight through to deepinfra:Qwen/Qwen3.5-4B at USD 0.03 per 1M input and USD 0.15 per 1M output in the frozen 25 Sep snapshot (the same rate aplomb-1 carries for the same base) x this row's own measured tokens. Output tokens are 0 in every answered record: one prefill, option-label readout, nothing generated. Estimated, not charged.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,500 of 1,500 answered, 0 invalid. A FOLLOW-UP OF THE PUBLISHED typecastlm ROW: the author asks to 'measure 1.4.0 to replace it' and to list it as 'TypeCastLM'; measured here as a CANDIDATE - replace-or-keep-both is the release lane's decision. Weights mihailgribov/typecastlm-qwen3.5-3.8b tag v1.4.0 = commit 0984c3e832f7c28495134bc1d22c66bb0e0a555d (GGUF files excluded), package typecastlm[server]==1.2.2 (latest on PyPI) + flash-linear-attention + fla-core, the model card's own install lines, in an isolated venv (torch 2.14.1+cu130, transformers 5.19.0). Served with the author's own `typecastlm-serve --model <local v1.4.0 snapshot> --host 127.0.0.1 --port 8000 --device cuda`; the UNCHANGED typesafe adapter with the published row's own wire model; all eight checks of the unchanged smoke gate passed on synthetic requests before any sealed item was sent. THE 23 LONG ITEMS ARE ANSWERED, folded in the middle to the server's own 32,768-token state context (its documented behaviour; max input 32,953 tokens observed). One RTX 5090 32 GB (Lium, France). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, credential-free. add-requests run 59, benchmarkheaven.com/submit ref 34CA4AED

Published source