Seb-9B (Qwen3.5-9B-shaped multimodal decision model, one forward pass, stock vLLM, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 66.8 · Not ranked in the current release

registry ranked=false or no composite

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: undisclosedsource

intelligence

43.5

calibration

90.1

speed

90.0

cost

—

Composite (secondary)

— · Not ranked on this board

Cost and measurement conditions

not published (not published)

UNPRICED, and deliberately not invented. The author declares NO base model anywhere in the repository, so pricing_v15.reference() has nothing to resolve: it returns None without raising, and the frozen module records the row unpriced. Complete but unpriced - shown, not ranked - exactly as run 52 left ines-1. RELEASE-LANE LINE: whether an UNDECLARED base that is architecturally identified may supply a size-class estimate. The repo ships Qwen3_5ForConditionalGeneration with 32 layers, hidden 4096 and the Qwen3.5 tokenizer under the name '9B', and qwen/qwen3.5-9b = (0.10, 0.15) is the hosted model of that exact shape; r56/receipts/COST-ARITHMETIC-R56.json carries the resulting cost per 1,000 over this row's own measured tokens, clearly labelled as OUR candidate reference and NOT the author's claim, so the release lane needs no re-measurement either way.

p50 latency: 0.3 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights ironbcc/seb-9b@ac57b36bbe978762032d0ad7197778b4bc4a0209, Apache-2.0, 18.0 GiB bf16. Served with the AUTHOR'S OWN PUBLISHED vLLM COMMAND, verbatim from their model card, on vLLM 0.30 - the version their card states it is tested with - at the 8,192 --max-model-len their own flag note names for long documents (the larger of the two values they give). r56/serve/seb_systemone_r56.py supplies TRANSPORT ONLY: it builds their own documented chat request per decision (their system string verbatim, max_tokens 1, temperature 0, logprobs true, top_logprobs = the option count, structured_outputs.choice = the option keys, enable_thinking false) inside their own user-message envelope {task,state,question,type,options}, and reads the probabilities exactly where their card says they are - exp of choices[0].logprobs.content[0].top_logprobs - renormalised over the option keys. Their own key convention is used unchanged (noul Y/N, choice A..T, score 0..4) and every criterion description travels verbatim; no prompt wording of theirs is altered and no file of theirs is edited. THE 23 LONG ITEMS are rejected server-side by vLLM because the author's own published command sets --max-model-len 8192; recorded invalid and recorded as a CONFIGURATION limit under their own published flags, NOT as a model refusal - their config declares 262,144 positions and raising the flag would be our change, not theirs. One H100 PCIe 80 GB (Lium). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 56, benchmarkheaven.com/submit ref C566D021

Published source