Messier One v0.2 (Qwen3.5-4B fine-tune, one prefill, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 58.1 · rank #36 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

cc-by-nc-4.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

42.4

calibration

73.9

speed

92.7

cost

64.8

Composite (secondary)

45.3 · Composite rank #21

Cost and measurement conditions

$0.015 per 1,000 decisions (not published)

ESTIMATE (base-model reference; self-hosted open weights). NO JUDGEMENT OF OURS: pricing_v15.reference() resolves Qwen/Qwen3.5-4B straight through to deepinfra:Qwen/Qwen3.5-4B at USD 0.03 per 1M input and USD 0.15 per 1M output in the frozen 25 Sep snapshot (the same rate aplomb-1 carries for the same base) x this row's own measured tokens. Estimated, not charged. THE FRONT END REPORTS NO TOKEN USAGE (usage {}), exactly as v0.1 did in run 54, so the counts come from run 54's pre-registered route unchanged (PREREG-R59-AMENDMENT-1 = R54-AMENDMENT-2): the author's own prompt.Prompter (v0.2 source, unmodified) recomputes each item's prompt token ids, output is fixed at 1 by their own max_tokens=1, and the totals were required to match vLLM's own counters or the row would get no cost: 914,665 input and 1,500 output tokens both ways (vllm prompt_tokens_total 915,349 - 684; generation_tokens_total 1,508 - 8). Counts sit in cost_estimate, never in usage.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,500 of 1,500 answered, 0 invalid. REPLACES the v0.1 request (GitHub #194, measured by run 54); the submission names it so. Weights agentmessier/messier-one tag v0.2 = 8bbddb84140e8fc732b35713f998ed628fc70306 (CC BY-NC 4.0, NON-COMMERCIAL), front end github.com/agentmessier-ai/messier-one tag v0.2 = 1b8c3b80 (Apache-2.0). README install line (vllm==0.30.0, transformers==5.18.0; torch 2.13.0+cu130 resolved by vllm) and the author's two serve commands verbatim; --gate off, their default and the configuration of their own reported single-read run. The author discloses that part of the training material was written in the families of JevBench's hard tier, with no JevBench item used. Their Prompter caps the state at 8,000 tokens (max_total 8,600) and splices long states around a readable marker, so the 23 long items are answered truncated (r54 s on disclosed truncation). UNCHANGED typesafe adapter; the smoke gate refused the first attempt at usage-input-tokens and passed with its existing pre-registered opt-in. One RTX 5090 32 GB (Lium, France). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, offline, credential-free. add-requests run 59, benchmarkheaven.com/submit ref 455A3944

Published source