APUS-OpenJev-v1-9B (merged bf16 checkpoint-3000 on Qwen3.5-9B, their own native runtime, full depth, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 56.9 · rank #42 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-9Bsource

intelligence

47.2

calibration

66.6

speed

82.2

cost

48.7

Composite (secondary)

49.2 · Composite rank #19

Cost and measurement conditions

$0.051 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights): the exact declared base Qwen/Qwen3.5-9B resolves in the frozen pricing snapshot through BASE_REFERENCES to qwen/qwen3.5-9b = USD 0.10 per 1M input and USD 0.15 per 1M output, read before the pod and asserted in make_registry_r56.py. Nothing provisional. x this row's own measured tokens. BECAUSE OUTPUT TOKENS ARE 0 BY CONSTRUCTION this axis rests on the input rate alone. Estimated, not charged.

p50 latency: 0.7 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights apus-ailab/APUS-OpenJev-v1-9B@82c9c56cfa9de8d36704ed91948d4726ef111635, Apache-2.0, 17.55 GiB bf16 merged weights, checkpoint-3000, declared base Qwen/Qwen3.5-9B, full depth 32 layers. STACK: the authors' own requirements.txt, unmodified - torch==2.8.0 and transformers==5.16.1. The transformers pin is THEIRS, not a workaround: OpenJet.__init__ raises 'Layer execution is audited for transformers==5.16.1' on any other version. Served behind a TRANSPORT SHIM ONLY (r56/serve/apus_systemone_r56.py): loaded with THEIR OpenJet.from_pretrained(device=cuda:0, dtype=bfloat16) and every request handed to THEIR decide(), returning THEIR format_response() output. No contributor file edited. effort='high', the full-depth path, their decide() default and the configuration their own headline Frozen80 number is quoted on; effort='low' is not measured and no routing between the two is invented. One H100 PCIe 80 GB (Lium). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. MAPPING (run 56's, not the authors'; pre-registered in PREREG-R56 s2.2 before any item was sent and repeated here because it is part of the number): their runtime accepts primitive in {choice, noul, score_level} with 2-16 candidates labelled A-P. choice -> their 'choice' primitive, candidates = the item's own labels as id and its own criteria[label] as description, in labels order; their format_response returns probabilities keyed by candidate id, i.e. by the JevBench label already, so nothing is rewritten. All 750 choice items carry 3-9 options, inside their 2-16 window. noul -> their 'noul' primitive; their contract REQUIRES the canonical yes/no criteria there, so the item's own true/false description strings are NOT delivered - the proposition itself travels verbatim in instructions, which is what their noul is built to judge. score -> their 'choice' primitive over the ordinal levels, because their own RUNTIME.md rules score_level out as an ordinal Score API ('score_level is the judgement of a single proposition, it cannot be treated as a complete ordinal Score API'); the 4-5 ordered levels become candidates with id = the level index and description = the item's own level text, which carries the order. Recording all 375 score items as misses instead would understate a system whose published interface answers them. CALIBRATION IS MEASURED AS THE AUTHORS SHIP IT AND THEY DISCLOSED IT IN ADVANCE: their response carries calibrated: False and both their card and RUNTIME.md state the probabilities are an uncalibrated relative softmax over the offered candidates. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, NOT BY ESTIMATE: decide() runs one forward pass with logits_to_keep=1 and reads candidate logits; nothing is generated anywhere in that path. Input tokens are their own prompt_tokens field. THE 23 LONG ITEMS: their compile() raises ValueError('input exceeds runtime limit; no truncation permitted') above 8,192 tokens and their RUNTIME.md states it in advance. Recorded invalid. This is the first row in the series whose refusal on those items is guaranteed in the author's own code with an explicit no-truncation promise. add-requests run 56, GitHub #106, benchmarkheaven.com/submit ref 2033065A

Published source