APUS-OpenJev-v1-35B-A3B (merged bf16 checkpoint-5949 MoE on Qwen3.5-35B-A3B, their own native runtime, full depth, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 60.4 · Not ranked in the current release

registry ranked=false or no composite

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-35B-A3Bsource

intelligence

52.8

calibration

67.9

speed

80.5

cost

—

Composite (secondary)

— · Not ranked on this board

Cost and measurement conditions

not published (not published)

UNPRICED, and deliberately not invented. pricing_v15.reference() RAISES on the exact declared base Qwen/Qwen3.5-35B-A3B ('base model lacks an explicit market-listing decision'). Run 46's bongard-mini rule applies unchanged: the frozen module raises rather than estimating, and the FROZEN-TABLE ROW is the release lane's (board #11 / #10132) - not this job's and not Florian's, because his standing 24 Sep 20:30 rule already decides the principle and the 26 Sep rule fixes the method. Complete but unpriced - shown, not ranked. RELEASE-LANE LINE, and a tight one: the dated snapshot ALREADY lists the exact base as qwen/qwen3.5-35b-a3b at (0.3125, 1.25), and BASE_REFERENCES ALREADY carries 'Qwen/Qwen3.5-35B-A3B-Base' -> 'qwen/qwen3.5-35b-a3b'. Only the non--Base (instruct) spelling of the same model is missing, so every Qwen3.5-35B-A3B derivative raises today. BECAUSE OUTPUT TOKENS ARE 0 BY CONSTRUCTION this row's Cost depends on the input rate alone, so the number is one multiplication over tokens the raw keeps; r56/receipts/COST-ARITHMETIC-R56.json carries it already computed.

p50 latency: 0.9 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights apus-ailab/APUS-OpenJev-v1-35B-A3B@090388abe2237117cc3b4df6f4250e81cbb36864, Apache-2.0, 66.99 GiB bf16 merged weights, checkpoint-5949, Qwen3_5MoeForConditionalGeneration (40 layers, 256 experts, 8 active), declared base Qwen/Qwen3.5-35B-A3B, full depth 40 layers (their depth_config declares exit_depth 20 / full_depth 40). accelerate>=1.10 additionally, per this repo's own requirements.txt. STACK: the authors' own requirements.txt, unmodified - torch==2.8.0 and transformers==5.16.1. The transformers pin is THEIRS, not a workaround: OpenJet.__init__ raises 'Layer execution is audited for transformers==5.16.1' on any other version. Served behind a TRANSPORT SHIM ONLY (r56/serve/apus_systemone_r56.py): loaded with THEIR OpenJet.from_pretrained(device=cuda:0, dtype=bfloat16) and every request handed to THEIR decide(), returning THEIR format_response() output. No contributor file edited. effort='high', the full-depth path, their decide() default and the configuration their own headline Frozen80 number is quoted on; effort='low' is not measured and no routing between the two is invented. One H100 PCIe 80 GB (Lium). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. MAPPING (run 56's, not the authors'; pre-registered in PREREG-R56 s2.2 before any item was sent and repeated here because it is part of the number): their runtime accepts primitive in {choice, noul, score_level} with 2-16 candidates labelled A-P. choice -> their 'choice' primitive, candidates = the item's own labels as id and its own criteria[label] as description, in labels order; their format_response returns probabilities keyed by candidate id, i.e. by the JevBench label already, so nothing is rewritten. All 750 choice items carry 3-9 options, inside their 2-16 window. noul -> their 'noul' primitive; their contract REQUIRES the canonical yes/no criteria there, so the item's own true/false description strings are NOT delivered - the proposition itself travels verbatim in instructions, which is what their noul is built to judge. score -> their 'choice' primitive over the ordinal levels, because their own RUNTIME.md rules score_level out as an ordinal Score API ('score_level is the judgement of a single proposition, it cannot be treated as a complete ordinal Score API'); the 4-5 ordered levels become candidates with id = the level index and description = the item's own level text, which carries the order. Recording all 375 score items as misses instead would understate a system whose published interface answers them. CALIBRATION IS MEASURED AS THE AUTHORS SHIP IT AND THEY DISCLOSED IT IN ADVANCE: their response carries calibrated: False and both their card and RUNTIME.md state the probabilities are an uncalibrated relative softmax over the offered candidates. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, NOT BY ESTIMATE: decide() runs one forward pass with logits_to_keep=1 and reads candidate logits; nothing is generated anywhere in that path. Input tokens are their own prompt_tokens field. THE 23 LONG ITEMS: their compile() raises ValueError('input exceeds runtime limit; no truncation permitted') above 8,192 tokens and their RUNTIME.md states it in advance. Recorded invalid. This is the first row in the series whose refusal on those items is guaranteed in the author's own code with an explicit no-truncation promise. add-requests run 56, GitHub #105, benchmarkheaven.com/submit ref AE709BBB

Published source