Aplomb 1 (5.3B decision model on Qwen3.5-4B, trained readout head, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 67.5 · rank #16 of 87 on the open-weights board

Last measured: 2026-10-06 · measurement release: v1.6.1

EmpirioLabs Model License 1.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

45.8

calibration

89.2

speed

88.9

cost

57.9

Composite (secondary)

54.6 · Composite rank #17

Cost and measurement conditions

$0.025 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights): the exact declared base Qwen/Qwen3.5-4B resolves in the frozen pricing snapshot through DEEPINFRA_BASE_REFERENCES to USD 0.03 per 1M input and USD 0.15 per 1M output, read before the pod and asserted in make_registry_r55.py. The same rate run 54 gave messier-one, the other self-hosted Qwen3.5-4B derivative. DELIBERATELY NOT the vendor's own cheaper public tariff: Aplomb 1 is also sold on EmpirioLabs' hosted API at USD 0.02 per 1M input with output free (GitHub #195), which would LOWER this row's cost - but all 106 self-hosted rows in the registry carry cost_kind `estimate` and not one is priced at a vendor tariff, and creating that precedent for one row is the release lane's call, not this job's. The estimate can therefore only OVERSTATE this row, and it is reversible by arithmetic because the raw keeps the tokens. x this row's own measured tokens. INPUT TOKENS ARE EXACT, NOT ESTIMATED, and they include BOTH passes of the author's own --debias flag on every multi-option choice, because that is what the configuration consumes. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, not by estimate: the author's answer_logits runs the model with logits_to_keep=1, use_cache=False and reads label logits, and nothing is generated anywhere in their code path. Counts carried in the scorer's own cost_estimate field, never written into usage. Estimated, not charged

p50 latency: 0.3 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-06 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights empiriolabsai/aplomb-1@f8e2dc81f0505db21a3b7a933ea479658e4ebb8f, bf16, 5.3 B, base Qwen/Qwen3.5-4B plus the shipped 1.3 GB audio encoder and the trained DecisionHead; click-through gate accepted under the EmpirioLabs Model License 1.0 section 2.1, which grants evaluation and the publication of benchmark results without asking. Served behind the author's own TypeSafe wire format by a TRANSPORT SHIM ONLY: run_aplomb.main()'s loading block verbatim, every request handed to THEIR run_aplomb.decide() with THEIR edm_config.json temperatures and abstain_bias and debias=True - the configuration the author names as the reference in #195 and in the model card. Their own aplomb/schema.py documents the request schema as wire-compatible with TypeSafe /v1/systemone and their aplomb/readout.py already emits the answer dicts the unchanged typesafe adapter parses, so no answer field is rewritten and no contributor file is edited. STACK: the author's own documented install line, with transformers pinned to 5.17.0 - the version their own audio sidecar was saved with; their loader compares the saved config to the live one and 5.19.0 fails it on the transformers_version string alone. SPEED CAVEAT: their documented install line does not include causal_conv1d, so transformers falls back to a reference PyTorch implementation it itself calls much slower; the Speed axis is measured on the stack the author documents. One H100 80 GB (Lium). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. add-requests run 55, GitHub #195

Published source