Tacet Sonata (CodePawl, 144M packed encoder on mmBERT-small, option-marker softmax in one pass, in-process via the author's tacet package)
Data: JevBench v1.6.1, published 2026-10-06
Capability 24.3 · rank #81 of 87 on the open-weights board
Last measured: 2026-10-07 · measurement release: v1.6.1
apache-2.0
intelligence
2.2
calibration
46.4
speed
80.3
cost
89.0
Composite (secondary)
0.0 · Composite rank #114
Cost and measurement conditions
$0.0023 per 1,000 decisions (estimated)
ESTIMATE (size-class reference; self-hosted open weights, no public tariff). THE SAME CHOICE THE PUBLISHED quyet-1-0-tiny ROW CARRIES FOR THE IDENTICAL mmBERT-small BACKBONE: DeepInfra base-size encoders at USD 0.005 per 1M input, 0 output - deepinfra:thenlper/gte-base, deepinfra:BAAI/bge-base-en-v1.5, deepinfra:intfloat/e5-base-v2 and deepinfra:sentence-transformers/all-mpnet-base-v2 are all (0.005, 0.0) in the frozen 25 Sep snapshot - x this row's own measured usage.input_tokens (the package's packed length, capped at 4,096). Estimated, not charged. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, verified in the package source (tacet/decoding.py usage_of sets output_tokens 0; the readout is a softmax over option [MASK] markers from one encoder pass and nothing is generated), the 02 Oct deputy rule for classifier read-outs. RELEASE-LANE LINE, recorded and not acted on, AND IT IS THE QUYET ROWS' CARRIED LINE, NOT A NEW ONE: jhu-clsp/mmBERT-small still has no explicit market-listing decision (pricing_v15.reference() raises), so an UNLISTED_BASE_MODELS entry is the release lane's.
p50 latency: 0.6 s
x2 + 0.15 s (assumption, not measured)
measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...), 1,500 of 1,500 answered, 0 invalid. Weights codepawl/tacet-sonata @ 26f447c35881a80020fe557c95ceb8c61bc3fd12, Apache-2.0, public and ungated, loaded from a local snapshot (no network during the run). Runtime tacet==0.3.0 from PyPI, the author's own install line (`pip install "tacet>=0.3"`) in an isolated venv (torch 2.14.1+cpu, transformers 5.19.0). ADAPTER: the author's own PR fstandhartinger/jevbench#197 file tacet_local.py at head 76eb75da, byte-identical (sha256 c63a4fe0...), reviewed before running: it maps answers only (noul P(true) -> yes/no; choice and score probabilities used as returned, score keys '0'..'n-1') and reaches nothing but the Hub. It ran through the UNCHANGED v1.6 driver run_v16.py; the only addition, in a wrapper, registers the file as jevbench.adapters.tacet_local and lists it in IN_PROCESS_MODULES as laya_local is listed - lane verification and the runtime egress guard unchanged. Synthetic smoke gate passed before any sealed item was sent (PREREG-R59.md, receipts/SMOKE-tacet-sonata.json). THE 23 LONG ITEMS ARE TRUNCATED, NOT REFUSED, by the package's own code at its default and documented max_length 4,096 (tacet/packing.py keeps the questions and cuts the state; it refuses only states over 200,000 characters, which no item here reaches): exactly 23 records report input_tokens = 4,096 and they are answered and scored. RUN ON SANDY'S SHARED CPU (12-core Zen 2, AVX2, fp32 as the package runs off-accelerator, 8 threads), load average 6.6-24.2 (median ~14) logged every 60 s: the Speed axis is a DECLARED LOWER BOUND, and the author reports raw p50 0.042 s on an RTX 3060 against this run's 0.265 s. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per call, no retry, no batching, offline, credential-free. add-requests run 59, benchmarkheaven.com/submit ref 716D1957 (bh-submit 07d384eb-fc4e-4536-ba61-888567efdf34)