watt-flash-0.1 (Zaitgeist Labs, 140M mmBERT-small encoder with an option-marker scorer, one forward pass per request, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 41.6 · rank #66 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: jhu-clsp/mmBERT-smallsource

intelligence

3.8

calibration

79.4

speed

75.8

cost

89.0

Composite (secondary)

0.1 · Composite rank #110

Cost and measurement conditions

$0.0023 per 1,000 decisions (not published)

ESTIMATE (size-class reference; self-hosted open weights, no public tariff). THE SAME CHOICE THE PUBLISHED quyet-1-0-tiny ROW AND RUN 59'S tacet-sonata ROW CARRY FOR THE IDENTICAL mmBERT-small BACKBONE, carried unchanged and not a new judgement: DeepInfra base-size encoders at USD 0.005 per 1M input, 0 output - deepinfra:thenlper/gte-base, deepinfra:BAAI/bge-base-en-v1.5, deepinfra:intfloat/e5-base-v2 and deepinfra:sentence-transformers/all-mpnet-base-v2 are all (0.005, 0.0) in the frozen 25 Sep snapshot - x this row's own measured input tokens. Estimated, not charged. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, verified in the author's source: watt_flash/serve.py:answer() formats a softmax over the option markers of one encoder pass and the package contains no generation path at all (the 02 Oct deputy rule for classifier read-outs). THE SERVER REPORTS NO usage BLOCK OF ANY KIND, so the input counts come EXACTLY from the author's own unmodified watt_flash.layout.Layout.pack_request - the very number their own WattFlash.decide_batch returns as len(p) and only serve.py drops - under PREREG-R61-AMENDMENT-1, cross-checked three ways (independent reassembly of their documented layout; identity of their packer's Overflow set with the server's own 422 set), with no cost number at all on any disagreement; the counts sit in cost_estimate and never in usage. RELEASE-LANE LINE, recorded and not acted on, AND IT IS THE CARRIED quyet/tacet LINE, NOT A NEW ONE: jhu-clsp/mmBERT-small still has no explicit market-listing decision (pricing_v15.reference() raises), so an UNLISTED_BASE_MODELS entry is the release lane's.

p50 latency: 1.0 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights zaitlabs/watt-flash-0.1 @ d49a3a880b0201b325f2d2299d4c36ffc8735477 (Apache-2.0, public and ungated), model.safetensors sha256 099e611bd27c9542732b89be47bcace54964c00d9c7161ebccb061648c9780bf = THE SUBMITTER'S OWN DECLARED HASH, loaded from a local snapshot with the author's own --model flag (no network during the run). Code github.com/zaitgeist-charly/zaitlabs-watt-flash-0.1 @ e9863f06230ad874a5a2cfca3af922e45161a4d7 (Apache-2.0, = HEAD of main), installed with the author's own line `uv sync` against their own uv.lock in an isolated venv (torch 2.14.0, transformers 5.17.0, safetensors 0.8.0) and served with their own documented command `python -m watt_flash.serve --port 8942`. No file of theirs was edited. Their server answers POST /v1/systemone with Jev's shapes (noul -> {noul: P(true)}; choice probabilities keyed by the criteria keys with choice = argmax; score probabilities keyed '0'..'n-1'), so the UNCHANGED typesafe adapter applies and no mapping is ours; all eight checks of the unchanged smoke gate passed on synthetic requests before any sealed item was sent. THE AUTHOR'S FREE HOSTED ENDPOINT api.inference.zaitlabs.com WAS NEVER USED - the submission states it may train on the data, so no sealed item reached it. THE LONG ITEMS ARE REFUSED, NOT TRUNCATED, by the author's own code: watt_flash/layout.py raises Overflow past its documented 8,192-token window and serve.py answers that with HTTP 422, which their own commit e9863f06 says is deliberate ('answer an over-long request with 422, which JevBench scores as a refusal, not an outage'); those items are counted wrong and no other item failed. A RESUBMISSION of GitHub #113 via the form, which the author asks to dedupe with it. RUN ON SANDY'S SHARED CPU (12-core Zen 2, AVX2, fp32 - the package itself calls model.float() off-accelerator, 8 threads, nice), server and driver inside a loopback-only network namespace (unshare -rn) with no credential in the environment; load average logged every 60 s, so the Speed axis is a DECLARED LOWER BOUND. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, offline. add-requests run 61, benchmarkheaven.com/submit ref 0874B655 (bh-submit 4cf0bdfb-b6d7-4777-a8ac-350d29c5ac3f)

Published source