WaterSheep (Samrat Dutta, ModernBERT-base encoder with calibrated heads, served by the author's own `watersheep --serve`)
Data: JevBench v1.6.1, published 2026-10-06
Capability 35.6 · rank #73 of 87 on the open-weights board
Last measured: 2026-10-07 · measurement release: v1.6.1
apache-2.0
intelligence
5.3
calibration
65.8
speed
77.6
cost
68.7
Composite (secondary)
0.2 · Composite rank #102
Cost and measurement conditions
$0.011 per 1,000 decisions (estimated)
ESTIMATE, labelled, and RUN 48'S RATE FOR THIS EXACT ROW AND BASE, CARRIED UNCHANGED: answerdotai/ModernBERT-base IS in the frozen pricing_v15.UNLISTED_BASE_MODELS, so no exact base-model floor applies and the module takes the estimate rate from the row's own registry entry; USD 0.03/M input, 0 output is the rate run 47 (quyet-10-small-en) and run 48 (watersheep on v1.5) already assigned to this base. Output 0 is literal: watersheep/infer.py returns output_tokens 0 and the serving path has no generation step (02 Oct deputy rule). x this row's own measured input tokens. Estimated, not charged. No new decision.
p50 latency: 1.0 s
x2 + 0.15 s (assumption, not measured)
measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,500 of 1,500 answered, 0 invalid. RE-MEASUREMENT on v1.6 of the row run 48 measured on the retired v1.5 basis, same weights: samratduttaofficial/WaterSheep @ ebf3705d2d9dc648d5cc6b4cbca7646c6ab48998 (Apache-2.0), package github.com/SamratDuttaOfficial/WaterSheep @ 3ccf02b3 via the author's own install line in an isolated venv (torch 2.14.1+cpu, transformers 5.19.0), served with the author's own command `watersheep --model samratduttaofficial/WaterSheep@ebf3705d... --serve --device cpu`. Their server speaks POST /v1/systemone with Jev's shapes, so the UNCHANGED typesafe adapter applies; all eight checks of the unchanged smoke gate passed on synthetic requests before any sealed item was sent. Their own max_len 512 (watersheep.json) truncates every longer state, the 23 long items included, without refusing; not overridden. RUN ON SANDY'S SHARED CPU (12-core Zen 2, 8 threads) with server and driver inside a loopback-only network namespace and no credential in the environment; load average 8.6-15.4 (median 13.8) logged every 60 s, so the Speed axis is a DECLARED LOWER BOUND. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, offline. add-requests run 59, benchmarkheaven.com/submit ref 6F1561FC; previously GitHub #179