JevBench v1.5.4 · individual system

Deem 0.8B v1

system-one-open · by LibertAI · Code and weights marked open

Base model: Qwen/Qwen3.5-0.8Bsource

v1.5 roster addendum A4

Not independently recorded

JevBench v1.5.4 score

2.138

Option A: rank #75 of 106 ranked systems.

The three option scores and ranks are published independently; the headline is Option A.

Published JevBench option scores and ranks
OptionScoreRank
A · headline2.138#75 of 106
B1.664#75 of 106
C1.485#75 of 106

Published axes

intelligence
13.0
calibration
38.9
speed
84.3
cost
80.3

Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.

Run and cost evidence

Run status
complete · 1,624 decisions · 0 missing
Cost
$0.0045 per 1,000 decisions · estimate · ESTIMATE: documented hosted-model estimate. Frozen market reference deepinfra:Qwen/Qwen3.5-0.8B.
Median latency
0.501 seconds, adjusted · x2 + 0.15 s (assumption, not measured)
Endpoint condition
evaluator-owned H100 GPU
Published source
https://huggingface.co/LibertAIDAI/deem-0.8-v1
Model and serving disclosure

LibertAIDAI/deem-0.8-v1 @ 8cbabbb2c4a7ef13c6b43f0ef3ae4157983c6d21; root model.safetensors SHA-256 80438125177c0681e855158d46720cbbd4ff07ea620cd55b0f456aafbfcf778d. Author Python GPU server, default DEEM_N_ORDERS=1, raw softmax (temperature 1.0). A four-field shim translates the wire format. The submitted interface accepts choice labels without descriptions and noul instructions without criteria; score receives full level descriptions. The author Rust/AVX-512 latency claim was not measured.

Values come from the public v1.5.4 aggregate. Scores and ranks may change in a later release.

Read the full leaderboard, the v1.5.4 release page, and the published method.