JevBench v1.5.5 · individual system

lev (Interfaze AI, Qwen3.5-4B + LoRA r32, label-token readout + candidate-path head, per-bucket temperatures)

Jev-compatible decision model · by Interfaze · Openness: open weights

Base model: undisclosed

v1.5 roster addendum A6

apache-2.0

JevBench v1.5.5 score

69.061

Option A: rank #11 of 109 ranked systems.

The three option scores and ranks are published independently; the headline is Option A.

Published JevBench option scores and ranks
OptionScoreRank
A · headline69.061#11 of 109
B66.436#11 of 109
C63.800#8 of 109

Published axes

intelligence
57.7
calibration
77.9
speed
86.5
cost
61.8

Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.

Run and cost evidence

Run status
complete · 1,624 decisions · 0 missing
Cost
$0.019 per 1,000 decisions · estimate · documented hosted-model estimate; reference deepinfra:Qwen/Qwen3.5-4B frozen at the 25 Sep cut-off; the snapshot records deprecation on 11 Jun 2026 and replacement by Qwen/Qwen3.5-9B; the frozen snapshot rates remain in use (I-3); self-hosted estimate, not measured hosted billing
Median latency
0.403 seconds, adjusted · x2 + 0.15 s (assumption, not measured)
Endpoint condition
evaluator-owned Lium GPU pod (H100 80GB), offline (HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE), credential-free, server bound to 127.0.0.1, harness on the same pod over loopback HTTP; Lium pod 08a1ff4a-248a-4ff2-948e-ea558f99e808 (huid cosmic-hawk-4f, node calm-fox-65), 1x NVIDIA H100 80GB HBM3, United States, USD 1.30/h
Published source
https://huggingface.co/interfaze-ai/lev
Model and serving disclosure

Qwen/Qwen3.5-4B

Values come from the public v1.5.5 aggregate. Scores and ranks may change in a later release.

Read the full leaderboard, the v1.5.5 release page, and the published method.