JevBench v1.5.4 · individual system

Manchego v2.1

system-one-open · by oraculumai · Code and weights marked open

Base model: Qwen/Qwen3.5-4Bsource

v1.5 roster addendum A5

Apache-2.0 (code and weights, author-declared)

JevBench v1.5.4 score

68.808

Option A: rank #11 of 106 ranked systems.

The three option scores and ranks are published independently; the headline is Option A.

Published JevBench option scores and ranks
OptionScoreRank
A · headline68.808#11 of 106
B64.380#14 of 106
C50.108#20 of 106

Published axes

intelligence
51.2
calibration
84.9
speed
88.6
cost
64.3

Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.

Run and cost evidence

Run status
complete · 1,624 decisions · 0 missing
Cost
$0.015 per 1,000 decisions · estimate · ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes.
Median latency
0.319 seconds, adjusted · x2 + 0.15 s (assumption, not measured)
Endpoint condition
evaluator-owned offline RTX6000 GPU
Published source
https://github.com/nschlaepfer/manchego-serve
Model and serving disclosure

Author server manchego-serve v0.1.1 @ 0ee62188af474e3c1d11d47a9c4c8b3a08766db3; oraculumai/Manchego @ 77403228b7dfdf823af99a5f562bdcf80b708d4c. Evaluator-owned RTX6000, offline/read-only container, safetensors with verified weight and support-file checksums, bfloat16, sequential execution, temperature 1.0 and original option order. Stock TypeSafe adapter; all three types native. The author disclosed using public aggregate/family results during development, a prior Jev-assisted phrase filter on two template pools, and one Jev-labelled development group for checkpoint selection. These disclosures do not establish item-text leakage; this is not a claim of fully independent development. Cost is an estimate using measured token usage and the frozen Qwen3.5-4B reference, not a bookable Manchego tariff.

Values come from the public v1.5.4 aggregate. Scores and ranks may change in a later release.

Read the full leaderboard, the v1.5.4 release page, and the published method.