Manchego v2.1
system-one-open · by oraculumai · Code and weights marked open
Base model: Qwen/Qwen3.5-4Bsourcev1.5 roster addendum A5
Apache-2.0 (code and weights, author-declared)
JevBench v1.5.4 score
68.808
Option A: rank #11 of 106 ranked systems.
The three option scores and ranks are published independently; the headline is Option A.
| Option | Score | Rank |
|---|---|---|
| A · headline | 68.808 | #11 of 106 |
| B | 64.380 | #14 of 106 |
| C | 50.108 | #20 of 106 |
Published axes
- intelligence
- 51.2
- calibration
- 84.9
- speed
- 88.6
- cost
- 64.3
Bands use the published 0–100 axis values; the marked reference is Jev 1.13.0 when that axis is available.
Run and cost evidence
- Run status
- complete · 1,624 decisions · 0 missing
- Cost
- $0.015 per 1,000 decisions · estimate · ESTIMATE: frozen 25 Sep DeepInfra Qwen/Qwen3.5-4B market reference, already marked deprecated at that cutoff (replaced by Qwen/Qwen3.5-9B). $0.03/M input, zero generated output; 836,500 measured input tokens across 1,624 decisions. Frozen v1.5 base-model floor applied; no bookable Manchego tariff claimed. Re-score if the basis changes.
- Median latency
- 0.319 seconds, adjusted · x2 + 0.15 s (assumption, not measured)
- Endpoint condition
- evaluator-owned offline RTX6000 GPU
- Published source
- https://github.com/nschlaepfer/manchego-serve
Model and serving disclosure
Author server manchego-serve v0.1.1 @ 0ee62188af474e3c1d11d47a9c4c8b3a08766db3; oraculumai/Manchego @ 77403228b7dfdf823af99a5f562bdcf80b708d4c. Evaluator-owned RTX6000, offline/read-only container, safetensors with verified weight and support-file checksums, bfloat16, sequential execution, temperature 1.0 and original option order. Stock TypeSafe adapter; all three types native. The author disclosed using public aggregate/family results during development, a prior Jev-assisted phrase filter on two template pools, and one Jev-labelled development group for checkpoint selection. These disclosures do not establish item-text leakage; this is not a claim of fully independent development. Cost is an estimate using measured token usage and the frozen Qwen3.5-4B reference, not a bookable Manchego tariff.
Values come from the public v1.5.4 aggregate. Scores and ranks may change in a later release.
Read the full leaderboard, the v1.5.4 release page, and the published method.