jul fast (usejul/jul-decision-minicpm5-2b: MiniCPM5-2B fine-tune with a pointer head, option probabilities from one pass, self-hosted via `jul serve`)

Data: JevBench v1.6.1, published 2026-10-06

Capability 42.8 · rank #64 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: openbmb/MiniCPM5-2Bsource

intelligence

7.5

calibration

78.1

speed

90.7

cost

69.0

Composite (secondary)

0.5 · Composite rank #94

Cost and measurement conditions

$0.011 per 1,000 decisions (estimated)

ESTIMATE (size-class reference; self-hosted open weights, no bookable listing for this base). openbmb/MiniCPM5-2B is an EXPLICIT entry in the frozen module's UNLISTED_BASE_MODELS table, so pricing_v15.reference() returns None without raising: the decision that there is no bookable listing for this base is already taken, and the only MiniCPM key in either dated snapshot is deepinfra:openbmb/MiniCPM-Llama3-V-2_5 at (0.34, 0.34), a VISION model and not a comparable reference. The row therefore carries its own labelled size-class estimate: deepinfra:Qwen/Qwen3.5-2B at USD 0.02 per 1M input and USD 0.10 per 1M output in the frozen 25 Sep snapshot - a 2B dense instruct model, the same size class as this 2B base - x this row's own measured tokens. Estimated, not charged. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, not by estimate: the readout is a distribution over candidate labels from one forward pass and nothing is generated, so the row's cost rests on the input rate alone and any other decision is reversible by arithmetic from the raw. RELEASE-LANE LINE, recorded and not acted on: if a bookable MiniCPM5-2B text listing has appeared since 25 Sep, this row reprices by one multiplication from the raw. receipts/COST-ALTERNATIVES-R58.json has the arithmetic for the alternatives already done.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,498 of 1,500 answered. Library github.com/usejul/jul tag v0.5.1, PyPI jul==0.5.1, Apache-2.0; weights usejul/jul-decision-minicpm5-2b at f556a6936089693ea68bfed3080c637e4853a3ff, Apache-2.0; base openbmb/MiniCPM5-2B at 12a3808a. NO MAPPING DECISION OF OURS: `jul serve` speaks the TypeSafe /v1/systemone wire format, so the UNCHANGED jevbench.adapters.typesafe applies, and all eight checks of the unamended smoke gate passed on synthetic requests before any sealed item was sent. STACK: the author's own published install line from their request, `pip install "jul[torch]==0.5.1"` in an isolated venv (torch 2.14.1+cu130, transformers 5.19.0), served with their own published command verbatim (`jul serve --model jul-decision-minicpm5-2b --host 127.0.0.1 --port 8577`). One RTX 5090 32 GB (Lium, Paris FR). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. THE 23 LONG ITEMS: on this pool exactly 23 of 1,500 items are 77,857-82,318 tokens and the 24th largest is below 2,048 (run 57's own per-item counts on the identical pool, receipt LONGITEM-TOKENS-R58.json). They are also the 23 items the v1.6.1 cost rule excludes. THIS ROW IS THE FIRST THAT NEITHER REFUSES NOR UNIFORMLY TRUNCATES THEM, AND THE SPLIT IS BY QUESTION TYPE. The author declares a 2,048-token state budget with no refusal - 'the beginning of the text is kept and the end is cut, without an error'. Their server's own usage.input_tokens shows that budget applied to 17 of the 23 (counts 2,096-2,135) and NOT to the other 6, which were read at 107,331-107,997 tokens; those 6 are exactly the 6 score questions among the 23, while all 17 truncated ones are choice or noul. Two of the six full reads then failed with the server's own CUDA OutOfMemoryError on a 32 GB card ('Tried to allocate 5.59 GiB' / '5.41 GiB'), which is this row's only 2 invalid and which a larger card would most likely have answered: the cause is the score path not applying the author's own documented truncation, and the author's note that '24 GB is enough' holds only under that truncation. The 2 invalid are counted wrong, the conservative direction, and are 0.13 % of the pool. PROBABILITIES ARE RETURNED ROUNDED TO FOUR DECIMALS: 426 of 1,498 answered distributions sum to 1 +/- 2e-4 (observed range 0.9998-1.0002); none is off by more than that, so it is the server's rounding and not a partial readout. Serial, one decision per request, no retry, no batching, credential-free. add-requests run 58, GitHub fstandhartinger/jevbench#198

Published source