Image JevBench v0.3.0 · individual system

Liquid d1-3B (Liquid AI)

LLM decoderBase model: undisclosed

Published source

local/GPU evaluation

Image JevBench v0.3.0 composite score

4.096

Rank #39 of 54 ranked systems.

From the published full-benchmark aggregate. The composite combines intelligence, calibration, speed and cost.

Published axes

intelligence
16.3
calibration
79.3
speed
87.6
cost
54.1

Cost and speed evidence

Cost per 1,000 decisions
USD 0.033932
Cost basis
measured GPU seconds x $1.19/GPU-hour
Measured latency · p50 / p95
0.100 s / 0.173 s
Adjusted latency · p50 / p95
0.350 s / 0.497 s
Latency adjustment
v1.4 local latency adjustment
Public accuracy
43.4%
Sealed accuracy
43.7%
Measurement note
Official ImageJev v0.3 S1200+P300 pass 1 on 10 Oct 2026 (no new draw: the same 1,500 items as every other self-hosted row) on one RTX PRO 6000 with the same reviewed runtime as the other self-hosted rows, offline after the weights were staged and hash-verified. Model: LiquidAI/d1-3B at revision 051bcc46. Request: the NeoHorse-shaped single choice question through the model's own system_one() call with the item image; bfloat16 as recommended by the model authors. 1,500/1,500 ok, no retries. No context limit is applied by the model code (config 32,768 positions). Same vendor as the hosted API row liquid-d1 on JevBench (text); the two rows are not linked here because the page data has no family field.

Published 2026-10-10. Scores and ranks can change in a later release. See the method notes.