Liquid d1-3B (Liquid AI)
LLM decoderBase model: undisclosedlocal/GPU evaluation
Image JevBench v0.3.0 composite score
4.096
Rank #39 of 54 ranked systems.
From the published full-benchmark aggregate. The composite combines intelligence, calibration, speed and cost.
Published axes
- intelligence
- 16.3
- calibration
- 79.3
- speed
- 87.6
- cost
- 54.1
Cost and speed evidence
- Cost per 1,000 decisions
- USD 0.033932
- Cost basis
- measured GPU seconds x $1.19/GPU-hour
- Measured latency · p50 / p95
- 0.100 s / 0.173 s
- Adjusted latency · p50 / p95
- 0.350 s / 0.497 s
- Latency adjustment
- v1.4 local latency adjustment
- Public accuracy
- 43.4%
- Sealed accuracy
- 43.7%
- Measurement note
- Official ImageJev v0.3 S1200+P300 pass 1 on 10 Oct 2026 (no new draw: the same 1,500 items as every other self-hosted row) on one RTX PRO 6000 with the same reviewed runtime as the other self-hosted rows, offline after the weights were staged and hash-verified. Model: LiquidAI/d1-3B at revision 051bcc46. Request: the NeoHorse-shaped single choice question through the model's own system_one() call with the item image; bfloat16 as recommended by the model authors. 1,500/1,500 ok, no retries. No context limit is applied by the model code (config 32,768 positions). Same vendor as the hosted API row liquid-d1 on JevBench (text); the two rows are not linked here because the page data has no family field.
Published 2026-10-10. Scores and ranks can change in a later release. See the method notes.