CoCo-Decision-4B (corners-ai, LoRA on Qwen3.5-4B, option-label logits of one pass, served by oh-my-jev)

Data: JevBench v1.6.1, published 2026-10-06

Capability 56.9 · rank #41 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: Qwen/Qwen3.5-4Bsource

intelligence

32.7

calibration

81.2

speed

91.4

cost

65.2

Composite (secondary)

24.6 · Composite rank #47

Cost and measurement conditions

$0.014 per 1,000 decisions (estimated)

ESTIMATE (base-model reference; self-hosted open weights). NO JUDGEMENT OF OURS: pricing_v15.reference() resolves Qwen/Qwen3.5-4B straight through to deepinfra:Qwen/Qwen3.5-4B at USD 0.03 per 1M input and USD 0.15 per 1M output in the frozen 25 Sep snapshot (the same rate aplomb-1 carries for the same base) x this row's own measured tokens. Output tokens are 0 in every answered record: one prefill, option-label readout, nothing generated. Estimated, not charged.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,477 of 1,500 answered. LoRA adapter corners-ai/CoCo-Decision-4B-Ko tag v1.0.0 = commit b580ede993493690212a07ed0e6cb972868cc808 (the submission's 'Version 1.0.0', the card's own pin; Apache-2.0) on Qwen/Qwen3.5-4B @ 851bf6e8 as its own adapter_config.json names it. Toolkit github.com/iamupd/oh-my-jev @ f0b8c786 (Apache-2.0) installed with its README's `uv sync` plus the semif extra the README names for --adapter (torch 2.9.1+cu128, transformers 5.17.0, peft 0.21.0). Served by the author's own `omj serve --backend semif --adapter <adapter dir>`: the README documents that --adapter runs locally on semif and loads the adapter's base, in 4-bit only when it would not fit; on the 32 GB card omj itself chose bf16 ('Using Qwen/Qwen3.5-4B (the adapter's base model, bf16)'), the card's headline configuration. NO config.toml was authored and no setting of ours was chosen, which removes run 58's blocker. UNCHANGED typesafe adapter; all eight checks of the unchanged smoke gate passed on synthetic requests before any sealed item was sent. THE 23 LONG ITEMS ARE REFUSED with the documented 422 context_length_exceeded at omj serve's default 4,096-token limit (README 'Known limitations'); they are the system's own disclosed refusal, counted wrong, and no other item failed. One RTX 5090 32 GB (Lium, France). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. Serial, one decision per request, no retry, no batching, credential-free. add-requests run 59, benchmarkheaven.com/submit ref 619F3346 (bh-submit 0195e1e7)

Published source