EXAONE-4.0-1.2B-JEV v0.3 (unofficial fine-tune of LGAI-EXAONE/EXAONE-4.0-1.2B, native option-label softmax, two reads averaged, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 47.7 · rank #53 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

EXAONE AI Model License 1.2 (non-commercial)

Base model: LGAI-EXAONE/EXAONE-4.0-1.2Bsource

intelligence

14.8

calibration

80.5

speed

94.1

cost

57.9

Composite (secondary)

3.3 · Composite rank #87

Cost and measurement conditions

$0.025 per 1,000 decisions (estimated)

ESTIMATE (size-class reference; self-hosted open weights, no bookable listing for this base). THE RATE IS RUN 48'S, CARRIED UNCHANGED FOR THE SAME BASE: run 48 priced exaone-4.0-1.2b-jev (v0.2, identical base) at USD 0.02 per 1M input and 0 output, as the HIGHER of the two bracketing published rates, chosen in that direction because the rate was fixed after the measurement existed. Keeping it makes v0.3 directly comparable with v0.2 and takes no new decision. The frozen 25 Sep snapshot still contains NO listing for LGAI-EXAONE/EXAONE-4.0-1.2B (pricing_v15.reference() raises: 'base model lacks an explicit market-listing decision'), and still brackets 1.2B between deepinfra:Qwen/Qwen3.5-0.8B at (0.01, 0.05) and deepinfra:Qwen/Qwen3.5-2B at (0.02, 0.10). Estimated, not charged. RELEASE-LANE LINE, recorded and not acted on, AND IT IS RUN 48'S CARRIED LINE, NOT A NEW ONE: LGAI-EXAONE/EXAONE-4.0-1.2B must be adopted into the frozen UNLISTED_BASE_MODELS table before either EXAONE row may be published. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, not by estimate: the readout is a distribution over candidate labels from one forward pass and nothing is generated, so the row's cost rests on the input rate alone and any other decision is reversible by arithmetic from the raw.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...), 1,477 of 1,500 answered. NEW REVISION of the row run 48 measured: weights carrtesy/EXAONE-4.0-1.2B-JEV at TAG v0.3, server github.com/carrtesy/EXAONE-JEV at tag v0.3 (commit recorded in receipts/POD-R58.json), both the artefacts the author's own submit-form entry names. EXAONE AI Model License 1.2 - NC, NON-COMMERCIAL, and an unofficial personal project - the author states it is 'not an LG AI Research release'; both belong on the row. NO MAPPING DECISION OF OURS: their serve/systemone_server.py speaks the TypeSafe Jev POST /v1/systemone contract, so the UNCHANGED jevbench.adapters.typesafe emits exactly what it parses, and all eight checks of the unamended smoke gate passed on synthetic requests before any sealed item was sent. Read count (permute 2: options in the given and the reversed order, distributions averaged) and the per-type temperatures come from the weights' own jev_config.json, which the author's doc states needs no flags - their shipped default, not a setting of ours. STACK: the author's own docs/REPRODUCE.md install line in an isolated venv, vllm==0.30.0 letting it resolve its own stack (torch 2.13.0+cu130, transformers 5.19.0), served with their own published command verbatim (`python serve/systemone_server.py --model weights --port 8000 --gpu-util 0.9`) and VLLM_USE_FLASHINFER_SAMPLER=0 as their doc sets it. One RTX 5090 32 GB (Lium, Paris FR). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. THE 23 LONG ITEMS: on this pool exactly 23 of 1,500 items are 77,857-82,318 tokens and the 24th largest is below 2,048 (run 57's own per-item counts on the identical pool, receipt LONGITEM-TOKENS-R58.json). They are also the 23 items the v1.6.1 cost rule excludes. THIS ROW REFUSES THEM, at the limit the author DECLARES: their docs/REPRODUCE.md section 3 states a 65,536-token prompt limit answering 400, and all 23 came back as 'HTTP 400: prompt of 84996 tokens is longer than the maximum context length 65536'. The 23 invalid are therefore the system's own disclosed refusal, counted wrong, and no other item failed. Serial, one decision per request, no retry, no batching, offline, credential-free. add-requests run 58, benchmarkheaven.com/submit ref CA53D793 (bh-submit 5631d258-1eaa-4e2d-a658-f6417b85b351); previously GitHub issue #176

Published source