ClassOne Gemma 4 E2B (Gemma-4-E2B backbone with ClassOne decision heads, schema-constrained single forward pass, self-hosted)

Data: JevBench v1.6.1, published 2026-10-06

Capability 15.2 · rank #85 of 87 on the open-weights board

Last measured: 2026-10-07 · measurement release: v1.6.1

apache-2.0

Base model: google/gemma-4-E2B-itsource

intelligence

7.8

calibration

22.6

speed

93.3

cost

70.2

Composite (secondary)

0.5 · Composite rank #96

Cost and measurement conditions

$0.0098 per 1,000 decisions (estimated)

ESTIMATE (same-family reference, one size up; self-hosted open weights, no bookable listing for this base). pricing_v15.reference() RAISES on google/gemma-4-E2B and on google/gemma-4-e2b - the base is in none of the three frozen tables - and the published page carries no priced row on this base either: the one Gemma-4-E2B entry on the board, system-one-open, is listed not_measured. The row therefore carries its own labelled estimate from the nearest reference the frozen 25 Sep snapshot itself contains, deepinfra:google/gemma-4-E4B-it at USD 0.02 per 1M input and USD 0.10 per 1M output - the SAME Gemma-4 effective-parameter family, one size LARGER, so it is both the nearest available reference and the direction that cannot flatter this row - x this row's own measured tokens. Estimated, not charged. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, not by estimate: the readout is a distribution over candidate labels from one forward pass and nothing is generated, so the row's cost rests on the input rate alone and any other decision is reversible by arithmetic from the raw. RELEASE-LANE LINE, recorded and not acted on: google/gemma-4-E2B needs either a listing decision or an UNLISTED_BASE_MODELS entry in the frozen table before this row may be published; receipts/COST-ALTERNATIVES-R58.json has the alternatives' arithmetic already done.

p50 latency: 0.2 s

x2 + 0.15 s (assumption, not measured)

measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (sha256 901983ae...), 1,477 of 1,500 answered. Weights devops-thiago/classone-gemma4-e2b (sha cd502aa33d288e67f16937c2e125f2297ab7b4ad), Apache-2.0, public and ungated; code github.com/devops-thiago/class-one, Apache-2.0. ONE DOCUMENTED MAPPING OF OURS, AND IT IS THE ONLY ONE. The author's server keys SCORE probabilities '1'..'n' while JevBench's canonical keys are '0'..'n-1'. The UNAMENDED smoke gate REFUSED the first attempt at check score-keys-are-indices, before any sealed item was sent, and the server was torn down with the card verified free. The 1-based convention - which their own schemas.py documents only as 'Probability distribution across each ordered rubric level', without stating the convention - was then established on the AUTHOR'S OWN SERVER with SYNTHETIC requests across every rubric size their own validator allows (2..10): exactly the keys '1'..'n' in order every time, never '0' and never a key above n. It is corroborated independently by their own `score` field being a 1-based probability-weighted mean of the same distribution (1.0185 on a 2-level rubric, 1.6147 on a 10-level one, 3.1155 where the mass sits on key 3; a 0-based convention would have given 0.0185), and a reversed rubric moved the mass to a different key, so 'k' means the k-th level of the criteria array AS SENT - which is what JevBench's 'k-1' means. classone_reindex_r58.ClassOneReindexAdapter therefore relabels score keys k -> k-1 and NOTHING else: the request body, the answer mapping, the noul range check and the choice-must-be-one-of-the-item's-labels rule are the UNCHANGED vendored typesafe adapter's, reused by calling its own run(). An item whose score key set is not exactly {'1'..'n'} is REFUSED and recorded invalid rather than rescored on a guess; none was. choice and noul are untouched. The gate then ran AMENDED in exactly that one check, requiring the key set to be exactly {'1'..'n'} - stricter than the original, because it also pins the count - with the other seven byte-identical, and passed. PREREG-R58-AMENDMENT-1.md (sha256 in PREREG-R58-AMENDMENT-1.md.sha256), receipt PROBE-classone-e2b.json. No file of the author's was edited and no prompt of ours exists. THE SUBMIT FORM CARRIED NO DESCRIPTION FOR EITHER CLASSONE ROW, so everything about how it is run comes from the author's own repository: install line README Quickstart step 1 (`pip install -e ".[dev]"`, in an isolated venv - torch 2.14.1+cu130, transformers 5.19.0, fastapi 0.142.2), serve line step 3 (`uvicorn classone.server.app:app`), and the checkpoint selected the way their own src/classone/server/app.py reads it (CLASSONE_BASE_MODEL). THEIR SERVER REGISTERS POST /v1/systemone ITSELF (app.py:115, next to /v1/decide and /v1/classone) and their DecisionRequest/DecisionResponse in src/classone/schemas.py are the TypeSafe System One shapes, so no transport shim exists anywhere in the path. One RTX 5090 32 GB (Lium, Paris FR). Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. output_tokens is 0 by their own schema comment ('In System 1 models this is 0 because there is no autoregressive decode'). THE 23 LONG ITEMS: on this pool exactly 23 of 1,500 items are 77,857-82,318 tokens and the 24th largest is below 2,048 (run 57's own per-item counts on the identical pool, receipt LONGITEM-TOKENS-R58.json). They are also the 23 items the v1.6.1 cost rule excludes. THIS ROW REFUSES THEM at a limit far below them: all 23 came back as 'HTTP 400: Prompt length (81338 tokens) exceeds maximum limit (8192)', an 8,192-token ceiling the author's code enforces with an explicit error rather than truncating. The 23 invalid are the system's own refusal, counted wrong, and no other item failed. Serial, one decision per request, no retry, no batching, credential-free. add-requests run 58, benchmarkheaven.com/submit ref C4D7AACD (bh-submit 4567a065-e912-40df-9bc1-68fb37cc7bba)

Published source