Bobcat Flash 1.2 (Gemma-4-26B-A4B-it + two merged rank-64 LoRA adapters, typed-decision readout at the first answer position, FP8, self-hosted)
Data: JevBench v1.6.1, published 2026-10-06
Capability 73.5 · rank #6 of 92 on the open-weights board
Last measured: 2026-10-07 · measurement release: v1.6.1
apache-2.0
intelligence
65.8
calibration
81.2
speed
90.8
cost
49.7
Composite (secondary)
67.3 · Composite rank #14
Cost and measurement conditions
$0.048 per 1,000 decisions (estimated)
ESTIMATE (base-model reference; self-hosted open weights, no public tariff): OpenRouter google/gemma-4-26b-a4b-it in the frozen 25 Sep snapshot, USD 0.09 per 1M input and USD 0.3 per 1M output x this row's own measured tokens. Estimated, not charged. THIS IS THE REFERENCE, THE RATE AND THE LABEL THE PUBLISHED PAGE ALREADY CARRIES FOR THIS EXACT BASE on three rows published within the last day - xor-26b-a4b, surogate-rune-26b-a4b-v3-v16 and gev-26b-decide - so pricing this row any other way would make it inconsistent with the field it is placed in. RELEASE-LANE LINE, recorded and not acted on: pricing_v15.reference() RAISES on the author's exact declared spelling google/gemma-4-26B-A4B-it, because that string is in none of the three decision tables, while the lowercase google/gemma-4-26b-a4b-it IS in BASE_REFERENCES and in the dated snapshot at (0.09, 0.30). That is a CAPITALISATION gap - one notch tighter than run 56's -Base/instruct gap on Qwen3.5-35B-A3B - and the three published rows say in their own basis string that the frozen module has no explicit listing-decision line for this base. OUTPUT TOKENS ARE 0 BY CONSTRUCTION, not by estimate: their build_response returns usage.output_tokens == 0 and the readout takes candidate logits at the first answer position of one forward pass, so this row's cost depends on the input rate alone and any other decision is reversible by arithmetic from the raw.
p50 latency: 0.3 s
x2 + 0.15 s (assumption, not measured)
measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...). Weights sanghwa-na/bobcat-flash-1.2@58d15456937ce58625b13a5e3d982846554a3aef, Apache-2.0, 51.68 GB BF16 - a merge of two rank-64 LoRA adapters on google/gemma-4-26B-A4B-it@4d7ae4984b7db7de8f8457170b3f1a419ee76d52 (0.6*A + 0.4*B, B being the Flash 1.1 adapter), file-by-file verified against the Hub's own manifest before the engine was started. Code github.com/foxl-ai/bobcat@43387a09bbc0310edcf57e2b21fa59ea81ac4166, Apache-2.0. NO MAPPING DECISION OF OURS EXISTS FOR THIS ROW, which is why it is worth saying so explicitly. The author implements TypeSafe's published System One contract, and their own src/bobcat/protocol.py parse_request accepts the JevBench canonical record EXACTLY as the UNCHANGED jevbench.adapters.typesafe already emits it: choice -> a dict of 1-255 named criteria, so the item's own labels are the keys and its own criteria[label] strings travel verbatim as the descriptions (their protocol renders {name, description} per option); score -> the item's ordered array of levels, which their own parse_request labels '0'..'n-1' - the same keys JevBench uses - with the item's own level text as each description, so the ordinal order is carried; noul -> criteria restricted to {true,false}, and the item's OWN true/false description strings are delivered (unlike run 56's APUS rows, whose runtime required canonical criteria there), with the answer returned as noul = P(yes). Their answer() refuses to reply unless the probabilities sum to one. Nothing was rewritten, no option was dropped, no item type was recorded as a miss, and no shim of ours sits anywhere in the path. STACK: the author's own published install line, verbatim from their model card's Quickstart - a uv-managed Python 3.12 (their note: it ships the headers Triton compiles against) and vllm==0.30.0 LETTING IT RESOLVE ITS OWN STACK, which gave torch 2.13.0+cu130, transformers 5.19.0, tokenizers 0.23.2. Served by THEIR OWN SERVER, bobcat.flash_server, with THEIR OWN published command verbatim: --engine vllm --quantization fp8 (their online dynamic FP8 of these BF16 weights, the served configuration) --temperature choice=0.594604,noul=0.353553,score=0.529732 --max-num-seqs 64 --max-model-len 98368 --schedule all --engine-arg max_num_batched_tokens=16384 --engine-arg attention_backend=TRITON_ATTN, with VLLM_USE_FLASHINFER_SAMPLER=0 as their command sets it. --local only because the server is on loopback on our own pod. One H100 PCIe 80 GB (Lium) - the second of the two cards the author measured. Latency axis uses the standard self-hosted x2 + 0.15 s adjustment. No file of theirs was edited and no prompt of ours exists. CALIBRATION IS THEIR PUBLISHED DEFAULT AND WAS NOT CHOSEN BY US. Their card ships two calibrations: the per-type triple used here, which is the Quickstart's - i.e. the published serve command of this release - and a single product temperature 0.9232 fitted by log-loss on their six product tasks. The triple was bound in PREREG-R57 s3.3 before any item was sent and the alternative was deliberately NOT probed on our items, because the author has ALREADY PUBLISHED both configurations' numbers on JevBench's public 231 (per-type Capability 68.2 = 57.5/78.9; product 68.3 = 54.3/82.3), so there was nothing to learn from our pool and choosing between them from our own measurement would have been exactly the mistake run 56's probe was designed to avoid. The author states the triple was fitted on a held-out split of 8,906 JevBench-FORMAT decisions (generated and public) for the highest JevBench-style Capability, and separately that no JevBench item was used for training, calibration, selection or prompt design, with generated rows audited against JevBench's public items at 0 exact and 0 8-gram overlaps. THE 23 LONG ITEMS: the server's limit is --max-model-len 98368 (their compiler gets max_model_len - 64) against a pool whose longest item is 128,111 BYTES of canonical record, and their RequestLimitError answers 422 with 'Inputs are never truncated' rather than truncating. This row therefore reads every item in full. Serial, one decision per request, no retry, no batching, offline, credential-free. add-requests run 57, benchmarkheaven.com/submit ref 7A714784 (bh-submit e9d576d7-e1f1-49d8-8151-05cdbbdd7ed4)