Mercury Decide (Inception; System One decisions API, served free on OpenRouter as inception/mercury-decide:free)
Data: JevBench v1.6.1, published 2026-10-06
Capability 73.6 · API offering, ranked on the API leaderboard
Last measured: 2026-10-07 · measurement release: v1.6.1
Licence not stated
intelligence
65.9
calibration
81.3
speed
85.5
cost
62.1
Composite (secondary)
72.4 · Not ranked on this board
Cost and measurement conditions
$0.018 per 1,000 decisions (estimated)
LABELLED ESTIMATE, not charged - and nothing WAS charged: usage.cost is 0 on every one of the 1,500 calls. Inception publishes NO paid tariff for Mercury Decide: OpenRouter lists it free-only (prompt 0, completion 0) and the non-free slug inception/mercury-decide has no endpoints at all (404 'No endpoints found', read 7 Oct 2026), so there is nothing bookable. Because the system is API-only with undisclosed weights, the open rows' size-class estimate table cannot apply either. The row is therefore priced from THE SAME OPERATOR'S cheapest published per-token tariff on the SAME platform, inception/mercury-2.5 at USD 0.04 per 1M input and USD 0.15 per 1M output, which IS a key in the frozen 25 Sep snapshot - x this row's own measured tokens. Output tokens are 3 per decision on average (4,434 over 1,477 answers) because the probability is read from the model rather than written out, so the output rate is nearly inert and the Cost axis is set by the input rate alone. RELEASE-LANE LINE, recorded and not acted on, and it is RUN 57'S SHAPE EXACTLY - one mapping line, not a new price: the rate inception/mercury-2.5 (0.04, 0.15) is already IN the dated snapshot; what is missing is only a base-model decision line pointing at it, so pricing_v15.reference() raises on every inception/* string. receipts/COST-ALTERNATIVES-R58.json has the two alternatives' arithmetic already done.
p50 latency: 0.3 s
none (API)
measured 2026-10-07 on the live v1.6.0/v1.6.1 pool (inputs/input-selfhosted-1500.jsonl, sha256 901983ae...), COMPLETE: 1,500 of 1,500 dispatched, 1,477 answered, 23 invalid. Measured through the release lane's reviewed run_api_full_v161.py UNCHANGED, on Florian's 5 Oct amendment (hosted systems answer the same full S u P as self-hosted ones); provider inception, proxy openrouter, every sealed item exposure-logged before the first request. Florian work order of 6 Oct 2026 ('put it on JevBench', under https://x.com/airesearch12/status/2107569862486659111). NO MAPPING DECISION OF OURS: Mercury Decide publishes the same /v1/systemone contract as Jev and all three primitives come back in the canonical shape (noul a float; choice a label plus probabilities keyed by the item's own option labels; score probabilities keyed by level index '0'..'n-1' - 0-based, unlike ClassOne's), so the vendored typesafe adapter's request body, answer mapping and validations are reused UNCHANGED. or_decisions_r58.OpenRouterDecisionsAdapter rewrites ONLY the URL, because OpenRouter serves the contract at POST /api/alpha/decisions rather than /v1/systemone - its own 400 from chat/completions names that path - and it does so by calling the parent run(), so none of the reviewed mapping is duplicated. The pinned harness/vendor tree was NOT edited; the module is reached by PYTHONPATH. THE 23 LONG ITEMS: on this pool exactly 23 of 1,500 items are 77,857-82,318 tokens and the 24th largest is below 2,048 (run 57's own per-item counts on the identical pool, receipt LONGITEM-TOKENS-R58.json). They are also the 23 items the v1.6.1 cost rule excludes. THIS ROW REFUSES THEM: the context window is 32,768 tokens and all 23 came back as HTTP 422 with the provider's own 'upstream_error: Decision service could not complete the request' - a disclosed refusal, which the harness counts as the system's own rather than as an outage. The invalid set is EXACTLY those 23 and nothing else failed in 1,477 answers. RUN IN TWO PASSES BECAUSE OF A FREE-TIER BURST LIMIT, AND IT DOES NOT MOVE THE SPEED FIGURE: items 1-736 ran back to back at ~1.7 req/s until OpenRouter returned a header-less 429 ('rate_limit_exceeded: Request or spending limit exceeded'), which is run_v16's own stop rule; the row was resumed at 1 req/s and completed. The two passes' medians differ by 0.038 s (0.3624 s vs 0.3246 s over 1,477 answers; 0.3624 vs 0.3104 on the 124 items the Speed axis actually uses), and run_v16 sleeps the spacing BEFORE starting the timer, so the pacing is not in the measurement. Receipt LATENCY-SEGMENTS-mercury-R58.txt. Production API, so NO latency adjustment. Serial, one decision per request, no retry, no batching. add-requests run 58