SPX-CD Flash (SurdAI, hosted /v1/systemone, Oct-4 checkpoint)
Data: JevBench v1.6.1, published 2026-10-06
Capability 74.3 · API offering, ranked on the API leaderboard
Last measured: 2026-10-06 · measurement release: v1.6.1
Licence not stated
intelligence
58.4
calibration
90.2
speed
79.4
cost
52.1
Composite (secondary)
66.7 · Not ranked on this board
Cost and measurement conditions
$0.039 per 1,000 decisions (estimated)
Operator's own public-beta overage tariff, conservative reading USD 0.04 per 1M input tokens, output USD 0 (the provider's /v1/models payload reports output_price "0"): three dated primary readings of the operator's own published overage tariff: 24 Sep 2026 provider console (flash 0.025, pro 0.08 per 1M input), 1 Oct 2026 docs page (flash 0.04, pro 0.02), 6 Oct 2026 docs page AND /v1/models payload agreeing (flash 0.02, pro 0.04). Sources conflict, so the per-model MAXIMUM is used, which is the published fastino-gliner-2-5-decide precedent ("published sources conflict ... conservative higher"). Today's agreeing pair would be USD 0.02/M, i.e. half this row's cost. No base-model floor applies: the submitter states the base model stays sealed until a later open release, so no base model exists to reference. Priced on usage.input_tokens as the provider reports them. The provider also reports billable_input_tokens = input_tokens / 2 on every single item (297/297 of the public set, exactly 2.0): its default effort 2 renders the prompt twice and its docs bundle the second pass free during public beta ("default effort 2 is included at the single-pass input price"). That bundle is a public-beta promotion, and TASK.md forbids a free or promotional tier setting the price, so the two passes actually consumed are charged. The operator's own billed figure is exactly half of this row's cost
p50 latency: 1.0 s
none (API)
measured 2026-10-06 on the live v1.6.0 S u P (1,500 rows) under Florian's 5 Oct 2026 full-set amendment, through the release lane's own full-set API lane (run_api_full_v161.py, rotation.py expose --full-set logged 1,200 sealed items to provider 'surd' before the first request); stock jevbench.adapters.typesafe with an endpoint override only, reasoning at the provider default, one request at a time, no retries; add-requests run 53, benchmarkheaven.com/submit ref 3FF00545