JevBench API leaderboard: hosted decision APIs
JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
This board ranks hosted API offerings: decision APIs and models we reached through an endpoint we do not run. Cost uses each row's documented reference price (the provider's tariff, or for an API with a known base model the developer's own list price; estimates are marked). Open-weights models are ranked on the main JevBench board; every score is identical on both boards.
Release v1.6.1 · 1,500 decisions per self-hosted system and 1,500 per hosted API · 4 ranked API offerings · 15 systems retain a separately dated v1.5.x score · only system-level aggregates are published · aggregate results JSON · SHA-256 5d4567d5e5acd945d17dd082adbbd6188d38174ca523b0fa6b3b02c6cc2dc5b1
Share this board · Previous release: JevBench v1.6.0
Making decisions from images? Explore Image JevBench v0.1.5 and compare its systems.
JevBench v1.6.1 · headline
JevBench Capability Score
Capability ranking of Jev-class systems
Capability Score averages Intelligence and Calibration. Sage 1.3.0 leads the Jev-class systems with 78.6.
Jev-class means at most 2× the cost and median latency of Jev (v1.5 reference). How we choose ↘
Adjust cost / latency caps · 2× official
ScoreCCost
ScoreCap.Capability
Score$/1k$/1k decisions
- 1Sage 1.3.0API65.658.278.6$0.025*
- 2Jev 1.13.0API63.654.777.1$0.032
- 3Fastino GLiNER-2.5-DecideAPI7.045.834.6$0.064
Wide coloured bar = Capability Score (0–100). Thin bars: cost above, median latency below; shared log ratio scale 0.25× … 64× Jev (v1.5 reference), ticks at 1× and 2× (cap). Shorter is cheaper or faster. * = est. (estimated cost). # counts ranked Jev-class systems; “–” marks unranked or outside systems. Tap ⓘ or hover a row or thin bar for details.
Cost and latency correlate here (Spearman ρ = -0.20, n = 4).
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- green: ≤ reference
- amber: ≤ cap (2× reference)
- red: > cap
Show general-purpose LLMs and other systems outside the limits
Sorted by Capability Score, not numbered. Each row says which limit it misses, measured against Jev ($0.032 per 1,000 decisions, median 0.62 s).
- –wity-1API70.358.879.1$0.024Outside: latency 2.6× Jev (v1.5 reference)API price · eligibility checked at the developer's list price; at base-model pricing it would exceed the cost cap (2.70× Jev)
Jev-class = cost per decision at most 2× Jev's (≤ $0.065 per 1,000 decisions) and median latency at most 2× Jev's (≤ 1.23 s), the adjusted p50 — the same median the speed chart plots, not the four-axis Speed score. 3 of 4 systems qualify; the other 1, including the general-purpose LLMs, are listed below the divider in the ranking. The charts below show speed and cost beside Capability Score; the official JevBench Score weighs all four axes.
Capability against cost and speed
Jev-class systems are shown by default. Bubble size follows the official JevBench Score; official rank stays unchanged. The five most capable Jev-class systems are labelled.
Capability vs cost
Upper right is best: more capable and cheaper. Cost is USD per 1,000 decisions on a log scale. The dashed line is 2× Jev’s cost.
Capability vs speed
Upper right is best: more capable and faster. Speed here is the median-latency speed — the same adjusted median (p50) latency the Jev-class limit uses — a log scale, so each 20 points is 10× faster (median under the numbers). The dashed line is 2× Jev’s median latency.
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- faint = outside Jev-class
Model kind
Jev-class
Release
No system in this release reports an exact parameter count.
No system in this release reports a base-model reference price.
JevBench v1.6.1
JevBench Composite Score: 4 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · Method notes ↓
Adjust weights ↓
The official order weighs Intelligence, Calibration, Speed and Cost equally. Each button re-sorts the same systems by one axis, and the column headings sort too.
Greener = stronger within its column.
4 of 4 systems, sorted by official rank, #1 first.
- 1Sage 1.3.0APInewClosed APIBase model: undisclosed74.0I 66C 92S 94K 58est.$0.025
- 2Jev 1.13.0APIJev referenceBase model: undisclosed71.5I 64C 91S 91K 55$0.032
- 3wity-1APInewClosed APIBase model: undisclosedBase-model reference price (base undisclosed at the author's request): 44.0 (would be #3)70.8I 70C 88S 72K 59$0.024
- 4Fastino GLiNER-2.5-DecideAPInewClosed APIBase model: undisclosed · Hosted Fastino API (model id fastino/GLiNER-2.5-Decide). Checked fastino.ai home page and the GLiNER2.5-Decide announcement blog: no statement of the hosted endpoint's base model. Open-weight HF cards fastino/GLiNER2.5-Decide (base_model fastino/gliner2-large-v1) and fastino/GLiNER2.5-Decide-1B (base_model fastino/gliner2-xl-0111) exist, but nothing public says which one (or a different build) the hosted API serves, so no base is claimed.0.3I 7C 62S 85K 46$0.064
Adjust weights ↓
Weights are relative: each axis counts in proportion to its slider. The score stays a weighted harmonic mean with the low-axis gates; the gates still apply when an axis sits at 0. Only equal weights give the official JevBench Score and rank.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Jev — reference (TypeSafe, closed) (1)
- Closed API (weights not public) (3)
- Striped bar = same system under the labelled alternative price assumption
Compare two systems
Pick any two. Radars for the score axes, capability by subject topic and use cases, chance-corrected competence per request type on the open and sealed sets, and competence per tier on each set. Further out is better on every spoke; the link keeps the pair.
- A: Jev 1.13.0 — Jev — reference (TypeSafe, closed) · Score 71.5 (#2)
- B: Sage 1.3.0 — Closed API (weights not public) · Score 74.0 (#1)
The four score axes
Capability by subject topic
What each category means · items per category
- Rules, policy & law — applying written rules: company policies, contracts, regulations, eligibility. 763 items (151 open / 612 sealed)
- Coding & software — code, SQL, repositories, developer tools and IT systems. 302 items (62 open / 240 sealed)
- Math & numbers — a calculation decides the answer: arithmetic, word problems, probability, dates, units. 193 items (37 open / 156 sealed)
- Support & operations — support tickets, incidents, logistics, scheduling desks and routing work to a team. 93 items (21 open / 72 sealed)
- Finance & commerce — money: payments, refunds, invoices, orders, expenses, insurance payouts. 66 items (12 open / 54 sealed)
- Everyday language — short everyday messages: intents, assistant requests, reading a detail out of a text. 46 items (11 open / 35 sealed)
- Safety & security — untrusted or injected instructions, fraud, moderation, access and security triage. 37 items (6 open / 31 sealed)
Use cases (TypeSafe categories)
Low sample, n < 30 — indicative only
| Category (items) | A: Jev 1.13.0 | B: Sage 1.3.0 |
|---|---|---|
| Feature extraction for predictive modeling (19) | 22.5 n=19 | 34.4 n=19 |
| Graphs and knowledge graphs (18) | 49.8 n=18 | 74.1 n=18 |
| LLM guardrails (17) | 44.1 n=17 | 65.7 n=17 |
| Search and retrieval (17) | 80.5 n=17 | 83.3 n=17 |
| Lead generation (17) | 67.5 n=17 | 75.5 n=17 |
| Semantic code linting (16) | 12.2 n=16 | 22.9 n=16 |
| Demand forecasting (16) | 0.0 n=16 | 0.0 n=16 |
No published value for either system (under 15 answered items, or no per-category values): Gaming (14), Scientific discovery (14), Recruiting (13), Advertising (11).
Not drawn on the radar: with so few items a single answer moves a category score by several points, so these values are noise-prone. These values become spokes once the item pool reaches 30 items per category.
What each category means · items per category
- Model routing — deciding which LLM or model tier should handle a prompt; classifying a request's intent, domain, difficulty or risk to pick a model. 442 items (86 open / 356 sealed)
- Legal and compliance — contracts, policies, regulations, eligibility rules, checking documents against written requirements. 387 items (74 open / 313 sealed)
- Customer support — classifying tickets and customer intents, urgency, refunds, routing cases to a team, checking support replies. 149 items (37 open / 112 sealed)
- E-commerce marketplaces — product listings, product attributes, orders and deliveries, seller catalog policy. 48 items (9 open / 39 sealed)
- Risk assessment — estimating probabilities and severity of risks, incidents, vendor or underwriting risk. 45 items (8 open / 37 sealed)
- Insurance claims — first-notice-of-loss reports, claim complexity, missing information, claim triage. 41 items (9 open / 32 sealed)
- Financial crime — suspicious transactions, KYC, fraud alerts, entity matching for investigations. 39 items (9 open / 30 sealed)
- Moderation and trust and safety — user content moderation: toxicity, harassment, spam, unsafe advice, personal data exposure. 32 items (7 open / 25 sealed)
Not drawn: Other — none of the above: general reasoning, arithmetic or date puzzles, everyday assistant requests without a business workflow. 145 items (24 open / 121 sealed) — not a use case of its own, so it is counted but not drawn.
Competence per request type, open / sealed
Competence per tier — open set
Competence per tier — sealed set
How the categories were made
Chance-corrected competence (0 = chance, 100 = perfect) over the category's items the system saw, open and sealed pooled, tiers pooled: per request type the method's competence cell, then the mean over the request types present. Not part of the JevBench Score; never equated for API systems; compare systems within a category, not categories with each other. Values can be negative (below chance).
Family and language are authoring metadata of every item in the frozen v1.6 pool. Each of the 1,500 v1.6 items was also labelled with one subject topic (the v1.2 topic list, unchanged) and one TypeSafe use-case category by Winnow-12B Q8 on our own GPU pod (same model, questions and taxonomy as v1.5); where a uc1 item was written for a use case, that authoring use case is used instead. Sealed items stayed on our own infrastructure. 75 public items (5 %) were checked by hand; item-group rules fix the systematic misses (topic agreement before rules 93 %). All non-English uc1 items are machine-authored and not native-reviewed. Self-hosted systems saw S u P (1,500 items); hosted API systems saw only A u P (600 items), so their cells cover fewer items.
- Cells with fewer than 15 items are omitted; categories under 30 items are marked low-n.
- Core v1.6 items carry no use case; the use-case view covers the use-case (uc1) items only.
All values as a table
| Spoke | A: Jev 1.13.0 | B: Sage 1.3.0 |
|---|---|---|
| The four score axes | ||
| Intelligence | 63.6 | 65.6 |
| Calibration | 90.6 | 91.6 |
| Speed | 91.5 | 93.6 |
| Cost | 54.7 | 58.2 |
| Capability by subject topic | ||
| Rules, policy & law | 70.6 | 68.1 |
| Coding & software | 74.5 | 80.6 |
| Math & numbers | 30.5 | 37.9 |
| Support & operations | 34.9 | 60.9 |
| Finance & commerce | 60.3 | 66.3 |
| Everyday language | 51.5 | 47.4 |
| Safety & security | 35.2 | 65.3 |
| Use cases (TypeSafe categories) | ||
| Model routing | 78.2 | 83.2 |
| Legal & compliance | 81.3 | 70.3 |
| Customer support | 47.6 | 63.3 |
| E-commerce | 44.6 | 66.5 |
| Risk assessment | 42.1 | 33.7 |
| Insurance claims | 76.5 | 65.1 |
| Financial crime | 36.0 | 35.0 |
| Moderation | 53.4 | 70.5 |
| Competence per request type, open / sealed | ||
| Choice · open | 79.0 | 85.1 |
| Choice · sealed | 75.8 | 79.7 |
| Noul · open | 54.5 | 54.3 |
| Noul · sealed | 46.1 | 56.4 |
| Score · open | 63.3 | 57.7 |
| Score · sealed | 63.0 | 60.5 |
| Competence per tier — open set | ||
| Easy | 93.9 | 89.6 |
| Standard | 63.2 | 62.3 |
| Judge | 67.7 | 67.5 |
| Hard | 65.6 | 70.5 |
| Competence per tier — sealed set | ||
| Easy | 78.4 | 77.5 |
| Standard | 68.7 | 75.8 |
| Judge | 66.9 | 73.2 |
| Hard | 58.9 | 60.8 |
Self-hosted cells and hosted-API cells pool S 1,200 + P 300 (1,500 items); raw and unequated. Cells under 15 items are omitted. Sealed counts in the compare view refer to the sealed set S (1,200), which hosted APIs now answer in full.
Languages
Raw chance-corrected competence per item language (0 = chance, 100 = perfect; can be negative), from each system's own measured items: self-hosted systems over S 1,200 + P 300, and every hosted API that answered the full set the same way (1,500 items). Unequated and outside the Composite. Cells under 15 items are left empty. A dagger (†) marks every displayed cell with fewer than 30 answered items. English (1,160 items) is listed first; the other 21 languages and the mixed-language group share 340 items. In 14 further groups no system reaches the 15-item reporting minimum, so they get no column (items in the pool shown): Mixed-language (14), Hindi (13), Chinese (13), Turkish (12), Arabic (12), Ukrainian (12), Korean (11), Dutch (11), Czech (11), Swedish (10), Norwegian (9), Greek (9), Finnish (8), Indonesian (5) — 150 items, scored like every other item. This is the v1.6.1 main-pool breakdown. Per-language coverage grows with the expanded uc1.1 multilingual pool, a candidate for a later release that is not part of v1.6.1.
| System | en1160 | pl31 | es31 | pt†29 | de†27 | da†22 | it†17 | ja†17 | fr†16 |
|---|---|---|---|---|---|---|---|---|---|
| Sage 1.3.0 · API | 68 | 47 | 84 | 77† | 80† | 35† | 61† | 54† | 69† |
| Jev 1.13.0 · API | 67 | 28 | 80 | 76† | 72† | 48† | 48† | 33† | 76† |
| wity-1 · API | 75 | 55 | 78 | 79† | 85† | 44† | 67† | 30† | 45† |
| Fastino GLiNER-2.5-Decide · API | 0 | 0 | 20 | 0† | 7† | 0† | 0† | 0† | 0† |
Intelligence gate and Noul decisiveness
The Composite applies a soft gate below Intelligence 50 (× (I/50)²). Rows under it show their value with a "below gate" label. The decisive rate is the share of valid Noul answers with P(yes) ≥ 0.80 or ≤ 0.20 (for systems that return only a label, every valid label counts as decisive); only decisive answers count as answers. Accuracy among decisive answers is a diagnostic and does not enter any score.
| System | Intelligence | Choice | Noul | Score | Noul decisive rate | Accuracy among decisive |
|---|---|---|---|---|---|---|
| Sage 1.3.0 | 65.6 | native | native | native | 83 % | 95 % |
| Jev 1.13.0 | 63.6 | native | native | native | 79 % | 98 % |
| wity-1 | 70.3 | native | native | native | 91 % | 94 % |
| Fastino GLiNER-2.5-Decide | 7.0below gate | confidence | confidence | confidence | 100 % | 54 % |
All data (23 systems)
Not yet measured on v1.6 · 15 systems with a dated carried score
Every ranked system of the live v1.5.7 board that is not measured on the v1.6.0 pool keeps its last published score, marked with the release that first published that measurement and that release's publication day. Carried rows are listed separately and never ranked together with v1.6-measured rows. v1.5 protocol (1,624 decisions: 904 open + 720 sealed); scores are on the v1.5 scale and are not comparable with v1.6-measured rows.
Dates: v1.5.0 published 2026-09-28 (13) · v1.5.3 published 2026-09-29 (1) · v1.5.6 published 2026-10-03 (1). The date is the publication day of the release that first published the measurement, not a per-model measurement timestamp.
| System | Measured on | Capability (v1.5 scale) | v1.5 Composite | Cost / 1,000 | Median latency |
|---|---|---|---|---|---|
| Qwen3.8 27B (Chutes TEE) | measured on v1.5.0 (2026-09-28) | 96.8 | 0.0 · was #108 on v1.5.7 | $2.1784estimate | 6.49 s |
| GPT-6 Luna (default medium reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.9 | 38.8 · was #38 on v1.5.7 | $0.1138estimate | 1.56 s |
| DeepSeek V4.1 Flash (thinking default) | measured on v1.5.0 (2026-09-28) | 95.3 | 6.6 · was #72 on v1.5.7 | $0.4976estimate | 1.78 s |
| GPT-6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 95.1 | 40.5 · was #37 on v1.5.7 | $0.1075estimate | 1.58 s |
| GPT-5.6 Luna (low reasoning effort) | measured on v1.5.0 (2026-09-28) | 94.5 | 22.4 · was #51 on v1.5.7 | $0.2047tariff | 1.32 s |
| Autoloops – Gemma 4 31B IT | measured on v1.5.0 (2026-09-28) | 81.3 | 40.5 · was #36 on v1.5.7 | $0.1032tariff | 0.61 s |
| SimpleJev Qwen3.8-27B | measured on v1.5.0 (2026-09-28) | 80.0 | 3.4 · was #74 on v1.5.7 | $0.6868estimate | 1.68 s |
| Gemini 3.1 Flash-Lite | measured on v1.5.0 (2026-09-28) | 76.1 | 19.6 · was #54 on v1.5.7 | $0.2194tariff | 0.86 s |
| Instinct (ZooWork, Qwen3.8-27B) | measured on v1.5.0 (2026-09-28) | 73.9 | 18.3 · was #60 on v1.5.7 | $0.2298estimate | 0.52 s |
| Vansa-3.4 (Vansa, hosted System One API)Hosted API retained at its v1.5.6 score under the three-refresh exposure cadence. | measured on v1.5.6 (2026-10-03) | 72.8 | 71.6 · was #5 on v1.5.7 | $0.0193estimate | 0.14 s |
| openjev-sglang (Qwen3.6-35B-A3B on SGLang) | measured on v1.5.0 (2026-09-28) | 70.8 | 29.0 · was #45 on v1.5.7 | $0.1398estimate | 1.20 s |
| Instinct Dual 4B | measured on v1.5.3 (2026-09-29) | 65.7 | 47.0 · was #30 on v1.5.7 | $0.0212estimate | 0.51 s |
| system-one-open (Gemma 4 E2B LoRA on an L4) | measured on v1.5.0 (2026-09-28) | 56.6 | 42.4 · was #35 on v1.5.7 | $0.0114estimate | 1.15 s |
| decision-machine-1 (milliseconds.ai) | measured on v1.5.0 (2026-09-28) | 48.0 | 3.2 · was #75 on v1.5.7 | $0.0286estimate | 0.18 s |
| JevAct (einptein, jev1-2b-v2) | measured on v1.5.0 (2026-09-28) | 36.9 | 1.5 · was #81 on v1.5.7 | $0.0112estimate | 0.80 s |
Method · v1.6.1
Open weights and API offerings (board v1.7.0)
- The main board at /jev-models ranks open-weights systems: published weights that we ran ourselves, on GPU or CPU machines we rent and operate. A system we measured through an endpoint we do not run (vendor API, author-hosted or third-party-hosted endpoint, for example Qwen3.8 27B via Chutes) is an API offering, even when its base weights are open; API offerings are ranked on this board.
- Why separate boards: open weights can be compared on equal hosting terms, while an API price is a vendor decision that can be subsidised or raised later and is not reproducible by readers. Every score and measurement is the same on both boards; only the set of ranked rows differs, and ranks are the published order filtered to that set.
- Jev 1.13.0 is a hosted API. It stays on the open-weights board as the reference row (it defines the Jev-class cost and latency caps) and is not ranked there; it is ranked on the API leaderboard.
- Official cost basis is unchanged: the Cost axis keeps each row's documented reference price (see the cost notes below; APIs with a known base model are priced at the developer's own list price). The GPU cost calculator on the main board is a What-If for your own hosting and never changes a score or rank.
Revision history
- v1.7.0 (2026-10-06): Leaderboard split. /jev-models ranks open-weights systems we ran on our own hardware, with Jev 1.13.0 as an unranked reference row and a “Show API offerings” switch; hosted API offerings are ranked on the new /jev-models/api board. No score was recomputed: ranks are the published order filtered to each board. Adds the GPU cost What-If.
- v1.6.1 (2026-10-05): Hosted APIs (fastino-gliner-2-5-decide, jev-1.13.0, sage-1.3.0, wity-1, wity-1-always, wity-1-off) answer the full 1,500-item set instead of the 600-item API subset and are no longer equated; self-hosted rows unchanged; whole v1.6.0 sealed draw retired. Cost per 1,000 decisions is measured on one common item set (all items except the 23 outside the Jev reference input range) for every token-priced row.
- v1.6.0 (2026-10-05): Rotating item draw (1,200 sealed + 300 public); hosted APIs on API subsets equated to the self-hosted scale; /jev-models/v1.6.0 stays available unchanged.
- Earlier releases: v1.5.7, v1.5.6, v1.5.5 and the historical boards below.
System types (colours)
- Jev — reference (TypeSafe, closed) — The closed system JevBench is named after, shown as the reference.
- Closed API (weights not public) — Available through a hosted API; the weights cannot be downloaded.
- Open weights · LLM decoder — An autoregressive language model with public weights, including fine-tunes, merges and Jev rebuilds.
- Open weights · diffusion LM — A language model that generates by iterative denoising instead of token by token.
- Open weights · encoder / classifier — BERT-style encoders, NLI zero-shot classifiers and GLiNER-type models.
- Open weights · reranker — A cross-encoder or LLM reranker that scores options against the input.
- Base model control (no decision fine-tune, raw logits) — An official open checkpoint without any decision fine-tune, read out from raw logits, used as a floor.
- System (router / cascade / ensemble) — Several models combined at inference time.
Colours show architecture only; they never change a score or rank. Each class is assigned from cited evidence (config.json, model card, provider docs, or our own run record for hosted APIs).
Method amendments
- 5 Oct 2026, v1.6.1 (Florian's decision): hosted/API systems now answer the same full item set as self-hosted systems (1,200 sealed + 300 public). Their earlier answers on the API subsets were reused and only the missing sealed items were sent; every item was logged in the exposure ledger before it was sent. API rows are no longer equated. Consequence: every v1.6.0 sealed item has now been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases; v1.6.2 onward draws fresh items from the reserve.
- 5 Oct 2026, v1.6.1 cost rule (Florian's decision): cost per 1,000 decisions is now measured on one common item set for every system: all items except the 23 very long items that lie outside the Jev reference's accepted input range. Before, a system that processed those items paid for their tokens while a system that refused them did not. Refusals still count against Intelligence exactly as before. Hosted APIs with a per-token tariff are priced from their measured tokens on this common set; carried per-decision prices of self-hosted systems are unchanged. Effect: Sage 1.3.0 USD 0.0247 per 1,000 decisions on the common set (0.0650 on all 1,500 items); Jev 1.13.0 0.0323 (unchanged).
Rotating item sets
Each release draws fresh sealed decisions from a larger reserve. Self-hosted open-weights models (run offline on our own GPU pods or Sandy) answer S and P; externally hosted models answer the same full set (S and P) as the self-hosted systems since v1.6.1, so hosted APIs are no longer equated.
| Set | Items | Choice | Noul | Score | Answered by |
|---|---|---|---|---|---|
| S · Sealed release draw | 1,200 | 600 | 300 | 300 | self-hosted systems (run offline on our own GPU pods or Sandy) and, from v1.6.1, externally hosted APIs |
| A · API subset (part of S; used by the v1.6.0 hosted-API measurements) | 300 | 150 | 75 | 75 | every system (inside S since v1.6.1); retired for future draws |
| P · Public set | 300 | 150 | 75 | 75 | every system |
- Self-hosted systems: 1,500 items (S 1,200 + P 300). Hosted APIs: 1,500 items (S 1,200 + P 300), the same as self-hosted systems; A (300) is the part of S that earlier API measurements used.
- A sealed item is scored in at most three releases, then retired. An item used in one release is not drawn again in the next. If coverage minimums cannot be met, the release waits for newly reviewed items.
- The selection seed is committed (SHA-256) before any inference, and the draw is a deterministic function of policy, seed and item id.
API-exposure rule
- An external exposure is any sealed item sent to an endpoint we do not control: closed APIs, and open-weights models reached through third-party hosts, routers or a submitter's endpoint. Timeouts count as exposure.
- From v1.6.1 hosted models receive the full item set (S plus P); items they had already answered were reused and only the missing sealed items were sent. Hosted models are re-measured at most once every three refresh releases unless a verified new model version ships. Between measurements they keep their last score with its measurement date.
- Every externally sent sealed item is logged before dispatch. An item exposed to a provider is never scored again for that provider; once two different providers have received it, it retires for everyone. In v1.6.0 Jev 1.13.0 and Fastino GLiNER-2.5-Decide both received the same A, so A retires globally. With the v1.6.1 amendment every v1.6.0 sealed item has been seen by at least two providers, so the whole v1.6.0 sealed draw is retired for future releases.
Public-versus-sealed gap penalty
Intelligence is half public, half sealed. A system whose public score exceeds its sealed score by more than the field-median gap (G_med = 2.6 points) plus 8 points loses one Intelligence point per excess point. Hosted APIs are compared on P versus A against the same self-hosted systems' P-versus-A gap (-1.2 points).
Hosted APIs (not equated since v1.6.1)
Hosted APIs that answered the full set are scored exactly like self-hosted systems; no equating offset is applied to them. Category and language values stay raw.
Headline and Composite
Capability = mean(Intelligence, Calibration) for systems within twice the Jev 1.13.0 cost and median latency (the official caps; the sliders change only your view). The Composite (option A) is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost with the v1.5 low-axis gates; it remains secondary. Request types Choice, Noul and Score weigh equally; tiers weigh easy 0.1, standard 0.2, hard 0.4, judge 0.3.
Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s.
Costs of v1.6-measured systems carry each system's published v1.5.4 cost per 1,000 decisions (pricing rules unchanged; v1.6 item lengths differ) unless the row says otherwise; Fastino's is an estimate from its published tariff and measured tokens.
Noul decisiveness and Score baseline (addendum B)
Scored with method option B (scorer setting O1S), selected on 3 Oct 2026 after the v1.6 scores were known and disclosed as a post-results change. Each split × type competence is clipped at 0 before the type weighting, so a type answered no better than chance counts as chance instead of negative. The Score chance baseline is the error of always predicting the mid-scale level, so a flat know-nothing distribution earns about 0. A Score cell whose golds all sit at mid-scale keeps the v1.5 random-level baseline; this only occurs in small breakdown and bootstrap cells. Calibration is unchanged. This run: G_med = 2.57 points.
Failed requests and very long items
23 items of this draw are very long (about 77,000 to 81,000 input tokens). Systems with a shorter context window refuse them; Jev-Omni runs out of GPU memory on them on its listed RTX 6000 (48 GB) recipe; decider-4b v2's server truncates them to about 32,800 tokens and answers; Plumb-4B reads them in full. Each outcome is the system's own recipe and is scored as such. A failed, refused or unparseable answer counts as wrong for Intelligence and stays in the denominator; it does not enter Calibration (the v1.5 rule, applied to every system). Failed answers per system: Sage 1.3.0 0/1500 · Jev 1.13.0 23/1500 · wity-1 0/1500 · Fastino GLiNER-2.5-Decide 31/1500 · wity-1 0/1500 · wity-1 0/1500.
Calibration basis
Systems that return a full probability distribution are calibrated on all components (top-label error, plus distribution distance for Choice and ranked-probability error for Score). Fastino GLiNER-2.5-Decide returns a single confidence value, so its Calibration is the top-label error only and is not like-for-like with full-distribution systems.
Overnight full re-measure (4–5 Oct 2026)
Every system with a reproducible recipe was re-run on the v1.6.0 pool overnight with the same pinned inputs and scorer (method option B / O1S). This page uses scoring round score-v161-3 (v1.6.1) (2026-10-06 00:39:36 UTC). Only complete runs (1,500 items self-hosted, the full API input for hosted APIs) are ranked; partial runs are never ranked, and systems not yet re-measured keep their dated v1.5.x score in the separate table.
- Scores are the official v1.6.0 scorer (score_v16.py, method option B / O1S, bootstrap B = 1,000) over complete outputs only; each scoring round is kept separately.
- Hosted APIs are scored on the same full item set as self-hosted systems (S 1,200 + P 300, v1.6.1) and are not equated.
- GPU-class deviations: reproducible recipes used the hardware listed per row, including H100 for large fast-lane decoders and RTX 6000 for the baseline; hardware differences remain in measured latency. The standard x2 + 0.15 s self-hosted adjustment is an assumption, not a hardware normalization.
- Jev-class caps use a fixed reference: Jev 1.13.0 as measured in v1.5 (p50 0.62 s, USD 0.0323 per 1k answers); caps = 2x (1.23 s, USD 0.0646). Jev's own v1.6 p50 is 0.24 s. Wity auto remains outside the latency cap and stays ranked in Composite A. OFF and ALWAYS are unranked variants of the AUTO main row.
- 6 further measured candidates await a separate publication decision.
- Sage 1.3.0 (Levanto Labs) was measured on 5 Oct from Sandy (Helsinki), text only. Cost uses the Levanto list tariff (USD 0.05/M input, USD 10/M output; levanto.ai/pricing, read 5 Oct) over all 600 answered rows: USD 0.0766/1,000 decisions. It exceeds the frozen Jev-class cost cap and is excluded from the Capability headline.
- A fresh Monday Jev 1.13.0 re-check on A3 measured p50 0.239 s and equated Capability 77.5 versus the official 76.5, within the confidence interval; the official v1.6.0 Jev row is retained.
Supplementary API draws A2 and A3
History: hosted APIs were first measured on API subsets (A, then A2 and A3) and equated to the self-hosted scale in v1.6.0. From v1.6.1 they answer the same full set as self-hosted systems, so no equating is applied; the earlier subsets A, A2 and A3 are retired.
Per-model exposure counts (hosted and author-hosted endpoints)
| System (provider) | v1.6 sealed items sent | Scored sealed set | Status |
|---|---|---|---|
| wity-1-auto (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-off (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| wity-1-always (wity via railway) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| jev-1.13.0 (typesafe) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| fastino-gliner-2-5-decide (fastino) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
| gpt-6-luna (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gpt-6-luna-low (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gpt-5.6-luna (openai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| gemini-3.1-flash-lite (google) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| deepseek-flash (deepseek) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| qwen3.8-27b (chutes) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| decision-machine-1 (milliseconds.ai) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| classifier-dev-fast (classifier.dev) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| vansa-3.4 (vansa) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| kushal-gemma4-31b-it-autoloops (autoloops) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| system-one-open (modal via modal) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| openjev-sglang (modal via modal) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| instinct (zoowork) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| instinct-dual-4b (zoowork) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| simplejev-qwen3.8-27b (featherless) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| simplejev-qwen3.6-35b-a3b (featherless) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| jevact (jevact) | 0 | v1.5 sealed-720 (retired for scoring) | carried: last measured on v1.5 (API-exposure rule; not due / not in tonight's scope) |
| Sage 1.3.0 (Levanto Labs) | 1200 | S (1,200 sealed) + P (300 public) | full set since v1.6.1 (5 Oct 2026); earlier API-subset answers reused, only missing sealed items sent; whole v1.6.0 sealed draw retired |
Provenance
Measured on the v1.6 pool (run completion day, UTC): 2026-10-05: Sage 1.3.0, Jev 1.13.0, Fastino GLiNER-2.5-Decide · 2026-10-06: wity-1, wity-1, wity-1.
Aggregate files: results sha256 5d4567d5e5acd945d17dd082adbbd6188d38174ca523b0fa6b3b02c6cc2dc5b1 · categories sha256 db333c130a7b16beae0aacc153d3dd69cf8e54f277e790dad41769208570d6db · dated carry sha256 334e851944e92383ab3071e52b3012040c7023d0fe236395953169ebee8f0459. Scoring source sha256 e15347783b8bedeeddd9017a831608a957f5bf0386f05902c50161ce0468f911. The method, release data and carry artifact are independently hashable.
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
Open this section to load the earlier public-only board and diagnostics.