{"schema_version":1,"publisher":"Benchmark Heaven","definition":"JevBench by Benchmark Heaven is a benchmark for AI decision models, including Jev-compatible systems. It measures typed decision accuracy, probability calibration, latency and cost. Maintained by Benchmark Heaven, independently of TypeSafe AI.","canonical":"https://benchmarkheaven.com/jev-models","revision":"v1.6.1","published_at":"2026-10-06T00:39:36.000Z","version_page":"/jev-models/v1.6.1","artifact_sha256":"5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1","source_sha256":"5cd8c1332226ea9c179883d2285c0d8b8cc593ae38f2645197364a7ddd9c83c1","option":"A","amendments":[{"label":"a6","date":"2026-10-07","scorer_source_sha256":"9d512dc13d7d876b233560e8710ed5f64d83f555be90141fa62644f22d5c9421"},{"label":"a7","date":"2026-10-07","scorer_source_sha256":"9d512dc13d7d876b233560e8710ed5f64d83f555be90141fa62644f22d5c9421"},{"label":"a8","date":"2026-10-07","scorer_source_sha256":"7a5b4d4c7b7ce966c1d7048d5106d50a6a700c8d89372e223ac276f41540be9d"},{"label":"a11","date":"2026-10-08","scorer_source_sha256":"4d2833926c2daf284ce98e1fea6d9223dfc08341905778eb77391734638c27d2"},{"label":"a12","date":"2026-10-08","scorer_source_sha256":"4d2833926c2daf284ce98e1fea6d9223dfc08341905778eb77391734638c27d2"}],"amended_through":"2026-10-08","max_row_last_measured_on":"2026-10-07","scores":{"capability":{"definition":"Arithmetic mean of Intelligence and Calibration; ranked models within both official caps. Headline on the open-weights board.","reference_key":"jev-1.13.0@v1.5.7","cost_cap_usd_per_1000":0.06459465517241379,"median_latency_cap_s":1.2329566404223442,"factor":2},"composite":{"definition":"Weighted harmonic mean of Intelligence, Calibration, Speed and Cost, multiplied by squared penalties when Intelligence is below its configured floor, or Speed or Cost is below 50. Equal 25% weights do not mean an arithmetic average.","weights":{"calibration":25,"cost":25,"intelligence":25,"speed":25},"intelligence_floor":50,"speed_floor":50,"cost_floor":50,"zero_axis_score":0,"missing_axis_score":null}},"counts":{"A":300,"P":300,"S":1200,"api_input":1500,"api_smoke_public":5,"core":900,"selfhosted_input":1500},"current_board_note":"The live board may also include later dated API reruns and presentation revisions.","reproducibility":{"scorer":"https://github.com/fstandhartinger/jevbench","presentation":"https://github.com/fstandhartinger/model-market-comparison","scorer_sources_sha256":{"noul_method_v16.py":"119e59486ab3300b972fbcbadb0fe1dceceb124287607375aa61c66063c607d0","score_v15.py":"0931e05c6410b1b971201834397cb609cdb25f942723c50375632952f6dcc447","score_v15_headlineA.py":"826f2cbcb57007ec8a161e6eb0111182b8a656c0638027c6c78022a858fd8e2c","score_v16.py":"02132271117653e6302c509232a581d0094f0e78eac8b1bfc67596fd5d5097e9"},"results":"/api/jevbench/v1.6.1"},"image":{"revision":"v0.3.0","built_at":"2026-10-07T15:19:10.573549+00:00","artifact_sha256":"1f5014f77a9945c2098f50460872a3d5c193b3751a3fb44543bcaf0931365ffc","counts":{"reserve":2441,"real_hf":713,"synthetic":1728,"public":300,"sealed":1200,"api_subset":300,"equating_note":"Frozen v0.2 scoring math is applied to the new draw. Core is the headline; computer-use is separate. Pending field-median gap/API equating holds remain visible per row. Scores from v0.1.5 are not directly ranked on the v0.3 scale. GPT-6 Luna via OpenAI Decisions was measured on 6 Oct 2026 through the native Decisions choice endpoint (same option text, native option probabilities, images re-encoded losslessly as PNG, refusals and provider errors counted wrong, no retries); the GPT-6 Luna chat-completions row was re-measured separately on 7 Oct 2026 (see below). Wity-1 (AlphaNimble) was measured on 7 Oct 2026 on the same A300+P300 input through its native SystemOne choice endpoint (reasoning=auto, native option probabilities, no retries; failed requests and timeouts counted wrong). Images that would exceed Wity's 12 MB request limit were downscaled to 1,500 px on the long side, as Wity's docs advise (2 images). Latency is measured end to end; on this run large images took much longer on the Wity side than the model time Wity reports. A second Wity-1 row uses reasoning off (the server default); the first row uses reasoning auto. Dated-carry re-measurement (7 Oct 2026, requested by Florian): GPT-6 Luna (chat), GPT-5.6 Luna, Gemini 3.1 Flash Lite, Gemini 3.8 Flash and Autoloops Gemma 4 31B IT were re-measured on the same A300+P300 input as an owner exception to the N=3 API refresh cadence (no provider version change is claimed; exposure logged before dispatch). OpenRouter and Autoloops requests used the unchanged v0.2 prompts and parsers; images were sent with their real MIME type, 7 BMP images were re-encoded losslessly as PNG and 5 images over 6 MiB were downscaled to 2,048 px on the long side. Provider failures count wrong; no retries. Gemini 3.8 Flash reasons by default and ran out of the 384-token answer budget on 329 of 600 items, which count wrong under the unchanged method. NeoHorse Jev 4B was re-measured on its self-hosted runtime at Hugging Face revision 434cb21d (identical weights and code; the reviewed revision 56c36ae6 was removed by the author). The two JevAny rows remain carries because their source is unavailable.","passes":"Pass 1 is the official prediction pass; pass 2 is diagnostic and is never included in these rankings. Reproducibility: a full second pass of the 29 common self-hosted systems on a fresh pod returned the same answer for 1,500 of 1,500 items in every system (100 % agreement)."},"capability":{"anchor":"jev-1.13.0","anchor_source":"JevBench v1.5.4","factor":2,"cost_usd_per_1000_cap":0.06459465517241379,"latency_p50_s_cap":1.2329566404223442,"anchor_cost_usd_per_1000":0.032297327586206896,"anchor_latency_p50_s":0.6164783202111721,"rationale":"ImageJevBench has no Jev reference row because Jev does not accept images. Capability eligibility therefore uses the same absolute envelope as JevBench: twice Jev 1.13.0’s v1.5.4 cost and adjusted median latency. These anchor values are frozen here so later JevBench releases do not move the ImageJevBench cap."},"composite":{"definition":"Weighted harmonic mean of Intelligence, Calibration, Speed and Cost, multiplied by squared penalties when Intelligence is below its configured floor, or Speed or Cost is below 50. Equal 25% weights do not mean an arithmetic average.","weights":{"intelligence":25,"calibration":25,"speed":25,"cost":25},"intelligence_floor":50,"speed_floor":50,"cost_floor":50}},"historical_note":"Historical releases retain their original scoring definitions and measurement sets. Do not compare scores across method versions as if they were the same experiment."}