Image benchmark · v0.1

Image JevBench v0.1

A held-out comparison of systems that make decisions from images, from interface targets to everyday scenes.

The frozen benchmark has 684 scored items: 228 public and 456 sealed; 93 further items are retired and not scored. This page shows aggregate sealed results only. It contains no sealed task, image, answer key, or per-item prediction.

Caveat: The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems.

Composite score

Whole benchmark, 684 decisions. Bars include all 12 systems. Pink bars are hosted APIs; Gemma used Autoloops and no provider no-retention claim is made. A system whose Cost or Calibration axis falls under the gate scores 0; the row says which gate it is.

  1. 1. Jev-Omni73.10
  2. 2. Mapika decider-2b-vision BF1666.77
  3. 3. Reflex 4Breleased stable configuration65.38
  4. 4. djev-spark NVFP449.38
  5. 5. Autoloops – Gemma 4 31B ITAPI47.14
  6. 6. djev-dev BF1646.64
  7. 7. Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.26
  8. 8. GPT-6 Lunalow reasoning effortAPI16.06
  9. 9. GPT-5.6 LunaAPI11.05
  10. 10. Gemini 3.1 Flash LiteAPI9.58
  11. 11. Gemini 3.8 Flash0.0 · gated by Cost (USD 2.06 per 1,000 decisions)API0.00
  12. 12. OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.00

Full ranking

Ranked by the composite score: equal-weight Intelligence, Calibration, Speed and Cost axes, then the unchanged Jev-class gates. Matched gap is signed public-minus-sealed accuracy within the matched families. No system currently exceeds the 15 pp allowance. Hosted systems are marked API because their providers received sealed images and questions; the Gemma 4 endpoint is identified separately in the exposure note.

#SystemCompositeIntelligenceCalibrationSpeedCostGap (matched)Earlier splitPublic accuracySealed accuracyUSD / 1,000p50 / p95
1Jev-Omni73.1063.9389.9089.6859.52-1.0 pp#3 · 55.63153/228 · 67.1%368/456 · 80.7%USD 0.02240.076 s / 0.103 s
2Mapika decider-2b-vision BF1666.7752.8481.4688.1757.59-4.4 pp#1 · 70.02151/228 · 66.2%319/456 · 70.0%USD 0.02590.091 s / 0.154 s
3Reflex 4Breleased stable configuration65.3863.3483.6886.1449.35-0.6 pp#2 · 68.10184/228 · 80.7%333/456 · 73.0%USD 0.04880.154 s / 0.191 s
4djev-spark NVFP449.3849.1084.3083.5145.95+3.4 pp#5 · 41.56146/228 · 64.0%307/456 · 67.3%USD 0.06330.173 s / 0.374 s
5Autoloops – Gemma 4 31B ITAPI47.1479.2782.3873.8942.65+0.7 pp#4 · 45.91193/228 · 84.6%397/456 · 87.1%USD 0.08161.425 s / 2.868 s
6djev-dev BF1646.6449.6078.0283.0044.69+6.6 pp#6 · 38.76153/228 · 67.1%302/456 · 66.2%USD 0.06980.193 s / 0.392 s
7Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.2674.9190.7173.6229.43+1.9 pp#7 · 18.95191/228 · 83.8%379/456 · 83.1%USD 0.22520.799 s / 1.167 s
8GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 91.8% of calls

API
16.0681.1088.3165.7027.48+12.0 pp#8 · 16.51203/228 · 89.0%395/456 · 86.6%USD 0.21662.905 s / 9.256 s
9GPT-5.6 Luna

Cost receipts cover 99.6% of calls

API
11.0591.8195.7862.5823.49-0.8 pp#9 · 12.57211/228 · 92.5%436/456 · 95.6%USD 0.35242.878 s / 19.176 s
10Gemini 3.1 Flash LiteAPI9.5884.2891.4464.4922.31-0.2 pp#10 · 9.24194/228 · 85.1%419/456 · 91.9%USD 0.38881.799 s / 19.775 s
11Gemini 3.8 Flash0.0 · gated by Cost (USD 2.06 per 1,000 decisions)API0.0088.4772.9858.780.58+0.8 pp#11 · 0.00202/228 · 88.6%430/456 · 94.3%USD 2.05994.832 s / 27.423 s
12OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0063.580.0079.5440.89+3.9 pp#12 · 0.00187/228 · 82.0%331/456 · 72.6%USD 0.09340.317 s / 0.634 s

Compare two systems

Search any two measured systems. The radar uses the same four 0–100 score axes as the ranking; further out is better.

  • A: Jev-Omni · composite 73.10
  • B: Mapika decider-2b-vision BF16 · composite 66.77
Image JevBench four-axis comparisonJev-Omni versus Mapika decider-2b-vision BF16. Intelligence: 63.9 versus 52.8; Calibration: 89.9 versus 81.5; Speed: 89.7 versus 88.2; Cost: 59.5 versus 57.6.50100Intelligence63.9 · 52.8Calibration89.9 · 81.5Speed89.7 · 88.2Cost59.5 · 57.6
Scores are from the frozen Image JevBench v0.1 aggregate; no item-level results are shown.
Values as a table
AxisJev-OmniMapika decider-2b-vision BF16
Intelligence63.952.8
Calibration89.981.5
Speed89.788.2
Cost59.557.6

Examples from public items

These eight examples are from the public split. Each card shows the image, question, options and correct answer; licensed sources are credited below their image.

A cardboard parcel with a visibly torn corner on a conveyor.

Everyday photo

Parcel condition

Question: Is the parcel visibly damaged?

  • A. yesCorrect
  • B. no

Correct answer: A. yes

Source: Synthetic image created for ImageJevBench · public item promo-parcel-01

A printed cafe receipt with item names, amounts, tax and total visible.

Everyday photo

Receipt legibility

Question: Is the receipt readable?

  • A. yesCorrect
  • B. no

Correct answer: A. yes

Source: Synthetic image created for ImageJevBench · public item promo-receipt-00

A spreadsheet with a number-format dialog and five labelled target markers.

Computer use · spreadsheet

Change a cell format

Question: Goal: change selected cells to type “Text”. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker BCorrect
  • C. Click marker C
  • D. Click marker D
  • E. Click marker E

Correct answer: B. Click marker B

Source: ScreenSpot-Pro · excel_macos_0 · MIT · 210e78d38442 · public item mm-195 · Licence text: MITAdapted: five labelled click markers were drawn on the source screenshot.

A desktop browser showing a sports page, browser tabs and five labelled markers.

Browser

Open a new tab

Question: Goal: open a new tab. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker B
  • C. Click marker C
  • D. Click marker D
  • E. Click marker ECorrect

Correct answer: E. Click marker E

Source: ScreenSpot · item 15 · Apache-2.0 · 0be08781e2e1 · public item mm-135 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A mobile translation interface with “Hello, world” and five labelled target markers.

Mobile app · translation

Use the translate control

Question: Goal: translate. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker B
  • C. Click marker C
  • D. Click marker DCorrect
  • E. Click marker E

Correct answer: D. Click marker D

Source: ScreenSpot · item 341 · Apache-2.0 · 0be08781e2e1 · public item mm-118 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A Windows file-manager context menu with five labelled target markers.

Computer use · file manager

Copy a file

Question: Goal: copy the file. Which labelled marker should be clicked?

  • A. Click marker ACorrect
  • B. Click marker B
  • C. Click marker C
  • D. Click marker D
  • E. Click marker E

Correct answer: A. Click marker A

Source: ScreenSpot · item 19 · Apache-2.0 · 0be08781e2e1 · public item mm-139 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A compact financial table with 2014 and 2015 net revenue values.

FinQA · financial table

Calculate a revenue change

Question: What is the net change in net revenue during 2015 for Entergy Corporation?

  • A. 103.4
  • B. 84.6
  • C. 94Correct
  • D. 112.8

Correct answer: C. 94

Source: FinQA · table item · MIT annotations; CDLA-Permissive-1.0 table data · 3d6a736bc67e · public item mm-061 · Licence text: MIT, CDLA-Permissive-1.0Table data: IBM FinTabNet (CDLA-Permissive-1.0). The table was re-rendered as an image by ImageJevBench; no original filing page is shown.

Screenshot images are adapted from the credited datasets (labelled markers added); the FinQA table is re-rendered by ImageJevBench. Licences: Apache-2.0 (ScreenSpot), MIT (ScreenSpot-Pro, Geometry3K, FinQA annotations), CDLA-Permissive-1.0 (FinTabNet table data). App and website content shown in screenshots belongs to its respective owners. No Mind2Web or Android-in-the-Wild image is shown.

Results by track

Each track is ranked on its own public and sealed items. The licensed core and synthetic everyday-photo results remain separately visible.

Core · 500 items

139 public · 361 sealed: 62 real-source items and 299 fresh synthetic pool items (documents, charts, inventory, safety).

#SystemCompositeIntelligenceCalibrationSpeedCostPublic accuracySealed accuracyUSD / 1,000
1Jev-Omni69.7155.9687.3289.6759.1771/139 · 51.1%276/361 · 76.5%USD 0.0230
2Reflex 4Breleased stable configuration65.7559.2282.5286.5049.90103/139 · 74.1%249/361 · 69.0%USD 0.0468
3Mapika decider-2b-vision BF1654.8346.2978.5988.5759.0975/139 · 54.0%234/361 · 64.8%USD 0.0231
4Autoloops – Gemma 4 31B ITAPI45.9975.1076.4975.0442.62107/139 · 77.0%305/361 · 84.5%USD 0.0818
5djev-spark NVFP432.1641.1180.2683.5545.8167/139 · 48.2%224/361 · 62.0%USD 0.0640
6djev-dev BF1629.9841.3672.4183.0444.5571/139 · 51.1%220/361 · 60.9%USD 0.0705
7Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj18.9568.8687.6774.1629.48103/139 · 74.1%286/361 · 79.2%USD 0.2243
8GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 90.6% of calls

API
17.1076.6284.8266.9228.33114/139 · 82.0%310/361 · 85.9%USD 0.1955
9GPT-5.6 Luna

Cost receipts cover 99.6% of calls

API
10.9089.5894.0063.0823.40122/139 · 87.8%342/361 · 94.7%USD 0.3550
10Gemini 3.1 Flash LiteAPI9.0679.4185.5665.1721.96105/139 · 75.5%324/361 · 89.8%USD 0.3994
11Gemini 3.8 Flash0.0 · gated by Cost (USD 2.24 per 1,000 decisions)API0.0085.1173.6759.410.00113/139 · 81.3%336/361 · 93.1%USD 2.2351
12OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0060.060.0080.2041.25104/139 · 74.8%251/361 · 69.5%USD 0.0909

Everyday photo decisions · 184 synthetic items

89 public images from the promo set; 95 sealed: 61 variants from the reviewed v0.1 candidate pool across 17 matched situations and 34 fresh pool photos. Ambiguous labels were dropped after visual, two-model and gold-blind human checks. No brands and no focused faces.

#SystemCompositeIntelligenceCalibrationSpeedCostPublic accuracySealed accuracyUSD / 1,000
1Jev-Omni79.8690.7787.7489.7460.5082/89 · 92.1%92/95 · 96.8%USD 0.0207
2Mapika decider-2b-vision BF1673.2377.1185.1587.2554.2076/89 · 85.4%85/95 · 89.5%USD 0.0336
3Reflex 4Breleased stable configuration64.1079.6180.7386.0647.9681/89 · 91.0%84/95 · 88.4%USD 0.0543
4djev-spark NVFP459.7276.7991.3883.4746.3479/89 · 88.8%83/95 · 87.4%USD 0.0615
5djev-dev BF1655.5277.7787.8682.9745.0582/89 · 92.1%82/95 · 86.3%USD 0.0679
6Autoloops – Gemma 4 31B ITAPI49.4493.8287.9173.2742.7386/89 · 96.6%92/95 · 96.8%USD 0.0811
7Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.8996.6493.5272.3829.2988/89 · 98.9%93/95 · 97.9%USD 0.2275
8GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 95.1% of calls

API
13.6487.0092.7961.9525.6889/89 · 100.0%85/95 · 89.5%USD 0.2712
9GPT-5.6 Luna

Cost receipts cover 99.5% of calls

API
11.4398.7096.5461.7623.7389/89 · 100.0%94/95 · 98.9%USD 0.3451
10Gemini 3.1 Flash LiteAPI10.89100.0092.9361.7923.3189/89 · 100.0%95/95 · 100.0%USD 0.3601
11Gemini 3.8 Flash0.0 · gated by Cost (USD 1.58 per 1,000 decisions)API0.0998.7070.4157.604.0189/89 · 100.0%94/95 · 98.9%USD 1.5839
12OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0075.930.0079.5339.9883/89 · 93.3%80/95 · 84.2%USD 0.1002

djev-spark sealed photo result: 83/95 sealed decisions · 87.4%. It saw the public promo images in an earlier inference-only video run, with no training; this sealed score is the independent measurement for it.

Split

The split is 228 public / 456 sealed (33.3% / 66.7%). The public part is unchanged. The sealed part is 123 never-exposed v0.1 items plus 333 fresh items from our private synthetic rotation pool, and all 12 systems were re-run on the fresh items with their original settings. 93 legacy items are retired and not scored. Split counts and hashes were frozen and posted before any system saw a fresh item.

Public / sealed split

228 / 456

33.3% public · 66.7% sealed · was 228 / 216

Licensed real-source items

201 items

139 public · 62 sealed · 93 retired

Fresh sealed items

333 items

Our own synthetic renders and photos · 299 core · 34 everyday photos

Synthetic share

70.6%

483/684 scored items are our own synthetic content

Public means an item was already exposed anywhere. That includes all 79 Mind2Web-derived items because Kev's training data overlaps Mind2Web; the 89 promo photos shown in videos; the 8 example cards on this page and in the status video; and 61 source rows in the public site repository since the 21 Sep preview. The split also counts 1 image asset already present in that repository. Every scored item that has never been exposed is sealed.

Sealed is 123 never-exposed v0.1 items plus 333 fresh items drawn from our private synthetic rotation pool (documents, charts, inventory and safety scenes, and everyday photos). The draw was stratified by family and difficulty with a seeded draw, and the split counts and hashes were frozen before any system saw a fresh item. Computer Use and Browser Use pool items were excluded because those tracks stay separate. Existing item-level outputs are reused for the public and kept sealed items; every system was run on the fresh items with the same code, pinned revisions, prompts and settings as its original run.

Retired. 93 items (ScreenSpot 45, ScreenSpot-Pro 31, Android-in-the-Wild 17) were moved from public to sealed on 24 Sep while their inputs and every system's predictions sat outside the sealed store, so they cannot count as an unseen holdout. They are not scored and not relabelled.

FamilyPublicSealedRetired
Android-in-the-Wild (AITW_Single mirror)01117
Everyday photo89950
FinQA2000
Geometry3K1900
Multimodal-Mind2Web7900
Pool: charts0790
Pool: documents01020
Pool: inventory0610
Pool: safety inspection0570
ScreenSpot202945
ScreenSpot-Pro12231
Total22845693

Computer Use and Browser Use tracks (preview)

Two planned tracks extend the image benchmark to computer and browser interactions.

Computer Use

400 items

147 public · 253 sealed

Public decision types

  • target location
  • element type
  • next action class

Sealed decision types

  • task completion
  • action success
  • dialog safety
  • blocker status

Sealed origin: original synthetic local fixtures

Browser Use

400 items

240 public · 160 sealed

Public decision types

  • target location
  • element type
  • next action class

Sealed decision types

  • task completion
  • action success
  • dialog safety
  • blocker status

Sealed origin: original synthetic local fixtures

Cross-track sealing rule. A public-origin Computer Use or Browser Use row that shares a source row or screenshot with a sealed core item is sealed too.

Kev / Mind2Web flag. 133 Mind2Web-derived Browser Use rows are public-only and reported separately for Kev, whose training data overlaps Mind2Web.

Not measured yet — no scores. Neither track has a reviewed system or reusable output. Computer Use has no reviewed candidate in the current queue. Browser Use has one candidate requested, not yet evaluated: kev-0.6b-browser-use. Its public, ungated HF adapter is pinned at 08414b0001f6eba32bc69372abfeac5b74718ac0 (model card license: Apache-2.0), with base jaredpalmer/kev-0.6b at dece6dba8d43f0f7ded45e9f5b9df12474d90843. The model takes textual DOM state and uses a pointer head; its serving path and safe head loader are not independently reviewed. Its Mind2Web-derived examples remain public-only. Both tracks still lack a ready matched-family split and measurements, so they remain separate, requested, and unscored outside the core ranking.

Method and limitations

Intelligence. Accuracy counts missing, invalid and unparseable answers as wrong. Each part is chance-corrected against its own average chance rate, then combined as 35% public and 65% sealed. Calibration uses the same weights.

Matched-family overfit penalty. The gap is public accuracy minus sealed accuracy within families that have at least 10 items on both sides: ScreenSpot and Everyday photo. If that matched gap is above 15 percentage points, Intelligence is multiplied by max(0, 1 − (gap − 15)/100). The same rule applies to every system. The raw overall gap is shown in the data but does not affect the score.

Calibration. Ten-bin top-label ECE is scaled by valid probability coverage. Label-only output receives zero calibration. OpenJev's NLI entailment values select an answer but are not treated as categorical probabilities.

Speed and cost. These rules are unchanged. Speed uses whole-call p50 and p95 latency; local latency uses the v1.4 2× plus 0.15-second adjustment. Hosted unit cost uses returned per-call usage receipts; missing receipts are not zero-filled, and the Cost axis is scaled by receipt coverage. Retry costs are tracked separately. Local cost uses measured GPU seconds at the recorded per-system GPU-hour rate and excludes loading, downloads, build, and idle time. For self-hosted systems, the original run's timings are combined with the fresh-item run on the same GPU type.

Composite and gates. These rules are unchanged. The four axes use an equal-weight harmonic mean, followed by the Jev-class Intelligence, Speed, and Cost gates below 50. Gemini 3.8 Flash's high raw accuracy but near-zero composite reflects its measured cost and the Cost gate; label-only systems have zero Calibration under the inherited convention.

Difficulty balance. The split follows exposure, not a stratified draw, so the parts differ in family mix: browser actions (Mind2Web), chart questions (FinQA) and geometry are public-only, while ScreenSpot-Pro, Android-in-the-Wild and the fresh pool families are sealed-only. This is why the overfit penalty compares only matched families (ScreenSpot and Everyday photo).

Fresh-item difficulty. The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems. Scores on this split are therefore not comparable with the earlier 228/216 preview.

Exposure. GPT-6 Luna and Gemini 3.8 Flash previously saw public promo-photo candidates and sealed-photo candidates in stateless label-check calls, including candidates later dropped. The checks showed no gold; human gold-blind adjudication decided inclusion. GPT-6 Luna is the saved low-reasoning-effort setting. Four OpenRouter systems used the requested no-retention route with fallback disabled; Gemma used the Autoloops endpoint and is API-flagged, with no no-retention claim made here. Local systems ran without network, credentials, or gold maps. The fresh pool items were authored with Claude and reviewed by OpenAI Codex as a blind critic, and the pool photos were generated with an OpenAI image model. This is disclosed for the two OpenAI rows. Self-hosted systems ran the fresh items on one rented GPU pod in offline containers with read-only inputs.

Requested and excluded candidates (37)

The ranking covers 12 measured configurations. Gemma 4 31B reuses completed saved results on the original items and was run on the fresh items; GPT-6 Luna is the saved low-reasoning-effort setting. Candidates below have no score unless listed in the ranking. Requested rows remain visible with the exact access or review blocker; exclusions describe the reviewed interface, license or duplicate status.

CandidateSource revisionAccessStatusReason
CUA-S1-4B-0.2 multimodal adapterHF 16818868bPublic, ungated; pinned revision resolvesrequested, not yet evaluatedPinned source and multimodal image path passed source review; the integrated runner’s exact-hash independent review is pending. No weights have been staged and no ImageJevBench inference has run.
Visual Jev 4B Answer-SFTHF 7a3f1bb0dPublic, ungated; pinned revision resolvesrequested, not yet evaluatedPinned source and image path passed source review; exact-hash runner review is pending. Training used two and four options while the benchmark includes up to five. No weights have been staged and no inference has run.
NeoHorse-Jev-4BHF 56c36ae62Public, ungated; pinned revision resolvesrequested, not yet evaluatedPinned self-contained multimodal source passed source review; the integrated runner’s exact-hash independent review is pending. No weights have been staged and no inference has run.
Jevify Qwen3-VL-2B Tier 2HF 46e8e72c2Public, ungated; pinned revision resolvesrequested, not yet evaluatedPinned Tier 2 vision adapter and image path passed source review; exact-hash runner review is pending. The loader rejects an optional pickle head. No weights have been staged and no inference has run.
Standard One 3BHF c0d23877ePublic, ungated; pinned revision resolvesrequested, not yet evaluatedPinned image-capable checkpoint and server path passed source review; exact-hash runner review is pending. No weights have been staged and no inference has run.
Standard One 8BHF 5d1285dd7Public, ungated; pinned revision resolvesrequested, not yet evaluatedPinned image-capable checkpoint and server path passed source review; exact-hash runner review and a fresh pod/disk resource gate are pending. The checkpoint is about 17.9 GB and will be staged only on an owned remote pod. No weights have been staged and no inference has run.
Sage1 / Levanto SageHosted imageAPI requires SAGE_API_KEY; no configured key or free grantrequested, not yet evaluatedNo-cost access is unavailable. No paid quota was purchased.
djev-distill-v426B-class chNo checkpoint stagedrequested, not yet evaluatedCustom vLLM backend and loader review are incomplete; no model download was attempted.
yah01/vjev-visionHF 2fa8b58e4Public, ungatedrequested, not yet evaluatedReviewed loader uses unrestricted torch.load(head.pt); safe loader review is incomplete.
WIlfLin/JEV-Qwen3.8-Flash-Next-Linear-RuntimeHF 1847787ffPublic model metadata; upstream weights about 184 GBrequested, not yet evaluatedLicense permission and a viable reviewed multi-GPU runtime remain unresolved; no weights were downloaded.
Visual-Jev generic Qwen3.5-4B scorerGitHub 2d68cPublic repositoryrequested, not yet evaluatedThis is distinct from Answer-SFT, but its exact base, runtime and optional image path lack pin-specific review.
divyanshx11/JEVisionHF 324698caaPublic, ungatedrequested, not yet evaluatedThe pointer-head .pt artifact and incomplete runtime need safe-loader and interface review.
SeanLiu/Jev-Vision 8BExact infereCheckpoint page is publicrequested, not yet evaluatedThe head is a .pt artifact and its exact inference loader has not passed review.
Fr0zencr4nE/jev-spatialHF 5727eade6Metadata public; previous anonymous weight-file inventory returned HTTP 401requested, not yet evaluatedNo keyless weight staging path is confirmed, and the spatial specialist's general typed-choice interface is unverified.
Sarashina 2.2 vision JEV mmproj 3BHF 09ce275e3Public, ungatedrequested, not yet evaluatedThe vision-language projection has no reviewed general typed-choice adapter/runtime.
OhtaMan Gemma 4 E2B IT choice-64HF 3d403b990Public, ungatedrequested, not yet evaluatedImage-text metadata is present, but its actual choice interface and inference code are unreviewed.
MetaSK-Jev 4B policy mixHF ea20fe85bPublic, ungatedrequested, not yet evaluatedImage-text metadata is present, but its actual image path and typed-choice runtime are unreviewed.
Jevify Gemma 4 E4BHF a6b5a716fPublic, ungatedrequested, not yet evaluatedDistinct from measured hosted Gemma 4 31B; local runtime, typed interface and license review are incomplete.
Jevify Gemma 4 26B-A4BHF d4c0d1d45Public, ungatedrequested, not yet evaluatedDistinct from measured hosted Gemma 4 31B; local runtime, typed interface and license review are incomplete.
JevAny-27B-SFTHF ad7b48b70Public, ungatedrequested, not yet evaluatedImage/video capability is reported, but the pinned base/runtime and training-data overlap review are incomplete.
JevAny-27B-RLCRHF 078883b2ePublic, ungatedrequested, not yet evaluatedDistinct adapter; its pinned base/runtime and training-data overlap review are incomplete.
LFM2.5-VL-3B-Decision-NVFP4HF 1ff8bf556Public, ungatedrequested, not yet evaluatedImage/video input is reported, but license terms, decision-head behavior and local runtime are unresolved.
yeyan00/Jev-DecisionGitHub eac9dPublic repository; exact checkpoint access unknownrequested, not yet evaluatedThe checkpoint, image interface and inference runtime are not pinned or reviewed.
OmniJev 4B / 2B / 0.8BDistinct pubPublic model pagesrequested, not yet evaluatedPublished loaders use unrestricted pickle .pt files; an approved safe loader is unavailable.
Bonsai-Llama-Jev submission (measured as Bonsai-2-27B v2)kyr0/Bonsai-Public pinned submission and model artifacts; completed item-level outputs reused from the Bonsai measurement jobincluded in ranking (rank 7/12)This submission maps to the existing Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj row (score 18.9529). The reviewed System One path accepts inline image data; offline_eval.image_body reads each local image and sends it as a data:image data URI in the request state. The final clean run returned 444/444 image predictions. The frozen 228-public / 216-sealed roster was rescored from these saved item outputs; no rerun was done. See receipts/BONSAI-ALIAS-AND-IMAGE-PATH.md.
AutoJev-27BGitHub denisPublic, ungated; Apache-2.0 weights / MIT coderequested, not yet evaluatedPinned upstream inference source supports images: DecisionInput.images is typed; model.open_image loads image objects, data URIs or paths; DecisionModel.prepare passes images to the processor; server EvaluationRequest.images flows into model.prepare. However, this continuation has no ImageJevBench offline, local-only, safetensors adapter/controller for AutoJev. The exact-hash-reviewed runner is frozen to six other candidates pending independent PASS. A separate adapter and exact-hash review are required. The public, ungated Apache-2.0 checkpoint is about 52.2 GB and may only be staged on a remote owned pod after fresh disk, watchdog, GPU and spend checks; the current USD 6 approval covers the six-candidate proposal and has not been extended to AutoJev. No weights or task inputs were staged and no predictions were made.
CLM-8BContrastive-Public checkpoint; text encoderexcluded from image rankingThe reviewed Engine.answer(state, {decision: question}) adapter and frozen Qwen3-8B encoder have no image argument or image processor.
Eikos 4B / 27BReviewed LetPublic checkpoint; source reviewedexcluded from image rankingLetterAdapter.dist(state_text, question, options) receives text only; conditional-generation architecture selection is not image input.
Main Jev / TypeSafeOfficial docOfficial API/interface; no native image fieldexcluded from image rankingThe documented state/decision interface accepts text, objects and arrays, with no supported image input. A wrapper field is insufficient.
OpenJev-VisionGitHub b83bePublic repositoryexcluded from core rankingThe published fixed classifier/head does not expose the general typed dynamic-option interface used by the benchmark.
hr98w/jev-visualGitHub 4382bPublic repositoryexcluded as a duplicate wrapperA request/runtime wrapper, not a distinct checkpoint or decision head.
Qevi-2B and Laya Vision 201M/256MCandidate swPublic model pages; non-commercial termsexcluded pending license clearanceExisting review reports non-commercial or non-commercial/share-alike terms; public-release permission is not established.
Jev-Omni reuploads and quantizationsCandidate swPublic reuploadsexcluded as duplicatesNo distinct model family/runtime is established beyond the already measured Jev-Omni row.
joyfox/Qwen3.5-0.8B-JEVCandidate swPublic checkpointexcluded from image rankingThe decision artifact removes the vision tower; text-only.
Hanno-Labs/bosun-v3.1 0.6B / 1.7BCandidate swPublic checkpointexcluded from image rankingQwen3 text decision models; no image path documented in the reviewed release.
autotrust/JEVCandidate swPublic sourceexcluded from image rankingTyped text/JSON decisions only; no image input documented in the reviewed interface.
Programalyst/realtime-vision-decision-agentCandidate swPublic repositoryexcluded as an application wrapperCombines vision detection with Jev but is not a distinct model checkpoint.

Image JevBench is a separate benchmark from the text-only JevBench Score. Sealed item-level content remains private. Public/sealed item counts, accuracy, track and score breakdowns are aggregates. Results describe these exact tested configurations and do not establish absence from model training data.