Archived release: ImageJevBench v0.1.4. See the current v0.2 results.

Image benchmark · v0.1.4

Image JevBench v0.1.4

A held-out comparison of systems that make decisions from images, from interface targets to everyday scenes.

The frozen benchmark has 684 scored items: 228 public and 456 sealed; 93 further items are retired and not scored. This page shows aggregate sealed results only. It contains no sealed task, image, answer key, or per-item prediction.

Wity-1 under author review. The run used the listed production endpoint, but its response did not identify the deployed build. The author is checking the build; this score may change after a full rerun.

Caveat: The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems.

Composite score

Whole benchmark, 684 decisions. Bars include all 50 systems. Pink bars are hosted APIs; Wity-1 and Gemma are hosted endpoints without a no-retention claim. A system whose Cost or Calibration axis falls under the gate scores 0; the row says which gate it is.

  1. 1. Imajev-4B76.39
  2. 2. Wity-1API74.36
  3. 3. Jev-Omni73.10
  4. 4. NeoHorse Jev 4B71.94
  5. 5. Visual-Jev 4B Answer-SFT69.81
  6. 6. JPT-4Bkirp / llm2jev69.55
  7. 7. imajev 2B68.72
  8. 8. Glancefrozen Qwen3-VL-4B66.86
  9. 9. AutoJev-27B66.85
  10. 10. Mapika decider-2b-vision BF1666.77
  11. 11. jev-spatialFr0zencr4nE65.87
  12. 12. JPT-9Bkirp / llm2jev65.82
  13. 13. OmniJev 4Btinnel12365.65
  14. 14. Reflex 4Breleased stable configuration65.38
  15. 15. OmniJev-Qwen3.5-9B-v4tzcfly65.14
  16. 16. Jevify Gemma 4 26B-A4B65.00
  17. 17. Surogate Rune 26B-A4B v363.82
  18. 18. JevAny-27B SFT63.79
  19. 19. JevAny-27B RLCR63.46
  20. 20. Jev-Vision 8BSeanLiu63.39
  21. 21. shisa-de-163.20
  22. 22. OmniJev 2Btinnel12362.84
  23. 23. CUA-S1 4B62.63
  24. 24. Standard One 8B59.80
  25. 25. Jevify Qwen3-VL-2B T259.02
  26. 26. imajev 9B58.33
  27. 27. Visual-Jev genericzero-shot Qwen3.5-4B baseline57.73
  28. 28. djev-distill-v4tarsur38550.04
  29. 29. djev-spark NVFP449.38
  30. 30. diffusiongemma-26b djev v10 step160snowicarus48.76
  31. 31. vjev-visionyah0147.67
  32. 32. Autoloops – Gemma 4 31B ITAPI47.14
  33. 33. djev-dev BF1646.64
  34. 34. Winnow-12BQ8_0 + F16 vision projector43.56
  35. 35. JPT-0.8Bkirp / llm2jev30.51
  36. 36. Jevify Gemma 4 E4B29.76
  37. 37. Standard One 3B29.01
  38. 38. Gevva E4Btext checkpoint, image input26.83
  39. 39. OmniJev 0.8Btinnel12326.81
  40. 40. Gevva E2B multimodal21.85
  41. 41. JEVisiondivyanshx11, visual route21.50
  42. 42. Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.26
  43. 43. Qevi-2BMeerDevelopment17.44
  44. 44. GPT-6 Lunalow reasoning effortAPI16.06
  45. 45. GPT-5.6 LunaAPI11.05
  46. 46. Gemini 3.1 Flash LiteAPI9.58
  47. 47. PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor)0.11
  48. 48. Gemini 3.8 Flash0.0 · gated by Cost (USD 2.06 per 1,000 decisions)API0.00
  49. 49. OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.00
  50. 50. OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported)0.00

Full ranking

Ranked by the composite score: equal-weight Intelligence, Calibration, Speed and Cost axes, then the unchanged Jev-class gates. Matched gap is signed public-minus-sealed accuracy within the matched families. No system currently exceeds the 15 pp allowance. Hosted systems are marked API because their providers received sealed images and questions; the Gemma 4 endpoint is identified separately in the exposure note.

#SystemCompositeIntelligenceCalibrationSpeedCostGap (matched)Earlier splitPublic accuracySealed accuracyUSD / 1,000p50 / p95
1Imajev-4B

Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132.

76.3973.7790.5287.5961.18-4.2 pp#1 · 76.39185/228 · 81.1%380/456 · 83.3%USD 0.01970.099 s / 0.174 s
2Wity-1

Under author review: server build ID was not recorded; this score may change after verification.

Setting: Wity SystemOne, reasoning=auto

API
74.3683.6887.7176.9557.33+2.1 pp—200/228 · 87.7%410/456 · 89.9%USD 0.02640.710 s / 2.844 s
3Jev-Omni73.1063.9389.9089.6859.52-1.0 pp#2 · 73.10153/228 · 67.1%368/456 · 80.7%USD 0.02240.076 s / 0.103 s
4NeoHorse Jev 4B71.9472.8991.2086.7651.56-0.8 pp#3 · 71.94184/228 · 80.7%377/456 · 82.7%USD 0.04120.140 s / 0.170 s
5Visual-Jev 4B Answer-SFT69.8165.4384.1387.5253.48-6.2 pp#4 · 69.81175/228 · 76.8%352/456 · 77.2%USD 0.03550.101 s / 0.176 s
6JPT-4Bkirp / llm2jev69.5566.9287.3186.5251.13+3.3 pp#5 · 69.55172/228 · 75.4%362/456 · 79.4%USD 0.04260.123 s / 0.206 s
7imajev 2B68.7257.7390.4687.8554.19-1.1 pp#6 · 68.72164/228 · 71.9%328/456 · 71.9%USD 0.03360.096 s / 0.164 s
8Glancefrozen Qwen3-VL-4B66.8659.3772.9788.7355.54-2.8 pp#7 · 66.86177/228 · 77.6%322/456 · 70.6%USD 0.03030.089 s / 0.129 s
9AutoJev-27B66.8569.3382.4386.7949.37+0.1 pp#8 · 66.85174/228 · 76.3%371/456 · 81.4%USD 0.04870.141 s / 0.167 s
10Mapika decider-2b-vision BF1666.7752.8481.4688.1757.59-4.4 pp#9 · 66.77151/228 · 66.2%319/456 · 70.0%USD 0.02590.091 s / 0.154 s
11jev-spatialFr0zencr4nE65.8752.0473.0490.5259.63-0.7 pp#10 · 65.87159/228 · 69.7%307/456 · 67.3%USD 0.02220.058 s / 0.092 s
12JPT-9Bkirp / llm2jev65.8275.7491.5485.1848.25+1.4 pp#11 · 65.82187/228 · 82.0%387/456 · 84.9%USD 0.05310.155 s / 0.255 s
13OmniJev 4Btinnel12365.6556.8582.6187.0650.65-3.0 pp#12 · 65.65163/228 · 71.5%325/456 · 71.3%USD 0.04420.131 s / 0.164 s
14Reflex 4Breleased stable configuration65.3863.3483.6886.1449.35-0.6 pp#13 · 65.38184/228 · 80.7%333/456 · 73.0%USD 0.04880.154 s / 0.191 s
15OmniJev-Qwen3.5-9B-v4tzcfly65.1452.1769.5990.3459.54+0.1 pp#14 · 65.14149/228 · 65.4%318/456 · 69.7%USD 0.02230.063 s / 0.093 s
16Jevify Gemma 4 26B-A4B65.0071.6487.9786.2448.37-1.2 pp#15 · 65.00189/228 · 82.9%366/456 · 80.3%USD 0.05260.156 s / 0.182 s
17Surogate Rune 26B-A4B v363.8274.9187.7785.9347.81+0.9 pp#16 · 63.82191/228 · 83.8%379/456 · 83.1%USD 0.05490.163 s / 0.193 s
18JevAny-27B SFT63.7969.5790.0286.1148.05+4.7 pp#17 · 63.79177/228 · 77.6%369/456 · 80.9%USD 0.05390.161 s / 0.185 s
19JevAny-27B RLCR63.4670.2389.9986.0647.90+4.7 pp#18 · 63.46178/228 · 78.1%371/456 · 81.4%USD 0.05450.162 s / 0.186 s
20Jev-Vision 8BSeanLiu63.3964.1863.1485.8649.99-1.0 pp#19 · 63.39181/228 · 79.4%340/456 · 74.6%USD 0.04650.149 s / 0.215 s
21shisa-de-163.2072.4491.8285.8547.60+1.7 pp#20 · 63.20182/228 · 79.8%377/456 · 82.7%USD 0.05580.165 s / 0.195 s
22OmniJev 2Btinnel12362.8449.4780.5388.5354.40-3.9 pp#21 · 62.84163/228 · 71.5%291/456 · 63.8%USD 0.03310.097 s / 0.130 s
23CUA-S1 4B62.6360.1788.5785.3948.55+0.5 pp#22 · 62.63170/228 · 74.6%333/456 · 73.0%USD 0.05190.150 s / 0.246 s
24Standard One 8B59.8052.7792.0684.9148.26-1.6 pp#23 · 59.80168/228 · 73.7%301/456 · 66.0%USD 0.05300.142 s / 0.298 s
25Jevify Qwen3-VL-2B T259.0247.3186.8389.5259.38+1.6 pp#24 · 59.02164/228 · 71.9%280/456 · 61.4%USD 0.02260.062 s / 0.129 s
26imajev 9B58.3374.7589.2784.1346.06+1.3 pp#25 · 58.33197/228 · 86.4%372/456 · 81.6%USD 0.06280.188 s / 0.292 s
27Visual-Jev genericzero-shot Qwen3.5-4B baseline57.7360.2488.0983.9846.97+0.3 pp#26 · 57.73178/228 · 78.1%325/456 · 71.3%USD 0.05860.179 s / 0.319 s
28djev-distill-v4tarsur38550.0452.5484.8182.9845.11+7.9 pp#27 · 50.04142/228 · 62.3%327/456 · 71.7%USD 0.06760.195 s / 0.392 s
29djev-spark NVFP449.3849.1084.3083.5145.95+3.4 pp#28 · 49.38146/228 · 64.0%307/456 · 67.3%USD 0.06330.173 s / 0.374 s
30diffusiongemma-26b djev v10 step160snowicarus48.7657.1974.3082.8044.65+2.6 pp#29 · 48.76152/228 · 66.7%338/456 · 74.1%USD 0.07000.203 s / 0.397 s
31vjev-visionyah0147.6743.1883.8090.6260.78-2.2 pp#30 · 47.67140/228 · 61.4%286/456 · 62.7%USD 0.02030.058 s / 0.088 s
32Autoloops – Gemma 4 31B ITAPI47.1479.2782.3873.8942.65+0.7 pp#31 · 47.14193/228 · 84.6%397/456 · 87.1%USD 0.08161.425 s / 2.868 s
33djev-dev BF1646.6449.6078.0283.0044.69+6.6 pp#32 · 46.64153/228 · 67.1%302/456 · 66.2%USD 0.06980.193 s / 0.392 s
34Winnow-12BQ8_0 + F16 vision projector43.5667.4086.2382.1241.35+2.9 pp#33 · 43.56177/228 · 77.6%359/456 · 78.7%USD 0.09020.268 s / 0.372 s
35JPT-0.8Bkirp / llm2jev30.5136.2778.0689.1457.50-0.3 pp#34 · 30.51120/228 · 52.6%275/456 · 60.3%USD 0.02610.074 s / 0.129 s
36Jevify Gemma 4 E4B29.7635.7076.6890.6760.86-1.7 pp#35 · 29.76129/228 · 56.6%263/456 · 57.7%USD 0.02020.057 s / 0.088 s
37Standard One 3B29.0135.7683.9086.5252.36-1.1 pp#36 · 29.01136/228 · 59.6%256/456 · 56.1%USD 0.03870.100 s / 0.244 s
38Gevva E4Btext checkpoint, image input26.8335.9469.9185.8249.06+1.4 pp#37 · 26.83132/228 · 57.9%261/456 · 57.2%USD 0.04990.158 s / 0.206 s
39OmniJev 0.8Btinnel12326.8134.5776.4788.9555.39+4.4 pp#38 · 26.81148/228 · 64.9%238/456 · 52.2%USD 0.03070.088 s / 0.120 s
40Gevva E2B multimodal21.8532.1075.9786.7751.03-3.5 pp#39 · 21.85115/228 · 50.4%261/456 · 57.2%USD 0.04290.132 s / 0.179 s
41JEVisiondivyanshx11, visual route21.5035.4344.4085.4547.30+2.3 pp#40 · 21.50123/228 · 53.9%268/456 · 58.8%USD 0.05710.170 s / 0.216 s
42Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.2674.9190.7173.6229.43+1.9 pp#41 · 19.26191/228 · 83.8%379/456 · 83.1%USD 0.22520.799 s / 1.167 s
43Qevi-2BMeerDevelopment17.4429.5455.6789.3758.64+0.1 pp#42 · 17.44144/228 · 63.2%219/456 · 48.0%USD 0.02390.068 s / 0.127 s
44GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 91.8% of calls

API
16.0681.1088.3165.7027.48+12.0 pp#43 · 16.06203/228 · 89.0%395/456 · 86.6%USD 0.21662.905 s / 9.256 s
45GPT-5.6 Luna

Cost receipts cover 99.6% of calls

API
11.0591.8195.7862.5823.49-0.8 pp#44 · 11.05211/228 · 92.5%436/456 · 95.6%USD 0.35242.878 s / 19.176 s
46Gemini 3.1 Flash LiteAPI9.5884.2891.4464.4922.31-0.2 pp#45 · 9.58194/228 · 85.1%419/456 · 91.9%USD 0.38881.799 s / 19.775 s
47PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor)0.114.4244.4789.6959.35-1.0 pp#46 · 0.1180/228 · 35.1%170/456 · 37.3%USD 0.02260.067 s / 0.115 s
48Gemini 3.8 Flash0.0 · gated by Cost (USD 2.06 per 1,000 decisions)API0.0088.4772.9858.780.58+0.8 pp#47 · 0.00202/228 · 88.6%430/456 · 94.3%USD 2.05994.832 s / 27.423 s
49OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0063.580.0079.5440.89+3.9 pp#48 · 0.00187/228 · 82.0%331/456 · 72.6%USD 0.09340.317 s / 0.634 s
50OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported)0.0067.210.0079.4941.09+0.4 pp#49 · 0.00180/228 · 78.9%355/456 · 77.9%USD 0.09200.321 s / 0.635 s

Compare two systems

Search any two measured systems. The radar uses the same four 0–100 score axes as the ranking; further out is better.

  • A: Imajev-4B · composite 76.39
  • B: Wity-1 · composite 74.36
Image JevBench four-axis comparisonImajev-4B versus Wity-1. Intelligence: 73.8 versus 83.7; Calibration: 90.5 versus 87.7; Speed: 87.6 versus 76.9; Cost: 61.2 versus 57.3.50100Intelligence73.8 · 83.7Calibration90.5 · 87.7Speed87.6 · 76.9Cost61.2 · 57.3
Scores are from the frozen Image JevBench aggregate; no item-level results are shown.
Values as a table
AxisImajev-4BWity-1
Intelligence73.883.7
Calibration90.587.7
Speed87.676.9
Cost61.257.3

Examples from public items

These eight examples are from the public split. Each card shows the image, question, options and correct answer; licensed sources are credited below their image.

A cardboard parcel with a visibly torn corner on a conveyor.

Everyday photo

Parcel condition

Question: Is the parcel visibly damaged?

  • A. yesCorrect
  • B. no

Correct answer: A. yes

Source: Synthetic image created for ImageJevBench · public item promo-parcel-01

A printed cafe receipt with item names, amounts, tax and total visible.

Everyday photo

Receipt legibility

Question: Is the receipt readable?

  • A. yesCorrect
  • B. no

Correct answer: A. yes

Source: Synthetic image created for ImageJevBench · public item promo-receipt-00

A spreadsheet with a number-format dialog and five labelled target markers.

Computer use · spreadsheet

Change a cell format

Question: Goal: change selected cells to type “Text”. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker BCorrect
  • C. Click marker C
  • D. Click marker D
  • E. Click marker E

Correct answer: B. Click marker B

Source: ScreenSpot-Pro · excel_macos_0 · MIT · 210e78d38442 · public item mm-195 · Licence text: MITAdapted: five labelled click markers were drawn on the source screenshot.

A desktop browser showing a sports page, browser tabs and five labelled markers.

Browser

Open a new tab

Question: Goal: open a new tab. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker B
  • C. Click marker C
  • D. Click marker D
  • E. Click marker ECorrect

Correct answer: E. Click marker E

Source: ScreenSpot · item 15 · Apache-2.0 · 0be08781e2e1 · public item mm-135 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A mobile translation interface with “Hello, world” and five labelled target markers.

Mobile app · translation

Use the translate control

Question: Goal: translate. Which labelled marker should be clicked?

  • A. Click marker A
  • B. Click marker B
  • C. Click marker C
  • D. Click marker DCorrect
  • E. Click marker E

Correct answer: D. Click marker D

Source: ScreenSpot · item 341 · Apache-2.0 · 0be08781e2e1 · public item mm-118 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A Windows file-manager context menu with five labelled target markers.

Computer use · file manager

Copy a file

Question: Goal: copy the file. Which labelled marker should be clicked?

  • A. Click marker ACorrect
  • B. Click marker B
  • C. Click marker C
  • D. Click marker D
  • E. Click marker E

Correct answer: A. Click marker A

Source: ScreenSpot · item 19 · Apache-2.0 · 0be08781e2e1 · public item mm-139 · Licence text: Apache-2.0Adapted: five labelled click markers were drawn on the source screenshot.

A compact financial table with 2014 and 2015 net revenue values.

FinQA · financial table

Calculate a revenue change

Question: What is the net change in net revenue during 2015 for Entergy Corporation?

  • A. 103.4
  • B. 84.6
  • C. 94Correct
  • D. 112.8

Correct answer: C. 94

Source: FinQA · table item · MIT annotations; CDLA-Permissive-1.0 table data · 3d6a736bc67e · public item mm-061 · Licence text: MIT, CDLA-Permissive-1.0Table data: IBM FinTabNet (CDLA-Permissive-1.0). The table was re-rendered as an image by ImageJevBench; no original filing page is shown.

Screenshot images are adapted from the credited datasets (labelled markers added); the FinQA table is re-rendered by ImageJevBench. Licences: Apache-2.0 (ScreenSpot), MIT (ScreenSpot-Pro, Geometry3K, FinQA annotations), CDLA-Permissive-1.0 (FinTabNet table data). App and website content shown in screenshots belongs to its respective owners. No Mind2Web or Android-in-the-Wild image is shown.

Results by track

Each track is ranked on its own public and sealed items. The licensed core and synthetic everyday-photo results remain separately visible.

Core · 500 items

139 public · 361 sealed: 62 real-source items and 299 fresh synthetic pool items (documents, charts, inventory, safety).

#SystemCompositeIntelligenceCalibrationSpeedCostPublic accuracySealed accuracyUSD / 1,000
1Imajev-4B

Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132.

75.4970.2386.0388.0263.44104/139 · 74.8%290/361 · 80.3%USD 0.0165
2Wity-1

Under author review: server build ID was not recorded; this score may change after verification.

Setting: Wity SystemOne, reasoning=auto

API
72.8181.5384.9476.4556.11114/139 · 82.0%321/361 · 88.9%USD 0.0290
3NeoHorse Jev 4B70.8569.6486.1587.2452.54103/139 · 74.1%289/361 · 80.1%USD 0.0382
4JPT-9Bkirp / llm2jev70.0569.9288.1185.7450.5499/139 · 71.2%295/361 · 81.7%USD 0.0445
5Visual-Jev 4B Answer-SFT69.7562.1084.0987.9655.6099/139 · 71.2%265/361 · 73.4%USD 0.0302
6Jev-Omni69.7155.9687.3289.6759.1771/139 · 51.1%276/361 · 76.5%USD 0.0230
7JPT-4Bkirp / llm2jev67.9660.3283.0986.9453.3487/139 · 62.6%273/361 · 75.6%USD 0.0359
8imajev 2B66.7051.8686.5888.2456.1785/139 · 61.2%243/361 · 67.3%USD 0.0289
9Reflex 4Breleased stable configuration65.7559.2282.5286.5049.90103/139 · 74.1%249/361 · 69.0%USD 0.0468
10Glancefrozen Qwen3-VL-4B65.0455.3270.4388.6355.7499/139 · 71.2%239/361 · 66.2%USD 0.0299
11CUA-S1 4B64.2752.9183.3285.8950.7991/139 · 65.5%245/361 · 67.9%USD 0.0437
12AutoJev-27B63.9662.2779.8086.6449.1789/139 · 64.0%278/361 · 77.0%USD 0.0495
13imajev 9B63.9669.8785.8984.7948.33111/139 · 79.9%280/361 · 77.6%USD 0.0528
14Jevify Gemma 4 26B-A4B63.2166.9084.2786.2348.32105/139 · 75.5%276/361 · 76.5%USD 0.0528
15OmniJev 4Btinnel12363.0750.4980.6287.0550.6984/139 · 60.4%239/361 · 66.2%USD 0.0440
16Surogate Rune 26B-A4B v362.2970.9383.8185.9047.77107/139 · 77.0%289/361 · 80.1%USD 0.0551
17Standard One 8B61.9049.6083.8485.5250.4595/139 · 68.3%222/361 · 61.5%USD 0.0448
18shisa-de-161.8868.3690.4685.8347.5299/139 · 71.2%289/361 · 80.1%USD 0.0562
19JevAny-27B SFT60.2656.9388.7686.1048.0390/139 · 64.7%276/361 · 76.5%USD 0.0540
20JevAny-27B RLCR59.9757.7188.1286.0547.8991/139 · 65.5%278/361 · 77.0%USD 0.0546
21Jev-Vision 8BSeanLiu59.8657.8153.9386.4751.4997/139 · 69.8%251/361 · 69.5%USD 0.0414
22Visual-Jev genericzero-shot Qwen3.5-4B baseline59.1055.0681.4584.5648.2499/139 · 71.2%238/361 · 65.9%USD 0.0531
23Mapika decider-2b-vision BF1654.8346.2978.5988.5759.0975/139 · 54.0%234/361 · 64.8%USD 0.0231
24OmniJev 2Btinnel12352.8646.0378.4688.5054.4492/139 · 66.2%212/361 · 58.7%USD 0.0330
25diffusiongemma-26b djev v10 step160snowicarus46.1550.2267.6882.8044.6671/139 · 51.1%254/361 · 70.4%USD 0.0699
26Autoloops – Gemma 4 31B ITAPI45.9975.1076.4975.0442.62107/139 · 77.0%305/361 · 84.5%USD 0.0818
27Jevify Qwen3-VL-2B T244.2541.8782.8389.8161.2788/139 · 63.3%201/361 · 55.7%USD 0.0195
28Winnow-12BQ8_0 + F16 vision projector42.8359.4080.3382.4941.8189/139 · 64.0%267/361 · 74.0%USD 0.0870
29jev-spatialFr0zencr4nE41.7441.8066.0090.4959.3274/139 · 53.2%218/361 · 60.4%USD 0.0227
30OmniJev-Qwen3.5-9B-v4tzcfly38.9740.9360.7390.3459.4864/139 · 46.0%227/361 · 62.9%USD 0.0224
31djev-distill-v4tarsur38536.0343.8878.3283.0045.1660/139 · 43.2%245/361 · 67.9%USD 0.0673
32djev-spark NVFP432.1641.1180.2683.5545.8167/139 · 48.2%224/361 · 62.0%USD 0.0640
33vjev-visionyah0130.7936.1079.4390.7060.9766/139 · 47.5%206/361 · 57.1%USD 0.0200
34djev-dev BF1629.9841.3672.4183.0444.5571/139 · 51.1%220/361 · 60.9%USD 0.0705
35Standard One 3B24.6833.3481.2186.9954.8572/139 · 51.8%188/361 · 52.1%USD 0.0320
36Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj18.9568.8687.6774.1629.48103/139 · 74.1%286/361 · 79.2%USD 0.2243
37JPT-0.8Bkirp / llm2jev18.2429.3276.0089.3758.8549/139 · 35.3%201/361 · 55.7%USD 0.0235
38Jevify Gemma 4 E4B17.7428.8975.5290.7161.0259/139 · 42.4%187/361 · 51.8%USD 0.0199
39GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 90.6% of calls

API
17.1076.6284.8266.9228.33114/139 · 82.0%310/361 · 85.9%USD 0.1955
40OmniJev 0.8Btinnel12316.6728.5173.4288.9655.3874/139 · 53.2%167/361 · 46.3%USD 0.0307
41Gevva E2B multimodal15.9328.2676.2686.4749.9553/139 · 38.1%192/361 · 53.2%USD 0.0466
42Gevva E4Btext checkpoint, image input15.6829.2772.2885.5847.9761/139 · 43.9%186/361 · 51.5%USD 0.0542
43Qevi-2BMeerDevelopment15.0727.7554.9589.8060.8783/139 · 59.7%153/361 · 42.4%USD 0.0201
44JEVisiondivyanshx11, visual route11.8927.8237.2585.6047.8550/139 · 36.0%194/361 · 53.7%USD 0.0548
45GPT-5.6 Luna

Cost receipts cover 99.6% of calls

API
10.9089.5894.0063.0823.40122/139 · 87.8%342/361 · 94.7%USD 0.3550
46Gemini 3.1 Flash LiteAPI9.0679.4185.5665.1721.96105/139 · 75.5%324/361 · 89.8%USD 0.3994
47PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor)0.043.0048.3590.1260.9932/139 · 23.0%121/361 · 33.5%USD 0.0200
48Gemini 3.8 Flash0.0 · gated by Cost (USD 2.24 per 1,000 decisions)API0.0085.1173.6759.410.00113/139 · 81.3%336/361 · 93.1%USD 2.2351
49OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0060.060.0080.2041.25104/139 · 74.8%251/361 · 69.5%USD 0.0909
50OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported)0.0061.390.0080.5741.5196/139 · 69.1%266/361 · 73.7%USD 0.0890

Everyday photo decisions · 184 synthetic items

89 public images from the promo set; 95 sealed: 61 variants from the reviewed v0.1 candidate pool across 17 matched situations and 34 fresh pool photos. Ambiguous labels were dropped after visual, two-model and gold-blind human checks. No brands and no focused faces.

#SystemCompositeIntelligenceCalibrationSpeedCostPublic accuracySealed accuracyUSD / 1,000
1OmniJev-Qwen3.5-9B-v4tzcfly80.4891.7690.9690.4059.7185/89 · 95.5%91/95 · 95.8%USD 0.0220
2jev-spatialFr0zencr4nE80.2589.1690.3890.5760.5285/89 · 95.5%89/95 · 93.7%USD 0.0207
3Jev-Omni79.8690.7787.7489.7460.5082/89 · 92.1%92/95 · 96.8%USD 0.0207
4Wity-1

Under author review: server build ID was not recorded; this score may change after verification.

Setting: Wity SystemOne, reasoning=auto

API
78.4489.9290.1380.1061.3886/89 · 96.6%89/95 · 93.7%USD 0.0194
5Imajev-4B

Setting: PyTorch; one rotation; --fast; --merge-lora; calibration.json; server 8501f5c3; adapter c9e5f132.

77.9087.4193.8486.5656.5081/89 · 91.0%90/95 · 94.7%USD 0.0282
6AutoJev-27B74.2494.3686.2186.9049.9385/89 · 95.5%93/95 · 97.9%USD 0.0467
7vjev-visionyah0173.3869.0980.6090.5560.2874/89 · 83.1%80/95 · 84.2%USD 0.0211
8Mapika decider-2b-vision BF1673.2377.1185.1587.2554.2076/89 · 85.4%85/95 · 89.5%USD 0.0336
9Jevify Qwen3-VL-2B T272.8869.3189.9888.8155.3076/89 · 85.4%79/95 · 83.2%USD 0.0309
10imajev 2B72.6779.3991.7986.9149.9879/89 · 88.8%85/95 · 89.5%USD 0.0465
11Glancefrozen Qwen3-VL-4B72.2576.0378.3388.7155.0278/89 · 87.6%83/95 · 87.4%USD 0.0316
12OmniJev 4Btinnel12371.8380.6983.2487.1050.5179/89 · 88.8%86/95 · 90.5%USD 0.0446
13NeoHorse Jev 4B71.1384.8192.3286.6549.2181/89 · 91.0%88/95 · 92.6%USD 0.0493
14Jevify Gemma 4 E4B69.2960.8472.8990.6260.4670/89 · 78.7%76/95 · 80.0%USD 0.0208
15Jevify Gemma 4 26B-A4B69.1389.7090.2086.2848.5084/89 · 94.4%90/95 · 94.7%USD 0.0521
16OmniJev 2Btinnel12368.8565.5076.2788.5854.2771/89 · 79.8%79/95 · 83.2%USD 0.0335
17JevAny-27B SFT68.0595.8886.6086.1548.0987/89 · 97.8%93/95 · 97.9%USD 0.0537
18Visual-Jev 4B Answer-SFT67.6979.7181.4486.4949.0276/89 · 85.4%87/95 · 91.6%USD 0.0501
19JevAny-27B RLCR67.6195.8887.2186.1147.9387/89 · 97.8%93/95 · 97.9%USD 0.0544
20JPT-0.8Bkirp / llm2jev67.4959.0079.0088.5654.4371/89 · 79.8%74/95 · 77.9%USD 0.0330
21Surogate Rune 26B-A4B v366.6989.7087.3185.9947.9284/89 · 94.4%90/95 · 94.7%USD 0.0544
22OmniJev 0.8Btinnel12366.3957.3973.8588.9655.4374/89 · 83.1%71/95 · 74.7%USD 0.0306
23shisa-de-166.1686.3389.7285.9347.8183/89 · 93.3%88/95 · 92.6%USD 0.0549
24Reflex 4Breleased stable configuration64.1079.6180.7386.0647.9681/89 · 91.0%84/95 · 88.4%USD 0.0543
25JPT-4Bkirp / llm2jev62.8489.1693.7285.4146.5285/89 · 95.5%89/95 · 93.7%USD 0.0606
26Gevva E4Btext checkpoint, image input62.6860.3059.6287.4252.5771/89 · 79.8%75/95 · 78.9%USD 0.0381
27Jev-Vision 8BSeanLiu62.1788.4087.6885.3346.6084/89 · 94.4%89/95 · 93.7%USD 0.0602
28djev-spark NVFP459.7276.7991.3883.4746.3479/89 · 88.8%83/95 · 87.4%USD 0.0615
29djev-dev BF1655.5277.7787.8682.9745.0582/89 · 92.1%82/95 · 86.3%USD 0.0679
30djev-distill-v4tarsur38555.0177.7786.1482.9244.9582/89 · 92.1%82/95 · 86.3%USD 0.0684
31diffusiongemma-26b djev v10 step160snowicarus54.3279.6186.4682.7944.6181/89 · 91.0%84/95 · 88.4%USD 0.0702
32JPT-9Bkirp / llm2jev53.7395.3490.6083.9943.5288/89 · 98.9%92/95 · 96.8%USD 0.0763
33Visual-Jev genericzero-shot Qwen3.5-4B baseline53.1181.9986.3683.5744.0479/89 · 88.8%87/95 · 91.6%USD 0.0733
34CUA-S1 4B52.2083.2980.5684.2143.9079/89 · 88.8%88/95 · 92.6%USD 0.0741
35JEVisiondivyanshx11, visual route51.3560.5363.8085.1045.9373/89 · 82.0%74/95 · 77.9%USD 0.0635
36Gevva E2B multimodal50.3945.6668.4988.2154.5162/89 · 69.7%69/95 · 72.6%USD 0.0328
37Autoloops – Gemma 4 31B ITAPI49.4493.8287.9173.2742.7386/89 · 96.6%92/95 · 96.8%USD 0.0811
38Standard One 8B49.1067.0380.4883.6543.6973/89 · 82.0%79/95 · 83.2%USD 0.0754
39imajev 9B47.4893.8293.3482.9241.3586/89 · 96.6%92/95 · 96.8%USD 0.0902
40Standard One 3B44.7945.8879.1385.2747.3064/89 · 71.9%68/95 · 71.6%USD 0.0571
41Winnow-12BQ8_0 + F16 vision projector44.4995.3496.4481.8840.1588/89 · 98.9%92/95 · 96.8%USD 0.0988
42Qevi-2BMeerDevelopment36.7241.0052.5788.5253.9961/89 · 68.5%66/95 · 69.5%USD 0.0342
43Bonsai-2-27B v2 PQ2_0 + Q8_0 MMProj19.8996.6493.5272.3829.2988/89 · 98.9%93/95 · 97.9%USD 0.2275
44GPT-6 Lunalow reasoning effort

Setting: OpenRouter reasoning.effort=low

Cost receipts cover 95.1% of calls

API
13.6487.0092.7961.9525.6889/89 · 100.0%85/95 · 89.5%USD 0.2712
45GPT-5.6 Luna

Cost receipts cover 99.5% of calls

API
11.4398.7096.5461.7623.7389/89 · 100.0%94/95 · 98.9%USD 0.3451
46Gemini 3.1 Flash LiteAPI10.89100.0092.9361.7923.3189/89 · 100.0%95/95 · 100.0%USD 0.3601
47PlayJev 0.8B0.0 · gated by Intelligence (below the Jev-class floor)0.779.0035.2389.0955.7348/89 · 53.9%49/95 · 51.6%USD 0.0299
48Gemini 3.8 Flash0.0 · gated by Cost (USD 1.58 per 1,000 decisions)API0.0998.7070.4157.604.0189/89 · 100.0%94/95 · 98.9%USD 1.5839
49OpenJev 4B NLI v2official image-premise path0.0 · gated by Calibration (no probabilities reported)0.0075.930.0079.5339.9883/89 · 93.3%80/95 · 84.2%USD 0.1002
50OpenJev 4B NLI v50.0 · gated by Calibration (no probabilities reported)0.0088.400.0079.4540.0084/89 · 94.4%89/95 · 93.7%USD 0.1000

djev-spark sealed photo result: 83/95 sealed decisions · 87.4%. It saw the public promo images in an earlier inference-only video run, with no training; this sealed score is the independent measurement for it.

Split

The split is 228 public / 456 sealed (33.3% / 66.7%). The public part is unchanged. The sealed part is 123 never-exposed v0.1 items plus 333 fresh items from our private synthetic rotation pool, and all 50 systems were re-run on the fresh items with their original settings. 93 legacy items are retired and not scored. Split counts and hashes were frozen and posted before any system saw a fresh item.

Public / sealed split

228 / 456

33.3% public · 66.7% sealed · was 228 / 216

Licensed real-source items

201 items

139 public · 62 sealed · 93 retired

Fresh sealed items

333 items

Our own synthetic renders and photos · 299 core · 34 everyday photos

Synthetic share

70.6%

483/684 scored items are our own synthetic content

Public means an item was already exposed anywhere. That includes all 79 Mind2Web-derived items because Kev's training data overlaps Mind2Web; the 89 promo photos shown in videos; the 8 example cards on this page and in the status video; and 61 source rows in the public site repository since the 21 Sep preview. The split also counts 1 image asset already present in that repository. Every scored item that has never been exposed is sealed.

Sealed is 123 never-exposed v0.1 items plus 333 fresh items drawn from our private synthetic rotation pool (documents, charts, inventory and safety scenes, and everyday photos). The draw was stratified by family and difficulty with a seeded draw, and the split counts and hashes were frozen before any system saw a fresh item. Computer Use and Browser Use pool items were excluded because those tracks stay separate. Existing item-level outputs are reused for the public and kept sealed items; every system was run on the fresh items with the same code, pinned revisions, prompts and settings as its original run.

Retired. 93 items (ScreenSpot 45, ScreenSpot-Pro 31, Android-in-the-Wild 17) were moved from public to sealed on 24 Sep while their inputs and every system's predictions sat outside the sealed store, so they cannot count as an unseen holdout. They are not scored and not relabelled.

FamilyPublicSealedRetired
Android-in-the-Wild (AITW_Single mirror)01117
Everyday photo89950
FinQA2000
Geometry3K1900
Multimodal-Mind2Web7900
Pool: charts0790
Pool: documents01020
Pool: inventory0610
Pool: safety inspection0570
ScreenSpot202945
ScreenSpot-Pro12231
Total22845693

Computer Use and Browser Use tracks (preview)

Two planned tracks extend the image benchmark to computer and browser interactions.

Computer Use

400 items

147 public · 253 sealed

Public decision types

  • target location
  • element type
  • next action class

Sealed decision types

  • task completion
  • action success
  • dialog safety
  • blocker status

Sealed origin: original synthetic local fixtures

Browser Use

400 items

240 public · 160 sealed

Public decision types

  • target location
  • element type
  • next action class

Sealed decision types

  • task completion
  • action success
  • dialog safety
  • blocker status

Sealed origin: original synthetic local fixtures

Cross-track sealing rule. A public-origin Computer Use or Browser Use row that shares a source row or screenshot with a sealed core item is sealed too.

Kev / Mind2Web flag. 133 Mind2Web-derived Browser Use rows are public-only and reported separately for Kev, whose training data overlaps Mind2Web.

Not measured yet — no scores. Neither track has a reviewed system or reusable output. Computer Use has no reviewed candidate in the current queue. Browser Use has one candidate requested, not yet evaluated: kev-0.6b-browser-use. Its public, ungated HF adapter is pinned at 08414b0001f6eba32bc69372abfeac5b74718ac0 (model card license: Apache-2.0), with base jaredpalmer/kev-0.6b at dece6dba8d43f0f7ded45e9f5b9df12474d90843. The model takes textual DOM state and uses a pointer head; its serving path and safe head loader are not independently reviewed. Its Mind2Web-derived examples remain public-only. Both tracks still lack a ready matched-family split and measurements, so they remain separate, requested, and unscored outside the core ranking.

Method and limitations

Intelligence. Accuracy counts missing, invalid and unparseable answers as wrong. Each part is chance-corrected against its own average chance rate, then combined as 35% public and 65% sealed. Calibration uses the same weights.

Matched-family overfit penalty. The gap is public accuracy minus sealed accuracy within families that have at least 10 items on both sides: ScreenSpot and Everyday photo. If that matched gap is above 15 percentage points, Intelligence is multiplied by max(0, 1 − (gap − 15)/100). The same rule applies to every system. The raw overall gap is shown in the data but does not affect the score.

Calibration. Ten-bin top-label ECE is scaled by valid probability coverage. Label-only output receives zero calibration. OpenJev's NLI entailment values select an answer but are not treated as categorical probabilities.

Speed and cost. Speed uses whole-call p50 and p95 latency; local latency uses the v1.4 2× plus 0.15-second adjustment. Hosted unit cost uses returned per-call usage receipts; missing receipts are not zero-filled, and the Cost axis is scaled by receipt coverage. Retry costs are tracked separately. Wity-1 uses an estimated base-model price of USD 0.15 per million input tokens and USD 1.00 per million output tokens, applied to its measured 120,542 input and zero output tokens. Florian chose this basis on 29 Sep because Wity's younger public tariff is below the base-model reference. The API does not quantify image or thinking work, so this estimate may understate full image-compute cost. Local cost uses measured GPU seconds at the recorded per-system GPU-hour rate and excludes loading, downloads, build, and idle time. For self-hosted systems, the original run's timings are combined with the fresh-item run on the same GPU type.

Composite and gates. These rules are unchanged. The four axes use an equal-weight harmonic mean, followed by the Jev-class Intelligence, Speed, and Cost gates below 50. Gemini 3.8 Flash's high raw accuracy but near-zero composite reflects its measured cost and the Cost gate; label-only systems have zero Calibration under the inherited convention.

Difficulty balance. The split follows exposure, not a stratified draw, so the parts differ in family mix: browser actions (Mind2Web), chart questions (FinQA) and geometry are public-only, while ScreenSpot-Pro, Android-in-the-Wild and the fresh pool families are sealed-only. This is why the overfit penalty compares only matched families (ScreenSpot and Everyday photo).

Fresh-item difficulty. The 333 fresh sealed items are our own synthetic images and renders. They are easier for frontier API models than the older real-source items, so sealed accuracy is higher than on the earlier split for most hosted systems. Scores on this split are therefore not comparable with the earlier 228/216 preview.

Exposure. GPT-6 Luna and Gemini 3.8 Flash previously saw public promo-photo candidates and sealed-photo candidates in stateless label-check calls, including candidates later dropped. The checks showed no gold; human gold-blind adjudication decided inclusion. GPT-6 Luna is the saved low-reasoning-effort setting. Four OpenRouter systems (GPT-6 Luna low, GPT-5.6 Luna, Gemini 3.1 Flash Lite and Gemini 3.8 Flash) received sealed images and questions through the requested no-retention route with provider fallback disabled. Gemma 4 31B used the Autoloops endpoint. Wity-1 was run on 29 Sep through the hosted Wity SystemOne endpoint with bearer authentication; no no-retention claim is made for Wity or Autoloops. These six hosted rows carry the API flag. Autoloops – Gemma 4 31B IT: results on the original items are reused from the completed Autoloops measurement; the 333 fresh sealed items were run on 25 Sep with the original Autoloops runner. Returned per-request token usage is costed at the published Gemma 4 31B rates. djev-spark saw the 100 public promo photos in the earlier inference-only video run, with no training. Its score on the 95 sealed photo items is the independent measure. Local systems ran without network, credentials, or gold maps. The fresh pool items were authored with Claude and reviewed by OpenAI Codex as a blind critic, and the pool photos were generated with an OpenAI image model. This is disclosed for the two OpenAI rows. Self-hosted systems ran the fresh items on one rented GPU pod in offline containers with read-only inputs.

Requested and excluded candidates (69)

The ranking covers 50 measured configurations. Wity-1 completed a full 684-decision API run. Its Cost axis uses the Qwen3.6-35B-A3B market reference applied to returned usage, by the 29 Sep price decision; the server build remains under author review. Candidates below have no score unless listed in the ranking. Requested rows remain visible with the exact access or review blocker; exclusions describe the reviewed interface, license or duplicate status.

CandidateSource revisionAccessStatusReason
CUA-S1-4B-0.2 multimodal adapterHF 16818868bPublic, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#23 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Visual Jev 4B Answer-SFTHF 7a3f1bb0dPublic, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#5 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
NeoHorse-Jev-4BHF 56c36ae62Public, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#4 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Jevify Qwen3-VL-2B Tier 2HF 46e8e72c2Public, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#25 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Standard One 3BHF c0d23877ePublic, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#37 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Standard One 8BHF 5d1285dd7Public, ungated; pinned revision resolvesincluded in v0.1.4 ranking (#24 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Sage1 / Levanto SageHosted imageAPI requires SAGE_API_KEY; no configured key or free grantrequested, not yet evaluatedNo-cost access is unavailable. No paid quota was purchased.
djev-distill-v426B-class chNo checkpoint stagedincluded in v0.1.4 ranking (#28 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
yah01/vjev-visionHF 2fa8b58e4Public, ungatedincluded in v0.1.4 ranking (#31 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
WIlfLin/JEV-Qwen3.8-Flash-Next-Linear-RuntimeHF 1847787ffPublic model metadata; upstream weights about 184 GBrequested, not yet evaluatedLicense permission and a viable reviewed multi-GPU runtime remain unresolved; no weights were downloaded.
Visual-Jev generic Qwen3.5-4B scorerGitHub 2d68cPublic repositoryincluded in v0.1.4 ranking (#27 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
divyanshx11/JEVisionHF 324698caaPublic, ungatedincluded in v0.1.4 ranking (#41 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
SeanLiu/Jev-Vision 8BExact infereCheckpoint page is publicincluded in v0.1.4 ranking (#20 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Fr0zencr4nE/jev-spatialHF 5727eade6Metadata public; previous anonymous weight-file inventory returned HTTP 401included in v0.1.4 ranking (#11 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Sarashina 2.2 vision JEV mmproj 3BHF 09ce275e3Public, ungatedrequested, not yet evaluatedThe vision-language projection has no reviewed general typed-choice adapter/runtime.
OhtaMan Gemma 4 E2B IT choice-64HF 3d403b990Public, ungatedrequested, not yet evaluatedImage-text metadata is present, but its actual choice interface and inference code are unreviewed.
MetaSK-Jev 4B policy mixHF ea20fe85bPublic, ungatedrequested, not yet evaluatedImage-text metadata is present, but its actual image path and typed-choice runtime are unreviewed.
Jevify Gemma 4 E4BHF a6b5a716fPublic, ungatedincluded in v0.1.4 ranking (#36 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Jevify Gemma 4 26B-A4BHF d4c0d1d45Public, ungatedincluded in v0.1.4 ranking (#16 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
JevAny-27B-SFTHF ad7b48b70Public, ungatedincluded in v0.1.4 ranking (#18 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
JevAny-27B-RLCRHF 078883b2ePublic, ungatedincluded in v0.1.4 ranking (#19 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
LFM2.5-VL-3B-Decision-NVFP4HF 1ff8bf556Public, ungatedrequested, not yet evaluatedImage/video input is reported, but license terms, decision-head behavior and local runtime are unresolved.
yeyan00/Jev-DecisionGitHub eac9dPublic repository; exact checkpoint access unknownrequested, not yet evaluatedThe checkpoint, image interface and inference runtime are not pinned or reviewed.
Bonsai-Llama-Jev submission (measured as Bonsai-2-27B v2)kyr0/Bonsai-Public pinned submission and model artifacts; completed item-level outputs reused from the Bonsai measurement jobincluded in v0.1.4 ranking (#42 of 50)Bonsai-Llama-Jev is retained from the live ImageJevBench roster; see its ranking row above.
AutoJev-27BGitHub denisPublic, ungated; Apache-2.0 weights / MIT codeincluded in v0.1.4 ranking (#9 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
CLM-8BContrastive-Public checkpoint; text encoderexcluded from image rankingThe reviewed Engine.answer(state, {decision: question}) adapter and frozen Qwen3-8B encoder have no image argument or image processor.
Eikos 4B / 27BReviewed LetPublic checkpoint; source reviewedexcluded from image rankingLetterAdapter.dist(state_text, question, options) receives text only; conditional-generation architecture selection is not image input.
Main Jev / TypeSafeOfficial docOfficial API/interface; no native image fieldexcluded from image rankingThe documented state/decision interface accepts text, objects and arrays, with no supported image input. A wrapper field is insufficient.
OpenJev-VisionGitHub b83bePublic repositoryexcluded from core rankingThe published fixed classifier/head does not expose the general typed dynamic-option interface used by the benchmark.
hr98w/jev-visualGitHub 4382bPublic repositoryexcluded as a duplicate wrapperA request/runtime wrapper, not a distinct checkpoint or decision head.
Jev-Omni reuploads and quantizationsCandidate swPublic reuploadsexcluded as duplicatesNo distinct model family/runtime is established beyond the already measured Jev-Omni row.
joyfox/Qwen3.5-0.8B-JEVCandidate swPublic checkpointexcluded from image rankingThe decision artifact removes the vision tower; text-only.
Hanno-Labs/bosun-v3.1 0.6B / 1.7BCandidate swPublic checkpointexcluded from image rankingQwen3 text decision models; no image path documented in the reviewed release.
autotrust/JEVCandidate swPublic sourceexcluded from image rankingTyped text/JSON decisions only; no image input documented in the reviewed interface.
Programalyst/realtime-vision-decision-agentCandidate swPublic repositoryexcluded as an application wrapperCombines vision detection with Jev but is not a distinct model checkpoint.
JPT-0.8BPinned measuPublic checkpoint; non-commercial termsincluded in v0.1.4 ranking (#35 of 50)Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above.
JPT-4BPinned measuPublic checkpoint; non-commercial termsincluded in v0.1.4 ranking (#6 of 50)Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above.
JPT-9BPinned measuPublic checkpoint; non-commercial termsincluded in v0.1.4 ranking (#12 of 50)Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above.
Qevi-2BPinned measuPublic checkpoint; non-commercial termsincluded in v0.1.4 ranking (#43 of 50)Included under Florian's 27 Sep decision despite the stated non-commercial license; see the measured row above.
OmniJev 0.8BPinned measuPublic aggregate rowincluded in v0.1.4 ranking (#39 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
OmniJev 2BPinned measuPublic aggregate rowincluded in v0.1.4 ranking (#22 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
OmniJev 4BPinned measuPublic aggregate rowincluded in v0.1.4 ranking (#13 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Imajev-4BImajev-4B faPublic request modelincluded in v0.1.4 ranking (#1 of 50)Remeasured on the unchanged v0.1 method and split with the pinned server and adapter, one rotation, --fast, and --merge-lora. Only aggregate results are published.
Glance frozen Qwen3-VL-4ByoheinakajimFrozen checkpoint; requested by Yoheiincluded in v0.1.4 ranking (#8 of 50)Retained from the reviewed ImageJev aggregate. The hold applies only to the separate Qwen3-VL-2B CUDA speedlab path.
Laya Visionthaitea/layaCustomer-requested evaluationheld at author's requestNot ranked or published while the author explores other options.
Glance speedlab Qwen3-VL-2Bglance.yoheiApple-only MLX path; measured CUDA port differsheld pending author discussionNot ranked or published until Yohei is asked about the separate CUDA port.
Winnow-12BEldanRing/wiReviewed source pin; independently scored ImageJev handoffincluded in v0.1.4 ranking (#34 of 50)Independent recomputation matched 21/21 metrics with zero delta. The pinned source and author build.py/runtime lock were used, but this run used an isolated server image rather than the separate multi-model load-test Docker context. Medium dependency reproducibility: apt packages resolved at build time and pip wheels lack hashes. All per-file hashes and image membership passed; a separate image-tree digest was not compared with the aggregate receipt (low assurance limitation).
Jev-OmniFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#3 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
imajev 2BFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#7 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Mapika decider-2b-vision BF16Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#10 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Reflex 4B (released stable configuration)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#14 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
OmniJev-Qwen3.5-9B-v4 (tzcfly)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#15 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Surogate Rune 26B-A4B v3Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#17 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
shisa-de-1Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#21 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
imajev 9BFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#26 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
djev-spark NVFP4Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#29 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
diffusiongemma-26b djev v10 step160 (snowicarus)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#30 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Autoloops – Gemma 4 31B ITFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#32 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
djev-dev BF16Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#33 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Gevva E4B (text checkpoint, image input)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#38 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Gevva E2B multimodalFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#40 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
GPT-6 Luna (low reasoning effort)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#44 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
GPT-5.6 LunaFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#45 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Gemini 3.1 Flash LiteFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#46 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
PlayJev 0.8BFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#47 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Gemini 3.8 FlashFrozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#48 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
OpenJev 4B NLI v2 (official image-premise path)Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#49 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
OpenJev 4B NLI v5Frozen v0.1.Public aggregate rowincluded in v0.1.4 ranking (#50 of 50)Measured aggregate included in the ranking under the frozen ImageJevBench v0.1 method; see its row above.
Wity-1Wity productHosted Wity SystemOne API; Cost estimated from Qwen3.6-35B-A3B market referenceincluded in v0.1.4 ranking (#2 of 50)Full 684-decision run with valid probabilities and complete input/output usage; base-model reference applied to measured tokens. Server build under author review.

Image JevBench is a separate benchmark from the text-only JevBench Score. Sealed item-level content remains private. Public/sealed item counts, accuracy, track and score breakdowns are aggregates. Results describe these exact tested configurations and do not establish absence from model training data.