Hosting = where inference runs; company = where the provider or lab is registered.
EU hosting means the route's inference runs inside the EU: an EU region, AWS Bedrock's EU cross-region (geo) profiles, Azure's Europe Data Zone, or a provider whose entire public fleet is documented as EU-hosted, each checked per model against the provider's documentation. Global deployments do not count, and neither does an EU billing region, an EU company or an EU control plane on its own. One disclosed company-policy exception stays in, marked “EU equivalent”.
Data confidentialitywhat the provider may do with your prompts
More settings
Evidence requirements (benchmark evidence, priced provider, measured task tokens) sit above the table they apply to.
Applies to price views & model offers; benchmark evidence stays unfiltered.
JevBench by Benchmark Heaven · released v1.4.2.2
Jev vs Plumb-4B: published benchmark comparison
A side-by-side view of Jev 1.13.0 and Plumb-4B (crh225, JevK5 v0.2 + LoRA) from the same hash-checked v1.4.2.2 aggregate. The overall score is a composite; compare the separate measures against your use case.
Four-radar comparison: Jev and Plumb-4B
Four radars compare this fixed pair across the score axes, accuracy per tier, and accuracy by family on the hard tier and sealed set. Further out is better on every spoke.
0–100, the values in the table. A label-only system has no calibration (counted as 0).
Accuracy per tier, incl. sealed
Share correct per tier; Sealed = the 308 private decisions, aggregate only.
Hard tier by family (v1.2 topics)
Plumb-4B has no published v1.2 hard-tier family breakdown.
Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out).
Sealed set by family
Share correct within each sealed family — system-level aggregates; the items stay private.
Published values
Plumb-4B (crh225, JevK5 v0.2 + LoRA) has the higher published JevBench Score. The displayed score uses one decimal place; the comparison above uses the same published release.
All values as a table
Measure
Jev 1.13.0 (TypeSafe AI)
Plumb-4B (crh225, JevK5 v0.2 + LoRA)
Published rank
#4
#2
JevBench Score
63.3
65.8
Sealed-set accuracy
36.7%
38.0%
Intelligence axis
53.1
53.0
Calibration axis
76.3
75.5
Speed axis
83.3
93.5
Cost axis
52.0
55.8
Cost evidence
measured
estimated
Cost per 1,000 decisions
$0.040 per 1,000 decisions
$0.030 per 1,000 decisions
Intelligence, Calibration, Speed and Cost — equal-weight harmonic mean, with generalization and Jev-class gates. Cost evidence is labeled per row. Speed is a benchmark axis; deployment latency depends on the endpoint and conditions.
public tariff x measured tokens (https://docs.typesafe.ai/models (output tokens not billed)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)
Code and weights marked open in the published row. License note: Apache-2.0 (weights and code; NOTICE credits JevK5, Qwen and SemIf); base Qwen3.5-4B Apache-2.0.
H100 80 GB (evaluator-owned Lium pod) · our evaluator-owned Lium GPU pod (H100 80 GB (evaluator-owned Lium pod)), offline read-only container, author's model served on loopback
Speed adjustment
x2 + 0.15 s (assumption, not measured)
Cost basis
ESTIMATE: EmpirioLabs qwen3-5-4b public pay-as-you-go list price $0.04/M input, $0.07/M output; exact Qwen3.5-4B base model, 25 Sep 2026 cutoff; no output generated by this system. https://empiriolabs.ai/models/qwen3-5-4b
In v1.4.2.2, Jev is rank 4 with a JevBench Score of 63.3; Plumb-4B (crh225, JevK5 v0.2 + LoRA) is rank 2 with a score of 65.8. The table shows their published axes and sealed-set accuracy separately.
Which system has higher sealed-set accuracy?
Plumb-4B (crh225, JevK5 v0.2 + LoRA) has the higher published sealed-set accuracy (38.0% versus 36.7%). This is separate from the composite JevBench Score.
Can I compare the cost values as actual bills?
Jev 1.13.0 (TypeSafe AI): measured at $0.040 per 1,000 decisions. Plumb-4B (crh225, JevK5 v0.2 + LoRA): estimated at $0.030 per 1,000 decisions. Estimated and announced bases are not measured charges; inspect the board’s full row disclosure.
Is Plumb-4B open source, and is Jev?
The published row lists Plumb-4B (crh225, JevK5 v0.2 + LoRA) with license note “Apache-2.0 (weights and code; NOTICE credits JevK5, Qwen and SemIf); base Qwen3.5-4B Apache-2.0”. Jev 1.13.0 is listed as “proprietary API”: TypeSafe AI serves it through its own API and has not published its weights. Check each linked source for the exact terms.