
JevBench v1.4.1: a reproducible benchmark for typed decision models
Benchmark Heaven has published JevBench v1.4.1, a reproducible benchmark for Jev-class decision models. 82 systems were scored on 534 public and 308 sealed decisions; Jev 1.13.0 leads the ranking with 63.3 points.
Benchmark Heaven has published JevBench v1.4.1, a reproducible benchmark for Jev-class decision models — systems that take a state and a bounded rubric and return a typed answer. The version scored on 23 September 2026 measures 82 systems on 534 public and 308 sealed decisions, and ranks 77 of them. TypeSafe AI's Jev 1.13.0 leads with 63.3 points, with the open rebuild JevK5 v0.2.0 just 1.3 points behind.
How the score works
JevBench combines four axes — Intelligence, Calibration, Speed and Cost — each scored from 0 to 100, into an equal-weight harmonic mean. Two gates then punish lopsided systems: Intelligence, Speed or Cost below 50 pulls the score down quadratically, and a public-to-sealed accuracy gap above 25 percentage points reduces Intelligence. The newest addition is the sealed set: 308 fresh private decisions contribute 20% of Intelligence, while only system-level aggregates are published — item text, answers and per-item results stay private and rotate between versions. Chance on the sealed set is 29.3%.
The ranking
Jev 1.13.0 scores 63.3 (Intelligence 53.1, Calibration 76.3, Speed 83.3, Cost 52.0) at $0.040 per 1,000 decisions. Behind it are JevK5 v0.2.0 (62.0, ~$0.022 estimated), Hopper (59.4), Winnow-12B Q8 (55.6), reflex 4B (54.0) and djev (52.2, $0.026 announced). The strongest general model, GPT-6 Luna, has the field's highest Intelligence, 97.4, and its best sealed accuracy, 95.5%, but places only 30th: weak Speed (72.6) and Cost (36.0) scores cannot be compensated in the harmonic mean.
What it says about Jev-class models
Public items are nearly saturated for the leaders — Jev 1.13.0 answers 86.6% of them correctly — while the sealed set is far harder, at 36.7%. Gaps of roughly 48–57 points across the top systems trigger the generalization penalty. The benchmark's own notes call the 534-decision set a pilot, note that it is English-only, and warn that held-out items are still sent to the evaluated services. Harness, public tasks and scoring rules are MIT-licensed. A service running Jev, classifier.dev, scored 70.8 but is not ranked, since ranking it would rank the same model twice; separately, an independent analysis published the same day argued that a decision model's probabilities cannot stay calibrated across arbitrary user distributions and should be treated as scores.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.