
Jeff: Jev-compatible 0.8B decision models trained on one local GPU
The open-source project Jeff fine-tunes Qwen3.5 and Gemma 4 into small zero-shot classifiers that return a calibrated probability per option in one forward pass, at about 22 ms per decision.
An open-source project called Jeff has appeared on GitHub: it turns small 0.8B and 2B models from Qwen3.5 and Gemma 4 into zero-shot classifiers that use the same request format as Jev. The caller describes a situation and lists the options in plain words, and Jeff returns a calibrated probability for each option from a single forward pass. No text is generated, so nothing has to be parsed afterwards. A decision takes about 22 ms on an NVIDIA RTX PRO 6000 and 28 ms on an Apple M4 Max through MLX.
What Jeff is
Zero-shot means the options can be anything: support queues, user intents, moderation labels or game moves. The categories do not have to appear in the training data, because the caller describes them in the request. Three models are published on Hugging Face: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B. One request can carry three question types: a choice between up to 255 options, a yes/no answer returned as a probability, and a score on a described scale.
Benchmark results
The models were tested on 4,599 questions from five public benchmarks, plus the 105-item hard tier of JevBench scored separately. On the overall score Jeff-Qwen3.5-2B reaches 83.1, Jeff-Gemma4-E2B 81.6 and Jeff-Qwen3.5-0.8B 79.1, against 83.0 for Jev's published figures. On Financial PhraseBank the small models climb to 96.4 against 77.0, while on reasoning-heavy tests they fall behind: 64.0 on BBH against 94.3. The authors write that the overall score comes from classification, where the small models match the large ones.

Trained on local hardware
The project was built on local hardware: the 0.8B model trains in about two hours on one RTX PRO 6000 workstation GPU, the 2B in three and a half, with synthetic data written by the open model Qwen3.8-Flash-Next on two DGX Sparks and testing on a MacBook. No cloud GPUs were used, and no closed-model output went into the training data. Jeff is independent of TypeSafe, the makers of Jev; the code is MIT-licensed and the weights are Apache 2.0.
Games and caveats
In a zero-shot test Jeff played three games, each turn described in words: 6.55 kills in Doom, the same as a hand-coded rule bot, 10.3 crossings in Frogger against 1.0 untrained, and 57.0 of 98 pellets in Pac-Man against 25.8, deciding in 29 to 49 ms per move on an M4 Max. The README also lists limits: small models do not reason, they work on English text only, benchmark scores do not predict game play, and the 2B plays worse than the 0.8B. A short fine-tune on your own examples goes further: a voice-navigation run on about 11,000 examples lifted held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.