Back
Turning Qwen3-1.7B into a Single-Pass Decision Model with Calibrated Probabilities
SiTech AI Team2 min read

Turning Qwen3-1.7B into a Single-Pass Decision Model with Calibrated Probabilities

A practical walkthrough shows how to constrain a language model's output to fixed answer choices, then evaluate it on CommonsenseQA and fix overconfident scores with temperature scaling.

Constrained decoding can turn an everyday language model into a decision engine that picks from a fixed set of options in a single forward pass, instead of generating free-form text token by token. The approach, demonstrated with Qwen/Qwen3-1.7B, masks the vocabulary down to a handful of allowed tokens so the model can only emit one of the predefined answers.

How the single-pass trick works

Standard System one decision models infer and respond with calibrated probabilities across every allowed answer. When a generic language model is asked for JSON output through Structured Output, it still has to run a pass for every token it generates; in the example shown, producing the final output took 11 passes. A decision model such as Jev assumes a fixed set of options, masks everything else out of the vocabulary, and reads the highest probability answer from the first pass. This does not guarantee a correct answer, and raw token probabilities may reflect confidence in the next token rather than the true probability of being correct.

CommonsenseQA results and overconfidence

On a random holdout sample of CommonsenseQA, the 1.7B model answered 725 of 1221 questions correctly, an accuracy of 0.5938 with a macro F1 of 0.5844. A quick finetune on the dataset improved this to 762 of 1221, or 0.6241 accuracy and a macro F1 of 0.6234. The problem surfaced on ambiguous questions. Asked where you would most likely find a bat, with options such as Cave, Baseball game, Attic, Zoo and Sporting goods store, the model picked Cave with 0.9978 probability even though no clear answer existed. Binned evaluation confirmed the overconfidence: in the 0.90 to 1.00 confidence bin, holding 809 samples, accuracy was only 0.7009, and in the 0.80 to 0.90 bin, with 121 samples, it dropped to 0.4711.

Calibration with temperature scaling

To align confidence with accuracy, the author applied temperature scaling, fitting the temperature parameter to the model's accuracy through curve fitting. The fitted value was 3.797280788421631. After scaling, binned confidence tracked accuracy far more closely: the top bin showed 0.9333 confidence against 0.9541 accuracy, and mid-range bins fell within a few percentage points of their measured accuracy. The author published a GitHub repository with scripts for building a dataset, evaluating, finetuning and calibrating a model, and encourages trying the workflow on larger models.

Sources: nishtahir.com

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.