
Chat Template Switches How LLMs Talk About Themselves, Study Finds
A new paper shows that open instruct models are far more likely to say “I’m just an AI” when the chat template is applied, and that a single direction in their activations can reproduce the effect.
Open-source instruction-tuned models describe themselves differently depending on whether the chat template is applied. With the template, they say “I’m just an AI” far more often and write “I feel” much less. That is the main finding of a paper by Jędrzej Maczan, posted to arXiv on 9 August 2026 and accepted to the COLM 2026 and KONVENS 2026 workshops.
How the experiment was run
The study covered eight pairs of base and instruct models from four families (Llama, Gemma, Mistral and Qwen), with 1 to 9 billion parameters. Each model answered under three conditions: base weights with plain text, instruct weights with plain text, and instruct weights with the chat template. Four categories of prompts were repeated ten times, giving 9,600 generations scored by an LLM judge and checked against 87 hand-labelled samples.
The switch in numbers
With the template, instruct models produced disclaimers in 53% of responses and experiential statements in 1%. Without it, disclaimers fell to 36% while the experiential voice rose to 15%; base models sit lower, at 12% and 5.5%. Self-reference intensity also followed the format: 1.90 with the template, 1.27 without it, 0.72 for base models.

Reproducing the switch inside the model
In three models (Qwen 2.5 7B, Llama 3.1 8B and Gemma 2 9B), the author isolated a disclaimer direction in the middle layer’s activations by comparing disclaiming and non-disclaiming generations. Adding the vector raised the disclaimer rate to 0.77 and subtracting it lowered it to 0.40, against 0.54 unsteered. A random direction of the same size had little effect on Llama and Gemma, and adding the direction to an instruct model without the template restored disclaimers to the template’s level or above. Linear probes decoded both voices from activations (ROC-AUC 0.82 and 0.81), and a cosine similarity of 0.17 to 0.44 suggests two separate “buttons” rather than one slider.

Why it matters
The author argues that what a model says about itself is partly a property of how it is deployed, not only of its weights. Studies of introspection or self-knowledge that run only on templated instruct models therefore measure both effects at once and need a control. The limitations are stated plainly: the models are open-source and no larger than 9B parameters, steering was tested on three models at one layer, and scoring relied on a single LLM judge.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.