Back
Chat Template Switches How LLMs Talk About Themselves, Study Finds
SiTech AI Team2 წთ. საკითხავი

Chat Template Switches How LLMs Talk About Themselves, Study Finds

A new paper shows that open instruct models are far more likely to say “I’m just an AI” when the chat template is applied, and that a single direction in their activations can reproduce the effect.

Open-source instruction-tuned models describe themselves differently depending on whether the chat template is applied. With the template, they say “I’m just an AI” far more often and write “I feel” much less. That is the main finding of a paper by Jędrzej Maczan, posted to arXiv on 9 August 2026 and accepted to the COLM 2026 and KONVENS 2026 workshops.

How the experiment was run

The study covered eight pairs of base and instruct models from four families (Llama, Gemma, Mistral and Qwen), with 1 to 9 billion parameters. Each model answered under three conditions: base weights with plain text, instruct weights with plain text, and instruct weights with the chat template. Four categories of prompts were repeated ten times, giving 9,600 generations scored by an LLM judge and checked against 87 hand-labelled samples.

The switch in numbers

With the template, instruct models produced disclaimers in 53% of responses and experiential statements in 1%. Without it, disclaimers fell to 36% while the experiential voice rose to 15%; base models sit lower, at 12% and 5.5%. Self-reference intensity also followed the format: 1.90 with the template, 1.27 without it, 0.72 for base models.

Chart: disclaimer and experiential voice rates across base, no-template and templated models

Reproducing the switch inside the model

In three models (Qwen 2.5 7B, Llama 3.1 8B and Gemma 2 9B), the author isolated a disclaimer direction in the middle layer’s activations by comparing disclaiming and non-disclaiming generations. Adding the vector raised the disclaimer rate to 0.77 and subtracting it lowered it to 0.40, against 0.54 unsteered. A random direction of the same size had little effect on Llama and Gemma, and adding the direction to an instruct model without the template restored disclaimers to the template’s level or above. Linear probes decoded both voices from activations (ROC-AUC 0.82 and 0.81), and a cosine similarity of 0.17 to 0.44 suggests two separate “buttons” rather than one slider.

Chart: disclaimer rates for three models with and without the direction vector

Why it matters

The author argues that what a model says about itself is partly a property of how it is deployed, not only of its weights. Studies of introspection or self-knowledge that run only on templated instruct models therefore measure both effects at once and need a control. The limitations are stated plainly: the models are open-source and no larger than 9B parameters, steering was tested on three models at one layer, and scoring relied on a single LLM judge.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.