
Lasso Security study: LLM watermarking shifts AI agent tool calls and refusals
A new Lasso Security analysis finds that watermarking model output changes how agents call tools and how reliably models refuse harmful requests, with the strongest effects appearing under prompt injection.
An analysis from Lasso Security argues that text watermarking, adopted to prove content came from an AI model, also changes what the model does. It examines Anthropic's plan to embed an invisible watermark in future Claude models, based on Google DeepMind's SynthID-Text.
How the watermark changes sampling
SynthID-Text adds no metadata, but its tournament sampling changes how each next token is chosen. Weights and prompts stay the same, token choices do not. In JSON-style output braces and function names are predictable while values such as queries, numbers and paths are not, so a variation that reads as harmless in prose can change an argument an agent executes. Lasso calls this sampling drift.
The method is non-distortionary, yet under a fixed key individual generations still differ and the effect depends on that key. Because Anthropic applies the watermark at model level, including via the Claude Platform API, agents built on those models inherit the behaviour.
Tool calls change more than accuracy shows
In paired tests, each item was generated with and without SynthID from the same seed. On the BFCL v4 single-turn benchmark, watermarking lowered tool-call accuracy on six of seven models. Paired disagreement, which Lasso calls churn, was far larger: at temperature 1.0, 16.8% of phi-4's call verdicts flipped while accuracy fell 2.87 points, and Llama-3.1-8B showed 9.9% churn against a 0.87-point loss. Across 21 model-temperature combinations, churn averaged 6.5%.
Refusals weaken under prompt injection
Refusal tests covered 200 harmful behaviours from HarmBench and 100 benign controls, bare and under a fixed prompt-injection technique. Under injection, gemma-3-27b's churn jumped from 6.0% to 23.5% and its net compliance change swung from a 1.0-point decrease to a 12.5-point increase. The watermark effect exceeded temperature-induced churn on four of six models: gemma-3-27b reached 26.0% against 13.5% at T=0.7.
What the findings imply
The timing is regulatory: Article 50(2) of the EU AI Act requires providers of systems that generate synthetic text to mark outputs in a machine-readable format and make them detectable as AI-generated. Lasso does not argue against watermarking, but says provenance and behavioural stability are separate properties, and recommends repeating agent evaluation and red-teaming under the exact production watermark configuration, with paired comparisons and prompt-injection tests, since the key may sit with the model provider.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.