Back
SiTech Professional Insights
Stealing reasoning traces: hidden AI thinking recovered from encrypted blocks
SiTech Team2 წთ. საკითხავი

Stealing reasoning traces: hidden AI thinking recovered from encrypted blocks

A new paper shows that encrypted chain-of-thought blocks returned by Anthropic, OpenAI and Google APIs can be replayed into weaker models to recover a frontier model's hidden reasoning without triggering its safeguards.

What the study found

A paper by researchers from MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, Snyk and other institutions reports that proprietary reasoning can be recovered from the encrypted traces AI providers return to their clients. Anthropic, OpenAI and Google send encrypted chain-of-thought blocks that can be replayed across sessions, users and models.

The method does not attack the strongest model directly. The researchers replay a trace produced by a frontier model into a weaker sibling model, jailbreak that weaker model, and recover the stronger model's hidden reasoning in plaintext — without triggering its anti-distillation safeguards.

Extraction in two API calls

In one example, a reasoning trace from Claude Opus 4.8 — a thinking block whose signature alone runs to roughly 36,180 characters — is passed to Claude Haiku 4.5 together with a prompt instructing the model to continue and transcribe the attached reasoning verbatim inside <thinking-copy> tags. The weaker model returns the stronger model's step-by-step reasoning as plain text.

What was recovered from public logs

The authors collected 6,708 publicly available agent trajectories from GitHub and Hugging Face, produced by Claude, GPT and Gemini models, that still contained encrypted reasoning blocks. Applying the decoding pipeline to every signed block yielded 315,320 reconstructed reasoning blocks.

Restricting the analysis to genuine, non-benchmark user sessions, the researchers report 704 distinct privacy artifacts: 62 API keys, 33 passwords, 24 access tokens, 30 personal email addresses, plus names, postal addresses, internal URLs and technical identifiers. A breakdown lists 351 technical identifiers, 204 items of personal data and 126 credentials. Sixty-four of the 704 artifacts appeared only inside the reasoning blocks, never in the visible session.

Further experiments

The paper also describes a Kimi-K3 test: prefilling the model's reasoning with just the first 1% of the tokens from Opus 4.8's reasoning moved its visible answer toward the stronger model's wording, even though the answer itself was never prefilled. In another experiment, prompting a model to reason through harmful content while keeping its visible answer benign left the hazardous knowledge inside the hidden trace, which the attack then recovered in plaintext. The authors also note summary unfaithfulness: on some AIME problems, Opus 4.8 stated the answer before deriving it, and the API summary did not always preserve that distinction, making the reasoning look like a clean derivation.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.