
OpenAI disrupts a coordinated model-distillation campaign
OpenAI says it identified and disrupted a coordinated campaign to extract protected reasoning from its models, and attributes a core cluster of the activity to people linked to Moonshot AI, the developer of Kimi.
OpenAI said it identified and disrupted a coordinated campaign that tried to extract protected reasoning from its models, with the earliest activity dating to the first week of July. The company describes the activity as adversarial distillation: the systematic, unauthorized use of one model's outputs or reasoning to train, reproduce or improve another model.
What OpenAI observed
According to the company, the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. Instead they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester, at scale, in violation of its terms of service. Among the techniques was copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content. Independent security researchers flagged related cross-model and conversation-compaction vulnerabilities through responsible disclosure, and OpenAI confirmed the attack paths were real.
The activity began on 1 July and stayed at low volume until spikes on 24 and 25 July, when OpenAI recorded 16,000 requests using a relevant extraction pattern from more than 4,000 users. Further investigation found related prompt-pattern activity across a cluster of more than 15,000 users, which the company says it fully disrupted by 28 July.
Who OpenAI attributes it to
OpenAI says it is unclear whether every operator came from a single actor. It attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
Why it matters
OpenAI frames adversarial distillation as a safety and national security risk. Extracted reasoning can be used to train another model without the safeguards applied to the original model's user-facing outputs, and at scale it can speed the transfer of advanced capabilities. The company stresses the problem is not unique to OpenAI and calls it a shared security challenge for the industry. It notes that the figures describe attempted, not necessarily successful, extractions.
How OpenAI responded
The company banned or restricted fraudulent accounts, tightened signup and infrastructure controls, and expanded monitoring for related networks. It also strengthened protections for hidden reasoning across users, workspaces, organizations and model families, closed a pathway that let someone replay another user's encrypted reasoning and recover its contents, and added checks to detect and hold streamed output that might expose reasoning. Findings were shared through the Frontier Model Forum and government information-sharing channels. Going forward, OpenAI says it will focus on stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing across industry and government.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.