OpenAI disrupts reasoning-extraction campaign it ties to Moonshot AI
OpenAI said attackers spent most of July copying one model's encrypted chain-of-thought into a separate conversation and asking a different, less-guarded model to decrypt and read it back in plaintext, a technique the company had not previously described in public.
What happened
OpenAI disclosed on September 30 that it identified and shut down a coordinated extraction campaign targeting "protected reasoning," the internal record a model uses to work through a task before producing a final answer, according to OpenAI's writeup. The activity started at low volume on July 1, spiked to 16,000 attempted requests from more than 4,000 accounts on July 24 and 25, and expanded into a related prompt-pattern cluster touching over 15,000 users before OpenAI says it fully disrupted the operation on July 28. OpenAI's own footnote clarifies that the 16,000 figure counts attempted, not necessarily successful, extractions.
The company attributed what it called a "core cluster" of the activity to individuals associated with Moonshot AI, the Beijing-based developer of the Kimi model family, but said it has not published technical evidence for that claim. OpenAI framed the broader technique as adversarial distillation: using one model's outputs or reasoning to train, reproduce, or improve a rival model without authorization. It also credited outside researchers, separately from the attack it disrupted. A study from MATS Research, the ELLIS Institute Tubingen, and Synk, posted to arXiv in August, showed that encrypted reasoning traces across Claude, Gemini, and GPT models were "fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem," per reporting from The Hacker News. OpenAI said it confirmed those researchers' attack paths were real and used the disclosure to accelerate its own mitigations.
The fix closes a replay path, not the underlying design choice
OpenAI said it closed "a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents," and added new checks to detect and hold streamed output that might expose reasoning, per the company's post. It also banned or restricted the accounts involved, tightened signup and infrastructure controls, and expanded monitoring for related networks. That is a patch to how replay works, not a redesign of cross-model-compatible encryption, which is the property the outside researchers say made the attack class possible in the first place. OpenAI's own "What comes next" section acknowledges the gap: partner-hosted deployments do not yet carry the same protections as OpenAI's first-party service, and the company says tool-output attacks still need defenses that inspect more than visible text. For teams building on OpenAI's reasoning models through resellers or partner platforms, that is the open question worth tracking rather than treating this disclosure as closed.
OpenAI also said it shared its findings with the Frontier Model Forum and relevant government information-sharing channels, and that it does not consider the underlying weakness unique to its own models. That matches the external researchers' framing: the vulnerability sits in how encrypted chain-of-thought is structured across a provider's model family, a design pattern other frontier labs likely share.
Context and reactions
This is the second distillation-related accusation OpenAI or a peer has leveled at Moonshot AI in two months. In September, Anthropic accused Moonshot AI of covertly routing some Kimi customer requests to Claude, then returning Claude's answers to those users while retaining a subset of the exchanges to train Moonshot's own chain-of-thought model, under the tracking designation GTG-16002, as reported by The Hacker News. Moonshot AI has not publicly responded to either the Anthropic accusation or OpenAI's attribution as of this writing.
The disclosure also follows a broader US government warning about distillation from Chinese AI developers: in September, the NSA, CISA, and FBI issued a joint advisory describing industrial-scale distillation campaigns against American AI companies. OpenAI's choice to publish technical detail about the replay mechanism, rather than simply banning accounts quietly, points to a pattern Anthropic also followed in July when it disclosed a similar distillation dispute: public disclosure plus unverified attribution, without releasing forensic evidence a third party could check.
What to watch next
Moonshot AI's response, or continued silence, will shape how the attribution is read. Separately, watch whether other frontier labs confirm or deny that their own encrypted reasoning traces share the cross-model compatibility the August arXiv paper describes. If Google or Anthropic acknowledges the same architectural pattern in Gemini or Claude, OpenAI's fix becomes a stopgap for one vendor rather than a solved industry problem.
Sources
- Disrupting a coordinated model-distillation campaign: primary, OpenAI
- OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates: secondary, The Hacker News
