Summary

  • OpenAI reports that a campaign initiated on July 1 resulted in 16,000 extraction attempts from over 4,000 users on July 24-25 alone, within a larger group exceeding 15,000 users.
  • The company identifies a core group of this activity as being linked to individuals associated with Moonshot AI, the creator of Kimi, although it remains uncertain if all participants were from a single source.
  • The objective was to access the internal "reasoning" of OpenAI's AI models, which could potentially allow for the training of another model devoid of the original's safeguards.

OpenAI has announced that it has dismantled a coordinated effort aimed at replicating the internal reasoning processes of its AI models. This initiative is believed to involve individuals connected to Moonshot AI, the Chinese company responsible for the Kimi chatbot.

The campaign reportedly commenced on July 1, with OpenAI recording 16,000 extraction requests from over 4,000 distinct users on July 24 and 25, part of a broader group exceeding 15,000 users. By July 28, OpenAI claims to have effectively halted the activity.

The focus of the campaign was not on obtaining answers but rather on the underlying processes that generate those answers.

Contemporary AI systems engage in a reasoning process before delivering responses, working through problems step-by-step in an internal format before presenting the final answer. OpenAI secures this internal format, asserting that unauthorized access could disclose information not included in the final output.

“The operators did not breach our encryption, compromise any databases, or gain direct access to stored user conversations,” OpenAI stated. “Instead, they manipulated model interactions to reproduce protected reasoning in a manner visible to requesters, violating our terms of service.”

One technique involved extracting encrypted reasoning from one conversation and requesting a model to decode it in another.

Although OpenAI's statement does not directly link this campaign to K3, it suggests some ambiguity. “It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi,” OpenAI noted.

Following the incident, OpenAI has closed a method that allowed individuals with access to another user's encrypted reasoning to replay it and retrieve its contents.

Why would anyone seek this information? Because distillation, which involves training a new AI on the outputs of a more powerful one, can yield better results from smaller models without extensive training.

When done without permission, OpenAI refers to this as adversarial distillation: "the systematic and unauthorized use of one model's outputs or reasoning to assist in training, reproducing, or enhancing another model."

This scenario is just one among numerous controversies surrounding AI companies. A prominent issue is the unauthorized use of copyrighted data for model training. Distillation, however, does not cross that line, as AI outputs are not subject to copyright, leading companies to implement prohibitions in their terms of service to prevent competitors from utilizing those outputs.

A Recurring Allegation

OpenAI has faced similar accusations in the past. In January 2025, the company indicated it was investigating signs that DeepSeek may have distilled its models, coinciding with concerns over national security.

In February, Anthropic accused Chinese laboratories of employing approximately 24,000 fake accounts to generate over 16 million interactions with its AI, Claude. Critics responded by pointing out that Claude itself was trained on publicly available internet data.

By April, the White House was asserting that foreign entities, predominantly from China, were conducting large-scale distillation campaigns. Shortly thereafter, Elon Musk acknowledged in court that xAI had utilized distillation on OpenAI models to develop Grok.

In June, Anthropic appealed to Congress, advocating for penalties against large-scale model extraction practices.

In August, researchers revealed that OpenAI, Anthropic, and Google had each secured their reasoning processes with a single company-wide encryption key, allowing attackers to manipulate models into revealing hidden reasoning in plaintext. Following the disclosure, all three companies implemented server-side security updates, although previously shared session logs remain decodable.

As of now, Moonshot has not responded to OpenAI's statements and is pursuing a $3 billion IPO in Hong Kong, aiming for a valuation of $50 billion.

Daily Debrief Newsletter

Stay updated with the latest news stories, original features, podcasts, videos, and more.