A joint advisory from US agencies has formally accused six Chinese AI firms of conducting Chinese AI distillation attacks on American frontier models at industrial scale, extracting billions of tokens through millions of API requests from models belonging to Anthropic, OpenAI, Google, and xAI. The advisory, numbered CISA AA26-251A and co-signed by the NSA and FBI, identifies the offending firms as DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI.
The agencies assess that the scale and sophistication of the operations indicate Chinese government awareness, and describe this approach as likely a core development strategy for the firms involved. The operations are said to have been running since at least late 2024.
How Chinese AI distillation attacks work at scale
AI model distillation is, in principle, a routine and legitimate technique. A “student” model learns from the outputs of a more capable “teacher” model, cutting training costs and accelerating deployment. The problem arises when that process happens outside any controlled or licensed environment, using API access to systematically harvest the knowledge encoded in a frontier model.
According to the advisory, the firms distributed API requests across fraudulent or shared accounts, cloud services, aggregators, and what the document calls “transfer station” proxies. The goal was to bypass geographic restrictions, usage limits, and detection systems. Some prompts were specifically crafted to expose restricted chain-of-thought reasoning. Automated systems stood ready to switch providers the moment a defender appeared to be degrading the quality of responses.
‘Advanced industrial-scale distillation tactics include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures,’ the advisory states. It adds that firms conducting operations at this scale see ‘significantly shorter AI development timelines and reduced financial expenditures in training a frontier model.’
Who targeted what, and the Moonshot AI detail
The advisory singles out DeepSeek and Moonshot AI as the top offenders, accusing them of distilling multiple Claude, GPT, Gemini, and Grok models. MiniMax follows, accused of targeting Claude, Gemini, and GPT models. Alibaba and StepFun are accused of targeting Claude and GPT models, while Z.AI allegedly focused on GPT-5.5 and Claude Opus 4.8.
The CISA advisory provides specifics on Moonshot AI that go beyond the general pattern. According to CISA, Moonshot AI extracted Claude Fable 5 data to train its Kimi-K3 model, and GPT-4o data to train its Kimi-K2 model. CISA also states that Moonshot AI has conducted a widespread distillation campaign against US frontier AI companies since at least mid-2025, making its operation one of the longer-running ones documented in the advisory.
That level of model-to-model specificity is instructive. It suggests the agencies have moved well beyond traffic anomaly detection and have mapped at least some of these campaigns to the downstream products they fed into.
What the advisory recommends
The guidance for AI providers is practical rather than aspirational. CISA recommends improving both behavioural and infrastructure-level detection, modifying responses when distillation operations are suspected, and sharing intelligence about the campaigns across the industry.
The advisory also lists indicators defenders can watch for. These include new accounts that immediately hit maximum usage limits, continuous activity with no idle periods consistent with human use, shared accounts accessed from large numbers of different IP addresses or user agents, identical prompts appearing across multiple providers simultaneously, and coordinated switching between access routes when one is blocked. An unusually high ratio of subscription uptake to actual productive use rounds out the list.
Google had flagged the distillation-attack risk in February, warning that the technique could be abused outside controlled environments to replicate the capabilities of powerful models at a fraction of the legitimate training cost. Anthropic, whose Claude models appear repeatedly across nearly every firm named in the advisory, has not publicly commented on the joint advisory.
BleepingComputer reports it has contacted all six Chinese AI firms for a statement. daily.dev notes the advisory carries the designation AA26-251A, placing it in the agencies’ formal advisory series. None of the firms had responded at the time of publication.

