Chinese AI firms are stealing proprietary capabilities belonging to US-based AI models via massive distillation campaigns, according to US government agencies.
Distillation is a common machine-learning practice in which mature “teacher” AI models are used to train “student” AI models. On its own, distillation is a widely accepted practice used for academic research, making models more efficient, and improving models with specialized use cases.
The FBI, the National Security Agency (NSA), and the Cybersecurity and Infrastructure Security Agency (CISA), however, published a joint advisory on Sept. 8 warning that Chinese AI firms were conducting “industrial-scale” efforts to extract capabilities from leading US AI models using distillation. The advisory accuses a number of AI vendors — including Alibaba, DeepSeek, MiniMax, Moonshot AI, StepFun, and Z.AI — of using US large language models (LLMs) to assist training and development of their own proprietary models.
These companies supposedly “extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024,” the advisory stated.
What makes this case different, the advisory explained, is that these firms are allegedly training on outputs obtained in violation of the terms of service, deliberately extracting a competitor’s proprietary capabilities, and using evasive techniques to avoid detection. The US agencies claim this is likely happening with the awareness of the Chinese government.
Inside China’s AI Distillery
The advisory outlines multiple techniques for accessing US frontier models at industrial scale.
“Advanced industrial-scale distillation tactics include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures,” the advisory read. “China-based AI companies that conduct industrial-scale distillation against U.S. AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model.”
To save money during this intensive process, China-based AI firms allegedly obtain bulk premium subscriptions for US AI models and share them across teams of developers.
And to avoid detection, the firms reportedly route distillation requests through native APIs, remote cloud providers, and third-party aggregators that automatically obfuscate user metadata. They also take advantage of “transfer stations,” a gray market of proxies used specifically to bypass US AI model geographic restrictions and evade safeguards.
The advisory made multiple claims specific to different companies. CISA said DeepSeek ran an organized distillation campaign against frontier models of US AI companies since at least late 2024 to “generate synthetic training data for its models.”
“The company targeted specific knowledge domains to extract proprietary functionality and reasoning capabilities to reduce their compute and research costs,” the advisory read. “DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.”
Moonshot AI, meanwhile, allegedly extracted “significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model.” Both firms supposedly plundered models from Anthropic, OpenAI, Google, and xAI as part of this campaign.
What AI Companies Should Do
The authoring agencies recommend that US AI companies implement comprehensive detection and mitigation measures to find anomalous and malicious behavior; share intelligence with other AI organizations to gain awareness of broader campaigns; and tune responses for suspected malicious attempts to reduce the effectiveness of these outputs.
Ismael Valenzuela, vice president of labs, threat research, and intelligence at Arctic Wolf, tells Dark Reading that model extraction and distillation “should be treated as a security event category” in itself, rather than treating it as API abuse or an intellectual property dispute.
“AI companies and organizations should prioritize 24/7 monitoring of AI API access patterns for automation at scale, anomalous prompt harvesting behavior, and distributed account creation,” Valenzuela says. “Organizations should work with model providers to share telemetry and indicators, rather than treating detection as something each organization has to solve alone.”

Comments are closed