CISA Warns Chinese AI Firms Stealing Billions of LLM Tokens
Key Takeaways Chinese AI firms are accused of orchestrating large-scale “knowledge distillation” attacks against leading U.S. large language models (LLMs). These firms, including...
Key Takeaways
- Chinese AI firms are accused of orchestrating large-scale “knowledge distillation” attacks against leading U.S. large language models (LLMs).
- These firms, including DeepSeek, Moonshot AI, and Alibaba, allegedly extracted billions of tokens from models like Claude, GPT, Gemini, and Grok since late 2024.
- The attacks leverage automated requests via APIs, cloud services, and proxy networks to harvest model outputs and accelerate the development of rival AI systems.
- U.S. government agencies, including CISA, NSA, and FBI, warn this activity poses significant economic and national security risks by bypassing costly and time-consuming AI development.
- Recommended mitigations for AI providers include enhanced identity verification, rate limiting, advanced logging, and collaborative threat intelligence sharing.
A recent advisory from the U.S. government has brought to light a significant concern regarding the illicit replication of advanced artificial intelligence capabilities. This activity, distinct from traditional cyberattacks, involves the systematic harvesting of outputs from sophisticated AI models on an unprecedented scale.
Table Of Content
The alleged operations employed massive volumes of automated queries, routed through application programming interfaces (APIs), various cloud platforms, data aggregators, and intricate proxy networks. The objective, as CISA said in a report, is to amass synthetic datasets. These datasets are then used to train competing AI systems, enabling them to mimic the advanced functionalities of their Western counterparts.
Analysts from the Cybersecurity and Infrastructure Security Agency (CISA), in collaboration with the National Security Agency (NSA) and the Federal Bureau of Investigation (FBI), assert that AI companies based in China have likely siphoned billions of tokens through millions of interactions with leading U.S. frontier models since late 2024. The advisory characterizes this as industrial-scale knowledge distillation, moving beyond the scope of legitimate AI research. This unauthorized extraction of reasoning, coding, agentic, and domain-specific capabilities could drastically reduce the expense and timeline required to develop competitive AI models, raising substantial economic and national security implications for the global AI landscape.
CISA Warns Chinese AI Firms Extract Billions of Tokens
CISA specifically identified several Chinese AI companies implicated in these campaigns, including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. These entities reportedly targeted various iterations of prominent models such as Claude, GPT, Gemini, and Grok. While the advisory did not claim direct governmental control over these operations, it noted that the activities were likely conducted with the awareness of the Chinese government.
Knowledge distillation is a recognized technique where a smaller, more efficient model learns from a larger, more complex one. However, the concern arises when companies allegedly obtain restricted outputs from competitors on a massive scale to replicate protected capabilities without authorization. This mirrors previously reported large-scale AI distillation attacks.
DeepSeek is reported to have engaged in organized data collection from at least late 2024 through mid-2025. Their alleged targets included reasoning capabilities, specialized optimizations, legal functions, and writing support for its R1 and V3 models. CISA noted that DeepSeek’s publicly stated training costs likely do not reflect the true value derived from these alleged distillation efforts.
Moonshot AI was also linked to extensive activity starting around mid-2025, reportedly extracting data from Claude Fable 5 for its Kimi-K3 model and GPT-4o data for its Kimi-K2 model. Other capabilities allegedly targeted included programming, mathematics, reinforcement learning, and software engineering functions. Alibaba was cited for using distillation to enhance its software engineering, customer service, character creation, and training workflows. Allegations of unauthorized Claude model extraction have previously highlighted the growing issue of model-output collection for AI providers.
Proxies and Prompt Attacks
According to the CISA advisory, these operations frequently utilized “transfer stations”—a grey market for API proxies. These intermediaries enable threat actors to obscure user metadata, bypass geographical restrictions, and mask the true origin of requests, making it challenging to link isolated accounts to a larger, coordinated campaign.
The observed tactics included the use of account pools, bulk premium subscriptions, and automated routing systems designed to switch between providers as access controls adapted. Behavioral patterns indicative of these malicious activities included continuous 24/7 activity, repeated usage from diverse geographic locations, immediate maximum usage upon new account activation, and synchronized timing across disparate access pathways.
Furthermore, some operators allegedly employed advanced prompt injection and “jailbreak” techniques to compel models to reveal their underlying chain-of-thought reasoning. This differs significantly from standard prompting, as the intent is to manipulate the model into exposing proprietary internal processes, a risk extensively covered in discussions on prompt injection attack techniques.
What You Should Do
- AI providers should implement robust identity verification measures and continuously monitor for unusual subscription-to-usage ratios.
- Apply stringent rate limits to API access and maintain comprehensive logs of all requests for detailed investigation.
- Share infrastructure and behavioral threat intelligence with cloud platforms and API aggregators to detect distributed campaigns that might not be visible from a single service.
- For high-confidence malicious requests, consider implementing targeted response changes such as reducing response fidelity or varying outputs subtly, without alerting the suspected operators.
- Enhance defenses with differential privacy, adversarial testing, stricter API controls, and mechanisms to limit prompt injection to fortify against extraction attempts.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



No Comment! Be the first one.