Hackers News Hackers News
  • CyberSecurity News
  • Threats
  • Attacks
  • Vulnerabilities
  • Breaches
  • Comparisons

Social Media

Hackers News Hackers News
  • CyberSecurity News
  • Threats
  • Attacks
  • Vulnerabilities
  • Breaches
  • Comparisons
Search the Site
Popular Searches:
technology Amazon AI
Recent Posts
Critical Windows Defender ShieldCrash 0-Day Lets Attackers Read Files as SYSTEM
September 9, 2026
Critical Windows BitLocker Flaw Lets Attackers Remotely Execute Code
September 9, 2026
Critical cPanel CVE-2024-XXXX Vulnerability Lets Attackers Gain Full Server Control
September 9, 2026
Home/Threats/CISA Warns Chinese AI Firms Stealing Billions of LLM Tokens
Threats

CISA Warns Chinese AI Firms Stealing Billions of LLM Tokens

Key Takeaways Chinese AI firms are accused of orchestrating large-scale “knowledge distillation” attacks against leading U.S. large language models (LLMs). These firms, including...

Marcus Rodriguez
Marcus Rodriguez
September 9, 2026 4 Min Read
2 0

Key Takeaways

  • Chinese AI firms are accused of orchestrating large-scale “knowledge distillation” attacks against leading U.S. large language models (LLMs).
  • These firms, including DeepSeek, Moonshot AI, and Alibaba, allegedly extracted billions of tokens from models like Claude, GPT, Gemini, and Grok since late 2024.
  • The attacks leverage automated requests via APIs, cloud services, and proxy networks to harvest model outputs and accelerate the development of rival AI systems.
  • U.S. government agencies, including CISA, NSA, and FBI, warn this activity poses significant economic and national security risks by bypassing costly and time-consuming AI development.
  • Recommended mitigations for AI providers include enhanced identity verification, rate limiting, advanced logging, and collaborative threat intelligence sharing.

A recent advisory from the U.S. government has brought to light a significant concern regarding the illicit replication of advanced artificial intelligence capabilities. This activity, distinct from traditional cyberattacks, involves the systematic harvesting of outputs from sophisticated AI models on an unprecedented scale.

Table Of Content

  • Key Takeaways
  • CISA Warns Chinese AI Firms Extract Billions of Tokens
  • Proxies and Prompt Attacks
  • What You Should Do

The alleged operations employed massive volumes of automated queries, routed through application programming interfaces (APIs), various cloud platforms, data aggregators, and intricate proxy networks. The objective, as CISA said in a report, is to amass synthetic datasets. These datasets are then used to train competing AI systems, enabling them to mimic the advanced functionalities of their Western counterparts.

Analysts from the Cybersecurity and Infrastructure Security Agency (CISA), in collaboration with the National Security Agency (NSA) and the Federal Bureau of Investigation (FBI), assert that AI companies based in China have likely siphoned billions of tokens through millions of interactions with leading U.S. frontier models since late 2024. The advisory characterizes this as industrial-scale knowledge distillation, moving beyond the scope of legitimate AI research. This unauthorized extraction of reasoning, coding, agentic, and domain-specific capabilities could drastically reduce the expense and timeline required to develop competitive AI models, raising substantial economic and national security implications for the global AI landscape.

CISA Warns Chinese AI Firms Extract Billions of Tokens

CISA specifically identified several Chinese AI companies implicated in these campaigns, including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. These entities reportedly targeted various iterations of prominent models such as Claude, GPT, Gemini, and Grok. While the advisory did not claim direct governmental control over these operations, it noted that the activities were likely conducted with the awareness of the Chinese government.

Knowledge distillation is a recognized technique where a smaller, more efficient model learns from a larger, more complex one. However, the concern arises when companies allegedly obtain restricted outputs from competitors on a massive scale to replicate protected capabilities without authorization. This mirrors previously reported large-scale AI distillation attacks.

DeepSeek is reported to have engaged in organized data collection from at least late 2024 through mid-2025. Their alleged targets included reasoning capabilities, specialized optimizations, legal functions, and writing support for its R1 and V3 models. CISA noted that DeepSeek’s publicly stated training costs likely do not reflect the true value derived from these alleged distillation efforts.

Moonshot AI was also linked to extensive activity starting around mid-2025, reportedly extracting data from Claude Fable 5 for its Kimi-K3 model and GPT-4o data for its Kimi-K2 model. Other capabilities allegedly targeted included programming, mathematics, reinforcement learning, and software engineering functions. Alibaba was cited for using distillation to enhance its software engineering, customer service, character creation, and training workflows. Allegations of unauthorized Claude model extraction have previously highlighted the growing issue of model-output collection for AI providers.

Proxies and Prompt Attacks

According to the CISA advisory, these operations frequently utilized “transfer stations”—a grey market for API proxies. These intermediaries enable threat actors to obscure user metadata, bypass geographical restrictions, and mask the true origin of requests, making it challenging to link isolated accounts to a larger, coordinated campaign.

The observed tactics included the use of account pools, bulk premium subscriptions, and automated routing systems designed to switch between providers as access controls adapted. Behavioral patterns indicative of these malicious activities included continuous 24/7 activity, repeated usage from diverse geographic locations, immediate maximum usage upon new account activation, and synchronized timing across disparate access pathways.

Furthermore, some operators allegedly employed advanced prompt injection and “jailbreak” techniques to compel models to reveal their underlying chain-of-thought reasoning. This differs significantly from standard prompting, as the intent is to manipulate the model into exposing proprietary internal processes, a risk extensively covered in discussions on prompt injection attack techniques.

What You Should Do

  • AI providers should implement robust identity verification measures and continuously monitor for unusual subscription-to-usage ratios.
  • Apply stringent rate limits to API access and maintain comprehensive logs of all requests for detailed investigation.
  • Share infrastructure and behavioral threat intelligence with cloud platforms and API aggregators to detect distributed campaigns that might not be visible from a single service.
  • For high-confidence malicious requests, consider implementing targeted response changes such as reducing response fidelity or varying outputs subtly, without alerting the suspected operators.
  • Enhance defenses with differential privacy, adversarial testing, stricter API controls, and mechanisms to limit prompt injection to fortify against extraction attempts.

Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.

Tags:

AttackCybersecurityMalwareSecurity

Share Article

Marcus Rodriguez

Marcus Rodriguez

Marcus is a security researcher and investigative journalist with expertise in vulnerability research, bug bounties, and cloud security. Since 2017, Marcus has been breaking stories on critical vulnerabilities affecting major platforms. His investigative work has led to the disclosure of numerous security flaws and improved defenses across the industry. Marcus is an active participant in bug bounty programs and has been recognized for responsible disclosure practices. He holds multiple security certifications and regularly speaks at industry events.

Previous Post

The Best Mobile Device Management (MDM) Solutions for 2026

Next Post

Critical cPanel CVE-2024-XXXX Vulnerability Lets Attackers Gain Full Server Control

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Popular Posts
Critical Redis Vulnerability Exploited in Widespread Cryptomining Attacks
September 9, 2026
10 Best Mobile Threat Defense Solutions for 2026
September 9, 2026
Best Patch Management Software for 2026
September 9, 2026
Top Authors
Marcus Rodriguez
Marcus Rodriguez
David kimber
David kimber
Jennifer sherman
Jennifer sherman
Let's Connect
156k
2.25m
285k

Related Posts

Jennifer sherman
By Jennifer sherman
Threats

GlassWorm Attacks macOS via Malicious VS Code…

January 1, 2026
Emy Elsamnoudy
By Emy Elsamnoudy
Attacks

ClickFix Attack Hides Malicious Code via Stegan Security

January 1, 2026
Sarah simpson
By Sarah simpson
Vulnerabilities

MongoBleed Detector Tool Released to Detect MongoDB Vulnerability(CVE-2025-14847)

January 1, 2026
Emy Elsamnoudy
By Emy Elsamnoudy
Breaches

Conti Ransomware Gang Leaders & Infrastructure Exposed

January 1, 2026
Hackers News Hackers News
  • [email protected]

Quick Links

  • Contact Us
  • Privacy Policy
  • Terms of service

Categories

Attacks
Breaches
Comparisons
CyberSecurity News
Threats
Vulnerabilities

Let's keep in touch

receive fresh updates and breaking cyber news every day and week!

All Rights Reserved by HackersRadar ©2026

Follow Us