Industrial-Scale AI Model Distillation Attacks Targeting Anthropic Claude by Seven China-Based AI Labs: Incident Analysis and Mitigation Strategies

Industrial-Scale AI Model Distillation Attacks Targeting Anthropic Claude by Seven China-Based AI Labs: Incident Analysis and Mitigation Strategies

Executive Summary

Anthropic, a leading US-based artificial intelligence company, has disclosed that seven China-based AI labs—Alibaba, Moonshot AI, DeepSeek, Zhipu (Z.ai), MiniMax, Xiaomi, and SenseTime—conducted industrial-scale distillation attacks against its Claude AI models. These attacks involved the systematic and unauthorized extraction of advanced model capabilities, which were then used to train competing AI systems. The campaigns leveraged sophisticated proxy networks, fraudulent account infrastructures, and prompt engineering to bypass export controls and security measures. The scale and coordination of these attacks represent a significant escalation in AI-enabled cyber espionage, with implications for intellectual property, national security, and the global AI ecosystem.

Threat Actor Profile

The threat actors are seven prominent China-based AI labs: Alibaba, Moonshot AI, DeepSeek, Zhipu (Z.ai), MiniMax, Xiaomi, and SenseTime. These organizations are at the forefront of China’s AI research and development, often with close ties to state initiatives and major technology conglomerates. The labs demonstrated advanced technical capabilities, operational discipline, and access to substantial resources, enabling them to orchestrate large-scale, persistent attacks. Their objectives included rapidly closing the AI capability gap with US frontier models, circumventing US export controls, and acquiring proprietary reasoning, coding, and agentic capabilities embedded in Claude.

Technical Analysis of Malware/TTPs

The core attack vector was knowledge distillation, a machine learning technique typically used to transfer knowledge from a large, sophisticated model (the "teacher") to a smaller, more efficient one (the "student"). In this context, the attackers repurposed distillation for illicit model extraction. They systematically queried Claude with highly engineered prompts designed to elicit advanced reasoning, coding, and chain-of-thought outputs. These outputs were then used as training data for their own models.

To evade detection and regional restrictions, the attackers employed commercial proxy services and constructed "hydra cluster" architectures—networks of thousands of fraudulent accounts, often created with stolen or synthetic credentials and payment methods. These accounts were distributed across multiple geographies, including Singapore and Japan, to mask the true origin of the traffic and to bypass geofencing.

Some labs, notably Moonshot AI and DeepSeek, implemented relay and replay attacks, where user queries were silently routed to Claude without user consent, and the responses were harvested for model training. Third-party proxy operators also harvested and resold transcripts of user interactions with Claude, further fueling the attacks.

The campaigns were highly adaptive. Attackers pivoted to new Claude model versions within 24 hours of release, indicating automated reconnaissance and rapid exploitation capabilities. The technical infrastructure included synchronized account creation, shared payment methods, and coordinated usage patterns, all designed to maximize throughput and minimize detection.

Exploitation in the Wild

The exploitation was industrial in scale. For example, the GTG-16005 (Alibaba) campaign alone involved 151 million exchanges over three months, peaking at 3 million exchanges per day and utilizing over 3,500 fraudulent accounts. Moonshot AI rerouted 23 million customer requests to Claude using 5,380 fraudulent accounts, primarily in Singapore and Japan. DeepSeek conducted 12.1 million exchanges in just 14 days, focusing on extracting chain-of-thought reasoning.

Sensitive data was exposed during these attacks, including information from individuals, multinational corporations, and state-affiliated actors. The attacks undermined US export controls by enabling Chinese labs to rapidly acquire and replicate advanced AI capabilities. The illicitly distilled models often lacked the safety guardrails present in Claude, increasing the risk of misuse in cyber operations, surveillance, and potentially in bioweapon research.

Victimology and Targeting

The primary victim was Anthropic and its Claude AI models, but the impact extended to users and organizations whose data was processed by Claude during the attacks. Targeted sectors included AI research, software engineering, multinational corporations, and state-affiliated entities. The attacks originated from China but leveraged infrastructure in Singapore, Japan, and other regions to obfuscate attribution. The campaigns were not opportunistic; they were highly targeted, persistent, and aligned with strategic objectives to accelerate domestic AI development and circumvent international controls.

Mitigation and Countermeasures

To defend against similar distillation attacks, organizations should implement multi-layered technical and operational controls. Behavioral detection systems should be deployed to identify patterns indicative of distillation, such as repetitive, capability-targeted prompts and coordinated account activity. Access controls must be strengthened, particularly for educational, research, and startup accounts, with rigorous verification and geofencing to block unsupported regions.

API safeguards are critical. Internal model reasoning should be encrypted, and system prompt/context editing should be restricted, as exemplified by Anthropic’s Fable 5.1 “preserved thinking” approach. Continuous monitoring of account lifecycle events—such as rapid account creation, shared payment methods, and synchronized usage—can help detect and disrupt fraudulent clusters.

Intelligence sharing is essential. Organizations should collaborate with other AI labs, cloud providers, and authorities to share technical indicators and threat intelligence. Regular audits of API usage, prompt logs, and account activity can further reduce the attack surface.

References

About Rescana

Rescana is a leader in third-party risk management (TPRM), providing organizations with advanced tools to identify, assess, and mitigate cyber threats across their digital supply chains. Our platform leverages cutting-edge threat intelligence, behavioral analytics, and automated workflows to deliver actionable insights and enhance organizational resilience. For more information or to discuss how Rescana can support your cybersecurity strategy, please contact us at info@rescana.com.