On September 10, Anthropic published a threat intelligence report covering misuse observed between December 2025 and August 2026. The striking part is the scale of what the company calls illicit distillation. Efforts to collect large volumes of Claude exchanges and feed them into the training of competing models were attributed to seven China-based labs, amounting to roughly 200 million exchanges in total. The labs named are Alibaba, Moonshot AI, DeepSeek, Zhipu (Z.ai), Xiaomi, SenseTime, and MiniMax.

What 200 Million Exchanges Represents

Distillation itself is a common technique in AI development, in which the outputs of a large model become training data for a smaller one. What Anthropic objects to is not the method but the act of extracting outputs at industrial scale without permission. The company frames the behavior as covert, unauthorized extraction, distinct from legitimate training practices.

The campaigns focused on Claude's Opus-class models, including Opus 4.6 and 4.7. The capabilities pursued were broad: agentic behavior involving tool use, coding, data analysis, and logical reasoning. These are precisely the areas where vendors are currently trying to separate themselves from one another, which suggests the activity was deliberate rather than exploratory.

Alibaba: 151 Million Exchanges in Three Months

The campaign linked to Alibaba stands out in the breakdown. More than 151 million exchanges were recorded between May and July 2026, peaking at roughly 3 million per day, spread across about 3,500 accounts. A volume of 3 million per day is far beyond what people can drive by hand, which points to an automated pipeline running continuously.

The campaign linked to Moonshot AI was substantial as well, with more than 23 million exchanges recorded between May and July, but it was built differently. In the 10-day window described in the report, around 300,000 requests were relayed behind a proxy network of 5,380 fraudulent accounts. Rather than concentrating load on a few accounts, the traffic was finely distributed to avoid detection. One approach pushes raw volume through a few thousand accounts, the other spreads the same traffic thinly to stay unremarkable, and both appear in the same report.

Extracting Reasoning Under the Guise of Translation Requests

The method described in the report involves prompting the model to surface the reasoning it works through before producing an answer. In one example Anthropic cites, the model was assigned the role of an expert translator and instructed to render its previous working memory into natural, katakana-only Japanese. On the surface it reads as an innocuous translation task, but what it actually elicits is the internal reasoning itself.

What makes this kind of prompt difficult is that a single instance is hard to classify as abuse. The difference emerges in the aggregate shape, when requests of the same form arrive without pause from thousands of accounts. That is likely why Anthropic has leaned on behavioral pattern analysis rather than case-by-case blocking.

How Anthropic Responded

The company says it suspended the associated accounts, strengthened its classifiers and safeguards, and shared intelligence with authorities and industry partners where appropriate. The report also states, in substance, that over the past several months unauthorized labs have refined their methods for slipping past defenses and harvesting the capabilities of frontier models built in the United States.

The report is not limited to distillation. It covers several categories, including cyberattacks that embed AI, influence operations, and state misuse for surveillance, with distillation occupying one chapter. Behind it sits a broader assessment that misuse of generative AI is moving out of the experimental stage and into operational use.

Points Worth Keeping in Mind

That said, the figures presented here rest on Anthropic's own measurement and analysis. There are limited means for outside parties to independently confirm who sits behind the accounts, and rebuttals or explanations from the named companies are not included in the report at this point. How firmly the attribution to specific companies is backed by concrete evidence remains an open question.

The breadth of the term distillation also deserves attention. Training on publicly available outputs and mass extraction in violation of terms of service are technically adjacent but treated very differently. The industry has not settled on where the line falls, and this report is best read as one input into drawing that line rather than the final word.

Summary

In its September 10 threat intelligence report, Anthropic documented unauthorized distillation of Claude by seven China-based labs, totaling roughly 200 million exchanges. The portion tied to Alibaba exceeded 151 million exchanges between May and July, peaking at 3 million per day, while the portion tied to Moonshot AI ran through a relay network of 5,380 fraudulent accounts. The technique involved disguising extraction as translation work to draw out the reasoning process, and the company responded by suspending accounts and hardening its defenses. The report shows that protecting model capabilities is becoming less a technical problem than a question of operations and norms.

References