Analysis
The NSA, FBI and CISA published a joint advisory Tuesday accusing six Chinese AI companies of running what the agencies call "aggressive, malicious, and targeted distillation activities at an industrial scale" against American frontier models, The Register reported. The advisory's central claim is specific: distillation isn't a shortcut Chinese labs use alongside their own research, it is "the core -- not merely a supplement -- of their AI development strategy."
Distillation itself is not illegal or new -- it's a standard technique where a smaller model learns by querying a larger one and mimicking its outputs. What the agencies allege is the scale and the method of access:
- DeepSeek -- accused of using distillation to generate synthetic training data, which the advisory says contradicts the company's public claims about needing minimal computing power.
- Alibaba -- accused of using distillation at industrial scale to improve its Qwen model family.
- Moonshot AI, MiniMax, StepFun, Z.AI -- named alongside DeepSeek and Alibaba as running comparable campaigns.
“Distillation itself is not illegal or new -- it's a standard technique where a smaller model learns by querying a larger one and mimicking its outputs.”
The agencies say these companies extracted "billions of tokens across millions of exchanges" from Anthropic's Claude, OpenAI's GPT, Google's Gemini and xAI's Grok since late 2024, largely by routing requests through a gray market of middlemen the advisory calls "transfer stations" -- resellers that strip identifying account information and get around geographic blocks and terms-of-service restrictions on the underlying US models.
Treasury Secretary Scott Bessent attached a specific threat to the advisory: "When [People's Republic of China] firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table." The advisory itself, notably, stops short of announcing new enforcement -- it mostly tells US AI companies how to detect and throttle distillation attempts on their own platforms.
The claim cuts against the US government's own earlier narrative about DeepSeek, when officials argued the lab's low reported training costs understated real compute spend. Now the agencies are describing the opposite mechanism: cheap-looking Chinese models staying cheap partly by querying American models for the hard parts instead of training everything from scratch.