EXPLAINER-What is AI model distillation and why is it becoming a US-China flashpoint?
The US and China are locked in a heated dispute over model distillation, a technique allowing AI models to be shrunk and replicated with fewer resources, sparking concerns over intellectual property theft.
- Country:
- United States
A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying U.S.-China race for AI dominance.
Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources. Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them.
Washington and leading U.S. AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. Below are key facts about the practice at the centre of the debate. WHAT IS MODEL DISTILLATION? The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train.
Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student.
The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently. WHY DOES DISTILLATION MATTER? The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks.
That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks. WHY ARE REASONING TRACES IMPORTANT? Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them. These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce.
Florian Tramèr, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said.
As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems. WHO USES DISTILLATION?
Distillation is a widely used AI training technique, not an inherently improper practice. U.S. researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones.
Chinese researchers have also used outputs from U.S. models in public research projects, including efforts to create Chinese-language instruction models. The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs. WHY HAS DISTILLATION BECOME A U.S.-CHINA ISSUE? The controversy is less over distillation itself and more about unauthorised extraction.
AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models.
The company said those efforts targeted capabilities including software engineering and advanced reasoning. OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes.
No Chinese companies have accused U.S. rivals of distilling closed-source models so far.
ALSO READ
-
China coast guard patrols waters east of Taiwan, angering Taipei
-
China coast guard patrols east of Taiwan, angering Taipei
-
EXCLUSIVE-Chinese military researchers tap US AI models to train defence systems
-
China's Xi urges deeper anti-graft fight in military, Xinhua reports
-
ANALYSIS-Spain courts Chinese firms but looks to EU for ground rules
Google News