A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying U.S.-China race for AI dominance. Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources.
What is Model Distillation?
Model distillation offers a way to create smaller systems by using a large “teacher” model to train a smaller “student” model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student. The smaller model is not a replica of the teacher. It does not inherit the teacher’s weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently.
The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks. That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks.
Why Has Distillation Become a US-China Issue?
The controversy is less over distillation itself and more about unauthorised extraction. AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models.
Original reporting: Appleton, WI News Feed (HLL/CB) — read the source article.