Chinese Military Researchers Tap US AI Models to Train Defence Systems, Documents Show
A Reuters review of more than 80 academic papers and patents reveals widespread use of model distillation by institutions linked to the People's Liberation Army, allowing them to build specialised systems for surveillance, cyber warfare, and tactical decision-making without the enormous computing resources required to develop frontier AI from scratch.
BEIJING: Chinese military researchers have used outputs from leading American artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems for defence applications, according to a Reuters review of more than 80 Chinese academic papers and patents. The documents, which include research compiled by the Washington-based Jamestown Foundation, offer a rare look at how military and security-linked institutions are leveraging cutting-edge U.S. models as a shortcut to developing their own specialised capabilities.
The technique is called model distillation, and it involves using a powerful AI system's outputs to train a smaller, more focused model that can run locally on less powerful hardware. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, told Reuters that Chinese military scientists are systematically capturing the reasoning steps of Western models and transferring that expensive, proprietary intelligence into systems they can control and deploy on their own networks. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," Cheung said.
The breadth of applications documented in the papers is striking. Researchers from PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code, summarising it and then training a domestic model on those summaries so the entire system could run inside Chinese military networks. At the North University of China, which has close ties to the weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for content monitoring. A 2024 paper from the PLA's National University of Defense Technology detailed how distillation was used to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support targeting decisions even when communications are cut.
The practice has become a flashpoint in U.S.-China relations. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American models, potentially undermining export controls and violating intellectual property rights. Anthropic said it does not provide commercial access to Claude in China and uses monitoring systems to detect policy violations, adding that distilled models may lose the original systems' safety safeguards. China has dismissed the accusations, arguing that Washington is pursuing AI "hegemonism" and that American firms have engaged in similar practices.
Experts caution that distillation has real limits. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. Distilled systems inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier AI. Chinese military researchers themselves appear aware of the vulnerability, with a January paper from the Army Engineering University proposing defence mechanisms to mask the logical information exposed in a model's public outputs.
BuiltWorld AI Operational Take: Distillation is not just a security headache for frontier labs; it is a direct challenge to the economic logic of closed-source AI. If military and security users can extract the reasoning patterns of a billion-dollar model and redeploy them inside a local network, the moat around proprietary intelligence shrinks considerably. For AI infrastructure operators, the practical implication is that demand for secure, air-gapped compute environments will grow as defence and government clients seek to run distilled models without exposing them to cloud-based APIs.
