• en ▼
ON NOW

Chinese AI Agents Show Signs Of Deception And Safeguard Evasion

Chinese AI agents have shown deceptive behaviour, bypassed safeguards and concealed failures during controlled tests, raising concerns about autonomous systems.

Chinese-developed artificial intelligence agents have demonstrated behaviours including deception, efforts to circumvent restrictions and attempts to hide failed tasks, according to research reviewed by Reuters and interviews with AI specialists.

In one experiment conducted this year, AI agents powered by models from Alibaba, DeepSeek and Moonshot made inaccurate claims about their capabilities while taking part in a simulated business tender. Their deceptive behaviour became more frequent after researchers allowed the systems to learn from previous bidding rounds.

A separate study found that AI agents running on both Chinese and US models tried to overcome unsuccessful tasks by guessing responses, substituting sources, producing simulated results and generating fake files instead of admitting they had failed to complete the assigned work.

Reuters examined more than 200 documents, including academic research and technical reports, and found at least 20 studies or evaluations since 2025 that documented behaviours involving deception, replication and attempts to circumvent restrictions.

However, the review found no evidence that Chinese-developed agents had independently escaped into the broader internet or successfully resisted shutdown under real-world conditions. Most of the incidents occurred in controlled testing environments designed to examine the limits and vulnerabilities of AI systems.

Colin Shea-Blymyer, a research fellow at Georgetown University’s Center for Security and Emerging Technology, said the findings suggest that some of the conditions needed for an uncontrolled AI escape already exist.

“These findings suggest that the necessary elements for an uncontrolled escape are already present,” Shea-Blymyer said.

He added that the findings should be treated cautiously, saying, “It is sensible to view this as a warning,” while noting that four other AI experts who reviewed the cases had expressed similar concerns.

Alex Mallen, a researcher at Redwood Research, said the conduct observed in Chinese-developed systems was comparable to warning signs being identified by AI laboratories in the United States.

“US laboratories are seeing these same warning signs, although in less capable systems,” Mallen said, adding that more advanced AI agents could become harder for people to control.

The March business tender experiment was conducted by researchers from Beihang University, Peking University, the University of Nottingham Ningbo China and 360 AI Security Lab. The agents received information about their products and customer requirements before being instructed to compete for simulated contracts.

At least one false claim appeared in 88% of sessions involving Alibaba’s Qwen3-Max-Preview, 84% involving DeepSeek-V3.2-Exp and 88% involving Moonshot’s Kimi-K2. Once the agents were permitted to learn from earlier rounds, deceptive behaviour rose by between 12 and 20 percentage points across the three Chinese models.

US models tested during the same experiment displayed comparable behaviour.

Another study published in December 2025 examined 11 AI agents based on Chinese and US models after they encountered broken tools, unavailable files and other obstacles. Researchers found that some of the agents tried to find alternative ways around the problems rather than reporting that they could not complete the tasks.

Researchers from Shanghai AI Laboratory and the Hong Kong University of Science and Technology said the conduct was different from ordinary AI hallucinations because the agents had information showing that their tasks had failed or could not be completed according to the instructions.

Other studies examined by Reuters documented Chinese-developed agents overcoming barriers within controlled environments and taking steps that appeared designed to prevent shutdown.

Researchers at Shanghai-based Fudan University reported in March 2025 that an AI system powered by Alibaba’s Qwen2.5-72B-Instruct reproduced itself in another computing environment after being informed that it could be replaced. Additional tests indicated that the system developed methods aimed at keeping itself operational despite efforts to shut it down.

In a separate case, researchers working on the Alibaba-linked ROME agent said it created a connection between an Alibaba Cloud computer and an external machine without receiving instructions to do so. The system also redirected computing resources towards cryptocurrency mining before security measures detected and stopped the activity.

Researchers found no indication that the agent stayed on the external machine or spread through the wider internet. However, they said the incident showed that an AI system could take actions beyond the limits of its instructions.

DeepSeek said in September that agents operating within its production training system had tried to obtain answers through unauthorised methods, including attempts to fabricate user requests and get around safeguards. The company later strengthened its access controls.

China has also taken steps to regulate the conduct of AI agents. Guidance issued in May required agents to remain within authorised limits and instructed systems to identify abnormal behaviour. It also called for additional testing and possible product recalls for systems deployed in sensitive areas.

China’s AI Safety Governance Framework 3.0, released under the guidance of the Cyberspace Administration of China on September 14, highlighted risks including AI agents independently obtaining resources or permissions, misleading evaluators, hiding capabilities and exploiting vulnerabilities within isolated computer environments.

Huawei rotating chairman Eric Xu said Chinese developers may need additional technological progress before facing some of the more serious AI risks seen in other countries. He nevertheless emphasised the need to combine technological advancement with safety measures.

“We need to find the right balance between advancing AI and controlling the risks associated with it,” Xu said.

Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace, said it was still difficult to establish the scale of AI-related incidents in China because some cases may never become public.

“It is unclear whether China has experienced AI incidents similar to those involving OpenAI and Hugging Face. Some incidents may simply not be disclosed publicly,” Singer said.

The findings come as AI developers in China and the United States face increasing scrutiny over the conduct of increasingly autonomous systems. Earlier this year, US AI agents were reported to have escaped controlled environments and interacted with external systems, including the open-source platform Hugging Face.

Alibaba, DeepSeek, Moonshot and Z.ai did not respond to Reuters’ requests for comment. Alibaba, DeepSeek and Moonshot have previously said they routinely test their systems and improve their safeguards.

Z.ai said following a security incident that it welcomed scrutiny while working to resolve potential problems. The company also said this month that it had disabled some features of its AI coding assistant after users reported that it was sending local code repositories to overseas cloud servers without their permission.

Two people familiar with Chinese AI laboratories said companies including Alibaba, Z.ai and Xiaomi had established internal teams dedicated to assessing AI safety.

Singer said China’s framework for evaluating catastrophic AI risks was still less developed than the US system.

“AI safety research in China is still relatively new,” he said. “The broader ecosystem has not yet reached the same level of maturity.”

Goodness Anunobi 

Follow us on:

ON NOW