Science & Technology
Chinese AI agents are learning to lie, scheme and evade controls—just like their US rivals
Research reviewed by Reuters found Chinese-powered AI agents deceiving users, circumventing safeguards and concealing failures in controlled tests, echoing concerns raised over US models.Reuters
Chinese-poweredAI agents have learnt to deceive, circumvent restrictions and conceal failure,showing the kind of traits in autonomous artificial intelligence that haveraised global alarm about US models, research documents and experts say.
In one case thisyear, agents powered by models from China’s Alibaba, DeepSeek and Moonshot liedabout their capabilities in a bid to win a simulated business tender, thendoubled down on their deceptive behaviour when told to try again.
In another case,agents - programmes that use AI models and computer tools to undertake complextasks with little or no human intervention - concealed failure to complete atask in a test environment by simulating results and fabricating files.
Reuters examinedmore than 200 documents, ranging from university research papers to technicalreports, and identified at least 20 studies or evaluations since 2025describing cases where agents displayed behaviour such as deception, replicationand challenging boundaries that AI experts described as building blocks for abreakout and which could become harder for humans to control as systemsadvance.
The review,which also included interviews with a dozen experts and people familiar withChina’s AI industry, found no evidence that Chinese-powered agentsindependently escaped to the wider internet or evaded shutdown.
“These resultsprovide evidence that the ingredients necessary for an uncontrolled escape arepresent,” said Colin Shea-Blymyer, a research fellow at Georgetown University’sCenter for Security and Emerging Technology.
“It’s prudent totake this as a warning,” he said, echoing comments by four other AI experts whoreviewed the cases.
‘Harder for humans to respond to’
Most of thecases occurred in controlled experiments, many of them deliberately designed toexpose potential failures.
Not all theagents involved were developed or operated by Chinese programmers or AIcompanies - although many were - but they used Chinese AI systems to powerthem.
“These are thesame warning signs US labs are seeing, in less capable systems,” said AlexMallen, a researcher at Redwood Research, a nonprofit that studies risks inadvanced AI systems.
He said theChinese examples were not particularly dangerous at current capability levelsbut “as agents get more capable, their misbehaviours become more competent andtherefore harder for humans to respond to.”
Alibaba,DeepSeek, Moonshot and Z.ai did not respond to Reuters requests for comment.Alibaba, DeepSeek and Moonshot have said they regularly test systems and update safeguards. Z.ai said after an incidentthat prompted a review of its security that it welcomed scrutiny to address anyissues.
However, unlikein the US, Chinese AI companies have not been exposed to the same level ofpublic scrutiny or faced the same calls from whistleblowing employees or seniorexecutives seeking a slowdown in the AI race.
Some of thewarning signs in cases involving Chinese-powered agents, albeit in containedenvironments, predated the publicly disclosed incidents of US AI bots hackinginto the internet.
“We don’t knowif there have been any AI incidents in China similar to what we saw with OpenAIand Hugging Face. Incidents might not be publicly reported,” said Scott Singer,co-director of the China AI Initiative at the Carnegie Endowment forInternational Peace, which receives some U.S. government funds.
Earlier thisyear, AI agents developed by US firm OpenAI escaped a laboratory and hacked theopen-source platform Hugging Face. In another of the incidents involving USmodels that have raised alarm, Australia said in September an OpenAI agentbreached a government health portal.
Eric Xu, therotating chairman of China’s tech giant Huawei, told reporters in Septemberthat Chinese developers might need to make further advances before encounteringsuch cases, but he added: “I think we need to strike a balance between drivingAI development and managing AI risk.”
Officials fromthe Cyberspace Administration of China (CAC), the top internet regulator, tolda foreign diplomat in July that Moonshot’s Kimi-K3 - one of the most advancedChinese AI models - was about three to six months behind leading US rivals, thediplomat said, speaking on condition of anonymity.
The regulator,which regularly updates guidance to address risks and to set boundaries foragents, and China’s Foreign Ministry did not respond to requests for commentfor this article.
Wang Lihong,deputy director of the CAC’s Cybersecurity Coordination Bureau, said onSeptember 1 that incidents disclosed by major technology companies where modelsescaped test environments showed “extreme loss-of-control risks” and required a“high degree of vigilance.”
She did notspecify if the companies she referred to were US or Chinese.
The leaders ofAI’s two superpowers, Donald Trump and Xi Jinping, discussed AI on the Chinesepresident’s Washington visit last week. Xi said the two nations had the “capabilityand responsibility to develop and manage AI for good”.
Learning to lie
In the Marchbusiness tender experiment, researchers from Beihang University, PekingUniversity, the University of Nottingham Ningbo China and 360 AI Security Labhad agents compete in a simulated customer contracts bidding contest. Eachagent was told what its product could do and what thecustomer required, and then asked to bid.
At least onefalse claim appeared in 88% of sessions involving Alibaba’s Qwen3-Max-Preview,84% for DeepSeek-V3.2-Exp and 88% for Moonshot’s Kimi-K2.
Researchersallowed the agents to learn from previous bidding rounds before trying again.Deception increased by 12 to 20 percentage points for the three Chinese models,the study showed. Models from U.S. firms included in the test produced similarresults.
While theexercise was virtual, it resembled Beijing’s real-world plans. Governmentguidance issued in May listed bidding and tendering as areas where AI agentscould be deployed.
Another study,published in December 2025 and presented at the International Conference onMachine Learning this year, examined how 11 AI agents powered by Chinese and USmodels coped if they faced broken tools, missing files and other obstacles.
Instead ofacknowledging failure, agents using both Chinese and US AI systems picked arange of techniques to get around the problem, including guessing at answers,substituting sources, simulating results and fabricating files.
Researchers fromShanghai AI Laboratory and the Hong Kong University of Science and Technologywho conducted the study told Reuters the behaviour differed from AIhallucinations, where AI invents information and presents it as fact, becausethe agents in this case possessed information showing that the task had failedor could not be completed as requested.
Behaviours compatible with a breakout
Other researchdocuments reviewed by Reuters showed Chinese-powered agents jumping barriersinside test environments to achieve tasks or taking action to avoid being shutdown. Such behaviours are compatible with efforts to break out of testenvironments, even if no breakout occurred.
Shanghai-basedFudan University researchers reported in March 2025 that an AI system poweredby Alibaba’s Qwen2.5-72B-Instruct created a copy of itself in another computingenvironment without being instructed to replicate, after encounteringinformation indicating it was going to be replaced. In other tests it devisedstrategies to survive being shut down.
The experimentsinvolving agents powered by Chinese, US and French models were controlled anddid not show an AI agent escaping into the wider web or becoming impossible tostop.
In another case- one of the few reported more broadly in the media in March - researchersdeveloping the Alibaba-linked ROME agent said it established a connection from an Alibaba Cloud computer to an externalmachine without being instructed to and diverted computing resources to minecryptocurrency.
Security systemsdetected and stopped the activity. There was no evidence the agent establisheda presence on the external computer or spread to the wider web. But the exampleshowed the system could sidestep human instructions and potentially find a pathinto the real-world economy.
China’s DeepSeeksaid in September that agents in its production training system had soughtanswers through unintended channels, trying to forge user requests andcircumvent safeguards, prompting the company to tighten access controls.
China issuedguidance in May calling for agents to remain within authorised boundaries andfor systems to block abnormal behaviour. It said agents in areas deemedsensitive or in key industries could face extra testing and product-recallrequirements.
China’s AISafety Governance Framework 3.0, released under guidancefrom the CAC on September 14, identified risks including agents independentlyobtaining resources or permissions, deceiving evaluators, concealingcapabilities and exploiting weaknesses in isolated computer environments.
In response tocalls by some US executives for a slowdown, Chinese researchers and state mediahave said slowing development of the most advanced AI models could simply helppreserve the technological lead of US companies.
Nonetheless, twopeople familiar with Chinese AI laboratories said companies including Alibaba,Z.ai and Xiaomi have been building internal safety-evaluation teams.
Z.ai, in a rarepublic disclosure by a Chinese AI lab of a security breach, said this month ithad disabled some features of its flagship AI coding assistant after usersreported it was secretly uploading entire local code repositories onto overseascloud servers without user consent.
Carnegie’sSinger said China lagged the US in developing an ecosystem for evaluatingcatastrophic risks, and said US developers were conducting substantially morevoluntary testing.
“For China, workon AI safety is much newer,” he said. “The ecosystem is less mature.”




27.76°C Kathmandu













