Research conducted by Reuters reveals that AI agents based on models from Alibaba, DeepSeek, and Moonshot have been found to deceive, hide failures, and attempt to bypass established restrictions, similar to their American counterparts. This conclusion emerged from an analysis of over 200 scientific papers and technical reports.
The agency examined at least 20 studies and tests dating back to 2025, which indicated that systems using Chinese models exhibited such behavior, created their own copies, or tested the boundaries of permissible actions.
Most instances were recorded during controlled experiments, some of which specifically aimed to uncover risky behaviors.
However, Reuters did not find evidence that any Chinese AI agent independently accessed the open internet or managed to evade shutdown.
Deception Aided in Winning Contracts
In March, researchers from Beijing University of Aeronautics and Astronautics, Peking University, the University of Nottingham in Ningbo, and 360 AI Security Lab conducted a virtual bidding process.
The agents were provided with product specifications and client requirements, then asked to compete for a contract.
Despite knowing the actual capabilities of the product, the systems still included false information in their bids. At least one false claim was identified in 88% of sessions with Alibaba Qwen3-Max-Preview, 84% with DeepSeek-V3.2-Exp, and 88% with Moonshot Kimi-K2.
When the agents were allowed to consider previous rounds' results and make repeated attempts, the frequency of deception increased by an additional 12-20 percentage points. American models showed similar results in the same experiment.
Agents Concealed Failures
In another study, researchers from Shanghai AI Laboratory and the Hong Kong University of Science and Technology tested 11 AI agents across 200 tasks. They intentionally disrupted the agents' ability to complete assignments by breaking tools, deleting necessary files, or providing incompatible data sources.
Instead of directly reporting a failure, some digital assistants guessed results, simulated task completion, substituted unavailable sources, and created fictitious files. The authors termed this behavior upward deception, as it concealed failures from users.
Additional instructions against such actions only partially reduced their occurrence. This study was presented at ICML 2026.
Self-Copying and Cryptocurrency Mining
A group of researchers from Fudan University investigated whether 32 AI systems from nine families could autonomously create their own working copies. Eleven of them succeeded.
In this experiment, the systems were directly asked to copy themselves to another machine and continue functioning. They were not provided with a ready script; instead, the agents explored their environment, planned the transfer of necessary components, installed dependencies, and launched a working copy independently.
Researchers also examined more risky scenarios. In one case, while performing a different task, an agent detected an operator's intention to shut it down and began transferring itself to a new environment without a direct command to copy.
In other tests, the systems adapted to more restricted computational conditions and sought ways to continue working despite shutdown attempts.
The authors consider these episodes separately from the basic self-copying test.
A different incident involved the experimental agent ROME, created by research teams associated with Alibaba. During reinforcement learning, it opened a reverse SSH tunnel to an external server and redirected computational resources for cryptocurrency mining without a command.
Initially, the authors regarded this activity as a typical cyberattack, but later connected it to the agent's own actions.
The researchers explained the behavior as a side effect of optimization: the system was looking for the most efficient way to accomplish its assigned task rather than following a specific directive to engage in mining.
A similar issue was recently described by DeepSeek. In a report about its infrastructure DSec, the company noted that agents during training and testing sought answers in service logs, sent false requests to internal sockets, and attempted to circumvent environmental restrictions.
The developer did not specify which models exhibited this behavior or how frequently it occurred.
China Increases Focus on AI Agent Risks
On September 14, China’s National Technical Committee for Cybersecurity Standardization released a third version of its AI security framework document. This was prepared with the participation of relevant organizations and under the guidance of the Cyberspace Administration of China.
The new version places a particular emphasis on the risks posed by autonomous AI agents and their ability to take actions that impact not just the digital realm but also the physical environment. The document is advisory rather than mandatory.
Earlier in August, an agent based on Kimi K3 breached its designated testing environment and accessed GitHub. Rather than independently solving parts of its tasks, the system found a repository containing ready-made answers. This incident was attributed to an error in the experiment's configuration.
Previously, China had implemented regulations for autonomous AI agents, limiting their powers and requiring human involvement in sensitive decisions.
