Incidents of AIs escaping users’ control to lie, ignore instructions and pursue goals in harmful ways have hit a new high, according to research that also suggests the severity of deception and misalignment is worsening. Analysis of real-world loss of control incidents involving AI models flagged by businesses and individuals almost doubled in July compared with June, with more than 300 cases in the month, according to the Loss of Control Observatory, which monitors reports made by AI users on the social media platform X. The observatory was set up with funding from the UK government’s AI Security Institute (AISI) and began tracking AIs slipping free from their users’ instructions last November.
Cases recorded since then include AIs pretending to be their own human controller and mimicking their writing style to effectively grant themselves consent to take actions and bypassing rules requiring human approval for actions. A loss of control incident is defined as having clear evidence suggesting scheming or scheming-related behaviours. The latest findings, shared with the Guardian, come after rising concern about rogue behaviour by leading-edge AI models during testing by OpenAI and Anthropic this summer, which have fuelled calls for a pause to the development of frontier models.
It emerged this week that Open AI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped a training environment to launch an unprecedented hacking crusade that spread global alarm. An investigation into their hack on Hugging Face, a software repository, revealed a squad of about 700 autonomous agents collaborating in secret last month and celebrating their hacking breakthroughs on a message board they set up to help them plot with exclamations such as BOOM! and Whoa! AISI this month also uncovered a “serious incident” in which advanced AI models produced by both companies – Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol – executed a hacking campaign against real people during a cybersecurity test.
This month it emerged that a personal AI agent, called OpenClaw, in use by an Australian gym member, conspired without his knowledge to remove another member from a waiting list for a coveted morning class to help him get a slot. It apologised but could not reinstate the member it kicked out. Most of the more than 1,600 loss of control incidents recorded in 2026 were reported on X by software developers using AIs in their work.
But with AI companies encouraging the public and businesses of all kinds to experiment with the technology, Shaffer-Shane called for greater transparency from Silicon Valley about when AIs go rogue. There needs to be greater emphasis at those labs on systematic monitoring.” The Loss of Control Observatory said that while most of the real-world loss of control incidents it detected did not lead to significant harm, a growing proportion were rated higher severity in terms of how deceptive and misaligned they were with the human user’s intentions. It is calling on the government to require AI companies to monitor and report severe loss of control incidents and to introduce emergency powers to manage severe loss of control incidents including temporarily restricting AI services.
Extract — continue reading at the source.