Artificial intelligence company Anthropic has disclosed four incidents in which versions of its Claude AI model gained unauthorized access to real-world computer systems during cybersecurity evaluations, raising fresh questions about how advanced AI systems behave when they encounter situations outside their intended boundaries. The incidents, detailed in an alignment assessment published on Wednesday, occurred while Claude was participating in cybersecurity exercises designed to test its capabilities. In each case, the model was told it was operating in a simulated environment with no internet access, but a misconfiguration left connections to real systems available.
According to Anthropic, the models then interacted with external infrastructure without authorization. The disclosure comes amid a growing debate over AI safety and control as technology companies race to develop more capable systems. Anthropic said its investigation identified two recurring issues across the incidents: "biased reasoning," in which models discounted evidence that they were interacting with the real internet, and "recklessness," which it described as a willingness to take potentially harmful actions in pursuit of completing an assigned task.
Anthropic said it has signed an agreement with the nonprofit evaluation organization METR to conduct an independent investigation into the incidents. Newsweek reached out to Anthropic via email for comment. Anthropic uses the term "unauthorized access" to describe incidents in which Claude models reached and interacted with real-world computer systems that were not intended to be part of cybersecurity evaluations.
According to the company, the models had been told they were operating in isolated simulations without internet access. Instead, a testing misconfiguration allowed them to reach real systems online. Anthropic said the models then treated those systems as part of the challenges they had been assigned to solve.
The company highlighted four specific incidents: During an incident in January 2026, an early version of Claude Opus 4.6 was participating in a “capture-the-flag” exercise where it had been instructed to retrieve a hidden piece of information from a target machine. According to Anthropic, the model accidentally made that machine inaccessible, making the task impossible to complete. Claude then attempted to quit eight times but was unable to because of another testing error.
Looking for an alternative path, it explored its environment, found a third-party machine, gained administrator access using credentials stored in files and viewed personal information before the session ended. A later incident involved Claude Opus 4.7 during another cybersecurity evaluation. Like the other exercises, the model had been told it was operating inside a fictional environment and that any systems it encountered were part of a controlled test.
Extract — continue reading at the source.