Tech
EN AZ
Why did an OpenAI system hack Australia's health system - and can it be stopped in the future?

Why did an OpenAI system hack Australia's health system - and can it be stopped in the future?

bbc.co.uk 24.09.2026 16:08 1 views
News that an automated AI agent hacked a government IT system raises big questions about regulating the tech.

An OpenAI agent has gone "rogue" and "infiltrated" an Australian government website in what cyber-security experts are calling the first hack of its kind. But why did it take the government months to discover what happened - and could it happen again? The hack was carried out by an AI agent - an autonomous computer program that uses AI to complete a task with minimal human oversight.

On 18 June one of OpenAI's agents went rogue during a test exercise - the company has said it was was supposed to "look up answers, and available statistics for questions about Australia during an internal evaluation". In the process it "infiltrated" a private statistics portal containing "non-sensitive" data from Australia's universal healthcare scheme Medicare, Prime Minister Anthony Albanese said. OpenAI said it only realised the breach had happened at all in August while reviewing "misaligned model activity", and the company sent an email to a generic Australian government inbox some weeks later.

That email seems to have gone unnoticed for five days before it was escalated to Australia's cyber-security experts on 10 September. The prime minister described the breach as "obviously unacceptable" and said OpenAI took "way too long" to inform Australian officials. Analysts have also raised concerns over OpenAI's almost three-month delay in noticing and reporting the breach via email.

"The way the notice arrived bothers me as much as the delay," chief data and AI officer Simon Liu from cyber-security firm TrustDecision told the BBC. Australia has said this incident is the first of its kind, and experts agree it might be. As far as we know, hacks carried out by AI agents are still quite rare occurrences - but then again, it is largely up to companies themselves to disclose them.

Hacks like this have happened before. In July, OpenAI agents went rogue during a test and infiltrated tech start-up Hugging Face's internal systems. The AI agents decided that ignoring the limits on what should be done to achieve their goal was the best course of action.

This is what the industry calls "misalignment" - broadly defined as when AI machines do not act in humanity's best interests, such as by bending the rules. It is a problem that is fundamental to making AI safe, and it is proving challenging. To put it simply, the type of AI models at play here - known as large language models - are designed to predict the likeliest output to a given input, rather than consider the consequences of that output as a human would.

Extract — continue reading at the source.

Read full story