Last week, when Australian Prime Minister Anthony Albanese told the world that an OpenAI agent had accessed “non-public files” from his country’s Medicare statistics portal during testing, his description of the incident was a little light on details. Today, we’re getting new information on just how far OpenAI’s overzealous agent went in attempting to satisfy a rather innocuous-sounding informational prompt. In a newly published blog post, OpenAI says the June incident started when the company asked “an experimental, internal-only OpenAI model” to research government spending statistics in the Australian state of Victoria.
When the model ran into trouble finding that data using the publicly published statistics that it was supposed to reference, “it took actions that we had not authorized it to take” to find an answer, OpenAI said. Those unauthorized actions included finding “a way to gain non-public access to the service” and using that access to view “technical system information and source code” alongside credentials and the aggregate statistics it was actually searching for, OpenAI said. In a newly published disclosure email that was sent to Australia’s Public Disclosure account earlier this month, OpenAI said its model had “identified a way to make the server carry out instructions sent through the public reporting interface, without a private account or password.” That unauthorized access let the agent “read portions of internal program files and settings, obtain a list of files, and create and read back a small test file on the server,” according to the email.
Extract — continue reading at the source.