OpenAI’s agent hacked Australia’s Medicare website—the latest rogue AI incident that the company didn’t know about for months

1 hour ago 1
Add to circle

An OpenAI agent infiltrated an Australian government website in June, Australian Prime Minister Anthony Albanese said when speaking to reporters at the United Nations General Assembly (UNGA) in New York.

OpenAI did not notify the Australian government until Sept. 10, Albanese said. The agent gained unauthorized access to the public-facing Medicare Statistics Reporting Service, enabling it to access both public and non-public files, and to write files to an internal server.

It’s the latest in a growing list of systems OpenAI’s agents have accessed without authorization, and largely without OpenAI or the victims knowing until weeks or months later. Meanwhile, public trust in AI safety is cratering; a recent survey by Politico found that two-thirds of Americans think there is at least a “moderate” risk that advanced AI could destroy humanity.

“This situation is obviously unacceptable,” Albanese said, according to the Sydney Morning Herald. “And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia’s extreme concern about this incident, and I also expressed my disappointment that it took the company way too long to inform the government what had occurred and the nature of the way that that notification occurred as well was unacceptable.”

An OpenAI spokesperson told Fortune that the company did not notify the government of the breach until three months later because it was not aware it had happened. The company discovered it in August as part of an “extensive review” of any cases in which its models behaved in unexpected, or “misaligned,” ways during training and evaluation.

“The information accessed included aggregate health statistics and internal file names,” OpenAI said. “We notified the organizations and are providing technical information to support their investigations and help address potential security vulnerabilities. Our overall review is ongoing, and we remain committed to transparency about these issues and to sharing what we learn as that work continues.” 

Albanese said the Australian government is investigating the impact of the incident, and so far has not found evidence that the agent accessed any personal information. OpenAI also said it dound “found no evidence of patient records being accessed.” Albanese said the government is also aware of three other government systems the agent may have reached, two additional health-related organizations, and one related to crime statistics and research.

Perhaps not coincidentally, OpenAI became aware of this incident in August, the same month it published its long-awaited review of the Hugging Face hack, which occurred in July. The Hugging Face hack may have prompted an internal review, during which OpenAI also discovered the Australian website breach, although the company did not explicitly link the two events in its statement. In its Hugging Face report, OpenAI likewise confirmed it did not know about the breach until after the fact because of poor agent montioring and alarms; OpenAI said it has since bolstered those safety mechanisms.

OpenAI CEO Sam Altman is also in New York this week, attending a United Nations Security Council meeting. In his remarks, he spoke about the “anxiety” surrounding powerful AI systems, particularly the possibility that “we could lose control of the future to AI.”

“The risk is that it moves so fast that people can no longer follow what’s happening or intervene when needed. This would obviously be terrible,” he added.

Altman called for international cooperation to create “standards for measuring capabilities, assessing risks, determining whether safeguards are sufficient, and preserving meaningful human oversight as systems become more autonomous.” He also called for more reliable incident reporting, yet OpenAI did not reveal its breach of the Australian government website when it revealed a framework for disclosing incidents on Sept. 16.

As part of that framework, it disclosed six examples. The decision to publish a framework was in response to another report of misaligned model behavior, this time by rogue agents that co-opted a German wikipedia page to use for a messaging board. In this case, OpenAI knew about the incident but did not disclose it for weeks.

This story was originally featured on Fortune.com

Read Entire Article