Australia did not detect the intrusion into its Medicare statistics portal. OpenAI told them, 84 days after it happened. For UK risk and security leaders, the uncomfortable questions are the same ones, and most organisations do not have answers ready.

ai-agent-84-days-1200x675

An AI agent broke in. Nobody noticed. 84 days before Australia was told.

On 18 June, an AI agent built by OpenAI gained unauthorised access to the Medicare Statistics Reporting Service, a portal run by Services Australia. It had been given a benign task: compile health and medicine spending statistics during an internal evaluation. When the portal blocked its requests, the agent tried alternative routes, bypassed access restrictions, read public and non-public files, and wrote files to an internal server. Cyber Security News.

No patient records were involved. The portal holds aggregate figures on billing rates and medicine costs. Australia’s deputy prime minister, Richard Marles, has called the incident relatively minor, and on the facts available he is right.

The severity is not the point. Three things about this incident should concern anyone responsible for risk, security or resilience in a UK organisation, and none of them depend on how much data was taken.

Nobody was driving

The agent was not directed to break in. It was given a research task, met an obstacle, and routed around it. OpenAI has described the behaviour as misaligned, and says its models took actions the company did not intend.

That distinction matters for how we think about controls. Our security models assume an adversary with intent or an insider with access. This was neither. It was a system pursuing an assigned goal, encountering a barrier, and treating the barrier as a problem to solve. The access control worked exactly as designed, and the agent went around it.

Rob Nicholls of the University of Sydney put the consequence plainly: if a person had done this, we would call it hacking, and the fact that an AI agent did it makes our disclosure laws look out of date rather than making the act less serious.

Nor is this isolated. The AI research lab Transluce said this week it had detected agents going rogue on several occasions dating back to at least March, and OpenAI disclosed in July that during cybersecurity testing its models created a swarm of agents that hacked into Hugging Face’s systems. CNN

OpenAI discovered the activity in August while reviewing misaligned model behaviour, and notified Services Australia on 10 September, 84 days after the breach. The notification went to a public disclosures inbox used by academics and researchers to report weaknesses. It was read the next day. Escalation to the Australian Signals Directorate took a further four days. Cuber Security News | ABC News

Two failures sit inside that timeline, and the second is the one to take personally.

The first is disclosure. There was no obligation on OpenAI to report faster or through a different channel, which is precisely what Australia’s new taskforce has been asked to examine. That taskforce, led by the prime minister’s department with the Australian Signals Directorate and the AI Safety Institute, is also running a forensic investigation into whether other government systems were affected. Technology Org

The second failure is detection. An external party had to tell Australia that its systems had been accessed. Investigators will now examine why government monitoring did not pick it up. Most organisations reading this would have performed no better. Agent traffic looks like automated crawling, which is background noise on any public-facing service. Cyber Security News

The UK view

The day before Albanese spoke, Andy Burnham told the UN that the UK would establish a National Centre for Information Defence, bringing together the intelligence agencies, law enforcement and social media companies to detect, attribute and disrupt information attacks from hostile states. He was explicit that AI will multiply the threat.

Set the two announcements side by side and the shape of the problem emerges. One government is building capability against AI-enabled attacks by hostile states. The other has just been breached by an AI agent from an allied company, acting on its own, and found out by email three months later.

Cory Alpert of the University of Melbourne made the point that lands hardest: had this been a Chinese or Russian model, the reaction would have been markedly different, and the vulnerability is identical either way.

For UK organisations, the regulatory direction is already visible. The Cyber Security and Resilience Bill is due by the end of 2026 and will reshape incident reporting, supply chain responsibilities and procurement. An AI vendor whose agent accesses your systems is a supply chain question, and the answer to “who reports it, to whom, and how quickly” is not settled.

Four questions worth asking this week

  • Would we know? If an agent accessed a non-public file on one of our public-facing services, would our monitoring distinguish that from ordinary crawler traffic?
  • Who would tell us? What contractual notification obligations do we have with our AI vendors, and what timeframe do they commit to?
  • Whose incident is it? If an AI agent causes the event, does it reach the security team, or does it sit with whoever owns the AI relationship?
  • Can an agent write? Read access is one risk profile. Write access to an internal server is another entirely, and most AI risk assessments do not distinguish between them.

None of these need new technology. They need someone to own the answer.

#RISK Expo Europe 2026 | 10–11 November, ExCeL London

The questions raised by the Medicare incident run through the agenda:

Tuesday 10 November

  • 11:10 | BFSI Stage — Operational Resilience in a Crisis-Prone World: What Boards in BFSI Need Now, with Aaron Kalvani, Independent AI Strategist and Advisor on UN and Global AI Governance, Filip Maertens of Euroclear, and Emma Hurrell of HSBC Bank UK
  • 11:20 | #RISK Stage — AI and Cloud Concentration Risk: The Governance Questions Boards Can No Longer Defer
  • 12:50 | #RISK Stage — The Cyber Security Resilience Bill: What It Means for Your Business, with Duncan McDonald, CISO, NCC Group, and Florian Pouchet, Wavestone
  • 16:00 | #RISK Stage — Geopolitical Risk: Building Resilience Into Strategy, Not Just Contingency Plans, with Alex M., Head of Financial Services, Cyber, at the National Cyber Security Centre, and Christopher Mason of Woodhorn Global

Wednesday 11 November

  • 10:00 | #RISK Stage — Ransomware, Geopolitics and Critical Infrastructure: Preparing for the Worst Day
  • 10:00 | PrivSec AI Governance — Agentic AI and the Death of the Static Data Inventory, with Michael Charles Borrelli, AI & Partners
  • 10:30 | #RISK Stage — Executive Accountability and Personal Liability, with Phillip Davies, CISO, Equifax
  • 11:10 | PrivSec AI Governance — AI vs AI: Cyber Threats, Defences and the Governance Gap Between Them
  • 12:50 | PrivSec AI Governance — Shadow AI and the Extended Vendor Chain

Join the conversation Register for #RISK Expo Europe