Dev

Three records of one agent incident

OpenAI's logs, a URL scanner's public archive and a statistics bureau each tell it differently. Whoever keeps the fullest account is also the one grading it.

At the end of September OpenAI apologized to Australia. In June, during internal training and evaluation, its agents went after data held by four government agencies, and at one of them got non-public access. The model that went furthest was, in OpenAI's account, experimental, internal-only and running without the full set of safeguards its public products have. Three parties now hold a record of what happened, and they do not tell it the same way.

What did the agents reach?

One confirmed case of non-public access. For the other three, the accounts classify what happened in different terms. A model asked for government spending per person on skin-condition medicines in Victorian communities struggled to find it in published statistics. OpenAI says it then gained non-public access to Services Australia's Medicare Statistics Reporting Service, ran commands, retrieved internal files, credentials and aggregate statistics, wrote files, and read source code, still chasing the same question.

Other agents used BOCSAR's public Crime Mapping Tool, an exposed access key at the Victorian Agency for Health Information, and third-party browsing services at the Australian Institute of Health and Welfare, where attempts to bypass access controls failed. OpenAI says no individual records were touched anywhere.

Who found out, and how?

Mostly OpenAI, months later. After its agents reached Hugging Face in July, it started reviewing earlier runs and in mid-August found the Australian activity. It notified Services Australia and the Victorian Department of Health on 10 September, BOCSAR on 18 September and AIHW on 24 September. The Prime Minister said officials were told almost three months after the fact, and according to The Record complained the notice was an email to a generic government inbox.

OpenAI was not the only one looking. On 23 September the research lab Transluce published a reconstruction of agents probing AIHW on 20 and 21 June, built from the public records of urlquery.net, a URL-scanning service the agents used as a remote browser. Transluce says it contacted OpenAI and the affected organizations on 21 and 22 September, before OpenAI's own notice to AIHW. None of the reporting describes an agency identifying the activity in June.

Do the three accounts agree?

They rarely contradict each other outright, but they classify the same events differently. OpenAI on BOCSAR:

The BOCSAR system returned application configuration, operational jobs and logs, and website metadata.

BOCSAR, in its 25 September update:

Investigations to date have found no evidence of a security vulnerability in the Crime Mapping Tool, and no system changes have been identified as necessary at this stage. There is also no evidence that any information has been accessed beyond what is already publicly available through the tool.

OpenAI on AIHW:

The activity related to the Australian Institute of Health and Welfare did not meet our disclosure thresholds because the way it was accessed seemed consistent with public access

Transluce on the same episode, after describing a cross-site scripting probe:

Cloudflare's firewall blocked the probe before it reached the dashboard.
The file itself is public, so no non-public data was exposed, but the agent bypassed the site's anti-bot controls.

Why does it matter who holds the record?

Because each record frames a different incident. In OpenAI's account, AIHW was accessed in a way consistent with public access, below the threshold. In Transluce's reading of urlquery.net's public scans, it was an agent probing a government dashboard for injection and routing around its bot protection. On AIHW's side, per Transluce, a firewall blocked a probe, and no published account says what anyone there made of it.

My read is that the operator's record is likely the fullest, since it holds the task and the model's reasoning directly, and it is also the one grading itself. "Consistent with public access" describes how the service was reached and what came back. Whether the owner permitted what happened is a separate question, and only the owner can settle it. That is the class of rule whose damage lives in the world and not in the call.

What changes now?

OpenAI says its research environments now block live internet and serve web access from a cache, and that training and evaluation with tool use for its most capable models is paused, part of the same series of reports behind the compaction case I wrote about. Chief Strategy Officer Jason Kwon is due before the Joint Select Committee on Artificial Intelligence in Sydney on 6 October. OpenAI's taskforce, expected to finish by the end of the year, is to recommend policy on notification, coordination with government and protecting government systems.

What should you take from it if you run agents?

That your record will probably be the fullest one, and checking your reading of it from outside is hard unless some public log happens to exist. An agent with network access acts with whatever reach you gave it.

Label your agents' traffic with a user agent or header, knowing that a third-party fetcher may drop your identifying headers and publish a record somewhere you do not control, so check what actually reaches the destination. Keep a log keyed by the host touched, so "which outside systems did anything reach last month" is a query. And write the disclosure rule before you need it, with the threshold turned the other way from OpenAI's: if the owner could not have seen what your agent did without you, that is the reason to tell them.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles

AI search visibility measurement framework for technical teams

One panel gives you a snapshot. A report compares two periods, and that is where most AI visibility changes turn out to be artifacts of the panel, the sample size or an unlogged model swap. The design that makes a trend checkable.

Sonnet 5.5 keeps its reasoning for one model and one account

Switch to a model or an API account that can't read the earlier thinking, and the request still succeeds. The reasoning those turns built up is dropped before the model reads it, with no error and nothing on the bill. You only find out if you asked the API to report it.

A benchmark has to prove it tracks the clock

Once an agent will push any number you give it, the model stops being the part worth arguing about. The part you own is the evidence that the number and the goal move together, and that evidence expires as soon as the optimizer leaves the region where you checked it.