AI Agent Incident Register
AI agent incidents: who was told, and when
By Michael K. Onyekwere, CIPP/E. Sources checked between 27 and 30 September 2026. Not legal advice.
This page covers the eight incidents, or groups of incidents, in this register in which an AI agent run by the company that built or tested it reached a third party's systems in 2026. For each one it records when it happened, when the responsible party found out, when the victim was told, when any government authority was told, and when it became public, from primary sources where they exist and otherwise from named reports. It then checks those dates against seven notification rules. Every date links to its source, and most carry the source's own words. Where a source gives no date, the page says so and does not estimate one.
What the dates show
- Most of the delay came before the responsible party found out. In at least four of the eight, weeks or months passed between the agent's activity and the responsible party finding out: OpenAI's Medicare access in June was found in mid-August, Google's May runs in late July, Anthropic's first incidents from April were found on 24 July, and its fourth, from January, in August or early September. OpenAI detected suspicious activity six to eight days after its agents compromised parts of Hugging Face's infrastructure. The UK AI Security Institute's monitoring raised its alert on 28 July, the day after the run behind the malicious code had finished.
- Where both dates are given, the victim was usually told within days. OpenAI told Hugging Face within a day of detecting the activity. Anthropic notified the organisations three days after it found its first three incidents, and had reached two of the three when it published. The AI Security Institute contacted GitHub, the platform its agent had used, four days after detection. OpenAI's Medicare case took longer: the victim was told three to four weeks after OpenAI says it became aware “in mid-August”, and 84 days after the access. OpenAI now says “we should have shared preliminary findings sooner”. In the other four incidents the sources do not date the notice to the victim, and in at least two of them the evaluation partner gave it.
- Some victims learned of it only from the party whose agent reached them. Anthropic says of the organisations it contacted: “The two organizations we were able to reach had not previously detected the activity or contacted us”. Others caught it themselves: Hugging Face detected the intrusion and disclosed it before OpenAI's public attribution, and in the AI Security Institute's case a human maintainer refused the malicious code.
- In the public record, government authorities were told by victims or by a government evaluator. Services Australia reported to the Australian Signals Directorate, Hugging Face reported to law enforcement, and the AI Security Institute told the UK's government cyber bodies the same day. SecurityWeek reports that OpenAI “said it worked closely with the Australian Signals Directorate throughout the process”. Reports under California's law and to the EU AI Office are confidential, so their absence from the public record shows nothing either way.
- Two became public through the victim or its government first. Hugging Face disclosed its intrusion five days before OpenAI's public attribution, and the Australian Government announced the Medicare incident before OpenAI commented through a spokesperson. OpenAI published its own account on 28 September. Google's and Meta's incidents became public through on-record statements to journalists, Google's in September for runs in May.
Eight is a small number, and it is this register's own sample. It describes these incidents and is not a rate for the industry.
The overview
| Incident | Happened | Found out | Victim told | Authority told | First public |
|---|---|---|---|---|---|
| OpenAI | 18 June 2026 | Mid-August 2026 | 10 September 2026 | By 15 September 2026: the victim reported it to the Australian Signals Directorate | 24 September 2026 (Australian time), by the Australian Government |
| OpenAI | 9 to 13 July 2026 | 19 July 2026 | 20 July 2026 | Undated, by the victim (law enforcement) | 16 July 2026, by Hugging Face |
| OpenAI | Not stated | 29 July 2026, told by its evaluation partner | Undated, by the evaluation partner | Not stated | 4 August 2026 |
| Anthropic | From April 2026 | 24 July 2026 | 27 July 2026 (two of three reached by 30 July) | Not stated | 30 July 2026 |
| Anthropic | January 2026 | Between August and 9 September 2026 | By 9 September 2026 | Not stated | 9 September 2026 |
| Meta | Early July 2026 | Not stated | Undated, by the evaluation partner | Not stated | 5 August 2026 |
| May 2026 | Late July 2026 | Undated | Not stated | 18 September 2026 | |
| UK AI Security Institute | 25 to 28 July 2026 | 28 July 2026 | 1 August 2026 (GitHub, the platform); the users afterwards, undated | 28 July 2026, the same day | 4 August 2026 |
Each incident, with its sources
OpenAI: Unauthorised access to Services Australia's Medicare Statistics Reporting Service portal during an internal evaluation
Analysed in AIR-2026-009
- Happened
- 18 June 2026“On June 18, OpenAI's research team used an internal model to conduct internet based research into public medicine spending.”
- Responsible party found out
- Mid-August 2026“In mid-August, that review identified activity affecting the Australian government websites below.”OpenAI's own account. Earlier, the Acting Prime Minister said OpenAI had advised that it became aware "in August".
- Victim told
- 10 September 2026“We notified Services Australia and the Victorian Department of Health on 10 September and the NSW Bureau of Crime Statistics and Research on 18 September.”The notice went to a public mailbox the agency uses for vulnerability reports from researchers (Minister for Government Services, 24 September). OpenAI now says: "we should have shared preliminary findings sooner".
- Government authority told
- By 15 September 2026: the victim reported it to the Australian Signals Directorate“By 15 September, once Services Australia had analysed the information in the email and made some checks, they notified the incident to ASD.”SecurityWeek: "The company said it worked closely with the Australian Signals Directorate throughout the process". No date is given for OpenAI's own first contact with the agency.
- Made public by the responsible party
- After the Government's announcement: a spokesperson's statement, then a first-party post on 28 September 2026 (US time)“In June, during internal training and evaluation our models accessed Australian government websites in ways they were not authorised to.”The spokesperson's words, "our models took actions we did not intend", were added to SecurityWeek's report in its update of 25 September.
- First public, by anyone
- 24 September 2026 (Australian time), by the Australian GovernmentThe Prime Minister spoke in New York, where it was still 23 September.
OpenAI: Evaluation agents reached Hugging Face's production systems
Analysed in AIR-2026-009
- Happened
- 9 to 13 July 2026“between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC”The span of attacker actions Hugging Face reconstructed, the first of them on a third-party launchpad. OpenAI's report puts the compromise of Hugging Face's infrastructure "between July 11 and July 13".
- Responsible party found out
- 19 July 2026“OpenAI detected suspicious internal activity on July 19.”
- Victim told
- 20 July 2026“On July 20, after further investigation and initial outreach to Hugging Face to rotate credentials discovered in that investigation, OpenAI determined the activity was likely related to the Hugging Face compromise, informed Hugging Face”The same sentence records earlier "outreach to Hugging Face to rotate credentials". Hugging Face had already detected the intrusion itself and disclosed it on 16 July.
- Government authority told
- Undated, by the victim (law enforcement)“we have also reported this incident to law enforcement agencies”
- Made public by the responsible party
- 21 July 2026“July 21, 2026: We disclose”
- First public, by anyone
- 16 July 2026, by Hugging Face“Published July 16, 2026”
OpenAI: An evaluation model exploited a real website in a partner's test environment
Analysed in AIR-2026-012
- Happened
- Not stated
- Responsible party found out
- 29 July 2026, told by its evaluation partner“On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models”
- Victim told
- Undated, by the evaluation partner“affected third parties were notified”OpenAI reports this of Irregular and gives no date.
- Government authority told
- Not stated
- Made public by the responsible party
- 4 August 2026
- First public, by anyone
- 4 August 2026
Anthropic: Three incidents in which Claude models reached real systems from a partner's test environment
Analysed in AIR-2026-012
- Happened
- From April 2026“The earliest incidents date to April.”
- Responsible party found out
- 24 July 2026“We identified all three incidents the next day.”The review began on Thursday 23 July.
- Victim told
- 27 July 2026 (two of three reached by 30 July)“We notified our evaluation partner Irregular and the three affected organizations on Monday, July 27.”Anthropic adds: "The two organizations we were able to reach had not previously detected the activity or contacted us".
- Government authority told
- Not stated
- Made public by the responsible party
- First public, by anyone
Anthropic: A fourth incident, found in a later search, in the same partner's environment
Analysed in AIR-2026-012
- Happened
- January 2026“a fourth incident, from January 2026”
- Responsible party found out
- Between August and 9 September 2026“we identified these in August while assembling transcripts to share with METR. We scanned these transcripts and identified a fourth incident”The transcripts were found in August, and the incident was identified when they were scanned. Anthropic does not date the scan.
- Victim told
- By 9 September 2026“We notified the affected party after we discovered this fourth incident.”The same post says "We have notified all affected parties."
- Government authority told
- Not stated
- Made public by the responsible party
- 9 September 2026
- First public, by anyone
- 9 September 2026
Meta: A pre-release Muse Spark 1.1 breached a real website in a partner's test environment
Analysed in AIR-2026-012
- Happened
- Early July 2026“In early July, Irregular began an exercise”
- Responsible party found out
- Not statedMeta gives no date. Its evaluation partner, Irregular, told CNBC: "All relevant labs were notified in late July, and affected entities were contacted as part of the investigation."
- Victim told
- Undated, by the evaluation partner“ensured that the affected party was also notified”Meta says this of Irregular. It also says it holds limited information about the affected company, because the evaluation ran on Irregular's infrastructure.
- Government authority told
- Not stated
- Made public by the responsible party
- 5 August 2026, statement to the mediaSecurityWeek, 6 August: Meta "said in a statement to the media on Wednesday", which was 5 August. Meta's own post followed on 14 August.
- First public, by anyone
Google: A Gemini model reached three companies' systems in a partner's test environment
Analysed in AIR-2026-012
- Happened
- May 2026“The hacks occurred in May”
- Responsible party found out
- Late July 2026“Google said the incident happened in May and it was notified by Irregular in late July.”Google's account as reported by CNBC, which also quotes Irregular: "All relevant labs were notified in late July".
- Victim told
- Undated“We ensured the three entities were made aware”
- Government authority told
- Not stated
- Made public by the responsible party
- 18 September 2026, on-record statement reported by the Wall Street JournalCNBC: "Google said on Friday", which was 18 September in the US. No first-party post by Google or Google DeepMind has been found.
- First public, by anyone
UK AI Security Institute (the evaluator; the models were from Anthropic and OpenAI): Test agents created fake identities to pressure a real maintainer into approving malicious code
Analysed in AIR-2026-013
- Happened
- 25 to 28 July 2026“from 25 to 28 July 2026”The run behind the malicious pull request ran from 26 July 12:45 to 27 July 23:15; it had finished when the alert was raised.
- Responsible party found out
- 28 July 2026“On the morning of 28th July, our security monitoring flagged data leaving one of our testing systems”
- Victim told
- 1 August 2026 (GitHub, the platform); the users afterwards, undated“GitHub was contacted on Saturday 1st August, 22:21 BST”The report says AISI "requested support in informing affected users". Its blog says the maintainer caught the code: "A human maintainer caught and refused to approve the malicious code."
- Government authority told
- 28 July 2026, the same day“By 18:00 BST we had informed the Government Cyber Coordination Centre (GC3), National Cyber Security Centre (NCSC)”AISI is itself a government body. It told the model developers and the US Center for AI Standards and Innovation on 3 August.
- Made public by the responsible party
- First public, by anyone
The notification rules this register checked
Each rule below was read in its official text on 27 or 28 September 2026, except Australia's scheme, which was read in the regulator's own guidance, and the EU Code of Practice, which was read in the official text published by the European Commission. The last column is this register's reading of whether it would have reached the incidents above.
The gap
On this register's reading, none of the rules it checked obliges a developer to tell the organisation its agent reached. The rules that do reach developers (California's law, New York's from 2027, the AI Act's duty for the most capable general-purpose models, and the EU Code of Practice that three of these developers have signed) require a report to a government body, not to the victim. The breach rules bind whoever holds the data, which here is usually the victim. What the victim hears, and when, is left to the developer's own policy.
OpenAI has published its policy. After its review found more cases of agents going beyond their tasks on third-party websites, it wrote that “Our goal is to give each organization the facts and defer to them on if and when to make the incident public.” Its reporting framework says its process runs “with deadlines for each step to ensure timely investigation and disclosure”, that staff will “assess whether any third party was affected and needs private notification before publication”, and that “When a third party is affected, our security, legal, and responsible disclosure obligations take precedence over this framework.” It does not publish those deadlines. After the Medicare case, OpenAI wrote that “If we identify any additional affected agencies, we will notify them promptly and directly with the information available and provide updates as further facts emerge.” That is a commitment in a company post, and it names no time.
The Australian Government's taskforce, announced on 24 September (Australian time), is to review “whether existing processes are appropriate to respond to AI-related cyber incidents”, and the Prime Minister said its report “will consider also possible law enforcement and legislative responses”. The dates on this page bear directly on that question.
Sources for this section: OpenAI, timeline page, update of 25 September 2026; OpenAI, “Our framework for reporting model misalignment”; OpenAI, “How we will do better for Australia”, 28 September 2026; Prime Minister of Australia, press conference, 24 September 2026.
Method and limits
Scope. An incident is included when an AI agent, run by the company that built it or by an organisation testing it, reached the systems of a third party, and the Register has analysed it. The Register's other incidents are excluded because their disclosure follows different routes: software vulnerabilities reported by researchers (EchoLeak, DuneSlide), a compromised product release (Amazon Q), a customer's own deployment (Replit), a breach of an agent vendor's credentials (Salesloft Drift), and decided cases.
Dates. Each date is as the primary source states it, at the precision it states it. “Not stated” means no source this register holds gives the date as at 30 September 2026. It does not mean the event did not happen. Where a date comes from a third party, such as a government or an evaluation partner, the page says whose account it is. The dates for Google and Meta come from news reports, which the page names. Two dates are worked out from a weekday in a dated report: Meta's statement “on Wednesday”, which was 5 August, and Google's “on Friday”, which was 18 September in the US. Dates in Australian sources are given in Australian time.
What the public record cannot show. Reports of critical safety incidents to California's Office of Emergency Services are exempt from the state's public records law, and reports to the EU AI Office are not published. A report missing from this page may still have been made.
Updates. The dates will change as companies publish more. Corrections follow the register's methodology: dated, and never silent.
Data
The full dataset, with every date, quotation and source, is at /api/disclosure-timelines as JSON, licensed CC BY 4.0. Cite it as: Onyekwere, Michael K., AI agent incident disclosure timelines, AI Agent Incident Register, CompanyScope, https://companyscope.io/register/ai-agent-incident-disclosure-timelines, as at 30 September 2026.
The liability analysis for each incident is in its Register entry, and the question it all comes back to is answered on who is liable when an AI agent causes harm.
Subscribe to the AI Agent Incident Register
Every new Register entry delivered with the legal analysis: the incident, the duty engaged, who is liable across the chain, and what governance would have prevented it. Written by Michael K. Onyekwere, CIPP/E. Free.
Subscribe - freeDelivered via Compliance Engineering on Substack, which handles your subscription and consent. Unsubscribe any time. Privacy notice.