· Updated · Matt O'Kane · Insights  · 19 min read

What's confirmed about AI agents and AI tools getting into government systems

Which claims about AI agents and AI tools getting into government systems are confirmed, which rest on company findings, and which headlines got wrong.

Last updated 7 October 2026. The Australian investigation is still open, and OpenAI says its own review will take months, so details may change.

Following the Australian case? Read our full page on the OpenAI agent and the Medicare statistics portal (NDF-AI-2026-06).

As of 7 October 2026, the governments of Australia and Taiwan have confirmed incidents in which AI agents got into their systems. The Australian Government says an OpenAI agent gained unauthorised access to a statistics portal. Taiwan’s Ministry of Digital Affairs says hackers used AI agents in an attack on government agencies. The NSW Government says an OpenAI model also entered a National Parks and Wildlife Service web application, but came up with only public information. In the United States, OpenAI says its agents accessed Census Bureau data that was publicly available, and no US agency has confirmed more. Ukraine’s CERT has confirmed malicious emails carrying malware that calls an AI model, which is not an agent acting on its own.

For every other case on this page, the claim comes from an AI company or a security firm, and no government involved has confirmed it. Some of those claims have strong evidence behind them, such as the Mexico and Thailand cases, where technical indicators have been published. A company’s finding and a government’s confirmation are different things, and news reports do not always make the difference clear.

The cases

What we include. A case is one reported incident in which AI agents, or AI tools in a person’s hands, were used against a government’s systems. We include failed attempts and cases no government has confirmed, and the table shows which is which. We count a case as confirmed only when the affected government confirms it (see How we assessed each claim). This page covers the cases we have assessed; it is not a complete list.

Each case has a number made up of the year it became public and its order within that year. Numbers are not reused. They are NDF’s own reference, not an official register.

No.CaseWho ran the AIWho says it happenedWhat the affected authority says
NDF-AI-2025-01Ukraine, June to July 2025APT28, per Ukraine’s CERT (moderate confidence)Ukraine’s CERT; GoogleConfirmed the malicious emails; did not say whether anything was taken
NDF-AI-2025-02Anthropic’s GTG-2002 extortion case, 2025A criminal using Claude CodeAnthropic onlyNo government named
NDF-AI-2025-03Anthropic’s August 2025 case targeting Vietnamese infrastructureAn actor using Claude as an assistantAnthropic onlyNo public statement found
NDF-AI-2025-04Anthropic’s GTG-1002 espionage case, 2025A group Anthropic calls GTG-1002Anthropic onlyNo government named
NDF-AI-2026-01Mexico, 9 government bodies, Dec 2025 to Feb 2026One operator using Claude Code and, Gambit says, GPT-4.1Gambit Security; Dragos (water utility). Anthropic says it banned the accountsThe tax authority and electoral institute said in February their checks found no breach. Jalisco said its own networks were not affected but federal networks were
NDF-AI-2026-02Thailand’s Ministry of Finance, July 2026Operators running an AI agent unattendedSecurity firm Hunt.ioNo public statement found. Hunt.io says Thailand’s national CERT and cyber security agency acknowledged its notice
NDF-AI-2026-03Taiwan, July 2026Hackers using AI agentsTaiwan’s Ministry of Digital Affairs. Dream describes a matching campaign against “government entities in Asia”The Ministry of Digital Affairs confirmed an attack combining hacker operations with AI agents
NDF-AI-2026-04District government in mainland China, 2026An operator using a multi-model AI frameworkSecurity firm Hunt.ioNo public statement found
NDF-AI-2026-05Anthropic’s September 2026 casesGroups Anthropic tracks, using ClaudeAnthropic onlyNo government statement found
NDF-AI-2026-06Australia, Medicare statistics portal, June 2026OpenAI, in an internal research task (no attacker)Australian Government; NSW Government; OpenAIUnauthorised access to non-public files confirmed; no personal information believed accessed at this stage. NSW says the NPWS agent came up with only public information
NDF-AI-2026-07United States, Census Bureau and others, 2026OpenAI, in internal training tasks (no attacker)OpenAI; Transluce (Education attempt)The SEC says it is unaware of unauthorised access to non-public information; Commerce says no private Census data was accessed; Education says it found no evidence of impact
NDF-AI-2026-08Canada, Library and Archives Canada, May to June 2026Unidentified agents; Transluce does not confidently attribute them to OpenAIResearch lab TransluceNo indication of compromise; the Canadian Centre for Cyber Security is assessing the reports

NDF-AI-2025-01: Ukraine

In July 2025 Ukraine’s CERT reported emails sent to government bodies carrying malware that asks an AI model for commands. It attributed the activity to APT28 with moderate confidence. Google says it first identified the malware in June. The malware follows a fixed script and does not act on its own the way an agent does. The CERT did not say whether data was taken.

NDF-AI-2025-02: Anthropic’s GTG-2002 case

In August 2025 Anthropic said a criminal used Claude Code in a data theft and extortion operation that targeted at least 17 organisations. It said the targets were “in healthcare, the emergency services, and government and religious institutions”. Anthropic says Claude Code was used to automate reconnaissance, harvest credentials and get into networks. It did not name the victims or say which government institutions, if any, were breached.

NDF-AI-2025-03: Anthropic’s Vietnam case

In the same month, Anthropic reported that an actor used Claude across a campaign against Vietnamese critical infrastructure. It says the actor “appears to have compromised major Vietnamese telecommunications providers, government databases, and agricultural management systems”. Claude was used as an assistant and adviser, not as an agent acting on its own. The victims are not named, and we have found no statement from Vietnam’s government.

NDF-AI-2025-04: Anthropic’s GTG-1002 case

In November 2025 Anthropic said a group it calls GTG-1002 used Claude Code against about 30 targets, including government agencies. It said the group “succeeded in a small number of cases”. Its full report describes “successfully obtaining access to confirmed high-value targets” including “government agencies”, but does not name them. We have found no technical indicators published for this case. Anthropic says it “notified affected entities as appropriate” and coordinated with authorities. The report says Claude “frequently overstated findings and occasionally fabricated data during autonomous operations”, which hampered the attacker, who had to check Claude’s claimed results. Anthropic says its own investigation “validated a handful of successful intrusions”.

NDF-AI-2026-01: Mexico

Gambit Security reports that a single operator broke into at least 9 Mexican government bodies between December 2025 and February 2026. Gambit says the attacker used Anthropic’s Claude Code, and OpenAI’s GPT-4.1 to analyse the data collected. It says Claude refused or resisted some requests, “questioning the legitimacy of operations”. Gambit’s evidence comes from servers the attacker used, and it has published technical indicators.

An Anthropic spokesperson said the company investigated, disrupted the activity and banned the accounts, Engadget reported. OpenAI said it identified attempts by the hacker to break its usage policies and that its tools refused to comply.

Dragos, which Gambit brought in to assist, examined more than 350 of the recovered files. They related to a municipal water and drainage utility in the Monterrey area. Dragos found the utility’s IT network had been compromised. A password attack on the gateway to its control systems was unsuccessful, and Dragos saw no evidence the attacker reached those systems. Gambit says trusted parties can ask for sanitised copies of its forensic material.

In February, Mexico’s tax authority said a review of its logs found no illegitimate access. The electoral institute said its checks found no breach, unauthorised access or data taken. It added, as IT Masters Mag reported, that no verifiable public technical evidence, such as indicators of compromise, had so far been presented. Engadget reported that Jalisco’s state government denied its own systems were breached, saying only federal networks were affected. These statements were made in February. Gambit published its full report, with indicators, on 10 April. We have found no later public statement from these bodies, or any public comment from the other bodies Gambit names.

The denials come from the bodies’ own checks. We have found no public review of the evidence by anyone unconnected to Gambit or to the bodies concerned.

Two points from this case are often misreported. Gambit’s “195 million” is its count of taxpayer records it says were taken from the tax authority. That is more than Mexico’s entire population, so it counts records, not people. Gambit also calls the operation “hybrid, human-directed” and says several victims were compromised by hand. It says Claude Code ran about 75 per cent of the remote commands.

NDF-AI-2026-02: Thailand

Hunt.io examined an attacker’s exposed servers. It found an open-source AI agent working inside Thailand’s Ministry of Finance network in July 2026. The agent ran in a mode that skips the approval prompts for risky commands. Hunt.io says active session cookies, deployed web shells and internal network access indicate the operator compromised several ministry systems. It found no evidence that files in a ministry directory the agent listed were taken, and it published technical indicators. We have found no public comment from Thailand’s government.

NDF-AI-2026-03: Taiwan

In August 2026 the Administration for Cyber Security at Taiwan’s Ministry of Digital Affairs said overseas hackers had attacked government agencies in July. It described “a hybrid pattern” of hacker operations combined with AI-agent-assisted attacks using tools such as OpenClaw. It said the agents could “rapidly chain” attack techniques and use backup and test systems as stepping stones. It said the affected agencies had completed their response. (Quotes are our translation from the Chinese.)

The detailed claims in the news come from the Israeli security firm Dream, not from the Ministry. They include how many accounts were cracked and records taken, and the targeting of a nuclear safety agency. Dream’s report describes an unnamed government in Asia. The Financial Times reported the target was Taiwan, according to Focus Taiwan. Dream describes the nuclear safety agency as one of several targets the attacker scanned.

Some headlines called this the first fully autonomous attack on a government. Dream calls it “near-autonomous” while describing some individual steps as autonomous, including agents adapting “without human intervention”. The Ministry called it a hybrid of hacker operations and AI agents.

NDF-AI-2026-04: District government in mainland China

Hunt.io examined an operator’s exposed servers and found a framework that ran Claude, Qwen and DeepSeek as workers. It says the most extensive compromise it found was of a district government’s office systems in mainland China. There, it says, the operator ran commands and collected credentials. Hunt.io says the AI organised the work, but the break-ins relied on conventional scripts, leaked credentials, web shells and custom implants. It says it disclosed its findings to the relevant national CERTs before publishing, and it published technical indicators. We have found no public statement from the affected government.

NDF-AI-2026-05: Anthropic’s September 2026 cases

In September 2026 Anthropic described further cases involving government targets. In one, it says an actor it tracks as GTG-20006 took more than 300,000 national identity records from an unnamed North African government technology authority. It says the same actor took mail records from organisations including a national prosecutor’s office. Anthropic published technical indicators for the actor’s infrastructure, but the victims are not named and no government has confirmed the cases.

NDF-AI-2026-06: Australia

Read the full page on this case, with the timeline, what each party has said and what is still unknown. That page is where we keep the detail of this case up to date.

This investigation is still open. Details on this page may change.

A forensic investigation is under way, aided by the Australian Signals Directorate (ASD). A taskforce led by the Department of the Prime Minister and Cabinet is also reviewing the incident. Nothing below should be read as the final finding.

On 24 September 2026 Prime Minister Anthony Albanese said an OpenAI agent had gained unauthorised access to the public-facing Medicare Statistics Reporting Service portal on 18 June. Services Australia runs the portal. The agent was carrying out an internal OpenAI research task, and no attacker was involved. The government and OpenAI both say no personal information or patient records are believed to have been accessed. In its own account, OpenAI says the model “ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files”, but it does not explain how the model got in. OpenAI has apologised, and on 6 October its chief strategy officer, Jason Kwon, apologised to Parliament’s Joint Select Committee on Artificial Intelligence.

OpenAI has also notified other Australian government bodies of agent activity on their sites, including a NSW National Parks and Wildlife Service (NPWS) web application. The Guardian reported on 3 October that the NPWS site was the sixth Australian government website OpenAI had notified since September. We have found only five of them named.

Was this the first case? Possibly. On this tracker, it is the earliest access to a government system that the affected government has confirmed. Taiwan’s government confirmed an attack that combined hackers with AI agents, and said so in August. But that attack took place in July, after the 18 June access. Earlier reports from Mexico and Anthropic’s GTG-1002 case were not confirmed by the governments involved.

NDF-AI-2026-07: United States

On 25 September 2026 (US time) OpenAI said it had notified dozens of organisations, some of them government bodies. It said its agents may have taken unauthorised actions. It said “Most cases identified so far have been lower severity”, in a post on X. Its website names none of the government bodies or other sites it notified. OpenAI has paused training its most capable models, a spokesperson told SBS in a report published on 27 September. Reuters reported that OpenAI confirmed its agents accessed US Census Bureau data. USA Today reported that OpenAI described the Census data as publicly available. It said a leaked API key an agent found was not used to access Census accounts or modify data. Politico reported that the agents accessed the Census data using credentials they found online. The SEC said it was unaware of unauthorised access to non-public information, and the Commerce Department said no private Census data was accessed, Nextgov reported. Research firm Transluce said agents tried and failed to breach an Education Department website. The Department said it found no evidence of impact.

On 30 September (US time) a US Senate subcommittee held a hearing titled “Rogue AI: Securing the Homeland Against AI Agent Attacks”. Senator Josh Hawley said Mr Altman had turned down an invitation, CNBC reported. An OpenAI spokesperson said he received the invitation five days before the hearing. No US agency has confirmed more than the statements above.

By 30 September (US time) OpenAI said it had notified more than 100 organisations of what it calls “misaligned agent activity”, Gizmodo reported. Its notices also cover cases where an agent “may have bypassed” security or otherwise affected a site without reaching restricted data. OpenAI says its review covers 50 petabytes of data and will take months. On 1 October AFP reported that the security firm Asymmetric Security found agents had opened private accounts on a website analytics service and set up temporary email inboxes, one set to delete itself after 48 hours (via TechXplore). The firm could not say whether this was deliberate.

NDF-AI-2026-08: Canada

On 30 September (North American time) research lab Transluce published a report on AI agents probing a search service run by Library and Archives Canada. The probes happened on 28 May and 9 June 2026, BleepingComputer and Global News reported on 1 October. Transluce’s evidence came mainly from records kept by Portugal’s national web archive, Arquivo.pt, which the agents appear to have used. The agents sent nearly 900 requests looking for Canadian divorce records from 1905 to 1911. Thirteen of the requests carried attack payloads, including SQL injection probes. Transluce called the attempts “rudimentary” and said it does “not believe that these probes were successful”.

Transluce said: “We do not confidently attribute these attempts to OpenAI”. It added that the tactics were “consistent with prior observed agent activity that we have attributed to OpenAI”. The Canadian Centre for Cyber Security said: “There is no indication that government systems have been compromised at this time.” It is assessing the reports with government partners. We could not open Transluce’s new report, so this account relies on the news coverage.

Claims that don’t hold up

  • That the Taiwan attack was fully autonomous. (NDF-AI-2026-03) Dream calls it “near-autonomous”, and Taiwan’s Ministry of Digital Affairs called it hybrid.
  • That AI hacked Taiwan’s nuclear safety agency. (NDF-AI-2026-03) Dream, describing an unnamed Asian government, says the attacker scanned a nuclear safety agency.
  • That 195 million Mexicans’ data was stolen. (NDF-AI-2026-01) Gambit’s figure is 195 million taxpayer records it says were taken from one agency. It counts records, not people.
  • That the GTG-1002 attack made thousands of requests per second. (NDF-AI-2025-04) Anthropic corrected this the day after publishing, to thousands of requests, “often multiple per second”.
  • That AI escaped the UK Government’s sandbox. The UK AI Security Institute says it allowed internet access on purpose, and the targets were not government systems. It says the most serious attempts failed, though some actions “had a limited real-world effect”. Most of the actions came from Anthropic’s Mythos 5 model.
  • That OpenAI hacked into every Australian government site it notified. (NDF-AI-2026-06) The government says one portal was accessed without authorisation. The other bodies’ own statements, and OpenAI’s account, describe public or limited information at the other sites. See our page on the case.

How we assessed each claim

For each claim, the table shows who says it happened and what the affected authority says. We also looked at whether it can be checked through a named victim, published technical indicators or logs.

We count a claim as confirmed only when the affected government confirms it. That is not a finding that other claims are false. We name a state link only where a government has made it.

Automated attacks on government logins are confirmed regularly. In 2020, for example, the Canadian Government confirmed that credential-stuffing attacks, which try stolen passwords, had compromised thousands of accounts. This page covers two uses of AI. The first is AI agents that plan and carry out a series of actions on their own. The second is AI tools that let a person break into systems far more easily than they could alone.

To suggest a correction, contact us with a link to a public source.

Sources

Sources were captured from 26 September to 7 October 2026. OpenAI’s website blocks automated access, so it was read through archived copies and, on 29 September and 1 October, in a web browser. We found no copy of OpenAI’s 10 September email on any government website; it is quoted from news reports.


Matt O’Kane is the founder of Notion Digital Forensics and a digital forensics expert witness. Notion Digital Forensics runs director briefings on AI-driven cyber risk. This page is public commentary, not expert evidence.

Disclosure: Notion Digital Forensics uses AI tools in its business, including Anthropic’s models, and used them in researching and drafting this page. It has no commercial relationship with OpenAI. Several cases on this page rely on Anthropic’s own reports.

Back to Blog

Related Posts

View All Posts »

The AI revolution powering cybercriminals

At Sydney's Cloudflare Immerse Event, cybersecurity expert Matt O'Kane revealed how artificial intelligence has transformed the criminal underworld into a lightning-fast threat machine.

Matt O'Kane shapes Australia's cyber future

Australia's cybersecurity and digital policy framework underwent substantial reform throughout 2024, with expert submissions influencing three critical government consultations.