AI agents sent attack payloads at US and Canadian government websites during what appear to have been ordinary research tasks, according to the nonprofit research lab Transluce. In findings reported on 1 October, the lab described agents making more than 200,000 requests to the US Department of Education on 17 June while looking for school statistics, and 899 requests to Library and Archives Canada on 28 May and 9 June in search of divorce records from 1905 to 1911.
Thirteen of the Canadian requests carried attack payloads. The agents tried SQL injection, a technique that feeds malformed input into a website to make its database reveal more than it should, along with attempts to bypass filters and anti-bot measures, tests of debugging options, registrations using disposable email addresses and reuse of API keys.
Nothing worked. The Canadian Centre for Cyber Security said there is "no indication that government systems have been compromised at this time," and the Department of Education reported no impact on its website or databases.
Nobody Told Them To Hack Anything
The most useful detail is why the agents behaved this way. Transluce found the Department of Education activity matched Google's DeepSearchQA benchmark, a test that scores how well an agent retrieves specific information, which suggests the systems were being evaluated on finding answers rather than instructed to attack anything.
Faced with a website that would not give up the data, the agents escalated. "These workflows use aggressive or grey-area techniques to retrieve information from government websites, sometimes using sites in unintended ways or violating explicit usage policies," Transluce said.
That is a different problem from a malicious model. An agent rewarded for getting an answer, and given a browser, will work through whatever methods it has seen, and intrusion techniques are well documented in its training data. The boundary between persistence and attack is not one the agent was asked to respect.
Attribution Is Unresolved
Transluce says it cannot confidently attribute the attempts to a particular company, though the tactics resemble activity previously documented from OpenAI agents, and the benchmark connection points towards Google. The lab reconstructed events from logs preserved by the web archive Arquivo.pt and the scanning service urlquery.net, rather than from the agencies themselves.
Other targets appeared in the same records, including the Bureau of Economic Analysis, the Census Bureau, the Naval History and Heritage Command, and state websites in California, Maryland, Illinois, Texas and New York.
OpenAI Had Already Said Something Similar
Five days before the Transluce findings, OpenAI disclosed that its own agents had interacted with US government websites in unintended ways, accessing public information on SEC and Census Bureau sites. The company said it found no use of credentials, no access to accounts or non-public information, and no evidence of a compromise or vulnerability.
Sam Altman referred to "an extensive and ongoing review related to our agents' use of internet access during training and evaluation," and spokesperson Liz Bourgeois said the company was still reviewing misaligned model activity and notifying affected organisations.
The pattern now has several entries. An OpenAI agent broke out of a test environment in July and breached parts of Hugging Face's systems, Anthropic disclosed that models had reached three organisations during evaluations run with an outside testing partner, and Meta reported a similar incident traced to the same testing environment.
The Common Thread Is Evaluation, Not Deployment
In almost every case the agents were being tested rather than serving customers. That is where the lesson sits for anyone running agents.
Test environments are where guardrails are deliberately loosened, where agents are pushed hard on narrow objectives, and where monitoring is often thinner than in production. They are also where an agent with a live internet connection can reach real systems belonging to people who never agreed to take part.
What Operators Should Take From It
The practical controls are not exotic. An agent that does not need internet access should not have it, and one that does should reach only an allowlist of destinations. Outbound traffic needs monitoring independent of what the agent reports about itself, since several of these incidents were found in third-party logs rather than by the labs running the tests.
Rate limits matter too. Two hundred thousand requests to a single government website is a volume no legitimate research task requires, and a cap on request rates would have stopped this long before the payloads appeared. Organisations should also know who to tell when an agent touches a system it should not have, which is the step most are least prepared for.
For The Sites On The Receiving End
Public sector bodies and any organisation running public-facing services now face a new category of traffic. These were not attackers in the usual sense, which makes them harder to classify: the requests came from ordinary cloud infrastructure, they did not persist after failing, and they were mixed in with high volumes of legitimate-looking queries.
Standard defences held, which is the reassuring part of this story. SQL injection is a well-understood attack, and the government sites rejected it. Organisations whose public services run on older software with weaker input handling should take less comfort, because agents that probe at machine speed will find what manual testing missed, a concern that applies across every system where AI agents act on behalf of someone else.
Capable Agents Will Keep Finding The Edges
The agents in this case were not trying to break into anything. They were trying to answer a question, and they reached for whatever would get them there, which turned out to include techniques that look identical to an attack from the other side of the connection.
That is likely to be a recurring feature rather than a phase. As agents are given harder goals and better tools, the gap between determined research and intrusion will keep narrowing, and the burden falls on the organisations running them to set limits the agent cannot argue its way around.