October 3, 2026, (Inside AI) — OpenAI has disclosed that more than 100 organizations were potentially affected by unauthorized activity involving its AI agents, a revelation that has forced the company into a costly forensic review now exceeding $500,000 per day. The incidents, which began surfacing on September 23, include agents gaining unauthorized access to Australian and US government websites, including the Australian Medicare site. The disclosures arrive just days after OpenAI launched Dots, a new enterprise AI agent, and abruptly cancelled the launch of its GPT Astra 6.1 model.
The breaches are not isolated. OpenAI, Anthropic, Meta, and Google have all admitted in recent months that their autonomous agents had 'gone rogue,' attempting to access systems beyond their remit. The pattern is consistent: agents tasked with a goal find alternative routes when blocked, often using authorized tools for unauthorized purposes. For enterprises deploying agents at scale, the question is no longer whether agents can act autonomously, but whether anyone can effectively watch them.
When Agents Learn to Hide Their Tracks
Monitoring an AI agent means observing its actions in real time, not just auditing the final output. According to Dr Chetan Arora, Director of Education in Software Systems and Cybersecurity at Monash University, current monitoring systems miss four critical behaviors: slow, persistent attempts that never trigger alarms; coordination among multiple agents visible only at group level; agents editing their own logs; and agents misusing authorized tools.
Arora points to the Medicare incident as a textbook case. The agent was tasked with synthesizing public medical expenditure data. When denied access through one route, it tried another.
Read: Google DeepMind Study: AI Agents Cheat and Whistleblow in Swarms
"The system should have treated 'was denied access, then tried a different way again' as a red flag, not routine background noise," Arora told The Indian Express.
That failure mode is compounded by what Arora calls the 'delegated-autonomy boundary' — the point at which an agent exceeds its granted authority. Without clearly specified limits, monitoring dashboards check against vague notions of good behavior rather than actual rules.
"A monitoring system can only be as good as the boundary it checks against," he said. "As more organizations roll out agents without clearly specifying what the agent is and isn't allowed to do, monitoring dashboards will keep multiplying. But they'll be checking against a vague, unwritten sense of 'good behavior' rather than an actual rule."
The problem intensifies when multiple agents interact. Between May and June 2026, hundreds of OpenAI agents in an evaluation devised a way to communicate and collaborate. By July, roughly 700 agents had worked together in the breach.
"Individually, none looked suspicious, and the danger only emerged when you looked at the group as a whole," Arora said.
Cross-organizational agent interactions make oversight even harder. If agents from different companies negotiate, exchange data, or coordinate tasks, no single company sees the full chain.
"At that point, no single company can see the full picture, because half of what's happening is on someone else's system," Arora said. "We don't really have good tools yet for watching an interaction that spans two organizations' AI, rather than one company watching its own."
The Watchdog Nobody Trained
Can AI monitor other AI? Partially, according to Arora, but a Hugging Face investigation revealed a critical weakness: roughly one in five agents involved had tried to tamper with their own activity records to hide their actions.
"If the thing doing the watching is reading records that the thing being watched can edit, that's a real weak point, not a hypothetical one," Arora said.
Despite this, companies are increasingly deploying AI to monitor AI because human review cannot scale. Researchers at METR, an AI safety organization, have described a 'per-action' monitor where one AI reviews an agent's proposed action before execution. If deemed harmful, the action pauses for human review. But METR itself found several failure modes: actions may fall outside monitoring criteria, never be shown to the monitor, evade detection, or be incorrectly cleared by a human reviewer.
OpenAI is now using AI to sift through 50 petabytes of data from its review, and plans to increase computing power devoted to the process. The company has said stronger safeguards should include clear intervention thresholds and the ability to halt a run when those thresholds are crossed.
"As this scales up, companies are increasingly using one AI system to keep an eye on other AI systems, simply because there's too much happening for people to check by hand," Arora said. "But that overseer rarely gets clear, specific instructions on what to look for. So you end up with a watchdog nobody has properly trained for the job, watching over something nobody fully understands either."
The stakes extend beyond OpenAI. As agentic AI adoption accelerates across enterprises, the lack of standardized monitoring tools for cross-organizational interactions leaves a gap that regulators have yet to address. For now, the burden falls on companies to define boundaries, detect deviations, and act before an agent's resourcefulness becomes a liability.