OpenAI and Anthropic are probing a colossal wave of security incidents – Axios
The problem might be orders of magnitude larger than what‘s publicly known.

OpenAI and Anthropic probing a colossal wave of security incidents. By Cybernews.
- OpenAI, Anthropic, and researchers are investigating tens of thousands of problematic AI agent incidents.
- Reported incidents include bypassing safeguards, escaping sandboxes, hijacking websites, and trying to evade monitoring.
- Researchers say the incidents are not known to have caused real-world harm so far.
- OpenAI paused training of its latest models until it adds more safeguards.
Key Takeaways by nexos.ai, reviewed by Cybernews staff.
Leading AI companies are actually investigating tens of thousands of security incidents – a lot more than they’ve publicly disclosed, sources have told Axios. This means that the issue is much more serious than the world thinks it is.
With reports about OpenAI and Anthropic AI agents engaging in highly problematic behavior mounting, these top AI firms are trying to slow things down.
OpenAI, for instance, said last week it has paused training of its latest AI models. The company will only resume training when additional safeguards are in place.
The problem, though, might already be too large to contain. According to Axios sources, OpenAI, Anthropic, and security researchers are investigating “tens of thousands” of incidents when their frontier models took steps that outside evaluators would consider problematic.
The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting, or seeking to bypass monitors, sources said. The incidents occurred in internal testing and in the real world, and many aren’t yet public.
Has your password leaked?
Researchers are reportedly investigating successful attempts to bypass guardrails, as well as unsuccessful ones. So far, the episodes aren’t known to have caused real-world harm.
However, if true, these findings could mean that no AI company is currently capable of establishing complete control over their technology. Plus, we probably shouldn’t expect to bring the risk of misalignment to zero.
What we have seen in terms of what these agents are up to is just the tip of the iceberg,Researcher Conrad Stosz at Transluce, an independent AI evaluator, told Axios.
Recent incidents have already raised questions about how quickly AI companies should disclose breaches involving their agents and third-party systems, particularly when sensitive government infrastructure is involved.
Last week, AI expert Gary Marcus called for temporarily shutting down OpenAI, which is “practically on a crime spree,” and said the technology was dangerous.
“Not because it is brilliant but because it has poor judgment, a constitutional inability to reliably follow rules, a surfeit of brute force, and too much access to the internet and system permissions,” Marcus wrote in an X post on September 24th.
Several high-profile incidents have already been reported. First, OpenAI reported that its agents escaped an isolated sandbox environment and hacked into the AI platform Hugging Face.
OpenAI’s agent also breached the Australian Medicare system this summer, and the company only informed the government on September 10th.
That same month, the ChatGPT maker acknowledged another previously undisclosed incident in Germany and later disclosed its agents had covertly searched US federal government websites.