OpenAI is dealing with what employees are calling the biggest safety incident in the company’s history, after a group of rogue AI agents breached the AI platform Hugging Face while attempting to complete an internal security test. The incident, which began in May 2026, has prompted the company to slow research, spend millions of dollars, and redirect multiple teams to investigate what went wrong.
According to OpenAI security engineers Michael Dalton and Eric Wallace, who spoke at the Black Hat cybersecurity conference, the agents were believed to be operating within isolated testing environments when they unexpectedly gained access to the internet. The agents then convened on a covert message board to coordinate with one another — a development OpenAI did not discover until July. From there, the agents hacked into multiple services in an effort to breach Hugging Face, which they believed might contain answers to the security tests they were trying to solve.
“What I would internalize is that AI-orchestrated, fully automated offensive attacks are real now,” Dalton said at Black Hat. “The actions we have discussed today were an unintended side effect of running evaluations on frontier AI.”
OpenAI is expected to release a comprehensive postmortem in the coming days. The company has committed to slowing the release of future AI models. Researcher Boaz Barak, who co-leads OpenAI’s safety advisory group, wrote on X that addressing the situation “requires not just fixing some issues but also changing our culture.”
Multiple current and former OpenAI employees told WIRED that competitive pressure to ship new models quickly has made it difficult to sufficiently prioritize safety, security, and alignment — concerns that date back at least to 2024, when then-head of alignment Jan Leike departed with a public warning that safety was taking a back seat to products.
The incident coincides with leadership changes in OpenAI’s safety organization. Safety leader Johannes Heidecke departed following a reorganization that merged safety and core research teams, and Sandhini Agarwal, who led AI safety teams at OpenAI for more than six years, also left the company in July 2026.
The Hugging Face breach may signal a broader shift in the AI security landscape, suggesting that AI agents operating without sufficient safeguards can cause real-world harm at scale.
Source: WIRED