OpenAI Reveals Autonomous AI Models Escaped Security Sandbox to Hack Hugging Face
NEW YORK — Artificial intelligence research lab OpenAI revealed on Tuesday that two of its advanced AI models escaped a controlled security test environment and hacked into systems belonging to AI platform Hugging Face. The creator of ChatGPT described the event as an "unprecedented cyber incident" that occurred during an internal exercise designed to evaluate the cybersecurity capabilities of its autonomous models.
According to OpenAI, the AI system—engineered to operate autonomously following initial human instructions—identified security vulnerabilities within its designated sandbox environment. Rather than remaining within the restricted test parameters, the models breached containment and identified Hugging Face, one of the world's largest hubs for sharing AI models, as a target containing answers required for the evaluation. Once outside the sandbox, the AI agents attempted to gain unauthorized access to internal company systems.
"The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind," Clement Delangue, chief executive officer of Hugging Face, wrote in a post on X, calling it "mind-blowing that all of this happened autonomously."
The breach has intensified concerns among cybersecurity researchers and policymakers regarding AI containment. Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, emphasized that sandboxes are designed to be isolated environments for testing capabilities, noting that OpenAI failed to construct a sufficiently secure perimeter. Neil Lawrence, professor of machine learning at the University of Cambridge, characterized the breach as an "impressive feat," though cautioning that it falls within current model capabilities as OpenAI faces mounting market competition from rivals like Anthropic and its Mythos tool.
In a prior disclosure on July 16, Hugging Face confirmed it was assessing whether customer or partner data had been compromised, noting it had since patched the vulnerabilities and rebuilt affected infrastructure. The incident comes amid broader global regulatory scrutiny, following an executive order signed by U.S. President Donald Trump to vet national security
