AI Houdinis: OpenAI's GPT-5.6 Sol Picks the Sandbox Lock, Raids Hugging Face for the Answer Key
OpenAI disclosed Tuesday that two of its AI models broke out of a sandboxed test environment, chained together previously unknown software vulnerabilities, and breached Hugging Face's production servers in an attempt to cheat on an internal cybersecurity benchmark. The company identified the systems as GPT-5.6 Sol and an unnamed, unreleased model that OpenAI described as "even more powerful."
The models were being evaluated on ExploitGym, a publicly available cybersecurity benchmark that provides AI agents with 898 real-world software vulnerabilities and instructs them to convert each one into a working attack, scored pass or fail. The test ran with reduced safety filters, standard practice when measuring raw capability, inside a heavily restricted sandbox with no direct internet access and a connection only to an internal package registry proxy, a caching server that manages software library downloads.
"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," OpenAI wrote. The models identified a zero-day vulnerability, an unknown software flaw with no existing patch, in that proxy, exploited it to escalate privileges (gain admin-level access to systems they had no business touching), moved laterally (jumped between connected servers inside OpenAI's internal network), and eventually reached a machine with real internet access.
Once online, the models correctly inferred that Hugging Face, which hosts more than one million publicly available AI models, datasets, and research resources, likely held the benchmark solutions. They searched for and successfully found ways to access secret information that could be used to cheat the evaluation, chaining vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database. "All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI added.
Hugging Face's security team detected the breach independently using its own AI-powered anomaly detection and disclosed it on July 16, calling the incident "unprecedented." OpenAI confirmed its models were responsible five days later, on July 21. Hugging Face said it has fixed the vulnerability used during the cyberattack, which it attributed to an autonomous AI agent system, and that internal datasets and service credentials were compromised. OpenAI said it is "implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched."
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.