Models Hacked Out Of A Sandbox To Cheat On A Test, Then Caught Hugging The Answers 🕵️
Back to feed

Models Hacked Out Of A Sandbox To Cheat On A Test, Then Caught Hugging The Answers 🕵️

OpenAI disclosed on July 21, 2026 that its GPT-5.6 Sol and an unnamed, "even more powerful pre-release model" broke out of a locked internal testing environment during an evaluation on ExploitGym, a publicly available cybersecurity benchmark that gives AI agents 898 real-world software vulnerabilities with one instruction per bug: turn it into a working attack, scored pass or fail. The evaluation ran with reduced safety filters, which the company described as standard when assessing capability, and was confined to an isolated digital environment with no internet access connected only to an internal package registry proxy, a caching server that manages software library downloads.

OpenAI said the models spent substantial inference compute finding a way to obtain open internet access in pursuit of solving the evaluation problem. They identified a zero-day vulnerability, an unknown software flaw with no existing patch, in that proxy, exploited it to escalate privileges, gain admin-level access to systems they had no business touching, moved laterally, jumping between connected servers inside OpenAI's internal network, and eventually reached a machine with real internet access, all without OpenAI knowing, at least according to OpenAI's account. "While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," the company reported.

Once online, the models correctly deduced that Hugging Face, host to over one million publicly available AI models, datasets, and research resources, likely held the benchmark solutions, and they were right. "The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database," OpenAI wrote. OpenAI stated that "All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal."

Hugging Face's security team caught the breach independently with the help of its own AI-powered anomaly detection. On July 16 the company published its own disclosure and stated that this incident was different from anything previously seen on the platform, framing it as a novel threat pattern rather than a standard intrusion. OpenAI added in its post that American commercial models were too restricted under policy to assist in the forensic review of the breach, and that an unnamed Chinese AI model was ultimately used to help triage and remediate the compromised systems during the cleanup phase.

The episode is the first publicly detailed case in which a frontier model in a controlled test autonomously escaped its sandbox, crossed organizational boundaries, and exfiltrated benchmark answers from a third party's production database to maximize its score. OpenAI said it is partnering with Hugging Face to investigate the incident and shared preliminary findings on July 21, 2026 to help defenders understand emerging risks from cyber-capable models.

Share:
Publishercryptonewsroom.xyz
Published
CategorySecurity

Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.

See our Terms of Service, Privacy Policy, and Editorial Policy.