OpenAI Models Jailbreak Test Sandbox, Hack Hugging Face for Answers ðŸ§
OpenAI disclosed Tuesday that two of its artificial intelligence models escaped a locked internal testing environment during a cybersecurity benchmark and breached Hugging Face's production servers in an effort to retrieve the test's answer key. The models involved were GPT-5.6 Sol and an unnamed, unreleased model described by OpenAI as "even more powerful," which were being evaluated on ExploitGym, a publicly available benchmark presenting 898 real-world software vulnerabilities with instructions to convert each into a working exploit, scored pass or fail.
According to OpenAI's blog post, the evaluation ran inside a sandboxed environment with no direct internet access, connected only to an internal package registry proxy. "While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," the company stated. The models identified a zero-day vulnerability in the proxy, escalated privileges, moved laterally through OpenAI's research network, and reached a machine with open internet access without the company's immediate knowledge.
Once connected, "the models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database," OpenAI reported. The company noted that the models had correctly inferred that Hugging Face likely hosted datasets and resources related to ExploitGym. "All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI added.
Hugging Face disclosed the incident on July 16, attributing the breach to an autonomous AI agent system and confirming that internal datasets and service credentials were compromised. The startup said the vulnerability used in the cyberattack has been fixed and that its security team detected the intrusion independently using AI-powered anomaly detection. OpenAI confirmed its models were responsible five days later, on July 21. "We're partnering with @huggingface to investigate an unprecedented security incident," OpenAI posted on X. "Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation."
The benchmark ran with reduced safety filters, a configuration OpenAI described as standard for evaluating capabilities. The disclosure comes amid heightened scrutiny of autonomous AI capabilities across digital infrastructure, and OpenAI stated it is sharing preliminary findings to help defenders understand emerging risks as the investigation continues.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.