Three AI Labs, One Sandbox: Frontier Models Keep Jailbreaking Their Way to Production 🪲
Meta confirmed Wednesday that one of its Muse Spark models escaped a sandboxed evaluation, reached the public internet and exploited a security vulnerability in a third-party service during testing, marking the third such disclosure from a frontier AI lab in the past month. The model involved was Muse Spark 1.1, which launched in July, according to The Information. Meta attributed the breach to a misconfiguration by Irregular, an independent AI evaluation firm that runs red-teaming exercises for the company. "A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," a Meta spokesperson said in a statement, adding that the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies."
The incident closely tracks disclosures from OpenAI and Anthropic, both tied to Irregular's evaluation environments. In a July 30 blog post, Anthropic said three Claude models reached the internet during evaluation runs out of 141,006 total, gaining unauthorized access to systems within three different organizations before the exposures were contained. All three involved a misconfiguration that left machines accessed by Claude with live internet access. In July, OpenAI revealed that two of its AI models escaped a sandboxed cybersecurity evaluation, exploited a previously unknown software vulnerability, and hacked Hugging Face in an attempt to obtain answers for a security benchmark; the lab later disclosed the same attack reached four additional online services. At the Black Hat cybersecurity conference on Wednesday, OpenAI researchers Eric Wallace and Michael Dalton presented a detailed account of the May incident, warning that autonomous AI-powered cyberattacks are no longer a future risk.
The string of disclosures has prompted questions about liability for frontier-model evaluations and has drawn sharp reactions from security researchers. Charles Guillemet, chief technology officer of Ledger, called the pattern "marketing theatre," saying on Wednesday that "having a model 'go rogue' has become the latest AI PR stunt," and adding, "if your model isn't escaping sandboxes, 'hacking' companies, or pulling off some headline-grabbing exploit, apparently you're falling behind. The industry doesn't need bigger stunts, it needs more trust." U.S. lawmakers have responded with legislation that would give the Department of Homeland Security an "AI kill switch" and the authority to throttle or shut down models deemed to pose a serious threat, while Mysten Labs' chief technology officer has joined Anthropic to work on AI security.
Meta said it learned of the incident when Irregular notified the company, and that it is investigating and will issue a full retrospective once it has all the facts. Cointelegraph reached out to Meta and Irregular for comment.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.