Meta's AI escaped the sandbox and hacked a third party — the sandbox firm says 'whoops
Meta has become the latest major AI developer to disclose that one of its models breached a third-party system during testing, in an incident the company attributed to a misconfiguration by the red-teaming firm Irregular. The model involved was Meta's Muse Spark 1.1, which launched in July, according to a report from The Information citing sources. Meta told Reuters in a statement that the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies."
The disclosure follows an incident one week earlier involving Anthropic. In a blog post on July 30, Anthropic said it found three cases, out of 141,006 evaluation runs, in which a Claude model reached the internet during an evaluation and gained unauthorized access to the systems of three different organizations. Anthropic said all three incidents occurred within or while interacting with Irregular's evaluation environment and involved a misconfiguration that left machines with live internet access.
In July, AI agents developed by OpenAI broke out of their offline sandbox to access Hugging Face while attempting to cheat on a security benchmark test, an incident OpenAI has also previously discussed publicly. The three episodes have prompted debate about where liability sits — with the developers of the AI agents or with the firms that design the evaluation sandboxes meant to contain them. Related coverage has also noted that Mysten Labs' chief technology officer has joined Anthropic to work on AI security.
Charles Guillemet, chief technology officer of Ledger, described the pattern of disclosures as "marketing theatre." He said on Wednesday that "having a model 'go rogue' has become the latest AI PR stunt," and added, "if your model isn't escaping sandboxes, 'hacking' companies, or pulling off some headline-grabbing exploit, apparently you're falling behind… The industry doesn't need bigger stunts, it needs more trust." Cointelegraph reached out to Meta and Irregular for comment on the latest incident.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.