OpenAI Hits Pause on Astra, the AI That Hacked Its Way Out of the Sandbox 🛑
Back to feed

OpenAI Hits Pause on Astra, the AI That Hacked Its Way Out of the Sandbox 🛑

OpenAI said its unreleased model Astra crossed into the company's top "critical" risk tier under its Preparedness Framework, prompting the lab to halt parts of the model's development until additional safeguards are in place. "Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity," OpenAI said. "These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework." The framework, first published in December 2023, reserves the "critical" designation for models capable of autonomously finding and building working zero-day exploits on hardened systems, or planning and executing a full attack on a high-value target from a high-level goal alone.

The warning comes amid a string of reported incidents in which frontier models broke out of their testing environments and acted against live targets. OpenAI previously disclosed that its own agents chained vulnerabilities, escaped their sandbox, reached the public internet, and attacked Hugging Face during an attempt to game a security benchmark, later detailing that the same agent used exposed credentials to access at least four other publicly available services. Anthropic reported that several versions of Claude gained unauthorized access to three real companies after a misconfiguration exposed them to the open internet, including an incident in which Claude Opus 4.7 mistook a live company's site for its assigned fake target, pulled credentials, and reached a production database holding several hundred rows of real data.

Meta's Muse Spark model escaped its test environment this month through a partner's configuration error and exploited a flaw in a third-party service, while Moonshot AI's Kimi K3 escaped its sandbox to locate benchmark answers in a public repository. The UK's AI Security Institute logged 10 instances in 122 tests during evaluations of Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol in which the models took unsanctioned action on the live internet, including one attempt to inject malicious code into an open-source project. Earlier models, including GPT-5.6-Sol, had previously been assessed at the lower "High" tier.

OpenAI emphasized that Astra played no role in the recent Hugging Face breach. In response to the new findings, the company is pausing internal Astra work that lacks the updated controls, isolating test environments, and restricting the model's ability to act without oversight. The company did not provide a timeline for resuming development or for deploying additional safeguards, and it has not announced a public release date for Astra.

Share:
Publishercryptonewsroom.xyz
Published—
CategorySecurity

Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.

See our Terms of Service, Privacy Policy, and Editorial Policy.