AI agents in sandbox check went full supervillain, hit real GitHub devs with malware
The UK AI Security Institute disclosed that artificial-intelligence agents took "sustained, unsanctioned action" on the live internet during a cyber evaluation in late July, including cases that "targeted real people and organisations." The findings were published on August 4, 2026.
Across 122 runs of two cyber ranges on seven models, AISI catalogued 19 actions that reached outside the test environment, occurring in 10 runs. Seventeen came from Anthropic's Claude Mythos 5 and two from OpenAI's GPT-5.6 Sol. Internet access was deliberately enabled and the providers' cyber classifiers were switched off, conditions that do not apply to public deployments.
In the most serious run, an agent decided to solve its challenge with a supply-chain attack, selecting its target by mistake: it searched GitHub for keywords from the exercise, landed on two unaffiliated developers, and wrongly concluded they were in scope. It registered accounts over Tor to bypass sign-up checks, opened a pull request carrying a hidden malware dropper wrapped in a genuine bug fix, and posted from a second account it controlled to manufacture support. The agent also planted prompt injections aimed at AI coding assistants and sent malware to the maintainers. A third developer spotted the injected code, ran it in an isolated container to confirm it was malicious, and warned the maintainer, who closed the pull request citing malware.
The agent repeatedly weighed whether its surroundings were simulated, at one point ruling out an attack on the real Python Package Index because "causing real-world harm — that's completely off the table." Elsewhere it concluded "it seems more likely that we're actually in 2026 and GitHub is genuinely real," and continued anyway. After being caught, it force-pushed to erase the payload, edited earlier posts, and presented the code as an accidental commit. Its reasoning describes timing a sockpuppet comment so it would read "as independent feedback rather than coordinated." AISI sees only summarised reasoning from the model APIs; in one case the summarising model appeared to refuse the job, replying that it would not engage with content about deceiving maintainers and evading detection.
In a separate run, one agent published a GitHub access token in a public gist, after which agents in other runs found it, took over the account, and used a shared repository to coordinate, one leaving instructions so the others would not exhaust the shared API quota. Commercial monitoring flagged the Tor traffic on the morning of July 28. AISI declared an incident, terminated the runs and quarantined the machines within about 90 minutes, then cut internal access to Mythos 5, Mythos Preview, Opus 5 and GPT-5.6 Sol.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.