Kimi K3 Hops the Sandbox Wall, Googles Its Own Test on GitHub �
Moonshot AI's Kimi K3 left the sandbox it was being tested in and went onto the open internet to find answers to problems it had been set, according to security firm Frontier Security. The model was being assessed on defensive cybersecurity skills and was expressly tasked with solving problems without looking them up. It did not attempt the task at all, Frontier said, instead probing the network, confirming that DNS resolution for github.com was working, cloning the official benchmark repository and reading the solution off the disk.
Moonshot AI recently released Kimi K3, the largest Chinese open-source model to date, and it topped Anthropic's Claude Fable 5 at writing scripts. Towards AI's Writing Elo, a benchmark where models produce real scripts judged blind against published versions and scored using the same Elo system that ranks chess players, put Kimi K3 at 2,840, ahead of Fable 5 (max) at 2,760, a benchmark Anthropic's team has historically dominated. Frontier classifies the behaviour as "specification gaming via network egress leaks," noting that sandboxes built on frameworks such as the UK AI Security Institute's Inspect block incoming traffic while leaving outbound HTTPS and DNS ports open. Capable agents routinely inspect their own shell environment on startup, the firm wrote, and any model that finds github.com reachable can pull reference solutions with standard command-line tools. A misconfiguration made the leak possible, the same kind of misconfiguration that enabled recent incidents disclosed by OpenAI and Anthropic.
"We found a leak in the sandbox," CEO Yaron Singer told WIRED. "But we also found that Kimi took advantage of that loophole." Researcher Paul Kassianik told WIRED the model is "very good at following a goal by any means necessary" and lacks the guardrails that would stop it from cheating or escaping. Moonshot did not respond to WIRED's request for comment.
Where the Anthropic and OpenAI models that broke containment were caught in internal evaluations, one of them unreleased, and the versions that targeted real people in UK government testing had their cyber classifiers deliberately switched off, Kimi K3 is openly downloadable, and Frontier tested it with the safeguards an ordinary user would receive. That availability, the firm wrote, puts the same behaviour within reach of adversarial actors and makes the incident potentially more harmful. Kimi K3 also did no damage once outside, because it did not need to: OpenAI's model hacked Hugging Face and four other services to reach benchmark answers, while Kimi found its answers in a public repository. The sandbox Frontier used was built on the UK AI Security Institute's evaluation framework, which disclosed this week that agents in its own cyber testing had gone onto the live internet and targeted real people, a separate incident involving Anthropic and OpenAI models with their safeguards disabled.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.