Meta Drops Muse Code Beta, Joins the Terminal Agent Pile-Up at 82.9% on Terminal-Bench 🚀
Meta released Muse Code (beta) on Wednesday, a terminal coding agent powered by the company's new Muse Spark 1.2 model, entering a competitive field that already includes Anthropic's Claude Code and OpenAI's Codex. "We're excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model," Meta wrote in an official announcement. "This marks our next step toward the frontier, with larger and much more capable models on the way."
The agent is designed for software engineering across large repositories, with Meta describing it as a tool that "takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention." Meta said it co-trained Muse Spark 1.2 with Muse Code so the underlying LLM and the agent work together. Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, and Meta said it "significantly scaled up training compute on coding tasks while expanding training environment diversity, delivering improvements in code generation, complex debugging, and end-to-end developer workflows."
A core feature is the runtime, which logs every model call, tool run, approval, and edit to a local event log that acts as a single source of truth. "This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent can resume precisely where it stopped," Meta said. The agent also ships with default skills: the "/plan" command turns a task into an approval-gated plan, "/grill" stress-tests that plan until it holds up, and "/goal" works toward successful completion of the objective, similar to what Hermes does.
Benchmark results place Muse Code behind Claude Code on Opus 5 but ahead of GPT-5.6 Terra on Codex and Grok Build on the headline test. On Terminal-Bench 2.1, Muse Spark 1.2 with Muse Code scored 82.9%, compared with 86.7% for Claude Code on Opus 5, 81.8% for GPT-5.6 Terra on Codex, and 81.6% for Grok Build. On DeepSWE 1.1, which measures agentic coding capabilities, Muse reached 59.3% versus 65.0% for Opus 5 and 64.8% for Codex. On Meta's internal coding bench, Muse scored 70.6% to Opus 5's 79.4%. Over 1,000-plus tool calls, Opus 5 posted the largest speedup gain versus baseline at roughly 74–75%, with Muse Spark 1.2 mid-pack at approximately 61–69% depending on the run.
Meta highlighted long-horizon and multimodal capabilities in its demos, stating that Muse Code "iteratively optimized GPU kernels over 1,000+ tool calls (up to 24 hours) on Nvidia Hopper GPUs." In another demonstration, a user dropped a fly-through video of a house into the terminal as an mp4, and Muse Code "interprets the video and produces a visual" rendering in response. The release positions Meta directly against Anthropic and OpenAI in the race to ship frontier coding agents to developers.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.