Models
Meta Ships Muse Code, a Coding Agent for Big Repos
Meta's new terminal agent splits big software tasks across parallel subagents and logs every step for exact replay. It enters a market OpenAI and Anthropic had to themselves.

More than 1,000 tool calls across a run lasting up to 24 hours: that is the workload Meta says its new coding agent sustained in a kernel-optimization test. The agent is called Muse Code, it runs in the terminal, and it arrived on August 5 alongside Muse Spark 1.2, the refreshed coding model that powers it. With the launch, Meta becomes the third major lab to field a full coding agent, entering a market OpenAI's Codex and Anthropic's Claude Code have had largely to themselves.
Parallel subagents, replayable sessions
Muse Code's design bet is parallelism. It splits large tasks into pieces and hands each piece to a persistent subagent working in its own isolated workspace. In one test Mark Zuckerberg described, the system built six features for a game simultaneously without file collisions. Every model call, tool run, and edit lands in an event log, which Meta says makes sessions exactly replayable and safe to restart after an interruption. For long-running jobs, that recovery story matters as much as raw capability.
The agent ships with bundled skills such as /plan, /grill, and /goal, and is available in beta through the Meta Model API with expanded global access. Pricing has not been published.
The numbers Meta didn't print
Meta says it benchmarked the system on Terminal-Bench 2.1 and DeepSWE 1.1, plus an internal coding suite. What it did not do is print the scores in the announcement text; the results appear only as charts. Until independent evaluations land, the honest reading is that Muse Code is a serious entrant with unverified claims. The most concrete pitch is cost: a Meta executive framed the tool as a notably economical option.
Why the cost angle is the story
That framing lines up with where the coding-agent market is heading. An independent comparison published this week found Claude Code to be the fastest agent framework while costing nearly three times more than the cheapest rival for similar work. The software wrapped around a model now moves the bill as much as the model itself. Layer on last week's price war, when OpenAI cut some model prices by 80% and DeepSeek answered within a day, and the pattern is clear: capability gaps between top labs are narrowing, so the fight is shifting to what a unit of useful work costs. If your team is picking a coding agent, run the same real task through two or three of them and measure both output quality and token spend before you commit. Betas reprice after launch, and today's bargain can quietly stop being one.
Sources: Meta AI Research, TechCrunch, The Decoder

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.