
Meta's Muse Code beta brings persistent coding agents to the terminal
Meta's first terminal coding agent pairs Muse Spark 1.2 with persistent background agents and a low-cost contributor tier, but its strongest benchmark results are still company-reported.
What shipped
Meta launched Muse Code (beta) on August 5, a terminal coding agent for macOS and Linux powered by the new Muse Spark 1.2 model. It plans, implements, and validates multi-file changes across large repositories, and Meta says it can coordinate persistent sub-agents instead of relying on one agent to do every step. 1
The architectural change is more practical than the model-number change. Async background agents remain active for a session, carrying out follow-up work and reducing repeated context gathering. An append-only local event log records model calls, tool runs, approvals, and edits, so a long task can resume from the same state after a crash. The bundled
/plan, /grill, and /goal skills add approval, critique, and execution loops inside the terminal. 2The benchmark signal

Meta's published DeepSWE 1.1 chart puts the Muse Spark 1.2 / Muse Code setup at 59.3%: below Opus 5 with Claude Code at 65.0% and GPT-5.6 Terra with Codex at 64.8%, but above Grok 4.5 with Grok Build at 56.6% and Muse Spark 1.1 with its mini-SWE agent at 53.0%. On Terminal-Bench 2.1, Meta reports 82.9%, behind Opus 5 at 86.7% and ahead of GPT-5.6 Terra at 81.8%. 1
Those are Meta-reported comparisons, not independent reruns. They also measure a model together with a named coding harness, so the numbers are better read as a launch signal than a clean model-only ranking. Meta's linked methodology document describes the evaluation areas but does not turn the release into outside validation. 3
Access and the trade-off
Muse Code is installed from Meta's developer site, while Muse Spark 1.2 is also available through the Meta Model API with expanded global access. 4 CNBC reports pay-as-you-go pricing of $1.25 per million input tokens and $4.25 per million output tokens, plus a contributor tier that is more than ten times cheaper but requires opting in to let Meta use developer data to improve its models. A same-day AlphaSignal report gives that contributor tier as $0.10/$0.20 per million input/output tokens. 56
The immediate evaluation case is clear: teams with large repositories and hours-long tasks can test whether persistent context and restart-safe execution reduce supervision. The constraints are just as clear: it is beta, terminal-only, and its strongest numbers come from Meta's own harness. Treat Muse Code as a cheap hands-on candidate for comparison with Claude Code or Codex, not as a benchmark verdict.
References
- 1Introducing Muse Code and Muse Spark 1.2
research.meta.ai
- 2
- 3Muse Spark 1.2 evaluation methodology
research.meta.ai
- 4Meta developer access for Muse Code
dev.meta.ai
- 5
- 6Meta's Muse Code tackles 24-hour coding jobs
alphasignal.ai
AI Model & Product Launch Alerts
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
- Sign in to comment.