~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / muse-code-meta-terminal-bench $
News

Muse Code: Meta Beat Codex on Terminal-Bench — and Charges $0.30 to Read Your Repository

LS Lucas Souza · · 11 min read
Muse Code: Meta Beat Codex on Terminal-Bench — and Charges $0.30 to Read Your Repository

Meta's Muse Code just joined the terminal agent fight. And it came in swinging on price.

On August 5, 2026, Meta Superintelligence Labs shipped Muse Code in beta: Meta's terminal agent running the Muse Spark 1.2 model, built to go straight at Claude Code and Codex territory. You install it with a curl, it runs on macOS and Linux, and there's no GUI and no IDE extension. It's pure terminal.

The number that made the rounds on Twitter was Terminal-Bench: 82.9%. Above GPT-5.6 Terra. Plus the "contributor" tier pricing — $0.10 per million input tokens and $0.20 per million output — which makes Muse Spark 1.2 up to 20x cheaper than the competition.

Except both of those sentences are hiding something. Let's crack them open.

TL;DR

  • What it is: Muse Code, Meta's terminal agent in beta, running the proprietary Muse Spark 1.2 model.
  • Stack/Models: Muse Spark 1.2 (closed-weight), 1 million tokens of context, sub-agents in isolated git worktrees, append-only event log.
  • Cost/Access: Standard tier $1.25/M input and $4.25/M output. Contributor tier $0.10/M input and $0.20/M output — with Meta training models on your code.
  • The benchmark: 82.9% on Terminal-Bench 2.1. Beat GPT-5.6 Terra (81.8%). Lost to Claude Opus 5 (86.7%) on all three benchmarks Meta itself chose to show.
  • Useful link: dev.meta.ai

The context: why Meta needed Muse Code

Until this week, the terminal agent race was Anthropic versus OpenAI. Claude Code on one side, Codex on the other. Meta was on the sidelines — and not for lack of money.

Muse Code is the first serious attempt to change that, and it comes with a shift in posture worth paying attention to: Muse Spark is closed-weight. No weights on Hugging Face, no permissive license, no "download it and run it locally." After years of building a reputation on open Llama, Meta shipped its coding model closed. Zuckerberg left the door open to opening it up later, but today it isn't open.

The choice makes sense once you look at the product. Muse Code isn't a chat that spits out a function. It was designed for the kind of task that breaks most tools: big repository, changes across multiple files, work that takes hours instead of seconds. Meta showed off a kernel optimization case on NVIDIA Hopper GPUs with more than a thousand tool calls over up to 24 hours.

That's the part that matters if you build things: the fight moved from "which model writes the prettiest function" to "which harness can run for 24 hours without losing the plot." It's an engineering problem, not a model problem.

And that's where the annoying part lives. Every week a new agent shows up promising to replace the last one, and testing each of them on a real repository is the work nobody has time to do alone — which is exactly what we do together, live, every week in the Clã Beer and Code. It's paid, it's a subscription, and it's the environment this post describes.

The Muse Code benchmark: Meta beat Codex, but which Codex?

This is where the headline needs surgery.

The chart Meta presented at launch has these numbers on Terminal-Bench 2.1 (89 tasks, pass@1 over five attempts):

Model Terminal-Bench 2.1
Claude Opus 5 86.7%
Muse Spark 1.2 82.9%
GPT-5.6 Terra 81.8%
Grok 4.5 81.6%

Yes, Muse Spark 1.2 landed above an OpenAI model. It landed 1.1 points above GPT-5.6 Terra.

The problem: Codex doesn't run Terra. Ever since GPT-5.6 went generally available on July 9, 2026, Codex has defaulted to GPT-5.6 Sol — which on the independent Artificial Analysis leaderboard scores 89.5% on the same Terminal-Bench v2.1, the top of the table. Claude Opus 5 at max effort shows up right behind it, at 89.1%.

In other words: Meta chose to compare against the middle model in OpenAI's lineup, not the one the competing product actually uses. The Register noticed the same thing and described Zuckerberg's chart as "vague," with tightly packed scores where Muse Spark looks competitive but never comes in first.

And it doesn't stop at Terminal-Bench. On the other two benchmarks from the announcement itself:

Benchmark Muse Spark 1.2 Claude Opus 5
DeepSWE 1.1 59.3% 65.0%
Meta Internal Coding Bench 70.6% 79.4%

Read that second row again. On Meta's internal benchmark — 440 tasks pulled from real pull requests, built by Meta, picked by Meta to go on the slide — Claude Opus 5 won by almost 9 points.

Is that honesty or a lack of options? Probably both. But it's a data point most headlines didn't carry.

The gap between measurements is worth noting too: Meta puts Opus 5 at 86.7% on Terminal-Bench, Artificial Analysis puts it at 89.1%. Effort setting, harness, number of attempts — all of that moves the number. It's the same pattern we already saw with DeepSeek V4 Flash 0731, which scored 82.7 on the vendor's chart and 79% in the independent measurement. A vendor benchmark is marketing material with a methodology attached. Treat it that way.

What's actually left standing: Muse Spark 1.2 climbed 6.7 points over 1.1 (which scored 76.2%) on Terminal-Bench and 6.3 on DeepSWE. That jump between versions is real and it's fast. Meta showed up late, but it didn't show up lost.

The architecture: where Muse Code is genuinely different

If the model is second or third best, why look at it? Because the harness makes engineering decisions worth studying even if you never install the tool.

Sub-agents in isolated worktrees. When the task is big, Muse Code spins up sub-agents in separate git worktrees. In Zuckerberg's words: "When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees." In the demo, six features of a game got built at the same time, with no collisions and without touching your working copy. Anyone who's tried running two agents in the same checkout knows exactly what pain this solves.

Persistent background agents. Instead of being born and dying per task, the sub-agents stay alive for the entire session. That changes the cost of coordination: the agent doesn't rebuild context from scratch on every delegation.

Append-only event log. This is the best part. Muse Code writes a local append-only log of everything: model call, tool execution, approval, edit. The runtime is described as replay-exact and restart-safe — if the process dies in the middle of a 24-hour task, it picks up where it left off instead of starting over.

This isn't a frill. It's the difference between a demo agent and an agent you actually leave running. If you're building your own agent and still don't have an event log, that's the item to copy from Muse Code — no matter which model you run underneath. We've already written here about agents that keep state between runs; the event log is the disciplined version of that.

On top of that, Muse Code ships with default skills — /plan, /grill and /goal — in the same slash command spirit Claude Code popularized.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

$0.30 per million: what you're really paying

Now the price, which is the real weapon in this launch.

Tier padrão
  entrada:          US$ 1,25 / 1M tokens
  entrada cacheada: US$ 0,15 / 1M tokens
  saída:            US$ 4,25 / 1M tokens

Tier contributor
  entrada:          US$ 0,10 / 1M tokens
  entrada cacheada: US$ 0,002 / 1M tokens
  saída:            US$ 0,20 / 1M tokens

Add up contributor input and output and you get the famous $0.30 per million tokens. That's 12.5x cheaper on input and 21x cheaper on output than the standard tier. On cached input the gap is 75x.

Alexandr Wang, Meta's head of AI, sold exactly that: the tool "can be an incredibly good option, especially from a cost standpoint."

And what does Meta ask for in return? The right to use your prompts and completions to train its models.

Pause here, because this deserves more than a shrug.

A terminal agent doesn't read "prompts." It reads the repository. To work, Muse Code needs to load your source code, your folder structure, your database schemas, your business rules, your internal comments, your commit messages — and, if you're careless, your .env files and your keys. With 1 million tokens of context, a lot fits.

So the contributor tier math isn't "how much do I save per month." It's "how much is the intellectual property in my code worth as training material." And that call doesn't belong to the finance team. It belongs to security, to legal and — if you do client work — to your client, who probably has a confidentiality clause that says something about this.

For a personal project, open-source code or a throwaway prototype? $0.30 per million is an absurdly good price and I'd run it without thinking twice. For the monolith at the company that pays your salary? You don't have the authority to accept those terms on your own.

The discount isn't generosity. It's the value Meta assigns to the data — and it's probably getting it cheap.

Limitations and things to watch

Beta, and a real beta. macOS and Linux only. No Windows, no GUI, no IDE integration. If your workflow depends on VS Code or JetBrains, there's no path today.

Closed model. No weights, no running locally, no auditing. You depend on Meta's API and its uptime. Anyone who picked Llama precisely for the openness gets nothing here.

Installation via curl | bash. The published command is curl -fsSL https://dev.meta.ai/install.sh | bash. It works, but it's a remote script executing directly in your shell. Download it, read it, then run it. That goes for any installer, not just Meta's.

Contributor pricing is still reported secondhand. The figures circulated through public comments and press coverage; the official pricing page is the source that counts for a purchasing decision. Check dev.meta.ai before you lock in a budget.

A demo is a demo. The thousand tool calls in 24 hours optimizing a kernel are the vendor's number, on the vendor's hardware, on a task picked by the vendor. That's not independent validation.

Quick FAQ

Does Muse Code replace Claude Code or Codex today? By the numbers, no. Muse Spark 1.2 loses to Claude Opus 5 on all three benchmarks in Meta's own announcement, and Codex runs GPT-5.6 Sol, which leads the independent Terminal-Bench. Muse Code's pitch today is cost and long-running architecture, not raw capability.

Can I use the contributor tier on client code? Assume no until someone with contractual authority says yes. The tier trades a discount for permission to train models on whatever you send, and a terminal agent sends the whole repository. A contract with a confidentiality clause usually prohibits that.

Does it run locally or offline? No. Muse Spark is closed-weight and served through an API. Without public weights, there's no local execution.

Muse Code vs Claude Code: which one for a big repository? If the criterion is output quality, Claude Code — Opus 5 wins on all three benchmarks. If the criterion is the cost of long tasks on code you're allowed to share, Muse Code gets competitive. And the tiebreaker is rarely the model: it's the harness you build around it.

Is it worth testing anyway? Yes — on a repository you don't mind sharing, on the standard tier if the code is sensitive. What's worth more is studying the design: isolated worktrees per sub-agent and an append-only event log are patterns you can implement in your own harness today.

The takeaway

Meta didn't beat the state of the art. It beat GPT-5.6 Terra by 1.1 points, lost to Opus 5 on everything — including the benchmark it built itself — and chose not to compare against the model Codex actually uses. A "Meta won" headline doesn't survive five minutes of reading the chart.

But two things are real and they matter. The 6.7-point jump between Muse Spark 1.1 and 1.2 shows iteration speed. And the $0.30 per million tier puts on price pressure that Anthropic and OpenAI will have to answer one way or another.

The predictable next move is the price war dropping down to the paid tier with no data trade. Until that happens, the real cost of cheap Muse Code keeps getting paid in the currency Meta wants: your code.

Read the chart before you switch tools. And read the terms before you switch tiers.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing