#Codex
Meta launched Muse Code in beta, a terminal agent running Muse Spark 1.2. The chart says 82.9% on Terminal-Bench and a win over Codex. I went and read the chart: Meta beat GPT-5.6 Terra, not the GPT-5.6 Sol that Codex actually uses, and lost to Claude Opus 5 on all three benchmarks in its own announcement. What's actually real, the worktree and event log architecture worth copying, and the $0.30 per million price you pay for with your code.
You send a one-line question and /usage reports a whole day's worth of consumption. Saving tokens in a coding assistant has nothing to do with prompt size: it's about prefix caching. How it works in Claude Code, Codex, and Cursor, the seven actions that invalidate it without you noticing, how to measure it with cache_read vs cache_creation, and eight levers to stretch the session.
We combed through the new official Codex documentation (learn.chatgpt.com) and boiled it down to the 10 best practices OpenAI itself recommends: outcome-first prompting, a living AGENTS.md, restrictive permissions, model selection, skills, cloud, and automation.
GPT-5.6 has been in Codex since July 9, across all three tiers. Ultra mode coordinates four subagents that cooperate during the task and buys 3 points on Terminal-Bench (88.8% to 91.9%), at a much higher token cost. And CodeRabbit's independent test kills the reflex of picking the cheap tier: Terra burned 2.6x more tokens than Sol to pass 40.7% of tasks versus 63.7%. Price per token is not price per task.
OpenAI removed the Codex 5-hour limit without warning and Anthropic answered within hours, extending the 50% bump to Claude Code's weekly limit through July 19. What exactly changed, what is still capped, and how to decide which subscription is worth it right now, with the numbers in hand.
Claude Code or Codex? The answer comes with numbers: Codex opens a 13-point lead on Terminal-Bench (82.7% vs 69.4%) and Claude Code opens a 10-point lead on SWE-bench Pro (69.2% vs 58.6%), which is the benchmark for real multi-file problems. On SWE-bench Verified they tie. Here is the verdict by scenario, the real cost per dev, and the criterion that matters more than quality: how much control you want during the task.