~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / best-ai-for-coding-2026-claude-code-vs-codex $
News

Claude Code vs Codex: Codex Wins the Terminal by 13 Points, Claude Wins the Hard Repo by 10

LS Lucas Souza · · 10 min read
Claude Code vs Codex: Codex Wins the Terminal by 13 Points, Claude Wins the Hard Repo by 10

Updated on August 15, 2026. The benchmarks below were measured with Claude Opus 4.8 and GPT-5.5, the current versions when this comparison came out. Since then, Anthropic has released Opus 5 and OpenAI has released the GPT-5.6 Sol, Terra and Luna family. The individual numbers have changed; the structural difference between the two tools, which is what this post is about, has not.

Short answer, before anything else: for terminal and shell work, Codex wins by 13 points. For the hard repository problem, the one that touches five files and breaks a test somewhere else, Claude Code wins by 10. For the well-defined bug, they tie. The choice isn't about quality. It's about which bottleneck hurts you the most.

At the end of my model testing, two were left standing for real work: Claude and Codex. I dropped the rest along the way: one hallucinated function names, another choked on large repos, another was great in chat and terrible in the terminal. But these two I couldn't take out of my workflow.

In this post I compare the two on real agentic coding: strength in a real repo, behavior in the terminal, cost, and execution model. No marketing leaderboard. This is what was left after running both on projects that pay the bills.

TL;DR

  • What it is: a practical comparison between Claude Code (Anthropic) and Codex (OpenAI) for agentic coding.
  • Models measured: Claude Code with Claude Opus 4.8; Codex with the GPT-5.5/5.4 family. See the update note at the top.
  • Terminal-Bench: Codex 82.7% vs Claude Code's 69.4%. Codex ahead by 13 points.
  • SWE-bench Pro: Claude Code 69.2% vs Codex's 58.6%. Claude ahead by 10 points.
  • SWE-bench Verified: a statistical tie, 88.7% vs 88.6%.
  • Cost/Access: both start at $20/month; Claude Code scales up through Max ($100/$200), Codex through Plus/Pro ($20 to $200).

The context: two different takes on a "coding agent"

Before comparing numbers, you need to understand that Claude Code and Codex are not the same tool with a different logo. They start from different philosophies about how an agent should work.

Claude Code is a supervised pair programmer, in your terminal. It reads your local filesystem, runs commands in your shell, uses your git, and calls the Anthropic API only to reason. You're right there, watching every step, approving what matters. It's assisted engineering, not outsourced engineering (Builder.io). If you haven't gotten the hang of it yet, it's worth reading what Claude Code is and how it works first.

Codex is an autonomous executor, in the cloud. It runs tasks in sandboxed containers on OpenAI's infrastructure, in parallel, often far from your terminal. You hand it the task, it goes off and does it, and you review the result afterward (Developers Digest).

That difference is not a detail. It explains almost everything that follows: why one is better in the terminal, why the other is better at delegating batch work, and why most experienced devs ended up running both.

Adoption tells the same story. In February 2026, Claude Code was already behind roughly 4% of all public GitHub commits, something like 135,000 commits a day, with a peak of 326,000 in a single day in March, and a SemiAnalysis projection of passing 20% by the end of 2026 (SemiAnalysis via Composio). Codex, on the other side, became a coding command center: a VS Code extension with nearly 10 million installs, a CLI with more than 88,000 GitHub stars, a web app, iOS, and, since June 2026, availability on Amazon Bedrock.

Both took off. The fight isn't "who survives." It's "who's better at what."

And notice what the update note at the top of this post means: in eight weeks, both models cited here have already been replaced. If you pick your tool by the benchmark of the month, you redo that choice every month. If you have an eval for your own domain, you swap models in one line of config and move on. That's the kind of system we build together, live, every week in the Clã Beer and Code.

Benchmarks: where each one wins

This is where the conversation gets concrete. And the result is more split than the hype suggests.

On Terminal-Bench, which measures shell work, commands, and terminal automation, Codex with GPT-5.5 opens a clear lead: 82.7% vs 69.4% for Claude Code. Thirteen points. That's not a statistical tie, it's a difference you feel day to day if you live in the terminal (morphllm).

On SWE-bench, which measures fixing bugs and implementing features in real repositories, things flip and tighten:

  • SWE-bench Verified: GPT-5.5 leads by a hair, 88.7% vs 88.6% for Opus 4.8. A statistical tie of 0.1 point.
  • SWE-bench Pro (the hardest, multi-file problems): Opus 4.8 pulls ahead, 69.2% vs 58.6%.

Read that slowly, because it's the heart of the post. On the truly hard problem, the one that touches five files, breaks a test somewhere else, and requires understanding the whole repo, Claude Code is consistently better. On terminal and shell work, Codex dominates. And on "fix a well-defined bug," they tie.

In other words: there is no winner in the abstract. There's the best one for your bottleneck.

Claude Code or Codex, by scenario

Enough benchmarks. Let me translate this into a practical decision, the way I actually use them.

Scenario 1: heavy refactor in a large repo

You need to rename a concept that shows up in 30 files, or migrate an entire layer without breaking the tests. This is where Claude Code shines. The supervised loop, plan mode before expensive operations, and the SWE-bench Pro lead make a real difference when a mistake is costly and needs to be caught before it spreads.

# Claude Code: you review the plan before it touches 30 files
claude
> /plan migrate the payment layer from Stripe v1 to v2 without breaking the tests

Scenario 2: terminal work and automation

Shell scripts, CI wrangling, tasks that live in the terminal. Codex delivers 13 more points on Terminal-Bench, and its sandboxed design fits tasks that boil down to "run this command, give me back the result."

Scenario 3: delegating batch work, in parallel

You have five independent tasks and want to fire them all off and review later. Codex's autonomous cloud execution model was built for this: agents working in parallel across multiple projects. In fact, both already have real multi-agent support: Codex took subagents to GA on March 14, 2026, and Claude Code has Agent Teams evolving in the same direction.

The golden tip: the criterion isn't quality, it's control

The right question isn't "which one is smarter." Both are absurdly capable. The question is: how much control do you want during the task?

Want to be right there, seeing every decision, approving changes to sensitive code? Claude Code. Want to delegate and review the finished result? Codex. It's the same difference as pair programming versus opening a PR for someone to solve on their own.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Cost: what you actually pay

Both start in the same place: $20/month. But what you get for that money differs.

  • Claude Code: Pro at $20 (tighter quota, burns fast), Max 5x at $100, and Max 20x at $200 for heavy users. In practice, the average cost lands around $13 per dev per active day, with 90% of users under $30/day (CloudZero).
  • Codex: Free, Plus at $20 (15 to 80 messages per 5-hour window), and Pro starting at $100. In April 2026, OpenAI switched from per-message billing to token-based credits.

The reality: for the same $20, Codex tends to give you more agent execution time, and Claude Code gives you a tighter quota that you burn through quickly if you're a heavy user. If you work with the agent all day, the step up to Max ($100) or to Codex's Pro is almost inevitable.

Limitations and things to watch out for

Neither one is magic. Here's where you'll get burned:

  • The numbers have a date. They were measured with Opus 4.8 and GPT-5.5. Both companies have already shipped a new generation. Treat the scoreboard as a direction, not as today's truth.
  • Claude Code depends on your local environment. If your terminal, git, or filesystem is a mess, it inherits the mess. It's power and responsibility in the same package.
  • Codex is more opaque by design. Running in a cloud container brings isolation and parallelism, but makes browser automation and tasks that need your immediate local context harder.
  • Quota disappears fast under heavy use. Both of them. The $20 plan is an entry point, not a production plan for someone who lives in the agent.
  • A benchmark is not your repo. SWE-bench and Terminal-Bench are proxies. Your legacy code, with all its quirks, can flip any of these numbers. Test on your own project before making up your mind.

Quick FAQ

Claude Code or Codex: which one should I pick if I can only have one? If your work is mostly changing a real repo, with refactors and multi-file changes, go with Claude Code (a 10-point lead on SWE-bench Pro). If it's mostly terminal, automation, and batch delegation, go with Codex (13 points on Terminal-Bench). When in doubt, Claude Code has the edge on the hard problem.

What's the best AI for coding in a large Laravel/PHP project? Going by the SWE-bench Pro numbers, Claude Code has the advantage on changes that touch many files, which is exactly the case in a mature Laravel monolith. But test both on your repo: legacy code is unpredictable.

Can I use both at the same time? Yes, and that's what most experienced devs do. Claude Code for focused, supervised implementation; Codex for firing off parallel work in the background. They aren't mutually exclusive.

Does the tie on SWE-bench Verified mean it doesn't matter which one I pick? No. Verified measures well-defined bugs, where they tie within 0.1 point. The difference shows up on Pro (Claude ahead by 10) and on Terminal-Bench (Codex ahead by 13). Look at the benchmark that reflects your work.

Conclusion

Claude Code wins in the real repo, on the hard refactor, on step-by-step control. Codex wins in the terminal, on autonomous execution, on batch delegation. And the two tie where the problem is well defined. The dev who understands this stops asking "which one is better" and starts asking "which one fits this bottleneck."

The next step in this game isn't choosing the agent. It's learning to build around it: the harness, the context, the limits. Knowing how to operate an agent is different from knowing how to put together an agentic system that holds up in production.

In the end, the tool is just a tool. And it changes names every eight weeks, as the note at the top of this post proves. The differentiator is still the dev who understands the problem, models the solution, and knows when to trust the agent, and when to keep a tight rein.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing