~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / 10-codex-best-practices $
Tutorials

10 Codex Best Practices (Straight from the Official Documentation)

LS Lucas Souza · · 8 min read
10 Codex Best Practices (Straight from the Official Documentation)

Everybody installed Codex. Very few people read the documentation.

And it changed, both its address and its philosophy. The official docs left developers.openai.com and now live at learn.chatgpt.com, rewritten for an era where Codex is an agent you hand a task to, not an autocomplete you hit tab on.

I went through all of it: prompting, permissions, models, skills, cloud, changelog. This post boils it down to the 10 best practices OpenAI itself recommends, each with its source and with how to apply it in your workflow today.

TL;DR

  • What it is: the 10 practices recommended in the official Codex docs, broken down with practical examples.
  • Stack: Codex CLI, IDE extension (VS Code, Cursor, Windsurf, JetBrains, Xcode), desktop app, and Codex cloud.
  • Cost/Access: included in paid ChatGPT plans (Plus/Pro and above).
  • Useful link: learn.chatgpt.com/docs.

The context: the Codex docs moved house (and changed philosophy)

If you bookmarked the old link, it redirects: developers.openai.com/codex now points to learn.chatgpt.com. It wasn't just a domain change. The new docs frame Codex as one agent with four surfaces: the CLI in the terminal, the IDE extension, the ChatGPT desktop app (an integration that landed in July, according to the changelog), and Codex cloud, which runs tasks in isolated environments while your machine stays free.

Under the hood, it's the GPT-5.6 family in three flavors (Sol, Terra, and Luna) plus reasoning levels that go from Low to Max. The implicit message of the entire docs: the bottleneck isn't the model, it's how you set up context, permissions, and verification around it.

Reading the docs is half the journey. The other half is seeing these practices run on a real project, with someone next to you reviewing what you built. That's exactly the environment the Clã Beer and Code offers, every week, live.

The 10 Codex best practices

1. Ask for the result, not the step-by-step

The first piece of guidance in the prompting guide is literal: "start with the result, not a detailed list of steps". Describe what needs to exist at the end (behavior, format, constraints) and let the agent decide how to get there. Micromanaging every step wastes the model's main skill: planning.

Bad:   "Open UserController, add a method, then create the route..."
Good:  "I need a GET /users/{id}/invoices endpoint that returns the
        user's paginated invoices, with a test covering the empty case."

If you want to take prompt-as-contract further, we've already broken down how to use goals in the Codex CLI to guide the agent without micromanaging.

2. Give the context that changes the answer

The docs recommend attaching everything that alters the result: files, screenshots, links. They also recommend enabling web search when the problem involves recent information. Context you didn't give is context the model will make up. In the IDE extension, reference open files and selections right in the composer instead of re-explaining the problem.

3. Don't chase the perfect prompt

Also literal in the docs: "your first prompt doesn't need to be perfect". The recommended flow is to iterate: run it, look at it, ask for the specific adjustment in the next message. Starting from scratch on every attempt throws away the context the session has already built up. In the CLI, codex resume brings back old sessions for exactly this reason.

4. Treat AGENTS.md as living documentation

AGENTS.md is where the project's persistent instructions live: conventions, build commands, what the agent must never touch. Run /init to generate the first one. The hierarchy is global → repository → directory, with the most specific one winning. Golden rule: every correction you repeat to Codex twice is a candidate to become a line in AGENTS.md, and the only things that belong there are the ones it can't infer by reading the code.

5. Run with the most restrictive permission that gets the job done

The permissions docs define three profiles: :read-only (inspection only), :workspace (writes inside the workspace), and :danger-full-access, where the name is already the warning. The official recommendation: pick the most restrictive profile that still completes the task. Codebase analysis? Read-only. Feature? Workspace. Full access is a deliberate exception, never a default because you're too lazy to approve commands.

6. Pick model and reasoning by the size of the problem

The models page is direct: Sol for ambiguous, high-value work, Terra for day-to-day, Luna for repeated, well-defined tasks. And use the lowest reasoning effort that produces an acceptable result. Max on a trivial task is a bigger bill and a slower answer, with no gain. Start at the default (Sol, medium effort) and calibrate to what the task calls for, via /model.

7. Turn everything you repeat into a skill

Skills package a workflow (instructions, templates, examples) into something reusable that Codex recognizes on its own. The official path: take a focused, repetitive task, describe the flow to $skill-creator, test it with a real case, share it with the team. It's tacit knowledge becoming a versionable artifact. Plugins come in when the skill needs an external service via MCP.

8. Delegate the long stuff to the cloud

A task that will take an hour doesn't need your machine, or you watching it. Codex cloud runs tasks in isolated, parallel environments, integrated with GitHub. The best practice from the docs: configure the repository environment (dependencies, setup, variables) before delegating. An agent in an environment that doesn't build produces a PR that doesn't build.

9. Git checkpoint before, /review after

The IDE docs recommend creating git checkpoints before and after each task, so reverting becomes a git reset, not archaeology. And before committing, /review (or codex review against the base branch) has the agent inspect its own diff for problems. It doesn't replace your eyes: it organizes what reaches them.

10. Automate the repetitive: codex exec and scheduled tasks

codex exec runs Codex in non-interactive mode, ready for scripts and CI. And the July changelog brought scheduled tasks (recurring work running on its own, locally or in worktrees) and Record & Replay, which converts a demonstration of yours into a skill. Issue triage, changelog updates, dependency checks: if it happens every week, it doesn't need you every week.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

Limitations and things to watch out for

danger-full-access is what the name says. A profile with broad write access changes scripts, hooks, and shared files persistently, and a wildcard network rule stays open even with an allowlist. Use it when it's intentional, never as a shortcut.

High reasoning isn't free quality. It's cost and latency. The docs themselves tell you to calibrate down whenever the result holds up.

Human review is still mandatory. /review organizes the diff, cloud opens the PR, but you're the one who merges. No page in the docs promises otherwise, and be suspicious of anyone who does.

Docs in motion. The migration to learn.chatgpt.com is recent and new features ship every month. Check the changelog before assuming something doesn't exist.

Quick FAQ

Is the Codex CLI paid? The CLI is open source, but using it consumes your ChatGPT plan limits (Plus/Pro and above). Installation: curl -fsSL https://chatgpt.com/codex/install.sh | sh.

Does it work in Cursor and JetBrains? Yes. The extension covers VS Code and its forks (Cursor, Windsurf) and has native integration with JetBrains and Xcode.

Does AGENTS.md apply to the cloud too? It does. The file lives in the repository, so any surface that clones the repo, including Codex cloud environments, reads the same instructions. That's why it's the most durable personalization layer.

What's the difference between a skill and a plugin? A skill is a packaged workflow (instructions + resources) for a focused task. A plugin is the installable package that can combine skills and MCP connectors. Use it when you need an external tool, not just instructions.

Conclusion

The 10 practices fit in one sentence: state the result, give context, persist what repeats (AGENTS.md, skills), restrict what's dangerous (permissions), calibrate what's expensive (model, reasoning), and verify everything (git, /review). None of this is a secret. It's published, for free, in the official docs. The difference is who applies it.

The provocation: with scheduled tasks and Record & Replay, Codex is going from a tool you open to an asynchronous coworker that works while you sleep. Whoever structures repository, permissions, and verification first reaps the rewards first. And if you want to compare the other side of the ring, see how the HumanLayer manifesto criticizes "software factories". The harness discussion applies to both agents.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing