~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / 5-prompt-patterns-llm-code-review $
Tutorials

Code Review Prompt: 5 Patterns That Raise Signal from 12% to 67%

LS Lucas Souza · · 15 min read
Code Review Prompt: 5 Patterns That Raise Signal from 12% to 67%

Your company's code review bot has become a meme. Every PR that gets opened receives the same comment: "consider adding tests for this method." Nobody reads it anymore. The team learned that 4 out of 5 findings are pollution, so it started treating the 1 in 5 that matters as noise too.

That is the fastest way to kill a good code review tool. It's not a model problem. It's a prompt problem, and a problem with the architecture around the prompt.

The good news is that you can raise the rate of useful comments without switching stacks. The labs that do this well (Anthropic, Cursor) converged on five patterns that work together. In this post we walk through each one, with a real prompt and a Laravel workflow ready to plug into GitHub Actions.

TL;DR

  • What it is: five prompt patterns for LLM code review that raise the signal ratio from ~20% to above 60%.
  • Stack: Claude (Opus 4.7 or Haiku 4.5), GitHub Actions, REVIEW.md in the repository.
  • Cost/Access: works with any Anthropic key; Anthropic's managed review costs between US$ 15 and 25 per PR.
  • Repository/Useful link: official Code Review docs and Claude Code Action.

Quick answer: the code review prompt that works has 5 parts: (1) it reviews only the lines added in the diff, (2) it defines in REVIEW.md what is Important and what is a Nit, with a cap of 3 nits, (3) it forces the model to read the file/grep before asserting behavior, (4) it requires file:line + a literal quote in every finding, (5) it scores each finding and only posts above 0.7. The complete REVIEW.md and the claude-review.yml are in the Ready-to-use Laravel workflow section.

Why code review bots turn into noise so fast

There's one number that explains everything. An analysis of 22,000 review comments across 178 repositories shows that concise comments are 3x more likely to be applied than long ones. And that traditional AI review tooling sits at a signal ratio (critical or important findings / total) of around 21%. Almost four out of every five comments are style suggestions, micro-optimizations, or subjective opinion.

The team learns fast. After two weeks of seeing "consider extracting this method into a helper" all the time, the developer scrolls past the batch of comments and clicks "resolve all."

When that happens, the real bug that was sitting in the middle slips through too.

The labs that measure this properly got to much better numbers. Anthropic published that Claude's Code Review has fewer than 1% of findings marked as incorrect, and that internally it raised the fraction of PRs with a substantive comment from 16% to 54%. Cursor showed that Bugbot went from 52% to almost 80% of bugs resolved between July 2025 and now, surpassing Greptile (63%) and CodeRabbit (49%).

The difference between the two groups isn't the model. It's what surrounds the model: the prompt, the verification, the severity calibration.

Prerequisites

  • [ ] An Anthropic API key (or a Team/Enterprise account to use the managed Code Review).
  • [ ] A repository where you can add REVIEW.md and .github/workflows/.
  • [ ] PHP 8.3+ if you're going to use the Laravel example.
  • [ ] Basic knowledge of GitHub Actions.

The patterns below work with any provider (OpenAI, Gemini, a local model). The examples use Claude because it currently has the best public documentation specifically about code review, but the concept isn't tied to Anthropic.

Pattern 1: diff-anchored instead of file-scoped

This is the most common mistake. The team dumps the entire file into the context and asks "review this code." The model finds ten problems, nine of which were already there before the PR.

The right pattern is the opposite. The prompt context delivers the diff as the main entity and the full file as reference. The citation anchors are lines of the diff, not lines of the final file.

<diff>
diff --git a/app/Http/Controllers/OrderController.php
@@ -42,8 +42,15 @@ public function store(Request $request)
+    $order = Order::create([
+        'user_id' => $request->user_id,
+        'total' => $request->total,
+    ]);
</diff>

<file_context>
// full file, only to understand the context
</file_context>

Review only the lines ADDED in the <diff>. Lines that already existed
must not generate findings, even if they are bad: they are
pre-existing and were approved in other PRs.

This alone solves half the problem. The model stops commenting on legacy code nobody asked it to touch. It's exactly what Anthropic's Code Review does internally: it splits findings into three severities, and the "Pre-existing" class is marked in gray so it doesn't compete with what the PR brought in.

The practical consequence: the developer opens the PR and sees two comments, not fifteen. Two comments get read.

Pattern 2: a severity gate calibrated for your repository

The default definition of "critical bug" rarely matches your team's. For a payment service, a query without WHERE tenant_id is catastrophic. For a docs repository, that never even shows up, but a typo in a bash command is a P0.

Anthropic solves this with a REVIEW.md file at the root of the repository, injected verbatim into the system prompt of every agent in the review pipeline. It doesn't matter which stack you use: any LLM review improves a lot when you explicitly define the tiers.

Example for a Laravel API:

# Review instructions

## What counts as Important

Reserve Important for findings that would break behavior,
leak data, or prevent a rollback:

- Eloquent query without a tenant scope (`->where('tenant_id', ...)`)
- Log with PII (email, CPF [Brazilian tax ID], phone, request body)
- Migration that is not backward-compatible (column drop,
  rename without alias, alter type without default)
- Job without `tries`/`backoff` configured
- Public endpoint without a rate limit

## What is a Nit at most

- Naming, style, refactor
- Method or trait extraction
- Test suggestion without evidence of a bug

## Nit cap

Report at most 3 nits per review. If there are more, write
"+N similar nits" in the summary. If EVERYTHING you found is a nit,
open the summary with "No blocking issues."

## Do not report

- Lint, format, type errors (Pint, PHPStan, Larastan already cover them)
- Files under `database/factories/` and `tests/Pest.php`
- Test suggestion without pointing to a concrete scenario that breaks

The detail a lot of people miss: the nit cap matters as much as the definition of Important. A review that fires off 20 style suggestions drowns the 2 real bugs that came with the PR. Limiting nits to 3 or 5 forces the model to choose.

Pattern 3: tool use before guessing

The model is looking at a 30-line diff and wants to comment "this calculateTotal function can throw a null pointer." Looking at that snippet, it seems reasonable. But the real calculateTotal lives in another file, was updated last month to return 0.0 instead of null, and the "bug" doesn't exist.

This is the scenario where the model invents a bug. And it invents one because the only thing it has is 30 lines.

The solution is to give it tool use before asking for the verdict. The agent opens the referenced file, greps for the function, reads the related test, and only then decides whether the finding is real.

Prompt outline in pseudo-code:

You are a code reviewer. Available tools:

- read_file(path: str) -> str
- grep(pattern: str, path: str) -> list[match]
- run_tests(path: str) -> str

Before reporting ANY finding about the behavior of a function,
class, or method that is NOT entirely in the diff, you MUST:

1. read_file on the file with the definition
2. grep for the existing callers
3. Validate that the scenario you were going to report actually
   manifests, not that it "could theoretically manifest"

If you cannot verify it, do NOT report it. Note it under
"unverified_concerns" for the author to review manually.

It's exactly the "verification step" Anthropic describes in Code Review: multiple agents in parallel generate candidates, and a verification step checks each candidate against the actual behavior of the code before it becomes a comment. That is what backs the "fewer than 1% of findings marked as incorrect."

The right model for this stage matters. Claude Opus 4.7 follows severity instructions much better than previous versions: when the prompt says "only report high severity," it investigates with the same depth but filters before posting. On weaker models, the same instruction falls into the "I'll comment on everything just to be safe" trap.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

Pattern 4: mandatory citation

This pattern is short, but it's the one that cuts hallucination the most.

Every claim about code behavior has to come with a file:line pointing to the evidence. If the model writes "this function can throw an unhandled exception," it needs to attach where in the code that exception is raised.

Rule in REVIEW.md:

## Mandatory verification

Every finding that makes a claim about behavior (it will break,
it will leak, it will time out) must include:

- file:line of the code that supports the claim
- Literal quote of the snippet (up to 3 lines)
- Concrete scenario: "if the input is X, line Y returns Z, which
  makes line W (in caller A) throw exception B"

Findings derived from inference by function name
("it looks like `processPayment` doesn't handle errors") without a citation
must NOT be posted.

This changes the game. The model can no longer post an educated guess. Either it has concrete evidence, or it stays quiet. Anthropic's REVIEW.md gives exactly this example: "behavior claims need a file:line citation in the source, not an inference from naming".

A nice side effect: when the finding comes with a citation, the developer doesn't have to decide whether to trust the bot. They click the link, read the 3 lines, and decide in 10 seconds.

Pattern 5: self-grading with a threshold

The four previous patterns reduce noise on the way in. The fifth is the output filter.

Before posting any finding, a second step (LLM-as-a-judge) scores each one. Findings below a confidence threshold are discarded, not sent.

For each generated finding, evaluate it on three dimensions (0.0 to 1.0):

- confidence: how sure you are that this is a real bug,
  not an inference. Strong citation = high. Inference by name = low.
- severity: impact if the bug manifests in production.
- actionability: can the author act on this finding within
  10 minutes without asking for context?

Final score = weighted average (0.5 * confidence + 0.3 * severity
+ 0.2 * actionability).

Post only findings with score >= 0.7.

Findings between 0.5 and 0.7 go to the "unverified" section of the summary
(no inline comment). Below 0.5, discard.

A 0.7 threshold is where most teams converge. Lower, and the noise comes back. Higher, and you lose genuine bugs. It's worth running a calibration round on your own repository with 30–50 old PRs: apply the pipeline in dry-run mode and compare it with the actual human decision (who merged, who rejected, who asked for a fix).

This is the pattern that Cursor implemented across 40 experiments to discover the counterintuitive insight: being more aggressive on detection and at the same time stricter on the filter reduced false positives. The secret isn't the model seeing fewer bugs. It's the model posting less.

Ready-to-use Laravel workflow

Put it all together in a .github/workflows/claude-review.yml:

name: Claude Code Review

on:
  pull_request:
    types: [opened, synchronize, ready_for_review]

jobs:
  review:
    if: github.event.pull_request.draft == false
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write
      issues: write

    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - name: Run Claude code review
        uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          model: claude-opus-4-7
          mode: review
          # file read as the highest-priority instruction
          review_instructions_path: REVIEW.md
          # restrict to the diff
          diff_only: true
          # self-grading threshold
          min_finding_score: 0.7
          # cap on comments per PR
          max_inline_comments: 8

And the matching REVIEW.md, combining the five patterns for a Laravel project:

# Review instructions

## Scope

Review ONLY lines added in the diff. Pre-existing lines
do not generate findings.

## Severity

Important (blocks merge in the summary):
- Eloquent query without a tenant scope
- Log with PII (email, CPF, phone, body)
- Non-backward-compatible migration
- Public endpoint without throttle middleware
- Job without `$tries` or `$backoff`

Nit (up to 3 per review):
- Naming, refactor, extraction
- Missing type hint on a public method

Pre-existing: do not report, even if you see it.

## Mandatory verification

Every claim about behavior needs file:line +
literal quote + concrete scenario. No citation, no post.

## Self-grading

Score each finding on confidence/severity/actionability.
Post threshold = 0.7. Below that, it goes to the summary
as "unverified".

## Do not report

- Lint/format/types (Pint + Larastan cover them)
- Test suggestion without a concrete scenario
- Files in `database/factories/`, `bootstrap/`, `vendor/`

For the first week with this turned on, it will look like the bot got lazy. PRs that used to come with 12 comments now come with 2 or 3. That's the expected effect. What changes is that those 2 or 3 get read, and the bugs in them get fixed before the merge.

Limitations and caveats

This pipeline does not replace human code review. It filters noise so the human can focus on architecture, business context, and product decisions, things the model still has no way of knowing. Treat the bot as a first filter, not as the reviewer.

Watch out for over-filtering. If you calibrate the threshold too high, the pipeline starts missing subtle bugs that would have been worth the comment. Keep an "unverified" section in the review summary, so the borderline findings stay visible without becoming noisy inline comments.

Cost scales with PR size. Anthropic's managed Code Review costs US$ 15 to 25 per PR on average, with large PRs (>1000 lines) pushing toward the ceiling. If you use the open-source action with your own key, the cost drops quite a bit, but you inherit the complexity of calibrating the prompt, tool use, and verification. It's worth measuring the ROI: 33 hours/dev/month lost filtering bad comments is a number that justifies the investment for almost any team above 5 devs.

Finally: REVIEW.md ages. Every time the team changes stack or learns an expensive lesson in production, update the file. Patterns 2, 4, and 5 work because the content of REVIEW.md is up to date, not because of the tool itself.

Quick FAQ

Why does the bot keep commenting "consider adding tests" even with REVIEW.md forbidding it? The REVIEW.md is probably in the wrong directory, or the prompt of the bot you use doesn't inject the file. In Anthropic's managed Code Review, REVIEW.md needs to be at the root of the repository. In custom pipelines, confirm that the content is being concatenated to the system prompt before each agent. It's not enough for the file to exist; someone has to read it.

Can I use Haiku 4.5 instead of Opus 4.7 to save money? Yes, especially on small PRs. Haiku 4.5 is good at following severity instructions and works well with simple tool use. On large PRs (>500 lines) or in code with a lot of domain logic, Opus 4.7 still pays off through the quality of the findings.

How do I know if my signal ratio is good? Take 20 PRs reviewed in the last week, count how many findings the bot posted and how many were actually applied (a follow-up commit, with a message referencing the comment). Signal ratio = applied / total. Above 60% is healthy. Below 30%, recalibrate REVIEW.md before continuing.

What if I don't use Claude, does the pattern still hold? It does. The five patterns are prompt architecture + verification, not Anthropic-specific. They work with OpenAI, Gemini, a local model (any model with decent tool use). What changes is the quality of the filter: models with more mature tool use deliver fewer false positives.

Wrapping up

LLM code review doesn't fail because the model is bad. It fails because nobody calibrated what counts as a bug, nobody gave it a tool to verify the hypothesis, and nobody filtered the output before posting. The five patterns here are exactly what separates a pipeline the team respects from a bot that would become a meme within one sprint.

If you're just starting, take the REVIEW.md from this post, adjust the rules for your stack, and turn it on in a new repository. Measure the signal ratio in two weeks. Calibrate. The hardest gain, getting the team to trust the bot again, comes quickly once the comments start making sense.

And if you're building this kind of pipeline for real and want to compare notes with people who are debugging the same prompts in production, the conversation happens at Beer and Code, the best AI engineering community in Portuguese, with a WhatsApp group open to anyone building AI in production.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing