Code Review Prompt: 5 Patterns That Raise Signal from 12% to 67%
Your company's code review bot has become a meme. Every PR that gets opened receives the same comment: "consider adding tests for this method." Nobody reads it anymore. The team learned that 4 out of 5 findings are pollution, so it started treating the 1 in 5 that matters as noise too.
That is the fastest way to kill a good code review tool. It's not a model problem. It's a prompt problem, and a problem with the architecture around the prompt.
The good news is that you can raise the rate of useful comments without switching stacks. The labs that do this well (Anthropic, Cursor) converged on five patterns that work together. In this post we walk through each one, with a real prompt and a Laravel workflow ready to plug into GitHub Actions.
TL;DR
- What it is: five prompt patterns for LLM code review that raise the signal ratio from ~20% to above 60%.
- Stack: Claude (Opus 4.7 or Haiku 4.5), GitHub Actions,
REVIEW.mdin the repository. - Cost/Access: works with any Anthropic key; Anthropic's managed review costs between US$ 15 and 25 per PR.
- Repository/Useful link: official Code Review docs and Claude Code Action.
Quick answer: the code review prompt that works has 5 parts: (1) it reviews only the lines added in the diff, (2) it defines in REVIEW.md what is Important and what is a Nit, with a cap of 3 nits, (3) it forces the model to read the file/grep before asserting behavior, (4) it requires file:line + a literal quote in every finding, (5) it scores each finding and only posts above 0.7. The complete REVIEW.md and the claude-review.yml are in the Ready-to-use Laravel workflow section.
Why code review bots turn into noise so fast
There's one number that explains everything. An analysis of 22,000 review comments across 178 repositories shows that concise comments are 3x more likely to be applied than long ones. And that traditional AI review tooling sits at a signal ratio (critical or important findings / total) of around 21%. Almost four out of every five comments are style suggestions, micro-optimizations, or subjective opinion.
The team learns fast. After two weeks of seeing "consider extracting this method into a helper" all the time, the developer scrolls past the batch of comments and clicks "resolve all."
When that happens, the real bug that was sitting in the middle slips through too.
The labs that measure this properly got to much better numbers. Anthropic published that Claude's Code Review has fewer than 1% of findings marked as incorrect, and that internally it raised the fraction of PRs with a substantive comment from 16% to 54%. Cursor showed that Bugbot went from 52% to almost 80% of bugs resolved between July 2025 and now, surpassing Greptile (63%) and CodeRabbit (49%).
The difference between the two groups isn't the model. It's what surrounds the model: the prompt, the verification, the severity calibration.
Prerequisites
- [ ] An Anthropic API key (or a Team/Enterprise account to use the managed Code Review).
- [ ] A repository where you can add
REVIEW.mdand.github/workflows/. - [ ] PHP 8.3+ if you're going to use the Laravel example.
- [ ] Basic knowledge of GitHub Actions.
The patterns below work with any provider (OpenAI, Gemini, a local model). The examples use Claude because it currently has the best public documentation specifically about code review, but the concept isn't tied to Anthropic.
Pattern 1: diff-anchored instead of file-scoped
This is the most common mistake. The team dumps the entire file into the context and asks "review this code." The model finds ten problems, nine of which were already there before the PR.
The right pattern is the opposite. The prompt context delivers the diff as the main entity and the full file as reference. The citation anchors are lines of the diff, not lines of the final file.
<diff>
diff --git a/app/Http/Controllers/OrderController.php
@@ -42,8 +42,15 @@ public function store(Request $request)
+ $order = Order::create([
+ 'user_id' => $request->user_id,
+ 'total' => $request->total,
+ ]);
</diff>
<file_context>
// full file, only to understand the context
</file_context>
Review only the lines ADDED in the <diff>. Lines that already existed
must not generate findings, even if they are bad: they are
pre-existing and were approved in other PRs.
This alone solves half the problem. The model stops commenting on legacy code nobody asked it to touch. It's exactly what Anthropic's Code Review does internally: it splits findings into three severities, and the "Pre-existing" class is marked in gray so it doesn't compete with what the PR brought in.
The practical consequence: the developer opens the PR and sees two comments, not fifteen. Two comments get read.
Pattern 2: a severity gate calibrated for your repository
The default definition of "critical bug" rarely matches your team's. For a payment service, a query without WHERE tenant_id is catastrophic. For a docs repository, that never even shows up, but a typo in a bash command is a P0.
Anthropic solves this with a REVIEW.md file at the root of the repository, injected verbatim into the system prompt of every agent in the review pipeline. It doesn't matter which stack you use: any LLM review improves a lot when you explicitly define the tiers.
Example for a Laravel API:
# Review instructions
## What counts as Important
Reserve Important for findings that would break behavior,
leak data, or prevent a rollback:
- Eloquent query without a tenant scope (`->where('tenant_id', ...)`)
- Log with PII (email, CPF [Brazilian tax ID], phone, request body)
- Migration that is not backward-compatible (column drop,
rename without alias, alter type without default)
- Job without `tries`/`backoff` configured
- Public endpoint without a rate limit
## What is a Nit at most
- Naming, style, refactor
- Method or trait extraction
- Test suggestion without evidence of a bug
## Nit cap
Report at most 3 nits per review. If there are more, write
"+N similar nits" in the summary. If EVERYTHING you found is a nit,
open the summary with "No blocking issues."
## Do not report
- Lint, format, type errors (Pint, PHPStan, Larastan already cover them)
- Files under `database/factories/` and `tests/Pest.php`
- Test suggestion without pointing to a concrete scenario that breaks
The detail a lot of people miss: the nit cap matters as much as the definition of Important. A review that fires off 20 style suggestions drowns the 2 real bugs that came with the PR. Limiting nits to 3 or 5 forces the model to choose.
Pattern 3: tool use before guessing
The model is looking at a 30-line diff and wants to comment "this calculateTotal function can throw a null pointer." Looking at that snippet, it seems reasonable. But the real calculateTotal lives in another file, was updated last month to return 0.0 instead of null, and the "bug" doesn't exist.
This is the scenario where the model invents a bug. And it invents one because the only thing it has is 30 lines.
The solution is to give it tool use before asking for the verdict. The agent opens the referenced file, greps for the function, reads the related test, and only then decides whether the finding is real.
Prompt outline in pseudo-code:
You are a code reviewer. Available tools:
- read_file(path: str) -> str
- grep(pattern: str, path: str) -> list[match]
- run_tests(path: str) -> str
Before reporting ANY finding about the behavior of a function,
class, or method that is NOT entirely in the diff, you MUST:
1. read_file on the file with the definition
2. grep for the existing callers
3. Validate that the scenario you were going to report actually
manifests, not that it "could theoretically manifest"
If you cannot verify it, do NOT report it. Note it under
"unverified_concerns" for the author to review manually.
It's exactly the "verification step" Anthropic describes in Code Review: multiple agents in parallel generate candidates, and a verification step checks each candidate against the actual behavior of the code before it becomes a comment. That is what backs the "fewer than 1% of findings marked as incorrect."
The right model for this stage matters. Claude Opus 4.7 follows severity instructions much better than previous versions: when the prompt says "only report high severity," it investigates with the same depth but filters before posting. On weaker models, the same instruction falls into the "I'll comment on everything just to be safe" trap.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãPattern 4: mandatory citation
This pattern is short, but it's the one that cuts hallucination the most.
Every claim about code behavior has to come with a file:line pointing to the evidence. If the model writes "this function can throw an unhandled exception," it needs to attach where in the code that exception is raised.
Rule in REVIEW.md:
## Mandatory verification
Every finding that makes a claim about behavior (it will break,
it will leak, it will time out) must include:
- file:line of the code that supports the claim
- Literal quote of the snippet (up to 3 lines)
- Concrete scenario: "if the input is X, line Y returns Z, which
makes line W (in caller A) throw exception B"
Findings derived from inference by function name
("it looks like `processPayment` doesn't handle errors") without a citation
must NOT be posted.
This changes the game. The model can no longer post an educated guess. Either it has concrete evidence, or it stays quiet. Anthropic's REVIEW.md gives exactly this example: "behavior claims need a file:line citation in the source, not an inference from naming".
A nice side effect: when the finding comes with a citation, the developer doesn't have to decide whether to trust the bot. They click the link, read the 3 lines, and decide in 10 seconds.
Pattern 5: self-grading with a threshold
The four previous patterns reduce noise on the way in. The fifth is the output filter.
Before posting any finding, a second step (LLM-as-a-judge) scores each one. Findings below a confidence threshold are discarded, not sent.
For each generated finding, evaluate it on three dimensions (0.0 to 1.0):
- confidence: how sure you are that this is a real bug,
not an inference. Strong citation = high. Inference by name = low.
- severity: impact if the bug manifests in production.
- actionability: can the author act on this finding within
10 minutes without asking for context?
Final score = weighted average (0.5 * confidence + 0.3 * severity
+ 0.2 * actionability).
Post only findings with score >= 0.7.
Findings between 0.5 and 0.7 go to the "unverified" section of the summary
(no inline comment). Below 0.5, discard.
A 0.7 threshold is where most teams converge. Lower, and the noise comes back. Higher, and you lose genuine bugs. It's worth running a calibration round on your own repository with 30–50 old PRs: apply the pipeline in dry-run mode and compare it with the actual human decision (who merged, who rejected, who asked for a fix).
This is the pattern that Cursor implemented across 40 experiments to discover the counterintuitive insight: being more aggressive on detection and at the same time stricter on the filter reduced false positives. The secret isn't the model seeing fewer bugs. It's the model posting less.
Ready-to-use Laravel workflow
Put it all together in a .github/workflows/claude-review.yml:
name: Claude Code Review
on:
pull_request:
types: [opened, synchronize, ready_for_review]
jobs:
review:
if: github.event.pull_request.draft == false
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
issues: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Run Claude code review
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
model: claude-opus-4-7
mode: review
# file read as the highest-priority instruction
review_instructions_path: REVIEW.md
# restrict to the diff
diff_only: true
# self-grading threshold
min_finding_score: 0.7
# cap on comments per PR
max_inline_comments: 8
And the matching REVIEW.md, combining the five patterns for a Laravel project:
# Review instructions
## Scope
Review ONLY lines added in the diff. Pre-existing lines
do not generate findings.
## Severity
Important (blocks merge in the summary):
- Eloquent query without a tenant scope
- Log with PII (email, CPF, phone, body)
- Non-backward-compatible migration
- Public endpoint without throttle middleware
- Job without `$tries` or `$backoff`
Nit (up to 3 per review):
- Naming, refactor, extraction
- Missing type hint on a public method
Pre-existing: do not report, even if you see it.
## Mandatory verification
Every claim about behavior needs file:line +
literal quote + concrete scenario. No citation, no post.
## Self-grading
Score each finding on confidence/severity/actionability.
Post threshold = 0.7. Below that, it goes to the summary
as "unverified".
## Do not report
- Lint/format/types (Pint + Larastan cover them)
- Test suggestion without a concrete scenario
- Files in `database/factories/`, `bootstrap/`, `vendor/`
For the first week with this turned on, it will look like the bot got lazy. PRs that used to come with 12 comments now come with 2 or 3. That's the expected effect. What changes is that those 2 or 3 get read, and the bugs in them get fixed before the merge.
Limitations and caveats
This pipeline does not replace human code review. It filters noise so the human can focus on architecture, business context, and product decisions, things the model still has no way of knowing. Treat the bot as a first filter, not as the reviewer.
Watch out for over-filtering. If you calibrate the threshold too high, the pipeline starts missing subtle bugs that would have been worth the comment. Keep an "unverified" section in the review summary, so the borderline findings stay visible without becoming noisy inline comments.
Cost scales with PR size. Anthropic's managed Code Review costs US$ 15 to 25 per PR on average, with large PRs (>1000 lines) pushing toward the ceiling. If you use the open-source action with your own key, the cost drops quite a bit, but you inherit the complexity of calibrating the prompt, tool use, and verification. It's worth measuring the ROI: 33 hours/dev/month lost filtering bad comments is a number that justifies the investment for almost any team above 5 devs.
Finally: REVIEW.md ages. Every time the team changes stack or learns an expensive lesson in production, update the file. Patterns 2, 4, and 5 work because the content of REVIEW.md is up to date, not because of the tool itself.
Quick FAQ
Why does the bot keep commenting "consider adding tests" even with REVIEW.md forbidding it?
The REVIEW.md is probably in the wrong directory, or the prompt of the bot you use doesn't inject the file. In Anthropic's managed Code Review, REVIEW.md needs to be at the root of the repository. In custom pipelines, confirm that the content is being concatenated to the system prompt before each agent. It's not enough for the file to exist; someone has to read it.
Can I use Haiku 4.5 instead of Opus 4.7 to save money? Yes, especially on small PRs. Haiku 4.5 is good at following severity instructions and works well with simple tool use. On large PRs (>500 lines) or in code with a lot of domain logic, Opus 4.7 still pays off through the quality of the findings.
How do I know if my signal ratio is good?
Take 20 PRs reviewed in the last week, count how many findings the bot posted and how many were actually applied (a follow-up commit, with a message referencing the comment). Signal ratio = applied / total. Above 60% is healthy. Below 30%, recalibrate REVIEW.md before continuing.
What if I don't use Claude, does the pattern still hold? It does. The five patterns are prompt architecture + verification, not Anthropic-specific. They work with OpenAI, Gemini, a local model (any model with decent tool use). What changes is the quality of the filter: models with more mature tool use deliver fewer false positives.
Wrapping up
LLM code review doesn't fail because the model is bad. It fails because nobody calibrated what counts as a bug, nobody gave it a tool to verify the hypothesis, and nobody filtered the output before posting. The five patterns here are exactly what separates a pipeline the team respects from a bot that would become a meme within one sprint.
If you're just starting, take the REVIEW.md from this post, adjust the rules for your stack, and turn it on in a new repository. Measure the signal ratio in two weeks. Calibrate. The hardest gain, getting the team to trust the bot again, comes quickly once the comments start making sense.
And if you're building this kind of pipeline for real and want to compare notes with people who are debugging the same prompts in production, the conversation happens at Beer and Code, the best AI engineering community in Portuguese, with a WhatsApp group open to anyone building AI in production.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã