Code review's purpose is two-fold: catch bugs before merge, and share knowledge so the team grows together. AI is good at the first; it cannot replace the second. The right model: AI as a first-pass reviewer that catches the easy stuff so humans can focus on the architectural and contextual concerns.
This page is the working pattern.
In good evaluations:
What it misses:
This is why AI review is "first pass" not "the review."
Bot runs on every PR. Posts comments with suggestions. Author reviews; takes valid ones; ignores false positives.
On PR open:
ai_review(diff, repo_context) -> [list of comments]
for each comment: post inline on PR
summary comment: "Reviewed by [bot]. N suggestions. Examine before merge."
Tools that do this:
Runs locally before commit. Catches things before the PR.
git pre-push hook:
diff = git diff origin/main..HEAD
ai_review(diff) -> findings
prompt user to address before pushing
Trade-off: more friction in dev loop; catches issues earlier; reduces PR noise.
Human is reviewing a PR; asks AI to summarise the changes, flag risky parts, suggest what to focus on:
You: What changed in this PR? What should I be careful about?
AI: This PR adds support for async refunds. The risky parts:
1. The refund retry logic in payment_service.go uses a fixed
2-second delay which doesn't match the rest of the codebase's
exponential backoff pattern.
2. The new database column is NOT NULL but the migration doesn't
backfill existing rows — this will fail in environments with
existing data.
3. The test in test_refunds.py tests the happy path but no error
paths.
This pattern amplifies human review without replacing it. The human still decides; the AI helps them allocate attention.
A naïve prompt ("review this code") gets generic feedback. Better:
You are reviewing code for a financial-services backend. Focus on:
- Correctness of the business logic
- Security vulnerabilities (esp. around money handling)
- Concurrency / race conditions
- Match with existing patterns in the codebase
Here is the diff:
{diff}
Here are related files for context:
{neighbouring code}
Output:
- Severity: blocking | should-fix | optional
- Line: specific line in the diff
- Concern: one paragraph
- Suggestion: concrete code change if applicable
Don't comment on style or formatting (handled separately).
Don't comment on things you can't see (architecture not in this diff).
If everything looks good, say so.
Categorising by severity matters — without it, every observation gets flagged equally and the human can't filter.
Including context (neighbouring code) is essential. Without it, the model can't tell whether a function "should" be called by the caller or whether the caller is wrong.
AI reviewers produce false positives. Common categories:
Mitigations:
A reviewer that produces 30% noise gets ignored. Tune toward fewer high-confidence findings.
A team that fully delegates code review to AI will produce more bugs than a team that uses it as one signal. Specific risks:
The discipline: AI review is part of the review process. Human review is also part of the review process. The goal is faster + better, not faster instead of better.
For security-sensitive PRs, dedicated security-focused prompts catch more:
Review this code for security issues. Focus on:
- Input validation and sanitisation
- Authentication and authorisation
- Injection (SQL, command, template, XSS)
- Cryptographic mistakes
- Secret management
- Race conditions in security-relevant code
- Path traversal, SSRF
Pair with SAST tools (Semgrep, CodeQL). The combination catches more than either alone.
Teams that have adopted AI code review well report:
Teams that adopted AI code review poorly report:
The difference is workflow design, not tool choice.