Using Agents for Code Review

Code review is high-leverage agent work — if set up right. What agents catch well, what they fake, and how to build reviews that find real issues.

Published 2026-09-29 · 10 min read

Code review is one of the highest-leverage places to deploy an AI agent: the task is well-defined, the input is structured, and the output is verifiable. But have the agent review my code covers a wide range of setups, from genuinely useful to actively harmful. An agent that rubber-stamps everything is worse than no review at all, because it creates the illusion of scrutiny while adding latency. This guide covers what agents actually catch, how to set up reviews that find real issues, and where human reviewers remain irreplaceable.

What agents catch well

  • Style and convention drift: naming inconsistencies, formatting deviations, and patterns that violate the project's established conventions.
  • Obvious bugs: null dereferences, off-by-one errors in simple loops, unused variables, and unreachable code that a tired human eye skips.
  • Missing error handling: unhandled promise rejections, unchecked return values, and happy-path-only logic in code that clearly needs failure paths.
  • Security basics: hardcoded secrets, string-concatenated queries, missing input validation, and overly broad permissions on new endpoints.
  • Documentation drift: comments describing behavior the code no longer has, stale README examples, and public APIs whose docs were not updated.
  • Test coverage gaps: new branches without tests, tests that do not assert anything meaningful, and edge cases the test suite ignores.
  • API misuse: deprecated calls, wrong argument orders, and library patterns used in ways the documentation warns against.

What they miss or fake

The dangerous category is not what agents miss — it is what they pretend to have checked. Agents are fluent at producing review-sounding prose about code they did not truly analyze. Architectural problems sail through: the change works locally but fights the system's design, and the agent, lacking the design context, approves. Business logic correctness is the biggest gap — does this change do what the ticket actually asked? The agent can read the diff but rarely the intent. Subtle concurrency issues, race conditions, and performance cliffs in hot paths get vague approval. And the most insidious failure: code that reads well but is wrong. Agents are pattern-matchers trained on plausible code; plausible-but-wrong is their blind spot, and they will praise it confidently.

Setting up reviews that find real issues

  • Give it the diff and the ticket. A diff without the intent behind it can only be judged on style. Include the PR description, linked issue, and what the change is supposed to accomplish.
  • Ask for specific categories. Correctness, security, performance, and maintainability as explicit headings force the agent past generic observations.
  • Require file and line references. Every finding must cite its location. Findings without locations are usually findings without substance.
  • Demand severity ratings. Critical, major, and nitpick. If everything is a nitpick, the review added nothing; if everything is critical, the calibration is broken.
  • Forbid vague praise. Ban phrases like looks good overall and clean implementation unless attached to something specific. Praise without specifics is the tell of a fake review.
  • Set a findings quota floor, not a ceiling. Ask for at least three substantive observations — it counteracts the agent's bias toward agreeable approval.

A review prompt that works

Structure beats cleverness. A prompt that consistently produces useful reviews looks like this: state the change's intent in one paragraph (paste the ticket summary), then instruct the agent to review the diff for correctness against that intent first, security issues second, and style last — in that priority order. Require every finding to include the file, the line, a severity, and a concrete suggestion, not just a complaint. Explicitly invite the agent to say when it is uncertain rather than guessing, and to flag anything it cannot verify from the diff alone. End with: list the three riskiest aspects of this change, even if you found nothing wrong. That final instruction is the most valuable line — it forces the agent to think adversarially instead of defaulting to approval.

Where humans stay in the loop

  • Architecture decisions. Whether the approach is right is a judgment call that needs system context no diff contains.
  • Security-sensitive changes. Agent review is a useful first pass, but authentication, authorization, and data-handling changes need human eyes before merging.
  • Anything the agent flagged as uncertain. Uncertainty flags are the system working as designed — treat them as mandatory human review triggers.
  • First contributions from new team members. The review is also mentoring; an agent cannot calibrate feedback to someone's experience level.
  • Changes where the tests are also AI-generated. AI reviewing AI-tested AI-written code is a closed loop with no ground truth — a human must break the circle.

The right mental model is an indefatigable junior reviewer: fast, thorough on the mechanical stuff, honest about uncertainty, and never the final word on anything important. Set it up with the ticket context, demand specifics, track whether it actually flags things, and keep humans on the decisions that matter. Done this way, agent code review removes the most tedious part of the process and leaves the interesting part — the thinking — to people.

Calibrating the reviewer with examples

A code-review agent with no examples of what good feedback looks like will produce generic commentary — the equivalent of looks good to me with extra words. Calibrate it with few-shot examples: two or three real diffs from your codebase paired with the review comments a senior engineer actually left. Include one example where the right response was approval with no comments, so the agent learns that silence is sometimes correct. Include a subtle bug the reviewer should catch and a style nit it should ignore, to set the threshold. These examples do more than any instruction paragraph to establish the bar, because they show rather than tell. Refresh the examples quarterly: as the codebase evolves, the interesting bug classes change, and a reviewer calibrated on last year's patterns misses this year's.

  • Use real diffs from your own codebase — generic examples teach generic reviewing.
  • Show the full range: approval with no comments, minor nits, and serious catches.
  • Include a deliberate false-positive example: a comment a human would not leave, marked as over-eager.
  • Keep examples short. Three tight examples beat ten sprawling ones the agent will skim.
  • Version your examples alongside the review prompt so improvements are traceable.

Triaging the reviewer's output

Even a calibrated reviewer produces false positives, and how you handle them determines whether the team keeps listening. Route findings into three buckets: definite issues (the code is wrong — these block), probable issues (worth a human look — these get a quick triage pass), and style observations (never block; batch them or feed them to a formatter). Track the false-positive rate per bucket: if probable issues are wrong more than half the time, developers will stop reading them, and the bucket becomes noise. The fix is usually tightening the prompt's confidence language — instruct the agent to label its certainty and to stay silent below a threshold — rather than adding more review categories. A reviewer that flags three real bugs per week and nothing else is infinitely more valuable than one that flags thirty things, twenty-seven of them noise.

Security-focused review passes

General review prompts underperform on security because security bugs require adversarial thinking, not pattern matching. Run a dedicated security pass with a different prompt: one that asks what an attacker could do with this code, not whether the code looks right. Give it your threat model — what data matters, where trust boundaries lie, which inputs are attacker-controlled — because without that context the agent reviews against a generic checklist. Scope the pass to changed code that touches trust boundaries: authentication, authorization, input handling, cryptography, and data access. Full-codebase security review by an agent produces overwhelming noise; targeted passes on risky diffs produce findings worth reading. And always have a human confirm security findings before acting — the cost of a false alarm is an hour, but the cost of a missed real issue is the incident.

Integrating into CI without noise fatigue

The graveyard of review automation is full of bots that commented on every pull request until the team muted them. Integrate surgically: run the agent on pull requests above a size threshold or touching sensitive paths, not on every typo fix. Post findings as review comments on the relevant lines, not as a single wall-of-text summary — inline comments get read, summaries get skimmed. Deduplicate against existing findings: if the same issue was flagged and dismissed before, do not flag it again without new information. Give developers a one-click way to mark findings as false positives, and feed those back into calibration. Measure the metric that matters: the fraction of flagged issues that developers actually fix. If that number is healthy, the integration is working; if it decays, the bot is becoming wallpaper and needs recalibration, not more rules.

Reviewing non-code artifacts

The review agent's skills transfer to everything engineers produce that is not code: configuration files, infrastructure definitions, database migrations, documentation, and — recursively — prompts themselves. Configuration review catches the misconfigured timeout, the overly permissive security rule, the environment variable pointing at production. Migration review checks for destructive operations, missing rollbacks, and lock contention on large tables. Documentation review verifies that examples actually work and that the docs describe the current behavior rather than last year's. Prompt review — using a calibrated agent to review another agent's instructions — catches ambiguity, conflicting directives, and missing edge cases before they become production behavior. Each artifact type needs its own calibration examples, because the bug classes differ completely. But the infrastructure is shared: the same CI integration, the same triage buckets, the same fix-rate metric. Extend the reviewer's jurisdiction gradually, one artifact type at a time, calibrating each before moving on.

  • Configuration: misconfigured timeouts, permissive rules, wrong-environment pointers.
  • Migrations: destructive operations, missing rollbacks, lock risks on large tables.
  • Documentation: examples that fail, descriptions of behavior that changed.
  • Prompts themselves: ambiguity, conflicting directives, missing edge cases — review recursively.
  • Calibrate each artifact type separately; bug classes differ completely across types.

Keep reading