Engineering · AI
AI Code Review: What It Catches and What It Misses
Mihajlo Petrović5 min read
An honest breakdown of where an AI reviewer genuinely outperforms a tired human, where it is blind by construction, and the three configuration mistakes that make a team mute the bot in week two.
Running an AI reviewer over a pull request is one of the easiest wins available right now. It's also one of the easiest ways to make a team ignore code review entirely, if you set it up carelessly.
I've been running agents over my own diffs for a while — both on side projects where I'm the only reviewer, and as a first pass before human review at work. Here's an honest breakdown of the two categories, and how to configure it so people don't mute the bot in week two.
What It Reliably Catches
The failure paths you didn't think about. Null and undefined access, an unhandled promise rejection, a catch block that logs and continues as if nothing happened, an array index that's fine until the list is empty. This is the bread and butter, and it's genuinely good at it.
Inconsistency with your own codebase. Given the surrounding files, it will notice that every other service returns a typed result object and this one returns a raw response. Humans miss this constantly — especially in a large PR, especially at 5pm on a Friday.
Duplication. "This does what formatCurrency in shared/utils already does." A reviewer who has read the whole repo in the last thirty seconds has an advantage over one who wrote that helper eight months ago.
Security smells. String-concatenated SQL, a secret in a config file, a missing authorisation check on a route where its siblings all have one, unescaped user input rendered as HTML. It won't find your subtle auth bypass, but it catches the ones that show up in incident reports.
Tests that assert the implementation. A test that mocks the thing under test and asserts the mock was called is worse than no test — it fails when you refactor and passes when the behaviour breaks. AI reviewers are unusually good at spotting this, maybe because it's a pattern rather than a judgement.
The boring stuff nobody wants to comment on. Missing await, inconsistent naming, dead code, a TODO left in. Low value individually; not free when a human has to write them all out.
What It Misses
Whether this should exist at all. The most valuable review comment is sometimes "we already have a service that does this, and it belongs there". That requires knowing where the codebase is going, not where it is.
Cross-service invariants. "This endpoint now returns null for deleted users, and the mobile app crashes on null here." Nothing in the diff says that. It lives in the heads of the people who've been on the team a year.
Real performance. It'll flag an O(n²) loop. It will not tell you that this particular query runs on every page load for a table with 40 million rows, or that the index you rely on doesn't exist in production.
Product intent. The code is correct and implements the wrong requirement. No amount of reading the diff reveals this.
Domain-specific correctness. In banking: is this rounding rule right for this product? Does this comply with the process the regulator signed off on? Should this field be logged at all? These are the questions that actually matter in my day job, and they're exactly the ones an agent has no way to answer.
Notice the shape: it's strong on local, mechanical, pattern-shaped issues and weak on global, contextual, intent-shaped ones. Which is the same distribution as everything else in this space.
Making It Useful Instead of Annoying
Three failure modes kill AI review in practice, and all three are configuration problems.
1. Noise
Twenty comments on a five-line PR, most of them nitpicks, and the team stops reading. Fixes, in order of impact:
- Set a severity floor. Only report things you'd actually block a merge on, or that carry a real risk. "Consider extracting this into a variable" is not review, it's chatter.
- Cap the count. Ten findings maximum, ranked. If there are genuinely more, the PR is too big — which is itself the finding.
- Let the formatter own formatting. Style belongs to a deterministic tool. Never to a reviewer, human or otherwise.
2. No context
An agent that sees only the diff will confidently suggest things your conventions forbid. Give it what a human reviewer has: the project instructions file, the surrounding files (not just changed lines), and the PR description. Quality goes up sharply and false positives drop.
3. Treating it as a gate
The moment the bot can block a merge, people optimise for satisfying the bot. It's a first pass, not an approver: it clears the mechanical layer so that human review can spend its attention on design, intent, and the things only a human knows.
Where I Actually Run It
Before I push. The highest-value moment by far, and a fixed step in my daily workflow. Reviewing my own diff before anyone sees it means the embarrassing stuff never gets a public comment, and the loop is seconds instead of hours.
On the PR, as a comment rather than a review — findings ranked, low-confidence ones excluded.
Never on someone else's PR without saying so. Pasting an agent's findings as if they were your own review is a bad look, and if one is wrong you'll be defending an opinion you never held.
The Uncomfortable Bit
The most common use of AI review right now is reviewing code that an AI wrote. That's worth being honest about: two systems with correlated blind spots checking each other. It catches real bugs — but it is not a substitute for someone who understands the system reading the diff and asking "wait, why?"
Which is the same conclusion as everywhere else in this stack. The tooling raises the floor. It doesn't move the ceiling, and it doesn't transfer the accountability.
- #code review
- #ai agents
- #pull requests
- #quality
- #ai
Written by
Mihajlo Petrović
Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.