Development

AI Code Reviewers Can Coach Attackers to Approval

A 4 October 2026 paper shows an AI pull-request reviewer's own comments guide an attacker to approval with the exploit intact, just as GitHub makes Copilot's approval count.

Anyone who has had a pull request reviewed knows the routine: you read the comments, fix exactly what they point at, push again and ask for another look. The attack in a 4 October 2026 paper from Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova and Meng Wang is that routine, carried out by someone who wants the bug to survive. The clearer an AI reviewer is about what is wrong, the easier it becomes to edit the change until the reviewer approves it while the security hole stays open.

An attack that uses no trick at all

The authors built AFCRA-Bench, 159 pull requests drawn from real CVEs across 132 repositories, each paired with a working proof-of-concept exploit. An attack counts as a success only if the reviewer approves and the exploit still works, with up to five reviewer interactions allowed.

Against Claude Code running Sonnet 5 as the reviewer, the adaptive revision attack succeeded 33 times in 159 (20.8%). Against Codex running GPT-5.5 it succeeded 25 times (15.7%). The strongest text- or comment-based baseline managed 13 and 2, which is where the paper’s 2.5x and 12.5x come from.

The paper’s Figure 1 shows how. A vulnerable PR to a Node.js file-serving package was rejected. The reviewer asked for a directory-boundary check, and the attacker added one that blocks sibling-directory traversal. File access still followed symlinks, so an outside file stayed reachable, and the reviewer approved. In the authors’ words, “The repair satisfies the reported concern at the string-checking layer while the exploit survives at the filesystem layer.”

The mechanism is simple: “Feedback identifies a repair target, which the attacker fixes while preserving harmful behavior elsewhere.” And the more the reviewer says, the better it works. With Sonnet 5, a verdict alone gave the attacker 11.9%, adding locations 15.7%, and a full explanation 20.8%.

Attack success rate against Sonnet 5 by reviewer feedback level (%)
Attack success rate against Sonnet 5 by reviewer feedback level (%)
CategoryAttack success rate (%)
Verdict alone11.9
Locations added15.7
Full explanation20.8
Source: Gong, Zhang, Jahangirova and Wang, arXiv, 4 October 2026
The property that makes AI review worth paying for, specific and actionable feedback, doubles as a search signal for anyone iterating against it.

Why GitHub’s recent changes raise the stakes

Nobody tested Copilot. The paper measures two frontier models inside agent harnesses set up as reviewers, defending against an attacker who already holds an exploit for a known flaw. It names neither Copilot code review nor CodeRabbit.

Read across to shipped products, though, the mechanism transfers and the surroundings have changed. Since 1 September, GitHub says , a Copilot approval “counts toward the repository’s required-approvals rule” when an admin enables it. New commits dismiss that approval, after which a fresh review can be requested. Since 2 October, reviews can also be requested through the REST and GraphQL APIs , so a script can run the loop.

That is the same feedback-and-resubmit cycle, with a binding approval at the end of it. We covered the approval change in our look at Copilot approving pull requests ; this is its adversarial sequel. Maintainers and repo admins deciding whether to let that approval count, or to wire AI review into a script, are also deciding how many free, explained attempts an outside contributor gets against their security gate.

It also differs from GhostCommit , where an instruction hid inside an image. Here there is no hidden input to scan for. Every push is a legitimate fix.

The attacker is strong, and the defenses held

The counterpoint deserves its full weight. The paper’s attacker is handed a documented flaw and a working exploit, and gets one trajectory per condition. Every one of the 33 successful attacks on Sonnet 5 was caught by at least one tested defense, with the supplied exploit run catching 29 of them. None of the defenses rejected a genuine fix. The price: the paper reports $50 to $96 for the tested defenses against $12 for default review, a gap of roughly four to eight times, though what each figure covers is unclear.

There is a baseline argument too. The BleepingComputer report on GhostCommit cites UMKC researchers who found 73% of merged PRs in 300 active repositories got no substantive human or bot review. A gameable reviewer may still beat none.

Both points hold. Neither answers the new risk, which is that approval now carries weight. A reviewer nobody trusts costs little when fooled; one whose approval satisfies a merge rule does not. The property that makes AI review worth paying for, specific and actionable feedback, doubles as a search signal for anyone iterating against it. The paper’s own advice is to “give full feedback first, then withhold locations and explanations for security-sensitive issues”. That is an awkward bargain, since the explanations are what you bought.

A softer extension is pure speculation: the same habit of satisfying the comment rather than the problem might surface without malice when an authoring agent iterates against a reviewing one. The paper does not test that.

What it does show is narrower and harder to wave off. The review comment was always a map of what the reviewer cared about, and a contributor who wants the merge can follow it to the stamp.

This article was written by AI. How Pipeline works.