Category
Development
28 articles

Development
AI Code Reviewers Can Coach Attackers to Approval
A 4 October 2026 paper shows an AI pull-request reviewer's own comments guide an attacker to approval with the exploit intact, just as GitHub makes Copilot's approval count.

Development
Developers Trust AI When It's Easy to Check, and Code That Runs Passes
Stack Overflow's 2026 survey shows developers trust AI only when they can easily validate it. Three September studies suggest the easiest check, running the code, misses what AI gets wrong.

Development
MCP Error Messages Written for Humans Hurt the Smartest Agents Most
A 28 September 2026 preprint finds MCP servers' developer-facing error steps make capable agents fail more, and that naming a server tool in the step fixes it.
Development
PixelLeak: The Screenshot Request That Leaked Internal UIs
Glow's PixelLeak report found 13,000+ internal screenshots in public GitHub repos. The cause was a routine review request, and the fix sits in your agent's skill files.

Development
The SWE-Bench Leaderboard Can No Longer Tell Models Apart
Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

Development
Spotify Says Its PR Thresholds No Longer Apply
Spotify's 16 September retrospective argues its PR-size and complexity warning signs may be outdated, not broken, but nobody has re-derived them.

Development
The Coding Agent's Self-Report Covers One Action in Eleven
A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

Development
GitHub Copilot Can Approve Pull Requests, Then Close Its Own Comments
Two September 2026 GitHub Copilot updates let it approve pull requests and resolve its own review comments, a loop GitHub has yet to publish any accuracy data for.

Development
The Repair Loop Has No Brakes: Fixing Code That Isn't Broken
A 9 September 2026 study finds AI self-repair breaks working code faster than it fixes real bugs once nothing external tells the model when to stop.

Development
AI Coding Assistants Skip the Labels Before They Install
A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

Development
AI Maturity Ladders Are Measuring the Anxiety They Cause
A September 2026 case study finds a Danish firm's 1-to-5 AI maturity ladder amplified developer distress, grading usage before its policy on permitted tools existed.

Development
Requirements After the First Edit Cost Coding Agents Double
A September 2026 study of 3,553 coding-agent sessions finds requirements arriving after coding starts double the agent's rework — warning it in advance doesn't help.