Category

Development

A bold vintage screen-print illustration in cream, olive green and brick red of a printed code listing with a large red checkmark and a round approval stamp.

Development

AI Code Reviewers Can Coach Attackers to Approval

A 4 October 2026 paper shows an AI pull-request reviewer's own comments guide an attacker to approval with the exploit intact, just as GitHub makes Copilot's approval count.

A bold vintage screen-print illustration in deep green-teal, mustard and salmon-cream of a large magnifying glass over a paper checklist with tick-boxes.

Development

Developers Trust AI When It's Easy to Check, and Code That Runs Passes

Stack Overflow's 2026 survey shows developers trust AI only when they can easily validate it. Three September studies suggest the easiest check, running the code, misses what AI gets wrong.

A bold vintage screen-print illustration in cream, slate indigo and rust red of a large geometric padlock with a blank sticky note on its body and a small key lying out of reach beside it.

Development

MCP Error Messages Written for Humans Hurt the Smartest Agents Most

A 28 September 2026 preprint finds MCP servers' developer-facing error steps make capable agents fail more, and that naming a server tool in the step fixes it.

A bold vintage screen-print illustration in cream, deep teal and ochre of a bulletin board with two paper printouts, a before and an after sheet, each held by a round push pin.

Development

PixelLeak: The Screenshot Request That Leaked Internal UIs

Glow's PixelLeak report found 13,000+ internal screenshots in public GitHub repos. The cause was a routine review request, and the fix sits in your agent's skill files.

A bold vintage screen-print illustration in pale cream, deep pine green and rust orange of a large calibration dial gauge crossed by a socket wrench, set against a scoreboard-tile band and diagonal stripes.

Development

The SWE-Bench Leaderboard Can No Longer Tell Models Apart

Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

A bold vintage screen-print illustration in cream, deep teal and brick red of a raised diagonal-striped boom barrier gate pivoting open on a circular hinge post above a tall stack of printed code-listing pages.

Development

Spotify Says Its PR Thresholds No Longer Apply

Spotify's 16 September retrospective argues its PR-size and complexity warning signs may be outdated, not broken, but nobody has re-derived them.

A bold vintage screen-print illustration in terracotta, dark ink and warm ochre of a long paper receipt unrolling from a geometric spool, torn off after only a short filled-in stub while the rest of the ribbon runs blank.

Development

The Coding Agent's Self-Report Covers One Action in Eleven

A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

A bold vintage screen-print illustration in cream, deep navy and muted sage green of a large identification badge on a lanyard, its checkmark medallion pinned to a nameplate left deliberately blank.

Development

GitHub Copilot Can Approve Pull Requests, Then Close Its Own Comments

Two September 2026 GitHub Copilot updates let it approve pull requests and resolve its own review comments, a loop GitHub has yet to publish any accuracy data for.

A vintage screen-print illustration in cream, charcoal and dusty rose of an oversized faceted pencil-eraser wheel with a snapped brake arm, grinding a hole through a printed manuscript page.

Development

The Repair Loop Has No Brakes: Fixing Code That Isn't Broken

A 9 September 2026 study finds AI self-repair breaks working code faster than it fixes real bugs once nothing external tells the model when to stop.

A bold vintage screen-print illustration in kraft-gold, deep aubergine and muted sage green of a tall stack of cardboard shipping parcels bound in crossing packing tape, the topmost parcel bearing a large round provenance seal-label still fully intact and unopened.

Development

AI Coding Assistants Skip the Labels Before They Install

A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

A bold vintage screen-print illustration in parchment cream, deep slate-teal and muted ochre of a large geometric ladder climbing diagonally across the frame, with a small checklist clipboard at its top rung.

Development

AI Maturity Ladders Are Measuring the Anxiety They Cause

A September 2026 case study finds a Danish firm's 1-to-5 AI maturity ladder amplified developer distress, grading usage before its policy on permitted tools existed.

A vintage screen-print illustration in charcoal, aged parchment cream and brick red of a tilted printed specification document struck by a jagged red pen line that cracks open a stack of code blocks below it.

Development

Requirements After the First Edit Cost Coding Agents Double

A September 2026 study of 3,553 coding-agent sessions finds requirements arriving after coding starts double the agent's rework — warning it in advance doesn't help.