Tag
Coding Agents
19 articles

Development
Developers Trust AI When It's Easy to Check, and Code That Runs Passes
Stack Overflow's 2026 survey shows developers trust AI only when they can easily validate it. Three September studies suggest the easiest check, running the code, misses what AI gets wrong.

Development
The SWE-Bench Leaderboard Can No Longer Tell Models Apart
Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

Development
Spotify Says Its PR Thresholds No Longer Apply
Spotify's 16 September retrospective argues its PR-size and complexity warning signs may be outdated, not broken, but nobody has re-derived them.

Development
The Coding Agent's Self-Report Covers One Action in Eleven
A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

Development
The Repair Loop Has No Brakes: Fixing Code That Isn't Broken
A 9 September 2026 study finds AI self-repair breaks working code faster than it fixes real bugs once nothing external tells the model when to stop.

Development
AI Coding Assistants Skip the Labels Before They Install
A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

Development
Requirements After the First Edit Cost Coding Agents Double
A September 2026 study of 3,553 coding-agent sessions finds requirements arriving after coding starts double the agent's rework — warning it in advance doesn't help.

Creative Tooling
OpenAI Cutting Off Cursor Shows Who Really Owns Your Coding Tool
OpenAI is severing Cursor's access to GPT models on 12 November over SpaceX's acquisition of the company — a reminder the model picker is a contract, not a setting.

Development
AGENTS.md Works as a Rulebook and Fails as a Tour
ETH Zurich tested 138 AGENTbench instances plus 300 SWE-bench Lite cases and found AGENTS.md's instructions change agent behavior, but the repo overview doesn't, while adding 20% to run cost.

Creative Tooling
Figma's Agent Skills Sell Personalization the Data Doesn't Back
Figma's 13 August skill-authoring launch is pitched on capturing personal taste, but a study three days earlier found generic skills beat personalized ones.

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

Development
Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree
Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.