Category
Development
16 articles

Development
AGENTS.md Works as a Rulebook and Fails as a Tour
ETH Zurich tested 138 AGENTbench instances plus 300 SWE-bench Lite cases and found AGENTS.md's instructions change agent behavior, but the repo overview doesn't, while adding 20% to run cost.

Development
Multi-Agent Coding Teams Don't Need a Boss, a Study Finds
A 1,902-run study of Claude Code agent teams found naming a coordinator adds no measurable benefit, while shared-file versus messaging coordination swings token costs by up to 42%.

Development
AgenTag: AI Pull Request Tells Are in the Prose, Not the Code
AgenTag's 2 August 2026 study found AI-authorship signal in pull requests comes almost entirely from PR descriptions, not code diffs — and rewriting the text defeats it.

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

Development
Wealthfront's AI Code Reviewer Costs $4 a Pull Request to Say Nothing
Wealthfront rebuilt its code reviewer so three AI models argue before any comment reaches a human, and now counts silence on most pull requests a win.

Development
Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree
Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

Development
GitClear's Maintainability Gap Puts a Price on AI's PR Boom
GitClear and GitKraken's 623-million-change study finds AI coding lifts pull requests but pushes code duplication up 81% as legacy maintenance falls 74% since 2023.

Development
GitHub Copilot's Two Modes Work Great Separately, Badly Together
A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

Development
Spec-Driven Development Can't Spec the Governance That Matters
A 420-KLOC case study challenges spec-driven coding's premise, arguing real governance is discovered from failures mid-project, not written into a spec upfront.

Development
GhostCommit Shows AI Reviewers and Agents Don't See Alike
GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.

Development
Your Coding Agent Trips the Same Alarms as an Intruder
Sophos telemetry from June 2026 shows Claude Code, Cursor and OpenAI Codex tripping the same rules built to catch attackers, just as GitHub ships an auto-approve mode.

Development
Amazon and Meta Killed Their AI Coding Leaderboards
Amazon's KiroRank and Meta's Claudeonomics ranked engineers by AI tokens consumed, until gamed usage inflated costs and both companies quietly shut the boards down.