Category

Development

A bold vintage screen-print illustration in warm parchment, deep teal and muted rust of a large faceted judge's gavel poised above its block, with three tiny orbs clustered near the hammer head.

Development

Wealthfront's AI Code Reviewer Costs $4 a Pull Request to Say Nothing

Wealthfront rebuilt its code reviewer so three AI models argue before any comment reaches a human, and now counts silence on most pull requests a win.

A bold vintage screen-print illustration in cream, deep teal-charcoal and dusty slate blue of a large mechanical split-flap scoreboard with mismatched abstract flip-panel tiles and a small stopwatch beside it.

Development

Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree

Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

A bold vintage screen-print illustration in aged cream, dusty slate-blue and rust terracotta of a boxy photocopier spilling an identical fan of duplicated paper sheets onto a filing cabinet.

Development

GitClear's Maintainability Gap Puts a Price on AI's PR Boom

GitClear and GitKraken's 623-million-change study finds AI coding lifts pull requests but pushes code duplication up 81% as legacy maintenance falls 74% since 2023.

A bold vintage screen-print illustration in mustard amber, deep plum and muted teal-green of a railway track switch splitting into two diverging rails beside a faceted semaphore signal.

Development

GitHub Copilot's Two Modes Work Great Separately, Badly Together

A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

A bold vintage screen-print illustration in aged paper cream, charcoal and red pen crimson of a tilted printed specification document marked up with a heavy diagonal strike-through, a circled correction, and a solid red sticky note.

Development

Spec-Driven Development Can't Spec the Governance That Matters

A 420-KLOC case study challenges spec-driven coding's premise, arguing real governance is discovered from failures mid-project, not written into a spec upfront.

A bold vintage screen-print illustration in indigo-charcoal, warm parchment and amber-gold of a large picture frame holding a faceted geometric ghost silhouette, set against horizontal code-printout bands.

Development

GhostCommit Shows AI Reviewers and Agents Don't See Alike

GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.

A bold vintage screen-print illustration in cream, deep charcoal-navy and alarm red of a circular smoke-detector alarm with concentric rings casting a triangular beam of red warning light down onto a laptop keyboard.

Development

Your Coding Agent Trips the Same Alarms as an Intruder

Sophos telemetry from June 2026 shows Claude Code, Cursor and OpenAI Codex tripping the same rules built to catch attackers, just as GitHub ships an auto-approve mode.

A constructivist vintage screen-print cover in cream, slate-blue and terracotta, with a large geometric bar chart whose tallest bars topple and shatter.

Development

Amazon and Meta Killed Their AI Coding Leaderboards

Amazon's KiroRank and Meta's Claudeonomics ranked engineers by AI tokens consumed, until gamed usage inflated costs and both companies quietly shut the boards down.

A constructivist vintage screen-print cover in dusty blue, navy and ochre, with a large geometric stopwatch.

Development

The Study METR Couldn't Run: What a Failed Control Group Reveals About AI Coding

METR tried to repeat its AI-productivity study in 2026 and couldn't recruit developers willing to work without AI — a methodological failure that may say more than any number could.

A constructivist vintage screen-print cover in oat, oxblood and teal, with a large geometric kill-switch toggle.

Development

Open Source's No-More-Pull-Requests Moment

Ladybird, tldraw, and the whole Jazzband collective have stopped taking public pull requests. It isn't a verdict on AI code quality — it's open source rebuilding its trust model from scratch.

A constructivist vintage screen-print cover in muted mint-grey, charcoal-teal and amber, with a large geometric magnifying glass.

Development

The End of Code Review, or Just Its Relocation?

A provocative paper declares human code review obsolete now that agents can do it faster. The evidence suggests something narrower and more interesting is actually happening.

A constructivist vintage screen-print cover in sand, forest green and rust, with a geometric git-merge symbol of branch lines converging.

Development

FrontierCode: The Benchmark That Asks Whether AI Code Is Ready to Merge

A new benchmark built with more than 20 open-source maintainers deflates the record-breaking numbers behind coding agents: even the best model clears only 13% of the hardest tasks.