The Pipeline Mag Podcast

37signals' Basecamp 5 Passed Every Pull Request Review, Then Broke

At 37signals, AI-generated pull requests for Basecamp 5 each passed code review individually, but together they wrecked the system’s architecture, into what DHH himself described as “a little like a Swiss cheese.” He calls the fix a better model; critics, including keynote observer Jared Smith and independent account from Caio Bianchi, argue the real gap is that nobody owned the aggregate — a diff-by-diff review process was never built to answer a portfolio question.

The hosts trace that argument through a wider pattern: half of designers surveyed for the AI in Design Report 2026 already push AI-generated code to production, a study of 6,774 merged agent pull requests found they need follow-up fixes at 1.62 times the rate of human ones, and Spotify’s own review thresholds stopped correlating with risk once agents did the typing. A separate study found that simply naming a coordinator changes nothing without real authority attached — the same missing piece DHH’s optimism about better models still doesn’t supply.

This episode was made from the article 37signals' Basecamp 5 Passed Every Pull Request Review, Then Broke.