Design Engineering

37signals' Basecamp 5 Passed Every Pull Request Review, Then Broke

At 37signals, AI-agent pull requests for Basecamp 5 each passed review, but together they wrecked the architecture, exposing who wasn't watching the whole system.

“It’s solved,” DHH told the Lex Fridman Podcast earlier this year, describing the plan for 37signals’ final Basecamp 5 sprint. “We can just have the designers do the programming. They know what features they want. They know what shape they want it to take. Let them vibe.” Each pull request the designers’ AI agents opened looked defensible on its own, but stacked together, DHH said, “we ended up with a lot of PRs that individually perhaps could have been justified for a hot moment, but taken all together, destroyed the architecture of the system.” Programmers mopped it up by hand. Nothing that got reviewed was wrong; the thing that broke was never reviewed at all.

The Swiss cheese problem no single pull request could show

In late September, at Rails World 2026 in Austin, DHH returned to the story as proof of concept, not confession. Writing afterward, Jared Smith reports the architecture ended up, in DHH’s words, “a little like a Swiss cheese,” because “each pull request looked reasonable on its own. Twenty or thirty of them together left the architecture” full of holes. Caio Bianchi independently describes designers who “brought in dozens of agent-generated pull requests that looked reasonable separately but left the architecture looking like ‘Swiss cheese’ together.”

The pattern isn’t a one-company curiosity: half of all designers surveyed for Designer Fund and Foundation Capital’s AI in Design Report 2026 already say they’ve pushed AI-generated code to production, which means what broke at 37signals is already the default at half the organizations the survey reached.

Nothing that got reviewed was wrong, and the thing that broke was never reviewed at all.

A better model answers the wrong question

DHH’s own reading is the strongest objection to treating this as a process failure. He calls concluding that “the technology wasn’t ready” the wrong lesson, per Dealroom News’s coverage of the keynote — “things are quite different now,” he told the Lex Fridman Podcast — and per that same Dealroom coverage, the agent-accelerated Basecamp 5 has since shipped. Smith, the keynote’s sharpest critic, grants a real point too: Rails’ unusually strong conventions give agents an obvious way to do things, which he says makes a pull request cheaper for an agent to generate and cheaper for a human to review.

But Smith’s rebuttal lands harder than his concession. “A stronger model writing each pull request doesn’t touch the actual failure,” he argues, “which is that nobody owned the aggregate.” A cleaner diff is still a diff, reviewed one at a time; code review was built to catch a diff’s own bugs, not answer a portfolio question. Rails’ conventions and a better model can both narrow the damage a batch of pull requests does, without settling who was watching the batch as a whole.

Who gets left holding the aggregate

The clearest evidence for how narrow that unit of review is comes from outside design entirely. A study of 6,774 merged agent pull requests , against 5,044 human ones in the same repositories, found agent PRs needed a verified follow-up fix at 1.62 times the odds of human ones, with the same agent responsible for nearly 70% of those fixes. That measures pull requests catching their own mistakes later, nothing about architecture or designers. Read across to 37signals, though, the pattern suggests the failure isn’t one diff’s code quality; it’s what fix-at-a-time accounting can never catch, because coherence across dozens of diffs isn’t a property any single one carries.

Spotify’s own retrospective on reviewing agent-authored code found the same blind spot from inside: size and complexity thresholds that used to flag risky diffs stopped correlating with anything once agents did the typing, because the risk had moved to a level those thresholds never watched. It isn’t just a staffing gap — a study of multi-agent coding teams found naming a coordinator changed nothing measurable, since a title with no attached duty to approve or reject buys nothing. Ownership of the aggregate has to be an assigned job with real authority, not a role anyone backs into.

Picture two designers, a week apart, each asking an agent for a notification feature. One tucks a helper inside a view component; the other writes a near-identical one inside a controller. Both diffs are small, readable, and pass review without a raised eyebrow. By the time a third feature needs notifications, there are two incompatible patterns to copy, and no pull request anywhere contains the mistake — it exists only in the sum. That’s the mechanism Smith names directly: unmaintainable code, he writes, “doesn’t look unmaintainable at pull-request size.” It shows up at architecture size, weeks later, to whoever is staring at the whole file tree that day.

A product designer who opens pull requests through an AI coding agent gets each one approved on its own merits. The engineer beside them inherits the job of noticing, weeks later, that twenty approved changes no longer fit together — a job nobody gave them, the same shift the survey of 900-plus designers behind Pipeline’s earlier reporting found happening without anyone redrawing who’s accountable. DHH’s optimism isn’t wrong, just incomplete: better models will keep making each pull request harder to fault, which is exactly why the aggregate still needs someone whose job is to fault it anyway. Basecamp 5 shipped. The mop, evidently, still needed a human holding it.

This article was written by AI. How Pipeline works.