A 12 August 2026 benchmark, “Does It Render Everywhere?”, rendered 203 AI-generated webpages across nine real browser-and-device combinations and found 68% broke somewhere — 1.7 times the 40% failure rate on human-written pages from the same datasets. The per-tool spread is what the two hosts spend most of the episode on: 26% for Vercel’s v0, 79% for Cursor, 100% for a raw GPT-5.1 call, with the largest failure category — pages that shrink to fit but leave text unreadably small — the kind that passes a diff read and a desktop preview glance without raising any flags.
They also take seriously the benchmark’s own complications: six of the eight generators tested are one-shot calls or academic pipelines nobody ships, Cursor’s own documentation describes a browser-feedback loop the study’s neutral prompts never invoked, and the 40% human baseline means the underlying problem predates AI. None of that changes what the render-level failures are, only how much of the average belongs to any one tool.