All articles
64 published

Prototyping
AI Prototypes Lost the Signal That Said They Were Disposable
AI prototypes now look shipped, not sketched, which studies from Mozilla and Purdue show gets users to react honestly but stops teams asking for changes.

Development
The SWE-Bench Leaderboard Can No Longer Tell Models Apart
Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

Development
Spotify Says Its PR Thresholds No Longer Apply
Spotify's 16 September retrospective argues its PR-size and complexity warning signs may be outdated, not broken, but nobody has re-derived them.

Design Engineering
WebMCP's Second Interface: Nothing Keeps It True to the First
Chrome's WebMCP trial and Shopify's default agent tools give sites a second, agent-only interface, but nothing enforces that it still matches the one people click.

Development
The Coding Agent's Self-Report Covers One Action in Eleven
A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

Development
GitHub Copilot Can Approve Pull Requests, Then Close Its Own Comments
Two September 2026 GitHub Copilot updates let it approve pull requests and resolve its own review comments, a loop GitHub has yet to publish any accuracy data for.

Design Engineering
Axe-core Was Built Never to Be Wrong. That's Its Blind Spot
A 8 September 2026 study finds axe-core misses most real WCAG violations because it is tuned never to raise a false alarm, and Figma's checker inherits the same blind spot.

Development
The Repair Loop Has No Brakes: Fixing Code That Isn't Broken
A 9 September 2026 study finds AI self-repair breaks working code faster than it fixes real bugs once nothing external tells the model when to stop.

Development
AI Coding Assistants Skip the Labels Before They Install
A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

Product Design
The AI Plan You Approve Is the One You Check Least
A 8 September 2026 study of six AI reasoning formats found the plan-and-decomposition style users prefer nearly triples the false alarms of a plain step-by-step trace.

Prototyping
Generative UI Tools Are Benchmarked on a Turn, Used in a Session
Two 2026 studies (EvoGenUI-Bench, Maru) find generative-UI sessions degrade because each fix silently undoes an earlier one — not because models misread later prompts.

Development
AI Maturity Ladders Are Measuring the Anxiety They Cause
A September 2026 case study finds a Danish firm's 1-to-5 AI maturity ladder amplified developer distress, grading usage before its policy on permitted tools existed.