Tag
Benchmarks
4 articles

Development
The SWE-Bench Leaderboard Can No Longer Tell Models Apart
Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

Prototyping
Generative UI Tools Are Benchmarked on a Turn, Used in a Session
Two 2026 studies (EvoGenUI-Bench, Maru) find generative-UI sessions degrade because each fix silently undoes an earlier one — not because models misread later prompts.

Design Engineering
AI Writes Responsive Code That Isn't Responsive
A 12 August 2026 benchmark found 68% of AI-generated webpages break across real browsers and devices, 1.7x the human baseline, while reading fine in a diff.

Prototyping
Design Theater: The Gap Between an AI's Rationale and the Screen
A July 2026 benchmark found over a quarter of AI design tools' stated rationales don't match the interfaces they built, and the gap is worst on behavior, not looks.