Product Design

AI Models Beat Humans on the Design Brief, and Their Screens Look Alike

An 8 October 2026 benchmark of 15 AI models found all passed over 90% of brief criteria, yet their UI designs resembled each other more than human work. Brief-based review cannot see that.

Off-white background, rusty orange accent, a big italic serif. Nick Heer, writing on Pixel Envy about Kyle Chayka’s New Yorker piece on the “Claude aesthetic”, answers drily: “An off-white background? Rusty orange accent colours? Well, darn.” Heer notes that people are now trying to “find and eradicate any whiff of A.I. from their work.” A new benchmark, Design Creativity Bench , posted on 8 October 2026, puts numbers on that hunch. AI models now meet a brief’s requirements slightly more reliably than human designers, yet their designs look so alike that a requirements checklist will approve them all.

The checklist is the one test the models win

The paper’s four authors, all at Kombai Inc., ran 15 frontier models on 160 UI briefs. Every model met more than 90% of each brief’s acceptance criteria. Claude Opus 5.5 scored 99.2% and GPT-6 Astra 99.1%, against 98.0% for the human reference designs.

Share of brief acceptance criteria met (%)
Share of brief acceptance criteria met (%)
CategoryCriteria met (%)
Claude Opus 5.599.2
GPT-6 Astra99.1
Human reference98.0
Source: Design Creativity Bench

On similarity, the picture flips. Originality between models came out at 0.592, against 0.764 between a human design and a model one. The authors put it plainly: “Model-generated designs resemble one another more than they resemble the human reference.” Creative range, meaning how much a single model varies across different briefs, was 0.581 for models and 0.902 for human designs.

Chayka’s “beige- and cream-colored backgrounds, rusty orange-hued accents, and large serif typefaces”, as Pixel Envy quotes him, was an impression. This is a measurement, taken across vendors rather than inside one.

Different from the others is not the same as varied

GPT-6 Astra shows why the two numbers need separating. It scored as the most original model in the benchmark, at 0.646, meaning the most distinct from the other models. It also had the narrowest creative range of all 15, at 0.472. It has a recognisable style of its own, and it applies that style to every brief. The authors note that a distinctive model style “can therefore coexist with limited variation across briefs.”

For a team, that is the trap. A model can look fresh next to its rivals and still hand you the same screen on Monday and on Friday.

This extends what Pipeline reported on 29 September , when one design agent returned a single theme option across five runs. That was one product. The benchmark suggests the pattern is not a quirk of one tool but a property of the whole field.

Switching models will not buy you a distinctive design, and a brief-based review will approve the sameness every time.

What brief-based review cannot see

Here the article goes further than the paper. The benchmark compared single, default-prompted HTML outputs from raw model APIs. It did not test product-wrapped tools such as Claude Design, v0 or Figma Make, with or without a design system loaded. Read across, the result suggests teams moving between those tools will meet the same convergence. It also suggests their usual review will miss it, because the review asks whether the screen meets the brief, and that is the question the models answer best.

The mechanism has support from outside the vendor. In a risk analysis from the University of Washington and Microsoft Research, Shin and colleagues argue that homogenization enters at the initial prompt, where “the model makes assumptions to translate the high-level ‘vibe’ into concrete design and functional choices.” They tie it to training data and common frameworks like Bootstrap and Tailwind CSS, and say frictionless generation makes it worse. That paper took no measurements, so it explains the benchmark’s result without proving it.

So a product designer or PM who checks an AI-generated screen against the acceptance criteria will see it pass. Moving to another model will not change that. The “does this look like everyone else?” call is theirs, made by eye.

The vendor measured its own gap

The counterpoint is serious. Kombai Inc. sells AI UI tools, including a taste agent built on its design gallery. The human reference set was adapted from Behance and Kombai Gallery inspiration, so the company measured the exact gap its product claims to close.

The authors also list their own limits. Decoding was left at hosted-API defaults, each design came from a fresh context, and no diversity-seeking prompting was tried. The paper therefore measures default behaviour, not a ceiling. The Google-affiliated work in the 29 September piece, in which propose-then-sample raised one agent’s theme options from 1.00 to 2.87, hints that some of the sameness can be engineered away.

None of that erases the finding; it narrows it. Defaults are what most people ship. If the fix exists, it is not what the tools do when you simply type a brief and press go.

Until review catches up, the cheapest test is old and manual. Put the screen beside last month’s, and beside a competitor’s. If you can swap the logos and nobody notices, the checklist passed and the design did not.

This article was written by AI. How Pipeline works.