Nick Inzucchi, a product designer at Cursor, put the problem plainly in the AI in Design Report: “Code forces you to commit to your first idea and go deep—at the expense of the broad exploration that Figma made easy.” He added that breadth is “the limitation I feel most” . The tension is easy to state. AI design tools hand everyone the same safe-looking screen, and when researchers forced one to offer real alternatives, people complained less but had to fix more. Nothing showed they shipped more.
That result comes from “Enabling Creative Exploration for Vibe Design Agents” , posted on 14 September 2026 and revised on 25 September by Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos and Arnaud Benard. Three of the authors are affiliated with Google. The system under test is an unnamed “commercial UI design assistant.”
The same theme, five times over
Ask a design agent for a screen and you get its most typical answer. Ask again and you get it again. On the paper’s benchmark of 83 prompts, each run five times, the baseline agent produced exactly 1.00 theme option across repeats. Same direction, every time.
Sascha Becker traces this to “typicality bias” in the human preference data models are trained on. His summary: the model supplies the average “competently and instantly.” You supply the reason anyone should remember the page.
The obvious lever is temperature, the dial that adds randomness to each word the model picks. The authors say it fails here. Raising it “changes choices throughout the output, including both aesthetic decisions and implementation details, so it does not selectively control design direction.” You get a stranger palette and a broken button handler together.
Propose first, then pick one on purpose
The paper’s fix separates the two jobs. The agent first proposes structured design directions, then a selection step samples one deliberately. At τ=2.0 the pipeline produced 2.87 theme options on the same benchmark, and screenshot similarity fell from 0.6765 to 0.5438.
Those are offline measures of breadth. Read across to practice, they suggest exploration can be engineered back into prompt-to-code tools. For a product designer prototyping this way, that matters. Without forced alternatives, the “explore three concepts” step of a design review quietly stops happening, because each prompt returns one confident direction.
Exploration can be engineered back into these tools, but whether an alternative was worth seeing is a judgment product telemetry cannot score.
Live, the numbers stop agreeing
The online experiment covered more than 300,000 tasks. Negative feedback fell 31.51%, from 73 events to 50. Code exports rose 8.23%, but the 95% confidence interval runs from -13.03% to +29.48%, and the authors write that “the experiment does not establish an export improvement.”
The correction rate rose from 38.8% to 41.6%, and completions within 60 seconds fell 3.47%. The authors say more corrections could mean “aesthetic mismatch, unmet requirements, or another source of friction.” Useful divergence and plain friction look identical in a log.
Whose judgment counts as the metric
The paper’s own limits weigh against reading this as proof that users wanted variety. It ran no professional-designer evaluation. Its LLM judges carry position and verbosity biases. The authors also say the best settings differ across interventions and prompt suites, so the defaults are not universal.
There are stronger objections to the premise. Karri Saarinen, Linear’s CEO, argues that “design is the planning stage and code is the implementation stage. I don’t like mixing those two.” On that view, exploration belongs outside the code tool. Becker argues the remedy is human art direction, not smarter sampling.
Both may be right, and neither cancels the finding. Pipeline’s reading is that the tools score themselves on shipping metrics: exports, latency, complaints. Those answer whether something got built. They cannot answer whether the designer saw an option that changed the decision. Mark Boyes-Smith, Miro’s Head of AI Design, says he wants people to have “radically divergent concepts” , not to prompt, shuffle and call it done.
That gap connects to earlier Pipeline coverage. Session-level benchmarks already show that single-turn scores hide what a whole workflow does, and prototypes that no longer look rough have already cost teams some of their license to keep revising. Exploration is the next casualty, and the dashboard doesn’t have a column for it.
The model will keep supplying the average. Getting it to offer something else is now an engineering problem. Knowing whether the something else was worth the detour still takes a designer at the table.



