Two studies published days apart score generative-UI tools the way vendors do — on a single turn — and then measure what that misses. EvoGenUI-Bench finds a passing revision survives the next prompt only 66.5% of the time; Maru finds approval collapsing from 71% to 33% across a session, recovering to 61% only when the tool persists the interface’s structure instead of rebuilding it from scratch each turn.
We walk through both papers’ numbers, the case study where a UI updates correctly on screen while the logic underneath stays frozen, and the strongest objection — that commercial tools already edit a persistent file — along with the limits Maru’s own data puts on that fix.