All articles
64 published

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.
Product Design
The AI Sparkle Icon Meets Europe's New Disclosure Law
EU AI Act Article 50, effective 2 August 2026, needs a persistent AI-made label — a job the industry's shared sparkle icon, unrecognizable as AI in testing, was never built for.

Development
Wealthfront's AI Code Reviewer Costs $4 a Pull Request to Say Nothing
Wealthfront rebuilt its code reviewer so three AI models argue before any comment reaches a human, and now counts silence on most pull requests a win.

Creative Tooling
AI Coding Tools Helped Blind Developers. Now Their Interfaces Are the Barrier
A 5 August 2026 study finds AI coding interfaces are a new accessibility barrier, and maintainer attention — not model quality — decides which tool a blind developer can use.

Development
Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree
Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

Design Engineering
MCP Apps' Final Spec Turns Design Systems Into Graceful Degradation
MCP's 28 July 2026 spec graduated MCP Apps to eleven clients, letting hosts theme a team's UI with tokens that carry no guarantee of arriving at all.

Product Design
UX.md and DESIGN.md Reveal What AI-Ready Documents Leave Out
Nielsen Norman Group published contradicting July 2026 essays on AI design documents, exposing the alignment work curated context alone doesn't do.

Design Engineering
Design Systems Need Evals to Check if AI Agents Obey Them
A practitioner argues design systems need CI-run evals, since evidence on AI instruction-following suggests agents may ignore documented rules more than teams assume.

Development
GitClear's Maintainability Gap Puts a Price on AI's PR Boom
GitClear and GitKraken's 623-million-change study finds AI coding lifts pull requests but pushes code duplication up 81% as legacy maintenance falls 74% since 2023.

Development
GitHub Copilot's Two Modes Work Great Separately, Badly Together
A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

Development
Spec-Driven Development Can't Spec the Governance That Matters
A 420-KLOC case study challenges spec-driven coding's premise, arguing real governance is discovered from failures mid-project, not written into a spec upfront.

Development
GhostCommit Shows AI Reviewers and Agents Don't See Alike
GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.