Tag

Coding Agents

A bold vintage screen-print illustration in cream, deep slate-blue and red of an open instruction manual, its ruled checklist page intact with red-pen checkmarks while its facing map page shatters into scattered geometric fragments.

Development

AGENTS.md Works as a Rulebook and Fails as a Tour

ETH Zurich tested 138 AGENTbench instances plus 300 SWE-bench Lite cases and found AGENTS.md's instructions change agent behavior, but the repo overview doesn't, while adding 20% to run cost.

A bold vintage screen-print illustration in cream, deep brown and rust red of an open recipe card file box holding a row of identical printed cards, with one handwritten card fallen outside the box.

Creative Tooling

Figma's Agent Skills Sell Personalization the Data Doesn't Back

Figma's 13 August skill-authoring launch is pitched on capturing personal taste, but a study three days earlier found generic skills beat personalized ones.

A bold vintage screen-print illustration in parchment cream, deep navy and terracotta red of two crossing pens over a manuscript page, trailing two overlapping zigzag correction lines that visibly contradict each other.

Development

SWE-Touch: The Edit You Make While an Agent Still Runs

SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

A bold vintage screen-print illustration in cream, deep teal-charcoal and dusty slate blue of a large mechanical split-flap scoreboard with mismatched abstract flip-panel tiles and a small stopwatch beside it.

Development

Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree

Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

A bold vintage screen-print illustration in parchment cream, deep forest green and red-pen accent of a large circular rubber stamp bearing a checkmark, pressed onto a printed grid of UI component swatch cards.

Design Engineering

Design Systems Need Evals to Check if AI Agents Obey Them

A practitioner argues design systems need CI-run evals, since evidence on AI instruction-following suggests agents may ignore documented rules more than teams assume.

A bold vintage screen-print illustration in mustard amber, deep plum and muted teal-green of a railway track switch splitting into two diverging rails beside a faceted semaphore signal.

Development

GitHub Copilot's Two Modes Work Great Separately, Badly Together

A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

A bold vintage screen-print illustration in indigo-charcoal, warm parchment and amber-gold of a large picture frame holding a faceted geometric ghost silhouette, set against horizontal code-printout bands.

Development

GhostCommit Shows AI Reviewers and Agents Don't See Alike

GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.

A bold vintage screen-print illustration in warm cream, deep aubergine and mustard gold of an open ring-binder style guide with a fan of color-swatch chips spreading from its spine.

Creative Tooling

DESIGN.md Turns Brand Identity Into a Forkable File

Community projects now package Apple, Stripe and Nike's visual identity into MIT-licensed DESIGN.md files that any coding agent can install to generate on-brand UI.

A constructivist vintage screen-print cover in oat, oxblood and teal, with a large geometric kill-switch toggle.

Development

Open Source's No-More-Pull-Requests Moment

Ladybird, tldraw, and the whole Jazzband collective have stopped taking public pull requests. It isn't a verdict on AI code quality — it's open source rebuilding its trust model from scratch.

A constructivist vintage screen-print cover in muted mint-grey, charcoal-teal and amber, with a large geometric magnifying glass.

Development

The End of Code Review, or Just Its Relocation?

A provocative paper declares human code review obsolete now that agents can do it faster. The evidence suggests something narrower and more interesting is actually happening.

A constructivist vintage screen-print cover in sand, forest green and rust, with a geometric git-merge symbol of branch lines converging.

Development

FrontierCode: The Benchmark That Asks Whether AI Code Is Ready to Merge

A new benchmark built with more than 20 open-source maintainers deflates the record-breaking numbers behind coding agents: even the best model clears only 13% of the hardest tasks.