Tag

Coding Agents

A bold vintage screen-print illustration in deep green-teal, mustard and salmon-cream of a large magnifying glass over a paper checklist with tick-boxes.

Development

Developers Trust AI When It's Easy to Check, and Code That Runs Passes

Stack Overflow's 2026 survey shows developers trust AI only when they can easily validate it. Three September studies suggest the easiest check, running the code, misses what AI gets wrong.

A bold vintage screen-print illustration in pale cream, deep pine green and rust orange of a large calibration dial gauge crossed by a socket wrench, set against a scoreboard-tile band and diagonal stripes.

Development

The SWE-Bench Leaderboard Can No Longer Tell Models Apart

Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

A bold vintage screen-print illustration in cream, deep teal and brick red of a raised diagonal-striped boom barrier gate pivoting open on a circular hinge post above a tall stack of printed code-listing pages.

Development

Spotify Says Its PR Thresholds No Longer Apply

Spotify's 16 September retrospective argues its PR-size and complexity warning signs may be outdated, not broken, but nobody has re-derived them.

A bold vintage screen-print illustration in terracotta, dark ink and warm ochre of a long paper receipt unrolling from a geometric spool, torn off after only a short filled-in stub while the rest of the ribbon runs blank.

Development

The Coding Agent's Self-Report Covers One Action in Eleven

A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

A vintage screen-print illustration in cream, charcoal and dusty rose of an oversized faceted pencil-eraser wheel with a snapped brake arm, grinding a hole through a printed manuscript page.

Development

The Repair Loop Has No Brakes: Fixing Code That Isn't Broken

A 9 September 2026 study finds AI self-repair breaks working code faster than it fixes real bugs once nothing external tells the model when to stop.

A bold vintage screen-print illustration in kraft-gold, deep aubergine and muted sage green of a tall stack of cardboard shipping parcels bound in crossing packing tape, the topmost parcel bearing a large round provenance seal-label still fully intact and unopened.

Development

AI Coding Assistants Skip the Labels Before They Install

A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

A vintage screen-print illustration in charcoal, aged parchment cream and brick red of a tilted printed specification document struck by a jagged red pen line that cracks open a stack of code blocks below it.

Development

Requirements After the First Edit Cost Coding Agents Double

A September 2026 study of 3,553 coding-agent sessions finds requirements arriving after coding starts double the agent's rework — warning it in advance doesn't help.

A bold vintage screen-print illustration in slate-navy, cream and burnt terracotta of an oversized mechanical keyboard with one central key pulled from its slot and struck through by a diagonal cross.

Creative Tooling

OpenAI Cutting Off Cursor Shows Who Really Owns Your Coding Tool

OpenAI is severing Cursor's access to GPT models on 12 November over SpaceX's acquisition of the company — a reminder the model picker is a contract, not a setting.

A bold vintage screen-print illustration in cream, deep slate-blue and red of an open instruction manual, its ruled checklist page intact with red-pen checkmarks while its facing map page shatters into scattered geometric fragments.

Development

AGENTS.md Works as a Rulebook and Fails as a Tour

ETH Zurich tested 138 AGENTbench instances plus 300 SWE-bench Lite cases and found AGENTS.md's instructions change agent behavior, but the repo overview doesn't, while adding 20% to run cost.

A bold vintage screen-print illustration in cream, deep brown and rust red of an open recipe card file box holding a row of identical printed cards, with one handwritten card fallen outside the box.

Creative Tooling

Figma's Agent Skills Sell Personalization the Data Doesn't Back

Figma's 13 August skill-authoring launch is pitched on capturing personal taste, but a study three days earlier found generic skills beat personalized ones.

A bold vintage screen-print illustration in parchment cream, deep navy and terracotta red of two crossing pens over a manuscript page, trailing two overlapping zigzag correction lines that visibly contradict each other.

Development

SWE-Touch: The Edit You Make While an Agent Still Runs

SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

A bold vintage screen-print illustration in cream, deep teal-charcoal and dusty slate blue of a large mechanical split-flap scoreboard with mismatched abstract flip-panel tiles and a small stopwatch beside it.

Development

Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree

Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.