Tag

Claude Code

A bold vintage screen-print illustration in cream, olive green and brick red of a printed code listing with a large red checkmark and a round approval stamp.

Development

AI Code Reviewers Can Coach Attackers to Approval

A 4 October 2026 paper shows an AI pull-request reviewer's own comments guide an attacker to approval with the exploit intact, just as GitHub makes Copilot's approval count.

A bold vintage screen-print illustration in pale sage, deep ink blue and vermilion of a large rubber approval stamp pressing a checkmark mark onto a form sheet.

Creative Tooling

Claude Code Mods Can Approve Commands Before the Classifier Sees Them

Anthropic's new Claude Code mods can approve a tool call after rules and hooks decide, skipping the auto-mode classifier. The real safety layer is now plugin vetting, and solo developers carry it.

A bold vintage screen-print illustration in cream, deep teal and ochre of a bulletin board with two paper printouts, a before and an after sheet, each held by a round push pin.

Development

PixelLeak: The Screenshot Request That Leaked Internal UIs

Glow's PixelLeak report found 13,000+ internal screenshots in public GitHub repos. The cause was a routine review request, and the fix sits in your agent's skill files.

A bold vintage screen-print illustration in pale cream, deep pine green and rust orange of a large calibration dial gauge crossed by a socket wrench, set against a scoreboard-tile band and diagonal stripes.

Development

The SWE-Bench Leaderboard Can No Longer Tell Models Apart

Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

A bold vintage screen-print illustration in terracotta, dark ink and warm ochre of a long paper receipt unrolling from a geometric spool, torn off after only a short filled-in stub while the rest of the ribbon runs blank.

Development

The Coding Agent's Self-Report Covers One Action in Eleven

A 5,851-session study finds coding agents' self-written wrap-ups cite about one action in eleven, and drift back toward the plan exactly when it was abandoned.

A bold vintage screen-print illustration in kraft-gold, deep aubergine and muted sage green of a tall stack of cardboard shipping parcels bound in crossing packing tape, the topmost parcel bearing a large round provenance seal-label still fully intact and unopened.

Development

AI Coding Assistants Skip the Labels Before They Install

A pre-registered audit of 1,920 trials finds AI coding assistants open a provenance signal before installing 0.5% of the time, and never once verify one.

A bold vintage screen-print illustration in cream, deep indigo navy and dusty teal of a large two-button computer mouse centered between blocky choropleth-map quadrants, with a tiny closed chat bubble icon left unused in the corner.

Product Design

The Prompt Box Lost 94% of the Time to the Ordinary Mouse

An OOPSLA 2026 study put a chat box next to click-and-drag controls in the same editor and found users typed prompts for only 6% of their edits.

A bold vintage screen-print illustration in cream, deep slate-blue and red of an open instruction manual, its ruled checklist page intact with red-pen checkmarks while its facing map page shatters into scattered geometric fragments.

Development

AGENTS.md Works as a Rulebook and Fails as a Tour

ETH Zurich tested 138 AGENTbench instances plus 300 SWE-bench Lite cases and found AGENTS.md's instructions change agent behavior, but the repo overview doesn't, while adding 20% to run cost.

A bold vintage screen-print illustration in cork tan, deep pine green and rust red of a corkboard grid of pinned index cards linked by a dense crisscrossing web of red strings and pushpins, with one card slot left empty at the top where a boss card would sit.

Development

Multi-Agent Coding Teams Don't Need a Boss, a Study Finds

A 1,902-run study of Claude Code agent teams found naming a coordinator adds no measurable benefit, while shared-file versus messaging coordination swings token costs by up to 42%.

A bold vintage screen-print illustration in deep navy, warm cream and dusty rose of a large refreshable braille display angled diagonally, its row of raised and flat braille dot cells crossed by a keyboard band and a headphone-arc crescent.

Creative Tooling

AI Coding Tools Helped Blind Developers. Now Their Interfaces Are the Barrier

A 5 August 2026 study finds AI coding interfaces are a new accessibility barrier, and maintainer attention — not model quality — decides which tool a blind developer can use.

A bold vintage screen-print illustration in cream, deep charcoal-navy and alarm red of a circular smoke-detector alarm with concentric rings casting a triangular beam of red warning light down onto a laptop keyboard.

Development

Your Coding Agent Trips the Same Alarms as an Intruder

Sophos telemetry from June 2026 shows Claude Code, Cursor and OpenAI Codex tripping the same rules built to catch attackers, just as GitHub ships an auto-approve mode.