Tag
SWE-Bench
3 articles

Development
The SWE-Bench Leaderboard Can No Longer Tell Models Apart
Two September 2026 papers, days apart, find SWE-bench Verified can't statistically separate its top coding agents, and the harness deciding the score resets with every model swap.

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

Development
FrontierCode: The Benchmark That Asks Whether AI Code Is Ready to Merge
A new benchmark built with more than 20 open-source maintainers deflates the record-breaking numbers behind coding agents: even the best model clears only 13% of the hardest tasks.