Two hosts trace two September 2026 papers that land two days apart: an audit of 254 SWE-bench Verified submissions that finds exact paired statistical tests can no longer separate the top thirty leaderboard entries, and a 176-configuration study across four models showing that planning, tool access and context strategy each help or hurt depending on the model underneath, not on any fixed best practice. Together they explain why a harness built out of scar tissue for one model quietly goes stale the week a team swaps in a better one.
They also hold onto the honest limits of that story: the same audit’s larger Test split still separates most of its pairs, the harness paper’s own findings suggest the effect may just shrink to a cost question as models improve, and a rival framework treats the harness as model-independent scaffolding on purpose. The asymmetry that’s left standing is the one that matters — the leaderboard has lost its resolving power exactly where teams need it, and nobody has written down when a harness needs recalibrating.