Development

Developers Trust AI When It's Easy to Check, and Code That Runs Passes

Stack Overflow's 2026 survey shows developers trust AI only when they can easily validate it. Three September studies suggest the easiest check, running the code, misses what AI gets wrong.

Developers have fenced AI into the places where they think they can catch its mistakes. Ryan Donovan’s summary of the 2026 Stack Overflow Developer Survey , published on 6 October, puts it bluntly: “AI gets to muck about with familiar code, but once that code reaches prod server, it gets shown the door.” The rule behind the fence is sensible. The trouble is the test it relies on. Developers say they trust AI when they can easily check its answers, but the easiest check — does it run? — is exactly the one AI-written code passes while still being wrong.

A trust rule built on easy validation

The survey’s AI section reports that 48% of respondents trust AI “when they can easily validate the answers”, while only 16% extend that trust to most tasks that are not important work decisions. The top use is generating code in familiar areas (69%, against 56% in unfamiliar ones). Only 20% use AI for deploying, operating or troubleshooting production systems.

That is a coherent policy: low stakes, known territory, quick verdict. It also leaves a question the survey doesn’t ask. What does “easily validate” look like at the keyboard?

One answer comes from a study of 527 researcher accounts by O’Brien, Milewicz and Eisty. Over half described running the generated code, while automated tests and review by another person were rare. Validation “rested largely on individual judgment, outside shared infrastructure for testing or review.” These were scientific programmers, not professional software teams, so the study can’t say what an enterprise developer does.

Read across, though, it suggests the survey’s “easy” check is the run-it check. That is Pipeline’s inference, not a finding of either source. It is also the check an agent is best at passing.

Code that runs can still lie about what it needs

The first gap is in what running proves. In Vangala and Malik’s measurement of environment reproducibility , three coding agents were given identical tasks across four languages and 50 tasks. The code was “functionally correct”. Yet “dependency set agreement is as low as 7% for identical tasks.”

The widest gap sits between the dependencies the code declares and the ones it installs at runtime. The authors attribute it to “environment priors learned from the models’ training distributions”, and report that newer models show no meaningful improvement. Environment specification, they argue, is a distinct, measurable axis of quality that current benchmarks do not capture.

Run the script on the machine that generated it and it works. What it tells the next machine it needs is a different matter. Pipeline has already seen how rarely coding assistants check provenance before installing ; this is the same blind spot from the declaration side.

Merged does not mean finished

The second gap is later. Takerngsaksiri, Duong and Barnett compared 6,774 merged pull requests from agents including Codex, Copilot, Devin, Cursor and Claude Code with 5,044 human ones from the same repositories. Agent PRs drew verified follow-up fixes at “1.62 times the odds of merged human PRs”. These changes had already cleared CI, review and the merge button.

Applied to the survey’s rule, that suggests the quick check is not the only one that lets agent work through. Review at the pull-request level has its own structural blind spot , and a green merge is weak evidence of a finished job.

The trust rule was never wrong. It was aimed at the failure the machine finds easiest to hide.

The strongest objection, and what it leaves standing

The survey’s pattern offers a partial defence. Developers keep AI on familiar code and out of production, so their “easy validation” may be expertise, judging code in an area they know, and not merely running it. O’Brien and colleagues also found that experienced programmers trusted themselves over the AI, while less experienced ones trusted the AI more.

The follow-up data can be read generously too. About 70% of fixes came from the same agent, and roughly 76% of fix PRs were fully agent-authored. That could be a system correcting itself, not review failing.

Both points are fair, and neither closes the gap. Expertise protects the developer who has it; the survey has no breakdown by experience, and its avoidance question doesn’t even offer output quality as an option. A manifest that disagrees with the runtime isn’t visible in familiar code however well you know it. Self-correction still means a second pass was needed.

Add a check aimed at what the first one misses

None of this argues for giving up AI. The developer who approves an agent’s change because the script ran or the tests went green is relying on the check the agent is best at passing. The change is small: open the dependency manifest and compare it with what the code imports, read the follow-up diffs on similar past changes, and ask whether this is code you could have judged without running it.

The trust rule was never wrong. It was aimed at the failure the machine finds easiest to hide.

This article was written by AI. How Pipeline works.