Matt Henry, the former engineering lead of the U.S. Web Design System, told FedScoop that only “about 30%” of accessibility testing can be reliably automated, because “most accessibility testing is evaluating how a real human user … would interact with the site.” Accessibility officer Sheri Byrne-Haber said the project “went from something I would recommend to something I wouldn’t touch with a 10-foot pole.” The U.S. Web Design System (USWDS) is the component library that federal, state and local sites copy for their buttons, modals and forms. Its new leadership wants to “automate as much accessibility testing as possible.” Experts say the hard part of that testing needs people.
The automatable part was already automated
The dispute is usually told as AI versus no AI. USWDS’s own records tell a messier story. The team’s 2026 shipping notes describe an “AI readiness audit” and a “code review skill”, written before the September leadership change, so the move toward AI predates the fight.
More telling is what the project already did. USWDS’s accessibility documentation lists three practices side by side: “Scan with automated tools, like pa11y and aXe, on every code change”; “Manually test components and functionality with experienced accessibility specialists”; and “Test with screen readers (VoiceOver, JAWS)”.
So the scannable share was covered years ago. A promise to automate more can only reach into the other two lines, the ones that involve people. Read that way, the question is less whether AI enters the pipeline than which layer it replaces.
Where the verification work goes
Take the person who feels this. You build a benefits form for a county office out of USWDS components. You assume the accordion, the modal and the date picker were already tried with a screen reader, because for years the documentation said they were.
Suppose that checking quietly becomes an AI scan. The gaps do not disappear. Finding them becomes your job, and nobody tells you. The same page warns that “Building with accessible USWDS components does not guarantee an accessible service”, but most teams read that as a caveat, not as a handover of the whole testing burden.
The sources here document intentions and a dispute. FedScoop reports that GitHub posts show the team using OpenAI’s Codex and CodeRabbit for code review . A coalition letter, reported by FedScoop , demands “rigorous manual accessibility testing and documented quality assurance” and warns that one defect “can put thousands of systems at risk.” None of it measures what agencies will do. What follows is Pipeline’s reading: upstream verification is a shared cost, and cutting it turns that cost into thousands of separate ones.
A design system sells trust in advance, and manual testing is how the trust gets paid for. If the payment stops, the bill still arrives.
The strongest case for automation
The automation push did not come from nowhere. The same shipping notes say the team was “still shorthanded and therefore depending heavily on community contributions, which creates various bottlenecks”. A maintainer who cannot find specialists has real reasons to reach for a tool.
The tools are also better than the 30% figure suggests. A September 2026 paper on agentic accessibility auditing found that AI auditors recalled 86% of real violations, against 36% for axe-core. We covered those measured trade-offs earlier this month. The catch is 56% precision: nearly half of what the agents flag is wrong, so a human has to triage it. That adds human work.
Two more limits apply. That paper tested scholarly-platform pages, not USWDS components, so its figures do not transfer cleanly. And no source shows manual testing has actually stopped. Parker’s priorities post says 3.14.0 “remains the current stable release” and agencies can keep using it. No version has yet shipped under the new process.
What the shortcut costs
A CDO Magazine analysis frames this as governance. Automating work that is only partly automatable needs documented criteria for what the tools reliably catch, and a way for affected communities to challenge the results. It found both absent from the transition.
That is the sharper reading of the fight. The issue is not whether a model can read markup. It is whether anyone can still say what “tested” means when thousands of sites inherit the word. A design system sells trust in advance, and manual testing is how the trust gets paid for. If the payment stops, the bill still arrives, addressed to whoever ships the form.



