Deque analyzed more than 13,000 pages and nearly 300,000 accessibility issues to answer a question every team asks and few can answer precisely: how much can automated evaluation catch?
The answer was 57.38 percent of issues, considerably better than the 20 to 30 percent figure that circulated for years. Then comes the part that matters more. Those detected issues map to only 16 of the 50 WCAG success criteria.
Automation is genuinely strong across roughly a third of the standard and structurally blind to the rest. It reliably finds missing alt text and contrast failures. It cannot tell you whether the alt text describes the right thing, whether focus order makes sense to someone navigating by keyboard, or whether a customer using a screen reader can actually complete a purchase.
That ceiling is measured, and no model improvement moves it, because the remaining criteria require knowing what the interface is for.
Which frames the real question about AI-assisted evaluation. Not whether it can produce a heuristic review, because it can produce a plausible-looking one in ninety seconds. The question is what kind of evaluator it is, and where that places it in a process.
Coverage is the real gain, and it is a large one
Human expert review does not scale to surface area. A team evaluating a large property picks representative journeys and high-traffic screens, because that is what the budget allows. Everything else goes unlooked-at, sometimes for years. Legacy sections, low-traffic account flows, localized variants, error and empty states, the fourth step of a form nobody demos.
That unexamined surface is where AI earns its keep. It does not get bored on page 400. It applies the same criteria to a rarely-visited returns flow as to the homepage. For organizations running multiple products, extensive localization, distributed content ownership, or frequent releases, this is a genuine capability change rather than a marginal speedup.
There is also a methodological reason to welcome an additional evaluator, and it predates AI by thirty years. Heuristic evaluation has always been a plural method, because any single evaluator misses most of what is there. Jakob Nielsen's original measurement across six projects put a single evaluator at roughly 35 percent of a product's usability problems, which is why the standard recommendation has always been three to five people rather than one expert with good instincts.
The reason is blind spots, and specifically that different evaluators have different ones. That logic does not change when one of the evaluators is a model. It does mean the model's blind spots need to be understood as carefully as a human's, and they are less familiar.
The failure mode nobody flags in the sales demo
An AI evaluator will hand you eighty findings. Most will be technically accurate. A meaningful number will be irrelevant, and a few will be actively wrong in a way that costs you credibility if you forward them unfiltered.
Three patterns recur.
It misjudges importance. AI can correctly identify that a step in a flow is unusual and have no idea that the step exists because of a regulatory requirement, a contractual obligation, or a deliberate friction decision that took nine months to negotiate. It will recommend simplification with total confidence. An expert evaluator asks why the requirement exists and then decides whether the right fix lives in the interface, the process, the platform, or the policy.
It flattens context. A workflow that looks clean on screen may be genuinely difficult for someone under time pressure, switching between systems, using unfamiliar vocabulary, or performing this task once a year. Frequency of occurrence is not the same as severity of impact, and AI defaults to frequency because that is what it can see.
It cannot detect the problems that are about feeling. Customers routinely understand exactly how to complete a task and still abandon it, because the language felt legalistic, the progress was unclear, or the moment asked for trust the interface had not earned. These are among the most commercially expensive experience problems in existence and they are close to invisible to pattern matching.
The same weakness shows up when these tools generate rather than evaluate. Independent testing against real design scenarios has documented outputs with no clear visual hierarchy, related information separated by excessive spacing, patterns imported from the wrong context, and recurring contrast problems. The instinct that misapplies a pattern is the instinct that misjudges one.
Why this is worth doing anyway, right now
Because the alternative is that nobody looks at all, and the evidence says the situation is deteriorating.
The 2026 WebAIM Million found that 95.9 percent of the top million home pages had detectable WCAG 2 failures, up from 94.8 percent in 2025, with average errors per page rising 10.1 percent to 56.1. That reversed six consecutive years of gradual improvement. WebAIM points to increased reliance on third-party frameworks and "automated or AI-assisted coding practices," alongside page complexity growing 22.5 percent in a single year.
Meanwhile the regulatory floor moved. European Accessibility Act enforcement began June 28, 2025. The harmonized standard, EN 301 549 v3.2.1, incorporates WCAG 2.1 Level AA in full, and a version incorporating WCAG 2.2 is expected this year. Penalties are set nationally and are not symbolic: Sweden caps around €900,000, Spain at €600,000, and France allows up to 5 percent of annual turnover for serious breaches. Any organization selling to EU consumers is in scope regardless of where it is headquartered.
More output, generated faster, against a weaker set of standards, into a tightening enforcement environment. That is the actual risk position of most large digital estates in 2026.
How to run this so the findings survive contact with your team
Treat AI as one evaluator among three to five, not as the panel.
Give it real criteria. A generic prompt returns generic findings. Feed it your design system standards, your accessibility target, your brand voice guidance, your known technical constraints, and the audience definition for the journey in question.
Have a human triage before anyone else sees the list. This is the step teams skip and the step that determines whether the output builds or destroys trust. One experienced practitioner sorting eighty raw findings into confirmed, context-dependent, and wrong takes a couple of hours and turns a liability into a work plan.
Prioritize on impact rather than count. Consider frequency, whether the issue blocks task completion, legal and accessibility exposure, effect on business outcomes, and remediation cost. A single finding that blocks screen reader users from checkout outranks forty low-contrast footer links.
Then validate the ones that matter with actual customers. Heuristic analysis produces hypotheses about what is wrong. It has never produced proof, and adding AI to it does not change that.
The honest positioning
An AI-assisted heuristic and accessibility sweep is a good first AI engagement precisely because it is small. The scope is definable, the current state is measurable, the output is a prioritized remediation list with owners and dates, and the return is defensible to a CFO who has stopped being impressed by demos.
It also teaches your organization something true and specific: where AI is reliable, where it needs a human check, and what your review process actually looks like under load. That knowledge is worth more than the findings, and it is the thing you cannot get from a pilot that stays theoretical.
Start there. The operating model conversation goes much better afterward.
![]() | Schedule an AI-Enabled Experience Design Working Session |
Sources
- Jakob Nielsen, The Theory Behind Heuristic Evaluations, Nielsen Norman Group.
- Jakob Nielsen, 10 Usability Heuristics for User Interface Design, Nielsen Norman Group.
- Deque, The Automated Accessibility Coverage Report.
- WebAIM, The WebAIM Million, 2026.
- Level Access, European Accessibility Act compliance overview.


