How to Choose a WCAG Checker: What Each Tool Detects, What It Misses, and How to Build a Tool Stack
There is a procurement conversation that happens in almost every organisation approaching the European Accessibility Act for the first time. Someone asks which accessibility checker to buy. Someone else suggests running several, on the reasonable-sounding theory that more scanners means more coverage.
Both questions skip the one fact that determines the answer: most of the tools on the shortlist are the same tool. Choosing a stack means understanding what is genuinely different between them, what none of them can see, and which of those gaps a regulator will actually ask you about.
One engine under many logos
axe-core, Deque's open-source rules engine, has been downloaded over four billion times and powers a large share of the accessibility tooling on the market - including Lighthouse's accessibility audit and Pa11y, among many commercial products.
This is the single most important thing to know when evaluating tools. If you run axe DevTools, then Lighthouse, then a commercial scanner that wraps axe-core, you have not triangulated anything. You have run largely the same rule set three times and paid for it twice. Adding tools that share an engine increases your report volume, not your coverage.
The practical rule: diversify by engine, not by vendor. Ask every prospective supplier which engine their scanner uses. If the answer is axe-core - and it frequently is - treat their differentiation as being about workflow, reporting, crawling, integrations and support, not detection. Those are real things to pay for. They are just not coverage.
What the tools actually differ on
Where genuinely different engines exist, the meaningful axis is not "how many issues does it find" but the trade-off between recall and false positives. Published comparisons put the main free tools roughly here:
| Tool | Detection | False positives | Real strength |
|---|---|---|---|
| axe DevTools (Deque) | ~78% of issues in its scope | Zero-false-positive policy | Trustworthy output; the reference engine; strong dev workflow |
| WAVE (WebAIM) | ~71% of violations | ~8% | Visual in-page overlay showing exactly where each issue sits |
| Lighthouse (Google) | ~52% of violations | Low (axe-based) | Free, in every Chrome install, single-page only, narrower rule set |
| Pa11y | Inherits its runner | Inherits its runner | CLI and CI automation; 9.1 ships axe-core 4.11 with Puppeteer 24 |
Read the false-positive column as carefully as the detection column, because the two failure modes cost you differently.
axe DevTools' zero-false-positive policy means when it reports a violation, it is a violation. Anything it cannot be certain about is returned as incomplete - flagged for human review rather than asserted. That distinction is the most underused feature in accessibility tooling. Teams routinely clear the violations list, declare the page clean, and never open the incomplete queue, which is precisely where the judgement calls live.
WAVE is noisier by design, and that is a feature in the right context. Its visual overlay renders icons directly onto the page, so a content author or designer can see that this heading is out of order and that image has no alternative. For teaching and for manual review it is often more useful than a cleaner list. For a build gate it is unsuitable, because an 8% false-positive rate means blocking deploys on things that are fine.
Lighthouse is the trap. It is free, built into Chrome, produces a confident-looking score out of 100, and audits a materially smaller rule set than dedicated scanners on a single page with no crawl. A Lighthouse accessibility score of 100 is a common and entirely false basis for believing a site is compliant. It is a useful smoke test and nothing more.
Pa11y is infrastructure, not an engine. Its value is scriptability - CI pipelines, scheduled crawls, machine-readable output. Its detection is whatever runner you point it at.
The ceiling no tool crosses
This is the number to put in front of whoever believes a scanner subscription constitutes an EAA programme.
Automated tooling detects roughly 30-40% of WCAG issues; the remaining 60-70% require human testing. Even on the most generous framing - axe-core is documented as finding on average 57% of WCAG issues by volume - you are looking at half the problem at best. And by volume is the crucial qualifier: a scanner finds many instances of a few highly detectable criteria. Measured by how many success criteria it can evaluate at all, automated coverage is considerably lower. Running two engines rather than one has been measured as lifting detection from around 27% to 35% of known issue types - real improvement, nowhere near sufficient.
The ceiling is structural, not a temporary state of the art. Deque has said that target-size is likely the only WCAG 2.2 rule it will add to axe-core, because the remaining criteria generate too many false positives without human review. The vendor with the most capable engine and the strongest commercial incentive is telling you the automation stops here.
What no scanner will ever tell you:
- Whether your alt text is correct. A scanner confirms the attribute exists. "image1.jpg" passes.
- Whether your heading structure is meaningful, as opposed to merely present and sequential.
- Whether a keyboard user can complete your checkout. Focus order, keyboard traps and modal focus management are largely manual findings.
- Whether your error messages let someone actually recover.
- Whether the flow is usable at screen-reader pace - the WCAG "accessibility supported" condition, which requires real assistive-technology testing.
- Whether a complete process conforms end to end. WCAG's conformance rules mean one broken step invalidates an entire checkout, and no page-level scanner reasons about journeys.
This is also the honest explanation for a statistic that otherwise looks like collective indifference. The WebAIM Million 2026 report found detected WCAG failures on 95.9% of the top million home pages, up from 94.8%, averaging 56.1 errors per page - a reversal after six years of gradual improvement, driven substantially by pages growing 22.5% more complex in a single year. Those are automatically detectable errors, on pages that overwhelmingly have access to free scanners. Detection was never the bottleneck.
Where overlays fit: they do not
One category deserves explicit exclusion. Overlay widgets are sold in the same market and sometimes in the same sales conversation as genuine checkers, and they are a different product with a different risk profile. A JavaScript widget that modifies your page at runtime is not a testing tool, does not produce evidence, and has been rejected as a compliance defence in multiple jurisdictions. If a vendor's pitch moves from "we will show you your issues" to "we will fix them with one line of code," you have left the tooling market.
Assembling a stack
Four layers, each doing a job the others cannot.
1. A CI gate - axe-core, via Pa11y or a direct integration. Fail the build on axe-core violations. Do not fail on incompletes; route them to review. This is your regression floor: it stops re-introducing what you already fixed, which is most of what goes wrong in a codebase that ships weekly.
2. A developer tool at the desk - axe DevTools browser extension. Catching an issue in a branch costs a fraction of catching it in an audit. This is where the cost curve actually bends.
3. A crawl for breadth - any reputable site-wide scanner. Its job is inventory and trend, not adjudication: which templates are worst, whether the number is moving, which page types you have never looked at. Buy this layer on crawl quality, deduplication and reporting, since the engine is probably axe-core regardless.
4. Manual and assistive-technology testing on a defined sample - the layer that is not a tool. Every complete process end to end, plus one instance of each page template. Keyboard-only, then screen reader across a documented reader-and-browser support matrix. For EU-facing products, WCAG-EM gives you a structured sampling methodology that an auditor will recognise, which matters more than it might seem.
Budget the fourth layer first. It is the expensive one, it is the one that finds the issues that block people, and it is the only one that produces evidence a market surveillance authority will accept.
Evaluating a vendor: the questions that separate them
Ignore the detection-rate claim on the website - it is measured against the vendor's own corpus. Ask instead:
- Which engine do you use? If axe-core, what rules have you added, and what is your false-positive policy?
- Do you distinguish violations from items needing review? A tool that only reports certainties hides your judgement calls; a tool that reports everything as a violation cannot gate a build.
- Can it authenticate and crawl behind login? Checkout, account management and application flows are your highest-risk surfaces and are usually gated. A scanner that only sees marketing pages is auditing the wrong site.
- Does it test rendered state? Modals, expanded menus, validated forms, and post-interaction states are where component failures live. Static crawls miss them.
- What does it export? You need per-issue records mapped to success criteria, with URLs, selectors and timestamps - the raw material of a remediation log.
- Does it map to EN 301 549, not just WCAG? EN 301 549 contains requirements beyond the WCAG criteria, and its clause numbering is what EU conformance documentation is written against.
On that last point, note the live standards position: ETSI published EN 301 549 v4.1.1 on 2 September 2026, referencing WCAG 2.2 AA and adding an Annex ZB mapping to the EAA. It is not yet the legal reference standard - until the European Commission cites it in the Official Journal, v3.2.1 (2021) and WCAG 2.1 AA remain the operative reference. A tool that reports against WCAG 2.2 is useful and forward-looking; a vendor who tells you WCAG 2.2 is already legally mandatory under the EAA is either imprecise or selling urgency.
The report that survives scrutiny
Enforcement is live. Sweden's PTS has published a list of 200 e-commerce platforms to audit by Q3 2026, and German and Dutch authorities have confirmed enforcement activity is underway. What gets asked for in that situation is not a dashboard screenshot.
It is a documented evaluation: defined scope and sample, the methodology used, dated results, the tools and the manual methods applied, findings mapped to standard clauses, and a remediation plan with owners and dates for what remains. A scanner produces one input to that document. Buying a better scanner does not produce more of it.
If you are choosing your first tool: install the axe DevTools extension and Chrome's Lighthouse, both free, and spend the budget you were going to spend on a licence on a manual audit of your checkout instead. That is where the compliance risk is, and no scanner is going to find it.
Related reading
WCAG 2.5.8 Target Size (Minimum): The 24px Rule, Its Five Exceptions, and the Dragging Fix Most Teams Skip
WCAG 2.2 added a 24 by 24 CSS pixel minimum for pointer targets and a single-pointer alternative for dragging. What 2.5.8 and 2.5.7 require, how the exceptions work, CSS patterns that pass, and how to apply the thresholds on iOS and Android.
Chatbot and AI Assistant Accessibility Under the EAA: Streaming Replies, Focus Management and the Widget Nobody Tested
Chat widgets and AI assistants sit inside checkout, banking and support flows the EAA covers, yet streaming responses, focus handling and launcher buttons routinely fail screen reader and keyboard users. What WCAG AA requires, how EU AI Act transparency duties interact, and a build-and-test checklist.
WCAG 1.4.10 Reflow: How to Pass the 400% Zoom Test Without Breaking Your Layout
WCAG 1.4.10 Reflow asks one thing: at 320 CSS pixels wide (400% zoom on a 1280px screen), can people read and use your page without scrolling in two directions? What the criterion actually requires, the exceptions teams misread, the layout patterns that fail, the CSS that fixes them, and a test method you can run in five minutes.