Turn accessibility testing into a distributed delivery program, not a specialist inspection waiting at the release gate. Before work begins, identify what is changing and assign every applicable check a trigger, an earliest useful stage, a trained performer, an accountable accepter, a retained evidence record, a blocking rule and a retest owner. That structure prevents an automated report from becoming false reassurance when nobody has tested the actual journey by keyboard, at reflow, with assistive technology or with disabled people.
The operating rules
Accessibility testing is distributed delivery evidence, not a specialist audit added at the end.
Every method needs a trigger, stage, performer, accountable accepter, evidence record, blocking rule and retest owner.
Automation, manual review, assistive-technology testing and evaluation with disabled people answer different questions.
Higher-risk change demands deeper evidence, while lower risk never excuses a known barrier.
A release exception records an authorised risk decision; it does not establish conformance.
What turns accessibility testing into an ongoing program?
An accessibility testing program distributes distinct forms of evidence across design, content, development, QA and release while preserving accountable acceptance and specialist assurance. W3C advises evaluating early and throughout development or redesign, when problems are easier to address, and says no tool alone can determine whether a site meets accessibility standards. Public-sector role matrices from Section508.gov also demonstrate how accessibility work can sit across product, design, development, QA, content, specialist and oversight roles, although their exact assignments are not private-sector or Australian requirements.
Automated detection finds repeatable, programmatically detectable conditions within its configured scope.
Manual conformance checks examine behaviour and meaning that require informed human judgement.
Assistive-technology checks examine compatibility while a trained tester completes representative tasks.
Evaluation with disabled people investigates usability, barriers and needs that the other methods may not reveal.
Creators remain responsible for accessible decisions and implementation; the accessibility lead owns policy, coaching, methods and difficult interpretations.
How should test depth change with the release?
Test depth should increase with interaction, reuse, novelty, journey criticality and potential user impact, but every change still needs applicable evidence. Start by inventorying affected journeys, components, templates, content types, documents, media, controls and supported technologies. Section508.gov guidance illustrates depth ranging from automation and spot checks to component and comprehensive testing, with monitoring and disabled-user evaluation as separate activities. For a broader conformance evaluation, WCAG-EM defines scope, explores key views and functions, selects representative coverage where necessary, evaluates it and reports the findings.
Content-only change: human content review and applicable automation; add structural, keyboard, zoom or assistive-technology checks when meaning, media, documents or controls change.
Visual or layout change: design review, applicable automation, zoom and reflow checks, plus focus review wherever interaction is affected.
Component or interaction change: acceptance criteria before build, developer checks, independent QA, relevant states and tasks, and regression coverage where the component is reused.
New template, critical journey or major release: all applicable layers, trained assistive-technology testing, representative coverage, sampled conformance evaluation and disabled-user evaluation while findings can still affect the work.
What belongs in the test-ownership matrix?
The matrix should name the method, trigger and scope, earliest useful stage, responsible performer, accountable accepter, required expertise and environment, retained evidence, blocking rule, and remediation or retest owner. Section508.gov examples connect design review to UX, implementation to developers, accessible content to authors, test planning and independent execution to QA, and sign-off to named accountable roles. They also show accessibility criteria, regression, defects, evidence and feedback within ordinary delivery artefacts. Adapt that distribution principle rather than importing its federal RACI unchanged.
Accessibility stops being somebody else’s final check when every change carries named evidence, ownership and a retest path.
A worked accessibility test-ownership matrix
Test layer, trigger and scope
Earliest stage, performer, expertise and environment
Accountable accepter and retained evidence
Release effect and retest owner
Automated checks: every relevant code, template or content change; record configured scope.
During authoring or development; author or developer, then pipeline; tool configuration and supported environment required.
QA accepts the run record, ruleset, scope, results and reviewed false positives.
Configured blocking failures stop progression; creator fixes and reruns, with QA confirming.
What should the core accessibility checks examine?
Core checks should test representative tasks and meaningful outcomes, not merely confirm that a tool ran or a control received focus. WCAG 2.2 requires keyboard operation subject to its path-dependent input exception, the ability to move focus away from components, visible focus and protection from focus being entirely obscured by author-created content. It also requires text, with stated exceptions, to resize to 200 per cent without losing content or functionality. Human review remains necessary because titles, headings, labels, link purpose and input instructions must communicate useful meaning.
For keyboard review, operate every relevant control, follow the expected focus order, verify visible and unobscured focus, enter and leave components, observe state changes, recover from errors and complete the task.
For reflow, check the 320 CSS-pixel width equivalent without horizontal scrolling and the 256 CSS-pixel height equivalent without vertical scrolling, subject to the exception for layouts requiring two dimensions for meaning or use.
During zoom and reflow review, look for lost, hidden or obstructed information and functionality, off-viewport changes, hidden focus and disallowed scrolling rather than demanding pixel-for-pixel similarity.
For content review, judge whether headings, labels, links, instructions, errors, alternatives, captions and transcripts convey the intended meaning in their actual page and journey context.
When should specialist and disabled-user evidence be added?
Add trained assistive-technology testing for representative tasks, relevant states and higher-risk changes, and add disabled-user evaluation when prototypes or critical journeys can still be changed. GOV.UK guidance recommends task-based assistive-technology testing throughout development, especially after significant features or major changes. Choose current browser and technology combinations from audience evidence, product technology, support commitments and known risks. A screen reader is one compatibility method, not a simulation of every blind person’s experience or proof of conformance.
Record the affected task, user impact, expected behaviour, browser, operating system, assistive technology and version, supporting evidence, remediation owner and retest result.
Address significant obvious barriers before planned sessions so disabled participants can also reveal deeper usability concerns, without postponing useful early prototype input.
Do not generalise one participant’s experience to a disability population; an individual finding can still expose a serious barrier worth investigating.
Combine disabled-user evaluation with standards-based conformance evaluation: W3C says the former can reveal issues the latter misses, but cannot alone determine whether a website is accessible.
Remember that even the highest WCAG conformance level does not make content accessible to every individual with every type, degree or combination of disability.
How should evidence control release and improvement?
Release should depend on the evidence required for the affected change class, not one global score. Applicable checks must be complete, blocking findings resolved and retested, and retained records must identify scope, method, environment, result, owner, disposition and retest status. Section508.gov guidance illustrates a repeatable methodology spanning depth, timing, roles, qualifications, environments, tracking and pre-deployment requirements, while its agile matrix connects plans, defects, keyboard and screen-reader review, reports, logs, release readiness and post-release feedback.
Pilot the matrix on one critical journey and make its evidence trail complete.
Train the named role owners and provide concise evidence templates.
Add appropriate automation without treating its result as a conformance decision.
Calibrate blocking rules through the organisation’s authorised decision process.
Review recurring defects, reported barriers and retest failures, then improve training, templates and regression coverage.
Expand to further journeys and change classes once ownership and evidence are working.
The authorised product or release owner makes the acceptance decision; QA and accessibility specialists provide independent evidence, while creators remain responsible for remediation. Where organisational policy permits an exception, record its authorised owner, rationale, affected users, mitigation, expiry and follow-up. The exception does not change the failed result or establish conformance. Bring in a trained evaluator for complex interactions, assistive-technology behaviour, representative conformance coverage or disputed findings, an experienced researcher for studies with disabled people, and qualified legal counsel for jurisdiction-specific interpretations or claims.
Accessibility testing program FAQs
How do you create a web accessibility testing program?
Define the journeys and change classes in scope, distinguish the evidence types, and build a matrix covering triggers, stages, performers, accountable accepters, evidence, blocking rules and retesting. Train the named owners, pilot one critical journey, then expand using recurring findings and reported barriers.
Who is responsible for accessibility testing?
Responsibility is distributed: designers own accessible design decisions, content teams own meaningful content, developers own implementation and local checks, QA owns the test plan and independent execution, and researchers own ethical studies. Accessibility specialists steward policy and difficult methods, while an authorised product or release owner accepts the release decision.
Can automated accessibility testing prove WCAG conformance?
No. Automation can efficiently detect repeatable programmatic conditions within its configured scope, but a clean scan cannot determine whether a site meets accessibility standards. Knowledgeable human evaluation and other applicable methods are still required.
When should a team test with screen readers and disabled users?
Use trained, task-based screen-reader or selected assistive-technology testing for representative and higher-risk changes throughout development. Involve disabled people with prototypes and critical journeys while findings can influence decisions. The methods are complementary: compatibility testing is not disabled-user research, and neither alone proves conformance.
What accessibility findings should block a release?
Each organisation must define authorised blocking rules appropriate to its work, obligations and affected users. Release evidence should show that every required check is complete and blocking findings are resolved and retested. Any permitted exception should remain explicit, authorised and time-bound, and must not be presented as conformance.
References & Sources
This article was researched using the following sources:
We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.
Build an eight-domain website governance model that gives each decision an owner, clear delegation, required input, an escalation route and a durable record.
Build a journey-linked register connecting third-party website services to owners, information flows, measured costs, failure effects and review triggers.