Turn accessibility testing into a delivery programme, not a specialist inspection bolted onto release week. Before work begins, identify the affected journeys and change class, then assign each applicable check a trigger, earliest useful stage, trained performer, accountable accepter, evidence record, blocking rule and retest owner. That structure prevents an automated scan from becoming false reassurance while keeping the accessibility lead available for policy, coaching and difficult judgements instead of making one person the queue for every decision.
Key decisions
Accessibility testing works best as distributed delivery evidence, not as a specialist audit added at the end.
Every testing method needs a trigger, stage, responsible performer, accountable accepter, evidence record, blocking rule and retest owner.
Automation, manual review, assistive-technology testing and evaluation with disabled people answer different questions.
Higher risk should increase test depth, never excuse a known barrier or support a claim about an untested path.
A release exception records an authorised risk decision; it does not make a failed result conformant.
What turns accessibility testing into a programme rather than a final audit?
An accessibility testing programme distributes distinct evidence throughout design, content production, development, QA and release preparation while preserving accountable acceptance. W3C advises evaluating early and throughout development or redesign, when problems are easier to address. Its guidance also says no evaluation tool alone can determine whether a website meets accessibility standards. Repeatable automation is valuable, but knowledgeable human evaluation remains necessary for behaviour, context and meaning.
Responsibility should remain close to the work. Designers own accessible design decisions, authors and editors own meaningful content, developers own implementation and local checks, and QA owns the test plan and independent execution. The accessibility lead stewards policy, methods, training and complex interpretations. Section508.gov offers a public-sector example of this distribution across a delivery lifecycle; use the principle, not its exact US federal assignments, as the model for an Irish organisation.
Automated detection finds repeatable conditions that can be checked programmatically within a stated scope.
Manual conformance checks examine operation, structure, context and meaning that require informed human judgement.
Assistive-technology checks examine compatibility while a trained tester completes representative tasks and relevant states.
Evaluation with disabled people investigates usability and unmet needs that standards-based checks may not reveal.
How should test depth change with the work being released?
Test depth should rise with interaction, reuse, novelty, journey importance and potential user impact, without allowing lower-risk work to skip all evidence. Start by inventorying the affected journeys, components, templates, content types, documents, media, controls and supported technologies. This ladder is an adaptable operating model rather than an official risk standard, fixed score or basis for claiming that an untested route conforms.
For broader conformance work, WCAG-EM provides a different but useful discipline: define the scope and goal, explore key views and functionality, select representative coverage where full evaluation is infeasible, evaluate it and report the findings. Sampling manages a defined evaluation scope; it does not prove anything about excluded paths. Record the reasoning behind both the change class and the selected coverage.
Content-only change: require human content review and applicable automated checks; add structural, keyboard, zoom or assistive-technology review when meaning, media, documents or controls change.
Visual or layout change: add design review and zoom or reflow checks, plus focus review wherever the change affects interaction.
Component or interaction change: define acceptance criteria before build, require developer checks and independent QA, exercise relevant states and add regression coverage for reuse.
New template, critical journey or major release: use all applicable layers, trained assistive-technology testing, representative coverage, sampled conformance evaluation and timely evaluation with disabled people.
What belongs in the test-ownership matrix, and who owns each handoff?
The test-ownership matrix should make every required handoff explicit before delivery starts. Give each test layer a trigger and scope, earliest useful stage, responsible performer, accountable accepter, required expertise and environment, retained evidence, blocking rule, and remediation or retest owner. Keep the performer separate from the person authorised to accept release risk, particularly where the change affects a reused component or critical journey.
One person may wear several hats in a small team, but the record should name the hat being worn. Preserve independent review for higher-risk work and leave creators responsible for fixing what they create. Public-sector RACI and agile examples show how acceptance criteria, local checks, regression, defect records, release evidence and post-release feedback can sit within ordinary delivery artefacts, although their precise ceremonies are not universal requirements.
Accessibility stops being somebody else's final check when every change arrives with named evidence, ownership and a retest path.
A practical seven-layer test-ownership matrix
Test layer, trigger and scope
Earliest stage, performer, expertise and environment
Accountable accepter and retained evidence
Release effect and retest owner
Automated checks: every relevant code, template or content change; record the scanned scope.
During authoring or development; author or developer using an approved, maintained workflow.
QA accepts a scoped report with tool version, result and exclusions.
Defined blocking findings return to the creator; automation runs again after remediation.
Content review: changed titles, headings, labels, links, instructions, errors, alternatives or media text.
During drafting and design; trained author or editor reviewing meaning in context.
Content owner accepts the reviewed copy and records decisions or unresolved dependencies.
Misleading or missing essential meaning blocks publication; the content team retests.
Keyboard review: new or changed controls, interactions, navigation, forms and task states.
During component development, then independent QA; use a physical or equivalent keyboard interface.
QA accepts task notes covering operation, order, focus, state changes and recovery.
A blocking task failure returns to development; QA repeats the complete affected task.
Zoom and reflow: layout, typography, viewport, overlay or responsive-behaviour changes.
At design review and working-build QA; tester understands the applicable WCAG conditions.
QA accepts evidence from affected views, states and orientations where relevant.
Lost or obstructed content or functionality blocks under policy; development retests after correction.
Selected assistive technology: representative tasks, complex components and higher-risk changes.
From working prototypes onwards; trained tester using a justified browser, operating system and technology combination.
QA or accessibility lead accepts reproducible task, environment, impact and output evidence.
Applicable blocking failures return to development; a trained tester verifies the fix.
Disabled-user evaluation: prototypes, critical journeys and questions about usability or unmet needs.
While findings can influence the work; experienced researcher with accessible, ethical study arrangements.
Product owner accepts the research record, limitations and resulting decisions.
Findings inform remediation and priorities; named owners verify changes without overgeneralising participants.
Sampled conformance evaluation: new templates, major releases or a defined assurance need.
After representative functionality is stable; trained, suitably independent evaluator using a declared scope.
Authorised owner accepts the scope, sample, methods, findings and limitations.
Blocking findings require remediation and evaluator retest before the defined decision.
What should each core accessibility check examine?
Each core check should exercise meaningful content, controls, states and representative tasks rather than merely record that a tool ran. Automated checks belong in relevant local workflows and integration pipelines, with their scope and exclusions retained. Human content review asks whether titles, headings, labels, link names, instructions, errors, captions, transcripts and text alternatives communicate useful meaning in context; their presence alone does not settle that question.
Keyboard review should operate every relevant control, follow the expected focus order, confirm visible and unobscured focus, enter and leave components, observe state changes and recover from errors. WCAG 2.2 requires keyboard operability subject to its path-dependent input exception and requires users to move focus away from focused components. At Level AA, focus must be visible and must not be entirely obscured by author-created content.
Resize text to 200 per cent under the applicable WCAG criterion, preserving its stated exceptions.
For reflow, test the 320 CSS-pixel width equivalent without horizontal scrolling where required.
Separately test the 256 CSS-pixel height equivalent without vertical scrolling where required.
Preserve the exception for layouts that require two dimensions for meaning or use.
Look for lost information, blocked functionality, hidden focus and unexpected off-viewport changes rather than pixel similarity.
When should teams add assistive-technology, disabled-user and conformance evaluation?
Teams should add trained assistive-technology testing for representative tasks, relevant states and higher-risk changes, and add disabled-user evaluation while findings can still influence design. GOV.UK guidance recommends task-based assistive-technology testing throughout development, particularly after significant features or major changes. A screen reader is one compatibility method, not a simulation of every blind person's experience and not proof of overall accessibility or WCAG conformance.
Choose current browser, operating-system and assistive-technology combinations from audience evidence, product technology, support commitments and known risks instead of copying a universal matrix. Record the task, user impact, expected behaviour, environment, technology and version, supporting evidence, remediation owner and retest result. This makes a finding reproducible and keeps an old or irrelevant test combination from becoming policy by habit.
Address significant obvious barriers before a research session, without delaying useful early feedback on prototypes.
Use an experienced researcher to plan accessible and ethical evaluation with disabled people.
Do not generalise one participant's experience to a disability population, though an individual finding may expose a serious barrier.
Combine disabled-user evaluation with standards-based conformance evaluation because the methods answer different questions.
Evaluation with disabled people may range from informal prototype feedback to formal task-based research suited to the project stage. W3C says it can reveal usability issues that conformance evaluation alone may miss, but cannot by itself determine whether a website is accessible. Conversely, even the highest WCAG conformance level does not make content accessible to every individual with every type, degree or combination of disability.
How should evidence control release and improve the programme over time?
Release should depend on the evidence required for the affected change class, not one global score. The authorised product or release owner decides whether to accept release, using independent evidence from QA and accessibility specialists while creators remain responsible for remediation. Applicable checks must be complete, blocking findings resolved and retested, and each retained result must identify its scope, method, environment, finding, owner, disposition and retest status.
If organisational policy permits an exception, record its authorised owner, rationale, affected users, mitigation, expiry and follow-up. The exception does not alter the result or establish conformance. After release, feed reported barriers and recurring defects back into the matrix, regression coverage, training, templates and future change classes. Bring in trained evaluators for complex or disputed findings, experienced researchers for disabled-user studies, and qualified legal counsel for jurisdiction-specific compliance interpretations.
Pilot the matrix on one critical journey.
Train each named role owner.
Add concise evidence and defect templates.
Introduce appropriate automation and regression checks.
Calibrate blocking rules from real findings.
Review recurring defects, then expand coverage.
Frequently asked questions
How do you create a web accessibility testing programme?
Define the journeys and change classes in scope, distinguish the evidence types, and build a matrix covering triggers, stages, owners, environments, records, blocking rules and retesting. Train the named performers, pilot one critical journey and expand using recurring findings.
Who is responsible for accessibility testing?
Responsibility is distributed: designers, content teams and developers remain accountable for accessible work, while QA independently executes the plan and researchers lead disabled-user studies. The accessibility lead stewards policy and methods, and an authorised owner makes the release decision.
Can automated accessibility testing prove WCAG conformance?
No. Automation can detect repeatable programmatic conditions within its scanned scope, but no evaluation tool alone determines whether a website meets accessibility standards. Knowledgeable human evaluation and other applicable methods are still required.
When should a team test with screen readers and disabled users?
Use trained, task-based assistive-technology checks for representative states and higher-risk changes throughout development. Involve disabled people in prototypes and critical journeys while their findings can influence decisions, and combine that usability evidence with standards-based evaluation.
What accessibility findings should block a release?
Each organisation must authorise its own blocking rules, but required checks should be complete and blocking findings resolved and retested before release. Any permitted exception should remain explicit, time-bound and separate from a claim of conformance.
References and Sources
This article was researched using the following sources:
We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.