Turn accessibility testing into a delivery system, not a final audit. Before work starts, identify what is changing and assign every applicable check a trigger, earliest useful stage, trained performer, accountable accepter, evidence record, blocking rule and retest owner. That prevents a familiar release-review failure: automated scan results are attached, yet nobody has completed the journey by keyboard, checked reflow, judged whether labels make sense or confirmed who will retest the fixes.
What to put into practice
Accessibility testing works best as distributed delivery evidence, not as a specialist audit added at the end.
Every test method needs a trigger, stage, performer, accountable accepter, evidence record, blocking rule and retest owner.
Automation, manual checks, assistive-technology testing and evaluation with disabled people provide different evidence.
Higher-risk changes need deeper testing, but lower risk never excuses a known barrier or an unsupported conformance claim.
A release exception records an authorised risk decision; it does not turn a failed result into conformance.
What turns accessibility testing into an operating programme?
A testing programme distributes distinct evidence across design, content, development, QA and release preparation while preserving accountable acceptance and specialist assurance. W3C recommends evaluating accessibility early and throughout development or redesign, when problems are easier to address. The useful shift is from asking whether the accessibility lead has checked the finished site to asking whether each team has produced the evidence expected at the point where it can still change the work.
Creators remain responsible for the accessibility of what they create. Designers make accessible interaction and visual decisions, authors and editors make content meaningful, developers implement and check behaviour, and QA plans and performs independent tests. Section508.gov offers a public-sector example of distributing comparable activities across delivery roles. Adapt that principle rather than importing its exact RACI assignments or treating United States guidance as a requirement for New Zealand businesses.
The accessibility lead should own policy, coaching, methods and difficult interpretations, not become the default operator for every check. One person may wear several hats in a small team, provided the record names which role they are performing. For higher-risk work, retain meaningful separation between creating the change, checking it independently and accepting any remaining release risk.
Automated detection finds repeatable conditions that can be identified programmatically, but no tool alone can determine whether a site meets accessibility standards.
Manual conformance checks examine behaviour and meaning, including keyboard operation, focus, structure, labels and instructions.
Assistive-technology checks examine compatibility while a trained tester completes representative tasks in selected environments.
Evaluation with disabled people investigates usability, unmet needs and barriers that standards checks may not reveal.
How should test depth change with the release?
Test depth should rise with interaction, reuse, novelty, journey criticality and likely user impact, while every change still produces appropriate evidence. Start by inventorying the affected journeys, components, templates, content types, documents, media, controls and supported technologies. The change class then selects the minimum test layers; it does not excuse a known barrier, guarantee an untested path or operate as an official risk standard.
This proportionate approach reflects established guidance without copying a fixed score. Section508.gov describes documented depth ranging from automated and spot checks through component and comprehensive testing, alongside continuous monitoring and usability evaluation. For broader conformance work, WCAG-EM calls for a defined scope, exploration of key views and functions, representative sampling where necessary, evaluation of that coverage and a report of findings.
Content-only change: complete human content review and applicable automated checks; add structural, keyboard, zoom or assistive-technology review when meaning, media, documents or controls change.
Visual or layout change: add design review and zoom or reflow checks, plus focus review wherever interactive behaviour is affected.
Component or interaction change: define acceptance criteria before build, require developer checks and independent QA, cover relevant states and tasks, and add regression checks when reused.
New template, critical journey or major release: use all applicable layers, trained assistive-technology testing, representative coverage, sampled conformance evaluation and disabled-user evaluation while findings can influence the work.
What belongs in the test-ownership matrix?
The matrix should give every test layer a change trigger and scope, earliest useful stage, responsible performer, accountable accepter, required expertise and environment, retained evidence, blocking rule, and remediation or retest owner. Complete it during planning, not at release review. That makes the expected evidence visible before estimates, acceptance criteria and research plans harden around an incomplete testing approach.
Role boundaries should follow the work. Designers own accessible design decisions; authors or editors own meaningful content; developers own implementation and local checks; QA owns the test plan and independent execution; user researchers own ethical studies with disabled people; and an authorised product or release owner makes the release decision. The accessibility lead provides policy, coaching and specialist judgement while trained independent evaluation adds assurance.
Public-sector matrices show that accessibility criteria, manual checks, automated regression, defect records, release readiness, evidence logs and post-release feedback can sit inside ordinary delivery artefacts. Private teams can use tickets, design reviews, content workflows and deployment records that already exist. The important feature is traceability: a result must lead to a named decision, remediation owner and retest path.
Accessibility stops being somebody else’s final check when every change arrives with named evidence, ownership and a retest path.
A worked ownership matrix to adapt to your team, technology and change classes
Test layer, trigger and scope
Earliest stage, performer, expertise and environment
Accountable accepter and retained evidence
Release effect and retest owner
Automated checks — every relevant code, template or content change; record the scanned scope and configured rules.
During authoring or development; author or developer locally, then QA or the delivery pipeline in a controlled environment.
QA accepts the run record, scope, configuration, result and linked defects; the creator owns fixes.
Configured blocking findings stop progression; the creator remediates and the same check is rerun.
During design and authoring; the author or editor applies human judgement in the rendered context.
The content owner accepts the reviewed item, rationale and any defect record; specialist advice is retained where needed.
Meaningful content failures block affected publishing; the author fixes them and an editor retests.
Keyboard review — new or changed controls, interactions, states and representative tasks.
From interactive prototype through QA; developer checks locally and QA independently operates the task by keyboard.
QA accepts task steps, focus observations, environment, outcomes and defects.
Inability to complete a required task or escape a component follows the defined blocking rule; developer fixes and QA retests.
Zoom and reflow — layout, typography, navigation, overlays, focus or responsive behaviour changes.
From design review through QA; designer and QA use the relevant viewport and resizing conditions.
QA accepts affected views, settings, screenshots where useful, observations and defect disposition.
Lost or obstructed information or functionality follows the blocking rule; designer or developer fixes and QA retests.
Selected assistive technology — higher-risk interactions, major changes and representative tasks.
Prototype, development and QA; a trained tester uses selected current browser, operating-system and assistive-technology combinations.
QA or the accessibility lead accepts task, impact, expected behaviour, full environment, evidence and result.
Blocking compatibility failures return to the creator; the trained tester repeats the same task after remediation.
Disabled-user evaluation — prototypes, unfamiliar patterns and critical journeys where usability evidence can alter decisions.
Before design hardens; an experienced user researcher conducts an ethical study with suitable participants and accessible materials.
The product owner accepts the research record, findings, limitations and planned response.
Findings inform redesign and prioritisation under programme rules; research and product owners confirm follow-up.
Sampled conformance evaluation — new templates, critical journeys, major releases or wider assurance needs.
Before release, with enough time to remediate; a trained evaluator defines scope and representative coverage.
The authorised owner accepts scope, method, sample, criterion-level findings, limitations and report.
Blocking findings must be resolved and independently retested, or handled through an explicit authorised exception process.
What should the core accessibility checks examine?
Each core check should complete a defined task or answer a specific conformance question, not merely confirm that a tool ran. Automated checks should record their scope and detect configured programmatic conditions. Content review needs human judgement because the presence of a title, heading, label, link, instruction, error message, caption, transcript or text alternative does not establish that it communicates useful meaning in context.
Keyboard review should operate every relevant control, follow the expected focus order, confirm visible and unobscured focus, enter and leave components, observe state changes, recover from errors and complete representative tasks. WCAG 2.2 requires keyboard operation subject to its path-dependent input exception, permits focus to move away from focused components, and at Level AA requires visible focus that author-created content does not entirely obscure.
Resize text to 200 per cent under the applicable WCAG criterion and look for lost content or functionality, while retaining its stated exceptions.
For reflow at the 320 CSS-pixel width equivalent, check for disallowed horizontal scrolling and any loss of information or functionality.
For reflow at the 256 CSS-pixel height equivalent, check the distinct condition for disallowed vertical scrolling.
Retain the exception for layouts that require two dimensions for meaning or use.
Look for obstructed content, hidden focus and unexpected changes outside the visible area rather than demanding pixel similarity.
When is specialist and disabled-user evidence needed?
Add specialist assistive-technology testing for representative tasks, meaningful state changes and higher-risk work, especially after significant features or major changes. A screen reader is one compatibility method, not a simulation of every blind person’s experience or proof of conformance. Select current browser, operating-system and assistive-technology combinations from audience evidence, product technology, support commitments and known risks rather than copying a universal matrix.
Make every finding reproducible. Record the affected task, user impact, expected behaviour, browser, operating system, assistive technology and version, supporting evidence, remediation owner and retest result. Spoken output, interaction steps and state changes may matter more than a screenshot. A trained tester should repeat the original task after remediation so the record shows whether the barrier was removed in the same environment.
Evaluation with disabled people answers a different question: whether people can understand and use the experience, including concerns a conformance review may miss. Plan it for prototypes and critical journeys while findings can influence decisions, and address significant obvious barriers without postponing useful early input. Do not generalise one participant’s experience to a disability population. Combine this work with standards evaluation, because neither substitutes for the other and even the highest WCAG conformance level cannot meet every individual need.
How should evidence control release and improvement?
Release should proceed only when the evidence required for the affected change class is complete, blocking findings are resolved and retested, and the retained record identifies scope, method, environment, result, owner, disposition and retest status. Use those records rather than one global accessibility score. The authorised product or release owner accepts the decision; QA and specialists supply independent evidence, while creators remain responsible for remediation.
If organisational policy permits an exception, record its authorised owner, rationale, affected users, mitigation, expiry and follow-up. Keep it visible and time-bound. An exception records a risk decision; it does not alter the underlying test result or establish conformance. Where the correct legal interpretation, certification position or jurisdiction-specific claim is uncertain, seek qualified legal advice rather than treating this operating model as a compliance determination.
After release, connect reported barriers and recurring defects back to the matrix, regression checks, training, templates and future change classes. Section508.gov’s agile example links test plans, defect records, keyboard and screen-reader review, reports, logs, release readiness and feedback, while W3C recommends evaluation throughout development. The practical lesson is continuous improvement: evidence should change the next decision, not disappear into an audit archive.
Pilot the matrix on one critical journey and make its evidence trail complete.
Train each named role owner in the checks and decisions they perform.
Add reusable evidence templates and appropriate automation to existing workflows.
Calibrate blocking rules against real findings without inventing a universal severity score.
Review recurring defect patterns and improve design, content, code and regression coverage.
Expand to further journeys once the handoffs, records and retest loop are working.
Accessibility testing programme FAQs
How do you create a web accessibility testing programme?
Define the affected journeys and change classes, distinguish the evidence types, and build a matrix naming the trigger, stage, performer, accountable accepter, evidence, blocking rule and retest owner for every method. Train the role owners, pilot one critical journey, then expand using recurring findings.
Who is responsible for accessibility testing?
Responsibility is distributed: designers, content teams and developers remain accountable for accessible work, while QA performs independent execution and researchers lead studies with disabled people. The accessibility lead stewards policy and complex methods, and an authorised product or release owner makes the release decision.
Can automated accessibility testing prove WCAG conformance?
No. Automation can repeatedly detect configured programmatic conditions, but no evaluation tool alone determines whether a site meets accessibility standards. Knowledgeable human evaluation and other methods applicable to the scope are still required.
When should teams test with screen readers and disabled users?
Use trained, task-based screen-reader or other assistive-technology testing for representative and higher-risk work throughout development, especially after major changes. Involve disabled people in prototypes and critical journeys while their findings can influence the design; use standards evaluation alongside that research.
What accessibility findings should block a release?
Each organisation must authorise its own blocking rules, but the gate should show that required checks are complete and blocking findings have been resolved and retested. Any permitted exception should remain explicit, time-bound and separate from a conformance claim.
We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.
Build a practical website governance model that names who decides, defines delegated limits, sets escalation triggers and keeps durable decision records.
Build a journey-linked register connecting third-party services to owners, information flows, measured costs, failure effects, fallbacks and review triggers.