Run the web as a business system.

Search strategy, design or web operations...
Toggle menu

Web Accessibility

Build a Role-Based Web Accessibility Testing Programme

Build a role-based accessibility testing programme that scales evidence to change risk, clarifies ownership and makes release decisions traceable.

Five colleagues gather around a wooden table as a standing man places a blank card on a wall grid beside accessibility testing equipment.

A workable accessibility testing programme assigns evidence before work starts, not after a release candidate reaches QA. For every affected journey, component, template or content type, decide which checks apply, when they first become useful, who performs them, who accepts the result and who retests a fix. This prevents an automated scan from becoming false reassurance and stops the accessibility lead from becoming the person everyone waits for. The release owner receives a traceable decision record; creators retain responsibility for the quality of what they design, write and build.

Operating principles

  • Accessibility testing works as distributed delivery evidence, not as a specialist audit added at the end.
  • Every method needs a trigger, stage, performer, accountable accepter, evidence record, blocking rule and retest owner.
  • Automation, manual checks, assistive-technology testing and evaluation with disabled people answer different questions.
  • Higher-risk work needs deeper evidence, but lower risk never converts an untested path into a conformance claim.
  • A release exception records an authorised risk decision; it does not make a failed result conformant.

What turns accessibility testing into an ongoing programme?

Four colleagues sort blank dark-blue and amber cards into shallow trays at a wooden table with a keyboard, headset and folders.

Accessibility testing becomes a programme when distinct evidence is produced throughout design, development, content production, QA and release preparation, with one named role accountable for accepting the evidence. W3C advises early and continuing evaluation because problems are easier to address while work can still change. A current Section508.gov RACI offers a public-sector example of distributing activities across product, design, development, content, QA, specialist and oversight roles; a business can adapt that principle without importing the source's federal duties or exact assignments.

  • Automated detection finds repeatable conditions that tools can identify programmatically.
  • Manual conformance checks examine operation, structure, meaning and behaviour against applicable requirements.
  • Assistive-technology checks examine compatibility while trained testers complete representative tasks.
  • Evaluation with disabled people investigates usability, expectations and needs that standards checks may not expose.

These forms of evidence complement one another. A clean automated result is useful, but W3C states that no evaluation tool alone can determine whether a site meets accessibility standards; knowledgeable human evaluation remains necessary. Designers should resolve accessible interaction decisions, editors should judge whether content communicates meaning, developers should verify implementation locally, and QA should conduct independent checks. The accessibility lead owns policy, coaching, methods and difficult interpretations, stepping into execution where specialist skill is genuinely required rather than performing every routine check.

How should testing depth reflect the change being released?

Two adults hold four increasingly large stacks of blank cards beside a keyboard, headphones, magnifying glass and Braille display.

Testing depth should grow with interaction, reuse, novelty, journey criticality and potential user impact, while every change still produces the evidence applicable to its scope. Begin by inventorying the affected journeys, views, components, templates, documents, media, controls and supported technologies. Section508.gov programme guidance illustrates documented depth ranging from automated checks and spot checks to component and comprehensive testing, with continuous monitoring and disabled-user evaluation treated separately. The ladder below is an editorial operating model, not an official risk standard, fixed score or basis for claiming that an untested route conforms.

  • Content-only change: require human content review and applicable automation; add structural, keyboard, zoom or assistive-technology review when meaning, media, documents or controls change.
  • Visual or layout change: add design review and zoom or reflow checks, plus focus review wherever interactive behaviour is affected.
  • Component or interaction change: define acceptance criteria before build; require developer checks, independent QA, relevant states and tasks, and regression coverage when the component is reused.
  • New template, critical journey or major release: use all applicable layers, trained assistive-technology testing, representative coverage, sampled conformance evaluation and disabled-user evaluation while findings can still influence the work.

Broader conformance work needs a defined scope rather than an informal collection of convenient pages. WCAG-EM describes a process that sets the evaluation goal and scope, explores key views and functions, selects representative coverage when full evaluation is infeasible, evaluates that sample and reports the findings. Sampling supports a bounded conclusion about the declared scope; it does not justify assumptions about journeys that were excluded. Record those boundaries beside the result so reviewers can see precisely what was and was not evaluated.

What should the ownership matrix record, and who owns each handoff?

Three colleagues place dark-blue cards across a five-column wall matrix while evidence sleeves and testing devices cover the table.

The ownership matrix should give every test layer a trigger and scope, earliest useful stage, responsible performer, accountable accepter, required skill and environment, retained evidence, release effect, remediation owner and retest owner. Section508.gov guidance similarly recommends defining timing, performers, depth, qualifications, environments, methods and result tracking, while its RACI and agile matrices show design, implementation, content, independent testing, defects and release evidence moving through ordinary delivery roles and artefacts. Those examples inform the handoffs below; they do not establish universal private-sector obligations.

Creators remain answerable for accessible work: designers for decisions, editors for meaningful content and developers for implementation and local checks. QA owns the test plan and independent execution, user researchers own ethical studies with disabled people, and the authorised product or release owner accepts or rejects the release. The accessibility lead stewards policy and complex interpretation. In a small team, one person may wear several hats, but the record should state which hat is active and retain trained or independent review for higher-risk work.

Accessibility stops being somebody else's final check when every change arrives with named evidence, ownership and a retest path.

Worked accessibility test-ownership matrix
Test layer, trigger and scopeEarliest stage, performer, expertise and environmentAccountable accepter and retained evidenceRelease effect and retest owner
Automated checks — every relevant code, template or content change; record the scanned scope.During authoring and build; author or developer locally, then QA or the pipeline; configured rules and representative rendered states.QA accepts coverage; retain tool version, rules, scope, result and exclusions.Defined blocking detections stop progression; creator fixes and reruns, with QA confirming.
Content review — new or changed titles, headings, labels, links, instructions, errors, alternatives, captions, transcripts and documents.During design and editorial review; trained author or editor using the content in its journey context.Content owner accepts meaning; retain reviewed items, findings, decisions and revisions.Missing or misleading essential content blocks release; editor remediates and an independent reviewer retests.
Keyboard review — new or changed controls, components, focus behaviour and representative tasks.From interactive prototype through QA; developer checks locally, QA executes independently with a physical keyboard.QA accepts task evidence; retain steps, states, focus observations, environment and defects.A blocking task or focus failure returns to development; QA repeats the complete affected task.
Zoom and reflow — layout, typography, navigation, overlays, responsive views or focus presentation changes.From design review through QA; designer and QA use applicable viewport, resizing and zoom conditions.QA accepts affected-view coverage; retain conditions, screenshots where useful, observations and exceptions.Loss or obstruction of required information or function blocks under policy; designer or developer fixes, then QA retests.
Selected assistive technology — higher-risk interactions, critical journeys and material component changes.At working prototype and build stages; trained tester uses selected current browser, operating system and technology combinations.Accessibility lead or QA lead accepts method; retain task, versions, output, impact, expected behaviour and result.Applicable blocking incompatibilities return to the creator; the trained tester repeats the affected task and states.
Evaluation with disabled users — prototypes, unfamiliar patterns and critical journeys where usability evidence can change decisions.Before obvious barriers consume the session; experienced user researcher works with appropriately supported participants.Research or product owner accepts the study record; retain scope, consent-safe findings, limitations and decisions.Findings inform redesign and prioritisation under policy; responsible creators act and research or QA verifies the response.
Sampled conformance evaluation — new templates, major releases, critical journeys or a defined wider evaluation scope.When representative views and functions are stable enough to assess; trained, suitably independent evaluator follows a declared method.Authorised owner accepts the bounded report; retain scope, sample, method, results, limitations and unresolved findings.Applicable blocking failures prevent acceptance; creators remediate and the evaluator retests affected and related coverage.

What should the core accessibility checks examine?

Two colleagues sit at a testing desk as a man uses a keyboard behind a turned-away monitor and a woman adjusts a video magnifier.

Core checks should complete meaningful tasks and inspect relevant states, not merely confirm that a tool ran or an element exists. Automation belongs in local workflows and appropriate integration pipelines, with its rules, version, scope and exclusions recorded. Human content review then judges whether page titles, headings, labels, links, instructions, error messages, captions, transcripts and text alternatives communicate useful meaning in context. WCAG 2.2 contains requirements for descriptive titles, headings and labels, link purpose and input instructions, illustrating why presence alone cannot settle content quality.

  • Complete each representative task by keyboard, operate every relevant control, follow the expected focus order, enter and leave components, observe state changes and recover from errors.
  • Confirm that focus is visible and not entirely obscured by author-created content; WCAG 2.2 also requires keyboard operation, subject to its path-dependent input exception, and a way to move focus away.
  • Resize text to 200 per cent under the applicable WCAG criterion and its stated exceptions, checking for lost content or functionality.
  • For reflow, test the 320 CSS-pixel width equivalent without horizontal scrolling and the 256 CSS-pixel height equivalent without vertical scrolling, subject to the exception for layouts requiring two dimensions for meaning or use.
  • Look for obstructed information, hidden focus, controls pushed off view and unexpected off-viewport changes instead of demanding pixel-for-pixel similarity.

The finding should preserve the task context. Record the starting state, action, expected behaviour, observed behaviour, affected users, environment, evidence and owner; after remediation, repeat the affected task and nearby states that may share the same component. This makes a keyboard, content or reflow failure reproducible for the creator and reviewable by the release owner. It also prevents a narrow retest of one corrected screen from concealing a regression introduced in another state or reusable instance.

When should specialist, assistive-technology and disabled-user evidence be added?

A blind man wearing headphones uses a Braille display and compact keyboard while a woman researcher observes and holds a blank task card.

Add specialist evidence when representative tasks, complex interactions or higher-risk changes exceed the team's routine capability, and add disabled-user evaluation while findings can still alter the work. GOV.UK guidance recommends task-based assistive-technology testing throughout development, particularly after significant features or major changes. Select current browser, operating-system and assistive-technology combinations from audience evidence, product technology, support commitments and known risks instead of copying a universal matrix. A screen reader is one compatibility method, not a simulation of every blind person's experience or proof of conformance.

  • Record the affected task and users, expected behaviour, browser, operating system, assistive technology and version, observed output, supporting evidence, remediation owner and retest result.
  • Address significant obvious barriers before a disabled-user session so participants can reveal deeper usability concerns, while still seeking early prototype input where it can shape the design.
  • Match the research method to the project stage, from focused prototype feedback to structured task-based evaluation, without generalising one participant's experience to a disability population.
  • Combine disabled-user evaluation with standards-based conformance evaluation because each answers a different question.

W3C explains that evaluation with disabled people can expose usability issues that conformance evaluation alone may miss, but cannot by itself determine whether a website is accessible. WCAG also notes that even its highest conformance level does not make content accessible to every individual with every type, degree or combination of disability. Neither point weakens standards evaluation. Together, they show why a mature programme needs both bounded conformance evidence and direct evidence about real tasks, expectations and unmet needs.

How should evidence govern release and strengthen the programme?

Three colleagues review evidence sleeves and status cards as one moves an amber card beside a keyboard and headphones for retesting.

Release should depend on the evidence required for the affected change class, not on one global accessibility score. The authorised product or release owner decides whether to proceed after confirming that applicable checks are complete, blocking findings are resolved and retested, and each result identifies scope, method, environment, outcome, owner, disposition and retest status. QA and accessibility specialists provide independent evidence; designers, editors and developers remain responsible for remediation. This mirrors the broader principle in Section508.gov examples that testing records, defects, release readiness and post-release feedback belong in the delivery system.

If organisational policy allows an exception, record its authorised owner, rationale, affected users, mitigation, expiry and follow-up separately. The exception does not change the failed result or establish conformance. After release, route reported barriers and recurring defects back into the ownership matrix, regression coverage, training, templates and future change classes. W3C's advice to evaluate throughout development supports this continuing loop; it does not prescribe a single score, cadence or maturity model. Programme measures should therefore illuminate incomplete evidence, repeated barriers and slow remediation rather than disguise them behind an aggregate pass rate.

  1. Pilot the matrix on one critical journey and complete its evidence trail.
  2. Train the named role owners and supply concise evidence templates.
  3. Add appropriate automation without treating it as the acceptance decision.
  4. Calibrate blocking rules against real findings and authorised organisational policy.
  5. Review recurring defect patterns, improve shared components and expand coverage.
  6. Use a trained evaluator for complex or disputed findings, an experienced researcher for disabled-user studies and qualified legal counsel for jurisdiction-specific interpretations.

Frequently asked questions

How do you create a web accessibility testing programme?

Define the journeys, components and change classes in scope, then separate automated, manual, assistive-technology and disabled-user evidence. Build an ownership matrix naming each method's trigger, stage, performer, accepter, retained record, blocking rule and retest owner. Train those owners, pilot one critical journey and expand the programme using recurring findings.

Who is responsible for accessibility testing in a web team?

Responsibility is distributed: designers own accessible decisions, content teams own meaningful content, developers own implementation and local checks, QA owns the test plan and independent execution, and researchers own ethical disabled-user studies. The accessibility lead stewards policy and complex methods, while the authorised product or release owner makes the release decision.

Can automated accessibility testing prove WCAG conformance?

No. Automated testing can repeatedly detect some programmatic conditions, but W3C states that no tool alone can determine whether a site meets accessibility standards. Knowledgeable human evaluation and every other method applicable to the declared scope are still required.

When should teams test with screen readers and disabled users?

Use trained screen-reader or selected assistive-technology testing for representative tasks, complex interactions and higher-risk changes throughout development. Involve disabled people with prototypes and critical journeys while their findings can influence decisions, preferably after significant obvious barriers have been addressed. User evaluation complements rather than replaces conformance evaluation.

Which accessibility findings should block a release?

Each organisation must define authorised blocking rules for its change classes and context. Release evidence should show that required checks are complete and blocking findings have been resolved and retested. Any permitted exception must remain explicit and time-bound, and it must never be presented as changing the result or proving conformance.

WebChorus logo

WebChorus Editorial Team

We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.