Audit representative tasks and every plausible route to their outcomes before auditing menus or sketching a new sitemap. A crowded navigation bar, a high-exit page or complaints about findability may indicate a genuine structural problem, but none identifies the cause on its own. The defect could instead be missing content, an unclear label, a weak contextual link, poor search results or an unusable control. Task-level evidence keeps a redesign or migration decision tied to journeys that matter.
Key decisions
Make representative user tasks and their plausible routes the unit of the audit.
Treat analytics, search logs, support contacts and expert reviews as signals that require interpretation.
Use card sorting for grouping, tree testing for hierarchy and labels, and usability testing for rendered routes.
Classify the actual failure before recommending a structural change.
Make the smallest evidence-supported repair, then retest the affected task.
What decision should your audit inform?
The audit should inform one bounded decision, such as whether to repair a section, change labels, prepare content for migration or investigate a broader redesign. Information architecture encompasses organisation, labels and navigation that help people find information, understand their location and options, and complete intended tasks. Define the decision before opening the sitemap, or the work can easily expand into a general critique with no agreed standard for action.
Record the audiences, goals, starting contexts, page types, devices, locales, permissions and journey states included. A customer arriving from a search engine may have a different route from a staff member entering an authenticated portal, while mobile navigation may expose choices differently from desktop navigation. Conclusions should therefore apply to the specified contexts, not to an imagined average user or to the entire estate without evidence.
Known fact: verified information about the website, content or operating context.
Observed behaviour: something recorded during research or testing with users.
Expert inspection: a likely issue identified through professional review but not yet observed with users.
Hypothesis: an explanation or proposed remedy that still needs evidence.
Keep the boundary explicit. A task-based IA audit diagnoses whether the existing route system supports selected tasks; it is not a content inventory, technical SEO audit, accessibility conformance assessment or redesign exercise. Those activities may become necessary, but combining them prematurely makes provenance unclear. A finding about an inaccessible control, for example, should trigger appropriate accessibility evaluation rather than being relabelled as proof that the entire hierarchy is wrong.
How do you build a representative task set from evidence?
Build the task set from evidence about what people are trying to achieve, how they currently attempt it, the problems they encounter and the outcome they need. Write each task in language the intended audience would recognise, without revealing a destination label or suggesting the navigation answer. “Find the support arrangements for an existing contract” tests an outcome; “Go to Customer Support” mainly tests whether a participant follows the wording supplied.
Useful inputs include prior user research, analytics, internal-search queries, support demand, feedback, interviews, observation and colleagues who work directly with users. Treat behavioural and operational records as signals, not self-explanatory findings. An exit may mean success, confusion or a move to another channel. A search query may reveal preferred language, but it does not by itself explain intent or identify the right structural remedy. Stakeholder assertions remain hypotheses until user evidence supports them.
Include frequent tasks, but also consequential, difficult and underserved tasks.
Name the relevant audience and the trigger that brings the task into view.
Describe realistic starting contexts rather than assuming a home-page visit.
Define the successful content, action or state and record the evidence source.
Note important device, locale, permission or journey-state variations.
Review the set for coverage before testing. A documented Digital.gov study derived realistic scenarios from earlier research and checked their coverage, but its particular task and participant counts belonged to that case. Your task set should instead match the decision and the diversity of relevant journeys. Preserve provenance beside every task so a later challenge can distinguish genuine user evidence from operational inference, expert judgement or an untested internal request.
What should the task-to-route worksheet capture?
The worksheet should connect each evidenced task to its successful outcome, plausible routes, inspected cues, observed behaviour, diagnosis, proposed change and retest. Begin with realistic entry points: an external landing page, a section hub, a saved deep link, an authenticated area or internal search. Then map browse, contextual-link and search routes. Do not assume that every journey begins on the home page or that one successful path serves every audience and context.
At each decision point, record the cue a person can actually perceive, the expectation it creates, the destination reached and the available recovery route. Capture the label, heading, grouping, breadcrumb or location cue, related link and search query involved. The WCAG 2.2 Success Criterion on Multiple Ways requires more than one way to locate pages within a set, subject to exceptions for process steps and results; documented approaches include related links, site maps, search and comprehensive navigation.
A task-based IA audit asks whether people can reach a needed outcome through realistic routes, with evidence the organisation can act on.
A compact task-to-route record
Task, audience, trigger, outcome and source evidence
Starting contexts, plausible routes and inspected cues
Observed behaviour, measures, failure mode and evidence strength
Smallest proposed change, owner and retest
State the outcome in user language; name the audience, trigger, successful destination and provenance.
List external entry, browse, contextual-link and search routes; record labels, groupings and orientation cues.
Capture completion, assistance, wrong turns, backtracking, reformulation and reasoning; classify the evidence.
Specify a bounded repair, accountable owner, affected routes and a task-level retest.
Keep device, locale, permission and journey-state variations where they materially alter the task.
Record the promise made at each decision point, the destination reached and the recovery available.
Distinguish observed failure from expert concern and note whether an alternate route was valid.
Define what evidence would show improvement without inventing a universal success threshold.
Use the same record from initial inspection through testing, prioritisation, ownership and retesting. That continuity prevents a common handover problem: the recommendation survives, but the evidence and affected task disappear. It also exposes gaps. If a proposed navigation change has no recorded failure, or a severe finding has no audience, context or observed behaviour attached, the team can resolve the weakness before approving costly work.
How do you inspect the complete route rather than menus alone?
Inspect every material component between a realistic starting context and successful completion. That includes external entry pages, global and local navigation, hubs or indexes, page groupings, headings, breadcrumbs or other location cues, contextual links, internal search, and the final content or action. A menu can appear orderly while the destination makes no sense, a crucial next step is absent, or a search result sends the user to an outdated page.
Judge labels by the expectation they create in context. Microsoft guidance recommends planning navigation around users' perspectives, common tasks and mental models, with labels that are accurate, familiar, concise, scannable and distinguishable. Short wording is not automatically clear: neighbouring choices must remain meaningfully different, and the destination must fulfil the promise. The WCAG 2.2 Success Criterion on Headings and Labels also requires provided headings and labels to describe their topic or purpose.
Coverage: the necessary content, action or state is absent or incomplete.
Entry: a likely starting context offers no plausible route.
Label or grouping: the cue misleads, or the destination sits where people do not expect it.
Orientation or cross-link: people cannot place themselves or find the next relevant step.
Search: relevant queries return poor, missing, misleading or difficult results.
Consistency or interaction: repeated mechanisms shift unexpectedly, or a rendered control prevents use.
Repeat important tasks across page types, devices, locales, permissions and states when those contexts materially alter available routes. For repeated navigation mechanisms, the WCAG 2.2 Success Criterion on Consistent Navigation concerns consistent relative order unless the user initiates a change; it does not ban local or secondary navigation. Internal search can also be a valid or preferred alternate route, so its use alone is not proof that navigation has failed.
Which research method should validate each uncertain path?
Choose the method that answers the unresolved evidence question. Expert inspection and existing behavioural data can identify likely defects, but an inspected concern is not an observed user failure. State what remains uncertain before recruiting participants or building a study. This avoids using a convenient method merely because the team knows it, then claiming results about parts of the journey that the method never exposed.
Use card sorting when the question concerns expected grouping or category language; open sorts also reveal labels participants create.
Use tree testing when the question is whether hierarchy and labels allow people to find a destination without page design.
Use task-based usability testing when the question spans rendered navigation, page cues, controls, contextual links, search, recovery or completion.
Use inspection and existing evidence to focus research, while keeping hypotheses visibly separate from observed behaviour.
NIST describes usability testing as representative users performing representative tasks, with possible evidence including completion, errors, time, qualitative comments and satisfaction. Select only measures that inform the decision. Assistance, wrong turns, backtracking, search reformulations, destination confidence and participant reasoning may be valuable too. Divergent paths can reveal ambiguous groupings or labels even when some participants arrive successfully, although a different path may also be a perfectly valid route.
Card sorting, tree testing and rendered-site usability testing are complementary, not interchangeable. Card sorting does not validate a live route. Tree testing isolates hierarchy and labels but removes much of the rendered interface, including page design, controls, contextual links and search-result behaviour. Usability testing can cover those elements, but its findings still apply to the tasks, participants and contexts studied rather than proving that every journey across the website works.
How do you turn findings into bounded changes or a justified redesign case?
Turn findings into action by classifying the failure, exposing the prioritisation inputs and selecting the smallest change supported by the record. Do not convert every symptom into a navigation recommendation. A failed task may require a content correction, clearer label, better contextual link, regrouping, search tuning, a repaired interaction, section-level restructuring or, only in stronger cases, broader redesign. The diagnosis should explain why the proposed intervention matches the observed breakdown.
Task importance and the consequence of failure
Audiences and contexts affected
Observed frequency and consistency of the failure
Strength and provenance of the evidence
Dependencies, ownership and practical remediation constraints
The task and route that will be retested
Keep those inputs visible instead of compressing them into a universal weighted IA score. A numerical total can conceal weak evidence or make unlike consequences appear comparable. Record the judgement in plain language: which task is affected, who encountered the problem, where it occurred, how confident the team is and what must change first. Retest the affected tasks and routes before declaring success, using measures suited to the original decision.
Escalate to broad structural change when important task failures are repeated across relevant contexts, supported by observation, structural rather than local, and not reasonably repairable through bounded interventions. Finish with that decision, not a new sitemap by default. Bring in an experienced information architect or UX researcher when the task set, research design or trade-offs exceed the team's capability. If findings raise accessibility questions, involve a qualified accessibility specialist: selected WCAG checks do not establish whole-site conformance.
Task-based IA audit questions
What is included in an information architecture audit?
A task-based audit traces evidenced tasks through entry points, navigation, labels, groupings, orientation cues, contextual links, internal search and destination completion. It records observed behaviour, diagnoses the failure and connects any recommendation to a retest. It remains distinct from a content inventory, technical SEO audit, accessibility conformance assessment or redesign.
How many users or tasks are needed for an IA audit?
There is no universal participant or task count in the cited authorities. Scope the study around the decision, the diversity of relevant audiences and contexts, task consequence, uncertainty and the strength of evidence required. Counts reported in an individual case study should not be adopted as general thresholds.
Can analytics identify website navigation problems?
Analytics, internal-search logs, exits and support contacts can identify journeys worth investigating. They do not independently prove user intent, the cause of a behaviour or the correct structural remedy. Combine them with qualitative evidence, observation or testing before diagnosing a navigation failure.
Does internal search use mean the navigation has failed?
No. Search can be a valid or preferred alternate route, and providing multiple ways to locate content can support different needs. Examine query reformulation, result relevance, destination confidence and task completion before deciding whether the problem lies in navigation, search or neither.
When does an IA audit justify a website redesign?
A redesign case is strongest when important task failures recur across relevant contexts, are supported by observed evidence, are structural rather than local, and cannot reasonably be repaired through bounded changes. Test label, link, content, grouping, search and interaction repairs where appropriate, then retest the affected tasks before escalating.
References and Sources
This article was researched using the following sources:
We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.