Use requirements to screen the CMS market, then decide among shortlisted platforms by having representative users run the same buyer-owned publishing scenarios. A polished demonstration may show that a page can be published, but it rarely reveals which revision was approved, what happens when a locale is incomplete, how a failed schedule is recovered or which dependencies sit outside the licence. Comparable tests make those operational conditions visible before selection, when the team can still price, contract for or reject them.
The evaluation method at a glance
Screen candidates against mandatory constraints before investing in detailed scenario testing.
Give every candidate the same versioned samples, actors, starting states, expected outcomes and failure variations.
Test authoring, review, localisation, reuse, permissions, scheduling, correction, archiving, integration and recovery.
Record demonstrated outcomes separately from effort, configuration, plan tier, extensions, custom code and external services.
Treat a successful proof of concept as bounded evidence, not certification of security, accessibility, compliance, scalability or continuity.
How should a CMS evaluation move from screening to operational proof?
Use non-negotiable requirements to narrow the field, then use comparable buyer-run scenarios to produce decision evidence. Architectural fit, security controls, accessibility obligations, data handling, legal constraints, commercial terms and support coverage remain gates; there is little value in testing a candidate that cannot meet them. Government Digital Service guidance similarly recommends understanding the service context and using prototypes to test users, interfaces, data, compliance, security considerations and technical constraints before a long-term commitment.
Confirm each candidate passes the mandatory gates using documentary and specialist review.
Freeze one version of the samples, roles, starting states, variations and outcomes for the shortlist.
Let representative users attempt the normal path before the vendor explains configuration or alternatives.
Repeat the same failure and exception conditions, then retain the evidence in a common format.
The ten scenario families used here are an adaptable editorial framework, not an official standard or a universal prescription. Tailor their depth to the organisation's publishing risks, but preserve comparability: changing the input, user or expected result for one vendor turns the exercise into a set of anecdotes. Feature lists and vendor demonstrations still help with discovery. They simply answer a different question from whether the operating team can achieve the required outcome safely and predictably.
What must every repeatable CMS scenario specify?
A repeatable scenario must fix the conditions, observable outcome, captured evidence and failure rule before testing begins. Prototype-based evaluation can expose assumptions about users, interfaces, data, compliance, security and technical constraints, but only if participants know what they are trying to prove. This scenario card is a buyer-oriented method for creating that discipline; adapt it to local governance rather than presenting it as a consensus standard.
Purpose: the operational question and risk the test is intended to expose.
Actors and starting state: who performs the task and the exact initial workflow, permission and environment conditions.
Task and variation: the normal end-to-end job plus one meaningful complication, exception or failure.
Expected outcome: what must be observable, including what must not occur.
Evidence: rendered pages, screens, audit entries, API responses, exports, timestamps, notifications and participant observations.
Effort: elapsed time, steps, hand-offs, training prompts, configuration, extensions, custom code and external help.
Dependencies and failure condition: every prerequisite, plus the outcome that makes the result failed or still open.
Keep success separate from the work required to produce it. A scenario can reach its visible endpoint yet still reveal a material plan-tier restriction, manual control, implementation dependency or unsafe privilege. Mark it failed or open when a must-pass outcome is missed, the state remains ambiguous, required evidence is absent or unresolved work is essential to the result. Record the explanation alongside the observation rather than allowing it to disappear into presentation notes.
A feature claim says a CMS can do something; a representative scenario shows what your organisation must do to make the outcome happen.
How should authoring and review scenarios expose everyday workflow risk?
Authoring and review scenarios should prove that representative people can create accessible, structured content and publish the intended revision without disturbing the current live version. W3C's ATAG overview covers both accessibility of the authoring interface for disabled authors and support for producing accessible web content. A bounded task can reveal barriers and support behaviour, but it cannot establish ATAG or WCAG conformance. Reserve that judgement for a properly scoped accessibility evaluation.
Ask a frequent author and an occasional author to create the same article with headings, links, an image and alternative text, metadata, a related-content reference and responsive previews.
Add a keyboard-only critical path and one validation or accessibility mistake that each author must identify and correct.
Keep the published version live while an author submits a revision, a reviewer comments and returns it, and an authorised publisher releases the corrected revision.
Create a newer parallel draft during review, then verify exactly which revision was approved, what each role could change and what the history retained.
Drupal documents one moderation model in which a published version remains live while a working revision moves through states and transitions. WordPress revision endpoints illustrate how systems can expose identifiable prior records containing content, authors, timestamps and statuses. Together, those behaviours make a parallel-draft collision worth testing, although the test is an operational recommendation rather than a universal CMS requirement. Workflow labels alone do not prove revision identity, separation of duties or the public result.
How can localisation and reuse tests reveal hidden content dependencies?
Localisation and reuse tests should expose which source, state and dependency reaches each destination when content changes. Drupal documents separately moderated translations and notes that a new translation can begin from the published source rather than the latest working revision. Contentful documents requested locales, a default locale and configured fallback values for missing localised content. These are product-specific examples, so the evaluation should observe the candidate's actual behaviour instead of assuming that a locale field represents a complete translation operation.
Create, review, preview and publish a secondary-locale edition independently, then change the source after translation has started.
Leave one localised field empty and inspect the delivered response, fallback value, metadata, permissions and publication state.
Reference one governed fact, disclaimer, profile or contact block from several destinations, update it once and preview every dependent use.
Give one destination a different context or release time and test whether the exception stays explicit without creating silent divergence.
Contentful also documents references that link one reusable entry from multiple destinations, with published changes appearing in entries that reuse it. That establishes a reuse mechanism, not the behaviour of every frontend, cache, release or rollback path. Capture the dependency graph, impacted previews, publication order, cache result and reversal steps. The decisive question is whether editors can understand and control propagation when one destination needs an exception, not whether the product displays a reference field.
What should permissions, scheduling and correction scenarios prove?
These scenarios should prove effective authority at real boundaries, actual timed-release behaviour and an accountable route through urgent correction and reversal. WordPress documents capabilities that distinguish reading, editing one's own or others' content, publishing, importing, exporting and administration. That model illustrates why role names are insufficient evidence. Test what each user can do through visible controls, direct routes and relevant APIs, including denied actions at content-type, locale, field or transition boundaries where the shortlisted product supports them.
Assign least-privilege author, reviewer, translator, publisher and administrator roles, then attempt defined allowed and forbidden actions.
Schedule a coordinated publish and later unpublish action in a named IANA time zone, including referenced content and assets.
Introduce a validation failure or last-minute time change and capture preflight results, notifications, partial-release behaviour, public state and recovery.
Correct a material live error, verify each delivery channel and cache, then restore the prior approved revision while preserving who changed what, when and why.
Contentful documents scheduled actions with dates, IANA time zones, permissions, notifications, validation failures and product-specific limits. WordPress revision endpoints expose prior records, but stored revisions alone do not prove safe rollback, approval handling, audit completeness or cache convergence. Record actual execution timestamps and delivered content rather than accepting a scheduled-status label. For urgent correction, measure the smallest authorised approval path and verify that the restored public state matches the intended approved revision.
How should archiving, integration and recovery be tested without overstating the result?
Test archiving as a defined public outcome, integration beyond the happy path and recovery as a bounded restore exercise. GOV.UK guidance distinguishes withdrawal, which can retain a URL with an explanation, from unpublishing, which removes content and can establish a redirect. Those are platform examples rather than universal business rules, but they show why archive, withdraw, unpublish, restrict, delete and redirect should not be collapsed into one checkbox. Define the organisation's intended state before the task begins.
Retire and then reinstate representative content, checking URL responses, links, search, feeds, APIs, attachments, permissions, history and downstream effects.
Create or update realistic content through the intended API or connector, then send invalid input, repeat a request and delay or fail the consumer.
Capture authentication boundaries, error detail, replay, ordering, duplicate behaviour, logs, limits and manual recovery rather than recording only a successful call.
Export agreed content, assets, models, relationships, identifiers, redirects and relevant operational state, then restore a representative set in isolation.
The WordPress REST API documents anonymous public resources and authenticated private management actions, but endpoint availability does not establish an organisation's security, mapping, observability, ordering, retry behaviour or scale. OASIS CMIS likewise defines a common repository model and bindings without exposing every repository capability. NIST describes contingency planning as coordinated plans, procedures and technical measures for recovering systems, operations and data after disruption. A trial restore informs that work; it does not certify production recovery or business continuity.
Ten adaptable scenario families and the evidence that makes each test useful
Scenario family and buyer-run task
Expected observable outcome
Decisive evidence
Failure or exception variation
Authoring: create a structured article
Valid content and previews retain required structure and accessibility fields
Rendered preview, keyboard path, validation result, time and assistance
Occasional author corrects an introduced accessibility or validation mistake
Review: approve and publish a revision
The intended revision publishes while the previous version remains live until release
Revision identity, comments, transitions, timestamps and audit history
A newer parallel draft appears during approval
Localisation: publish a secondary-locale edition
Locale state, metadata and delivery remain independently understandable
Source-change signal, fallback result, permissions and delivery response
Source changes and one localised field is missing
Reuse: update a governed shared item
Only the intended dependent destinations receive the approved change
Dependency view, impact previews, release order, cache state and rollback
One destination needs different context or timing
Permissions: perform allowed and forbidden actions
Authorised work succeeds and restricted work is denied at each boundary
Interface behaviour, direct-route result, API response and audit identity
One content type, field, locale or transition is restricted
Scheduling: coordinate publish and unpublish actions
Dependencies change at the intended local time with an observable final state
Stored time zone, preflight, timestamps, notifications and public result
Validation fails or the release time changes late
Correction: repair and reverse a live error
The approved correction reaches every channel and can be safely reversed
Revision comparison, approvals, timestamps, caches and audit record
The correction proves wrong and the prior approved revision is restored
Archiving: retire and reinstate content
The URL and downstream systems show the organisation's defined retirement state
Response, explanation or redirect, search, APIs, assets and history
The retirement decision is reversed
Integration: create or update content through an interface
Identifiers, status and delivered content reconcile across systems
Requests, responses, events, logs, duplicates, ordering and recovery
Input is invalid, repeated or delivered to a failed consumer
Recovery: export and restore a representative set
Agreed content, assets and relationships return with gaps made explicit
Export contents, restore log, validation, elapsed effort and dependencies
Deletion, corruption or platform unavailability changes the starting state
How should teams turn scenario evidence into a defensible CMS decision?
Turn the evidence into a decision by separating mandatory gates, observed outcomes, operational effort, dependencies and unresolved risks. Do not let an aggregate score conceal a failed gate. This is a procurement method rather than a formal scoring standard, so each organisation must set its own gates, weights and thresholds before testing. Government Digital Service guidance considers adaptability, control of stored data, security risk and total ownership cost, but it does not provide a universal CMS scoring formula.
Record each must-pass result independently from usability observations and elapsed effort.
Attribute every outcome to native capability, configuration, plan tier, add-on, extension, custom code, partner service, external system or roadmap commitment.
Convert unresolved migration, integration, testing, training, configuration and manual-control work into scope, cost, contract terms, explicit risk or rejection.
Retain versioned scenario cards, samples, participant roles, observations, timestamps, screens, API records, exports, assumptions, gate results and the decision log.
Carry that evidence packet into procurement and implementation so promised dependencies can be verified rather than rediscovered. Scenario success remains bounded: it does not certify accessibility conformance, security, privacy, legal compliance, scalability, disaster recovery, business continuity or total cost. Bring qualified accessibility, security, privacy, legal, data, infrastructure and continuity professionals into decisions requiring specialist judgement. The strongest shortlist decision is not the neatest demonstration; it is the one whose outcomes, effort, dependencies and open risks the organisation can explain and govern.
CMS evaluation questions
How do you evaluate a CMS?
First screen candidates against mandatory architectural, security, accessibility, data, legal, commercial and support constraints. Then have representative users run the same predefined, buyer-owned publishing scenarios in every shortlisted platform. Compare observable outcomes, effort, dependencies and unresolved risks.
What should a CMS proof of concept include?
Include a realistic sample, named actors, an exact starting state, the normal task, a meaningful failure variation and a predefined observable outcome. Capture rendered pages, screens, audit records, API responses, exports, timestamps and participant observations. Record time, steps, training, configuration and external dependencies separately from success.
What should a CMS vendor demonstration prove?
It should show whether representative users can complete buyer-owned scenarios under agreed conditions. Let users attempt the default path before the vendor explains configuration, add-ons or alternatives. A demonstration should expose prerequisites and failure behaviour, not merely confirm that a feature exists.
Which publishing scenarios should an enterprise CMS evaluation test?
A practical adaptable set covers authoring, review, localisation, reuse, permissions, scheduling, correction, archiving, integration and recovery. Select realistic samples and exceptions for each family. Treat the set as an editorial framework, not a universal product specification.
How should CMS evaluation results be scored?
Keep must-pass gates separate so a weighted total cannot conceal a disqualifying failure. Score or describe observed outcomes, usability and effort independently, then record every plan, configuration and delivery dependency. Set weights and thresholds before testing because no universal scoring formula suits every organisation.
References & Sources
This article was researched using the following sources:
We cover the decisions that shape a website long after launch. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial standards. We disclose commercial relationships wherever they exist.
Use a nine-field page-purpose brief to test website requests, choose whether to update or create content, and give Australian teams a focused writing contract.
Build an eight-domain website governance model that gives each decision an owner, clear delegation, required input, an escalation route and a durable record.