Summary

  • W3C published WCAG Evaluation Methodology 2.0 as a Group Note on 23 July 2026. It is endorsed by the Accessibility Guidelines Working Group, not by W3C or its Members, and it does not add to or replace WCAG requirements.
  • The methodology first builds a structured sample around common views, essential functions, content types, technologies and other relevant material. Its random set is then 10% of that structured set—not 10% of the digital product.
  • W3C’s example begins with 80 structured samples and adds eight random samples, making 88. The draw need not meet strictly scientific criteria, but it must span the full product scope, avoid a predictable pattern and have its method recorded.
  • Complete processes are a separate safeguard: when a sampled view belongs to a process, all pages or views in the relevant default and critical branch sequences must be included.
  • In most uses, sampled evaluation under WCAG-EM alone cannot support a WCAG conformance claim for the whole product. Every decision surface should therefore carry a sample-to-decision receipt that identifies the denominator, selection rationale, scope, responsible parties and exact statement type.

The denominator sits inside the method

On 23 July, the W3C Accessibility Guidelines Working Group published WCAG-EM 2.0 as a Group Note. Its full name—WCAG Evaluation Methodology—can make it sound like another layer of the accessibility standard. The document is more carefully bounded. It is informative guidance for evaluating how well a representative sample of a digital product conforms to WCAG 2. It defines no new success criteria, adds no conformance requirement and neither replaces nor supersedes WCAG.

The publication status draws a second boundary. The Accessibility Guidelines Working Group endorses the Note, but W3C and its Members do not. A Group Note is a stable reference rather than a Recommendation-track standard, and the W3C patent policy brings no licensing commitment to it. That does not make the method unimportant. It means its authority should come from the quality and declared use of the method, not from an imagined W3C certification.

The method is substantial. Evaluators define the product scope, target WCAG version and level, and accessibility support baseline. They explore common views, essential functionality, varieties of content and technology. They then assemble a representative sample, evaluate it, compare two selection paths and record the comparison.

The memorable percentage appears only after that work. Step 3.1 creates a structured sample selected to reflect what the evaluator learned about the product. Step 3.2 adds a random set whose size is 10% of the structured set. W3C gives an unambiguous example: if the structured set contains 80 samples, add eight random ones, for 88 in total.

The denominator is therefore 80, not the total population of pages, screens, documents or states. A product might contain hundreds, thousands or dynamically generated millions of possible views. The formula says nothing about the share of that population tested. Calling an 80-plus-eight evaluation “a 10% product audit” would change the claim.

The two samples do different jobs

The structured set is purposive. It should represent common views, essential functionality, different sample types, relied-upon technologies and other relevant material. A carefully chosen page may cover several criteria at once. Product size matters, but so do age, complexity, implementation variety, consistency, evaluator knowledge and the confidence required. WCAG-EM does not turn those judgments into one universal sample-size table.

The random set tests the structured set. It is intended to reveal content types or findings that the evaluator’s exploration and purposive selection missed. Step 4.3 requires the evaluator to compare the outcomes. If the random set uncovers a new type of content or a new finding, the structured set was not sufficiently representative. The evaluator must return to exploration and selection, add samples and repeat the comparison.

This is a feedback control, not a one-time statistical blessing. W3C says the random sample need not be selected according to strictly scientific criteria. It may be generated through a crawler, a list, logs, search or other practical methods. But the eligible selection must span the entire product scope, choices must not follow a predictable pattern and the method should be recorded for reliability and replicability.

Those qualifications resist two opposite mistakes. The random set is not meaningless just because it lacks a formal confidence interval. It can expose blind spots in an expert-built sample. But the word “random” does not convert eight views into an estimate about every unseen view with a declared margin of error. Its authority comes from the specified diagnostic role.

A complete process cannot be sliced at the convenient screen

Step 3.3 adds a different protection. If a selected sample belongs to a complete process, all pages or views in the process must enter the sample set. Evaluators identify the starting point, the default sequence and commonly used or critical branches. They also record the actions needed to move from one state to the next, because a URL alone may not reproduce a login, purchase, booking or account-creation path.

This prevents a favorable static screen from standing in for the interaction that makes the product useful. A checkout is not only its basket page. An application is not only its first form. Errors, confirmations, dialogs, preference states and changed content can determine whether the process conforms. WCAG itself treats complete processes as a conformance requirement; WCAG-EM translates that requirement into sample construction and evaluation work.

The structured, random and process rules should not be flattened into a single coverage number. The structured set represents mapped product variety. The random set probes for omissions. Complete-process expansion preserves interaction continuity. Each answers a different evidentiary question.

A sample finding is not a whole-product claim

WCAG-EM states the limit unusually plainly. WCAG conformance claims for an entire website cannot be based only on evaluation of a selected subset of pages and functionality, because unidentified errors may remain. In most uses of the methodology, evaluators inspect a sample. In most such situations, using WCAG-EM alone therefore does not enable a WCAG conformance claim for the target product.

This is not a confession that sampled evaluation is weak. A sample can inform remediation, procurement diligence, redesign, monitoring and risk management. The distinction concerns the scope of the conclusion. “Every sampled view met the target in this evaluation” is evidence. “The product conforms” is a broader proposition whose WCAG requirements apply to every page in scope or to a process that assures each page satisfies them.

WCAG-EM offers an optional evaluation statement for the narrower result. The statement identifies its date, WCAG version, target level, product scope, relied-upon technologies and accessibility support baseline. It may also describe partial conformance. The product owner must commit to the validity and continued accuracy of the statement.

That maintenance duty matters because digital products change. A sound finding about a dated release can become stale after a component library, payment flow, content system or third-party integration changes. A public sentence without a product version or re-evaluation trigger can preserve its grammar after its evidence has expired.

The report already contains most of the receipt

The Group Note’s mandatory reporting step is stronger than a percentage headline. It calls for the evaluator, commissioner, date, scope, conformance target, accessibility support baseline, technologies, structured samples, random samples and selection method, complete processes and outcomes. It also recommends issue descriptions, reproduction steps and examples for requirements not met. Optional records can archive samples, browsers, assistive technologies, tools and test methods, subject to security and privacy controls.

The missing governance move is to bind that record to the decision surface. A procurement memo, public accessibility page or executive dashboard may quote a conclusion while the full report remains confidential or several links away. The compact statement should therefore carry a sample-to-decision receipt.

At minimum, the receipt should identify the exact product and version, the defined scope, WCAG target and accessibility support baseline. It should state the structured-set size and rationale, the random-set count and its structured-set denominator, the random-selection method, complete processes included and whether the random comparison forced another sampling round. It should name the evaluator, commissioner and responsible product owner, state the evaluation date and expiry or change trigger, and label the output as an internal finding, a WCAG-EM evaluation statement or a WCAG conformance claim.

If an optional aggregate score appears, its formula and limitations belong in the same receipt. WCAG-EM says no single metric is currently known to combine the required reliability, accuracy and practicality, warns that aggregate scores can mislead and notes that WCAG provides no rating scheme. A score is not forbidden; its method must be documented and made available to the commissioner. The governance rule is simpler: never let the score outrank the findings or the percentage outrank its denominator.

What the evidence does not establish

Nothing in the Note says that ten percent of a structured set is correct for every product regardless of context. The absolute structured-set size is a professional judgment informed by product conditions and required confidence. Nor does the document promise formal statistical randomness. The method deliberately permits practical selection techniques while requiring broad scope, unpredictability and documentation.

The whole-product warning does not invalidate representative evaluation. It prevents a scope jump. An organization can responsibly say what its sampled evidence found, use those findings to repair systems and publish a bounded evaluation statement. It should not transform that result into W3C approval or universal product conformance.

This Article identifies no misleading vendor, deficient evaluator or inaccessible product. It reviews a public methodology and the incentives around its reuse. The proposed receipt is a governance extension built from the Note’s reporting logic, not an existing W3C badge or requirement.

Sources

  1. W3C News — Group Note: WCAG Evaluation Methodology (WCAG-EM) 2.0
  2. W3C — WCAG Evaluation Methodology (WCAG-EM) 2.0, Group Note of 23 July 2026
  3. W3C — WCAG-EM 2.0 publication history
  4. W3C — Standards and other document types
  5. W3C — Process Document, 18 August 2025
  6. W3C — Web Content Accessibility Guidelines (WCAG) 2.2
  7. W3C WAI — WCAG-EM Report Tool
  8. Heng Lu — On the Agency Problem at the Core of Internet Governance