Summary
- W3C’s Sustainable Web Interest Group published the first Group Note Draft of Web Sustainability Guidelines Impact Ratings on 18 August 2026. It is not endorsed by W3C or its Members and is not a Recommendation.
- A 6 August repository change renamed “Impact Measurement” as “Impact Ratings,” saying the change followed meeting consensus. The published method nevertheless converts three ordinal impact judgments and one timeframe into one additive score.
- In the 1 September public JSON snapshot, 71 guidelines produce 35 distinct four-field point vectors but only 10 totals. Twenty-three guidelines have all three impact dimensions marked Indeterminate; 56 have at least one Indeterminate field.
- Six different rating vectors produce a score of 7. The number can therefore support sorting, but it cannot preserve which impact, evidence gap or timeframe produced the rank.
- W3C should keep the total only as a declared view over a durable impact-vector receipt containing the original ratings, rationale, evidence, expert and consensus provenance, uncertainty, objections, formula version, use-specific weights and correction history.
The rename was a useful admission of method
The public sequence begins before the Note itself. On 6 August, pull request 16 in the W3C sustainableweb-impact repository changed the work’s title from “Web Sustainability Guidelines Impact Measurement” to “Web Sustainability Guidelines Impact Ratings.” The pull request says the change followed consensus at that day’s meeting. The associated commit says the same thing in shorter form: “Changed Measurement > Ratings per Consensus.” It was merged within minutes.
On 18 August, the Sustainable Web Interest Group published the first Group Note Draft under the new title. Its status matters. A Group Note Draft is not endorsed by W3C or its Members, and the Interest Group does not produce Recommendation-track documents. The publication is a public work product and an invitation to scrutiny, not a certified sustainability standard.
The new noun fits the disclosed method better. Each guideline receives an ordinal judgment—Indeterminate, Low, Medium or High—for three categories: People, Planet and Prosperity. It also receives an expected timeframe: Long, Medium or Short. W3C’s announcement says that when quantitative measurement is not possible, ratings draw on available research and expert input. Each row includes a rationale and, where appropriate, suggested metrics.
That is a legitimate form of editorial and expert synthesis. It is not the same thing as measuring tonnes of carbon, response time, wages or injury rates. Calling it a rating makes the epistemic boundary visible. The problem begins one step later, when four judgments are made to behave like one neutral quantity.
Four ordinal fields become one total
The published mapping is simple. Indeterminate, Low, Medium and High receive zero, one, two and three points. Long, Medium and Short timeframes receive one, two and three points. The points for People, Planet, Prosperity and timeframe are then added. Higher totals are described as greater impact achieved in less time.
The Note offers the score as a way to assess relative impact and prioritize which guidelines to act on. It also says ratings and scores may support sustainability reporting and assessment, and may be referenced alongside established frameworks such as GRI where relevant. A small sorting device can therefore travel into decisions with larger economic and reputational consequences.
Addition looks transparent because anyone can recompute it. But reproducible arithmetic is not the same as a neutral comparison rule. The sum makes at least four policy choices.
First, it treats People, Planet and Prosperity as commensurable on the same three-point interval. Second, it gives the three impact dimensions equal weight. Third, it allows one field to compensate for another: a short timeframe can raise a guideline whose impact is unknown, while a strong rating in one dimension can offset an Indeterminate rating in another. Fourth, it treats Indeterminate as zero points even though “not known” is not equivalent to “no impact.”
None of those choices is automatically wrong. A generic navigation view needs some rule. But the choices belong to governance, not arithmetic. A public total should state whose decision it serves, what trade-offs it permits and which information it discards.
The public JSON shows what the sum erases
BTW calculated the following figures directly from the public impact.json retrieved on 1 September. They are reproducible observations from that snapshot, not claims made by W3C. The dataset contains 71 guidelines: 17 in User Experience Design, 16 in Web Development, 12 in Hosting, Infrastructure and Systems, and 26 in Business Strategy and Product Management.
Those 71 rows contain 35 distinct combinations of People, Planet, Prosperity and timeframe points. The additive rule reduces them to 10 observed totals, from 1 to 10. Score 5 is the most common, with 14 rows; scores 3 and 7 follow with 12 and 10. The distribution is not itself a defect. It shows the compression ratio: many different impact stories occupy the same scalar bucket.
Score 7 makes the problem concrete. Six different rating vectors reach it:
- High People, Low Planet, Indeterminate Prosperity, Short timeframe;
- Medium People, Medium Planet, Medium Prosperity, Long timeframe;
- Medium People, Indeterminate Planet, Medium Prosperity, Short timeframe;
- Medium People, Low Planet, Low Prosperity, Short timeframe;
- Indeterminate People, High Planet, Low Prosperity, Short timeframe; and
- High People, Indeterminate Planet, High Prosperity, Long timeframe.
These are not interchangeable policy objects. A public-sector accessibility team, a hosting buyer and a product finance committee could rationally rank them differently. The total says only that one fixed addition rule produced seven.
The Indeterminate fields sharpen the distinction. Twenty-three of the 71 guidelines have Indeterminate ratings for People, Planet and Prosperity simultaneously. Their totals—one, two or three—come entirely from the timeframe field. Fifty-six guidelines have at least one Indeterminate impact dimension. Under the formula, every Indeterminate contributes zero, yet every timeframe contributes at least one.
That does not mean W3C claims unknown impact is zero. The Note’s prose defines Indeterminate as an impact that cannot be determined. The point is that the exported number cannot preserve that semantic distinction on its own. A consumer who sees only “3” cannot tell whether it means three low impact ratings plus a long horizon, three unknown impact ratings plus a short horizon, or another vector.
A correction shows why lineage must travel with the score
GitHub issue 12 supplies a narrow but useful operational example. On 29 July, a contributor reported that “Adopt organizational philanthropy practices” showed zero points for People, Planet and Prosperity and two for timeframe, but carried a total of five. The reporter said the rating-to-points mapping was correct for all 71 rows and that this was the only sum mismatch. The report also noted the same value in one downstream copy.
A maintainer confirmed the mismatch and corrected it to two that day. This is evidence of a functioning public correction path, not evidence of systematic failure. It also reveals why a score needs an edition and change record. HTML and JSON can agree with each other and still agree on the wrong value. A downstream copy can preserve a superseded total after the source is repaired.
The repository’s contributing guide is candid about the maintenance model. The HTML and JSON contain duplicated material and are not currently autogenerated from one source. If a contributor updates only one file, a chair or editor updates the others during approval to restore parity. The README says the JSON API is kept in sync with the specification.
Human review can be entirely appropriate. But a machine-readable surface invites reuse at machine speed. A consumer needs to know which dataset edition, formula and row revision produced a score, and whether a correction has superseded it. “Current” is not enough provenance for a number that may enter a report or procurement model.
Preserve the vector as the record
The safe architecture is not to abolish the total. It is to reverse the hierarchy. The durable record should be the impact vector; the number should be one named view calculated from it.
For every guideline, an impact-vector receipt should preserve the three ratings and timeframe separately. Each dimension should link to its rationale, cited research and suggested metric. “Indeterminate” should carry a reason code: evidence missing, evidence disputed, effect indirect, category not applicable or assessment not yet performed. Those states demand different next actions and should not collapse into the same zero.
The receipt should identify the editor or review group, decision date and call-for-consensus record. It should preserve objections, minority views and confidence statements where they exist. The formula version, weights and intended decision context belong beside the result, not in a distant methodological page. HTML, JSON and exports should always carry the raw vector and displayed total together.
Version fields should include the exact WSG edition, row revision, change history, correction notice and a machine-readable checksum. If a rating changes, the prior state should remain discoverable. If a formula changes, old totals should not silently acquire new meaning.
This structure follows a basic principle in Heng Lu’s agency analysis: a delegated institution earns trust by making the record of authority and accountability visible. The four ratings, their evidence and the process that accepted them are the accountable record. The sum is an interface choice. Treating the interface as the record reverses that relationship.
Different decisions need different views
One generic score may be adequate for browsing a long list. It is not automatically adequate for procurement, sustainability disclosure, product planning and research.
A purchasing authority might declare minimum constraints: no high-confidence harm to People may be offset by a short timeframe elsewhere. A climate-focused operator might weight Planet more heavily while requiring accessibility and labor safeguards as non-compensable floors. A reporting team might refuse to turn guideline ratings into claims about measured organizational outcomes at all. A researcher may need uncertainty distributions instead of integer points.
Those are policy choices. They should be explicit, named and reviewable. An adopter can publish its own weighting profile and compute a use-specific view from the same underlying vector. That approach allows disagreement without corrupting the common record.
It also preserves reversibility. If new evidence changes one rating, every use profile can be recomputed. If only the total survives, no reader can recover which dimension changed or test whether a different set of priorities would have produced another decision.
What the evidence does not establish
Nothing in this review shows that the 71 judgments are false, careless or captured. It does not reconstruct the full 6 August deliberation or claim the rename was a concession of error. “Ratings” may simply be the group’s preferred description after ordinary editorial discussion.
The arithmetic is not defective because it is simple, and equal weights are not invalid for every purpose. The question is whether a consumer can distinguish a convenient generic sorting rule from an authorized decision rule. The current Note provides rationales and a public dataset; those are useful foundations for the stronger receipt proposed here.
Issue 12 is one corrected mistake, explicitly described by its reporter as the only summation mismatch in the 71-row dataset reviewed at the time. It should not be inflated into a quality verdict. Nor does the mention of one downstream copy prove reliance, loss or harm.
The Group Note Draft can change. A later publication may add confidence, provenance, versioning or decision-specific profiles. The present opportunity is therefore constructive: preserve the simple score, but make it impossible for the score to travel without the judgments and authority that produced it.
Sources
- W3C News — Group Note Draft: Web Sustainability Guidelines Impact Ratings
- W3C — WSG Impact Ratings, Group Note Draft of 18 August 2026
- W3C Sustainable Web Interest Group — public Impact JSON
- W3C GitHub — pull request 16, Measurement > Ratings
- W3C GitHub — issue 12, corrected impact-score mismatch
- W3C GitHub — contributing guide
- W3C GitHub — repository README and JSON API
- W3C — Sustainable Web Interest Group Charter
- W3C — Sustainable Web Interest Group
- Heng Lu — On the Agency Problem at the Core of Internet Governance
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

