Summary

  • The 2025 IETF Community Survey contacted 52,660 mailing-list addresses, received 1,716 responses and retained 1,480 valid responses. The resulting 2.81% response rate does not by itself prove serious bias, but it makes cohort-specific validation important.
  • The report projects 4,460–7,136 regular posters from respondent self-classification and compares that range with 4,457 addresses that sent mail to IETF lists in 2025. The near match supports a bounded estimate of active posting addresses; it does not show that respondents and nonrespondents share the same views or validate every other participation cohort.
  • The report deserves credit for disclosing its frame, exclusions, missing-data choices, instrument changes and AI-assisted treatment of open comments. Its ±2.54% figure remains a sampling-precision quantity, not a measure of coverage, self-selection, wording, nonresponse or coding uncertainty.
  • Survey results can inform IETF LLC and IESG priorities. They are not a vote, a rough-consensus call or authority borrowed from people who did not answer. A cohort-by-claim evidence receipt would preserve each useful finding at its true scale.

The contradiction is only resolved by narrowing the claim

The final survey report sets out a clear ambition. It seeks a demographic picture of the IETF, evidence for leadership about important issues, and a time series against which programs and institutional changes can be assessed. That is a legitimate administrative instrument. The IETF LLC Community Engagement Policy expressly treats surveys as inputs to decision making and expects their results to be published while respondent anonymity is protected.

The report is also candid about the hardest limitation. Because people chose whether to respond, selection bias is possible. It says the information required for a full cross-check was unavailable, so that possibility remains. This is the correct starting point: invitation breadth is not the same as response probability, and response probability is not necessarily independent of the attitudes being measured.

Five pages later, the language becomes stronger. Regular posters made up 11.01% of respondents. Applied to the 52,660-address frame and its stated sampling calculation, that produces an estimated range of 4,460 to 7,136. A separate analysis found 4,457 addresses that had sent messages to IETF mailing lists during 2025. Because the external count sits almost exactly at the projected lower bound, the report concludes that regular participants are reasonably represented and describes selection bias as negligible.

There is no need to choose one passage and discard the other. The external comparison is real evidence. It tells the reader that the survey did not produce an obviously impossible active-poster share. It may support a cautious estimate of the number of sending addresses classified as regular posters. It does not test the whole survey. “Negligible” becomes defensible only if its object is written beside it: negligible discrepancy for this one external count, under these address-level definitions. Detached from that object, the word travels farther than the test.

Addresses are the frame, not verified people

The unit matters before any opinion is interpreted. The survey distribution process combined memberships of active IETF mailing lists, removed duplicates using plus-address notation, excluded prior opt-outs and known undeliverable contacts, and sent invitations to 52,660 addresses. That is broad and operationally sensible. It is not a census of 52,660 unique human participants.

The report says why. One person can subscribe to different lists with different addresses. A role address may represent a function rather than an individual. An internal expander can deliver one subscription to several people. These issues were not corrected because the necessary data were not available. The external benchmark is also expressed in sending addresses, not verified people. Matching one address estimate to another avoids some unit confusion, but it does not make the underlying population humanly unique.

This distinction is not pedantry. A person who posts from two addresses may appear twice in the sender count. A person who reads through an internal distribution list may be absent as an individual from the invitation frame. Silent subscribers may include automated accounts, archives, former participants and people who follow only one narrow topic. The frame is best described exactly as the report constructs it: contactable mailing-list addresses after specified cleaning.

That frame can support valuable statements. It can show what respondents linked to that frame said. It can estimate address-level patterns under an explicit model. It cannot silently become “the IETF community” each time a percentage enters a governance sentence. The transformation from addresses to people, people to participants and participants to an institution’s principal requires separate evidence at every step.

One calibration does not test attitude bias

Suppose the 4,457 external count perfectly captured every unique regular poster and the respondent classification was flawless. The match would still answer only a composition question: does the survey’s estimated proportion of regular posters resemble an external total? It would not answer the more important inferential question: are regular posters who replied similar to regular posters who did not reply on quality, openness, corporate influence, leadership, barriers or process speed?

Those two forms of representation are different. A sample can contain the right share of a cohort while attracting a distinctive subset within it. The most satisfied regular posters may answer because they value the institution. The most frustrated may answer because they want change. Busy participants may ignore the invitation. People who distrust surveys may be systematically absent. A correct cohort count cannot reveal which of these response mechanisms operated.

Nor does the regular-poster check validate other cohorts. Occasional posters, readers, monitors, former participants and newcomers experience different costs and have different reasons to respond. The report’s own results show that perceptions vary with participation intensity and leadership experience. That variation is precisely why one active cohort cannot serve as a universal calibration weight.

The match is therefore a good canary, not a clearance certificate. A failed match would warn that the sample composition is problematic. A passed match removes that one warning. It does not prove that every untested path is sound. Governance fails when the absence of one detected defect is rewritten as the absence of the class of defect.

A response rate is a diagnostic, not a verdict

The 2.81% response rate must be handled with equal discipline. A low rate can coexist with accurate estimates if response propensity is unrelated to the measured outcomes or if a credible adjustment model accounts for the difference. A high rate can still produce bias if the remaining nonrespondents differ systematically on the question that matters. The rate alone neither convicts nor acquits the survey.

The AAPOR Standard Definitions make this distinction useful: an outcome rate is a critical diagnostic for possible nonresponse error, but it does not by itself establish whether that error exists or how large it is. The correct question is not “Is 2.81% too low?” It is “What evidence connects response propensity to each claim we want to make?”

For the IETF survey, that evidence could include comparisons between respondents and known frame attributes, response waves, early and late respondents, meeting registration or Datatracker indicators under privacy controls, and stable external counts for several participation categories. None needs to identify an individual publicly. Aggregate checks would be enough to show where the sample aligns, where it does not, and where no benchmark exists.

This would also improve the interpretation of people who never post. The survey projects large populations of readers and monitors, but list traffic cannot externally validate silent activity as easily as it validates sending. The absence of a comparable trace should widen the uncertainty label, not make the cohort disappear. Quiet participation is socially real even when it leaves fewer observable events.

The ±2.54% label covers less than it appears to cover

The report describes 1,480 valid responses as yielding a maximum margin of error of ±2.54%. The arithmetic resembles the familiar simple-random-sample precision calculation for a proportion near 50%, adjusted only slightly for a finite population of 52,660. Readers recognize the label and may treat it as a compact uncertainty warranty.

It is not a warranty for total survey error. It does not include the possibility that subscribers outside the frame differ from those inside it, that volunteers differ from nonrespondents, that one address is not one person, that wording shifts an answer, that optional questions have different denominators, that exclusions alter a subgroup, or that an AI-assisted thematic summary groups comments differently from another coding process.

The AAPOR Transparency Initiative draws a practical line. For a non-probability or volunteer sample, a precision measure should be accompanied by the model, assumptions, validation and calculation that give it meaning; otherwise the report should state the non-probability basis and explain the limit on inference. The IETF need not adopt AAPOR as law to benefit from the distinction.

A precise label would say what the number does and does not describe. For example: conventional sampling precision conditional on a simple-random-response model; excludes coverage, voluntary-response, measurement and coding effects. That sentence does not make the report weaker. It prevents a mathematically tidy number from absorbing uncertainty it was never built to measure.

Open comments show themes, not prevalence

Question 54 illustrates a better instinct. Only 241 of the 1,480 valid respondents supplied open-ended comments. The report states that these comments come from a smaller cohort and should be read as actionable focus areas rather than a broad definition of systemic conditions. That is the right boundary.

The comments raised behavioural barriers, corporate capture, process delays, access and global representation, and leadership accountability. These are legitimate warnings. Their frequency within a self-selected open-comment subset cannot tell leadership how prevalent each concern is across the contact frame. Nor should a positive mean on a structured question be used automatically to neutralize the substance of a minority warning. A low-prevalence problem can still be severe, and a general satisfaction item may measure something different from the incident described in free text.

The report discloses that Claude Sonnet 4.6 generated the thematic summary. Naming the tool and task is better than hiding automation. The next reproducibility step would record the prompt family, codebook or theme definitions, human validation procedure, treatment of multi-theme comments, and stability check across a second coding pass. Verbatim comments can remain private. What matters publicly is the route from protected testimony to aggregate category.

AI assistance is not the central defect here. The governance issue is the same for human and machine coders: a summary must carry its method and denominator, and a theme must not impersonate a prevalence estimate.

A time series needs instrument versions

Annual repetition is one of the survey’s strongest features. A stable series can distinguish a persistent concern from a loud week. But continuity is not created by placing years beside one another. It depends on whether the population, question, response scale, coding and exclusions remain comparable.

The 2025 report documents several changes. Questions 25 and 26 shifted from frequency wording to agreement wording and changed their scales. Other questions were rephrased, removed or consolidated. These may be sensible improvements, especially after review with speakers of English as an additional language. They also create a measurement boundary.

Any trend chart crossing that boundary should carry an instrument-version marker. A difference after rewording may reflect a changed perception, a changed question or both. The 2024 publication notice also records corrected charts after the initial report. That correction is evidence of a functioning revision process, provided later analysis cites the corrected version and preserves the change log.

The durable object is therefore not a floating percentage. It is a versioned observation: question text, scale, eligible cohort, denominator, missing-answer treatment, field dates, processing rules, report revision and permitted comparison. Without those joins, a five-year series can look more continuous than its measurement actually is.

Survey evidence is not IETF consensus

The governance boundary becomes sharper when the result is used. The survey report says the IESG and IETF LLC will reference the findings in planning, directing attention toward evidence-backed priorities and away from matters the evidence suggests are not concerns. The first half is sound. Survey results can select questions for investigation, identify service problems, test whether interventions changed reported experience and reveal groups leadership rarely hears.

The second half needs caution. A self-selected survey can show that a concern was uncommon among respondents to a particular item. It cannot prove that the concern is not real, not severe or not present among people least likely to answer. Evidence that lowers a hypothesis is not authority to close it.

The distinction is institutional as well as statistical. RFC 8711 gives the IETF LLC administrative support responsibilities and denies it authority over standards development. RFC 7282 explains why rough consensus is not a vote or a count of raised hands: the substance of unresolved objections matters. A community survey can inform the people operating these processes. It cannot replace the process or grant an administrative result technical authority.

That limit follows the deeper principle developed in Lu Heng’s critique of attendance becoming mandate: participation is evidence, not authorization. The 1,480 respondents deserve to be heard in their own capacity. They should not be conscripted as representatives of 52,660 addresses, still less of everyone affected by Internet standards. Nonrespondents should not be imagined to agree, disagree or abstain. They are absent from the evidence, not present as a silent principal.

A cohort-by-claim evidence receipt

Daniel Kade’s proposal is a thin public concordance attached to consequential survey claims. It is analysis, not an existing IETF or AAPOR rule. It would preserve the survey’s usefulness while stopping one validation from travelling into unrelated conclusions.

For every headline finding, the receipt should name the target quantity and unit: respondent, response, address, inferred person, active sender, meeting participant or comment. It should show the invitation frame, eligible cohort, question-level denominator, missing-answer rule, response mode, field dates, instrument version, weighting or absence of weighting, and report revision.

It should then list validation separately. Which external benchmark was used? Does it share the same unit and period? Which cohort and variable does it test? What discrepancy was observed? What remains untested? A check against active sending addresses would be bound to active-poster composition, not inherited by leadership sentiment or reader experience.

For inference, the receipt should state whether the result is descriptive of respondents, modelled to a frame or used only as a hypothesis trigger. Any precision figure should name its model and excluded error components. AI-assisted qualitative coding should identify the model, task, coding instructions, human review and stability test without releasing private text.

Finally, the receipt should record the authority destination. An LLC service decision, an IESG management inquiry, a working-group consensus call and a future survey redesign are different downstream acts. The survey can enter all four as evidence, but each act must obtain legitimacy from its own process. The receipt makes that handoff visible.

Keep the instrument; reduce the borrowed certainty

The IETF should continue the annual survey. Few technical institutions publish this much methodological detail, expose awkward findings and preserve raw question counts in one report. The constructive response to an overbroad sentence is not to discard the evidence. It is to make the sentence as precise as the evidence.

The 4,457-address comparison is valuable because it demonstrates the form of validation the report needs more of. Add benchmarks for other observable cohorts. Separate address counts from people estimates. Label the model behind precision. Preserve question versions. Treat open comments as protected signals. Record which decision used which finding and what additional evidence was sought.

Then the central claim becomes both smaller and stronger: the 2025 survey reached a broad mailing-list frame; 1,480 valid respondents supplied a rich account; one active-poster estimate aligns with one external address count; many other inferential paths remain untested. That is enough to inform leadership. It is not enough to manufacture a community mandate—and it does not need to be.

Evidence limits

This analysis uses the published report, official IETF policy and RFCs, prior-year publication history and AAPOR professional guidance. It does not have respondent-level data, invitation-delivery logs, the separate sender-analysis dataset, individual comments or the AI-coding prompt. It does not calculate alternative estimates, determine the direction or size of nonresponse bias, or allege that any named IETF decision was improperly made. It does not dispute individual demographic or perception findings. The cohort-by-claim receipt is a governance recommendation, not a current institutional requirement.

Sources

  1. IETF Community Survey 2025 — final report
  2. IETF Community Survey 2025 — publication post
  3. IETF reports and surveys
  4. IETF LLC Community Engagement Policy
  5. RFC 8711 — Structure of the IETF Administrative Support Activity 2.0
  6. RFC 3935 — A Mission Statement for the IETF
  7. RFC 7282 — On Consensus and Humming in the IETF
  8. AAPOR Transparency Initiative
  9. AAPOR Standard Definitions
  10. IETF Community Survey 2024 — publication and correction notice
  11. Launch of the IETF Community Survey 2025