Summary

  • NANOG's published rules expose vendor presentations to useful scrutiny: vendors may submit, but employer promotion is constrained, confidentiality notices are prohibited, public slides are required, multi-vendor examples are preferred, and proposals pass through review and shepherding.
  • The NANOG 96 presentation IP Geofeeds: Trust, Accuracy, and Abuse disclosed substantial methods and denominators. Yet the page headed Accuracy reported 92.0% consistency, a distinction that later operator criticism made impossible to ignore.
  • The official mixed-subject list container reports 122 comments and 44 entities; exhaustive subject-line classification found 107 geofeed-subject replies from 35 displayed names. Neither count is a vote, a consensus measure or an experimental verdict.
  • A proportionate post-event record would link a material claim to its slide, method, denominator, critique, response, concession or correction, and artifact version—without requiring publication of customer data, GPS traces, device identifiers, proprietary code or private correspondence.

The noun beneath the number

Page 21 of the 39-page NANOG 96 slide deck carries a one-word heading: Accuracy. The country-level result below it uses a different noun. It reports 92.0% consistency, divided into 96.6% IPv4 and 79.6% IPv6. The city result is 79.6% consistency, with 88.3% IPv4 and 56.2% IPv6. The page also gives a median mismatch distance of 759 km.

That lexical gap is small enough to disappear in conversation and large enough to change the claim. Accuracy suggests comparison with a sufficiently independent account of where an address is used. Consistency can mean that two constructed observations agree, even if neither independently establishes location. The result is informative only while its method and noun travel with it. Once the qualifier is dropped and the figure is recast as an accuracy score, a measured relationship becomes a verdict the presentation did not independently establish.

The session was titled IP Geofeeds: Trust, Accuracy, and Abuse. Calvin Ardi of IPinfo presented it; the slides also name Oliver Gasser and William Leung. NANOG's agenda retains the abstract, video and slides. That archive made the later argument possible because entities could inspect the artifact rather than dispute a recollection. It also defines the limit of what agenda inclusion means. The session was selected and published. The record does not say that NANOG adopted its findings, that the Program Committee reproduced its measurements, or that the meeting presentation underwent scholarly peer review.

Important scrutiny arrived later on the public NANOG list. Entities argued over independent ground truth, the conversion of network delay into geographic distance, sample coverage and the route from a detected conflict to a correction. An IPinfo representative defended the work, welcomed an independent benchmark, acknowledged an outreach limitation and later said that contradictory wording on a correction form had been changed. These exchanges enlarged the public record. They did not settle every technical question.

The institutional issue begins there. A vendor's presence is not the defect, and a critical list message is not a verdict. The problem is that a permanent session page and a consequential post-session argument remain separate records. Someone arriving through the agenda sees the presentation, but no visible route from that item to the later challenge, response, concession and claimed correction.

Exposure is a better vendor rule than exclusion

An easy policy would bar speakers whose employers might gain from their claims. It would also deprive an operator forum of engineers who control large-scale instruments, implementation knowledge and production experience. NANOG chooses a more useful approach. Its current presentation guidelines welcome attendees and vendors to submit work in several formats. Employment by a supplier is not an automatic disqualification.

The rules instead constrain how commercial expertise enters the room. Presentation material may not carry confidentiality notices. Speakers are told not to use a talk to promote their employer, and employer logos are generally limited to the first and last slide. Multi-vendor configuration examples are encouraged. A single-vendor example is not forbidden, but it receives lower priority. NANOG says a PDF is posted on the event agenda on the day of the presentation. Its Code of Conduct separately prohibits aggressively pushing products or services.

Those rules cannot establish whether a method is sound. They do something prior and necessary: they keep the artifact public enough to quote, inspect and challenge. The organisation's presentation tips reinforce that operating emphasis by favouring case studies, anomalies, specific expertise and a proposed action for the audience.

The published review sequence adds an abstract, initial Program Committee consideration, an assigned shepherd, working slides, further slide review, selection, agenda publication and final materials. Submissions that include slides are generally favoured because reviewers can better see the intended presentation. The Program Committee describes itself as responsible for meeting programs, including recruiting and voting on submissions.

None of those stages carries the meaning of another. Selection determines that a session is worth putting before the community. Publication preserves what was said. A shepherd can improve scope, clarity and claim discipline. Scholarly peer review is a different form of scrutiny. Independent validation requires an outside test. Operational truth is harder still because it depends on performance across networks, time, failure modes and uses. Treating these as one ladder of increasing approval would give agenda acceptance an authority it does not possess.

The Program Committee page also lists “mutually rewarding agreements with sponsors and presenters” among its strategic goals. That wording makes a visible boundary valuable; it is not evidence that a sponsor or presenter influenced this session. The appropriate response is not suspicion attached to an employer name. It is a technical record in which commercial interest is visible and the supporting evidence remains contestable.

NANOG has debated that balance for years. In a 2008 list exchange, one entity criticised confidential or proprietary black-box content, missing alternatives and inadequate methodological detail. Others defended vendor work that was empirical and operationally useful, noted that methods may embody legitimate commercial value, and stressed the role of questions from the floor. A 2016 call for presentations again welcomed vendor submissions while rejecting promotional or proprietary talks. The recurring test is therefore substantive: what must a knowledgeable, interested actor expose so that operators can assess the claim?

What the slides exposed

The geofeed presentation supplied enough detail to put that test into practice. Its first approach used round-trip-time measurements from a probe network. The method constructed distance polygons from the measurements and compared their intersection with the location in a geofeed. Page 17 reports 27.7M round-trip-time measurements, 205.0k IPv4 prefixes and 74.5k IPv6 prefixes. The deck describes more than 1,300 probes in more than 500 cities, 143 countries and more than 400 autonomous systems.

These denominators matter. They distinguish a large measurement exercise from a collection of anecdotes, reveal separate IPv4 and IPv6 populations, and give critics something definite to examine. They do not answer every validity question. Measurement volume alone cannot establish that latency yields independent geographic truth. A wide probe footprint does not prove uniform coverage. A polygon inferred from network observation still depends on assumptions about routing, delay, propagation and probe placement.

The second approach used mobile devices equipped with GPS. The slides report 169 devices in 24 countries and 196 cities, covering 135 IPv4 and 97 IPv6 prefixes. For this method, the country result is 84.5% consistency; the city result is 29.9% consistency. The contrast is useful. The mobile sample differs in kind from the active-measurement set and covers far fewer prefixes. The country and city results also diverge sharply. A single headline score would conceal both facts.

The deck does not present the operational design as finished. Its closing pages propose an updated format and mechanisms for validation and feedback to operators. That is significant because it shows the presenters treating format, validation and feedback as unresolved surfaces. It remains a proposal, not evidence that the mechanisms were later implemented or that they solved the disputes.

RFC 8805 provides the technical setting. It defines a format for self-published IP geolocation feeds. Self-publication lets a network state where it understands a prefix to be used, but format conformance does not prove that a record is true, current or suitable for every downstream decision. At least three evidence layers therefore coexist: the network's assertion, an observer's method for evaluating it, and the decision a user makes from the result.

The NANOG 96 slides made the assertion and evaluation layers visible enough for an argument. They did not establish the quality of an entire commercial service, nor could one conference method settle every address range or operating condition. Their defensible contribution was narrower and still valuable: stated methods, denominators and results that operators could interrogate.

What the list added

The interrogation unfolded in an untidy place. The official list container reports 122 comments and 44 entities, but the conversation began with BGP communities and then changed subjects. Calling all 122 comments a geofeed debate would assign a precision the container does not have.

Exhaustive classification of the official reply subjects found 107 replies containing geofeed or geofeeds, posted under 35 distinct displayed names. That editorial count is useful for locating the visible branch and nothing more. A subject string does not prove that every message body analysed the NANOG 96 presentation. Displayed names are not a measure of expertise or unique human identity. Neither 107 nor 35 measures agreement, and neither can be converted into NANOG-wide consensus.

Tom Beecher made a direct methodological criticism. He argued that the active-measurement work showed agreement between probe-derived polygons and geofeeds rather than comparison with independent geographic truth. He challenged assumptions connecting network timing to straight-line or propagation distance, objected to reading the 92% figure as accuracy and raised coverage concerns.

The criticism goes to the relationship between the method and the headline. It remains criticism, not an adjudicated refutation. The reviewed public record contains competing arguments and operator examples, not an independent replication that resolves every assumption. Proper attribution is not a courtesy here; it is the line between preserving a challenge and pretending the challenge has already won.

An IPinfo representative pointed entities to the NANOG 96 talk and to a separate peer-reviewed paper, while saying an independent academic benchmark would be welcome. The separate paper's review status does not migrate to the conference presentation. Nor does the citation itself validate the exact result on page 21. IPinfo's later public explanation elaborated its own method and position, but it remains a vendor-controlled account rather than external corroboration.

The discussion then reached the correction mechanism. The representative said conflicts reported by an internet service provider are investigated and that systematic issues across several prefixes can prompt outreach. The same representative also said the company does not proactively contact every autonomous system about every individual-prefix mismatch when no complaint has been made, called that limitation real and proposed opt-in reports for high-confidence conflicts.

That is a meaningful concession about the described correction loop. Measurement at scale and notification at scale are different undertakings. The statement does not independently establish how every case is handled, whether the proposed reports were built or how many corrections reached a final outcome.

In another exchange, the representative described contradictory internal results associated with an operator and asked to continue privately. Private handling may have been appropriate, but it leaves the public outcome unknown. Later, after a entity cited correction-form wording that appeared inconsistent with claims in the thread, the representative accepted responsibility for the contradiction and said the form had been updated to describe a multi-source investigation rather than automatic preference for one source.

That statement belongs in the record as a claimed correction. Its limits belong beside it. The earlier form was not reconstructed, and the full history of the change was not independently verified. A public correction entry point existed during the research review; its existence reveals neither response times nor acceptance rates nor whether an eventual correction is accurate.

The list performed a function the stage could not. It surfaced assumptions, counterexamples and incentives, elicited a concession, and produced a public statement about changed wording. Its weakness is structural. The exchange grew inside a mixed-subject chain, sprawled over many messages and moved some investigation into private correspondence. A reader starting from the agenda must already know the later discussion exists.

Mailing lists organise conversation by reply and subject, not by the claims under examination. One message may challenge the measurement model, another may describe an affected prefix, and a third may revise the description of a correction practice. Chronology preserves the exchange but does not tell a later reader which statement answers which slide. Subject changes complicate discovery further: the container total overstates the geofeed branch, while a subject-only filter still cannot distinguish methodological analysis from a passing reference.

The list is therefore an essential evidence surface and an inefficient index to its own evidence. Linking it from the session record would preserve the openness of the discussion while reducing the chance that only specialists who followed the exchange in real time can reconstruct it.

Five kinds of authority

This case becomes easier to read when five different events keep their own names.

Selection means that a Program Committee judged a session suitable for the program under its published process. Publication means that the audience can inspect the artifact. Peer review describes a scholarly procedure that applies here to the separate paper cited in the list, not automatically to the NANOG presentation. Independent validation requires an outside test of the relevant method and claim. Operational truth concerns performance in the environments where people rely on the result.

One event can support another without substituting for it. A useful conference observation may never enter a journal. A reviewed method may behave differently as networks change. A mailing-list counterexample may expose a serious failure mode without estimating its prevalence. A vendor may concede a limitation without abandoning the larger analysis. The public record becomes trustworthy when it preserves these differences instead of allowing each to borrow authority from the next.

Geoff Huston's APNIC review of NANOG 96 adds an outside perspective. It identifies both geofeed sessions and describes IP-based geolocation as coarse, gameable and often hard to verify. That commentary places the dispute in a difficult operational field; it is not a benchmark of IPinfo's service.

Operator counterexamples require the same care. One can reveal that an aggregate score masks a consequential case or that a correction route is hard to use. It does not necessarily measure how common the failure is. The aggregate result, in turn, does not answer the operator whose prefix is wrong. Evidence governance must make both scales legible without forcing either into the role of universal verdict.

NANOG does not need to decide which correspondent won. It can ensure that readers can discover the disagreement, see what each side claimed, and learn whether the public artifact changed. That is an institutional task, not a product judgment.

The legitimate limits of disclosure

Complete reproducibility is an attractive demand until it reaches a production system. Vendor research may depend on licensed sources, customer-supplied records, protected infrastructure or methods built through costly experimentation. Mobile validation can implicate location traces and device identifiers. Operator investigations may expose address use, internal contacts or security-sensitive context. Unrestricted release can invade privacy, make systems easier to game, damage customer trust and destroy legitimate commercial value.

The resulting loss would extend beyond suppliers. If presenting at NANOG required releasing every observation and every line of proprietary code, some engineers with the best access to large-scale operational evidence would stay away. Operators would receive fewer measurements, fewer implementation lessons and fewer opportunities to question people responsible for the systems. A NOG meeting is neither a court nor an academic journal. Its strength is rapid exchange among people who operate networks.

NANOG's present arrangements already answer part of this objection. Public slides and the ban on confidentiality notices create an inspectable artifact. Limits on promotion separate a contribution from a sales pitch. Shepherding provides a pre-event review relationship. The open list provides a place for challenge. Private follow-up can protect details that should not be broadcast.

The proportionate standard is bounded disclosure. No claim-level record should compel publication of customer records, raw GPS traces, device identifiers, proprietary source code, private operator correspondence or exploit-relevant infrastructure. Nor should a critic have to expose a customer to prove that a failure occurred. The public entry can identify the class of withheld evidence, give the reason for sensitivity, and say what kind of independent or trusted test would be possible under controlled conditions.

Methods can also be described at several levels. A presenter may disclose measurement class, important assumptions, denominators, exclusions, uncertainty and plausible alternatives without releasing a complete implementation. A correction note can identify what claim or wording changed, and when, without exposing the case that prompted it. Private work can remain private while the public status says unresolved, confirmed, not reproduced or wording corrected.

Confidentiality and contestability are therefore not opposites. The governance obligation is to make their boundary visible.

The smallest useful addendum

The remedy need not become a second proceedings system. For a material post-event challenge, NANOG could place a concise claim-level addendum beneath the existing video and slide links. Its purpose would be discovery and attribution, not adjudication.

For this case, the first entry would identify page 21: heading Accuracy; country bullet 92.0% consistency; IPv4 and IPv6 breakdown. It would name the active-measurement method and denominators—round-trip-time polygons, 27.7M measurements, 205.0k IPv4 prefixes and 74.5k IPv6 prefixes—and separately identify the mobile-GPS method and its much smaller device and prefix counts.

A challenge field would link Beecher's criticism and describe it as his: questions about independent ground truth, geometry, coverage and the conversion of consistency into accuracy. A response field would point to the representative's defence, the separate research reference and the invitation for an independent benchmark. A limitation field would preserve the statement that individual-prefix mismatches do not all trigger proactive outreach. A correction field would record the representative's claim that contradictory form wording was updated.

The record would then show artifact status. Did a slide change? Was explanatory material added? Was public wording revised? Was a proposed conflict report implemented? Did an operator case continue privately with no publishable outcome? Unknown is an honest status. The function of the addendum is not to manufacture closure but to prevent separation by time and interface from looking like settlement.

Versioning is essential. A link to a live form cannot establish what that form said during the dispute. A dated note—presenter states that wording changed; earlier public version not preserved here—would say precisely what the record supports. When NANOG holds no artifact, it can state that absence rather than imply verification.

The threshold should remain high enough to avoid annotating routine disagreement. A central numerical claim, a consequential limitation, a presenter concession, a material erratum or a stated change to a public artifact would qualify. Stylistic objections and product preferences would not. The Program Committee or a small records function could curate links and labels without deciding the underlying science.

The original presentation should remain untouched. An addendum is most useful when it preserves the historical claim and places later evidence beside it, rather than silently replacing a slide or retroactively polishing its wording. Each entry could carry a date, a stable link, the name and role of the person making the statement, and a short evidence label such as critique, presenter response, stated limitation, claimed correction or independent replication. The label would describe the artifact's function, not certify its truth. If a link later disappears, the record could retain its citation and mark availability as changed. That modest discipline would let a reader reconstruct sequence without converting NANOG into an arbiter of every technical disagreement.

Neutral structure protects presenters as well as critics. A later reader who encounters an accusation elsewhere should also be able to find the response. Putting challenge, defence, concession and correction in one map reduces both uncritical testimonial and decontextualised blame.

No claim-level mapping was found on the public pages reviewed for this case. The captured NANOG 96 agenda links the session, video and slides, but it does not visibly link the later geofeed branch, Beecher's criticism, the representative's outreach limitation or the claimed form correction. That is a bounded observation about the reviewed record. It implies no intent and does not establish that NANOG lacks a private or internal practice.

Responsibility before and after the room

Post-event traceability should not become an impossible pre-event warranty. The public record does not disclose the Program Committee's proposal-level reasoning for this session, its exact reviewer questions, the changes made during shepherding or whether reviewers raised ground truth, geometry, coverage or wording. There is no responsible basis for inventing that history.

A Program Committee should not be expected to reproduce every measurement before scheduling a talk. Contested work can be valuable precisely because operator scrutiny improves it. If selection becomes a guarantee of correctness, committees will favour safe summaries over ambitious evidence. If every later dispute is labelled a selection failure, presenters gain an incentive to conceal uncertainty.

A clearer division of labour follows the life of the claim. Before the event, reviewers can test relevance, disclosure, method visibility, denominators, limitations and the distinction between employer value and audience value. During the event, operators can press assumptions and offer counterexamples. After it, the permanent record can connect material challenges, responses and artifact changes to the session. Independent researchers remain responsible for replication, and operators remain responsible for decisions in their own environments.

This sequence also prevents accountability from becoming an excuse for delay. NANOG would not have to hold a useful session until every outside test was complete. Publication would remain prompt, while later evidence could improve the record without erasing the original presentation. A presenter could answer a criticism without conceding the whole result; a critic could document a failure mode without claiming prevalence; and the agenda could show both positions without issuing a technical ruling. The forum's contribution would be continuity: keeping a public claim reachable as its evidentiary status changes.

That division gives shepherding a demanding but realistic role. A shepherd need not certify truth to ask whether a slide headed Accuracy reports a measurement called consistency; whether a sample denominator is visible; whether IPv4 and IPv6 results are separated; or whether a single-vendor result acknowledges limits and alternatives. These are questions of claim discipline. They make later scrutiny possible without pretending to complete it.

The 2008 debate already contained the central tradeoff. Enough method must be available to test a claim, yet some methods carry legitimate commercial sensitivity. Questions from the floor can expose weakness, yet they may leave no lasting map. The NANOG 96 case adds a temporal lesson: the consequential question may arrive weeks later and in a different medium.

The official list's scale demonstrates both the strength and the cost of open operator scrutiny. A long mixed container, dozens of displayed names and a substantial geofeed-subject branch show sustained attention. They do not prove consensus, representativeness or experimental validity. They show that relevant challenge-and-response material exists—and that reconstructing it requires work most agenda readers will never know to perform.

A linked record would convert that work into operating memory. It would keep Beecher's criticism distinct from an adjudicated result; a representative's concession distinct from observed system behaviour; a stated wording correction distinct from reconstructed version history; and a private continuation distinct from a public outcome. Those distinctions are not editorial decoration. They are the substance of a technical record that can survive disagreement.

The remaining unknowns should stay visible. The reviewed material contains no independent benchmark of the full commercial service. It does not reveal correction volumes or response times, whether the proposed opt-in report was built, what happened in private operator follow-ups, or whether the presentation's method was independently reproduced. These gaps delimit the case; they are not invitations to speculate.

That restraint is the durable vendor test. Employer identity remains visible but is not treated as guilt. Presenters receive credit for disclosing denominators, answering criticism, conceding limits and correcting public wording. Critics receive a lasting route to their objections without becoming final arbiters. NANOG's rules already provide a strong opening through public artifacts, restrained promotion, review and shepherding. Its list provides a serious challenge surface. What is missing from the reviewed public record is the join between them.

A talk archive tells readers what reached the stage. An evidence record also tells them what the claim measured, who challenged it, how the presenter answered, what limitation was conceded, whether an artifact changed and what remains unknown. The NANOG 96 geofeed dispute names no product winner. It shows why the evidentiary life of a talk cannot end with its acceptance.

Sources