Summary

  • ARIN’s ACSP 2023.1 records a publisher using the XML API for AS13335:AS-CUSTOMERS; the resulting RPSL placed many members on one physical members: line.
  • The submitter reproduced a recursive query against a mirror he believed to run IRRd v3. It returned only seven ASNs. ARIN agreed that limiting members per line would improve AS-SET query results and kept the suggestion open pending implementation.
  • Logical membership, RPSL attribute semantics, serialized bytes, mirror import, parsed object state, recursive query output and a generated route filter are separate artifacts. A valid first artifact does not prove the rest are complete.
  • A cross-mirror serialization receipt should bind a canonical member count and digest to exact exported bytes, line limits, import identity, parsed count and query-result digest, with an explicit truncation or rejection state.

Seven ASNs at the end of a much longer chain

The most revealing part of ARIN ACSP Suggestion 2023.1 is not the size of the object. It is the mismatch between two answers that appeared to describe the same routing-policy chain.

Joe Abley wrote that Cloudflare published AS13335:AS-CUSTOMERS automatically by sending XML payloads through ARIN’s API. The RPSL displayed in the suggestion contained a members: attribute beginning with a long sequence of AS numbers and continuing through what the page abbreviates as many more. ARIN’s serializer had placed that sequence on a single physical line.

The submitter then showed the complete response from a recursive query to a separate mirror, identified cautiously as an instance he thought ran IRRd v3. The query !iAS-CLOUDFLARE,1 returned seven ASNs and the protocol terminator. The page called the answer profoundly wrong. It did not publish a packet capture, parser trace or mirror database dump, so the exact failure component remains unproven. What it did preserve was enough to locate the class of problem: a large valid object left its authoritative publication path and a downstream resolver exposed only a small derived result.

ARIN’s response was similarly bounded. On 7 February 2023 it said that limiting the number of items on single members lines would produce better results for AS-SET queries. It would investigate the requirements, schedule the feature for future development and leave the suggestion open until implementation. The page remains marked Open today. That status is evidence about the public ticket, not proof that ARIN’s present output is unchanged or that no operational mitigation has occurred.

This article does not re-test the historical mirror, which may no longer be available, and it does not assert anything about Cloudflare’s current set, route filters or BGP. The durable lesson is the boundary the record exposes.

One set has several representations

An AS-SET is first a logical claim: a named collection whose explicit members may be AS numbers or other AS-SET names. RFC 2622 defines members as optional and multi-valued. It also makes an important distinction between a list-valued attribute and an attribute that occurs more than once. Either dimension can exist without the other.

The text representation adds another dimension. Each attribute-value pair begins on a physical line. A value may continue on following lines when those lines begin with a space, tab or plus sign. Multiple members: attributes are also valid because the attribute is multi-valued. These choices can preserve the same logical membership while producing very different byte sequences, line counts and maximum line lengths.

That difference is normally dismissed as formatting. ACSP 2023.1 shows why it can become an interoperability contract. A modern parser may accept a line measured in many kilobytes. An older input reader may have a smaller fixed buffer, a field-length assumption or a recovery rule that preserves the beginning while losing the remainder. A transport can deliver every byte while an importer rejects one attribute. A parser can store a partial value while still exposing an object whose key and surrounding fields look normal.

The public record does not tell us which of those mechanisms applied to the example mirror. Naming one would turn a bounded interoperability report into an unsupported diagnosis. The engineering consequence does not depend on choosing: physical serialization is a state worth measuring whenever independently operated systems exchange RPSL.

The serializer is part of the operational interface

ARIN’s current IRR overview says that a simple object submitted as XML is converted in the back end to RPSL. The record other users retrieve remains RPSL. The RESTful API guide exposes both XML and RPSL representations for supported operations and describes members as the list of ASNs or AS-SET objects belonging to the set.

That architecture creates two correct but different responsibilities. The publisher controls the member collection in its request. ARIN controls how an XML representation becomes physical RPSL. In the historical suggestion, the XML schema did not let the publisher select the number of members per output line; the page says ARIN’s construction code made that choice.

This is why “the input was valid” is an incomplete defence. The consumer does not read the publisher’s in-memory collection. It reads bytes emitted by a serializer, perhaps through a downloaded database, NRTM updates or Whois. Compatibility depends on the output profile as well as the abstract data model.

Shorter lines are a sensible mitigation because they remove one source of stress from legacy readers. RFC 2622 gives more than one legal way to do it: repeat the multi-valued attribute or continue a value across physical lines. But the mitigation should be tested as a byte-level change, not inferred from a pretty-printed screen. A serializer can wrap a display while an export still emits the long form. A mirror can accept the new lines but apply a separate member-count limit. A recursive query can be incomplete for source-selection or recursion reasons even when the object itself is intact.

A mirror is a new custody boundary

ARIN’s current documentation describes publication through NRTM, downloadable database files and Whois, and says its present server is IRRd Version 4. Current IRRd mirroring documentation describes full snapshots and incremental updates as data that a mirror retrieves, parses, validates and writes into its own database. A mirror is not a transparent pane. It is a separately operated copy with software versions, configuration, source policy, import history and error handling of its own.

The distinction matters even when transport is perfect. The source can publish an exact byte stream. The mirror can receive that stream. Yet the stored object can differ if validation rejects an attribute, parsing stops early or an earlier object remains after an import failure. Conversely, a mirror may store the full object while a query path later applies a depth, source or output limit.

Modern IRRd documentation says full reloads and updates are transactional, so users should not see a half-finished import. That is valuable, but atomicity is not the same as semantic completeness. An atomically committed partial parse would still be partial. An old mirror’s behavior also cannot be inferred from current IRRd documentation. The historical example must remain historical.

A complete custody chain therefore needs at least three comparisons. Did the mirror receive the exact exported bytes? Did it store a parsed object with the same normalized members? Did the query engine return the result expected from that stored state under explicit recursion and source rules? Skipping the middle comparison leaves operators unable to tell whether a bad answer began in transport, storage or resolution.

Recursive output is not the source object

IRRd’s Whois-query documentation describes !i as a processed query. With the recursive flag, it follows nested sets and returns resolved members as a space-separated result. That output is not raw RPSL. It is a computation over the mirror’s local object graph.

The difference is operationally important. A raw-object query can show whether the local copy contains every members value. A recursive query can show what the resolver derives after looking up nested sets. A filter generator may then convert that derived AS list into prefix queries and router policy. Each step can be internally valid while inheriting an omission from the previous one.

The final router configuration is farther still from the source. Operators may combine several IRR sources, prefer one source over another, exclude sets, cap recursion or merge RPKI and private policy. Nothing in ACSP 2023.1 proves that a particular router installed a partial filter. The example establishes that a partial query answer existed in the submitted evidence and that ARIN regarded line limiting as capable of improving query results.

That is enough to justify a pre-filter completeness check. A filter builder should not treat a syntactically successful AS-SET expansion as proof that the expansion is whole.

A receipt that compares meaning with bytes

ARIN and mirror operators could make this boundary observable with a cross-mirror serialization receipt. This is an editorial recommendation, not a published ARIN or IRRd commitment.

At the source, the receipt would record the AS-SET key, object version or publication serial, normalized member count and a digest over a documented canonical ordering. It would then identify the serializer and record the exact exported RPSL byte digest, total bytes, physical line count and maximum line length.

At the mirror, a companion record would identify the source, snapshot or NRTM serial range, import time, parser version and import outcome. It would compare the received-byte digest with the source receipt, then record the stored object digest, parsed member count and normalized member digest. Failure must be typed: rejected, truncated, partially parsed, stale after failed import, or complete. Silence is not an acceptable truncation signal.

At query time, the operator would add the query string, selected sources, recursion flag or depth, excluded sets, execution timestamp, result count and result digest. A downstream filter generator could optionally bind its own artifact digest to that query receipt without publishing confidential router policy.

The receipt does not require every operator to use the same software or routing policy. It lets them disagree visibly. If the source count is 10,000, the mirror parsed count is 10,000 and the recursive result differs because of nested-set rules, attention moves to resolution. If the received bytes match but the parsed count is smaller, attention moves to the importer. If the byte digests differ, transport or export custody becomes the first question.

The result is a narrow accountability mechanism: not a central routing authority, and not a claim that an IRR record proves live BGP, but a way to stop a quiet representation loss from masquerading as a complete answer.

Sources