Summary

  • APNIC’s captured 8 September 14:00–15:00 UTC snapshot reports 3,192,583 WHOIS queries, but its displayed query types sum to 3,183,463: a positive gap of 9,120. The same file’s RDAP total reconciles exactly.
  • In all 1,690 archived hours containing WHOIS traffic between the launch on 30 June and the evidence cutoff, total_queries differs from the type sum. The gap is positive in 1,011 files and negative in 679; an omitted non-negative “unknown” category cannot explain both signs.
  • APNIC’s posted APNIC 62 implementation deck says prop-167 remains in implementation because some data-accuracy issues are being resolved. It does not identify this arithmetic. A privacy-safe reconciliation ledger should name counting stages, excluded classifications, validation results and corrections without exposing raw queries.

The last equation in a public statistics file should not require a theory of the institution that produced it. APNIC’s latest snapshot at the evidence cutoff covers the hour from 14:00 to 15:00 UTC on 8 September 2026. The WHOIS section gives total_queries as 3,192,583. Its visible distribution contains 1,670,618 inetnum, 1,284,709 route, 93,537 aut-num, 88,530 domain, 31,500 organisation, 13,888 as-set and 721 inet6num queries. Those seven integers add to 3,183,463.

The difference is 9,120 queries, or 0.2857% of the published total. The adjacent RDAP section provides the useful control. Its five displayed query classes add to 588,346, precisely the value in its total_queries field. The snapshot’s MD5 sidecar also matches the downloaded JSON. The file arrived intact; one service’s arithmetic reconciles and the other’s does not.

This is a small discrepancy in one busy hour. It becomes an institutional question because the hour is not exceptional.

Every non-zero WHOIS hour has the gap

APNIC started the archive on 30 June, the date its prop-167 page records as “Implementation Complete”. I enumerated the monthly indexes and checked every available compressed hourly file through the archive ending at 14:00 UTC on 8 September. There were 1,691 expected intervals and 1,691 files. The filenames were continuous. Every gzip stream decoded, every document parsed as JSON, and every embedded half-open time range matched its filename. No two uncompressed files had the same fingerprint.

For each service and hour, I then summed the values in query_type_distribution and subtracted that sum from total_queries. I performed the same operation on every displayed ASN row: query_count minus the sum of query_count_by_type.

RDAP reconciled in all 1,691 files, at both levels. Its whole-service signed gap was always zero. None of its displayed ASN rows had a type sum different from the row’s query count.

WHOIS reconciled in one file. That exception was 27 August from 09:00 to 10:00 UTC, when the WHOIS section contained zero total queries, zero ASNs, an empty distribution and no rows. In other words, each of the 1,690 archived files containing WHOIS traffic had two different whole-service totals. Those same 1,690 files contained at least one displayed ASN row whose query_count did not equal its per-type sum; across the archive, 347,490 displayed rows had that property.

The first archived hour already contained the pattern. From 03:00 to 04:00 UTC on 30 June, WHOIS reported 4,774,145 total queries while its types summed to 4,801,599. The typed number was 27,454 higher. In the last archived hour, from 13:00 to 14:00 UTC on 8 September, the reported total was 2,989,539 and the type sum 2,987,334. This time the total was 2,205 higher.

The sign reversal matters more than the accumulated difference. In 1,011 files, total_queries was larger than the type sum. A plausible explanation could be requests that were accepted by the service but could not be assigned to one of the named object types. Add an unknown or other bucket, and those hours might close.

But in 679 files the direction runs the other way: the named type counts add to more than the declared total. On 14 July from 05:00 to 06:00 UTC, the file reports 3,288,941 WHOIS queries and a type sum of 3,652,538. The sum exceeds the total by 363,597, or 11.0551%. A missing non-negative category cannot subtract that amount. On 20 August from 03:00 to 04:00 UTC, the total exceeds the type sum by 80,961, or 2.3823%. One fixed arithmetic repair cannot resolve both examples.

Two honest counting stages can still make one misleading file

The strongest countercase is not that arithmetic is optional. It is that the fields may answer different questions.

total_queries might count requests accepted at a service boundary, while query_type_distribution counts objects after parsing. The type counters might be emitted by a later component, use a different event time, include retried internal work, exclude malformed commands or settle after the total is closed. ASN enrichment may succeed for one record and fail for another. A query can also ask for more than one class, depending on how the service records command syntax. Any of these arrangements could be operationally legitimate.

The public README does not say that this is the arrangement. It says each WHOIS and RDAP section contains a total number of queries and a mapping from query type to query count. It defines the one-hour time window and the per-ASN fields, but it does not name separate counting boundaries, late-arrival rules, retry treatment, exclusions or a multi-class rule. Nor does it provide an unclassified field, a validation result or a pointer from a corrected file to the version it replaces.

That omission is why the article cannot choose a winner. The evidence does not show whether 3,192,583 or 3,183,463 is the right answer to some properly specified question. It shows that a reader is offered both without the specification needed to distinguish the questions.

It also explains why checksums are insufficient. A checksum proves that the bytes a reader downloaded are the bytes APNIC published. It cannot prove that two fields inside those bytes share a denominator. File integrity and metric coherence are separate properties.

APNIC’s own status now contains an unresolved layer

The policy record and the conference material should be read carefully together. The prop-167 page still says Implemented and records “Implementation Complete” on 30 June. A presentation posted for the APNIC 62 Open Policy Meeting scheduled on 10 September uses different language. Its prop-167 slide says “Status: In Implementation”, notes that the statistics were initially published on 30 June, and adds: “Some issues with data accuracy – Currently being resolved”. It expects completion by the end of Q3.

That deck was available before its scheduled session. It is evidence of a prepared APNIC document, not evidence that Dave Phelan had already delivered the words to the Policy SIG. More importantly, the slide does not identify which data problem it means. It would be irresponsible to claim that APNIC has diagnosed the arithmetic documented here, or that the discrepancy caused the status change.

The slide does establish a narrower fact: APNIC itself no longer presents initial publication as the end of the implementation story. The public policy page and the posted implementation update now describe different layers of completion. One records that a service was launched. The other leaves data accuracy open.

That is a useful distinction. A feed can be live, hourly, continuous and checksummed while its published fields still need semantic reconciliation. “Implemented” can truthfully describe availability; “in implementation” can truthfully describe assurance. The mistake would be forcing both into one undifferentiated status.

The denominator changes what the archive can prove

The feed has real value. It makes aggregate WHOIS and RDAP use visible without publishing source IP addresses or the records being queried. Researchers can see service volumes, object-type mixes, origin-ASN concentration and distinct source-IP counts. Operators can compare hours and investigate exceptional patterns. APNIC deserves credit for making the result downloadable and preserving hourly files.

Yet each reuse chooses a denominator, whether the analyst acknowledges that choice or not. A concentration share calculated from an ASN row divided by total_queries is different from a share divided by the type sum. A trend line using the total can move differently from one built by adding types. If a negative gap is caused by multi-class counting, summing types may deliberately overstate unique requests. If a positive gap is caused by unclassified requests, the type mix may omit part of the service load. Without stage definitions, a perfectly accurate chart can answer an unstated question.

The per-ASN rows add a second boundary. Each service may report thousands of origin ASNs, while the asns array contains at most 1,000 rows in the observed files. That may be an intentional top-list limit, but the README does not label it as such. The sum of displayed rows therefore need not equal the service total even after their internal type arithmetic is repaired. A complete public contract needs to distinguish a truncated ranking from a population total.

None of this proves a service incident. It does not show that a query was dropped, duplicated or answered incorrectly. It does not identify a faulty ASN, reveal a source address or imply malicious automation. The zero-query hour is not, by itself, proof of an outage. The findings concern the evidence surface: what a third party can reconstruct from the published file.

Publish the bridge and the correction chain

APNIC does not need to disclose raw logs to repair that evidence surface. It needs a small, privacy-safe reconciliation record for each hour.

The record should begin with the counting equation. Name the number of requests accepted at the external boundary, the number successfully parsed and classified, and the counts excluded as malformed, unsupported, timed out or otherwise unresolved. If one query can contribute to more than one type, state that rule and publish the resulting multi-class increment. If totals and types close at different times, give both cutoffs and the late-arrival treatment.

Next, identify the transformation. Publish a run identifier, ingestion version, classifier version and the result of a machine check that recomputes every displayed equation. Report source-ASN enrichment successes and unknowns as aggregates, not addresses. Label the 1,000-row array as a ranking or full population, and state the selection rule.

Finally, make correction a first-class event. Each hourly file needs a quality state—provisional, validated, corrected or withdrawn—plus its fingerprint. A replacement should name the superseded fingerprint, correction time, reason family and affected fields. The old version should remain discoverable. A reader could then tell whether a historical chart used the initial publication or the repaired measure.

The design is modest because the question is modest. It does not ask APNIC to expose who looked up what. It asks the publisher of two totals to state how they join. The archive has already done the expensive institutional work: regular generation, public access, preservation and integrity checks. Reconciliation is the receipt that turns those files from observable output into reusable evidence.

Sources