Summary

  • RIPE NCC reported on 8 September that a Kafka-partition processing step timed out under unusually large record volume at startup; RRC24 files had been published.
  • The operator now considers a connection to the earlier RRC18 incident unlikely. Neither a precise repair nor a tested startup capacity was disclosed in the statement.

A service can be good at keeping up and poor at getting started. RIPE NCC's latest explanation for missing RRC24 routing-data files puts that distinction at the centre of this incident. It points to work concentrated at startup, rather than establishing a failure of ordinary publishing throughput.

At 16:01 CEST on 8 September, the operator's status page attributed the timeout to unusually many records handled by a step processing a Kafka partition at startup. It also reported publication of the updates and bviews. A day earlier, the operator had suspected a common cause with RRC18; its revised assessment says that connection is unlikely and further collectors are not expected to be affected. That is a materially narrower assessment, not a guarantee about every collector.

The distinction matters to users of routing evidence. RIS documentation describes per-collector MRT files: dumps record routing state, while updates record changes over an interval. Their documented generation intervals are eight hours and five minutes respectively. Delayed publication can delay an analyst's evidence without proving that the networks being observed stopped forwarding traffic. This report has not audited the recovered archive for completeness.

Collector metadata places RRC24 in Montevideo, Uruguay, as a multihop collector serving the LACNIC region, with LACNIC as sponsor. Multihop sessions need not share the collector's local exchange LAN. Neither that geography nor the sponsorship assigns responsibility for the processing timeout to LACNIC.

The new cause statement supports a practical inference: startup deserves its own workload test. A publication schedule describes the output normally expected; it does not specify how many records a starting processing step can complete before its deadline. The status update gives neither that deadline nor the record count. It also does not identify which software change, if any, restored publication. Calling this a Kafka product defect, a BGP outage or a proven data-loss event would go beyond the evidence.

There is a fair case for the operator's limited statement. An incident notice can identify a useful cause without serving as a full engineering report. RIPE NCC also revised an earlier hypothesis instead of preserving it for consistency. The unresolved question is narrower: what evidence should a service owner require before accepting the next startup as adequately tested?

This reading follows Lu Heng's distinction between documenting reality and advocacy. That is an editorial principle, not technical evidence about RRC24. The reported timeout is a fact attributed to the operator; the proposed acceptance test is our analysis.