Summary

  • At about 02:24 UTC on 6 November 2012, Cloudflare observed that Google services were unreachable from parts of the Internet and traced routes for Google addresses through Moratel, AS23947, and PCCW, AS3491, before the legitimate Google origin, AS15169 [1].
  • Cloudflare described the disruption as limited and lasting about 27 minutes. It estimated that 3-5 percent of the Internet population may have been affected, with heavier impact around Hong Kong. Those figures remain Cloudflare's estimates, not a complete independent census [1].
  • RFC 7908 later listed the Moratel-PCCW incident as a Type 4 route leak: peer prefixes were leaked toward a transit provider and propagated beyond their intended scope [2].
  • Cloudflare initially assessed that an erroneous announcement was likely. Its update recorded Moratel's statement that an unexpected hardware failure created the abnormal condition and that the event was not malicious [1]. The public record does not disclose enough internal evidence to decide the precise failure mode.
  • The accountability surface spans both export and import. Moratel needed to prevent unauthorized routes from leaving the relevant session; PCCW needed to distinguish acceptable customer routes from routes learned through another relationship. Detection, withdrawal, evidence retention and verified repair were shared operational duties.

What the public record shows

Cloudflare reported that its staff noticed Google's services were offline from its network at about 02:24 UTC. The investigation began with failed DNS reachability, including an inability to reach Google's public resolver at 8.8.8.8. A trace then showed traffic moving through a Moratel address in Indonesia, an unexpected path for traffic from Cloudflare's California network toward Google [1].

The route observed by Cloudflare for a Google prefix included the sequence AS4436, AS3491, AS23947 and AS15169. In that sequence, AS15169 remained the Google origin. The abnormality was in the path used to reach it: Moratel appeared between PCCW and Google in a route that Cloudflare did not expect to receive through that relationship [1]. Cloudflare observed the same pattern for Google's public DNS prefix.

That distinction matters. The event was not presented as an unknown network pretending to originate Google's address space. It was a propagation failure in which a route associated with the legitimate origin travelled across the wrong relationship boundary. The destination identity at the end of the path could still be correct while the path in the middle sent traffic toward a network that was not prepared or authorized to carry it.

Cloudflare said it contacted a colleague at Moratel. The announcement was corrected at about 02:50 UTC, and routing returned to normal roughly three minutes later [1]. The write-up estimated a 27-minute limited outage and a larger effect in parts of Asia, particularly around Hong Kong. The article does not turn that estimate into a universal outage count. Cloudflare had a useful operational vantage point, but it was not observing every user, every network or every Google service.

The same source also preserves an important correction. Its original analysis said someone at Moratel likely made an erroneous route change. A later update recorded Moratel's statement that an unexpected hardware failure caused the abnormal condition, that the event was not malicious, and that Moratel shut down the relevant BGP peering while examining the failure [1]. Without Moratel's router logs, hardware alarms, configuration history and incident timeline, the public cannot determine whether the first control failure was hardware, software, configuration, failover behavior or an interaction among them.

Why RFC 7908 calls it a Type 4 route leak

The Internet is divided into autonomous systems. Each autonomous system, identified by an autonomous system number, applies its own routing policy and exchanges reachability information with neighboring networks using the Border Gateway Protocol, or BGP. A route announcement carries an address prefix, an AS path and other attributes. Operators then decide which routes to accept, prefer and advertise onward.

Those decisions are not only technical. They encode business and operational relationships. A customer normally pays a provider for reachability to the wider Internet. Peers normally exchange routes for their own customers, not full transit for every route they have learned. A route server at an exchange has another role. The same prefix can therefore be valid in one session and invalid when exported across another.

RFC 7908 defines a route leak as propagation of routing announcements beyond their intended scope. Its Type 4 category covers a network leaking routes learned from a lateral peer to its own transit provider. The RFC names the Moratel-PCCW leak of Google prefixes as an example and says the event caused Google's services to go offline [2].

That classification gives the incident a precise accountability frame. The critical question is not merely whether AS15169 was allowed to originate the Google prefixes. It is whether AS23947 was allowed to export the learned Google path toward AS3491, and whether AS3491 should have accepted and propagated it under the relevant customer or transit policy.

Origin validation alone does not answer that relationship question. A route can end at the correct origin ASN and still contain an unauthorized provider, customer or peer transition. The 2012 event therefore belongs to the path-policy layer. Registry records and origin authorization remain useful, but they cannot replace session roles, export policy, import filters and route-level evidence.

The export-side evidence duty

The first evidence duty sits at the exporting network. Moratel should have been able to identify the exact BGP session involved, the role assigned to its neighbor, the routes learned on that session, the routes selected locally, and the routes advertised toward PCCW. That record should distinguish self-originated prefixes, customer prefixes, peer-learned routes and provider-learned routes.

An incident closeout should then connect the running route to an approved policy. If a hardware failure caused the abnormal condition, the useful evidence would show how that failure changed the routing state. Did a line card or route processor restart with an incomplete policy? Did a failover path activate a different configuration? Did a damaged or stale policy object replace a relationship-specific filter? Did the router advertise a broader table while converging? The public source does not answer those questions.

The difference between an explanation and evidence is important. "Hardware failure" identifies a class of trigger. It does not show why the route export was able to fail open. A robust export policy should make an unexpected device state less able to transform a peer-learned route into a customer route. That requires explicit policy at the session boundary, limits on the routes that can leave it, and monitoring of the actual advertised-route set.

The exporting operator should also retain a before-and-after record. A configuration diff is useful, but it is not enough by itself. Operators need the route information base, the routes advertised to the neighbor, relevant communities or attributes, session resets, hardware events, timestamps and the withdrawal sequence. Those records reveal whether the repair changed only the symptom or closed the control path that allowed the leak.

The upstream acceptance duty

The second evidence duty sits at the receiving network. Cloudflare described PCCW as Moratel's upstream provider and said PCCW trusted and propagated the routes Moratel sent [1]. Current RDAP identifies AS3491 as PCCWG-APAC-HK and associates it with PCCW Global (HK) Limited [6]. That current identity evidence supports directory continuity; it does not by itself prove every commercial term or session role that existed in 2012.

An upstream is not a passive pipe. It chooses which customer announcements to accept, how to classify them and where to propagate them. For a customer that may legitimately carry third-party routes, a simple origin-only allowlist can be difficult. The difficulty does not remove the acceptance duty. It makes the relationship record and provisioning system more important.

The receiving operator should be able to show what prefixes and paths the customer was authorized to announce, how that authorization was generated, how often it was reconciled, what maximum-prefix or route-count controls applied, which unusual origin or path patterns triggered alerts, and who had authority to override a block. If the accepted route was outside the intended customer cone, the closeout should identify why the import policy did not reject it.

Route leaks become shared incidents because each propagation step increases the blast radius. A customer can make the first erroneous export, but a large upstream can turn a local fault into a broadly visible path. Accountability therefore cannot end with "the customer sent it." The upstream's product is controlled propagation. Its evidence should show the point at which an abnormal announcement was accepted, selected and advertised onward.

What current registry records can and cannot prove

Current APNIC RDAP identifies AS23947 as MORATELINDONAP-AS-ID and associates operational contacts with PT. Mora Telematika Indonesia [5]. Current RDAP for AS3491 identifies PCCWG-APAC-HK and PCCW Global (HK) Limited [6]. These records help bind the autonomous system numbers to current network identities.

They do not reconstruct the complete 2012 event. Registry data can record who holds or operates a number resource and how to contact the responsible organization. It does not show which routes a router accepted at 02:24 UTC, which commercial relationship applied to a session, which filter version was active, or which hardware state changed.

This is the Heng.lu surface of the article. A registry is a ledger and recordkeeper, not a substitute for the running network. Accurate ASN records matter because operators need durable identity, contact and transfer continuity. The accountability verdict still comes from reconciling those records with the route that actually ran, the policy that should have constrained it and the evidence retained after repair.

RIPEstat's historical routing response for 8.8.8.0/24 shows visibility of AS15169 as origin during the requested day [4]. Its returned granularity is eight hours, so it cannot prove the short-lived leaked AS path or its minute-by-minute withdrawal. This article therefore relies on Cloudflare's contemporary path observation and RFC 7908's later classification for the event path, while treating the RIPEstat response only as bounded historical context.

Controls that should fail closed

The first control is explicit import and export policy. Every external BGP session should have a defined role and a route policy that reflects it. A missing policy should not default to accepting or advertising the full table. A route learned from a peer should not silently become exportable toward a provider.

The second control is a provisioning ledger tied to running configuration. The ledger should identify the neighboring ASN, relationship role, authorized prefixes or customer cone, permitted origins, expected path patterns, route limits, change owner and approval record. Automation should compile that intent into device policy and verify that the deployed policy matches the approved record.

The third control is route-set monitoring. Operators should compare the routes advertised to each neighbor against an expected set. A sudden appearance of many third-party prefixes, an unexpected upstream in an AS path, or a customer advertising routes outside its authorized relationship should trigger containment before broad propagation.

The fourth control is relationship-aware signaling. RFC 9234, published long after the incident, specifies BGP Roles and the Only-to-Customer attribute. Those mechanisms can make relationship boundaries more machine-readable and help detect paths that violate expected propagation rules [3]. They are useful control context, not evidence that either operator used them in 2012.

The fifth control is independent route observation. External collectors and network vantage points can show that a route escaped the intended boundary even when both adjacent operators believe their local configuration is correct. Collector evidence should be retained with timestamps and compared against router logs. It should not replace the operators' own Adj-RIB-In, local decision and Adj-RIB-Out evidence.

The sixth control is a tested withdrawal path. Recovery depends on more than stopping a bad announcement. Operators need to know how quickly withdrawals propagate, whether stale routes persist, how caches and control-plane convergence affect user reachability, and which external observers confirm that the route has disappeared.

A useful public closeout

A useful public closeout would begin with a bounded timeline: first observed symptom, first internal alert, first confirmed abnormal route, first contact between operators, export stop, withdrawal visibility and stable recovery. It would use UTC and distinguish observation time from diagnosis time.

It would then state the affected route class. The closeout need not publish sensitive full configurations. It can say whether peer-learned routes were exported toward a transit provider, whether the active policy was missing, stale or bypassed, and whether the receiving network's customer filter accepted a route outside the approved set.

The closeout should preserve uncertainty. If a hardware failure triggered the incident, the operator should explain the control effect without claiming more precision than the evidence supports. If the internal logs no longer exist, that absence is itself a continuity finding. It limits the operator's ability to prove that the repair addressed the cause rather than only withdrawing the visible route.

Finally, the closeout should state the durable change and how it was tested. A statement that "filters were updated" is weaker than evidence showing the intended relationship, generated policy, deployed policy hash, rejected test route, alert threshold, rollback owner and external route-view verification. The goal is not to expose sensitive network details. It is to make recurrence claims auditable.

Why the incident still matters

The event is old, but the control problem remains current. BGP still joins independently operated networks whose local policies can impose costs on remote users. Route-origin security has improved, and relationship-aware standards have developed, yet a correct origin does not automatically make the path legitimate.

The incident also shows why short outages deserve disciplined evidence. A 27-minute event can be dismissed as transient after routes recover. For affected users, the service was still unreachable. For operators, the short duration may be the only window in which the abnormal path, hardware state, route export and neighbor acceptance can be captured together.

The strongest conclusion is narrow. Public evidence supports that a Moratel-PCCW route leak affected reachability to Google services on 6 November 2012, that Cloudflare observed and helped escalate the abnormal path, and that RFC 7908 later classified the event as a peer-to-transit leak [1][2]. Public evidence does not settle the internal failure sequence or divide final responsibility. That gap defines the accountability duty: preserve enough running-state evidence that the next explanation can be tested rather than merely accepted.

Sources

  1. https://blog.cloudflare.com/why-google-went-offline-today-and-a-bit-about/
  2. https://www.rfc-editor.org/rfc/rfc7908.txt
  3. https://www.rfc-editor.org/rfc/rfc9234.txt
  4. https://stat.ripe.net/data/routing-history/data.json?resource=8.8.8.0/24&starttime=2012-11-06T00:00:00&endtime=2012-11-07T00:00:00
  5. https://rdap.apnic.net/autnum/23947
  6. https://rdap.arin.net/registry/autnum/3491