Summary
- BGPMon's public incident account said that Google service reachability problems on 12 March 2015 were tied to a route leak in which Google-originated prefixes were learned by Hathway AS17488 and leaked to Airtel AS9498, which then propagated the announcements to peers [1].
- The observed window was short, from 08:58 UTC to 09:14 UTC, but BGPMon said 336 Google IPv4 prefixes were affected, including prefixes associated with several Google autonomous systems [1].
- RFC 7908 later used the Hathway-Airtel leak as an example of a route leak involving peer-prefix propagation to a transit provider, causing widespread interruption of Google services in Europe and Asia [2].
- The article does not claim malicious intent, exact customer loss, or a full internal root cause. The accountable surface is narrower: route acceptance from a customer, export toward transit and peers, relationship validation, route-leak alarms and retained evidence.
What happened
BGPMon described the public signal as a Google service interruption. Its account began with user reports and a press report that European and Indian users were primarily affected. BGPMon then used its BGP data to identify a routing change: many Google prefix paths shifted to include Airtel AS9498 in India between 08:58 UTC and 09:14 UTC [1].
The AS-path detail is the important part. BGPMon wrote that the prefixes were still originated by Google AS15169. That means the event was not a simple route-origin hijack in which a different network pretended to be Google. Instead, BGPMon said Google peered with Hathway AS17488, that Hathway leaked the routes to its transit provider Airtel AS9498, and that Airtel propagated the announcements to peers at Internet exchanges [1].
For readers outside routing operations, this distinction matters. BGP is the routing protocol that lets autonomous systems advertise reachability to IP prefixes. A route can have the correct origin while still traveling through an incorrect relationship path. A network may legitimately learn a route from a peer or customer for one purpose, but it should not export that route into a different relationship where it becomes a misleading transit path.
RFC 7908 gives the vocabulary for that failure class. It defines a route leak as routing announcements propagated beyond their intended scope, where the intended scope is shaped by import and export policies between autonomous systems. The RFC listed the Hathway-Airtel leak of 336 Google prefixes as an observed example and placed it in the route-leak taxonomy, not merely in a generic outage category [2].
The public evidence also has boundaries. BGPMon's data supports the path, prefix count and time window it observed. It does not publish Hathway's route maps, Airtel's import policy, Google internal telemetry, every peer's decision process or any legal allocation of responsibility. A Daniel article should therefore avoid overclaiming. The strong claim is that public routing evidence showed a relationship and propagation control failure. The weaker, unsupported claim would be that every affected user, business loss and operator decision can be reconstructed from public BGP data alone.
Why it matters
This incident is still useful because it shows how a cloud-scale platform can be affected by a small number of routing decisions outside the platform's own network. Google owned the service and originated the prefixes. Yet if a peer-learned route was passed to a transit provider and then preferred by other networks, users could lose normal reachability without Google making the original routing mistake.
That makes the accountable chain multi-party. Hathway's control surface was the route it learned, accepted and exported. Airtel's control surface was whether it accepted the route from Hathway as customer or peer reachability, and whether it then exported it to peers. Other networks' control surface was whether they preferred the path through Airtel over their normal Google route. Google's control surface was monitoring and coordination with peers, but the route leak itself sat at the customer-transit boundary described by BGPMon [1].
The economic issue is cost transfer. A permissive customer route filter can be operationally convenient because it reduces manual exceptions. It also lets one network's route mistake impose debugging, support and availability costs on many unrelated users and networks. If a large transit provider amplifies the route, the cost moves from the local relationship into a wider public dependency.
The incident also shows why route security cannot stop at RPKI origin validation. If the origin remains Google AS15169, an origin-only check can pass while the path is still wrong. Route-leak accountability is about the path and the relationship, not only the prefix holder. The relevant questions are whether the route was learned from the right neighbor, whether it was allowed to leave that neighbor relationship, and whether downstream networks had enough evidence to reject the path.
The technical layer
The simplest model is a route-policy ledger. A network should know which prefixes each customer, peer and provider may send, and where those routes may be exported. If a customer sends its own prefix, the provider may carry it. If a peer sends a route, the receiving network generally should not turn that route into transit for another provider or peer. The exact commercial relationship can be complex, but the router policy still has to encode it.
BGPMon's account said the affected route path involved Google AS15169, Hathway AS17488 and Airtel AS9498. It said Airtel propagated the announcements to peers at Internet exchanges, and that some networks may have preferred the Airtel path because customer routes can be preferred over peering routes [1]. That observation puts the incident at the boundary between local preference, customer/provider economics and route-leak filtering.
The first control is prefix authorization. A provider should know what a customer is authorized to originate or transit. If Hathway was not authorized to provide transit for those Google prefixes toward Airtel, the route should have been rejected or at least held for review.
The second control is relationship tagging. A route learned from a peer should carry a provenance tag that prevents it from being exported as if it were customer reachability. A route learned from a customer should be checked against expected prefix ownership, route objects, historic paths and explicit customer authorization.
The third control is export policy. Even if a route is accepted for local use, it should not automatically be sent to every peer. Export policy must answer a narrower question: is this route appropriate for this neighbor under this relationship?
The fourth control is route-leak detection. A short leak can still be measured. The useful record includes first seen, last seen, AS path, affected prefixes, route collectors, local preference changes, withdrawals and post-repair collector checks. BGPMon's 08:58 to 09:14 UTC window is exactly the kind of public timestamp that operators should be able to reconcile with internal logs [1].
The fifth control is customer and peer communication. When a leak affects a major platform, private coordination can shorten the event. But private coordination is not the same as public accountability. Customers and peers need enough detail afterward to know whether the failed route class has been blocked.
Who was affected
BGPMon's article said reports indicated European and Indian users were primarily affected, and it concluded that the leak was picked up only in Europe, matching the Twitter reports it referenced [1]. RFC 7908 used broader language about interruption of Google services in Europe and Asia [2]. Vice's later narrative described the event as a globe-spanning Google outage and used the incident to explain how the global routing system can turn route choices into user-visible outages [3].
These sources justify a reachability-impact framing. They do not justify invented user counts, compensation figures, internal Google losses or a claim that every Google service was down globally. The safe formulation is that public route observers and public reporting linked the event to Google service reachability problems, with the clearest technical evidence coming from the route path and prefix data.
The affected groups were therefore not just Google users. Google had to diagnose and coordinate around a path it did not originate as an error. Airtel peers had to decide whether to accept and prefer the path. Network operators using public collectors gained evidence but not internal causation. Customers experienced the result as a service interruption, even though the practical control surface sat in route policy among autonomous systems.
Accountability standard
For Hathway, the closeout evidence would start with the session that learned Google routes and the session that exported them. It should show route maps, prefix lists, relationship classification, exception records, route-object checks and the withdrawal timeline.
For Airtel, the closeout evidence would start with import policy from Hathway, local preference handling, export policy to peers and route-leak alarms. If Airtel treated the route as customer reachability and exported it broadly, the evidence should show why that classification was allowed and what changed after the incident.
For affected peers, the evidence would include why the Airtel path was preferred over direct or normal Google paths. BGPMon suggested local preference for customer routes as one possible reason [1]. That is a business-routing interaction, not a software bug by itself. The accountability issue is whether preference policy can amplify an abnormal path faster than monitoring can catch it.
For Google, the evidence would include external route monitoring, contact paths with upstream and peer networks, and user-impact correlation. Google may not own the leaking policy, but a major platform has a resilience interest in detecting when its prefixes are carried through unexpected paths.
The Heng.lu doctrine surface is BGP/routing and network-resource evidence. The registry can say who holds a prefix or autonomous system, but running route policy decides where traffic actually moves. The route ledger must therefore connect registry rights, AS relationships, router policy and operational continuity. A correct origin is not enough if the path violates the intended relationship.
What to watch next
The first signal is whether operators can name the failed route class. A vague statement that routing was fixed is not enough. The public record should say whether the problem was customer-prefix filtering, peer-route export, local preference, route-object generation, max-prefix configuration, or a temporary exception that was not bounded.
The second signal is whether relationship-aware controls are deployed. RFC 7908 described the problem definition. Later controls such as BGP Roles and Only-to-Customer, and ASPA-style path authorization, move in the direction this incident requires: route policy that can represent provider/customer/peer scope instead of relying only on memory and manual filters.
The third signal is whether public collectors and internal telemetry can be reconciled. If an event is visible for sixteen minutes, the operator should be able to compare collector paths with route logs, flow data, interface counters, alerts and customer reports. If the public route disappears but support impact continues, the route leak was not the whole incident. If the public route appears but no traffic selected it, the severity should be lowered.
The fourth signal is whether repeated events produce stronger controls. Route leaks are common enough that each event should improve route hygiene. The dangerous pattern is not one short incident. It is repeated permissive acceptance, broad propagation, no public route evidence and no testable repair.
Scenario changes
The assessment becomes more severe if evidence shows that the route was accepted despite clear prefix authorization mismatches, that alarms fired and were ignored, that the same relationship class leaked again, or that public communications omitted the AS path and prefix evidence after user-visible harm.
The assessment becomes less severe if operators can show a short duration, limited selected traffic, automatic suppression, route-map repair, clear peer notification and external collector confirmation that the failed path class no longer propagates.
The assessment would fail if it turned into a generic Google outage story. The route-leak record is valuable precisely because it keeps responsibility tied to the network-control surface: which AS accepted the route, which AS exported it, which peers preferred it, and which evidence proves the next equivalent route will fail closed.
Sources
[1] BGPMon, "What caused the Google service interruption?", https://www.bgpmon.net/what-caused-the-google-service-interruption/
[2] RFC 7908, "Problem Definition and Classification of BGP Route Leaks," https://www.rfc-editor.org/rfc/rfc7908.txt
[3] Vice, "Anatomy of a Globe-Spanning Google Outage," https://www.vice.com/en/article/anatomy-of-a-globe-spanning-google-outage/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
