Summary

  • Independent monitoring places the start of a large BGP announcement burst from Telekom Malaysia's AS4788 at about 08:43 UTC on 12 June 2015. BGPMon reported roughly 179,000 announced prefixes, while RFC 7908 later cited the event as a major route-leak example in which Level 3 accepted and propagated about 179,000 prefixes. [2][10]
  • Different analyses report different route counts because they use different collectors, time windows and definitions. Geoff Huston's analysis discussed approximately 2,500 newly visible routes and examined a set of 22,577 affected routes. Those figures cannot responsibly be merged into one false precision. [1][2]
  • Level 3's AS3549 did not merely observe the AS4788 announcements. It accepted and propagated them, extending the blast radius through a major global transit network. ThousandEyes measured severe packet loss and terminal paths at multiple Level 3 points of presence. [2][3]
  • The evidence supports a relationship-policy leak: routes learned from peers appear to have been re-advertised toward upstream transit providers. The exact Telekom Malaysia router configuration, route map, command, software and approval chain are not public in this packet. [1]
  • Accountability is divided by control. AS4788 controlled its export policy and route-map lifecycle. AS3549 controlled what it accepted from a customer, the volume and path checks applied, and whether accepted routes were propagated. Downstream networks controlled their own import policy and monitoring.
  • Ordinary RPKI Route Origin Validation is not a complete answer to this event. Most leaked paths retained legitimate origins. Origin authorization can be valid while a path violates customer, peer or provider export intent. [1][15][16]
  • Later standards clarify possible controls. RFC 8212 makes explicit import and export policy the default requirement for eBGP, while RFC 9234 adds relationship-aware BGP Roles and the Only to Customer attribute. They are retrospective control guidance, not proof of a 2015 compliance violation. [11][12]
  • A credible repair claim requires more than restoration. It needs a frozen reconstruction of the announcement set, the intended session policy, before-and-after configuration evidence, tests against the same leak class, independent route observation, and proof that both the exporting and accepting sides can contain recurrence.

A route announcement became authority over other people's traffic

BGP is often introduced as the protocol that tells the Internet where networks are. That description is accurate but incomplete. A BGP announcement is also a claim of operational authority.

When one autonomous system tells another that a prefix is reachable through it, the recipient can prefer that path and advertise it onward. Other networks can then direct traffic toward the advertised path. The route does not carry a warranty that the announcing network has enough capacity, that the path conforms to commercial relationships, or that every intermediate operator intended to provide transit. BGP distributes reachability information and AS-path data; policy determines which claims a network accepts and repeats. [8]

That policy layer is where the 2015 Telekom Malaysia event became globally consequential.

Independent sources say a large announcement burst began at about 08:43 UTC on 12 June. Telekom Malaysia's AS4788 advertised a huge set of routes to Level 3's AS3549. Level 3 accepted those routes and propagated them to peers and customers. Traffic followed the changed paths. The path through AS3549 and AS4788 could appear attractive under routing policy even though the interconnections could not safely carry the resulting volume. Packet loss and latency rose, and services far beyond Telekom Malaysia's own customer base became difficult or impossible to reach. [2][3]

The event was not a forged certificate, a compromise of a domain name, or a fabricated route origin in the simplest sense. Many affected prefixes still ended at their legitimate origin autonomous systems. The harmful change was that AS4788 inserted itself as transit for routes it was not expected to export in that direction. A path can be syntactically valid, loop-free and origin-valid while violating the economic and operational relationship under which it was learned.

This distinction matters for accountability. If the problem is described only as "Telekom Malaysia leaked routes," responsibility appears to end at the exporting network. But BGP propagation is bilateral at every session. One side announces; the other side decides what to accept, prefer and advertise onward. A large transit network has greater propagation power than an isolated customer. That power creates a corresponding filtering and evidence duty.

The central question is not which operator first made the mistake. It is how many independent controls had the practical ability to stop the mistake before it became other networks' outage.

The chronology is clear at the edges and incomplete inside the networks

BGPMon reported that AS4788 began announcing a massive route set at 08:43 UTC. Its monitoring saw a sharp increase in BGP update messages at the same time that packet loss began. The analysis described approximately 179,000 announced prefixes and gave an affected Facebook prefix as an example of a path that ran through AS3549 and AS4788 before reaching the legitimate origin. [2]

ThousandEyes independently described the same broad sequence. It observed new paths through Telekom Malaysia and Level 3, severe packet loss, and terminal routes in locations including Amsterdam, Chicago, Frankfurt, London, Los Angeles, Seattle and Washington. It said Level 3 stopped accepting the routes at about 10:45 UTC and service began returning to normal. [3]

BGPMon reported improvement around 10:40 and broader clearing around 11:15. Those times should remain attributed. A route collector, active measurement platform, transit operator and end user do not observe the same event at the same instant. A filter can stop new advertisements while stale paths remain selected elsewhere. Withdrawals can propagate unevenly. Congestion can persist after the control-plane trigger is removed. Recovery is therefore a sequence, not one universal timestamp.

RIPE Atlas analysis later used the event to examine how large failures in core infrastructure affect end-to-end connectivity. It found evidence both of traffic routing around infrastructure under stress and of end-to-end failures. The authors were explicit about representativeness: even a diverse global measurement system observes only a finite set of paths and destinations. [4]

The public chronology has a strong external record and a weak internal record.

External observers can identify the approximate start, paths that changed, route-update volume, packet loss, latency and broad recovery. They cannot see the private route map, the operator terminal, the approved change, the configured prefix limit, the alert queue, or the decision conversation between Telekom Malaysia and Level 3.

That gap should shape the language of the article.

It is defensible to say AS4788 emitted the routes, AS3549 accepted and propagated them, and global reachability suffered. It is defensible to say the pattern is consistent with peer-learned routes being exported toward an upstream provider. It is not defensible to identify the exact command, router, employee, change ticket or software defect without an authenticated operator record.

Geoff Huston's analysis uses probabilistic wording about a route-policy failure and contains apparent AS-number typographical variants in some passages. Telekom Malaysia's network is AS4788. The article should not turn apparent AS4877 or AS4778 references into additional actors or use them to manufacture certainty about an internal device. [1]

A complete accountability record would connect the external timeline to internal evidence:

  • the last known-good export policy;
  • the proposed and normalized change;
  • the time each router or session received it;
  • the number and type of prefixes selected for export;
  • alerts for route volume and relationship violations;
  • acceptance and maximum-prefix state at AS3549;
  • escalation contacts and messages;
  • the command or automated action that stopped propagation;
  • collector evidence showing withdrawal and convergence;
  • tests proving the repaired policy rejects the same route class.

Without that chain, restoration is visible but institutional learning remains difficult to verify.

The prefix counts describe different views, not one disputed fact

Large Internet incidents attract a single memorable number. Here, that instinct can make the record less accurate.

BGPMon wrote that AS4788 began announcing about 179,000 prefixes and later referred to approximately 176,000 leaked prefixes. RFC 7908 cites the "massive Telekom Malaysia route leak" of about 179,000 prefixes. ThousandEyes described a large portion of the global routing table. [2][3][10]

Huston's analysis used different views of the event. It showed a net routing-table change involving thousands of newly visible and withdrawn routes, then examined 22,577 routes in a specific affected set. [1]

These figures can coexist because a BGP event does not have only one natural unit.

An observer can count every UPDATE message, every unique prefix announced through an unexpected path, every prefix newly visible at a collector, every changed best path, every more-specific route, every route still present at a selected time, or every affected origin. Collectors receive different feeds. A route can be announced, withdrawn and re-announced. Some paths are visible at one collector and not another. A full table and a filtered affected set answer different questions.

The accountable editorial choice is to preserve the measurement definition.

The article may say that BGPMon and RFC 7908 described roughly 179,000 leaked announcements or prefixes in their reconstructions. It may say that Huston's separate analysis examined a 22,577-route set and observed thousands of table additions and withdrawals. It should not average the values, choose the largest for drama, or present one as a complete census of user impact.

The same discipline applies to affected services. BGPMon and ThousandEyes identified examples involving major platforms and financial services. Those examples demonstrate scope and collateral effects. They do not establish that every prefix experienced the same packet loss, every service became unavailable, or every user routed through AS4788.

Counts become accountability evidence when their definitions are retained:

  • announcement count tests whether export volume was anomalous;
  • unique prefix count tests the breadth of routing authority claimed;
  • changed best-path count tests how many networks selected the leak;
  • collector visibility tests propagation;
  • traffic volume and packet loss test operational harm;
  • affected customer and application count tests business impact;
  • withdrawal duration tests containment.

Each measure should have an owner, threshold and retained record. A transit provider might accept a customer's normal set of a few thousand prefixes yet quarantine a sudden order-of-magnitude change. A route-monitoring system might detect paths that violate customer-cone expectations even when raw prefix volume stays below a static limit. A public postmortem might explain both measures rather than offering one headline total.

False precision is not only a writing problem. It can hide which control failed.

Relationship policy is the invisible structure behind BGP reachability

The Internet is not a flat mesh in which every autonomous system offers free transit to every other system. Networks buy transit, sell transit and peer under relationships that shape route policy.

A simplified operational rule works like this:

  • customer-learned routes may be advertised to customers, peers and providers;
  • peer-learned routes may be advertised to customers, but ordinarily not to another peer or provider;
  • provider-learned routes may be advertised to customers, but ordinarily not to another provider or peer.

These rules produce the familiar "valley-free" model. A path can climb from customers toward providers, cross at most one peer relationship, and descend toward customers. A path that descends and then climbs again can indicate that a network is providing unintended transit. [1][10]

Real commercial relationships are more complicated. Two networks can have different roles in different places, address families or services. Partial transit, paid peering, route servers and regional arrangements do not always fit a single label. That complexity is a reason to document and test policy, not a reason to omit it.

Huston's reconstruction says AS4788 appeared to collect routes from exchange-point peers and re-advertise them to upstream transit networks. In that model, routes learned laterally were exported "uphill." Level 3 then accepted and propagated the paths. [1]

The protocol itself cannot infer every private business relationship from the AS path. A sequence of legitimate AS numbers does not say whether a route was contractually and operationally permitted to travel through them. That knowledge must be encoded in local policy, published routing entities, negotiated roles, communities, customer-cone data, or another validation system.

This is why route leaks remain difficult. A router can receive a valid BGP UPDATE from an authenticated neighbor, see a legitimate origin, construct a loop-free AS path, and still accept a route that violates the intended relationship.

Operational accountability therefore requires networks to make their expectations machine-checkable where possible:

  • classify each eBGP session and each exceptional policy;
  • define the prefixes and customer paths expected from the neighbor;
  • constrain exports according to how routes were learned;
  • compare a proposed policy with the intended relationship;
  • reject or quarantine unexplained expansion;
  • retain a human-readable explanation for exceptions;
  • test the policy against representative full-table conditions.

The public event shows what happens when relationship intent remains implicit or enforcement is ineffective. A route can cross one session and become a global claim before any human reads a ticket.

AS4788 controlled export, but AS3549 controlled acceptance and propagation

Telekom Malaysia had the most direct control over the announcement set that left AS4788. An exporting network should know which routes it originated, which it learned from customers, which it learned from peers or providers, and which classes may be sent to each neighbor.

That control begins before configuration activation.

A change should be compiled into the actual prefix and AS-path policy that a router will enforce. A review should compare the result with expected customer cones, route counts and relationship rules. A test environment or offline evaluator should feed representative routes through the policy and show what would be exported. An independent check should flag peer- or provider-learned routes selected for another non-customer session.

The public record does not establish whether such controls existed at AS4788, whether a routine configuration changed, or whether a latent state was triggered. It establishes the output: a large, unsafe route set was exported.

Level 3's control boundary is separate and equally important to global propagation.

AS3549 chose whether routes received from AS4788 were eligible, how they were preferred, and where they were advertised. A major transit provider has customer-specific knowledge that arbitrary third parties do not. It can know the expected prefix count, registered customer routes, observed history, customer-cone relationships and session purpose. It can apply:

  • explicit import policy;
  • prefix lists derived from authenticated routing data;
  • AS-path and customer-cone constraints;
  • maximum-prefix thresholds;
  • route-length and bogon checks;
  • relationship-aware leak detection;
  • quarantine or lower-preference policy for anomalies;
  • human approval for exceptional expansion.

BGPMon's account says Level 3 accepted the announcements and advertised them to peers and customers. The routes then attracted traffic and contributed to congestion in Level 3 and major peering locations. [2]

This does not mean an upstream can guarantee every customer route is correct. Static filters can become stale. Multi-homed customers can legitimately change announcements. Emergency routing can expand a set. Complex policy can make customer cones difficult to calculate. A filter that is too strict can cause an outage of its own.

But those costs do not erase the provider's agency. They define the engineering problem.

A transit provider with global propagation power should be able to answer:

  • What range of route count and path shape was normal for this customer?
  • Which changes required pre-coordination?
  • Did the accepted set include routes with other major peers or providers behind the customer?
  • Did a maximum-prefix threshold exist, and was it set against a realistic baseline?
  • Did the session have an exception that disabled or weakened checks?
  • Which alert fired first?
  • Who could suppress the routes without waiting for the customer?
  • How was collateral traffic protected during investigation?

Accountability follows this practical capacity to limit harm. AS4788's export error and AS3549's acceptance are not mutually exclusive explanations. They are successive control failures in the same propagation chain.

Transit concentration turned policy error into shared harm

Not every route leak causes a global incident. Blast radius depends on where the leak is accepted, how attractive the path becomes, how widely it is propagated, and whether the receiving networks have capacity to carry the redirected traffic.

Level 3 was a major global transit provider. Once AS3549 propagated the paths, networks and customers far from Malaysia could select them. Traffic that normally followed direct, regional or better-provisioned routes was pulled toward a path through Level 3 and AS4788. [2][3]

Two harm mechanisms followed.

The first was direct path diversion. A destination prefix could acquire a selected path through AS3549 and AS4788. Packets then traveled toward Telekom Malaysia even though it was not intended to provide global transit for that destination. The interconnection could saturate, packets could be dropped, and latency could rise.

The second was collateral congestion. A service did not need to select a leaked path itself to suffer. If it relied on Level 3 capacity or a congested point of presence, the extraordinary traffic load could impair its normal route. ThousandEyes described examples in which a service's own route remained unchanged but congestion inside Level 3 reduced availability. [3]

This second mechanism is important because it broadens the accountability lens beyond a list of leaked prefixes. Shared transit infrastructure can transmit harm to customers whose routing policy is not directly wrong. Capacity, isolation and traffic engineering become part of the containment problem.

Networks cannot provision every link for an arbitrary fraction of the global table suddenly choosing it. Economic limits are real. Yet a transit provider can design controls so an anomalous route set does not acquire that traffic authority in the first place.

The event therefore connects routing security with concentration risk. A highly connected transit network improves reachability under normal conditions. The same connectivity amplifies a policy failure when unsafe routes are accepted and propagated. Scale is both a resilience asset and a blast-radius multiplier.

Responsible operation should treat propagation reach as a risk variable:

  • a small local customer announcement can use normal automated handling;
  • a sudden customer announcement of routes from many unrelated large networks should require quarantine or validation;
  • a change that would alter paths across many regions should trigger outside-in measurement;
  • a provider should know which points of presence and interconnections would receive redirected traffic;
  • containment should be possible without disabling healthy customer routes unnecessarily.

The goal is not to eliminate automation. It is to make automation proportional to the authority it grants.

Maximum-prefix controls help, but they are not a complete policy

Maximum-prefix limits are an intuitive defense against a huge leak. If a customer normally announces a bounded set and suddenly sends an enormous table, the provider can warn, reject new routes or shut the session.

BGPMon suggested that the abnormal route volume could also cause maximum-prefix limits on Level 3's sessions with other large networks to trigger, producing further churn and path changes. [2]

That observation reveals both the value and danger of simple thresholds.

At the customer edge, a well-calibrated maximum-prefix control can stop an implausible expansion before broad propagation. At downstream sessions, the same mechanism can react after the bad routes have already entered a major provider, potentially dropping an entire session and shifting traffic elsewhere. A limit can contain one path while destabilizing another.

Effective limits require context:

  • the customer's normal aggregate and more-specific prefixes;
  • expected growth;
  • maintenance and emergency scenarios;
  • separate IPv4 and IPv6 behavior;
  • whether rejected routes fail closed or keep the last known-good set;
  • alert escalation before a hard stop;
  • a safe override process with expiration;
  • testing of the response under realistic traffic.

Route count also cannot detect every leak. A customer might leak a small number of highly attractive more-specific routes. It might export routes from a powerful peer without increasing total volume much. It might replace legitimate customer routes with a similar-sized unauthorized set.

Maximum-prefix is therefore one layer. Prefix ownership, customer-cone validation, AS-path relationships, route-source tags and anomaly detection address different failure shapes.

A post-incident record should say which layers existed, not merely that "filters were improved." A maximum-prefix threshold added after the event would be meaningful evidence if the operator published the baseline, threshold logic, response mode and a test using the reconstructed announcement set.

RPKI origin validation would not have solved the path-policy failure

Routing-security discussions often use RPKI as a general answer to BGP incidents. That shorthand is dangerous here.

The Resource Public Key Infrastructure allows holders of Internet number resources to create cryptographically verifiable statements. A Route Origin Authorization identifies which autonomous system is authorized to originate a prefix, subject to the authorization's prefix-length rules. Route Origin Validation can classify a received announcement by comparing its prefix and origin AS with those authorizations. [15][16]

The 2015 AS4788 event was largely a path-policy leak, not a simple unauthorized origin.

For many leaked routes, the legitimate origin remained at the end of the AS path. AS4788 inserted itself as transit and advertised the route to a relationship where it was not expected. An origin validator could see an authorized origin and classify it as valid even though the route violated peer/provider export intent.

Huston's analysis made this point directly. In the route set he examined, only a small minority involved AS4788 appearing as the origin in a way ordinary ROA filtering might address. Most of the problem involved transit information. [1]

This does not make RPKI unimportant. Origin validation can stop unauthorized origins, accidental mis-originations and many hijacks. It can reduce one class of false reachability. It also supplies authenticated resource information that can support broader controls.

It does mean the control claim must be precise.

"We deployed ROV" does not prove protection against routes that have valid origins but invalid relationship paths. A network needs additional information about who may provide transit for whom and which paths are consistent with policy. RPSL, customer-cone data, communities, BGP Roles, the Only to Customer attribute, ASPA-related work and operator-specific filters address parts of that problem at different levels of maturity.

The accountable message is layered:

  • RPKI validates origin authority;
  • explicit import and export policy constrains sessions;
  • relationship-aware controls constrain path propagation;
  • monitoring detects anomalies that static data misses;
  • operational coordination contains what prevention does not stop.

Conflating those layers produces false assurance and weak post-incident learning.

Route registries can publish intent, but stale intent is not control

Routing Policy Specification Language was designed to describe routing policy in Internet Routing Registries. RPSL and RPSLng can express import and export policy, autonomous-system sets, route sets and related intentions. [13][14]

In principle, a provider can use authenticated and maintained policy data to generate filters for a customer. A customer can publish the prefixes and AS relationships it expects to advertise. Peers can compare observed routes with declared intent.

Huston's analysis explains the attraction and the limitations. Registry data can be incomplete, stale, duplicated across databases or too coarse for session-specific relationships. Complex policy can be difficult to express and maintain. Some registries historically allowed third-party entries with weak authority. [1]

The wrong lesson is that route registries are useless. The right lesson is that a registry entity is evidence only when its ownership, freshness, scope and use are verifiable.

A mature filtering pipeline should record:

  • the registry and entities used;
  • authentication and maintenance authority;
  • the last successful refresh;
  • expansion of AS sets into concrete prefixes and paths;
  • conflicts between registries;
  • local exceptions;
  • the generated filter diff;
  • the router deployment result;
  • monitoring for divergence between published and observed policy.

MANRS frames routing security as collective operational responsibility. Its operator actions emphasize filtering announcements, maintaining coordination contacts and publishing information that others can validate. The current implementation guide discusses prefix and AS-path granularity and recommends controls that prevent customer-learned or intermediate routes from being exported to inappropriate non-customer peers. [17][18]

Those current documents postdate the 2015 event in their present form. They should be used as a control framework, not retroactive legal evidence.

The event shows why the framework matters. A policy known only to one router configuration is hard for another network to validate. A policy published but never compiled into filters is only documentation. A filter compiled from stale data can reject valid routes or accept invalid ones. Accountability requires the chain from declared intent to deployed behavior to observed routes.

Default reject changes the failure mode

RFC 8212, published in 2017, updates BGP behavior so routes on an eBGP session are neither imported nor exported unless explicit policy has been configured. [11]

This is a deceptively important design choice.

A permissive default makes reachability easy during initial configuration. It also means a missing policy can silently become "accept everything" or "announce everything." An operator must remember to add every protective rule before the session carries routes.

A default-reject posture changes the failure mode. Missing policy produces no route exchange, which is visible and local, rather than unintended global propagation. Operators still can write an incorrect explicit policy. RFC 8212 says so. The control does not solve semantic mistakes, stale filters or intentional exceptions.

It does, however, encode a sound accountability principle: global reachability should require an affirmative policy decision.

For a customer-transit session, that decision should be reviewable:

  • which prefixes may be accepted;
  • which origins and customer paths are expected;
  • which routes may be exported back;
  • how exceptions are approved;
  • what happens when policy data is unavailable;
  • which system owns rollback;
  • what evidence proves deployment.

Had every relevant eBGP edge used a strict default with correct explicit policy, a missing filter would have failed closed. The public evidence cannot show whether RFC 8212-like behavior would have prevented this exact incident because it does not expose the actual 2015 configurations. The RFC remains a useful retrospective test: did route exchange require explicit, bounded authority at both sides?

BGP Roles and Only to Customer address relationship information

RFC 9234, published in 2022, standardizes BGP Roles and the Only to Customer attribute. Neighbors can negotiate roles such as provider, customer, peer, route server and route-server client. Propagated routes can carry information that helps enforce expected relationship direction and detect leaks. [12]

This mechanism targets the gap visible in the AS4788 event. A legitimate origin and loop-free path do not reveal whether a route learned from a peer may be sent to a provider. Relationship information makes that policy more explicit in the protocol exchange.

The standard still depends on correct configuration and deployment. Networks must assign roles accurately. Complex relationships require care. Partial adoption limits protection. Legacy routes and equipment remain. No protocol feature eliminates the need for monitoring and operational coordination.

The value is that both sides can compare expectations. A unilateral local label can be wrong without immediate feedback. A negotiated role can fail session establishment or mark a path when the two ends disagree. The Only to Customer attribute can help identify routes that should not travel to another provider or peer.

Again, this is later guidance. It would be historically inaccurate to say AS4788 or AS3549 failed to use a 2022 standard in 2015.

The incident instead supplies the test case:

  • Can a route learned from a peer be exported to an upstream without a detectable policy violation?
  • Can the upstream identify that the customer's path includes routes outside the expected customer relationship?
  • Can either side stop the route before global propagation?
  • Does the evidence distinguish a policy exception from an accidental leak?

Modern role-aware mechanisms should be evaluated against a reconstructed AS4788-like route set, not only against synthetic examples that match a clean topology.

Monitoring must compare routes with intent, not just availability

Availability monitoring detects the harm after users begin to lose reachability. Route monitoring can identify the control-plane anomaly earlier.

The 2015 public record was preserved by several forms of observation:

  • BGPMon processed update streams and identified the announcement burst;
  • RouteViews and RIPE RIS retained raw BGP archives;
  • ThousandEyes combined route and network measurements;
  • RIPE Atlas supplied active end-to-end measurements;
  • independent analysts compared paths, prefix counts and timing. [2][3][4][5][6]

These systems saw different slices. That diversity is a strength. A single provider's internal view can miss how its routes appear elsewhere. An active probe can see packet loss but not the policy that caused it. A route collector can see an AS path but not every traffic path or private session.

Modern detection systems can look for route leaks using topology, AS relationships, route history and abnormal propagation. Cloudflare describes public route-leak detection as a way to surface anomalous paths, while RIPE Atlas documentation supports reproducible measurement from distributed probes. [19][20]

Detection should be tied to action.

An alert that says "route count increased" is weak if no one owns the threshold or can suppress the route. A useful incident path defines:

  1. the expected relationship and route set;
  2. the anomaly condition;
  3. confidence and false-positive handling;
  4. the operator authorized to quarantine;
  5. a safe containment action;
  6. external confirmation;
  7. evidence retention;
  8. post-event review.

The first response does not always need to drop the entire session. A provider can lower preference, quarantine unexpected routes, preserve the last known-good accepted set, or reject only paths outside the customer cone. The right action depends on router capability and customer design.

Monitoring also should distinguish prevention from detection. Publishing a route-leak alert after global propagation is valuable public evidence. It does not prove that the provider had a pre-propagation control. Accountability reports should say which stage detected the event and which stage stopped it.

A safe route-policy change needs current-byte evidence

Routing configuration often passes through templates, databases, automation, policy compilers and vendor-specific syntax before it reaches a router. A human reviewer can approve one representation while the device receives another.

The evidence chain should bind the current bytes at each stage:

  • source policy or change request;
  • normalized relationship and prefix data;
  • generated route map or policy language;
  • device-specific configuration;
  • candidate configuration diff;
  • committed configuration hash;
  • resulting advertised and accepted route set;
  • external collector observation.

This matters because "the policy was reviewed" is ambiguous. Which version was reviewed? Did an automation job expand an AS set after approval? Did a stale registry snapshot produce the filter? Did a manual emergency command bypass the normal pipeline? Did all routers receive the same output?

An accountable change system should fail if those bindings diverge.

Before deployment, it should replay representative routes through the compiled policy. For AS4788-like conditions, tests should include:

  • customer-originated routes;
  • customer-cone routes;
  • peer-learned routes;
  • provider-learned routes;
  • routes containing large transit networks;
  • unexpected more specifics;
  • a sudden full-table-scale input;
  • mixed valid and invalid announcements.

The test should assert both positive and negative behavior. Valid customer routes must continue to pass. Peer- and provider-learned routes must not escape toward an upstream. The accepting provider should independently reject paths inconsistent with the customer's expected role.

After deployment, route collectors or looking glasses should verify the observable outcome. A configuration hash alone does not prove the router advertised only the intended routes. Control-plane state, device bugs and interaction with other policy can change effective behavior.

This current-byte discipline is not bureaucracy for its own sake. It is how an organization proves that the code, policy and routes being discussed are the same entities that produced or prevented harm.

Restoration is not the same as verified repair

The public sources show that routes were withdrawn or stopped being accepted and service recovered over the following hours. That is operational restoration.

Repair asks a harder question: could the same class of route escape again?

A credible remediation program would freeze a representative incident set from RouteViews, RIPE RIS and internal logs. It would identify the intended relationship for each route and reproduce the export and import decisions in a test environment.

For AS4788, the test would verify that routes learned from peers or providers cannot be selected for export toward AS3549 unless an explicit, reviewed exception applies. For AS3549, it would verify that a customer cannot announce paths outside the expected customer cone or exceed a justified volume without quarantine.

The program would then generate evidence:

  • failed tests before the fix;
  • policy or system changes;
  • passing tests after the fix;
  • device and software versions;
  • deployment coverage;
  • alert and containment exercises;
  • external route observations;
  • exception inventory and expiry;
  • ownership of continued monitoring.

The repair should also test degraded conditions. What happens if registry data is unavailable? Does the system fail closed, use a last known-good set, or accept everything? What happens if the anomaly detector is down? Can an operator isolate the session through an independent management path? Does a maximum-prefix stop preserve critical customer routes or drop them all?

Public disclosure need not expose private commercial terms or exploitable configuration. It can state the failure class, affected policy boundary, controls added, test method, deployment coverage and verification date.

Without that evidence, "we fixed the filter" is a claim about intent. With it, customers and peers can evaluate whether the operator changed the system that allowed global propagation.

Accountability should not collapse into personal blame

An Internet route leak often becomes a story about one engineer entering a bad command. The public record here does not establish that story. Even if a single action triggered the event, the global impact required multiple systems and organizational decisions.

An operator designs the configuration interface. It chooses whether changes are generated or handwritten. It defines peer and provider relationships. It decides which tests are mandatory, whether a second reviewer is required, how quickly policy propagates, and whether rollback is independent.

A transit provider decides how much trust to place in a customer announcement, which filters are economically and operationally feasible, and what anomaly will trigger containment. Leadership decides whether routing-security work has staffing, maintenance windows and authority to interrupt revenue traffic.

Personal blame can obscure these controls. It can also discourage disclosure. A better accountability model asks:

  • Who had the capability to prevent the route from leaving?
  • Who had the capability to reject it?
  • Who had the capability to limit its propagation?
  • Who could detect the harm independently?
  • Who could withdraw or quarantine?
  • Who retained evidence?
  • Who had authority to fund and verify remediation?

These questions can identify responsibility without claiming intent or negligence that public sources do not prove.

They also prevent responsibility from dissolving into "the Internet is decentralized." Decentralization means no single operator controls every path. It does not mean each operator lacks control over its own announcements, sessions and propagation decisions.

What customers and peers can reasonably demand

Most customers cannot audit a transit provider's routers. Peers cannot see every private change process. They can still demand evidence appropriate to the dependency.

Before an incident, an operator can publish:

  • accurate routing contacts;
  • registered prefixes and autonomous systems;
  • route and AS sets;
  • a high-level peering and filtering policy;
  • RPKI coverage;
  • support for relevant role and validation mechanisms;
  • status and incident channels.

During an incident, it can communicate:

  • the affected route or session class;
  • whether announcements are still propagating;
  • containment action;
  • known regions and services;
  • measurement uncertainty;
  • recovery evidence;
  • next update time.

After an incident, it can provide:

  • source and acceptance boundaries;
  • route-count definitions;
  • timeline with provenance;
  • controls that failed;
  • controls that contained harm;
  • testable remediation;
  • remaining limitations.

Customers should also test their own exposure. Multi-homing does not guarantee independence if both providers depend on the same upstream. A backup route can exist but lose under local preference. More-specific announcements can override intended diversity. Traffic can avoid a leaked path yet suffer congestion in a shared transit provider.

Independent route monitoring, RIPE Atlas measurements and looking-glass checks can reveal part of that exposure. [4][5][6][20]

The duty is proportional. A critical public service or financial platform should understand upstream concentration more deeply than a low-impact personal site. But no customer can compensate fully for a transit provider accepting and spreading a massive unsafe route set.

What the public record cannot prove

The source set supports a strong network-accountability analysis, but it does not support a complete internal postmortem.

It cannot prove:

  • the exact Telekom Malaysia router or location;
  • the exact configuration command or template;
  • whether the trigger was a planned change, stale state, automation fault or manual error;
  • the software or hardware version;
  • the private relationship terms between AS4788 and AS3549;
  • the exact import, export and maximum-prefix settings on either side;
  • the first internal alert and operator response;
  • private coordination messages;
  • one reconciled prefix count across all collectors;
  • a complete census of affected users or financial losses;
  • legal liability or contractual breach;
  • the durable remediation deployed by either operator.

The packet also should not turn later controls into historical requirements. RFC 8212 was published in 2017, RFC 9234 in 2022, and the current MANRS implementation guide reflects later operational work. They define useful present-day tests. They do not prove what configurations or obligations existed in 2015. [11][12][18]

Likewise, current RIPEstat data is current network-resource context, not a frozen 2015 registry snapshot. [7]

These boundaries make the conclusion more credible. The observable failure is enough to identify divided control. The missing internal record is itself an accountability gap, but it is not permission to invent one.

A reusable upstream-filtering accountability test

The event supports a practical test for any customer, transit provider or peer operating BGP at meaningful scale.

1. Define the relationship for each session.
Record provider, customer, peer, route-server and exceptional roles at the granularity where policy differs.

2. Bind intended routes to authenticated evidence.
Maintain prefixes, origins, customer cones, AS sets and exceptions with ownership, freshness and provenance.

3. Compile policy before deployment.
Show the concrete routes and paths that import and export policy will accept. Review effective behavior, not only template text.

4. Fail closed when explicit policy is missing.
No eBGP route should gain global authority because a filter was absent or data retrieval failed.

5. Test relationship violations.
Replay peer-learned and provider-learned routes against customer and upstream sessions. Verify that invalid direction is rejected at both exporter and receiver.

6. Calibrate volume controls.
Set maximum-prefix and anomaly thresholds against normal behavior, justified growth and emergency cases. Define safe containment rather than relying only on full session shutdown.

7. Separate origin and path validation.
Use RPKI for origin authority, but do not describe ROV as proof of relationship-valid propagation. Add path and customer-cone controls.

8. Monitor from outside.
Use independent collectors and active measurements to compare observed routes and reachability with intended policy.

9. Give containment an owner.
Identify who can quarantine routes, lower preference, restore a last known-good set or reset a session, including outside ordinary change windows.

10. Preserve current-byte evidence.
Bind approved policy, generated configuration, deployed bytes, route state, alerts, decisions and external observations.

11. Prove repair with the original event class.
Run the reconstructed leak through both sides of the session and show where it stops. Test semantic variants, not only one saved prefix list.

12. Publish enough for dependent networks to verify.
Explain the control boundary, route-count definitions, remediation and remaining uncertainty without exposing sensitive private terms.

This test does not promise that route leaks disappear. It makes prevention, containment and evidence duties explicit at every network that can grant the route further authority.

Conclusion

The 12 June 2015 Telekom Malaysia route leak demonstrated how quickly local routing policy can become global infrastructure harm.

AS4788 emitted a very large set of routes. AS3549 accepted and propagated them. Traffic shifted onto paths through Level 3 and Telekom Malaysia. Packet loss, latency and reachability failures spread across regions and affected both directly rerouted services and users exposed to congestion in a shared transit network. Independent route collectors and measurement platforms preserved the public outline. [1][2][3][4]

The event cannot be explained responsibly as one bad announcement by one network. Export and import are separate controls. A customer has a duty to advertise only authorized routes. A transit provider has a duty proportionate to its power to accept and spread those routes. Peers and downstream networks have additional monitoring and import controls. No layer can guarantee perfection, but each can prevent one mistake from acquiring more reach.

RPKI origin validation is valuable and limited public evidence for this path-policy failure. Route registries can publish intent and still become stale. Maximum-prefix controls can contain volume and still miss smaller leaks. Later default-reject and relationship-aware standards improve the control model but do not retroactively prove a 2015 violation. The durable answer is layered: explicit policy, authenticated route data, relationship checks, calibrated limits, independent monitoring, rapid containment and reproducible repair.

Risk follows the reach an announcement can acquire. Accountability follows who could have constrained that reach, who chose to propagate it, and who can prove that the same failure class will now stop before other people's traffic becomes the test.

Sources

  1. https://labs.ripe.net/author/gih/more-leaky-routes/
  2. https://www.bgpmon.net/massive-route-leak-cause-internet-slowdown/
  3. https://www.thousandeyes.com/blog/route-leak-causes-global-outage-level-3-network
  4. https://labs.ripe.net/author/emileaben/does-the-internet-route-around-damage-a-case-study-using-ripe-atlas/
  5. https://archive.routeviews.org/bgpdata/2015.06/UPDATES/
  6. https://data.ris.ripe.net/rrc00/2015.06/
  7. https://stat.ripe.net/AS4788
  8. https://www.rfc-editor.org/rfc/rfc4271.html
  9. https://www.rfc-editor.org/rfc/rfc7454.html
  10. https://www.rfc-editor.org/info/rfc7908
  11. https://www.rfc-editor.org/rfc/rfc8212.html
  12. https://www.rfc-editor.org/rfc/rfc9234.html
  13. https://www.rfc-editor.org/rfc/rfc2622.html
  14. https://www.rfc-editor.org/rfc/rfc4012.html
  15. https://www.rfc-editor.org/rfc/rfc6480.html
  16. https://www.rfc-editor.org/info/rfc6811
  17. https://manrs.org/netops/
  18. https://manrs.org/specifications/MANRS-007/01/
  19. https://blog.cloudflare.com/route-leak-detection-with-cloudflare-radar/
  20. https://atlas.ripe.net/docs/