Summary

  • AMS-IX reported two periods of platform instability on 22 and 23 November 2023. Customers connected to the Amsterdam peering infrastructure experienced flapping Link Aggregation Control Protocol, or LACP, and Border Gateway Protocol sessions. At the low point, platform traffic fell to 2.1 Tb/s. AMS-IX also reported IPv4 BGP sessions falling from 885 to 550 and IPv6 sessions from 800 to 450.[1] Those figures establish a severe exchange-level event. They do not establish that every connected network, application, or European user failed.

  • The initiating condition was not described as an ordinary physical link break. AMS-IX said LACP packets generated by customer equipment entered a Juniper provider-edge switch through a non-LACP connection and were propagated beyond the adjacent relationship where those packets had meaning. Other customers' link-aggregation groups reacted, their LACP and BGP sessions flapped, resources and buffers came under pressure, RSVP timeouts followed, and Path Error messages introduced further trouble on Extreme SLX switches.[1]

  • Accountability cannot stop at the device that emitted the first packet. The customer controlled the frame source. AMS-IX controlled the shared fabric, port policy, provisioning logic, ACL generation, monitoring, incident response, and public evidence. Juniper and Extreme controlled implementation behavior, subject to facts not yet public. Connected networks controlled whether they had alternate transit, remote peering, spare capacity, and authority to withdraw sessions. End users controlled none of those layers.

  • AMS-IX announced several repairs: applying LACP ACLs to non-LACP links, enhancing ACL creation in its provisioning stack, reviewing outbound LACP ACL behavior on both switch families, investigating Slow Protocol BPDU alerts, and revising technical-list communications.[1] Those steps align with the failure chain. They remain announced controls rather than independently verified proof that every relevant port, software release, and failover state now enforces the intended invariant.

The bounded event and the accountability question

This analysis is limited to the AMS-IX Amsterdam platform incidents on 22 and 23 November 2023. It does not merge them with the exchange's May 2015 switching-loop outage, unrelated failures at other Internet exchanges, or generic debates about whether Internet exchange points should use Layer 2, MPLS, or another architecture. Those comparisons can clarify resilience, but they cannot replace the event-specific evidence.

AMS-IX placed the first affected interval between 19:08 and 23:04 CET on 22 November. A second interval ran from 09:38 to 10:25 CET on 23 November.[1] The official report described active flapping of both LACP and BGP sessions. It recorded platform traffic falling to 2.1 Tb/s at the peak of disruption, with large reductions in visible IPv4 and IPv6 BGP sessions. Connected operators published their own, narrower observations.

Total Uptime said it shifted traffic away from the exchange and later restored it after observing stability.[9] NFOrce said it disabled direct-peer and route-server sessions to limit packet loss.[10] EDPnet reported intermittent reachability, alternate capacity, and recovery updates.[11]

The central question is not merely which packet arrived first. It is which actors had practical control over isolation, amplification, detection, mitigation, communication, and proof of repair. A shared peering platform exists to let many networks exchange traffic without treating every entity as part of one uncontrolled Ethernet segment. Its operator therefore carries a specific operational duty: control frames that belong to one adjacency should not be able to alter another customer's link state merely because the ports share switching infrastructure.

That duty is technical before it is legal. The public record does not establish negligence, breach of contract, or statutory liability. It does establish that AMS-IX designed and operated the boundary where a local control message acquired a wider effect. It also establishes that the operator had mitigation rules in place, but those rules did not cover or work across every relevant path. The accountability test is whether the intended boundary existed in running code and whether the repair can be demonstrated rather than only described.

This framing avoids two opposite errors. The first is to blame the unnamed customer for everything because its equipment generated the initiating packets. That would ignore the exchange's control over which frames can cross the shared fabric. The second is to treat the entire Internet as one centrally governed service for which the exchange alone guarantees every downstream outcome. Connected networks make their own routing, capacity, and resilience choices. The event crossed several control domains, and responsibility follows those domains.

The evidence supports a causal sequence with uncertainty preserved. A customer device emitted LACP packets. Those packets arrived on a port that AMS-IX characterized as non-LACP. A Juniper switch propagated the packets. Other LAGs reacted and flapped. BGP sessions riding on the affected links flapped. Resource starvation and full buffers contributed to RSVP timeouts. Path Error messages from affected Juniper provider edges caused further problems on Extreme SLX equipment. Operators routed around or withdrew from the exchange where they could.

AMS-IX isolated the problem and changed controls.[1] The exact commands, versions, packet captures, ACL rules, and vendor root-cause records are not public.

What the impact numbers prove, and what they do not

The traffic low of 2.1 Tb/s is the most visible measure of the incident. Contemporary technical commentary compared it with the exchange's normal multi-terabit load and described a loss or displacement of roughly eight terabits per second.[7][12] That is an extraordinary movement of traffic. It is not automatically an equivalent amount of lost user traffic.

Traffic can disappear from an exchange graph for several reasons. A BGP session may go down and leave the traffic without a usable path. A network may intentionally disable the exchange and move traffic to transit or remote peering. A destination may become unreachable. An application may reduce sending because earlier connections failed. Congestion elsewhere may limit replacement traffic. The graph records what crossed the AMS-IX platform, not the fate of every packet that otherwise would have crossed it.

RIPE NCC examined the incident using RIPE Atlas measurements. Its analysis asked whether source-destination pairs that normally traversed AMS-IX remained connected, moved to another path, or failed.[2] That approach is more informative than a single exchange aggregate because it separates successful rerouting from reachability loss. It still samples a set of probes and destinations. It cannot measure every private peering path, every service, or every user's application experience.

The BGP-session counts are similarly specific. A fall from 885 to 550 IPv4 sessions and from 800 to 450 IPv6 sessions shows that many control-plane adjacencies on the platform were not stable.[1] It does not mean that 335 IPv4 networks and 350 IPv6 networks were wholly offline. One network can maintain multiple sessions. A session can flap without causing complete customer loss if another path remains. A network can also keep a session established while suffering packet loss or congestion.

Downstream status evidence gives the aggregate numbers operational meaning. Total Uptime reported that it moved traffic to alternate providers and characterized the issue as no longer customer-impacting for its service while AMS-IX was still investigating.[9] NFOrce disabled exchange sessions to prevent further packet loss.[10] EDPnet said it had enough alternate bandwidth to serve its own customers, though it warned that Internet speed and stability could still be affected because other paths and networks were involved.[11]

These accounts show why an exchange outage has unequal consequences. A network with diverse transit, remote peering, automated steering, sufficient headroom, and on-call authority can escape quickly. A network that is heavily dependent on one exchange, has little spare capacity, or cannot make rapid routing changes may suffer longer. The exchange operator controls the common failure domain. Each connected network controls how much of its own service depends on that domain.

The impact record should therefore be stated in layers. Exchange traffic and session counts show platform severity. RIPE Atlas shows sampled end-to-end path outcomes. Operator notices show specific mitigation and customer observations. None alone measures total economic loss, emergency-service impact, enterprise disruption, or every user's downtime. Those remain evidence gaps, not an invitation to fill the record with a dramatic universal claim.

LACP is local by design

Link aggregation combines multiple physical links into one logical connection. It can increase capacity and preserve service when one member link fails. LACP coordinates which links belong to an aggregate and whether they are eligible to forward. The protocol's meaning is adjacency-bound: the systems at the two ends exchange control information about their shared bundle. IEEE 802.1AX is the authoritative standards family for link aggregation.[14]

AMS-IX's own documentation says it supports LACP on its connection types and warns that LACP can affect failover behavior. After a topology change, an LACP-enabled port can remain blocked until it receives a control frame, and that delay can cause BGP sessions to flap. AMS-IX recommends short LACP timers to reduce failover delay.[4] That guidance acknowledges a coupling between Layer-2 bundle state and Layer-3 routing sessions.

The 2023 event exposed a more fundamental boundary problem. According to AMS-IX, customer equipment sent LACP packets through a non-LACP connection. A Juniper provider-edge switch propagated those packets. Other customer links then treated the leaked frames as relevant to their own aggregation state.[1] A control message meant for one adjacent relationship had gained authority over another.

That is why the phrase "LACP leakage" matters. Ordinary customer data is expected to cross an exchange toward other entities according to forwarding state. Link-local control frames are different. Their scope is part of their safety model. If a shared fabric forwards them as though they were ordinary customer traffic, remote links can respond to an identity or bundle state that does not belong to their adjacency.

The safest operational invariant is simple to state: Slow Protocol frames from one customer-facing port must not change another customer's link-aggregation state. Enforcing it is less simple. The rule has to apply to ports that use dynamic LACP, static LAGs, and no aggregation. It has to survive provisioning changes, switch replacement, software upgrades, syntax changes, vendor differences, and topology failover. It must work in both ingress and egress directions where required by the architecture.

AMS-IX said mitigation ACLs existed on LACP links, but the packet source was connected through a non-LACP link.[1] That detail turns the incident from a single ACL failure into a completeness failure. A control was tied to the feature configuration where engineers expected the protocol. The unsafe traffic appeared on a port where the feature was not expected. Security boundaries must cover the traffic class, not only the configuration label that predicts it.

The lesson resembles input validation at any trust boundary. A system should not decide whether to validate a dangerous message only after assuming the sender will use the documented path. The unexpected path is precisely where validation matters. For an exchange, the port role and allowed frame set should be explicit. Provisioning should generate the filter. Independent conformance checks should read the deployed state. Packet tests should verify that prohibited frame types cannot exit toward other entities.

This is a stronger claim than saying a customer should configure its equipment correctly. Customer configuration matters. An operator can require entities to send only permitted traffic and can disconnect a source that violates the rules. But a shared platform is built on the assumption that one entity's mistake will not automatically become every entity's link-state instruction. Admission control and isolation exist because entity behavior cannot be assumed perfect.

The non-LACP port exposed a provisioning blind spot

The official postmortem says the LACP-generating customer equipment was connected to a non-LACP link. It also says the Juniper outbound LACP ACL was not fully operational.[1] Together, those statements identify a gap between intended policy and deployed coverage.

An intent-based provisioning system might represent a customer port with properties such as speed, VLAN, LAG mode, member interfaces, MAC limits, allowed ethertypes, and operational state. If ACL generation is conditional on lacp=true, it may protect dynamic bundles while leaving static or ordinary ports able to pass LACP frames. That logic is understandable as a feature configuration. It is unsafe as an isolation invariant.

AMS-IX announced that it enhanced ACL creation in the provisioning stack so newly created links would receive the relevant filter.[1] This is an important repair because it shifts control from an operator remembering a special case to a system generating a baseline. It also raises verification questions. Did the change update only newly created links, or did it reconcile existing ports? How did the operator prove coverage across every Juniper and Extreme port profile? Did it test inactive, quarantined, migrated, and failover states? Is a later software upgrade able to change ACL syntax without failing deployment?

Provisioning systems can create false confidence when they record intent but do not verify device state. A database may say a filter is attached. The switch may reject part of the syntax, place the rule in the wrong direction, reorder it behind a permit statement, or interpret a protocol match differently after an upgrade. An accountability-grade control needs three forms of evidence: intended policy, rendered configuration, and observed forwarding behavior.

The first is the declarative record: no customer port may transmit LACP control frames to another entity. The second is a device-specific proof: every relevant port has an effective rule matching the required destination MAC, ethertype, or slow-protocol semantics in the correct direction. The third is a packet test: a synthetic forbidden frame introduced in a safe test or quarantine environment does not appear on any other customer-facing port.

AMS-IX's quarantine VLAN documentation is relevant because it shows the operator already uses a separated environment to observe new customer ports before production.[4] A quarantine process can test MAC behavior, broadcast traffic, and prohibited protocols. It should not be treated as a one-time certification. Configuration changes, port migrations, switch replacements, and software upgrades can alter the same invariant after a port leaves quarantine.

The November incident also demonstrates why control coverage should be measured continuously. A conformance service could compare intended and deployed ACLs, alert on missing slow-protocol filters, and sample switch counters for blocked frames. A separate observation path could watch for Slow Protocol BPDUs on the peering fabric. AMS-IX said it was investigating such alerts.[1] The useful metric is not only whether an alert exists but whether it detects one leaked frame before other customers' LAGs react.

No public evidence shows whether a pre-incident conformance check existed, whether it missed the port, or whether the ACL was present but ineffective. The article therefore cannot claim a specific process failure beyond what AMS-IX disclosed. It can identify the proof that would resolve the question: configuration snapshots, provisioning templates, device commit results, test history, ACL counters, and packet captures.

Cross-vendor behavior turned a local leak into a cascade

AMS-IX operated Juniper provider-edge switches and Extreme SLX equipment in the affected fabric. The official report describes distinct behavior on both families. The Juniper switch propagated LACP packets from customer equipment. The Juniper outbound LACP ACL was not fully operational. The Extreme SLX outbound ACL did not work as anticipated, although AMS-IX said it had worked in the past. Engineers had not yet determined whether the difference reflected a bug or a syntax change after an SLX software upgrade.[1]

Those statements support scrutiny of cross-vendor conformance. They do not support assigning the entire outage to either vendor. The public report does not identify model numbers, operating-system releases, configuration snippets, defect identifiers, vendor advisories, or laboratory reproductions. Without that evidence, "vendor bug" is a hypothesis, not a finding.

The distinction matters because the same observed result can come from different failure classes. A switch could forward a reserved frame type by design under a particular bridging configuration. An ACL could fail to match because of syntax, direction, hardware offload, rule order, or platform limitations. A software upgrade could change parser behavior. A provisioning template could render a rule valid on one family and ineffective on another. A human could assume that a previously tested invariant remained true after a change.

Multi-vendor design can reduce common-mode dependence, but only if equivalent safety properties are tested across implementations. Diversity alone is not resilience. If two platforms react differently to the same control-plane stress, the interaction can create a new common failure domain. AMS-IX's postmortem describes exactly that concern: flapping and resource pressure on one side produced RSVP Path Error behavior that introduced further problems on the other.[1]

A strong cross-vendor verification program starts with properties rather than commands. Examples include: LACP frames cannot cross unrelated customer boundaries; loss of one LAG member does not flap BGP if another member remains usable; full buffers do not starve control-plane processing needed for recovery; malformed or repeated Path Error messages are rate-limited; and one device's failure notification cannot destabilize the entire alternate platform.

Those properties should be tested after software upgrades and provisioning changes. The test should include failure injection, not only steady-state forwarding. A rule that works under normal load may fail when buffers are full or the control plane is processing rapid state changes. A switch that blocks a frame on an ordinary port may handle it differently on a static LAG, dynamic LAG, virtual chassis, or failover path.

Vendor responsibility depends on evidence that is not public. If an implementation violated a documented behavior under a supported configuration, the vendor may control the defect and fix. If the operator used unsupported syntax or failed to apply a required filter, operational control lies elsewhere. If documentation was ambiguous or changed across releases, responsibility may be shared. The defensible conclusion is that AMS-IX, as integrator and operator of the shared fabric, controlled the acceptance test that should make those distinctions before production.

From LAG flapping to BGP session loss

LACP controls whether physical links participate in a logical aggregate. When aggregate membership changes repeatedly, the logical interface can lose forwarding continuity. BGP sessions carried over that interface may then reset or experience enough loss to time out. AMS-IX's documentation explicitly notes that LACP behavior after topology failover can flap BGP sessions.[4]

The 2023 incident was more severe than a single clean failover. Other customers' LAGs reacted to leaked frames, producing repeated state changes. BGP sessions dropped from the platform in large numbers.[1] Every session reset can withdraw routes, trigger best-path calculations, cause alternate advertisements, and shift traffic to other interconnections. Repeated resets create churn rather than one orderly convergence.

That churn affects both the exchange and its entities. The exchange fabric carries traffic among many autonomous systems. A entity whose session fails may move traffic to transit, another exchange, private interconnection, or remote peering. The alternate path may have higher latency or less capacity. The entity's neighbors may receive route withdrawals and replacements. Application sessions can break even if routing later converges.

RFC 4271 defines BGP's session and routing information exchange.[15] It does not promise that every network has an alternate path or enough capacity. BGP can select a route around a failed exchange when another route exists and policy permits it. The operational phrase "the Internet routes around damage" is therefore conditional. It routes around damage for a particular source and destination when topology, policy, capacity, and convergence make a usable alternative available.

RIPE Atlas measured both successful avoidance and failed connectivity during AMS-IX incidents.[2][20] Earlier research on the 2015 AMS-IX outage found that some paths moved to transit or other peering locations, some suffered failures, and control-plane restoration could outlast the short initiating event.[21] That comparison should not be used to claim the 2015 and 2023 causes were identical. It shows that exchange restoration and path restoration are separate measurements.

The relevant responsibility split is practical. AMS-IX controlled the stability of the common peering platform. Entities controlled whether they had alternate paths and how they reacted to session instability. Transit providers controlled the capacity and policy of escape paths. No single party controlled the entire end-to-end route, but each controlled a point where harm could be limited.

This makes BGP-session count a useful but incomplete recovery metric. AMS-IX should be able to show when LACP state stabilized, when BGP sessions stopped flapping, when the session count returned, and when platform traffic normalized. Entities should be able to show when their alternate paths became active, whether they congested, and when they restored exchange sessions. Independent measurements should test whether end-to-end reachability and latency returned, not assume them from the exchange graph.

RSVP errors were an amplifier, not the trigger

AMS-IX said LACP and BGP flapping led to resource starvation and full buffers. It then observed RSVP timeout errors. Affected Juniper provider-edge devices aggressively sent RSVP Path Error messages, which introduced further issues on Extreme SLX switches.[1] This part of the sequence is important because it shows failure crossing protocol layers and vendor boundaries.

RSVP-TE is used to signal traffic-engineered paths in MPLS networks. RFC 3209 defines RSVP extensions for establishing and maintaining label-switched paths, including Path and error processing.[16] RFC 4090 describes fast reroute mechanisms intended to protect MPLS traffic-engineering tunnels from link and node failures.[17] These protocols are recovery and path-control tools. Under resource pressure, their messages and state transitions can themselves become part of an amplification loop.

The public report does not say RSVP initiated the incident. LACP leakage and resulting flapping came first. It also does not establish that every RSVP implementation behaves the same way or that MPLS/VPLS architecture is inherently unsafe. RFC 4761 and related standards provide context for BGP-signaled VPLS and provider-edge tunnels.[18][19] The event-specific finding is narrower: the deployed interaction between flapping, buffers, Juniper Path Error behavior, and Extreme SLX processing expanded the disturbance.

Resource starvation can change the behavior of controls that appear independent at normal load. A router needs CPU and buffer capacity to process keepalives, link-state events, RSVP messages, BGP updates, management traffic, and telemetry. If repeated LAG state changes generate a storm of work, the control plane may delay the very messages needed to restore stability. Full buffers can cause loss, retries, and additional error signaling.

An accountability-grade postmortem would quantify that cascade. How many LACP state transitions occurred? How many BGP sessions reset, and how often? Which buffers filled? What RSVP timer expired first? How many Path Error messages were sent, at what rate, and to which devices? Did rate limits exist? Which Extreme resources became constrained? What change stopped the cycle?

AMS-IX's public summary gives the sequence but not those measurements. That is enough to identify the control domains, not enough to reconstruct every packet. Vendor blame should remain bounded until configurations, versions, captures, and defect findings are available.

The repair implication is cross-layer testing. A test that proves an ACL blocks one packet in steady state does not prove the fabric survives a burst of prohibited frames, repeated LAG transitions, BGP churn, buffer pressure, and RSVP errors. The invariant should be tested under stress, with telemetry showing that control-plane queues remain available and error messages cannot become a new denial mechanism.

Alternate paths limited harm but transferred load

Several connected networks responded by moving traffic away from AMS-IX. That was rational containment. It also transferred load to other interconnections. Resilience depends on whether those alternatives are genuinely independent and sufficiently provisioned.

Total Uptime said it shifted traffic to alternate providers and later restored service through AMS-IX after observing stability.[9] NFOrce disabled direct-peer and route-server sessions to prevent further packet loss.[10] EDPnet said it maintained enough bandwidth for its own customers without AMS-IX, while warning that wider Internet speed and stability could still be affected.[11] These statements illustrate three controls: path diversity, capacity headroom, and operational authority to change routing.

Path diversity on paper is not enough. Two logical services may share the same exchange, facility, fiber, switch family, route server, or control system. A backup transit session may exist but carry too little capacity for emergency load. A remote peering service may still traverse the affected platform. An automated policy may prefer a flapping route and repeatedly pull traffic back before the incident is stable.

The economic trade-off is visible. Spare transit and diverse peering cost money. They may be lightly used most of the year. During a common exchange failure, that unused capacity becomes the difference between a routing change and a customer outage. Operators decide how much resilience to buy, and customers often cannot observe the decision until a failure.

This does not mean every connected network should engineer for the complete loss of every exchange regardless of cost. It means the risk should be explicit. A network should know the share of traffic dependent on one peering fabric, the capacity of alternate paths, the time required to move traffic, the services that cannot tolerate the shift, and the decision owner who can disable unstable sessions.

RIPE Atlas provides external evidence about whether paths actually moved and whether destinations stayed reachable.[2] Flow records, BGP logs, and capacity graphs from each network would provide stronger local proof. An operator claiming successful resilience should show not only that BGP selected another path, but that packet loss, latency, and application success remained within defined bounds.

The shared exchange also has an interest in entities' escape paths. It cannot mandate every business model, but it can publish realistic failure scenarios, encourage remote and multi-exchange diversity, provide clean incident signals, and avoid communications that cause networks to oscillate between unstable and alternate paths. AMS-IX's planned revision to its technical-list communication policy recognizes that information cadence is itself a recovery control.[1]

Detection and communication are part of containment

The official report says AMS-IX planned to investigate alerts for Slow Protocol BPDUs on the platform.[1] Such an alert could detect the unsafe condition near its source. It is more specific than waiting for aggregate traffic to fall or BGP sessions to flap.

Detection should be layered. A port-level counter can record prohibited slow-protocol frames. A fabric observer can detect the same source control frame appearing on unrelated ports. LAG telemetry can identify synchronized state changes across customers. BGP session monitoring can detect a correlated fall. RSVP and buffer telemetry can warn that the incident is crossing into the MPLS control plane. External probes can show end-to-end harm.

Each signal answers a different question. The first prohibited frame is evidence of boundary failure. LAG flapping is evidence that another system acted on it. BGP resets show routing consequences. Traffic displacement shows operational severity. External failure shows customer harm. A response system should preserve the timestamps rather than compress them into one outage start.

Alert ownership matters as much as alert existence. Someone must be able to identify the source port, apply a filter, quarantine a link, disable an interface, or isolate a switch. The action must be safe under uncertainty. Disconnecting the wrong port or applying an untested global ACL could widen the outage.

Communication influences downstream decisions. Connected networks need to know whether the platform is stable, whether a mitigation is temporary, and when it is safe to restore sessions. Too little information delays containment. Premature reassurance can cause traffic to return to an unstable path. Overly frequent unstructured updates can create confusion.

AMS-IX said it would revise the tech-l communication policy, including rules for update frequency based on incident severity.[1] A useful policy should define message ownership, minimum fields, evidence confidence, next-update time, and a distinction between mitigation, platform stabilization, monitoring, and final resolution. It should also explain whether customers should keep sessions down or may safely restore them.

The downstream records show why this matters. Operators independently disabled sessions and watched the exchange graph.[9][10][11] They were making risk decisions from partial evidence. A common machine-readable incident feed with stable timestamps and state definitions could reduce inconsistent restoration.

The Heng.lu reality layer: records did not enforce the boundary

An Internet exchange maintains detailed records: entity identities, AS numbers, ports, VLANs, link-aggregation mode, switch location, configuration intent, and operational contacts. Those records are necessary accountability infrastructure. They identify who controlled each part of the system and what policy was supposed to apply.

They did not stop the 2023 incident. The relevant reality was the packet accepted by the switch, the ACL behavior on the actual port, the frame propagated to other links, the LAG state chosen by connected devices, the BGP sessions that remained established, and the alternate paths that carried traffic. This is running-code primacy in a direct network-control case.

The doctrine does not imply that records are irrelevant. Without accurate port and entity records, AMS-IX could not identify the source boundary, generate policy, contact operators, or prove coverage after repair. The point is narrower: a record that says a port is non-LACP does not make LACP packets harmless. In fact, that label appears to have contributed to the coverage gap if the safety ACL was associated only with expected LACP use.

The correct role of the ledger is to drive and audit enforcement. The source-of-truth should generate a deny rule for unsafe frame classes on every customer boundary. Deployment receipts should prove the rule reached the device. Readback should confirm effective state. Packet tests should confirm behavior. Incident evidence should link the failed port and rule to the same records without pretending that the database itself controlled the wire.

This is also why accountability should not be reduced to institutional authority. AMS-IX could publish rules and customer requirements, but permission language did not contain the frame. Vendor documentation could describe behavior, but the deployed cross-vendor system produced the actual cascade. A entity could promise correct configuration, but the exchange still needed isolation.

The reality layer therefore asks concrete questions. Which port received the frame? What policy should have blocked it? What configuration was rendered? What did the hardware actually do? Which links changed state? Which BGP sessions fell? Which alternate paths carried traffic? Which repair changed observed behavior? Those questions can be answered with logs, configurations, captures, measurements, and tests.

A responsibility map without speculative liability

The unnamed customer controlled its device and the LACP frames it generated. It may have had contractual duties to send only permitted traffic. The public record does not identify the customer, device, configuration, intent, or notice history. No conclusion about fault or liability can be drawn beyond control of the frame source.

AMS-IX controlled the shared peering fabric and the isolation boundary. It controlled port provisioning, ACL generation, switch integration, monitoring, incident response, communications, and the decision to apply mitigations. Its public report acknowledges that mitigation was incomplete or ineffective across relevant paths.[1] That supports operational accountability for containment and repair proof.

Juniper controlled the behavior and documentation of its software and hardware within supported configurations. Extreme controlled the same for SLX equipment. Whether either vendor supplied a defective implementation depends on versions, configurations, documented expectations, reproduction, and vendor findings that are not public.

Connected networks controlled their own route diversity, transit capacity, session policy, monitoring, and mitigation authority. Operators that moved traffic showed one form of resilience. That does not prove every packet or customer service remained healthy, and it does not erase the exchange failure.

Route-server operators, if involved in particular sessions, controlled route-server availability and policy but not entities' direct sessions or physical LAG state. The official public summary reports aggregate BGP effects but does not attribute the incident to route-server policy. The article should not invent that link.

End users and downstream organizations bore consequences without controlling the exchange fabric. They could retry applications or switch access providers only in limited cases. Assigning them prevention responsibility would confuse dependency with control.

Regulators and contractual counterparties may later examine availability claims, incident reporting, critical-service dependencies, or vendor obligations. The public sources used here do not establish a legal standard, breach, compensable loss, or enforcement finding. The evidence supports a control map and a set of verification requests.

What durable repair would look like

AMS-IX's announced actions map sensibly to the incident. Applying ACLs to non-LACP links addresses the unexpected ingress path. Enhancing provisioning logic reduces dependence on manual configuration. Reviewing outbound ACLs across Juniper and Extreme addresses cross-vendor enforcement. Slow Protocol alerts improve detection. Communication-policy changes support entity mitigation.[1]

Durable repair requires evidence at several levels.

First, define the invariant. Link-local LACP frames from one entity cannot leave that entity's permitted adjacency. The invariant applies to dynamic LAGs, static LAGs, ordinary ports, quarantine ports, migration states, and failover states.

Second, prove inventory completeness. Every customer-facing port in every relevant fabric should map to a policy profile. Unknown, orphaned, or manually configured ports should fail closed or generate an exception that is reviewed.

Third, prove configuration conformance. The provisioning system should render vendor-specific rules, verify commit success, read back effective state, and compare it with intent. A software upgrade that changes syntax should fail validation before production rather than silently weaken a filter.

Fourth, prove behavior. Safe packet-injection tests should show that LACP and other prohibited slow-protocol frames cannot cross the boundary. Tests should cover normal load and failover. Counters should show drops. Observers should confirm that no copy appears on other customer ports.

Fifth, test the cascade. Generate controlled LAG churn in a laboratory replica and measure BGP stability, buffer pressure, RSVP error rates, and cross-vendor behavior. Confirm that rate limits and control-plane protection preserve recovery capacity.

Sixth, test entity escape paths. Exchange members should measure whether alternate transit or remote peering can carry critical traffic under realistic load. They should record decision authority, automation thresholds, and restoration criteria.

Seventh, publish bounded evidence. A follow-up report need not expose customer secrets or exploitable configuration. It can state port coverage, test populations, software versions at a safe level of detail, vendor defect status, alert performance, and dates of conformance exercises. It should separate completed actions from planned work.

Eighth, preserve incident artifacts. Packet captures, configuration snapshots, ACL counters, BGP session logs, RSVP message rates, buffer telemetry, operator actions, and communication timestamps should share one timeline. That record allows later review without depending on memory or one aggregate graph.

The success metric is not the absence of another public incident for a few months. It is proof that the failed invariant now holds across current infrastructure and remains tested after change.

Evidence still missing

The official report is unusually specific for a shared-network incident, but major gaps remain.

The customer and device that generated LACP packets are not identified. That may be appropriate for confidentiality, but it limits independent understanding of the trigger. The exact port mode, packet contents, frame rate, and configuration are not public.

The Juniper and Extreme models, software releases, ACL syntax, rule order, hardware behavior, and upgrade history are not public. AMS-IX said the SLX ACL had worked previously and that it was unclear whether a bug or syntax change explained the difference.[1] A later vendor advisory or verified reproduction would materially change the attribution.

The public summary does not quantify LAG transitions, BGP reset rates, buffer occupancy, RSVP timeout counts, Path Error volume, or the precise mechanism by which Extreme devices were affected. Those measurements would distinguish resource exhaustion, protocol amplification, and implementation behavior.

Complete downstream impact is unknown. Exchange traffic, BGP sessions, Atlas probes, and operator notices provide strong samples. They do not measure every user, service, latency increase, packet-loss period, or economic loss.

The detection timeline is incomplete. The public record does not state whether a customer reported the first symptom, an internal platform alarm fired, an external measurement detected the event, or which signal identified the source port. It also does not give the decision sequence for mitigation.

Durable repair evidence is incomplete. AMS-IX listed follow-up actions, but public sources do not prove coverage across every port, successful cross-vendor regression tests, alert exercises, or the absence of later configuration drift.

These are not reasons to withhold all conclusions. They define the boundary of the conclusions. The event demonstrates a failed isolation property and a multi-layer cascade. It does not, on the current public record, prove a particular vendor defect, customer misconduct, negligence, legal breach, or exact loss.

Conclusion

The AMS-IX incidents of November 2023 began with a small-scope control message and became a large-scope infrastructure failure. LACP packets that should have mattered only to adjacent systems escaped one customer boundary. Other aggregated links reacted. BGP sessions flapped. Buffers and control-plane resources came under pressure. RSVP errors and cross-vendor behavior widened the disturbance. Connected networks moved traffic where they had usable alternatives.

The event is therefore not best understood as "a customer sent a bad packet" or "the exchange went down." It is an accountability chain. The customer controlled the source. AMS-IX controlled isolation, provisioning, monitoring, and incident response. Vendors controlled implementation behavior within supported configurations. Entities controlled alternate paths and capacity. Each layer could prevent or limit a different part of the harm.

The most important repair is an enforceable invariant: link-local control frames from one entity cannot change another entity's link state. The record of a port's intended mode is necessary, but running code must enforce the rule on every port. Provisioning must generate it, device readback must confirm it, packet tests must prove it, and upgrade tests must preserve it.

AMS-IX's announced actions point in that direction. Public accountability now depends on verification. The exchange and its vendors can close the evidence gap by showing that all current ports are covered, cross-vendor behavior is tested under stress, alerts detect prohibited frames before a cascade, alternate paths are measured honestly, and incident communications distinguish mitigation from restoration.

The broader lesson is specific rather than ideological. Internet exchanges reduce cost and improve reachability by sharing infrastructure. Shared infrastructure is valuable because it connects entities; it is safe only when it also isolates them. In November 2023, the connection worked more broadly than the isolation. The durable test is whether the repaired platform can prove the opposite.

Sources

  1. AMS-IX, "Outage on Amsterdam peering platform"
  2. RIPE Labs, "Does the Internet Route Around Damage? - Edition 2023"
  3. RIPE 87 archive, AMS-IX engineering postmortem presentation
  4. AMS-IX technical documentation, link aggregation and platform controls
  5. AMS-IX configuration guide
  6. AMS-IX port security documentation
  7. AMS-IX MPLS/VPLS documentation
  8. AMS-IX total platform statistics
  9. Total Uptime, AMS-IX incident status
  10. NFOrce NOC, AMS-IX issues status
  11. EDPnet, AMS-IX outage and mitigation updates
  12. ipSpace.net, "AMS-IX Outage: Layer-2 Strikes Again"
  13. PAM 2024, "Following the Data Trail: An Analysis of IXP Dependencies"
  14. IEEE 802.1AX Link Aggregation standard
  15. RFC 4271, A Border Gateway Protocol 4
  16. RFC 3209, RSVP-TE Extensions to RSVP for LSP Tunnels
  17. RFC 4090, Fast Reroute Extensions to RSVP-TE for LSP Tunnels
  18. RFC 4761, Virtual Private LAN Service Using BGP
  19. RFC 8614, Updated Processing of Control Flags for BGP VPLS
  20. RIPE Labs, earlier AMS-IX route-around-damage case study
  21. SIGCOMM 2017, "Detecting Peering Infrastructure Outages in the Wild"