Summary
- LINX reported that a dark-fibre failure at LON2 on 20 June 2023 reduced resilience without immediately affecting members. The following day brought Flowmon drops, inter-switch-link flaps, traffic loss and reachability problems across the peering LAN. [1][2]
- The operator said a newly installed router had previously been used in a lab. Its old OSPF and router-ID configuration had been removed and a new identity deployed, but the running OSPF process retained the old router ID because the process had not been cleared. [1]
- LINX also reported a mismatch in which MAC addresses appeared in a software table but not the hardware forwarding table. Clearing an L2VPN EVPN BGP session repopulated the hardware table for residual reachability cases. [1]
- Later June and November episodes show why an operator should not compress every symptom into one root cause. LINX documented further reachability problems, dark-fibre maintenance, a software change, rollback and continuing vendor investigation. [1]
- EVPN, VXLAN and OSPF are not the defendants. The accountable boundary is the admission and verification system: unique live identity, cleared process state, architecture-representative testing, degraded-path exercises and reconciliation of control-plane records with hardware forwarding.
- The annual report recorded LON2 availability of 99.997 percent, below LINX's 99.998 percent internal target. That aggregate matters, but it does not identify every member, prefix, session, packet or consequence. [2]
- Responsibility is distributed but not vague. LINX controlled device admission, automation, monitoring, maintenance sequencing, isolation, session clearing and member communications. Vendors controlled parts of fibre repair and defect investigation. Members controlled their own peering and continuity choices.
- The final accountability standard is running-code evidence. Intended configuration, deployment success and an annual percentage are useful records; safe operation requires proof that live identities are unique, protocol state is current, forwarding tables agree and member reachability survives the relevant failure conditions.
A Failure That Began as Reduced Resilience
The public chronology begins with a failure that LINX said did not immediately affect members. On 20 June 2023, a dark-fibre connection at LON2 failed. LINX's November technology update described the immediate effect as reduced resilience rather than a member-impacting outage. That distinction is important. A redundant network can continue forwarding after one path fails, yet the system is no longer in the same risk state. Maintenance decisions, device activation and later failures now operate against a smaller margin.
On 21 June, LINX reported Flowmon drops and flapping inter-switch links. Traffic loss and reachability problems appeared across the LON2 peering LAN. Engineers disabled links as part of restoration work and later disabled a link aggregation between core devices to restore service more fully. On 22 June, the broad incident had narrowed, but a smaller set of specific IP reachability problems remained. LINX said MAC addresses were present in an edge device's software MAC table but absent from its hardware table. Clearing the L2VPN EVPN BGP session between affected virtual tunnel endpoints repopulated the hardware table. [1]
The chronology is not simply a story about a broken fibre. Nor is it defensibly reduced to a single bad router identifier. It is an interaction among physical-path resilience, inter-switch behavior, protocol identity, EVPN state and the relationship between software control records and hardware forwarding. Each layer can appear healthy when viewed through one interface. A fibre path may have an administrative up/down state. An automated deployment may report success. A software table may contain a learned address. A BGP session may be established.
Yet packets can still fail if those records do not describe the state used by the forwarding hardware.
This is why the first event belongs in the accountability analysis even though LINX said it caused no immediate member impact. The fibre failure changed the system's operating envelope. Once resilience was reduced, later changes and anomalies had less room to remain local. An operator that treats the first alarm as closed because traffic still moves can miss the fact that the next routine action is entering a degraded system.
Accountability starts with an explicit declaration of that state. A degraded path should change the rules for production admission, maintenance and rollback. It may justify postponing a nonessential activation. It may require a smaller change set, stronger canary tests or direct executive acceptance of the remaining risk. The public record does not say which such decisions LINX made in June 2023. Those are controls the record invites operators to evaluate, not accusations about an undisclosed process.
The event also shows why terms such as redundancy and resilience need evidence. Redundancy is not the count of lines on an architecture diagram. It is the set of alternate paths that remain independent, configured, monitored and able to carry the expected traffic when a particular path fails. Once one dark-fibre path was unavailable, the relevant question was whether remaining inter-switch links, link aggregations, routing processes and forwarding tables had been tested in the exact topology that was now carrying production traffic.
The Lab-to-Production Boundary
LINX's most consequential disclosure concerned a newly installed router that had previously been used in a lab. The operator said the device originally shared a router ID with another production router. Before deployment, the old OSPF and router-ID configuration was removed and a new router ID was deployed through NETCONF. Nevertheless, the OSPF process continued using the old identity because the process had not been cleared. [1]
That distinction separates configuration intent from running state. A configuration database can show the desired identifier. A deployment transaction can report that the new value was accepted. A version-control record can preserve the exact change. None of those records proves that a long-running protocol process has discarded state derived from the earlier value.
OSPF uses a router ID as a stable identifier for the router within the routing domain. RFC 2328 defines the router ID as a 32-bit number that uniquely identifies a router in the autonomous system. [11] Duplicate identity can confuse the relationship between link-state advertisements and the routers that originate them. Vendor guidance classifies duplicate router IDs as a serious anomaly and emphasizes uniqueness, but the LINX evidence must remain attributed to LINX's own report. The public record does not justify assigning the event to a named hardware or software vendor merely because vendor documentation explains the protocol risk. [18]
The lab-to-production boundary therefore needs more than a configuration cleanup checklist. A device can carry residue in a running process, forwarding cache, adjacency database, management plane, credentials or test automation. The exact residue varies by platform and protocol. The general control is a state transition with evidence: reset or recreate the relevant process, boot from the intended production image and configuration, verify live identifiers and adjacencies, and reject admission if the device presents an identity already in use.
A useful admission gate would query the running network from both sides. The candidate router should report its own router ID, but the rest of the OSPF domain should also show which identity it observes. A duplicate-ID detector should operate before production traffic is enabled. The test should not merely compare the intended identifier with an inventory database; it should compare live protocol state, because the incident account demonstrates that intended and live values can diverge.
The same principle applies to interfaces and encapsulation state. A device entering an EVPN fabric should be checked for stale VTEP state, old MAC learning, unexpected VLAN or bridge-domain mappings, and any retained sessions from a lab topology. Its production neighbors should confirm expected adjacencies. The control is bidirectional because a local command can look correct while remote devices retain conflicting state.
Operators often rely on automation to make such transitions repeatable. Automation is valuable, but successful execution is not the same as successful effect. NETCONF can deliver a configuration change consistently, and the system can still require a process clear or reboot before the live protocol adopts the new identity. The correct lesson is not to distrust automation. It is to define automation success in terms of postconditions: the state observed by the running network must match the intended result.
This postcondition should be machine-testable. A deployment pipeline can query the candidate device, its neighbors and a source of truth. It can compare router IDs, adjacencies, route counts, VTEP reachability and selected forwarding entries. It can block activation if an identifier is duplicated or if a forwarding entry exists only in software. Such checks turn a lesson from a specific event into a durable control without pretending that every platform has the same implementation.
Why EVPN and VXLAN Matter to the Evidence
LINX described LON2 as an EVPN-over-VXLAN environment. In simplified terms, VXLAN allows Ethernet segments to extend across an IP underlay by encapsulating frames between tunnel endpoints. EVPN uses BGP to distribute reachability information for the overlay. RFC 7348 describes VXLAN; RFC 7432 defines BGP MPLS-based EVPN, while RFC 8365 describes how EVPN can support network virtualization overlays including VXLAN. [13][14][15]
These technologies create a deliberate separation between layers. The underlay must provide stable IP reachability among tunnel endpoints. The EVPN control plane distributes information that helps devices decide where a MAC or IP is reachable. The forwarding hardware must then program entries that move packets accordingly. This design can scale a peering fabric and reduce dependence on broad data-plane flooding. It also creates several places where an apparently valid record can diverge from packet behavior.
LINX's report that an address appeared in the software MAC table but not the hardware table is a direct example of that evidence problem. The software control plane had a record. The forwarding path that packets relied on did not have the corresponding programmed state. Clearing the L2VPN EVPN BGP session repopulated the hardware table in the cases LINX described. [1]
That observation does not prove a universal EVPN defect. It does not establish that BGP itself was incorrect or that every missing hardware entry had the same cause. It establishes a narrower point: a control-plane view was insufficient to prove forwarding readiness in this incident. A closeout process had to examine the data-plane result.
RFC 9062 is relevant because it describes operations, administration and maintenance requirements for EVPN, including mechanisms for checking connectivity and correlating control-plane and data-plane behavior. [17] The standard is not an incident report, and it does not identify LINX's implementation fault. It helps define the type of verification that an operator needs: tests that can reveal whether the distributed control view and actual forwarding path agree.
This is especially important at an Internet exchange. A peering LAN is shared infrastructure on which many autonomous systems establish bilateral sessions or use route servers. The exchange does not control every member's routing policy, but it does control the fabric that carries member traffic between ports. A forwarding inconsistency can therefore manifest as selective reachability: some addresses, paths or traffic classes fail while broad health indicators remain green.
Selective failure is difficult to close with aggregate metrics. A device can be reachable over management. Most MAC entries can be present. A route server can maintain BGP sessions. Overall traffic can remain high. Yet a subset of members can still lose reachability because a particular entry, path or hardware programming step is absent. The verification strategy must sample the network at the granularity where the failure can occur.
That means testing from more than the core. Member-facing probes should validate reachability across relevant leaf pairs and both peering LANs where applicable. Tests should cover known VTEP and bridge-domain boundaries. Operators should compare software and hardware tables for selected entries before and after a change. A canary should use the same forwarding path and control process as production rather than a simplified lab topology that omits the conditions most likely to expose stale state.
Identity Is an Operational Resource
The router-ID issue also illustrates a broader principle: network identifiers are operational resources, not decorative metadata. An identifier may look like an address, but its function is to distinguish one process or device from another. Accuracy and uniqueness affect the network's ability to reason about topology and state.
This aligns with a restrained Heng.lu doctrine frame. Registries and inventories are valuable ledgers. They record what an operator intends to exist and who is responsible for it. They do not enforce the running protocol. A source-of-truth system may assign a unique router ID, while a live process continues announcing an old one. The ledger is necessary but not sovereign over the running code.
The correct response is not to abandon records. It is to bind them to operational evidence. An inventory row should have a corresponding validation result from the device and its neighbors. A deployment should record the before and after identity, the process restart or clear event, and the absence of duplicates across the relevant domain. The record becomes authoritative because it is reconciled with the network, not because it was entered first.
Router identity is one surface among many. VTEP addresses, autonomous-system numbers, peering-LAN addresses, interface identifiers and MAC addresses each have uniqueness or ownership expectations. Errors can cause different symptoms, but the governance problem is related: an operator must know which system allocates the value, which system enforces it, how conflicts are detected, and who may override the result.
The lab makes this harder because it is designed for reuse. Devices are configured, reset, repurposed and connected to temporary topologies. That flexibility is useful for testing. It also means a lab-used device should be treated as carrying untrusted state until proved otherwise. Deleting visible configuration is only one part of sanitation. Persistent process state, startup files, cached forwarding information, management credentials and image differences must be included in the boundary appropriate to the platform.
A strong organization formalizes that boundary. It defines a production admission profile, a known reset procedure, a software and firmware baseline, identity allocation, cryptographic or checksum evidence for images and configuration, and a live-state acceptance test. The person performing the move should not be the only person able to attest that it passed. Independent verification can be automated or human, but it must check the postcondition rather than simply repeat the deployment command.
The public record does not say that LINX lacked all of these controls. It says enough to show that the old OSPF identity remained active after the intended change. The responsible analytical move is to focus on the failed postcondition and the evidence needed to prevent recurrence, not to invent an entire private process from one disclosure.
The June and November Episodes Must Stay Distinct
LINX's update records more than the 20-22 June sequence. It describes further reachability problems on 29 June and another episode on 5 November after vendor dark-fibre maintenance. LINX also discussed a software change intended to address hardware MAC-table behavior and a later rollback while a vendor investigation continued. [1]
These records strengthen the case for durable verification, but they do not permit the analyst to merge every symptom into one certain root cause. Similar reachability symptoms can arise from different combinations of physical, control-plane and forwarding failures. A repeated observation can indicate recurrence, an incomplete fix or merely a similar external effect. Root-cause confidence depends on evidence that connects the internal mechanism, not on the wording of the symptom alone.
The June router-ID disclosure is specific. The software-versus-hardware MAC-table observation is also specific. The November sequence includes vendor maintenance and a software rollback. An accountable incident record should preserve how each was linked, what remained a hypothesis, and which corrective action addressed which failure mode.
This matters because broad remediation can hide narrow uncertainty. If an operator reboots a device, clears sessions, changes software and adjusts fibre paths in the same recovery window, service may return without isolating the decisive action. Recovery is the immediate objective, but learning requires a separate record of what evidence supports each causal claim. Otherwise the next team may repeat a change without knowing whether it fixed the fault or merely reset the conditions.
The correct public language should mirror that uncertainty. LINX can report what it observed, what it changed and what a vendor continued to investigate. An analyst can explain the control implications. Neither should transform an unresolved vendor investigation into a definitive attribution. That boundary protects technical accuracy and legal fairness.
It also improves engineering. Teams are more likely to preserve competing hypotheses when they are not pressured to name a single cause prematurely. They can test whether stale OSPF identity explains topology changes, whether missing hardware MAC state explains selective reachability, whether the degraded fibre condition changed the blast radius, and whether a software defect recurs under a reproducible sequence. Each hypothesis has a different validation method.
Availability Percentages Are a Starting Point
LINX's 2023 annual report says LON2 achieved 99.997 percent availability, below the exchange's 99.998 percent internal target, and attributes the shortfall to several outages. [2] The difference appears small, but at network scale the operational question is not whether the percentage looks close to one hundred. It is what population, interval and service definition produced it.
An aggregate availability figure can summarize a service. It cannot, by itself, show which members could exchange traffic, which ports or VLANs were affected, whether reachability was symmetric, or whether a subset experienced a longer failure than the average. Selective forwarding failures are especially vulnerable to being diluted in a broad percentage.
The measurement method therefore belongs in the evidence package. The operator should define the denominator, excluded maintenance windows, probes, vantage points and threshold for an outage. If the metric is port-level, it may miss a path that is electrically up but unable to reach particular destinations. If it is based on aggregate traffic, unaffected members can mask failures among a smaller group. If it is based on alarms, the measurement inherits the blind spots of the monitoring system.
For an EVPN peering fabric, useful availability evidence includes member-to-member probes, route-server reachability, bilateral-session continuity, selected overlay connectivity checks and reconciliation of MAC and forwarding state. No single measure represents the entire service. The annual percentage should be traceable to lower-level events so that an auditor can understand why a particular incident counted as it did.
This does not mean every internal metric must be published. Detailed topology and member data can create security and confidentiality risks. Accountability can preserve restricted evidence while publishing bounded conclusions: the measurement scope, the affected service class, the time boundary, the confidence level and the remediation status. The objective is not maximal disclosure. It is a record strong enough to support the claim being made.
LINX's publication of technical detail is therefore valuable. The November update goes beyond a bare availability number and describes operational symptoms and recovery actions. The remaining gap is natural for a public report: it does not expose every member or private device record. A responsible article uses the disclosure to define controls without pretending the public has the complete incident file.
Degraded-State Change Control
The chronology suggests a general rule: a network in degraded redundancy should enter a different change-control mode. This is analogous to operating an aircraft with deferred equipment or a data center on a single power feed. The service may remain available, but the acceptable change envelope has narrowed.
The first requirement is visibility. The degraded condition should be represented in the operational source of truth, not only in an incident chat. Change tooling should know that a critical path is unavailable. A scheduled device activation or software change can then be evaluated against the actual topology rather than the nominal design.
The second requirement is dependency mapping. A dark-fibre path, link aggregation, core device and EVPN session can participate in the same end-to-end service even though different teams or vendors own them. The change record should identify which remaining components now carry the protected load and which test proves that the alternate path is functioning.
The third requirement is a canary that uses the degraded path. Testing through a fully healthy lab or an unaffected production segment can produce false confidence. If the risk is that traffic will traverse an alternate inter-switch link or a different VTEP, the canary must exercise that route. It should validate control-plane convergence and packet forwarding, not just management reachability.
The fourth requirement is a stop condition. Operators need predeclared evidence that ends the change: duplicate identity, unexpected adjacency, hardware-table mismatch, increased drops, link flaps or failed member probes. A stop condition makes rollback faster because the team does not need to debate whether the symptom is serious enough while the network is changing.
The fifth requirement is recovery ownership. Fibre repair may belong to a vendor, device state to the exchange, and member continuity to each connected network. The incident commander needs a record of who can take which action and which evidence marks completion. Distributed responsibility should produce explicit interfaces, not a gap in accountability.
The public material does not reveal whether all of these controls existed at LINX. They are derived requirements for operators facing the same class of risk. The evidence justifies them because it shows a degraded path, a production activation with retained process identity, forwarding inconsistency and later recurrence. It does not justify a finding about undisclosed internal governance.
What a Safe Production Admission Record Should Contain
A credible admission record begins before the device leaves the lab. It identifies the hardware, software image, configuration source, intended production role and identities that the device will use. It records whether the device has been sanitized or rebuilt and which persistent state is expected to survive. The objective is reproducibility: another qualified operator should be able to determine exactly what was admitted.
Next comes identity verification. The record should include the intended OSPF router ID, the live value reported by the running process and the values observed by neighbors. It should show a domain-wide duplicate check. If the identifier changes, the record should capture the process clear, restart or reboot required by that platform and the successful re-establishment of adjacencies under the new identity.
The underlay then needs validation. VTEP loopbacks must be reachable through the expected paths. Interface state, metrics and link aggregation should match the intended topology. The check should include the failure condition relevant to the design: if one dark-fibre path is unavailable, the remaining route must carry the expected traffic without creating a hidden single point of failure.
The overlay needs its own evidence. The device should advertise and learn the expected EVPN information. Route targets, bridge domains and VNI mappings should be checked against the production source of truth. Unexpected routes or MAC entries should block admission rather than be dismissed as harmless lab residue.
Forwarding reconciliation is the decisive step. Selected MAC and IP entries should be present in both software and hardware views where the platform exposes them. Operators should test real packet forwarding across the affected leaf and VTEP combinations. A software table alone is not proof because LINX's account describes exactly that divergence.
The record should also contain negative tests. There should be no duplicate router ID, no unexpected neighbor, no stale lab VTEP, no unapproved VLAN, no untracked route target and no forwarding entry confined to software. Negative evidence is often omitted because success checks are easier to automate. Yet a production admission gate exists partly to prove the absence of dangerous residue.
Monitoring must be bound before traffic is exposed. Flow telemetry, link state, control-plane session state, hardware programming alarms and member-facing probes should identify the new device and its dependencies. Alert ownership and rollback authority should be explicit. A monitor that generates an alert without a responsible responder is not a completed control.
The change should use a bounded canary. That may mean a limited set of paths, a maintenance segment or a low-risk period, depending on the exchange design. The canary must still be architecture-representative. It should not bypass the EVPN control plane, alternate fibre path or hardware tables whose behavior is under test.
Finally, the closeout should compare intended and observed state. It should preserve configuration checksums, live command output, neighbor observations, probe results, telemetry and the final decision. If an exception was accepted, the owner, duration and compensating control should be recorded. This is the evidence that transforms a successful deployment from a claim into an auditable event.
Restoration Evidence Must Reach the Member Edge
Core alarms can show that a link is stable. BGP can show that EVPN sessions are established. A hardware table can show that entries are programmed. None of those observations alone proves that a member can exchange traffic with the expected destination.
Restoration therefore needs an outside-in dimension. Probes should originate from member-facing or representative external vantage points and cross the paths that were affected. They should test bidirectional reachability and, where appropriate, more than one packet size or protocol. A single successful ping is weak evidence for a peering service that carries diverse traffic.
The result should be segmented. If only specific IP addresses remained unreachable on 22 June, a broad fabric health check could pass while those cases persisted. The closeout should identify the affected class and demonstrate that the exact failure is gone. This may require checking selected MAC entries, ARP or neighbor state, EVPN routes and packet forwarding together.
Member reports are also evidence, though not the only evidence. An exchange should correlate tickets and reports with telemetry rather than treating either source as dispositive. A member can observe a failure that central monitoring misses. Conversely, an application problem can resemble an exchange failure. Shared timestamps and path data help distinguish them.
The operator should preserve the gap between partial and full restoration. LINX's chronology indicates that links were disabled to restore service and that residual reachability problems continued into the next day. [1] That is a useful distinction. Reporting one restoration timestamp can hide the period during which most traffic worked but a subset remained affected.
An accountable statement should therefore say what was restored, for whom it was verified and what remained under investigation. It should avoid promising universal recovery based on a narrow sample. This is not rhetorical caution for its own sake; it prevents future investigators from treating an intermediate workaround as a final root-cause fix.
Allocating Responsibility Without Inventing Fault
LINX controlled the peering fabric and the process by which a router entered it. That includes device admission, configuration automation, monitoring, maintenance sequencing, link isolation, session clearing and communication about the exchange service. These are direct control surfaces even when a vendor supplies hardware or fibre.
Fibre providers controlled physical repair and maintenance within their contractual scope. Equipment and software vendors controlled product investigation, defect analysis and available fixes. Their responsibility cannot be inferred beyond what the public record attributes. A continuing vendor investigation is evidence of unresolved technical work, not proof of legal fault.
Members controlled their own connections, BGP sessions and continuity choices. Some may have used route servers, bilateral peering, both LINX peering LANs, other exchanges or transit. The public record does not reveal those designs for every member. It would be wrong to assume that all members had the same exposure or that a second contract necessarily supplied an independent working path.
Responsibility can overlap. LINX might operate the fabric, a vendor might maintain fibre, and a member might choose how to connect. The useful question is who held each decision and evidence item. Who knew the path was degraded? Who could postpone activation? Who verified the live router ID? Who could clear the process? Who compared software and hardware tables? Who tested member reachability? Who communicated residual risk?
This allocation is stronger than a search for a single culprit. Network incidents frequently cross organizational and technical boundaries. A control matrix can assign each action without asserting negligence. It also reveals gaps: if no party owns an end-to-end test, every component owner can close a ticket while the service remains impaired.
Public accountability should follow the same discipline. LINX's report supports attributed statements about what it observed and did. RFCs support explanations of protocol behavior. Vendor documents support general operational guidance. None of those sources supplies private contracts, complete customer harm or a legal conclusion. Keeping those layers separate makes the analysis more useful and more defensible.
An Evidence Model for Recurrence
The June and November records show why recurrence management needs a versioned evidence model. Every corrective action should identify the failure hypothesis it addresses, the expected observable change and the test that confirms the result. A reboot can clear state, but the record should say which state and why that matters. A software update can change behavior, but a successful installation does not prove the target symptom is gone.
For router identity, the hypothesis is that a stale running OSPF process continued using an old ID. The expected correction is unique live identity across the domain after the relevant process reset. Evidence includes local state, neighbor state and a duplicate scan.
For the software-versus-hardware MAC mismatch, the hypothesis concerns programming or synchronization between control and forwarding layers. The expected correction is that selected entries appear consistently and packets traverse the expected path. Evidence includes software tables, hardware tables, EVPN session state and packet tests.
For fibre degradation, the hypothesis concerns reduced path diversity and the load or topology carried by remaining links. The expected correction is restored physical diversity and successful failure testing. Evidence includes path maps, optical or link status, capacity measurements and controlled failover results.
For a vendor software change, the hypothesis should be tied to a defect or symptom. A rollback after an outage is important evidence, but it does not alone establish that the change caused every observed problem. The record should preserve software versions, activation times, symptoms, rollback time and post-rollback tests.
These evidence bundles should share a timeline. Without time synchronization, teams can mistakenly connect a control-plane event to a traffic symptom that occurred earlier or later. Devices, telemetry systems and ticket records need reliable timestamps and known clock behavior. The timeline should also distinguish observation time from event time and action time.
The operator can then compare incidents without flattening them. If November reproduces the same hardware-table divergence under a similar topology, confidence in a shared mechanism rises. If only the user-visible symptom recurs, confidence should remain lower. This preserves uncertainty while still enabling learning.
Governance Metrics That Reflect the Network
Traditional change metrics such as deployment success rate and mean time to recovery are useful but incomplete. A deployment can succeed in the automation system while leaving stale process state. Recovery can appear fast in aggregate while selected members remain unreachable.
A lab-to-production program should measure postcondition failures. How often does intended identity differ from live identity? How often does a device require an unplanned process clear or reboot? How many admissions reveal stale neighbors, routes or forwarding entries? How often do software and hardware tables disagree during a canary?
The exchange should also measure degraded-state exposure. How long does the network operate without the intended path diversity? How many changes occur during that interval? Which changes were explicitly risk-accepted? Did each have architecture-representative canary and rollback evidence?
Member-level verification deserves its own metric. For each consequential incident, what portion of the affected service was tested from representative member-facing vantage points? How long was the gap between broad restoration and resolution of residual cases? Were the failed probes and successful retests retained?
Evidence completeness is another operational metric. An incident closeout can score whether it contains synchronized telemetry, topology state, protocol identity, forwarding data, member probes, action ownership and bounded public claims. The purpose is not bureaucratic completeness. Missing evidence identifies where the operator could not prove its own conclusion.
These measures should not become a new layer of performance theater. A team can optimize the count of completed checks without improving the running network. Random sampling, independent review and direct comparison with incident outcomes help keep the metrics honest. The test remains whether the controls detect real mismatches before members do.
What the Public Record Does Not Establish
The available sources do not identify every affected member, prefix, session or traffic volume. They do not provide a complete duration for each reachability problem. They do not disclose the full device configuration, automation transaction, forwarding database or vendor defect record.
The public record does not establish that the duplicate OSPF router ID alone caused every June and November symptom. It does not prove that EVPN, VXLAN, OSPF, automation or disaggregated hardware is inherently unsafe. These are widely used mechanisms whose operational quality depends on implementation and control.
The annual availability figure does not establish a customer-by-customer harm total. A percentage can summarize a service under a chosen definition; it cannot be multiplied into an unsupported estimate of lost transactions or business damage.
The record also does not establish negligence, breach of contract or concealment by LINX, a vendor or an individual engineer. Technical accountability analysis can identify the control and evidence required without making a legal finding.
These limits are not footnotes to be discarded after the narrative becomes compelling. They define the confidence of every conclusion. The strongest claims are those directly attributed to LINX's technology update and annual report. Protocol explanations are strong within the scope of the RFCs. Claims about private causation, impact and responsibility must remain qualified.
The Running Network Is the Final Record
The enduring lesson from LON2 is not that laboratories are dangerous or that modern fabrics are too complex. It is that the transition from intended state to running state must be proved.
A configuration repository can assign a new router ID. NETCONF can deliver it. An inventory can show the right device and interface. An EVPN control plane can display a learned MAC address. An availability dashboard can remain close to one hundred percent. Each record is useful, but each can be true while the packet path is still wrong.
The final operational record is the running network: unique protocol identity observed across the domain, current adjacencies, coherent EVPN state, programmed hardware forwarding, stable physical paths and completed reachability tests. The evidence system must connect those observations to the change and incident timeline.
That standard also protects the value of automation. Automation becomes more trustworthy when it verifies postconditions and blocks on disagreement. The answer to a stale process is not more manual work. It is a deployment contract that defines when a process reset is required and proves the live result.
For Internet exchanges, the standard is especially important because the fabric sits between independent networks. The exchange cannot control every member's policy, but it can prove the state of the shared platform. It can identify degraded redundancy, reject duplicate identities, reconcile control and forwarding planes, test representative member paths and preserve an honest record of what remains unknown.
LINX's public account provides unusually useful material for that discipline. It identifies the reduced fibre state, the retained OSPF identity, the software-to-hardware MAC mismatch, the workarounds, later recurrence and continuing investigation. The accountable response is not to turn those details into a simplistic fault story. It is to turn them into admission gates and evidence that make the next lab-to-production transition safer.
Architecture names do not guarantee continuity. A redundant path that has not been exercised is a promise. A new identity that exists only in intended configuration is a promise. A MAC entry that exists only in software is a promise. Accountability begins when the operator can show, from the running network, that each promise became packet-forwarding reality.
Sources
- https://www.linx.net/wp-content/uploads/2022/07/LINX120-OpsRouteServers-AnneBatesTimPreston.pdf
- https://www.linx.net/wp-content/uploads/2024/05/Annual-Report-2023.pdf
- https://www.linx.net/news/world-first-as-linx-completes-migration-to-new-disaggregated-lon2-network-model-using-evpn-routing-technology-on-open-network-hardware/
- https://www.linx.net/lon2-and-the-linx-dual-lan-in-london/
- https://www.linx.net/wp-content/uploads/2021/02/DSLONA4v3-0920-1.pdf
- https://www.linx.net/wp-content/uploads/2021/04/LINX-2018-Annual-Report.pdf
- https://community.linx.net/exchange-docs-oo8vcsp0/post/linx-route-servers-information-xCXmq6SqZUpC80k
- https://www.linx.net/services/peering-services/
- https://www.linx.net/route-server-automation/
- https://www.linx.net/wp-content/uploads/2025/04/Peering-Bandwidth-Service-Terms-REDLINE-Draft-10-March-vs-22nd-April-1.pdf
- https://www.rfc-editor.org/rfc/rfc2328.html
- https://www.rfc-editor.org/rfc/rfc4271.html
- https://www.rfc-editor.org/rfc/rfc7348.html
- https://www.rfc-editor.org/rfc/rfc7432.html
- https://www.rfc-editor.org/rfc/rfc8365.html
- https://www.rfc-editor.org/rfc/rfc5880.html
- https://www.rfc-editor.org/rfc/rfc9062.html
- https://www.juniper.net/documentation/us/en/software/juniper-routing-director2.7.0/user-guide/topics/concept/igp-anomaly-detection-overview.html
- https://www.juniper.net/documentation/us/en/software/junos/bgp/topics/topic-map/troubleshooting-bgp-sessions.html
- https://www.peeringdb.com/ix/321
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
