Summary
- IBM said an external network provider flooded IBM Cloud with incorrect routing, causing severe congestion and affecting services and data centres. Contemporary reporting and Zello's affected-customer record show that the event also obstructed logins, active communications, administrative consoles and status visibility, so the practical failure was broader than a single application outage.
- The public record supports scrutiny of the rules that accepted provider routes, limits on unexpected route volume, alarms tied to customer reachability, and administration and status paths that do not depend on the failing cloud network. It does not identify the provider, disclose the precise contract or private topology—the arrangement of network devices and connections—or prove which control was absent or ineffective.
- IBM reported restoration, adjustment of routing policy—the rules governing which network paths are accepted and preferred—mitigation steps and no identified data loss or cybersecurity issue. Those are attributed statements, not proof of a malicious trigger, legal liability, a uniform worldwide recovery time or the present effectiveness of the remediation.
The outage customers could see
The first useful description of the incident is not a protocol diagram. It is the experience of people trying to use services. TechCrunch reported that problems appeared at about 14:30 Pacific Time on 9 June 2020 and developed into a broad networking disruption. Hosted services were affected. IBM's principal status page returned an internal-server error during the event, while a separately hosted status page for another IBM service showed widespread network impact. The live report did not know the trigger. That uncertainty is important: observation of a wide outage establishes impact, not cause.
Zello's incident record gives a more concrete downstream view. It opened its incident at 18:03 Central Daylight Time. Zello said most users could not log in or reconnect, people who had already connected lost communication, administrative consoles were unavailable, and four named Zello services were affected. Zello marked its incident resolved at 20:22 in the same time zone. These facts describe Zello's service and its users. They are not a complete count of IBM Cloud customers or a universal clock for every IBM region and product.
CRN, a technology news publication, reported similar control and visibility problems from IBM partners and customers. People could not reach environments, consoles or status screens. It placed IBM Cloud's first public tweet at 20:26 Eastern Time and a later all-services-restored tweet at 21:54. Those timestamps help reconstruct public communication, but they do not turn a complex recovery into one global start and finish. A cloud network can return in stages as routing policies change, congestion falls and different services re-establish connections.
The later IBM notice supplied the attributed cause. As preserved by The Register and CRN, IBM said an external network provider flooded IBM Cloud with incorrect routing. IBM said the resulting congestion affected cloud services and data centres. It also said network specialists adjusted routing policies, all services were restored, mitigation steps had been taken, and its root-cause work had not identified data loss or cybersecurity issues.
This is a strong operator statement, but it has limits. The public material does not include logs from IBM routers, the network devices that select and forward traffic paths, the unnamed provider's records or a conclusive independent reconstruction of the route exchange. The notice tells readers what IBM concluded and did at a high level. It does not disclose which individual messages were accepted, which engineer approved a change, which contractual party owned a route-acceptance rule, or whether a particular safeguard was missing.
Zello later described the mechanism from its affected-customer perspective. Its account says an external provider injected a large number of routes. Servers in IBM Cloud then sent outbound IPv4 traffic toward incorrect destinations, contributing to severe congestion and loss of service. Internet Protocol version 4 is the widely used addressing system in which each reachable device or service uses a numerical address.
Zello's description is useful because it connects routing state—the set of paths a network currently believes and uses—to an application consequence. It remains Zello's published account, not a substitute for the private records of IBM and the provider.
Taken together, the public evidence supports a disciplined narrative. A provider-originated routing event reached an IBM boundary. IBM attributed severe congestion and service impact to incorrect routing. Independent observation showed broad reachability and status problems. An affected customer documented login, communication and administration failures. Recovery was staged and included a reported routing-policy adjustment. Anything more precise must be labelled unknown.
How a routing message can become a cloud outage
BGP is often described as the routing system of the Internet. That phrase can sound as though one central service calculates every path. In reality, the Internet is a collection of independently operated networks. Each is commonly represented by an autonomous system, or AS, which is a network under one administrative policy. These systems exchange route advertisements so that each can decide where to send traffic for a network prefix, meaning a block of related addresses.
A route advertisement does not carry customer data. It belongs to the control plane, the part of a network that decides which paths should be used. Customer packets, the small units into which network data is divided, travel through the data plane, the equipment and links that forward the actual traffic after those decisions have been installed. If the control plane learns an incorrect path and installs it, the data plane can faithfully send traffic in the wrong direction. Running equipment can therefore be healthy in isolation while the service is unreachable because the installed path is wrong.
The public IBM account uses the broad phrase “incorrect routing.” Zello says a large number of routes were injected. Neither source publishes the exact prefixes, the private BGP session, the list of networks that originated the announcements, or the policy IBM used to evaluate them. It would be an overstatement to say that a full Internet routing table—the catalog of network paths known to a router—arrived, that one precise address block caused the incident, or that a particular maximum setting was disabled. The evidence establishes the category and outcome, not the private command history.
Volume matters because a route is both information and work. A router must receive a message, compare it with policy, update its stored view of reachable networks, choose a preferred path and, where appropriate, install forwarding instructions. It may also send changed information to other routers. A sudden set of unexpected routes can consume processor time, memory, update capacity or link bandwidth, the amount of data a connection can carry in a given time. Even if individual messages are syntactically valid, their number or content can cause damaging churn—the repeated recalculation and replacement of paths.
Congestion can then appear in more than one place. Routing processors may struggle to keep up with changes. Links can carry traffic toward unexpected destinations or exceed their available capacity. Traffic that should have used several paths may concentrate on fewer paths. Applications can retry when replies do not arrive, increasing demand. The IBM statement says severe congestion followed the incorrect routing, but it does not locate every bottleneck. A responsible explanation therefore describes the plausible chain without claiming a private measurement that was never published.
The term route convergence means the period in which routers process a change and settle on a consistent usable set of paths. Convergence is not always instantaneous. During it, different devices can temporarily hold different views. Some traffic may follow an old path, some a new one and some no usable path at all. That helps explain why a routing-policy adjustment can produce staged recovery. It does not prove the duration or internal sequence of IBM's 2020 convergence.
The incident should not automatically be called a route leak or a route hijack. A route leak is the accidental spread of reachability information beyond the relationship in which it was intended to be used. A route hijack is an unauthorized claim to reach addresses controlled by another network, sometimes malicious and sometimes caused by error. The admitted record does not establish either category. It says an external provider flooded IBM Cloud with incorrect routing. Preserving that wording prevents a familiar technical label from becoming invented evidence.
Nor should the event be called a distributed denial-of-service (DDoS) attack, a deliberate attempt to overwhelm a service with traffic from many sources. Congestion occurred, but congestion alone does not prove hostile intent. IBM said its work had not identified a cybersecurity issue. That is an attributed negative finding rather than proof about every possible security consequence, yet it is enough to reject an unsubstantiated attack narrative.
The provider boundary is a place to apply controls
A cloud company buys connectivity and exchanges routing information with other networks because no provider can operate the Internet alone. The relationship creates a boundary between organizations, but that boundary is not a place where operational responsibility disappears. It is a control surface: a point at which policies decide what information and traffic may enter, how much change is acceptable and what happens when the relationship behaves outside expectations.
The external provider in this case is unnamed. Its precise legal entity, contract, exact point of technical connection and internal cause are not public. The evidence does not show whether IBM bought direct Internet transit, a service in which another network carries its traffic onward, used an intermediary arrangement or delegated a particular decision. It also does not reveal who owned each route-acceptance rule, warning or emergency-contact step. Those omissions prevent a fair allocation of fault between the parties.
They do not prevent a control analysis. IBM Cloud operated the environment affected by the accepted routing state. It was positioned to define the conditions under which provider information entered its network, observe the health of customer reachability and maintain ways for engineers and customers to obtain information during a failure. The provider operated the system that sent the routing information. Each side therefore controlled different facts and actions, while customers controlled neither.
At a minimum, an operational relationship should have an expected-route model. That model identifies which address blocks a peer—a neighboring network that exchanges routes—or provider is reasonably expected to announce, how much variation is normal, which changes require approval and how exceptions expire. It should be based on the real commercial and network relationship, not on a generic list copied from another connection. The public record does not say whether IBM and its provider had such a model. The incident makes the question necessary.
The boundary also needs ownership that survives an emergency. One team must know who can stop accepting new route changes, who can withdraw an unsafe policy, who contacts the provider, who protects customer traffic while investigation continues and who decides when normal exchange may resume. A contract can describe responsibilities, but a name on paper is not enough. Contact paths, access rights and decision authority must work at the moment the shared network is congested.
This distinction matters because provider blame can become an excuse for weak internal containment. A cloud operator cannot guarantee that every partner will always send correct information. It can decide how much untrusted or unexpected state becomes active at its own boundary. Conversely, an operator should not use the ability to filter as proof that the provider bears no duty. The originating network must maintain accurate advertisements, safe change practices and rapid cooperation. Accountability is distributed, but not interchangeable.
The hardest test is whether the controls match the executed relationship. A policy document can say only approved routes are accepted while the configuration actually running permits much more. A contact list can name an emergency owner whose access depends on the failing network. A dashboard can appear green because it measures device health rather than whether customers can reach services. The reality of the service is the code, configuration, installed routing state and observed customer outcome at the time of the event.
What standards guidance can and cannot establish
Request for Comments (RFC) 7454 is a published Internet engineering guidance document about securing BGP operations. It recommends that network operators build explicit policy for routes received from and advertised to other networks. It discusses inbound prefix filtering, meaning rules that accept only allowed address blocks from a connected network, and outbound filtering, meaning rules that stop a network from advertising routes it should not export.
The guidance also discusses a maximum-prefix limit, a protective threshold on how many address blocks a neighbor may announce before the router warns, rejects additional routes or closes the exchange. The value should reflect the expected relationship and the capacity of the equipment. Set too low, it can interrupt a legitimate expansion. Set too high, it may fail to contain an abnormal flood before resources are strained. The threshold is therefore an operational decision supported by history, capacity tests and an exception process.
These controls are directly relevant to a reported flood of incorrect routing. They provide a vocabulary for asking whether the accepted set matched the provider relationship, whether the number of routes departed from an agreed range and whether abnormal change could be contained. They also show why both inbound and outbound discipline matter: one network's unsafe export becomes another network's unsafe input.
RFC 7454 does not tell the reader what IBM configured in June 2020. It does not prove that a specific filter or limit was absent, disabled or set incorrectly. It cannot establish that following every recommendation would have guaranteed prevention. Routing incidents can exploit unexpected combinations of valid information, equipment limits, operational error and dependencies that a single control does not cover.
A sound review would therefore treat the standard as a benchmark, not a verdict. Investigators could compare the live provider policy with the expected route set; examine whether an abnormal route count or content change crossed an alarm threshold; determine which actions the router and operators took; and test whether similar input is contained in a safe environment. The result should identify not only whether a rule existed, but whether it worked before customer reachability failed.
Filtering also requires reliable data. A list of permitted address blocks, often called a prefix list, can become stale when a customer legitimately adds, transfers or stops using address space. Registry data can help document allocation and routing intent, but a registry is a recordkeeper, not a machine that controls every live packet. Operators must reconcile records with contracts, authorized changes and the paths actually running. Accurate records support policy; they do not replace operational verification.
Another prospective control is rate-aware change handling. Not every router can or should process unlimited policy changes at once. An operator can warn on an unusual rate of additions and withdrawals, slow the spread of nonessential route changes, isolate a problematic connected network or require an explicit override for a sudden expansion. The exact design depends on equipment and service architecture. None of these measures is claimed to have been present or absent in IBM's network.
The boundary should also be tested for safe failure. If one provider session exceeds its accepted range, the response should not create a larger outage than the unsafe input would have caused. Tests should examine whether closing a session removes all reachability, whether alternate paths have enough capacity, how quickly services stabilize and whether engineers retain access. A safeguard that protects the control plane but strands customer traffic is incomplete.
Reachability, administration and status failed together
One of the most consequential observations is that customers and partners lost access not only to hosted applications but also to consoles and status information. A console is an administrative interface used to inspect or change a service. When it depends on the same network paths as the service it manages, a routing incident can remove both the workload and the means to diagnose or control it.
The same dependency can affect incident communication. IBM's main status page returned an error during the event, according to contemporaneous reporting. That does not prove there was no communication: separately hosted information and social-media updates existed. It does show why a primary status channel should not share all failure paths with the platform it reports on.
Out-of-band access is an administrative path designed not to depend on the ordinary production path that may be failing. It can use separate network connectivity, credentials, name resolution—the service that turns names into network addresses—routing and hosting. “Separate” must describe real dependencies rather than a different web address served by the same underlying network. The public record does not disclose IBM's management design, so this is a prospective continuity control, not a statement that IBM lacked one.
A cloud operator should map the dependency chain for each critical control surface. Can network engineers reach the relevant routers if the cloud backbone, the core network carrying cloud traffic, is congested? Can they authenticate if the identity service, the system that verifies users, is affected? Can they obtain current configuration and route state if the normal monitoring system is unreachable? Can customer support see impact without depending on the administrative console? Can status editors publish through a path that remains available outside the affected cloud region or network?
The answers must be tested. A tabletop exercise, in which teams discuss their response, can reveal unclear ownership. A live continuity test must go further by using the alternative access path and publishing a test status update while ordinary dependencies are deliberately unavailable. The goal is not to expose sensitive architecture. It is to establish that the supposedly independent route works when needed.
Status communication also needs a confidence model. During the first minutes, an operator can report confirmed symptoms and affected functions without claiming a final cause. It can distinguish observed reachability loss from an attributed provider explanation. It can publish times in one clear time zone and state that restoration is staged when different services return separately. This lets customers make decisions without forcing engineers to convert an active hypothesis into certainty.
The Zello record shows why this matters downstream. A communications service can depend on a cloud provider for login, active connections and administration. When all three become unavailable, its own responders need evidence and a customer channel that does not depend entirely on the failed upstream. Dependency owners should rehearse how they will operate when a provider's console and status page are impaired at the same time.
A timeline that preserves attribution
At about 14:30 Pacific Time, public observation identified a broad IBM Cloud network problem. This is the earliest time in the admitted record, not a universal outage start for every service. During the disruption, the main IBM status page returned an error and a separately hosted service page showed widespread impact.
Zello opened its customer incident at 18:03 Central Daylight Time. It documented widespread inability to log in or reconnect, loss of communication for users who had been connected, unavailable administrative consoles and four affected services. This provides an affected-customer window, not the total IBM incident denominator.
CRN reported an IBM public update at 20:26 Eastern Time and a restoration update at 21:54. Zello marked its own incident resolved at 20:22 Central Daylight Time. Converted between time zones, these records show overlapping but not identical observations. They should not be forced into a single minute-by-minute universal sequence because each publisher measured a different service or communication surface.
IBM's later notice attributed the event to the unnamed external provider and incorrect routing. It said network specialists adjusted routing policies, services were restored and mitigation steps were taken. This places route-policy action within recovery, but it does not reveal the commands, sequence of provider coordination or per-service return.
The most honest description is staged restoration. Some visibility returned before every customer application was necessarily stable. Different systems may have reconnected on their own schedules as paths settled. Applications with long-lived connections can recover differently from new logins. Administrative functions can return separately from workload traffic. The public sources support the fact of gradual recovery more strongly than a claim that one switch restored the entire cloud.
Chronology is part of accountability because decisions occur under changing evidence. Operations teams need to know when abnormal routing first arrived, when congestion warnings fired, when customer reachability fell, when the provider was contacted, when route policy changed and when each service objective—a measurable target such as availability or response time—recovered. Public reporting contains only fragments of that internal sequence. A complete post-incident record should align those events without rewriting uncertainty after the fact.
Evidence a provider-boundary review should preserve
Telemetry means measurements and event records collected while a system runs. For this class of incident, useful telemetry would include the number and rate of routes received from each connected network, the allowed and rejected route set, changes to preferred paths, processor and memory pressure, routing update queues—the backlogs of messages waiting to be processed—link utilization, meaning how much of each connection's capacity was in use, customer reachability tests and administrative-access health. The public evidence does not disclose these IBM measurements. They are the evidence needed to test the attributed cause.
The route record should preserve enough context to distinguish several possibilities. Did a provider add many new legitimate prefixes? Did it advertise paths outside the commercial relationship? Did a policy update cause previously rejected routes to become acceptable? Did the content remain stable while the rate of updates surged? Did IBM's response change preference, reject a set, close a session or move traffic to another provider? These are questions, not claims about what occurred.
Configuration history is equally important. A final configuration shows what remained after recovery, not necessarily what was active when impact began. Investigators need the version before the event, each change during response, who authorized it and what effect followed. A rollback is the deliberate return to a previously known configuration. It should be possible quickly, but it should also be recorded so that an emergency reversal does not erase evidence.
Customer reachability tests should come from outside the affected network as well as inside it. A synthetic probe is an automated test that behaves like a customer by attempting a connection, login or simple transaction. Device health can remain green while outside users cannot find a valid path. Probes from several independent networks can show whether the problem is regional, provider-specific or broad.
The review should also track blast radius, meaning the range of services, customers and regions affected by one failure. The public record proves broad impact and named downstream consequences but does not supply an audited denominator. Internally, service owners should reconcile failed logins, connection losses, console errors, network paths and support reports. Counts must distinguish attempts from people and affected services from affected customers.
Provider coordination creates its own evidence. The incident record should state when the provider acknowledged the problem, which route set each side believed was valid, which side took each containment action and how the parties decided restoration was safe. Sensitive commercial details can remain private. The operational allocation of action should not become unknowable merely because two companies share the boundary.
Evidence retention must also protect privacy and security. Router and application records may expose customer networks or sensitive topology. Access should be limited, retention periods justified and public disclosure aggregated. Those safeguards are compatible with a rigorous incident reconstruction. Privacy should shape evidence handling, not excuse the absence of evidence.
What the public record proves
The record proves that IBM Cloud experienced a broad networking disruption on 9 June 2020. It proves that hosted services, status visibility, environments and consoles were affected according to contemporaneous reporting. It proves that Zello documented login, reconnect, communication and administration failures in its own service.
The record also proves what IBM said. IBM attributed the event to an external network provider flooding IBM Cloud with incorrect routing and causing severe congestion. IBM said routing policies were adjusted, services were restored, mitigation steps were taken and its work had not identified data loss or cybersecurity issues. The operator attribution must remain visible in every retelling.
RFC 7454 establishes that explicit border policy, prefix filtering and peer-specific route-count limits are recognized operational safeguards. It provides an analytical standard for the questions an investigation should ask. It does not establish IBM's configuration or the provider's conduct.
The direct BTW Media directory, an organizational identity index, supports the identity link to International Business Machines Corporation. Event reporting identifies IBM Cloud as the affected operator. The directory does not establish which IBM legal entity signed the relevant provider contract, and it says nothing about the unnamed provider. Identity records help readers connect the organization; they do not reconstruct the running network.
What remains unknown
The provider remains unnamed. The public record does not identify its autonomous system number, the numerical identifier assigned to an independently operated network, or the exact IBM network that received the routes. It does not publish the prefixes, path attributes—the details attached to a route that influence path choice—policy code, private sessions or timestamps needed for packet-level forensics, a detailed reconstruction of network traffic.
The route category remains broad. The sources do not prove malicious intent, a hijack, a leak or a distributed attack. They do not identify an attacker. They do not show that one person made a negligent decision. The incident should remain an operational routing failure in public copy unless stronger authoritative evidence becomes available.
The full impact remains unknown. There is no audited count of IBM customers, affected regions, failed transactions, lost messages, financial loss or contractual credits. Zello proves a concrete downstream effect, but its users are not a proxy for the whole cloud. Partner reports establish meaningful access problems without creating a population denominator.
Data impact is bounded. IBM said its work had not identified data loss. That statement should not be reversed into a claim that data was lost. It also should not be expanded into an independent audit of every customer's application state. A service can lose reachability without losing stored data, while customers can still suffer interruption and uncertain outcomes.
Legal responsibility remains unknown. The public material does not establish breach of contract, negligence, regulatory violation, damages or personal liability. Provider-boundary accountability here means the ability to explain controls, decisions, evidence and continuity, not a legal judgment.
Remediation effectiveness remains unknown. IBM said mitigation steps were taken. The sources do not disclose the full measures, test results, recurrence criteria or current configuration. It would be unfair to say the measures failed, and unsupported to call them proven durable. Current assurance requires current evidence.
Turning unknowns into assurance questions
The absence of private detail should not produce a vague conclusion. It should produce exact assurance questions. What route set was the provider expected to send? What set arrived? Which routes were accepted, rejected or preferred? What route-count and change-rate thresholds existed? What alarm fired first? Which customer-reachability signal established impact? Who could contain the session, and how quickly?
The next questions concern continuity. Which administrative paths remained usable while ordinary cloud reachability failed? Which status system was hosted outside the affected dependency? Could support teams see the same impact as network engineers? Were alternative providers available, and could they carry the redirected load? A failover is a controlled switch to a backup path or system. It is useful only if the backup is truly independent enough and has sufficient capacity.
Tests should connect these questions. A provider-boundary exercise can replay a safe set of unexpected route announcements in an isolated environment. The expected result should include rejection or containment, a clear alarm, preserved management access and a visible customer-impact signal. A second exercise can simulate loss of the primary provider and measure whether traffic switches safely. A third can remove the normal status dependency and verify that incident communication still works.
Assurance should report outcomes rather than sensitive topology. An operator can state that peer-specific route rules were tested, abnormal volume was contained before customer objectives failed, administration remained reachable through an independent path and recovery completed within a tested range. It need not publish address lists or router names. Meaningful assurance is possible without exposing a blueprint.
The review should be repeated when the relationship changes. A new provider product, expanded address range, merger, router platform or traffic pattern can invalidate old assumptions. Exceptions should have an owner and expiry. A route allowed temporarily should not become permanent merely because the emergency ended.
Accountability without invented blame
The phrase “third-party provider” can make a failure sound external to the cloud operator. Operationally, the boundary is shared. One party sent information; another party decided how that information became active in its own network. The public record does not allocate fault, but it does show why both actions belong in the investigation.
IBM's duty to customers was not to make every upstream error impossible. That would be unrealistic. The practical duty was to operate a boundary with proportionate acceptance rules, monitor the customer result, contain abnormal behavior and preserve control and communication. The provider's practical duty was to maintain accurate routing, safe changes and rapid coordination. The contract should connect these duties instead of leaving a gap between organizations.
Customers need a different form of accountability. They need a timely description of what is affected, whether data integrity is implicated, how recovery is progressing and which dependencies they should avoid. They do not benefit from an unsupported technical label or a blame contest while the service remains unreachable.
Executives and public-sector customers should ask whether the cloud architecture treats network control as a critical service. A workload may run across several data centres yet remain dependent on a common routing boundary, console, identity service or status path. Logical distribution does not guarantee operational independence. The relevant question is which failure can still remove multiple apparently separate services at once.
This is the central reality-layer lesson. Registries, contracts and diagrams describe intended relationships. The running configuration and observed reachability determine whether the service works. Good records are necessary because they let operators compare intent with execution. They are not sovereign over the packets actually being forwarded.
Sources
- https://www.theregister.com/2020/06/11/ibm_blames_external_network_provider/
- https://techcrunch.com/2020/06/09/ibm-cloud-suffers-prolonged-outage/
- https://www.crn.com/news/cloud/ibm-blames-massive-cloud-outage-on-third-party-network-provider
- https://status.zellowork.com/issues/5ef227ab4a0ebd6da0cb113e
- https://www.rfc-editor.org/rfc/rfc7454.html
- https://btw.media/en/directory/international-business-machines-corporation-united-states-of-america-the
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
