Summary

  • A two-uplink failure in the early Firehose 1.0 design could leave machines reachable from others but unable to reach one another. That effort never carried production traffic.
  • Later accounts describe independently serviceable network blocks, but also show how shared software, control links and configuration procedures can defeat physical redundancy.
  • Urs Hölzle's documented infrastructure leadership and shared authorship provide a lens on this collective operating work, not evidence that he personally designed every mechanism.

Two machines can be connected to the network and still have no usable connection to each other. The contradiction was concrete in Google's early Firehose 1.0 design. If opposite uplinks on two top-of-rack switches failed within the same repair window, machines beneath those switches could reach other machines but not one another. Applications had difficulty handling this non-transitive connectivity. A simple picture of one isolated island would miss the fault.

This is a historical engineering account, not a report of a current Google outage. The original 2015 SIGCOMM paper says Firehose 1.0 never carried production traffic. Its value lies in the authors' willingness to describe an effort that failed operationally and informed the systems that followed.

A person's role, a team's evidence

Urs Hölzle is among the paper's many coauthors. Google Research's biography identifies him as a Google Fellow in Google Cloud and says he was Senior Vice President for Technical Infrastructure until 2023. That earlier remit covered designing, installing and operating the servers, networks and datacentres behind Google's services.

The connection matters because network architecture here was inseparable from physical construction and continued operation. It does not establish who proposed a particular topology or wrote a particular protocol. Reading the paper as the achievement of one celebrated engineer would obscure its strongest evidence: a large team learning which combinations of hardware, software and maintenance could actually work.

Google's 2013 computing announcement and Hölzle's March 2020 network account document the same infrastructure role at two dated moments. Neither turns the retrospective into a description of today's entire network. The latter also separates Google's infrastructure from the ISP last mile: the operating remit does not encompass every dependency a user encounters.

Reliability changed the traffic pattern

The paper explains why a datacentre needs traffic to move widely between groups of machines. Spreading jobs and storage across power and failure domains reduces exposure to one local interruption, but also weakens the locality that would otherwise keep traffic nearby. A resilience choice at the application level therefore creates a bandwidth requirement in the network.

A staged Clos topology offers multiple paths using many switching elements. That made a large fabric possible without relying only on a single enormous chassis. Yet more paths on a diagram do not automatically produce a repairable system. Firehose 1.0's low-radix uplinks illustrate how a particular combination of failures can remove the one relationship an application needs while leaving much of the surrounding connectivity intact.

The follow-on Firehose 1.1 changed both packaging and topology. Ordinary servers no longer housed the switching chips; dedicated enclosures, a separate control network, paired top-of-rack switches and a revised aggregation structure addressed lessons from the earlier effort. The account reports greater robustness to link failures. It also describes labour-intensive cabling and placement constraints. Architecture had to survive installation and replacement, not merely path calculation.

How much network leaves service?

The later Freedom architecture makes the maintenance question unusually tangible. Its typical connectivity layer had four independent blocks. Operators could withdraw traffic from one block and upgrade it, with aggregate capacity falling by 25%. The unit of intervention had become something smaller than the whole layer.

That number must stay attached to its example. It is not a guarantee that every workload preserves performance during a drain. Work still has to fit on the remaining resources, and aggregate capacity does not describe every path's load.

Nor is removing a quarter of the hardware always equivalent to losing a quarter of capacity. A separate upgrade illustration divides chassis across a Clos fabric into four sets. Disabling one set leaves 56.25% of capacity because the losses interact across stages. Eight sets make the procedure gentler but longer. These are different arrangements, not inconsistent measurements. Their common lesson is to choose maintenance groups by the surviving connectivity, not by counting boxes.

The shared machinery behind separate blocks

For Firehose, Watchtower and Saturn, the paper describes Firepath distributing a common topology and link-state view while switches compute forwarding locally. This was logically centralised coordination, not a controller deciding every packet's path. Redundant masters and a separate control network supported it. The paper explicitly leaves Jupiter's detailed control architecture outside its scope.

The intended topology also became an organising specification. Before Jupiter, a small set of cluster parameters generated materials lists, rack and cable plans, control-network details, monitoring information and common switch configuration. Limiting choices made repeated construction easier. It also made the quality of a shared specification important across many devices.

Operational examples show what physical separation cannot solve by itself. A simultaneous fabric restart exposed competition between liveness checks and route computation for limited switch CPU. Ageing links revealed weaknesses where standby and control-network health had not been actively monitored. In a Freedom configuration change, an unlocked concurrent read interacted with a write and produced a partial configuration. The authors describe reverting that change and hardening the tools.

These cases disclose mechanisms, not an outage-rate dataset. The passages do not establish durations or customer harm. They nevertheless explain why an independently serviceable block is an operating achievement rather than a hardware label: detection, software transition and change authority must respect the same boundaries.

What the retrospective leaves us

The 2015 publication record and 2016 CACM edition describe versions of the same work, not independent replications. The first-party paper is unusually useful for the detail it supplies, but cannot prove current implementation or universal availability.

Hölzle's significance in this account is the documented connection between infrastructure leadership and collective engineering that made a large network changeable in bounded pieces. The durable question is not how impressive the whole fabric looks. It is what still shares a failure or a change when one apparently independent piece is taken out.

Sources