Summary
- In October 1986, useful throughput on a short path between Lawrence Berkeley Laboratory and UC Berkeley fell from 32 Kbps to 40 bps: packets still moved, but retransmission feedback consumed the network's ability to deliver new data.
- Jacobson, Karels and their collaborators repaired 4BSD TCP by making senders obey packet conservation through ACK clocking, slow start, better timers, exponential backoff and a congestion window.
- RFC 1122 later made slow start and congestion avoidance mandatory for TCP hosts. The lesson is not rulelessness: decentralised operation survived because a narrow common rule prevented aggressive endpoints from exporting their costs to everyone else.
The 400-yard path that lost three orders of magnitude
In October 1986, two institutions separated by roughly 400 yards became endpoints of one of the Internet's most instructive failures. Van Jacobson and Michael J. Karels later reported that data throughput from Lawrence Berkeley Laboratory to the University of California, Berkeley, across two IMP hops, fell from 32 kilobits per second to 40 bits per second. The wire had not vanished. The systems continued sending. Almost all of the activity simply ceased to be useful.
That distinction matters. Congestion is often pictured as a crowded road where everybody advances slowly. Congestion collapse was stranger: the network remained busy while delivering a tiny fraction of its capacity as new data. Earlier TCP implementations could mistake delayed packets for lost ones. Retransmission injected copies into already full queues. Those copies increased delay and loss, which caused more retransmission. A recovery mechanism became positive feedback.
John Nagle had described this failure mode in RFC 896 in 1984 and gave it its durable name. The memo warned that pure datagram internetworks joining links of different speeds were susceptible to a stable collapse state. It was explicitly an invitation to discussion, not a finished standard. Diagnosis existed before the episode at Berkeley; what the 1986 path supplied was an intolerable, measurable encounter with it.
Jacobson and Karels asked two unfashionably practical questions. Was the 4.3BSD TCP implementation itself behaving badly? Could it be made to operate under abysmal network conditions? Their published answer to both was yes. The repair began with the running system, not a claim that an institution needed jurisdiction over all traffic.
Why more transmission produced less delivery
The original TCP flow-control window answered a receiver-side question: how much unacknowledged data could the destination buffer accept? It did not answer a different question: how much could the path between sender and receiver carry now? A fast host could therefore open a full receiver-advertised window across a much slower long-haul link and inject a burst that the first gateway could not absorb.
Under light load, crude timers could appear adequate. Under heavy load, round-trip time and its variation grew. A sender whose timeout estimate failed to follow that variation retransmitted data that was merely late. The network then spent scarce bandwidth transporting duplicates. Jacobson and Karels compared this to pouring gasoline on a fire: the moment the system struggled with useful work, its endpoints assigned it more useless work.
Their four-flow test made the waste visible. Without congestion avoidance, 4,000 of 11,000 transmitted packets were retransmissions. The bottleneck link could carry about 25 KB per second, yet six KB per second of useful capacity disappeared. With congestion avoidance, only 89 of 8,281 packets were retransmitted—about one per cent—and the measured flows accounted for the link's capacity. This was not a philosophical preference for elegance. It was an operational recovery of goodput.
The packet-conservation rule
The design was organised around a simple proposition. Once a connection reaches equilibrium with a full window in flight, a new packet should not enter the network until an old packet has left. Acknowledgements provide evidence of that departure. Because a receiver cannot generate acknowledgements faster than data passes through the bottleneck, ACK spacing returns a rough clock to the sender.
That clock let an endpoint adapt without knowing the path's topology, owning its gateways or consulting a central allocator. But a connection that had just started—or restarted after loss—did not yet possess a clock. Slow start created one. The sender began with a small congestion window and expanded it as acknowledgements arrived. Despite its name, the window opened exponentially over round trips; the restraint lay in probing capacity through returned evidence rather than dumping the receiver's full allowance into an unknown path.
Timers required similar discipline. Estimating not only average round-trip time but also its variation reduced false timeouts. Exponential backoff spaced repeated attempts ever further apart when delivery continued to fail. Congestion avoidance then added a second sender-side limit: the congestion window. On congestion, the implemented policy cut that window multiplicatively; as new acknowledgements arrived, it increased additively. Transmission was bounded by the smaller of the receiver's advertised window and the congestion window.
No single ingredient deserves to swallow the story. The 1988 paper says seven changes had entered 4BSD TCP, including work associated with Phil Karn, a receiver ACK policy and fast retransmit as well as the better-known slow start and window adjustment. It also credits John Nagle for the name “slow start” and notes that its window policy drew on Raj Jain's scheme. The historical achievement was a connected implementation programme and a measured control loop, not magic committed by a solitary hero.
A control loop at the edge
The architecture is striking because its main regulator lived in the sender. Existing loss and timeout signals could be used without modifying every gateway. Independently operated hosts could implement the same response while remaining free to choose operating systems, applications, routes and commercial relationships. The common layer specified behaviour at the point where one participant's choices could destabilise the shared path.
This was decentralisation with feedback, not decentralisation as indifference. A sender observed consequences it did not control, translated them into a local limit, and tried again cautiously. The network core did not schedule every flow by decree. Nor did the endpoint receive a right to ignore the state of the commons merely because its machine was privately administered.
Running code also disciplined the institutional sequence. The algorithms went into 4BSD TCP, were traced under load and circulated among implementers. Only afterward did RFC 1122, published in October 1989, state that TCP “MUST” implement the combined slow-start and congestion-avoidance algorithm. It also required exponential backoff for successive retransmission timeouts. The specification did not invent authority from rhetoric; it recognised a failure mechanism and a tested remedy that interoperating hosts needed from one another.
Freedom, reciprocity and the aggressive sender
That MUST complicates a romantic account of voluntary coordination. A conformant TCP implementation was no longer invited merely to consider restraint. Why was compulsion inside the specification legitimate here when thick Internet governance so often is not?
Because the scope was tied to a demonstrable shared invariant. A sender that refuses to back off does not keep the consequences within its own network. It fills common queues, raises loss for responsive flows and makes restraint irrational for others. RFC 2914 later described the resulting danger: vendors might market a more aggressive TCP as faster, applications might open more parallel connections, and the system could enter an arms race back toward chronic congestion.
The narrow requirement protected room for wider freedom. Hosts could innovate above and below it; new congestion algorithms could be developed; routers could add signals; operators could choose their paths. What they could not plausibly demand was an interoperability right to recreate the positive feedback that destroyed useful service. Reciprocity, in this setting, was not a political mandate over the endpoint. It was the technical price of sharing a bottleneck.
What endpoint control could not solve
Jacobson and Karels were explicit about the boundary. Endpoint algorithms could keep aggregate demand from persistently exceeding capacity, but they could not guarantee fair allocation. Gateways see flows converge and therefore possess information unavailable to any one sender. The paper treated gateway congestion detection as the next major step.
Later work added active queue management, explicit congestion notification and many alternative congestion-control algorithms. Loss is not always congestion, and paths with wireless errors, enormous bandwidth-delay products or novel application patterns complicate the original assumptions. RFC 2914 also made clear that responsive end-to-end control, though necessary, was not sufficient for every service and fairness problem.
That limitation strengthens the historical lesson. The 1988 repair did not pretend to settle every question because it did not need to. It isolated the shared failure mechanism, changed the smallest effective control surface, demonstrated the result and left adjacent decisions open. The common rule was powerful precisely because it remained specific.
Sources and evidence limits
The central measurements and algorithm descriptions come from Jacobson and Karels's Congestion Avoidance and Control, indexed by the LBNL Network Research Group. The earlier diagnosis is in RFC 896; the 1988 host-extension context appears in RFC 1072; the host requirement is RFC 1122. Later specifications clarify the lineage and limits: RFC 2001, RFC 2914 and RFC 5681. LBNL's contemporary email index records the implementation discussion.
The 32-Kbps-to-40-bps figure describes a documented path during a series of collapses, not a census of every Internet link. “Saved the Internet” is therefore too total a claim. Attribution must remain plural, and the loss-based 1988 mechanism should not be mistaken for the end of congestion-control research.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
