Summary

  • The 1998 Padhye–Firoiu–Towsley–Kurose model predicts steady-state sending throughput for a bulk, saturated TCP Reno flow from explicitly defined observations and assumptions.
  • Its throughput counts packets sent regardless of their eventual fate. The result is neither application goodput nor a measurement of link capacity, available bandwidth, fair share or authority.
  • The model’s value comes from preserving its provenance: observation window, loss-event grouping, RTT and timeout behavior, receiver window, ACK assumptions, TCP variant, validation regime and error.

The number that looked like a ceiling

Suppose a monitoring screen reports that a connection “should get” a particular number of bytes per second. The number is precise. It changes when loss rises or round-trip time falls. It may even track a long transfer rather well.

The temptation is to rename it. Predicted throughput becomes available bandwidth. Available bandwidth becomes capacity. Capacity becomes an entitlement: if the sender obtains less, somebody must have taken something away.

None of those steps is licensed by the equation.

In Modeling TCP Throughput: A Simple Model and its Empirical Validation, published at SIGCOMM in 1998, Jitendra Padhye, Victor Firoiu, Don Towsley and Jim Kurose studied the steady-state behavior of a bulk TCP transfer. The sender is saturated: it always has data to transmit. The transport is Reno in congestion avoidance. Loss indications and timeouts change its congestion window; round-trip time paces the cycles. The model predicts what that particular control process will send under stated conditions.

A bottleneck link is a different object. Other traffic may share it. Queueing changes over time. The receiver may constrain the flow. The application may stop producing bytes. Headers and retransmissions separate sending throughput from useful delivery. A policy or allocation system, not a formula, decides who is permitted to use a resource.

Precision does not erase those distinctions. It makes them more important.

Throughput was deliberately not goodput

The paper defines throughput as packets sent per unit time, regardless of their eventual fate. That phrase establishes a boundary that dashboards often remove.

A retransmitted packet contributes to sending throughput again. A packet later dropped has still occupied the sender’s transmission process. Application goodput asks another question: how much new, useful payload reached the relevant receiver over time? Capacity asks what a physical or logical resource could carry under a specified method. Available bandwidth asks what share is unused or obtainable during an interval. These quantities can influence one another without being synonyms.

The saturation assumption is equally consequential. A model of a sender that always has data cannot explain why an application-limited flow is quiet. If a telemetry system feeds the equation during idle periods and then labels the result “expected utilization,” it has mixed a counterfactual control rate with actual demand.

The model did not commit that error. It named its regime. Later users can create the error by dropping the regime from the record.

A loss event was not one lost packet

The loss input is not a context-free percentage. Reno reduces its congestion window in response to a loss indication. Multiple packet drops in one window can belong to one congestion episode rather than several independent control events.

The 1998 analysis treats losses as independent between rounds and correlated within a round, a pattern the authors relate to drop-tail queues. Its probability concerns the event that affects the window. Changing the grouping rule changes the input even when the same packet trace remains underneath it.

This matters operationally. One observer may count every missing sequence number. Another may group losses within an RTT. A third may use receiver reports with a different clock and incomplete reordering information. All can publish a decimal called “loss,” yet feed the equation materially different facts.

An auditable estimate therefore needs the raw marks or a durable reference to them, the observation interval, the event-grouping rule and the transport behavior that gives those events meaning. Without those records, two numbers cannot be compared merely because both use a percent sign.

Timeouts carried much of the explanation

A simpler model can explain Reno only through triple-duplicate acknowledgements. Padhye and his collaborators found that insufficient. Their traces contained retransmission timeouts, often more timeout events than fast-retransmit events. The fuller model and its approximation therefore account for both paths.

That inclusion changes more than a coefficient. A timeout pauses the sender and invokes a recovery process different from a fast retransmit. The retransmission timeout estimate, delayed acknowledgements and the maximum receiver window all shape the predicted rate. The familiar approximate equation compresses those behaviors, but it does not make them disappear.

The paper also states simplifying assumptions. A round lasts one RTT, and the current window can be sent inside that round. Slow start is negligible in the steady-state calculation. Not every subtlety of fast recovery is modeled. These are reasonable choices for an analytical instrument; they are not natural laws.

When the transport stack, acknowledgement policy or workload changes, the right question is not whether a famous formula remains famous. It is whether the evidence still fits the instrument.

Thirty-seven connections made a test, not a universe

The empirical validation was serious. The authors studied 37 TCP connections among 18 hosts in the United States and Europe. Twenty-four traces ran for an hour. Thirteen additional data sets consisted of serial 100-second connections. The traffic was unidirectional bulk transfer from an infinite source.

Across those observations, the model generally described throughput better than the version that considered only triple-duplicate loss recovery. The approximate equation followed the fuller model closely enough to be useful. That empirical work is a major reason the paper endured; ACM SIGCOMM lists it as the 2008 Test of Time Award recipient.

But the validation did not turn 37 connections into every path. A modem case exposed a mismatch: the dedicated buffer, round-trip time and window interacted in a way the model did not capture well. TCP implementations on Linux, Irix and SunOS also differed, while the model was not customized to every implementation.

The authors left further work visible, including fast-recovery detail, window evolution, loss distributions, slow links and implementation effects. This is what accountable modeling looks like. A failed fit is not an embarrassment to remove. It marks the edge of the instrument.

TFRC reused an estimate inside a control loop

RFC 5348, written by Sally Floyd, Mark Handley, J. Padhye and J. Widmer, later used a slightly simplified form of the Reno equation in TCP Friendly Rate Control. TFRC measures a loss-event rate and RTT, calculates an allowed sending rate and adjusts a smoother, rate-based sender.

The RFC is careful about the noun. The equation roughly describes TCP’s sending rate. It defines variables such as segment size, RTT, loss-event rate, timeout and packets per acknowledgement. It limits the calculated result relative to the measured receive rate. That additional cap is itself evidence that the equation is one component of a feedback system, not an oracle for the path.

TFRC sought reasonable coexistence with TCP. The RFC describes a rate generally within a factor of two of a conformant TCP flow under the same conditions, not numerical identity. It also describes the trade-off: smoother rate variation, slower reaction to changes in available bandwidth. TFRC is not a reliability protocol.

So the later standard does not enlarge the original estimate into capacity. It gives the estimate a job: constrain a sender in a defined control regime.

The name at the center belongs to four authors

Microsoft Research lists the 1998 publication under “Jitu Padhye,” with Firoiu, Towsley and Kurose. The paper itself uses Jitendra Padhye. An official 2010 Microsoft Research article identifies him in a group photograph and describes his role at that time. Those sources establish identity and dated context; they do not justify inventing a current title.

Padhye is the biographical center here because the model is widely attached to his name and because he later appears among the TFRC authors. Yet the equation commonly abbreviated PFTK carries four initials for a reason. The analytical model, its empirical validation and the paper’s limits are collective work.

That credit boundary is not separate from the article’s argument. A number without provenance can be made to speak for anything. A model without its authors, version and validation regime can be made to predict anything. Preserving attribution is one way of preserving scope.

The estimator needs its own ledger

A responsible deployment should be able to answer a sequence of questions:

  • Which raw observations formed the loss events, and over what interval?
  • Which RTT distribution and timeout estimate were used?
  • What segment size, ACK behavior, maximum receiver window and transport variant applied?
  • Was the sender actually saturated, or was the application limiting it?
  • Which equation and caps produced the result?
  • How did predicted sending throughput compare with actual sent rate and delivered goodput?
  • When the error grew, which assumption stopped holding?

The record should keep those answers together. Publishing only the final rate discards the means to reproduce or challenge it. Calling the result “capacity” then adds an unsupported meaning at exactly the moment the supporting evidence has disappeared.

Heng Lu’s Running-Code Primacy provides a useful present-day lens. A configured model or stored estimate is subordinate to observed transport behavior, while a live observation is itself bounded by its vantage point and measurement method. This is an editorial comparison, not a claim that the 1998 or 2008 authors followed a later doctrine.

The practical rule is modest: let the equation say what it was built to say. It predicts a long-run sending rate for a defined transport under defined evidence. Keep the inputs, the assumptions and the residual error. Measure capacity separately. Measure delivery separately. Record allocation separately. The model becomes more useful, not less, when it is not asked to impersonate them.

Sources