Summary

  • Dave Täht treated a network benchmark as a claim with inspectable conditions: mixed bidirectional load, sufficient duration, recorded latency and fairness, raw output, and explicit reasons to reject a run.
  • RRUL, the CeroWrt testbed and Flent shifted authority away from a single speed number toward an experiment another operator could repeat, contradict and improve.

The useful red trace

Picture two identical routers returning closely matched results while a third glows red because its processor is saturated. The red trace is not an embarrassment to delete. It says that the apparatus, rather than the queueing mechanism under study, has become the bottleneck. Keeping that run out of the comparison protects the conclusion. Keeping the reason preserves the lesson.

This was the kind of discipline Dave Täht brought to the campaign against bufferbloat. The familiar symptom was maddening: a link could report excellent throughput while a call stuttered, a game froze or a web request waited behind bulk traffic. A conventional speed score flattened those simultaneous experiences into one flattering number. Täht’s sharper intervention was methodological. Load the path in both directions, introduce unlike kinds of traffic, watch delay as well as delivery, and retain enough of the experiment for somebody else to challenge it.

A workload that refuses to behave politely

The Realtime Response Under Load specification, or RRUL, makes the network answer under stress. Its bulk TCP flows fill capacity in both directions while measurement streams observe latency and other traffic tests probe what happens to shorter or more time-sensitive work. The point is not to imitate every household or backbone. It is to deny a queue the easy conditions under which a defect remains invisible.

That changes what counts as a result. Throughput still matters, but it sits beside delay, jitter, fairness between flows, IPv4 and IPv6 behaviour, TCP and UDP, and the treatment of traffic classes. A device that moves many bits while trapping a small request behind them has not delivered an uncomplicated victory. Nor has a device proved itself if its initial acceleration vanishes once a sufficiently long run outlasts the boost.

RRUL also makes room for invalidity. If the test host runs out of processor, the observed ceiling may belong to the host rather than the router. If a run is too short, a transient vendor feature can masquerade as sustained capacity. Duration, load generators, direction and competing flows are therefore evidence, not laboratory trivia. They define the proposition that the graph can support.

The testbed as a public argument

CeroWrt gave the inquiry an inspectable place to happen. Built as open-source router software on known hardware, it made a test environment available for repeated examination rather than hiding a result inside an opaque appliance. Its project history also protects authorship from hero mythology: it credits CoDel to Kathleen Nichols and Van Jacobson, identifies Eric Dumazet’s flow-queueing work, and describes changes moving into Linux and OpenWrt. Täht was an organizer and experimental provocateur inside a collaborative system, not the solitary inventor of every mechanism it tested.

The fixed platform mattered because an experiment needs something that can be reconstructed. A software revision, configuration, interface, traffic source and hardware limit are parts of the result. When one changes, the comparison changes. CeroWrt was valuable not because fixed hardware represented the whole Internet, but because it reduced one region of uncertainty enough to let queueing behaviour become visible and debatable.

The raw result survives the plot

Flent completed another part of the chain. Developed from Toke Høiland-Jørgensen’s netperf-wrapper, it coordinates predefined and batched network tests, combines the measurements and renders them. More importantly, it stores data in compressed JSON that can be processed and replotted later. The polished image is therefore an interface to the evidence, not the only surviving evidence.

Täht’s 2014 SIGCOMM slides made the ethic explicit. Preserve raw results. Keep a virtual machine or testbed capable of repeating the work. When a fault reveals that old results cannot be trusted, discard them and rerun the experiment. Negative results belong in the record because they define where an explanation stops working. The distinction between repeating an experiment in the same apparatus and reproducing it through an independent reconstruction matters, but both demand more than a screenshot.

The approach turns disagreement into work. A critic can change the duration, isolate a CPU ceiling, compare another queue, add a competing flow or replay the stored data. A result no longer asks for confidence in the experimenter’s reputation. It offers handles by which confidence can be earned or withdrawn.

From experiment to shared mechanism

RFC 8290 shows one downstream form of that process. Published on the Experimental track in 2018, it describes FQ-CoDel, combining flow queueing with CoDel to separate flows and manage standing delay. Toke Høiland-Jørgensen heads the byline; Paul E. McKenney and Dave Taht follow, together with Jim Gettys and Eric Dumazet. Those five names are a reminder that operational techniques mature through distributed work. The RFC does not retroactively make every earlier benchmark correct. It makes a mechanism precise enough for more implementers to examine.

An APNIC account of Täht demonstrating latency under load captures why the public experiment mattered: people could feel a loaded connection become responsive, not just watch a score improve. Høiland-Jørgensen’s later remembrance adds the human practice behind the apparatus—teaching, mentoring, argument and an insistence on open inquiry. Those accounts are context, not substitutes for the specification, preserved data and repeatable procedure.

Sources