Summary
- RFC 9971 specifies MLRsearch, a laboratory benchmarking method that can search several stated loss-ratio goals and report goal-specific bounds for a stated system and traffic profile.
- A reported conditional throughput is evidence about the procedure that produced it: its system boundary, trial duration, offered load, loss goal, traffic profile and result width. It is not a production capacity promise or an application outcome.
- The sound operating pattern is to preserve the complete benchmark record, define the separate local decision it may inform, and require production evidence before turning a test result into a service, procurement or change claim.
The number needs its own passport
RFC 9971 calls its method Multiple Loss Ratio Search, or MLRsearch. The name is useful because it refuses a familiar simplification. The benchmark is not merely an exercise in discovering “the throughput” of a box. It searches one or more explicitly defined goals using trials at chosen offered loads and durations. Its purpose is to make results more repeatable and comparable while reducing search time, particularly for software data-plane functions running on commodity hardware.
That purpose matters. A result becomes comparable only when another team can see what it was comparing. Which system was stimulated? What traffic profile was sent? Were the trials short or long? Was the permitted loss exactly zero, low but non-zero, or defined in more than one goal? What result width was accepted? Which observations formed the lower and upper bounds? Without those fields, a capacity figure is not a result with provenance. It is a statement looking for an apparatus.
The RFC is deliberately modest about status. It is Informational, not an Internet Standard. Its use of BCP 14 language makes a claimed MLRsearch procedure unambiguous; it does not compel an operator to use MLRsearch or certify any product. The document also says that it is independent of RFC 2544: it neither changes nor obsoletes that older benchmark methodology. A procedure can be configured to satisfy RFC 2544 conditions in a particular case, but that requires its own documented conditions. A familiar RFC citation is not a substitute for the report.
The system boundary is part of the result
The distinction between a Device Under Test and a System Under Test is not bookkeeping. RFC 9971 adopts the classic terminology in which a DUT is the forwarding device or function being examined, while an SUT is the collective system to which traffic is offered and from which response is measured. On a software platform, the SUT can include the host, firmware, operating system, hypervisor, I/O resources and co-resident workloads. A nominally unchanged forwarding function can therefore produce different trials when its surrounding system changes.
The RFC calls the hard-to-separate effects noise. Some is plainly environmental: a shared CPU, an interrupt pattern, memory pressure, a scheduler decision or another workload. Some may arise inside the function. The method does not pretend a short experiment can always allocate each lost frame to its ultimate cause. Instead, it treats trial variability as an engineering fact that has to be described and bounded.
That is the first authority boundary. A benchmark may describe a configured SUT. It does not, merely by naming a DUT, establish an intrinsic and context-free performance property of software. It certainly does not establish the capacity of an entire production service whose traffic distribution, failure modes, routes, retries, storage paths, credentials and client behaviour were absent from the trial.
MLRsearch trades an undefined universal number for stated goals
The older benchmark vocabulary gives a sharp definition of throughput: the maximum rate at which no offered frames are dropped. That remains useful. It is also difficult to apply as though every modern software function had one stable zero-loss frontier. At high rates, a rare short fluctuation can cause a loss. Repeating a long trial can then expose different parts of the observed performance spectrum. A trial at a higher load may even report a lower loss ratio than an earlier trial at a lower load.
RFC 9971 addresses this not by claiming to discover the true essence of a machine, but by defining search terms and a reporting structure. A Search Goal has stated parameters, including a goal loss ratio and duration conditions. The Controller repeatedly selects Trial Inputs; the Measurer runs them; results classify a load relative to each goal. The Manager configures the exercise and produces the Test Report. Those are abstract roles. A single toolchain can embody all three. Their separation makes the decisions visible: who chose the experiment, what was measured, and what was reported.
For a regular result, relevant lower and upper bounds meet the stated goal width. The result is conditional in a precise sense. It says that this procedure found these bounds under this goal, not that reality now has an eternal capacity ceiling. The report is therefore not an administrative afterthought. It is the object that connects a number to its meaning.
MLRsearch can also carry multiple loss goals in one search. That can be more honest than forcing a single loss tolerance to impersonate every use case. Low non-zero frame loss may be meaningful in one controlled engineering question and intolerable in another. The RFC says there is no industry consensus on a best universal goal loss ratio, and notes that the relationship between trial loss and higher-layer performance is not simple. A TCP application, a real-time flow, a replicated storage protocol and a security control can have different consequences from the same observed loss ratio.
The correct conclusion is not that any loss budget is acceptable. It is that the party using the result must name its loss goal, justify it for its purpose and keep that justification with the result. A method provides a common grammar. It does not take ownership of the later value judgement.
Repeatability is evidence of procedure, not a warranty of the world
Repeatability is attractive because it reduces embarrassment. If two teams run the same procedure after a small change and see the same bound, they can detect a regression with more confidence. But a repeatable test can still have a narrow field of view. It can repeat an incomplete traffic profile, an unrepresentative SUT configuration or a laboratory topology that production will never resemble.
That is why RFC 9971's concern with comparability should be read as a discipline of disclosure, not an argument from laboratory neatness to operational certainty. The method leaves the Controller's internal load-selection heuristics implementation-specific. It does not prescribe how to handle every mismatch between intended and offered load. It does not set one universal Search Goal configuration. These are not omissions to fill with marketing language. They are local decisions that must remain attributable to the organisation that makes them.
Heng Lu's Running-Code Primacy points in the same direction. The meaningful evidence is the code and configuration that generated the trial, the traffic actually offered, the recorded frames and the precise report—not the prestige of a benchmark label. Minimum Initial Specification and Localized Future Decision add the governance rule: a standardised description lets different parties communicate; it does not export a local acceptance threshold or a local production decision to everyone else.
Build a bridge instead of smuggling an inference
An organisation can use a benchmark well by separating four records.
- The benchmark record names the software build, hardware and virtualisation boundary, traffic profile, frame sizes, direction, load units, trial durations, loss goals, goal widths, time, tool versions and raw results.
- The interpretation record says what narrow engineering proposition the result supports: for example, that a specified configuration did not regress against a defined laboratory baseline under the stated goal. It also records what it does not support.
- The production evidence record contains telemetry and service canaries from the actual deployment scope: queueing, CPU and I/O pressure, retries, route changes, error budgets, application transactions and a time-bounded comparison with the benchmarked configuration.
- The decision record identifies who can approve a release, buy capacity, alter a customer commitment or invoke a rollback; it states thresholds, exceptions, expiry and the evidence required to close the decision.
The bridge must be built deliberately. A benchmark can be a useful input to a procurement test, but procurement needs a workload and acceptance plan. It can inform capacity planning, but planning needs peak distributions, growth assumptions, fault margins and cost ownership. It can gate a release, but release safety also needs compatibility, security, rollout and rollback evidence. It can be considered in an SLA review, but an SLA is a promise about a defined service and period, not a remembered packet trial.
A conservative bound is not a production reservation
The tension is sharpest when a team needs a fast answer. A result with a conservative relevant lower bound feels ready for conversion into a limit: “we tested 40, therefore reserve 40.” But a capacity reservation has a different object. It binds future scheduling and the rights of other users. It has to address concurrent workloads, policies, maintenance, failure, dependency behaviour and the range of traffic actually admitted. RFC 9971 has not observed all those states just because it has searched carefully within one stated SUT.
The same warning applies in the other direction. A benchmark miss need not prove production failure. It can justify investigation, a workload comparison, a reproducibility check or a risk hold. It does not identify a supplier breach or an operational root cause without evidence that links the laboratory condition to the deployed effect.
This is not a plea to make tests bureaucratic. It is a way to keep them fast. The Measurer can run a controlled trial quickly. The Controller can converge on a declared goal. The Manager can publish a complete report. None must wait for a commercial committee. The wider decision can then move at its own appropriate speed, using the benchmark as a well-labeled input instead of a misleading conclusion.
Sources
- RFC 9971 — Multiple Loss Ratio Search
- RFC 9971 record
- IETF Datatracker — RFC 9971
- RFC 2544 — Benchmarking Methodology for Network Interconnect Devices
- RFC 1242 — Benchmarking Terminology for Network Interconnection Devices
- RFC 2285 — Benchmarking Terminology for LAN Switching Devices
- RFC 6349 — Framework for TCP Throughput Testing
- RFC 2119 — Key words for use in RFCs
- Heng Lu — Running-Code Primacy
- Heng Lu — Minimum Initial Specification and Localized Future Decision
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

