Summary
draft-ietf-bmwg-powerbench-03defines a controlled laboratory method, not an operational energy-management system or a product ranking. It joins an external input-power meter to named device, traffic and environmental conditions.- Its “Idle+” condition sends one bidirectional packet per second on every active interface. That deliberately small stimulus can wake the forwarding plane, exposing why Base, Idle, Idle+, Typical and loaded measurements are test conditions rather than universal device Power States.
- A comparable Energy Efficiency Ratio needs both sides of
T/P: predeclared traffic-load weights and successfully forwarded throughput in the numerator; measured power, meter accuracy, stabilization and averaging windows in the denominator.
The smallest trace exposes the largest ambiguity
A router is fully configured. Every interface is up. No user traffic crosses it. The power meter settles. That is PowerBench's Idle condition. Now the tester sends one packet per second in both directions on every active interface. The rate is intended to be too small to create measurable dynamic packet-processing power, yet enough to activate the forwarding plane. That is Idle+.
The difference can be only a handful of packets. Operationally, it can cross a hardware boundary. A component may leave a lower-power state; a forwarding pipeline may become active; optics, control processes or cooling behaviour may change. If two reports both say “idle” while one used no packets and the other used this minimum trace, their watt figures describe different propositions.
That is the most useful insight in revision 03 of Characterization and Benchmarking Methodology for Power in Networking Devices. The draft is not merely collecting energy numbers. It is trying to build a comparison contract around them.
The status matters. The Datatracker record identifies an active Benchmarking Methodology Working Group Internet-Draft, with revision 03 uploaded on 30 September 2026. The document header says intended status Standards Track; the frozen document API does not carry an intended-level value. It is not an RFC, an approved standard or proof that any named device has passed.
Five conditions are not five machine states
PowerBench defines Base, Idle, Idle+, Typical and Power with Traffic Load. Base starts from factory settings after boot, with boards and components active but no transceiver installed. Idle uses a fully configured forwarding device with all interfaces up and no traffic. Idle+ adds the one-packet-per-second trace. Typical uses a stated share of maximum throughput, such as 30 per cent, with an RFC 6985 IMIX packet-size distribution. The loaded test drives named ports or line cards at a specified percentage of maximum throughput.
Revision 03 adds an explicit boundary: these are measurement conditions, not Power States. A device can remain in the same internal state across two conditions, or transition between states inside one condition. Vendors may name and implement states differently. When a state is externally visible or deliberately configured, it should be reported together with the mechanism that configured or verified it. That state evidence supplements the benchmark; it does not replace the procedure.
This correction is more than vocabulary. Calling Idle a machine state invites a report to borrow authority from a label. Calling it a measurement condition asks what configuration, interfaces, background functions, traffic and time window actually existed. In the draft's Idle and Idle+ procedures, routing adjacencies, management processes, telemetry and thermal controls may remain active. Their operating condition must be disclosed because they can move the reading.
The difference also keeps this work separate from a sleep-control protocol. PowerBench does not command a link to sleep, preserve traffic-engineering state or prove a wake path. It observes a device under controlled stimuli. A state transition may explain a result; it is not itself the result.
The ratio has a negotiated numerator
The proposed Energy Efficiency Ratio is simple on paper: EER = T/P, expressed as Gbps per watt. The apparent simplicity disappears when T and P are unpacked.
T is a weighted sum of interface throughput across selected load levels. P is a weighted sum of the corresponding power readings. The load levels and weights must be established in advance. The draft gives examples such as 100, 30 and zero per cent load, with weights of 0.1, 0.8 and 0.1, while allowing different choices for access routers, core routers and data-centre switches.
Those choices are not bookkeeping around the answer. They construct the answer. A high-load-heavy weighting asks how well a device performs in one usage profile; an idle-heavy weighting asks another. An optional calculation can even substitute total interface capacity for weighted capacity. The result remains proportional, but it is not the same numerator. A procurement table that sorts all of these as one “Gbps/W” column can produce a precise ranking from incomparable contracts.
Useful work must also arrive. The draft insists that traffic leave the correct port. Dropping packets is an easy way to reduce power and an invalid way to win. The default is zero packet loss, aligned with the Non-Drop Rate tradition of RFC 2544. Any non-zero loss tolerance must be declared and justified. A cited procedure further requires the equipment to return to full NDR load; failure disqualifies the result.
That condition protects the denominator from consuming the numerator. A lower watt figure is not efficient if the DUT silently avoided the work being scored.
The meter has a jurisdiction
The test setup places an external meter at the DUT's power input. It measures the electrical power drawn there. It does not measure external cooling infrastructure. Heat can still matter because device fans and thermal controls consume input power, and temperature drift can alter their behaviour. The report therefore needs the environmental and thermal context without pretending that the meter covers the whole facility.
PowerBench prescribes a laboratory range of 23–27 °C, 25–75 per cent relative humidity and 812–1060 hPa atmospheric pressure. It also requires an applicable meter-accuracy specification appropriate to the measured range. Revision 03 adds both the setup requirement and the report field for that accuracy.
Accuracy is not calibration history, and neither is a guarantee that two results are close enough to rank. Suppose two devices differ by one watt while their meters, ranges or averaging methods leave a larger uncertainty. The table can still display different numbers. The evidence cannot yet sustain the ordering. A serious comparison keeps raw readings, instrument identity, accuracy specification, range and any calibration evidence available to the lab.
The scope clause added in revision 03 is equally important. This is a laboratory benchmark. It is not an operational energy-monitoring or management framework. The lab result can complement live measurements, but it cannot predict a fleet's workload mix, cooling overhead, redundancy policy, software drift or electricity source merely by travelling from a report into a dashboard.
Time is part of the measurement surface
Power changes after a traffic or configuration transition. Queues fill, caches warm, fans respond, clock rates move and control processes settle. Measuring immediately and averaging for a short interval can record a transient. Waiting longer can record another operating point. Neither number is self-explanatory.
The draft therefore requires two clocks. The stabilization interval runs from application of the load or configuration to the start of measurement. The measurement interval is the averaging window used to produce the reported input power. The averaging method must be named, and intervals must be reported per load level when they differ.
Revision 03 does not impose one universal duration. Device class, implementation and laboratory conditions vary. That choice improves feasibility and leaves uncertainty. Comparability then depends on disclosure: a report must say enough for another lab to understand whether its window captured the same behaviour.
This is where longitudinal benchmarking becomes valuable and dangerous. Repeating the test on the same chassis before and after a software upgrade can reveal a genuine change. But the comparison is only as clean as its retained configuration, trace, environment, firmware, transceivers, background functions and timing. “Version B used less power” is a claim about a controlled difference, not a property that follows automatically from the version label.
Programmable forwarding makes software part of the appliance
For a fixed-function device, hardware and software version were already reportable. Revision 03 goes further for programmable data planes. The installed program or workload, compiler or toolchain version, and a high-level description of tables or stateful operations should be recorded.
That addition recognizes an executable fact: the same physical switch can implement a different forwarding machine after recompilation. Table depth, match structure, counters, registers and stateful operations can change memory access and pipeline activity. A chassis name and port count no longer identify the tested object tightly enough.
The report therefore becomes a passport for one bounded execution context. It should name hardware, operating software, line cards, enabled and active ports, interface and transceiver types, port settings, utilization, trace, programmable workload, compiler, environmental conditions, meter, intervals and result. Remove enough of those fields and the benchmark becomes a marketing noun detached from the machine that produced it.
What revision 03 changed—and what it did not prove
The history shows revision 03 arriving three months after revision 02. The textual change is unusually coherent. It adds the laboratory-versus-operational boundary, meter accuracy, a clearer input-power boundary, programmable-data-plane context and the distinction between measurement condition and Power State. It also changes language that previously implied a known “low-power mode” into a more careful statement that the test may capture a transition.
Those edits improve the evidence grammar. They do not establish a winning device, a deployed implementation, an interoperability event, an achieved saving or a carbon outcome. The frozen sources contain none of those. The draft itself accepts a trade-off: many disclosed results with higher uncertainty may be more useful than a tiny number of perfectly prescribed tests that almost nobody can compare.
That choice makes disclosure the core control. PowerBench is strongest not when it promises one universal number, but when it makes the conditions behind each number difficult to hide.
Sources and limits
The analysis uses the frozen revision 03 text, HTML and XML, its Datatracker record and history, revision 02, RFC 2544, RFC 6985, RFC 6988, RFC 7460 and the referenced GREEN terminology draft. No private lab data or vendor evidence was used.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
