Summary

  • An RFC 5180 number is evidence about a declared laboratory profile: frame sizes, destination distribution, prefix lengths, neighbor mode, traffic direction, header chain, filters, offered load and the device build all determine what ran.
  • Hop-by-Hop traffic exposes the boundary sharply. RFC 5180 offers it at 1%, 10% and 50% while watching resources; the purpose is processing impact, not an ordinary throughput score. RFC 8200 later made actual processing configuration essential to the claim.

The benchmark sheet says the router reached line rate. The purchasing committee hears that the network will carry its projected IPv6 demand. Between those statements sit all the variables the headline removed: whether the frames were 64 or 1518 bytes, whether destinations were fixed or random, whether one port or the full chassis was active, whether traffic ran in both directions, whether filters had to find Layer 4 information, and whether a difficult packet stayed in hardware.

RFC 5180 does not make that ambiguity inevitable. It supplies a method for exposing it. Published in 2008 as an Informational companion to RFC 2544, the document describes tests for IPv6-capable interconnect devices. It asks laboratories to preserve repeatability, variance and the statistical meaning of small trial sets. The result is not a universal certification. It is a receipt for one controlled experiment.

The profile is part of the measurement

Frame size is not a formatting detail. At a fixed link rate, small frames impose far more packet-processing events per second than large ones. RFC 5180 therefore names an Ethernet series: 64, 128, 256, 512, 1024, 1280 and 1518 bytes. Selecting the largest frame and printing only gigabits per second may hide a packet-rate limit. Selecting the smallest and applying it to an ordinary traffic mix can understate useful capacity. The honest claim carries the entire series and the distribution relevant to the decision.

Destination selection matters for the same reason. A single source-destination pair can exercise a warm, narrow lookup path. Random destinations across a range ask a different question. The method calls for both and identifies /48, /64, /126 and /128 as useful prefix boundaries. A table that drops the prefix population and randomness has detached the answer from the work the device performed.

So has a table that drops scale. Single-port testing measures performance at an interface. Multi-port testing tests forwarding scale across the platform. A line card can meet its target while a shared fabric, lookup resource or control path becomes the chassis constraint. “Tested at 400 Gbit/s” is incomplete until the reader knows how many ports, which directions and which shared resources participated.

One green number can hide the weaker direction

RFC 5180 recommends bidirectional traffic for all its tests. That creates load in both directions, but its aggregate result reveals only the lower-performing direction. When the device or configuration is asymmetric, unidirectional trials are needed to locate the constraint.

This is not statistical fussiness. Production demand is often asymmetric. A network may receive many small requests and return fewer large responses, or enforce a policy in one direction that it does not enforce in the other. The same total load can execute different pipelines. A procurement threshold based on the lower direction is conservative only if the traffic mix and business question match. Otherwise it is merely opaque.

IPv4 and IPv6 coexistence adds another distribution. The method includes IPv4-only, IPv6-only and 90/10, 50/50 and 10/90 mixes. A pure IPv6 trial cannot testify for shared lookup, buffer and policy resources under a mixed load. The percentage is not a caption. It is an input to the machine.

Hop-by-Hop is a path test, not another bar on the chart

The most revealing part of RFC 5180 is its refusal to call every result throughput. For most extension headers, the document recommends individual tests and a separate chain of headers, with a common smallest frame size for fair comparison. For Hop-by-Hop traffic, it changes the objective. The tester sends traffic at 1%, 10% and 50% of interface bandwidth and monitors device resources. The question is how processing affects the router, not how much ordinary traffic can be forwarded losslessly.

The historical wording needs a current boundary. RFC 5180 described Hop-by-Hop processing under the RFC 2460 model. RFC 8200 later said nodes along the delivery path are expected to examine and process the header only when explicitly configured. RFC 7045 had already warned that high-performance routers may ignore it or send it to a slow path. RFC 9098 explains further implementation choices: limited lookup depth, recirculation through a forwarding engine, software forwarding, control-plane pressure or drop.

Those paths produce different measurements from identical-looking packets. A current report must therefore record the device configuration and collect resource evidence out of band. CPU and memory telemetry taken independently of the forwarding interfaces can help distinguish hardware handling from a software punt. Throughput alone cannot make that distinction. Nor can a lab result prove resistance to an attack; the offered percentages, duration, option content and protection policy bound the result.

Neighbor state and policy are executable inputs

Even apparently routine preparation changes the test. RFC 5180 allows static neighbors or dynamic Neighbor Discovery, while preferring dynamic interaction that keeps caches active. It models endpoints one hop beyond the device to avoid storms caused by Neighbor Unreachability Detection. A static table, a constantly refreshed cache and a cache-expiry event do not ask the forwarding system the same question.

Filters change it again. A device that must locate transport information beyond extension headers may use a deeper parse, a different hardware stage or a slower path. Routing-table size, the number of filters, control traffic and management traffic also belong in the profile. A result obtained with empty policy cannot settle a deployment whose value depends on policy enforcement.

Physical assumptions deserve the same discipline. The RFC's Ethernet maximums are theoretical and allow plus or minus 100 parts per million of clock variation. Packet over SONET can vary with bit stuffing. A percentage displayed to two decimal places is not automatically more precise than the clock, medium and trial design permit.

A laboratory is intentionally not production

RFC 5180 requires an independent benchmark topology. Test traffic must not leak into production or the management network. The measurement is black-box and externally observable. The device should not contain special benchmark-only capabilities.

These controls improve validity by isolating the mechanism. They also prevent the report from claiming that production was measured. A production network adds routing churn, uneven flows, failure domains, queue interactions, operational policy, software drift and users. The laboratory removes many of those variables on purpose.

The reserved address space illustrates that separation. RFC 5180's published prefix contained a verified technical error; the current IANA registry identifies 2001:2::/48 for benchmarking and marks it not globally reachable. Correcting the prefix is not clerical trivia. It is part of keeping test evidence from being mistaken for an operational route.

Scope has other edges. RFC 8219 says translation and encapsulation transition technologies need a complementary methodology, including extra state and overload questions. A dual-stack device can use RFC 2544 and RFC 5180, but NAT64 or an encapsulation gateway cannot be reduced to the same receipt. Method names do not erase state tables.

What leadership should require

A benchmark presented for procurement or capacity approval should arrive with an immutable profile identifier. That profile should name the device and software build, ports and media, topology, tester and clock assumptions, frame-size series, destination and prefix distribution, neighbor mode, extension-header content, filters and routes, direction, IPv4/IPv6 mix, trial duration and count, sample distribution, loss and latency, out-of-band resource evidence, and every deviation from the method.

Recovery also needs separate labels. RFC 5180 treats recovery from overload differently from recovery after a device or software reset. It declines to recommend the back-to-back-frames test because short-term processing variation produced significant variance. Refusing an unstable number is a sign of methodological control, not a missing feature.

The governance structure should be equally explicit. One owner defines the method, another maintains the test configuration, another signs the statistics, a release owner proves that the tested build matches the candidate build, and an operations owner conducts safe production observation. No one receipt substitutes for the next.

The decision is not whether to trust benchmarks. It is whether to trust them for the proposition they measured. A line-rate result can be excellent evidence and still be irrelevant to a mixed, filtered, asymmetric, multi-port production load. Preserving that boundary protects both engineering and capital allocation: the laboratory keeps its authority, and the unmeasured system does not inherit it.

Sources

  1. RFC 5180 — HTML
  2. RFC 5180 — plain text
  3. RFC Editor information page
  4. IETF Datatracker document page
  5. IETF Datatracker history
  6. IETF Datatracker references
  7. RFC 5180 errata
  8. RFC 2544
  9. RFC 1242
  10. RFC 8200
  11. RFC 7045
  12. RFC 9098
  13. RFC 4861
  14. RFC 8201
  15. RFC 6890
  16. IANA IPv6 Special-Purpose Address Registry
  17. RFC 8219
  18. Heng Lu — reality layers
  19. Heng Lu — minimum initial specification and voluntary adoption
  20. Heng Lu — running code is primary