Summary
- RFC 9411 connects security-effectiveness evaluation and performance benchmarking through the same testbed and device configuration. A shared model number does not supply that connection.
- A recommended protection can legitimately be omitted for a stated use case, but the reason and its performance implications must accompany the result. Disclosure and commercial acceptance remain separate decisions.
- Reusing a benchmark after a configuration or environment change requires an argument about what remains comparable, not an automatic assumption that the old result still covers the new system.
Two reports, one unresolved question
Consider a procurement scenario. A security team receives a report describing how an appliance handled a selected set of attacks. A network team receives a throughput report for the same model. Both reports look credible. The purchase table combines their conclusions into one row, and the row becomes evidence that the proposed system can provide the desired protection at the desired speed.
No number has necessarily been copied incorrectly. The difficulty is the word “and”. It joins two observations into a claim that the capabilities coexisted under relevant conditions. The model name cannot establish that. Neither can the fact that the documents came from a reputable laboratory. The configuration used for each observation has to make the connection.
This is a hypothetical purchasing problem, not a finding about a vendor or laboratory. Its importance is that an unsupported combined claim can emerge without either underlying report being false. One team may have answered its own question accurately while another assumes that the answer applies to a different question.
RFC 9411, an Informational IETF document published in March 2023, makes the connection unusually explicit. Its subject is performance benchmarking for next-generation network security devices. It replaced RFC 3511, but it is not an Internet Standards Track specification or a certificate that a particular product is secure.
The condition that travels with the result
Section 4.2 requires the same device or system configuration throughout the performance tests described in section 7. The device is to inspect traffic inline, with parameters and security features corresponding to an actual or typical deployment. Selected protections must remain consistently enabled across those tests, and a summary of the configuration, including enabled features, must accompany the results.
Appendix A carries this principle across the boundary between security and speed. Its effectiveness evaluation uses the same testbed as the performance tests. The device configuration must also remain the same. Matching client and server address ranges and relevant cipher settings further constrain the comparison.
The point is not that one immutable configuration will suit every buyer. It is that changing the work performed by the device changes the claim supported by a result. A forwarding path with a particular inspection policy is not interchangeable with a path that performs a different set of checks merely because both run inside the same enclosure.
“Security enabled” is too coarse a description to settle the matter. It may hide different feature selections, rule populations or treatment of encrypted traffic. The RFC acknowledges that implementations use different feature names and do not necessarily expose controls that align neatly with its taxonomy. A meaningful comparison therefore needs enough detail to establish what the selected functions actually did during the test.
An exception can be honest and still matter
The document does not require every available feature to be switched on indiscriminately. It distinguishes recommended and optional functions and allows a recommended feature to be absent when there is a reason. A particular deployment, for example, may not require a particular protection at that point in its architecture.
That flexibility is tied to reporting. The reason for omitting a recommended feature must be disclosed, along with the fact that the omission may affect performance. This is not an accusation that a faster configuration is improper. It is a limit on the comparisons that can be made from its result.
A buyer whose intended design genuinely matches the disclosed omission may find the result relevant. A buyer who intends to activate the missing function has another question to answer. The same report may be adequate evidence for the first decision and inadequate evidence for the second.
There is a useful distinction here between transparency and adequacy. A laboratory can make its scope perfectly transparent without deciding whether that scope is acceptable for a customer's operating needs. Publishing the exception transfers information; it does not transfer the customer's responsibility for accepting the difference.
The same reasoning applies to the document's explicit scope limitation. Its performance method is not intended for devices or systems relying on machine learning or behavioural analysis; it says such features should be disabled when present for this method. That is not advice to disable them in production. It means the resulting benchmark should not silently become a claim about a security stack whose behaviour depends on them.
A protection claim needs its own observation
RFC 9411 primarily addresses performance. It recommends validating the security-feature configuration through an effectiveness evaluation before performance benchmarking. If that evaluation is omitted, the report must explain the implications. A reference to the RFC alone therefore does not prove that every report includes that prerequisite assessment.
Appendix A observes more than a count of blocked attacks. Its measurements include unblocked vulnerabilities, background-traffic behaviour and the accuracy of the device's reporting. Its validation criteria include the absence of a false positive in the background traffic. This matters because a device that interrupts legitimate work creates a different operating burden from one whose controls perform as intended on that workload.
These observations are bounded by the selected test traffic and conditions. They are not evidence of protection against every attack, nor proof that a present deployment will produce no false positive. Even a correctly conducted effectiveness test cannot establish an unlimited security guarantee.
For the combined purchasing claim, however, they answer a narrower and essential question: was the configuration whose performance is being presented also the configuration subjected to the stated effectiveness evaluation? If the answer cannot be established, adding the reports together does not repair the gap.
The laboratory can become the bottleneck
A common device configuration is necessary, but it is not the only relevant condition. A test environment can constrain the measured result before the security device does. RFC 9411 therefore requires a reference test before the performance benchmarks and describes controls for the surrounding equipment.
The reference can use a trivial fast-forwarding setup or omit the device under test. It checks such matters as whether the generator can exceed the expected device performance and whether ancillary switching or routing introduces loss or delay. That stripped-down reference is a control for the testbed. It must not be confused with the security-enabled appliance's headline performance.
Virtualised environments make the attribution question especially visible. Host resources, interface acceleration and other workloads can affect the conditions in which the device runs. The document's reference-test guidance addresses stability during the session, including workload changes, virtual-machine movement and reduced processor performance under heat.
A changed host allocation might therefore explain why two tests differ without demonstrating that the security software itself improved or deteriorated. Conversely, a stable host does not prove that an altered protection policy is irrelevant. The questions are separate, and the report must retain enough context to distinguish them.
The practical value of this detail is not a longer appendix for its own sake. It is the ability to locate the unresolved comparison. If the generator constrained one run, investigating an inspection feature may not answer the problem. If the feature selection changed, a better generator does not restore configuration continuity.
The result also belongs to a workload
Endpoint behaviour and the offered traffic shape what the appliance is asked to do. RFC 9411 discusses transport parameters, application mix and encrypted-traffic settings because a device does not receive an abstract quantity called “load”. It receives particular connections, transactions and objects under particular conditions.
The method separates initialisation, ramp-up, sustain, ramp-down and collection. Measurements occur during sustain. The recommended minimum for that phase is 300 seconds, and the interval for collecting raw results and calculating statistics must be shorter than two seconds. Those are characteristics of the methodology, not a promise about how a customer application will behave during an operational transition.
The reporting guidance keeps benchmark indicators separate and asks for contextual information about the device, test equipment and workload. Even the protocol layer at which throughput is measured matters when results are compared. A number can be precise while the comparison made with it is poorly specified.
This does not make benchmarking futile. It makes benchmarking useful for a defined question. The more precisely the result describes the work performed, the less room there is for a later summary to turn it into a different claim.
The isolation boundary is equally practical. RFC 6815 explains why laboratory benchmark methods belong in an isolated test environment rather than on shared networks carrying live traffic. This article proposes no production stress test. A missing comparison is not permission to create an uncontrolled operational experiment.
Where the evidence stops
The analysis here rests on published methodology, not a newly measured product result. It does not identify a vendor that changed configuration between tests, quantify a speed penalty for enabling a feature, or rank appliances. It also does not establish how frequently buyers combine incompatible reports.
Its narrower conclusion is sufficient: the relationship between protection and performance is a property of the tested setup, not an attribute that a model name carries automatically.
That reading follows the distinction in Lu Heng's essay on reality rather than advocacy. The task is to describe the decision structure, not assign villains to it. His agency-problem essay supplies a question about who makes decisions and who bears their consequences; its claims about registry governance are not evidence against security laboratories.
A good report does not relieve a buyer of judgement. It makes the judgement possible. The critical handoff is the point where someone decides that the configuration represented by the evidence is sufficiently close to the configuration the organisation intends to operate.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
