Summary

  • draft-ietf-mpls-stamp-pw-21 says punt-path rate limiting of STAMP packets can be indistinguishable from actual network loss and can therefore be reported as packet loss.
  • A missing reply needs an attribution receipt covering encapsulation, LSP or pseudowire context, both endpoint punt paths, rate-limit counters, reflection, reverse capacity and MTU before it becomes a forwarding-failure claim.

The most dangerous alarm is not a false one. It is an alarm whose number is correct while its explanation is wrong.

Revision 21 of Encapsulation of Simple Two-Way Active Measurement Protocol for LSPs and Pseudowires in MPLS Networks extends STAMP into point-to-point MPLS label-switched paths and single-segment pseudowires. It defines packets with IP/UDP headers and packets without them, using MPLS Generic Associated Channels so that tests can follow the relevant forwarding context. The stated objective is familiar: measure delay, delay variation and loss in a way that resembles the traffic under observation.

The draft also exposes a harder operational truth. STAMP OAM packets are intercepted at provider-edge endpoints for exception processing. Each packet that reaches a Session-Sender or Session-Reflector can consume control-plane CPU and memory. Rate limiting is therefore not a defect; it is a required protection against denial of service and accidental overload.

Then comes the sentence that should change the incident workflow: rate limiting on the punt path is indistinguishable from actual loss in the network and can be reported as packet loss.

The counter has not lied. A packet was sent and the expected reflected packet was not accepted in time. What remains unknown is where the evidence chain broke. The forward LSP may have dropped it. The receiving PE may have punted it into a full queue. A limiter may have discarded it before the reflector ran. The reflector may have lacked reverse LSP or pseudowire context and therefore discarded the request without reply. The reverse direction may have had less capacity. The returned packet may have been limited at the sender.

Or the extra GAL and G-ACh overhead may have pushed the test above a directional MTU, where the draft says it is dropped rather than fragmented and can again appear as packet loss.

One symptom now represents several control surfaces.

Format 2 makes the identity problem especially visible. Without IP and UDP headers, the usual four-tuple is absent. A non-zero STAMP Session Identifier must be combined with the received LSP or pseudowire context, reverse-direction context and locally provisioned session parameters. The SSID is not a universal object name. It becomes meaningful inside the context held by each endpoint. If the provisioning records disagree, a valid-looking test packet can fail to identify a session and must be discarded.

The forwarding comparison has conditions too. The draft uses the same label stack and, where relevant, the control word so that the test should experience the same forwarding and ECMP behavior as the measured data. But Format 1 uses IP addresses that may differ from the data flow. Non-conforming equipment can include the GAL in a way that changes path selection. A correct encapsulation narrows the gap; it does not erase every device-dependent decision.

This Article does not repeat the broad warning that a probe is not automatically equivalent to a production flow. The narrower issue is attribution after the probe itself enters a protected processing path. The test may have followed the intended labels and still have been discarded because the diagnostic channel consumed its control-plane allowance.

The useful receipt begins before transmission. Record the configured test rate, packet size, format, label stack, channel type, SSID and exact LSP or pseudowire. Preserve forward and reverse MTU and allocated measurement bandwidth. At both endpoints, collect punt-queue admission, limiter policy, limiter occupancy and drops, reflector input, reverse-context lookup, reflector output and sender receive processing. Join those records by sequence and time window. Only then compare them with data-plane counters and the affected application.

The draft explicitly recommends that operators know when rate limiting was applied so alerting can correlate that fact with failure notifications. That requirement should not become another loose log line. It needs a causally ordered record. “Limiter dropped 23 packets” is still insufficient if the system cannot show that those packets belonged to the STAMP session that triggered the incident.

Rate policy also has two directions. A reflector produces a comparable return stream, and the reverse LSP or pseudowire may have less capacity than the forward direction. The packet size, including encapsulation overhead, must fit the MTU independently in both directions. A symmetric test payload does not imply symmetric network conditions.

This changes how automation should escalate. A missing reply can open an investigation. It should not immediately assert customer-traffic loss or a failed LSP. The first branch asks whether the test was admitted, processed and returned at both endpoints. Only after measurement-path causes are excluded should the system attach the stronger forwarding-failure label.

Revision 21 is dated 10 September 2026 and remains an active Standards Track Internet-Draft. It does not establish deployment or an observed incident. Its value is that it names the ambiguity before operators encode a loss counter as an unquestioned cause.

The attribution receipt

A defensible incident record separates observation from location. It carries the sender sequence, packet format and context; forward delivery evidence; reflector punt and limiter state; reflection decision; reverse context and transmit; sender punt and limiter state; directional MTU; and the final calculation. Missing links stay marked unknown.

Protection is part of the measurement system

Control-plane limiting cannot simply be disabled to make graphs cleaner. Without it, a diagnostic tool can become an overload vector. The correct design makes the protective action observable and keeps its threshold consistent with the bandwidth allocated to the test. A safe measurement is one whose own safeguards leave evidence.

Sources