Summary

  • draft-xiao-fann-fast-cnp-with-proxy-05 lets a congested node send one UDP notification to a proxy, which reconstructs a second congestion packet understood by a RoCEv2 sender.
  • The sender-facing CNP is useful feedback, but it does not disclose the proxy's flow/QP mapping version, the original congestion-node identity, or notifications the proxy dropped under local rate policy.

The sender sees a standard RoCEv2 Congestion Notification Packet and slows the affected queue pair. From its point of view, the signal is familiar. Somewhere upstream, however, no standard CNP was sent to it. A congested router emitted a different UDP message to a proxy. The proxy looked up state learned from traffic, chose a sender and Source QP, then created the packet the sender finally received.

That translation is the point of Fast Congestion Notification Packet with Proxy, revision 05, posted on 29 September. A VPN provider router may live in a different routing domain from the sender. It may also lack the Source-QP information needed to construct the ordinary RoCEv2 CNP the sender understands. The proxy bridges both gaps.

It also becomes a new evidence boundary. The second packet can be syntactically correct and operationally useful while revealing little about how the proxy produced it.

One congestion event becomes two packets

The design has a first leg and a second leg. On the first, the congested node sends a dedicated notification to a proxy inside the same controlled routing domain. On the second, the proxy sends a notification in a format supported by the traffic sender. Revision 05 narrows the text around the standard RoCEv2 case rather than claiming a universal translator for every sender.

Packet format 1 carries the data flow's IP five-tuple and 24-bit Destination QP to proposed UDP port TBD1. The proxy uses the source address to locate the sender. It uses protocol and destination port to identify RoCEv2, and must then create a standard RoCEv2 CNP. To identify the sender's Source QP, it consults a Source-QP/Destination-QP mapping maintained in source and destination address context.

That mapping is not configuration alone. The selected proxy must sit where it sees both forward and reverse RoCEv2 data traffic so it can learn the association. The packet received from the congestion point contains the Destination QP, not a ready-made Source QP. Translation therefore depends on retained observation state.

Packet format 2 carries the five-tuple and an NRP Selector ID to a different proposed UDP port, TBD2. In a RoCEv2 case, the selector helps identify Source QP. In a VPN case, it helps identify the VPN. The proxy now depends on Source-QP/NRP or VPN-ID/NRP mappings. The selector is a lookup key, not proof that the lookup table is current.

Capability is not current translation state

The congested node first needs to know which proxy covers the sender. The proposal advertises Proxy Node Capability with prefixes attached to traffic senders. IS-IS and OSPF use proposed P-flags. BGP uses a proposed Next Hop Dependent Characteristic TLV carrying the proxy's address.

Those signals answer a routing question: which node claims to provide proxy service for this prefix? They do not answer the data-plane questions that determine whether a translated CNP is correct. They do not show that the proxy observed both traffic directions, learned the relevant QP pair, retained the current NRP or VPN association, has rate-limit headroom, accepted the first notification or delivered the second.

This distinction matters because the advertisement can remain stable while the learned mapping ages, moves or collides. A capability bit is not a health check. Propagation across an area or IS-IS level proves that the advertised bit survived routing distribution; it does not certify the tables behind it.

The proxy may intentionally discard evidence

The draft says the proxy would normally send a second notification each time it receives a first one. It immediately adds an exception: if receipt frequency exceeds the proxy's limit, it can drop some first notifications under local policy.

Rate limiting is reasonable. An unauthenticated-looking flood of congestion packets could itself exhaust the control surface. The security section recommends rate limits, requires boundary filters, confines the mechanism to a controlled domain and makes the feature disabled by default.

But a protective drop changes the evidence available to the sender. Ten congestion-point notifications may become three sender-facing CNPs. The standard CNP contains no original event identifier, proxy drop counter, decision reason, mapping version or congestion-point identity. The sender can react to the three packets; it cannot reconstruct the seven omitted packets from their absence.

Silence becomes ambiguous. It may mean there was no congestion. It may mean the congested node chose the wrong proxy, the PNC route was stale, a QP mapping was absent, a border filter acted, the feature remained disabled, a first packet was rate-limited, or the second packet was lost. The wire format deliberately does not settle these possibilities.

A standard CNP remains useful without being a receipt

This is not an argument that the translated CNP is false. It can carry exactly the standard instruction a legacy sender needs: reduce the sending rate for the reconstructed Source QP. Compatibility is valuable precisely because the sender need not implement the new first-leg format.

The error is to promote that compatibility packet into a provenance receipt. UDP syntax and checksum do not authenticate the original congestion event. The standard CNP does not expose the first message bytes or when the proxy received them. It does not count all events, localize the bottleneck, or prove that the proxy's table selected the intended flow.

Operators need a joined evidence chain. At the congestion point, retain queue and ECN counters, first-message count and selected proxy. In routing, retain the exact PNC route used at the event time. At the proxy, retain mapping version, learning time, accepted and dropped counts, translation reason and emitted CNP count. At the sender, retain received CNPs, QP response and rate change. Then check independent queue recovery, completion time and application effect.

No one record supplies all of that. The useful control is the correlation among them.

Revision 05 narrows the promise

The new revision removes broader language about arbitrary non-RoCE senders and makes the standard RoCEv2 construction path clearer. It says explicitly that learning the Source-QP/Destination-QP relationship requires the proxy to sit on both forward and reverse RoCEv2 traffic. It also notes that some implementation uses an ECN-marked data packet as input to a proxy, while leaving that method outside this document.

That is an important maturity boundary. The note is not a defined third format or an interoperability report. The document remains an individual Internet-Draft with no stream, responsible AD or formal IETF standing. Its requested ports, routing bits and BGP characteristic code remain proposals. Nothing in the frozen packet demonstrates deployment, performance or a named conforming implementation.

The right leadership reading is therefore specific. The draft presents a plausible compatibility mechanism. It also places sender selection, flow identity and notification completeness inside a proxy whose decisions are invisible in the resulting standard packet. If that packet will drive automated rate changes or incident attribution, the translation must leave its own observable receipt.

Sources