Summary

  • RFC 9599 says a lower layer should create an explicit congestion mark only when the complete feedback loop can carry it through the egress and back to the load regulator. ECN capability belongs to the loop, not to one convenient bit.
  • Encapsulation must preserve accumulated congestion; decapsulation must combine inner and outer states. If a severely marked outer PDU contains a Not-ECN inner transport, dropping the packet is the only congestion signal that transport can understand.
  • Feed-forward-and-up, feed-up-and-forward and feed-backward modes solve different problems. A standard name, a marked frame or a successful packet delivery does not prove that the receiver reported congestion, the sender reduced load or service improved.

Picture one packet at a tunnel egress. Its outer header is marked Congestion Experienced. The inner IP header is still ECN-capable but unmarked. The decapsulator removes the outer header, forwards the inner packet unchanged and records a successful delivery. Every ordinary availability counter looks good: the packet survived, the tunnel did not drop it, and the destination can read it.

The control loop tells a different story. The only congestion evidence was discarded with the outer header. The receiver cannot report what it never sees. The sender cannot respond to feedback that never returns. More packets enter the same path, and congestion grows behind a dashboard that celebrates low loss.

RFC 9599 is guidance for preventing this exact category of false success. It asks designers of lower-layer and tunnelling protocols to propagate congestion consistently into IP so that IP can act as the portability layer between a non-IP-aware congested node and an end-to-end transport. Its object is neither a better queue algorithm nor a prescribed sender response. Its object is custody of the signal between those systems.

ECN capability belongs to a feedback loop

The RFC Editor record and IETF Datatracker identify RFC 9599 as an August 2024 Best Current Practice in BCP 89. It updates the subnetwork-design advice in RFC 3819. The base IP semantics come from RFC 3168, with experimentation boundaries later relaxed by RFC 8311.

The document's most useful abstraction is not a header layout. It is the distinction between an ECN-PDU and a Not-ECN-PDU. An ECN-PDU participates in a feedback loop in which every node required to carry a congestion indication back to the load regulator can do so. A Not-ECN-PDU belongs to a loop with at least one incapable node.

That makes ECN capability a property of a bound path of functions, not necessarily of the PDU in isolation. One protocol may self-describe capability in a field. Another may use a label, flow state or out-of-band control plane. The operational question remains the same: if an interior queue marks this PDU, will the signal survive every boundary and reach a function capable of controlling load?

A lower layer therefore should not mark traffic destined for a legacy L4 implementation that cannot understand ECN. Nor should it mark when a possible egress cannot propagate the indication upward. It should distinguish feedback loops that can carry marks from those that cannot. Different interior nodes may mix marking and dropping, but once any node can mark, every relevant egress must preserve or truthfully convert that mark.

This is a stricter test than “ECN enabled” on a switch. A switch can set a bit while the system remains Not-ECN. A decapsulator can understand a bit while the receiver does not. A receiver can report it while the sender ignores it. The label is earned by the loop.

Feed forward, then carry the evidence up

RFC 9599 names several operating modes. In feed-forward-and-up mode, a congested lower-layer node marks a lower-layer PDU; the mark travels forward to the subnet egress; and the egress transfers the indication into the higher-layer header before removing the outer one. The receiver's L4 protocol can then feed the evidence back to the source.

This is the established pattern for IP tunnels in RFC 6040 and for MPLS in RFC 5129. Related specifications apply it to TRILL in RFC 9600 and to IP headers separated by shim headers in RFC 9601. The encodings need not match across layers. What must persist is the congestion meaning.

At ingress, the new outer header must not erase congestion already accumulated upstream. Reinitializing the outer marking to zero makes the tunnel appear to begin a new reality. A monitor can no longer tell whether the outer observation represents the whole path or only the current subnet without private knowledge of the ingress behavior.

Preserving the baseline supports two different measurements. The outer header can represent congestion since the load regulator. The difference between outer and inner marking levels can estimate what the current subnet added. Neither measurement names the congested device or proves a queue length, but both remain interpretable because the ingress did not silently reset history.

Decapsulation is a calculation, not a copy

At egress, copying one field into another is too simple. The outgoing state must be calculated from the inner and outer states under ordered rules. When both are ECN-capable and express different severities, the more severe indication should survive. The egress is not choosing a preferred narrative; it is preventing a weaker inner state from overwriting stronger evidence accumulated outside it.

The hardest case reveals why drop remains necessary. Suppose the outer PDU carries the most severe explicit mark, but the inner packet is Not-ECN. Forwarding a marked inner packet would send a signal to a transport that cannot understand it. Forwarding it unmarked would erase congestion. The egress must drop it. Loss is not a failure of the ECN design in this case. It is the only feedback the legacy transport is capable of receiving.

Conversely, if the outer layer cannot signal congestion but the inner header is ECN-capable, the inner state should pass unchanged. Absence of a capability in one layer is not permission to clear evidence carried by another.

Some lower layers encode multiple congestion levels rather than one CE state. Their egress may need to translate a quantized level into a frequency of IP markings. Reframing makes the accounting harder: one IP packet can span many cells or frames, while one aggregated frame can carry several IP packets. Packet and byte congestion measures are not interchangeable. A counter must state which unit crossed the boundary.

A useful layering violation still needs a stopping rule

Feed-up-and-forward mode applies when the lower-layer header has no congestion field but a device can find and mark an encapsulated IP header. A so-called Layer 3 Ethernet switch may forward by Ethernet addresses while inspecting the IP ECN field inside the payload. A radio-access node can face a similar design.

This looks like a layering violation, but RFC 9599 permits it as a bounded optimization. It is not a universal substitute for native lower-layer notification. The payload may be encrypted. The next protocol may not be IP. Hardware may not parse deeply. Multiple nested headers may make the search expensive or ambiguous.

The search therefore needs a depth limit. If a recognizable IP header is not found soon enough, the device should use a built-in lower-layer mechanism or signal congestion through drop. It should not keep tunnelling through unknown headers until it finds two bits that resemble ECN. The first IP header found is sufficient; a later decapsulator can carry the mark into deeper headers.

The AQM recommendations in RFC 7567 and the benefits described by RFC 8087 explain why explicit notification is worth preserving: it can signal growing queues without waiting for loss. They do not authorize a parser to guess through encryption or unknown encapsulations. Accurate signalling includes knowing when explicit marking is unavailable.

Incremental deployment is an edge-coverage problem

MPLS demonstrates one safe, operator-controlled migration pattern. Interior label-switching nodes need not spend scarce header space indicating per-packet ECN capability. Instead, the operator upgrades every decapsulating edge in the domain once any interior node begins marking. An egress that finds a CE-marked outer PDU around a Not-ECN inner packet drops it on behalf of the earlier congested node.

This works because a professional operator can control the whole decapsulation set. It is not automatically safe in a plug-and-play environment. If devices can be assembled without a common configuration authority, an old egress may silently discard a new mark. A consumer link technology needs a fail-safe mechanism that prevents marks from entering such black holes under inconsistent deployment.

The unit of migration is therefore not “number of upgraded switches.” It is coverage of every node that may remove the marked header. One interior marker plus one legacy egress is enough to make an apparently modern domain lie. Inventory should be computed from possible decapsulation paths, failover paths and service-chain branches, not from an average firmware percentage.

Backward-only control moves the queue instead of closing the loop

Feed-backward mode sends lower-layer control toward the subnet ingress. It can regulate a stand-alone subnet efficiently when that ingress is also the original source. It couples poorly to an IP internetwork when the real sources sit farther away.

The subnet ingress can reduce its forwarding rate while the original IP sources continue at the old rate. The mismatch backs up an IP buffer at the ingress until that buffer experiences congestion and generates a separate forward signal. The lower layer has not delivered its original evidence end to end; it has relocated the pressure until another layer detects it.

RFC 9599 therefore says a technology intended to interface with IP should not be designed only around feed-backward mode, except for special cases. The history of ICMP Source Quench, now deprecated by RFC 6633, also shows the integrity problem: a source may be unable to distinguish a genuine path signal from a spoofed backward message.

A mutable mark must remain authentic as a permitted mutation

Adding a congestion field to a lower-layer header creates a security contract. Interior nodes are expected to modify it. If an authentication profile treats the entire header as immutable, every legitimate mark will invalidate the authenticator. The field must be declared mutable in transit, and the authenticated structure must reflect that fact.

Mutability does not make the signal truthful by itself. A node can suppress or fabricate a mark. RFC 9599 points to end-to-end approaches for detecting suppression and inadequate response rather than inventing a different hop-by-hop integrity scheme for every encapsulation. The design goal is not to make congestion evidence impossible to alter; it is to make the allowed alteration explicit and the feedback loop auditable.

Lu Heng's account of running-code primacy makes the required evidence concrete. “Supports RFC 9599” is not a receipt. A receipt joins the incoming inner state, outer state chosen by the encapsulator, interior mark, egress calculation, outgoing IP state, receiver feedback, load-regulator response and observed rate.

The minimum-initial-specification lens explains why one invariant can span heterogeneous technologies without dictating their internal design: never create an explicit signal unless the loop can carry it, and never erase congestion because a header is removed. The essay on reality layers supplies the final separation. Queue pressure, a lower-layer mark, an outgoing IP mark, receiver feedback, sender response and delivered service are related events, not one status light.

The tunnel that kept the packet did only half its job. Leadership should require proof that it kept the reason the sender needed to slow down.

Sources