Summary

  • RFC 3948 put IKE control messages and UDP-encapsulated ESP packets on the same NAT mapping, using four zero octets to mark IKE because a valid ESP SPI could not be zero.
  • The marker chose a parser, while SA lookup, anti-replay, cryptographic checks, inner-address policy and application delivery remained separate decisions; the one-octet NAT keepalive was not even a peer-liveness test.

One hole through the middlebox

By 2005, IPsec had a practical collision with the network around it. ESP protected payloads, but many deployed address translators recognized flows through transport ports. Native ESP did not provide the TCP or UDP shape on which those boxes commonly relied. Passing ESP through UDP gave the translator a familiar outer flow, yet it raised another question: IKE negotiation and ESP data now needed to reach the same endpoint through the same translated path.

RFC 3948 chose to share the port. Its reasons were concrete. One mapping scaled better than separate mappings, ordinary traffic on that mapping reduced the need for separate IKE keepalives, a firewall needed one port rule, and implementations had one external path to maintain. The companion RFC 3947 moved NAT-traversing IKE exchanges to UDP 4500 and supplied the negotiation around that choice.

Sharing did not merge the protocols. An arriving UDP payload still had to go either to IKE, which managed keys and associations, or to ESP, which carried protected traffic. A port number could no longer make that choice. RFC 3948 therefore turned the first 32 payload bits into a tiny classification namespace.

Zero was reserved so that zero could speak

An ESP header begins with a 32-bit Security Parameters Index. The SPI helps a receiver locate the Security Association under which the packet should be processed. RFC 2406, the ESP definition used by RFC 3948, reserves zero for local, implementation-specific use and excludes it from ordinary wire traffic. RFC 3948 made the operational rule explicit: an ESP SPI on this path must not be zero.

IKE traffic on port 4500 received the inverse shape. Four zero-valued octets—the Non-ESP Marker—were placed after UDP and before the IKE header. Those octets occupied exactly the position where an ESP SPI would have appeared. The receiver could inspect one word: zero led toward IKE processing; a nonzero word led toward ESP processing.

This was elegant because it reused an impossible value rather than adding another negotiated port or a larger envelope. It was also narrow. The marker did not say which peer sent the message. It did not identify an IKE Security Association, sign the header, authorize an exchange or prove that the bytes after it were well formed. It was a sign above two doors, not the guard behind either door.

The same limit applied on the ESP side. A nonzero first word could be treated as an SPI, but it was not automatically a valid one. The receiver still needed an installed association matching that SPI and destination context. It still had to process the sequence number, anti-replay state, cryptographic protection and traffic policy. Unknown SPI, replay failure, failed integrity or forbidden inner traffic could all end the packet after the first classification had worked perfectly.

A third message kept the mapping, not the conversation

The shared mapping also carried a remarkably small maintenance packet: one octet with value 0xFF. RFC 3948 called it a NAT keepalive. A sender could emit it after an idle interval to stop a translator from forgetting the UDP mapping. A receiver was advised to ignore it.

The specification drew an unusually useful negative boundary: receiving a NAT keepalive must not be used to detect whether the connection is live. The middlebox may refresh a timer because a datagram crossed it. That says nothing about an IKE state machine, an installed ESP association, successful decryption, permitted inner traffic or a functioning application. Even the peer application need not consume the octet.

The publication-era defaults—twenty seconds of outbound silence before a needed keepalive and up to five minutes after an association had existed—described configurable behavior, not timeless measurements. More importantly, the timer owners were different. The translator controlled mapping expiry. The endpoint controlled keepalive transmission. IKE and ESP controlled their own association lifetimes. No single tick proved the others.

Correct classification left harder policy behind

UDP encapsulation did not replace ESP. RFC 3948 inserted or removed the UDP header and then called for ordinary ESP processing. After decapsulation, the implementation still faced the address effects that NAT had created.

In tunnel mode, local policy could validate an inner source against an allowed range, compare it with an address assigned to the peer, or translate it for the local network. In transport mode, changed IP addresses could make the inner TCP or UDP checksum wrong; the receiver might adjust it using original-address information from IKE, recompute it, or take a narrowly bounded integrity-protected path described by the RFC. Passing the four-byte test decided none of these questions.

The security section made the remaining ambiguity visible. Two remote laptops behind different translators could both use the same private address. A gateway would then see several associations that appear to lead to the same inner destination. Several clients behind one public address could also propose overlapping traffic descriptions, leaving a simple outbound filter unable to choose the correct association. The outer address coordinated delivery only as far as the translator; it did not become a unique name for every host or policy behind it.

These were not theoretical reasons to reject NAT traversal. They were the price of making a minimal compatibility layer work with deployed boxes. The earlier RFC 3715 had catalogued the requirements and conflicts. RFC 3948 supplied a wire mechanism while preserving the responsibility of implementations to resolve collisions rather than pretending the wrapper had eliminated them.

The small rule survived because it stayed small

The RFC Editor record identifies RFC 3948 as a Standards Track document published in January 2005. Its accepted erratum corrects one cross-reference in the introduction; it does not alter the packet formats or the four-zero rule.

Later IKEv2 kept the arrangement. RFC 7296 says that IKE messages sent on UDP 4500 carry four leading zero octets, while an ESP header follows UDP directly and cannot validly begin with SPI zero. It also permits an initiator to use 4500 without first proving that a NAT exists. Port choice alone therefore cannot be backfilled into evidence of a translated path.

RFC 3948's durable lesson is not that four bytes secured IPsec. It is that a good coordination rule can do one modest job with precision. The marker made two parsers coexist. The SPI connected a candidate packet to receiver-side state. Cryptographic and replay processing tested that state. Local policy governed inner traffic. Application evidence established the outcome. Collapsing those receipts would make the system look simpler only by erasing where decisions actually occurred.

Sources