Summary

  • RFC 3234 made middlebox failure a distinct architectural problem: a replacement IP route could bypass a failed router, but it could not recreate state held by a failed intermediary.
  • The memo separated three recovery designs—soft state, stateful failover to a hot spare, and endpoint restart—and counted 16 of its 22 illustrative middlebox classes as hard-state and 21 as requiring session restart.

A path can come back while a session stays broken

The Internet’s simplest recovery picture starts with packets and routes. A router disappears; routing finds another path. The endpoints are still there, and the network again has a way to carry their packets.

RFC 3234, an Informational memo by Brian Carpenter and Scott Brim, drew attention to an intermediary that changes this picture. It defined a “middlebox” broadly: an on-path function beyond ordinary IP routing, whether a separate appliance or a virtual function inside another device. Such a box might translate addresses, enforce a firewall policy, proxy a connection, or perform another service. It was not the session’s final endpoint, but it could still be necessary for the session to work.

That distinction changes the failure unit. If an intermediary terminates one flow and originates another, the end-to-end packet path is no longer a single transparent chain. If it keeps a mapping or other session state, a routing protocol can restore packet reachability without restoring that mapping. The route is repaired; the existing conversation may still be unusable.

RFC 3234 offered a vocabulary for the difference. With soft state, the session can continue when the box disappears, perhaps less efficiently, while required information is reconstructed. A cache is the intuitive case: losing it should cost performance, not the application session. With hard state, losing the box’s state makes the function fail. A stateful hot spare can support rapid failover if it already has a copy. Alternatively, the two endpoints can detect the failure and restart through a spare, provided they retained enough information to do so.

These are not synonyms for “backup.” They place custody in different places. Soft state makes the intermediary’s data reconstructible. Failover puts a usable copy at the replacement before the incident. Restart asks the endpoints to recognize that their old session is gone and establish a new one. Each has a different interruption boundary and coordination burden.

A catalogue, not a census

The memo’s rough table classified 22 kinds of middlebox. Its authors marked 16 as hard-state and 21 as needing session restart after failure. Those figures make the architectural concern legible, but they are counts within an intentionally subjective catalogue—not a field survey, an estimate of Internet-wide prevalence, or a record of actual outage behavior. RFC 3234 explicitly declined to call its classifications complete or definitive.

The authors’ recommendation was narrower and stronger: middlebox design should include a clear failure mechanism. They also warned that failure coordination across protocol layers could not be assumed. A lower-layer device might have no way to know that an application-layer intermediary had failed, let alone move a session to the correct standby. Coordination has to be designed, not wished into existence by a diagram that draws the boxes on one path.

The later operator-perspective RFC 8517 describes a different slice of the environment: transport-aware functions and the flow visibility operators may use to troubleshoot application complaints. It confirms why responsibility and diagnosis can be shared across network providers and service platforms; it is not evidence that RFC 3234’s 2002 counts predicted current deployments. RFC 1958 supplies the architectural backdrop, and RFC 1812 describes the ordinary IPv4-router role against which the middlebox boundary was drawn.

The lasting lesson is not that intermediaries should disappear. RFC 3234 explicitly rejected a simple “good” or “evil” classification. It is that reachability and session continuity are separate claims. A path probe can show packets moving again; it cannot by itself show that the failed function’s state was recovered, replicated, or safely replaced by a new session.

Sources