Summary

  • RFC 8326 Graceful Shutdown deliberately lowers path preference before an EBGP session is removed; RFC 4724 Graceful Restart retains forwarding and stale routing state while BGP returns.
  • A defensible maintenance record must name the mechanism, state whether forwarding survives, prove the peer policy and convergence sequence, and define when to abort.

One word, two traffic instructions

The difference begins with the fate of the forwarding plane. If maintenance will take a port, line card or router out of service, traffic has to leave that path before the EBGP session disappears. RFC 8326 addresses that case. It standardizes the GRACEFUL_SHUTDOWN community and a sequence intended to reduce packet loss during deliberate session closure.

The sequence does not begin with a withdrawal. The initiator tags routes with the community; receiving policy lowers their LOCAL_PREF; both sides allow alternative paths to be selected and propagated; only after convergence does the operator close the session. The old path remains available temporarily, but at a preference below usable alternatives. RFC 8326 recommends a local preference of zero.

Graceful Restart starts from a different premise. RFC 4724 is for a BGP restart in which forwarding state can remain usable while control state is rebuilt. Peers negotiate a capability, retain relevant routes as stale, bound that retention with timers, re-establish the session and use End-of-RIB markers to complete replacement. The intended action is continuity, not drainage.

That makes the two mechanisms complements rather than substitutes. Shutdown says: prefer another path because this forwarding path is leaving. Restart says: keep using preserved forwarding while the control plane returns. A ticket that says only “enable graceful handling” has not made the core decision.

The community is an instruction, not an attestation

RFC 8326 assigns the well-known value 65535:0 to GRACEFUL_SHUTDOWN. RFC 1997 defines BGP communities as an optional transitive attribute and allows a receiving speaker to apply local policy. The signal can cross an interconnection, but its effect depends on the receiver having policy that matches it and lowers preference.

This division of control matters. The initiating network can announce its intent. It cannot force the neighbor to honor that intent, prove that the neighbor selected an alternate path, or create capacity on that path. Nor does the community prove that the maintained link will actually fail forwarding. It is a policy carrier.

LOCAL_PREF makes the boundary clearer. RFC 4271 defines it as an internal preference, with higher values preferred, and normally forbids sending it to external peers. The neighbor therefore translates an external community into its own internal route-selection decision. A successful drain is a bilateral operational outcome built from two independently controlled policies.

Restart depends on forwarding truth

Graceful Restart also contains a cross-boundary claim: forwarding entries remain viable while BGP restarts. RFC 4724 lets a restarting speaker indicate preserved forwarding state and tells a receiving speaker when to retain routes as stale. But retention is bounded. If the session does not return within the advertised restart time, or if the re-established session does not confirm the relevant forwarding state, the stale routes must be removed.

The specification acknowledges the central risk. If forwarding was not actually preserved, or topology changes before convergence completes, retained routes can contribute to transient loops or blackholes. A receiver that can determine that forwarding is no longer viable may delete stale routes early.

The maintenance controller therefore needs evidence beyond a BGP session alarm. It needs to know whether the data plane survived, which address families were covered, what restart and stale timers apply, whether the session returned, and whether End-of-RIB completed. “The peer came back” is not proof that every retained path was safe throughout the interval.

A change contract that can be tested

A useful maintenance record can be short, but it cannot be vague. It should state:

  1. Mode: Graceful Shutdown or Graceful Restart.
  2. Forwarding premise: the path will stop forwarding, or forwarding will remain intact while BGP restarts.
  3. Peer behavior: which community policy or restart capability has been verified on each side.
  4. Evidence: selected alternatives, FIB reachability, route re-advertisement, stale state, timers and End-of-RIB.
  5. Abort condition: no alternate path, insufficient capacity, failed forwarding check, unexpected community handling or expired restart interval.
  6. Rollback owner: the person or automation allowed to restore preference, clear stale state or cancel teardown.

This contract changes the economics of maintenance. The operator pays for preconfiguration, telemetry and a longer drain interval. The peer pays for policy and observation. Those costs are small compared with transferring an undefined failure to customers, who otherwise discover the difference through packet loss.

Sources and evidence boundary

The analysis uses RFC 8326, RFC 4724, RFC 1997 and RFC 4271. These documents establish protocol requirements and stated operational procedures. They do not establish deployment rates, vendor parity, spare-path capacity or correct operation in any named network.