Summary

  • BGP Graceful Restart lets a capable peer retain routes during a session restart, but retained control-plane state is not independent proof that the restarting router can still forward the affected traffic.
  • The operational safeguard is an address-family restart receipt that binds capability bits, retention timers, observed packet continuity, End-of-RIB completion and the final disposal of stale routes.

Consider a hypothetical planned maintenance window. A router restarts its BGP process, the neighbouring helper keeps the learned prefixes, and the routing dashboard shows an orderly recovery rather than an immediate withdrawal. Traffic to one address family nevertheless stops. The helper continues selecting the stale route until a timer or later protocol event removes it. A mechanism intended to mask a short control-plane interruption has made the data-plane failure last as long as the retention decision.

That outcome is not evidence that Graceful Restart is defective. It is evidence that its two sides were treated as one. The protocol coordinates control-plane recovery. Safe use also depends on forwarding state remaining valid while the peer acts on stale routing information.

What the capability actually says

RFC 4724 allows a BGP speaker to advertise Graceful Restart capability. Its Restart State bit tells a peer that the speaker has restarted. For each address family, the Forwarding State bit reports whether forwarding state was preserved during that restart. Those fields matter because restart handling is not one undifferentiated promise across the router.

When forwarding state was not preserved for an address family, RFC 4724 requires the receiving speaker to remove the stale routes for that family. When preservation was indicated, the helper may retain them while the BGP session is re-established and routing information is refreshed. The distinction is an explicit safety boundary, not an optional annotation.

The Restart Time estimates how long the restarting speaker expects to take to re-establish the session. A helper also has a locally configured limit for retaining stale routes. Neither number measures forwarding continuity. They bound how long the control plane is willing to rely on the preservation claim.

End-of-RIB closes a different question

After the session returns, the End-of-RIB marker tells the helper that the initial routing update for an address family is complete. Routes refreshed by the new exchange remain current; stale routes that were not refreshed can be removed. This closes the routing reconciliation phase. It still does not reconstruct packet delivery during the interval before the marker arrived.

That distinction changes the incident record. Session re-establishment, End-of-RIB receipt and successful traffic are three observations. Collapsing them into a single “restart succeeded” flag makes it impossible to tell whether the mechanism protected service or merely kept a route installed.

Longer retention enlarges the proof burden

RFC 9494 extends the model through Long-Lived Graceful Restart. It defines longer retention and the LLGR_STALE community, and recommends de-preferencing those routes so alternatives can win. This can be useful when forwarding truly survives a longer control-plane interruption. It also makes a bad assumption more durable when preservation has ended or was never present.

The policy question is therefore not whether a long timer is always wrong. It is whether the selected timer, preference treatment and alternate-path behaviour are supported by the platform and topology being operated. An operator needs evidence for the relevant address family, not confidence borrowed from a different family, line card or laboratory condition.

Notification support is not forwarding proof

RFC 8538 allows Graceful Restart procedures to be used with certain BGP NOTIFICATION events. That improves the precision of restart handling, but it does not change the evidence boundary. A notification can explain why a session was reset and which procedure the peer should follow. It cannot attest that packets crossed the restarting system throughout the event.

For the same reason, the absence of route withdrawal is not a delivery result. It may be exactly what the helper was instructed to produce. The question that remains is whether the forwarding plane honoured the state on which that instruction depended.

Build a restart receipt by address family

For every planned or accepted use of Graceful Restart, record the two peers, AFI/SAFI, advertised capability, Restart State and Forwarding State bits, Restart Time, local stale-path limit, any LLGR time and community treatment, and the available alternate path. Add timestamps for session loss, first successful packet observation, session return, End-of-RIB and stale-route removal or refresh.

The packet observation should come from a path and service cohort that matters. A loopback ping may prove that one destination remains reachable while customer prefixes or a different forwarding table fail. Likewise, a successful IPv4 family does not clear IPv6, and one line card does not clear another. The receipt must match the state the helper is retaining.

Sources