Summary
- RFC 9494 extends stale-route retention per address family, but an advertised Long-Lived Stale Time is a proposed budget, not evidence that the route still delivers traffic.
- A defensible LLGR deployment binds capability, timer, route preference and propagation to next-hop, forwarding and packet evidence, then removes unrefreshed state at EoR, expiry or an operator-controlled stop.
Consider an illustrative failure. A provider edge loses its BGP control session to a peer, yet one more-specific route learned from that peer remains installed. Ordinary Graceful Restart runs out. Long-Lived Graceful Restart keeps the route for another six hours, marks it stale and makes it least preferred. A live less-specific alternative exists, but destination matching still chooses the retained more-specific. The route remains visible while the service behind it has already gone dark.
Nothing in this scenario requires a malformed message or a broken implementation. It follows directly from the trade RFC 9494 makes available. The extension can preserve useful state when a control plane recovers slowly or when BGP carries information closer to configuration than hop-by-hop reachability. The same extension can preserve a black hole, an inconsistent choice or an obsolete service binding for much longer than ordinary Graceful Restart.
The leadership question is therefore not whether a router “supports LLGR.” It is who has authority to keep stale state alive, how that authority is bounded, and what evidence can defeat the timer before the timer defeats the network.
Two consecutive custody periods
RFC 4724 defines BGP Graceful Restart. A speaker can advertise capability 64, a Restart Time and per-address-family forwarding-state information. A receiving speaker can retain routes after the session fails while the restarting speaker rebuilds control-plane state. End-of-RIB markers identify completion of an initial update for each address family.
RFC 8538 extends the reset conditions that can invoke those procedures. Peers that exchange the Graceful Notification bit can retain one another’s routes after many NOTIFICATION resets or Hold Time expiry. A Cease with the Hard Reset subcode requests full termination instead. The cause and subcode of a reset therefore belong in the evidence record; “the session went down” is not a sufficiently precise explanation.
RFC 9494 adds capability 71. Its value contains tuples of address-family identifier, subsequent address-family identifier, flags and a 24-bit Long-Lived Stale Time. LLST is measured in seconds and has no standards-defined default. A sender can propose a different retention time for each AFI/SAFI, while local configuration may impose an upper bound, a lower bound or both.
LLGR does not replace Graceful Restart’s machinery. If a speaker advertises LLGR without the GR capability, the receiver must ignore LLGR. The two periods can run serially: ordinary GR first, then LLGR. The first period retains routes without changing their preference under GR. The second retains them as least preferred LLGR state.
Before the session is re-established, the maximum standards-level retention interval is the received Restart Time plus the received LLST, subject to local limits. Either period may be zero. This matters because the visible timer is not one universal promise. It is the product of a peer’s advertisement, a receiver’s acceptance, a specific address family and local operating policy.
A timer changes the consequence, not the truth
When the LLGR period begins, the helper starts the per-family stale timer and attaches the well-known LLGR_STALE community. IANA records that community as 0xFFFF0006, or 65535:6. The route must be treated as less preferred than every route that is not also least preferred. Normal tie-breaking applies only when the surviving candidates are all least preferred.
Depreference is important, but it is not a withdrawal. If the stale route is the only route for an exact prefix, it can remain best. If a fresh route covers only a less-specific prefix, longest-prefix forwarding can still choose the installed stale route. RFC 9494 warns expressly that conventional reachability may be lost in this condition.
The distinction is sharper in hop-by-hop routed networks. Different iBGP speakers can have different session state and therefore apply different preference to the same reachability. One may choose a fresh route while another retains the stale path. RFC 9494 shows how that disagreement can create a forwarding loop and does not recommend LLGR for such route sets. Tunnelling can constrain that particular risk, but it does not make stale service state true.
Base BGP resolvability still applies. A stale route commonly remains usable only while its recursive next hop remains resolvable. The standard notes that BFD may provide additional viability evidence in appropriate cases. Neither IGP resolution nor BFD alone proves end-to-end delivery: both are bounded observations about the path toward the next hop.
The correct operational statement is narrower. LLGR authorizes a helper to postpone withdrawal under defined conditions. It does not authenticate the route, validate its origin, certify preserved forwarding, or establish that an application is reachable.
Communities are instructions inside a relationship
RFC 9494 also defines NO_LLGR, registered by IANA as 0xFFFF0007, or 65535:7. A route carrying it must not enter long-lived retention. The sender can mark reachability that should disappear normally, and a receiver may apply local policy to add the exclusion.
That is a useful refusal path, not a universal enforcement mechanism. BGP communities are path attributes described by RFC 1997; their presence does not authenticate who attached them. An intermediary can have policy that changes community handling. Evidence must show the exact received and advertised attributes at each boundary that matters.
An LLGR-stale route should not be advertised to a neighbor that did not advertise LLGR. The restriction keeps stale state inside a perimeter whose speakers understand the least-preference rule. It also means capability visibility and export evidence are inseparable: retaining a route locally does not authorize announcing it everywhere.
RFC 9494 permits a narrow partial-deployment exception. Stale routes may be sent to an internal BGP or confederation neighbor without LLGR support only when NO_EXPORT is attached and LOCAL_PREF is set to zero. The same zero preference should be applied consistently throughout the AS. The purpose is not cosmetic signalling; inconsistent selection can create loops.
The F bit and the return of the peer
LLGR’s per-family F bit says whether state for that AFI/SAFI was preserved during the previous restart. It is not a general health flag. If a re-established session omits the family, clears its LLGR F bit, or omits the required LLGR and GR capabilities, the helper must immediately remove the retained stale routes for that family.
The timer does not restart freely when control connectivity flickers. Without manual operator intervention, a running LLST is not updated until the peer has established and synchronized a new session. Synchronization is per AFI/SAFI. It occurs when the helper receives End-of-RIB for that family or when the GR Selection_Deferral_Timer expires.
The LLST timer continues after session re-establishment until EoR. If it expires during synchronization, stale routes that the peer has not refreshed are removed. A newly green BGP session can therefore coexist with a shrinking stale set. The evidence question is not “did the peer return?” but “which exact routes were refreshed before the family’s synchronization boundary?”
EoR closes an update epoch; it does not certify delivery. A complete initial update can still describe a path whose next hop fails resolution, whose forwarding entry was not programmed, whose label state is obsolete or whose application is unavailable. Control-plane completion and service recovery remain separate facts.
Where extended retention is rational
RFC 9494’s motivation includes NLRIs that behave more like configuration than conventional reachability. Route Target constraints, FlowSpec state, discovery information and VPN control state can take a long time to rebuild. Immediate removal may cause expensive churn even when the retained state remains useful.
The standard nevertheless requires affirmative configuration per AFI/SAFI and forbids enabling the procedures by default. That requirement is a governance clue. LLGR is not a generic availability improvement. It is a deliberate exception whose safety depends on what the route means, how forwarding is built and how long dependent state remains valid.
The security boundary can extend beyond packets. In an MPLS VPN, an old stale route may continue to reference a label. If the egress reassigns that label to a different VPN before every ingress stops using the old route, traffic isolation can fail. RFC 9494 says the lower bound on label reuse should exceed the upper bound on LLST. A timer that preserves continuity in one ledger can outlive the safety epoch of another.
Current implementation documents reinforce the need for platform evidence. Cisco describes BGP persistence as entering after GR or immediately when GR is skipped, and exposes separate sent and accepted stale times. Juniper exposes receiver and restarter controls, policy exclusions, manual stale-route clearing and partial-deployment choices. Nokia documents supported families, displayed capabilities and operational timers. Their syntax, defaults and supported scopes differ; none may be treated as a protocol-wide fact.
An evidence ledger for one stale epoch
The first record is negotiation. Preserve both OPEN messages or authoritative neighbor output showing GR and LLGR capability, AFI/SAFI tuples, Restart Time, LLST, N bit and each relevant F bit. Store the configured accepted maximum beside the peer’s advertisement. A large proposed timer and a smaller local cap are two different facts.
The second record is entry. Preserve the reset cause, time, subcode and whether Hard Reset semantics applied. Mark the end of ordinary GR and the start of LLGR for each family. Record which routes received LLGR_STALE, which carried or acquired NO_LLGR, and which were removed.
The third record is selection and propagation. Compare received route state, Loc-RIB decisions and every material advertisement. Confirm that non-LLGR neighbors did not receive stale state, or document the exact internal partial-deployment safeguards. A displayed community without the route-selection consequence is incomplete evidence.
The fourth record is forwarding. Preserve recursive resolution, the installed FIB or service state, and packet or application probes during the stale interval. For VPN state, include the label or encapsulation binding and its reuse constraints. A selected route without an installed and working outcome is only control-plane intention.
The fifth record is exit. Capture EoR, LLST expiry or an operator clear, then identify every unrefreshed route removed. Prove clean re-advertisement where service returned, and prove withdrawal where it did not. Name the operator who can shorten the timer, clear stale state and reverse an unsafe rollout.
Sources
- RFC 9494 — Long-Lived Graceful Restart for BGP
- RFC 4724 — Graceful Restart Mechanism for BGP-4
- RFC 8538 — Notification Message Support for BGP Graceful Restart
- RFC 5492 — Capabilities Advertisement with BGP-4
- RFC 1997 — BGP Communities Attribute
- RFC 4271 — A Border Gateway Protocol 4
- IANA BGP Capability Codes
- IANA BGP Well-Known Communities
- Cisco 8000 BGP persistence guide
- Juniper — Understanding Graceful Restart for BGP
- Nokia — BGP Graceful Restart and Long-Lived Graceful Restart
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
