Summary
- The current LISP multicast draft says that changing an unreachable ingress tunnel router requires receiver sites to build state toward a new encapsulating root and tell it which
(S-EID,G)flow they want. - Anycast can preserve the routed locator and reduce underlay recovery to an RPF change, but the receiver-side ETR may not know that a different physical ITR now owns the address. The replacement may wait for the next periodic Join/Prune before it learns the receiver-specific state.
- Daniel Kade proposes a root-succession receipt and an explicit join-replay debt. These are operational governance tools, not new LISP or PIM protocol fields.
A green address can hide a silent tree
Picture a live briefing distributed from one source site to receiver sites in several networks. The source-side LISP ingress tunnel router fails. Routing moves the anycast locator to a second ITR. Probes to the shared address recover, and the multicast underlay can perform a reverse-path change. The incident board therefore records a short routing interruption.
That record may be accurate and still miss the service failure. The new physical ITR did not necessarily inherit the old ITR’s knowledge that particular receiver ETRs had joined particular source-and-group flows. The address survived as a routing destination; the (S-EID,G) state was held by a process that disappeared. Until receiver sites replay their intent, the replacement root can be reachable without knowing which traffic to encapsulate for whom.
This is not an indictment of anycast. Anycast offers one address from more than one location and lets routing choose an instance. It does not promise application or protocol-state replication between those instances. The governance error is made by the observer who treats restoration of the advertised address as proof that every stateful service behind it continued.
draft-ietf-lisp-rfc6831bis-07 makes the separation unusually visible. The draft is active IETF work in IESG Evaluation, intended for Proposed Standard and set to obsolete Experimental RFC 6831 if approved. Revision 07 was posted on 11 September 2026. The checked change from revision 06 is mainly normative wording, updated IGMP and MLD references, and editorial maintenance. The failover issue is part of the document’s continuing architecture, not a newly invented property of revision 07.
Two namespaces, several pieces of state
LISP separates endpoint identifiers from routing locators. A source and a receiver reason about source EIDs and multicast group addresses inside their sites. The underlay routes among RLOCs. That separation is useful because an endpoint identity can remain stable while the attachment location changes, but multicast makes the state boundary more demanding than ordinary unicast forwarding.
A receiver joins a source-specific flow as (S-EID,G) through IGMPv3 or MLDv2. Its egress tunnel router looks up the source EID and selects a source-site ITR RLOC using the mapping system’s priority and weight. With a multicast underlay, the ETR sends two distinct signals. One is a unicast-encapsulated PIM Join/Prune carrying (S-EID,G) to the selected ITR. The other builds (S-RLOC,G) state across the underlay. The first tells the encapsulator which EID flow is wanted; the second gives the underlay a locator-rooted distribution tree.
With a unicast underlay, the ETR sends the encapsulated join information to the ITR, and the ITR tracks a list of receiver ETR locators. The source-side device then performs head-end replication, producing an individual outer packet for each receiver ETR. In either case, a receiving ETR removes the LISP header and consults its inner (S-EID,G) multicast forwarding state before sending packets toward local receivers.
These are related but non-substitutable records. A Map-Reply can identify candidate locators without proving a chosen physical ITR has installed a flow. An RPF-successful underlay tree can lead toward the active locator without proving the root knows the inner EID source. A local receiver join can remain present while the remote encapsulator has forgotten it. A packet can reach an ETR through a shared (RLOC,G) tree and still be discarded because there is no matching inner state.
The practical object is therefore not simply “the multicast route.” It is a joined chain: receiver intent, ETR boundary state, selected root, root-side EID state, underlay or head-end replication state, ETR acceptance and receiver observation. Each link has a different owner and may age on a different clock.
When the physical root changes
Section 6 of the draft explains why an ITR failure is more than an ordinary RPF event. In unicast, local reachability can steer packets toward another locator with little ceremony. In LISP multicast, the selected ITR is the encapsulating root of a distribution tree. When that root becomes unreachable, joined ETRs need to create state toward a new root. They must trigger (S-RLOC,G) Join/Prune state toward the new ITR and send a unicast-encapsulated Join/Prune telling the new ITR which (S-EID,G) is being joined.
That is a state-succession obligation. It has a population—the affected receiver ETRs—and an object—the source/group state each one expects at the replacement. It also has a time dimension. Different receiver sites can detect, select and replay at different moments. A single “failover complete” timestamp cannot honestly describe the tree unless it is derived from those constituent transitions.
Anycast changes the visibility of the transition. If several ITRs use the same RLOC, routing can move the address and the underlay may only need to alter its reverse-path interface. But the ETR does not necessarily see an address change. It may therefore lack the event that would cause an immediate unicast-encapsulated replay of (S-EID,G). The draft says the new anycast ITR may receive that state only when the ETR sends it again through periodic procedures.
The resulting interval is not specified here as a universal number. Implementations and PIM timers matter, as do loss, convergence, refresh policy and the underlay mode. It would be wrong to infer a fixed outage duration or even to promise that every failover loses traffic. The durable conclusion is narrower: address recovery and state recovery have distinct evidence, and the latter may depend on a later refresh whose completion is not shown by the former.
Join-replay debt
I call the unresolved obligation join-replay debt. At the moment the active physical root changes, every receiver ETR whose required (S-EID,G) state has not been demonstrated at the replacement is an open item. The debt is not a packet counter and it is not a moral judgment. It is a way to stop a many-receiver transition from disappearing behind one green locator.
Debt should be scoped to a protected flow or a bounded class of flows. A useful entry records the receiver ETR identity, source EID, group, old selected ITR, expected new ITR, last join generation, last refresh, underlay mode and the evidence required to clear it. If an implementation exposes a state-install acknowledgement, that can contribute. If it does not, operators may need bounded root-side state inspection joined with packet observation at the ETR or a controlled receiver canary.
The debt is not cleared by a successful ping to the anycast RLOC, a route in the RIB, a mapping-cache entry, a PIM adjacency or a healthy process alone. None of those observations proves that this flow’s receiver intent crossed into the new physical root. Nor is it cleared merely because one receiver resumes; another ETR can be waiting on a later refresh.
Debt can also be retired without pretending that every desired receiver returned. A receiver may have deliberately left, a source may have stopped, or policy may have moved the flow elsewhere. The closure record should name the reason and authority: installed and observed, explicitly withdrawn, expired under a declared rule, or accepted as degraded by a named operator. Silent disappearance is not closure.
A root-succession receipt
The companion artifact is a root-succession receipt. This is Daniel Kade’s governance proposal, not a wire-format change to LISP, PIM, IGMP or MLD. It joins evidence already available at different operational surfaces so that one team cannot close an incident using only the layer it controls.
The receipt starts with identity. It records the old and new physical ITR instances, their RLOC or shared anycast RLOC, the detection source, the time and confidence of the transition, the mapping version used by receiver ETRs, and the operator who owns the decision. A shared address without instance identity is insufficient precisely because the article’s problem is hidden physical succession.
It then records control-plane progress per affected scope: the old root’s last known (S-EID,G) state, the new locator selection, underlay RPF convergence, the ETR’s next join generation, any encapsulated replay, and the replacement’s installed state when observable. Evidence should retain clock source and collection point. A timestamp from an ETR and one from an ITR cannot be ordered confidently if their clocks are not comparable.
Finally, it records data-plane consequence: first packet accepted at the replacement, first packet accepted by each sampled receiver ETR, sequence or application continuity where available, duplicate and loss window, unwanted traffic discarded because inner state did not match, and the time the remaining replay debt reached the declared threshold. The receipt must not call an application healthy when it has only router-level evidence.
Rollback belongs in the same record. The owner may restore a specific unicast locator, drain the replacement, force or accelerate controlled join refresh, move a flow to another ITR, or accept temporary head-end replication. The available action depends on the implementation and network. The governance requirement is that the action, scope and authority are explicit, not that this article invents a universal command.
Replication location moves the evidence
The draft permits replication at several places: inside a site on EID state, in the multicast underlay on RLOC state, at receiver ETRs and at source ITRs. That flexibility changes who can see the failure. Underlay replication concentrates tree state in transit routers. Unicast underlay mode trades that state for source-side copies and bandwidth. Multiple ITRs can serve different receiver sites according to Map-Reply priority and weight.
The accounting must follow the chosen topology. In head-end replication, an ITR list may show intended ETR destinations, but it does not show downstream delivery. In underlay replication, a healthy (S-RLOC,G) tree may combine traffic for different inner sources. The draft notes that a receiver site can receive traffic for an (S-EID,G) it did not join when flows share an underlay tree; the ETR should discard it on the inner lookup. That is correct forwarding behaviour, yet the unwanted packet already consumed capacity on the common path.
This matters for security as well as resilience. The draft describes a malicious receiver site causing legitimate sites on a shared underlay tree to receive unrequested traffic. A root-succession receipt should therefore distinguish missing expected traffic from present unwanted traffic. “Packets reached the ETR” can describe success, waste or attack depending on the inner state and intended receiver set.
The diagnostic boundary is still open
Operational proof is complicated by explicit scope limits. Detailed locator reachability, multicast Proxy-ITR behaviour and mtrace design are outside the draft’s scope. The mtrace section says design for LISP multicast remains to be determined, building on Mtrace Version 2. That is an honest gap, but it prevents operators from treating a generic path diagnostic as a complete view across the EID and RLOC namespaces.
Monitoring should therefore label the layer it observes. An underlay test can show the active RLOC path. Root-side inspection can show installed (S-EID,G) state. ETR counters can show decapsulation and inner lookup. Receiver instrumentation can show actual arrival and application acceptance. Correlating them may support a service claim; substituting one for all the others cannot.
The IETF draft is valuable because it specifies a mechanism and exposes its edges. It does not owe operators a universal observability product. Operators, vendors and reviewers do owe one another precision about what their evidence covers. The phrase “anycast failover succeeded” should be reserved for the address transition. A multicast service recovery claim needs the state and receiver chain too.
Sources
- Current LISP multicast draft
- Draft history
- Datatracker API record
- Revision 07 HTML
- Revision 07 text
- Revision 06 to 07 diff
- LISP Working Group
- RFC 6831: Experimental LISP Multicast
- RFC 9300: LISP data plane
- RFC 9301: LISP control plane
- RFC 8059: PIM Join Attributes for LISP
- RFC 7761: Protocol Independent Multicast
- RFC 8487: Mtrace Version 2
- RFC 9776: IGMPv3
- RFC 9777: MLDv2
- RFC 7799: Internet measurement terminology
- Revision 07 XML source
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
