Summary

  • RFC 9860 turns a TI-LFA repair list into hop-by-hop PIM RPF instructions so a secondary multicast tree can exist where a conventional LFA cannot be found.
  • The computed list, installed standby state and local UMH switch do not by themselves prove continuous packets, reachable receivers or a healthy application service.
  • Operators need joined evidence across topology generation, repair eligibility, PIM state, failure detection, forwarding counters, packet sequence and downstream experience.

The useful shortcut begins with a mismatch

A TI-LFA repair path and a multicast tree are not the same object. TI-LFA computes an outgoing interface and an ordered segment list that avoids a protected resource. Unicast traffic can be steered with that list. PIM, however, must establish state at every intermediate router: an incoming interface, an upstream neighbour and outgoing interfaces toward receivers. Tunnelling one Join directly to the far end would skip the routers that need that state.

RFC 9860 closes this gap without inventing a new PIM protocol. It resolves the Node SID and Adjacency SID information from the link-state database into IP addresses. A node address is carried in the RPF Vector Attribute defined by RFC 5496; an adjacency peer can be carried in the Explicit RPF Vector from RFC 7891. Each router processes the leading vector and passes the Join onward. The result can be a standby multicast tree built hop by hop along a path that ordinary LFA rules could not supply.

That is a meaningful extension to MoFRR. The earlier mechanism joins a primary and a diverse secondary upstream at a merge point, receives both streams and normally discards the secondary copy. After a detected primary-path failure, the router can promote the secondary upstream locally instead of waiting for the multicast protocol to construct a new tree from nothing.

Five records that should not be collapsed

The first record is topology. It says which LSDB generation, metrics, SIDs and protected link or node produced a repair list. RFC 9860 applies inside one link-state IGP area; domain exits and attached inter-domain links are outside its protection claim. Incorrect or stale routing information can produce a perfectly reproducible but wrong alternate.

The second record is eligibility. A calculated path is usable for this method only if the nodes along the backup Join path support RPF Vector processing and the relevant configuration permits the vector Join. RFC 9860 notes that when vector and non-vector Joins conflict, the non-vector option can take priority, leaving protection inactive. “A repair list exists” is therefore not the same as “this multicast flow has an armed standby tree.”

The third record is construction. Operators need to know that the secondary Join traversed the intended neighbours and that each hop created the expected (S,G) state. RFC 9860 explicitly recommends checking state along a manually specified path and points to Mtrace as a debugging aid. Mtrace can illuminate multicast routing state and path behaviour; it does not certify an application receiver, and RFC 7431 warns that ordinary ping and mtrace do not verify the secondary stream that MoFRR is receiving and discarding.

The fourth record is activation. Failure detection has its own clock and authority. Link status, BFD, missing multicast packets and IGP withdrawal can drive different responses. The secondary tree may be ready yet never selected; it may be selected late; or it may be released when the point of local repair converges while other routers still hold a different topology. RFC 9855 treats the TI-LFA interval as local, from failure detection to local IGP convergence, and keeps other micro-loop conditions separate. RFC 9860 likewise leaves PIM micro-loop prevention out of scope.

The fifth record is outcome. Sequence gaps, duplicates, reordering and forwarding counters show what packets crossed the repair junction. Receiver-side joins and probes show which branches remained reachable. Player buffers, decoder continuity, telemetry ingestion or another application measure shows whether the service met its objective. A switch in the upstream multicast hop is evidence of a control action, not a substitute for those outcomes.

Fifty milliseconds is a test result, not a product label

RFC 7431 says 50 ms switchover is possible in suitable designs and zero packet loss is possible for an RTP-based first-packet selection approach. It also describes detection based on an expected packet rate, where delay necessarily depends on the stream. None of these possibilities authorises a universal “sub-50-ms multicast recovery” claim.

A meaningful measurement needs a start event and an end event. Did the timer begin when light disappeared, when BFD changed state, when the primary UMH was demoted, or when the first packet was missing? Did it end at the first repaired packet, at stable sequence continuity, at the last receiver's recovery, or when the application returned to policy? Different endpoints answer different questions. Publishing only one duration hides the evidence layer that supplied it.

Capacity also changes the verdict. MoFRR obtains speed by maintaining the secondary tree and receiving the duplicate stream up to the merge point. The backup consumes links and state before it is needed. A path that is topologically diverse but insufficiently provisioned may activate correctly and still degrade a high-rate flow. The failover ledger therefore needs the offered rate, primary and secondary interface counters, queue loss and receiver observations for the same test interval.

A compact proof chain

The smallest defensible operational record links: topology generation and protected resource; computed repair list; vector-capable path and accepted configuration; secondary Join propagation and per-hop state; detected failure and selected UMH; packets entering and leaving the merge point; packet continuity at representative receivers; and service-level recovery. Each join should carry a common incident or test identifier and monotonic timestamps.

This is the practical value of Heng Lu's distinction between symbolic declarations and running reality. Running-code primacy does not reject specifications; it limits each specification to what implementations can demonstrate. The related account of reality layers explains why a configured feature, a packet event and a delivered service cannot inherit one another's proof by proximity.

RFC 9860 deserves credit for keeping its authority narrow. It supplies a way to construct a wider class of standby PIM trees. The operator still owns the evidence that the tree was built, activated and useful.

Sources