Summary
draft-ietf-bier-source-protection-11warns that a backup BFIR may lose the path by which it watches the selected BFIR even while the selected BFIR still reaches all or some BFERs. A takeover based on that observation is unnecessary and can duplicate packets.- A reverse-direction Ping has the same epistemic problem: unless its path is co-routed with the forward multicast path, it can fail while the service works or succeed while the service path is broken.
- Observation identity, detector result, standby mode, switching authority, duplicate or loss behavior, receiver continuity and application outcome need separate receipts. Revision 11 is an Informational working-group Internet-Draft, not proof of deployment.
At 03:17, the backup ingress stopped hearing the primary. Its timer expired. The local state machine did exactly what it had been configured to do: it promoted itself and began forwarding the multicast flow.
The primary had not failed. Neither had the paths from the primary to the receivers. Only the path between the two ingress routers was broken. For a short but consequential interval, both ingress routers sent the same stream into a network that had not been prepared to filter the duplicate.
That is the most important scene in BIER Redundant Ingress Router Failover, revision 11. The draft does not merely catalogue BFD and Ping options. It exposes an authority mistake: a detector observed one path, while the failover action changed another.
The mistake is easy to conceal behind a green word such as “protection.” Redundant ingress sounds like a single mechanism. In practice it is a chain of claims. Which path was sampled? What absence crossed which timer? Who may replace the selected ingress for which receivers? Did a second stream appear? Did any receiver lose, reorder or duplicate packets? Did the application above multicast continue correctly? No one bit answers all of them.
The topology has more than one failure surface
RFC 8279 gives BIER a distinctive architecture: multicast packets enter at a Bit-Forwarding Ingress Router, carry a bit string naming egress routers, and cross a domain whose intermediate routers do not keep per-flow multicast state. A BFER still needs an upstream multicast hop for each flow. Overlay protocols such as MVPN, MLD or PIM distribute the information used for that selection; BIER transports the packets.
Revision 11 calls the chosen ingress the S-BFIR and an alternate the B-BFIR. A source may attach to both. Some deployments make one ingress primary for every receiver. Others divide receiver subsets or flows between the two. “The primary” can therefore be a local fact, not a universal role.
The draft separates node monitoring from working-path monitoring. That distinction matters because the topology contains at least three directed surfaces:
- the B-BFIR-to-S-BFIR path used by a backup to judge the selected ingress;
- each S-BFIR-to-BFER forward distribution path carrying the service;
- any BFER-to-S-BFIR reverse probe path used by a receiver to solicit a response.
Those surfaces can share links, but the protocol cannot assume they do. A failure on the first surface says that the backup cannot currently observe the primary through that route. It does not, by itself, say that the primary is dead or that the second surface is impaired.
The document supplies the concrete counterexample. In warm standby, BFIR2 monitors BFIR1. Their inter-ingress path breaks. BFIR2 interprets the event as BFIR1 failure and takes over. Yet BFIR1's paths to all or some BFERs remain intact. The draft names the switch unnecessary and says it causes packet duplication in the network and at BFERs.
That is unusually valuable standards language. It makes the false positive part of the mechanism rather than an embarrassing implementation footnote.
A liveness bit needs a path identity
A BFD session is not an abstract health oracle. RFC 5880 detects failure on a particular bidirectional forwarding path. RFC 8562 extends BFD to multipoint networks. The current BIER BFD draft supplies BIER-specific bootstrap and active-tail notification mechanics.
Those tools can be fast and precise within their defined scope. The operational failure begins when a controller drops the scope.
| Recorded state | Defensible statement | Unsafe expansion |
|---|---|---|
| B-BFIR/S-BFIR BFD Down | The monitored session crossed its failure threshold | Every S-BFIR-to-BFER path failed |
| BFER missed reverse Ping replies | The chosen reverse probe did not complete | The forward multicast stream stopped |
| Multipoint BFD tail timed out | This tail lost the expected control packets for the Detection Time | All receivers lost the flow |
| BFER selected a new UMH | Overlay selection changed for that receiver and flow | The new ingress delivered continuous service |
| Backup began forwarding | A second source entered the forwarding system | Duplicates were suppressed everywhere |
| Application remained quiet | No application alarm was observed | No packets were lost, duplicated or reordered |
The receipt therefore needs the directed endpoints, subdomain, entropy or path-selection inputs, packet treatment, session discriminator, detection interval and traffic class. “BFD Down” without those joins is a detached verdict. It can be cryptographically logged and operationally misleading at the same time.
Reverse Ping can be wrong in both directions
The Ping example is even clearer. A BFER sends a request back toward BFIR1. After several missing responses within an expected period, it may conclude that BFIR1 is a failed upstream hop and select BFIR2.
But a BIER multicast packet travels from BFIR1 to the BFER. The echo request travels in the opposite direction. In a complex topology, routing need not be symmetric. The draft explicitly identifies both errors:
- the reverse Ping fails while the forward multicast path and flow remain healthy;
- the reverse Ping succeeds while the forward multicast path has a defect.
The document requires co-routing of the request with the monitored forward path to improve consistency. That requirement is a powerful clue about evidence design. The monitor must not merely reach the same endpoint. It must represent the same forwarding treatment that gives the service meaning.
Even co-routing remains a network-level receipt. A response does not prove that a receiver accepted an ordered, nonduplicated application stream. It only makes the detector's claim better aligned with the path it is supposed to describe.
Cold, warm and hot standby allocate risk differently
The companion multicast redundant-ingress draft describes three modes. They are not simply three speeds.
In cold standby, a receiver selects one ingress and signals the alternate only after failure. This saves duplicate bandwidth and limits routine state, but signaling and convergence can lose packets. The receiver owns more of the switching decision.
In warm standby, selected and backup ingresses already know the receiver demand, but only the selected ingress forwards. The backup can start faster. It also receives a dangerous institutional privilege: it may decide that the primary has failed even when it cannot directly see the protected receiver path.
In hot standby, both ingresses forward. The receiver discards the stream from the nonselected source until it switches. Detection no longer needs to awaken a dormant sender, but duplicate traffic is a permanent network cost and selective acceptance becomes a receiver obligation.
This is not a ranking from primitive to mature. Each mode moves loss, bandwidth, state synchronization, detector authority and duplicate suppression to a different actor. A procurement specification that asks only for “fast failover” has not described the control system it is buying.
Different receivers can hold different truths
The draft permits BFERs to choose the same or different selected BFIRs. Ping timeout periods may also differ by receiver performance requirements. One BFER can therefore switch while another continues to accept the original ingress. A backup may promote itself for one receiver subset while the selected ingress remains legitimate for another.
This is why a global event label such as primary_failed is too coarse. The smallest defensible key is closer to:
(flow, receiver-or-receiver-set, selected BFIR, observation path, detector, threshold, epoch).
The switch receipt must then say who acted, under which standby mode, for which BFER set, and with which prior selection state. Without that scope, operators cannot distinguish a split but valid local decision from an uncontrolled dual-primary condition.
The same principle governs recovery. Seeing packets from BFIR2 after a switch proves that BFIR2 sent something. It does not prove that every BFER switched, that BFIR1 stopped, that duplicate suppression worked, or that an application received a continuous sequence.
The draft's own limits belong in the decision
Revision 11 is a BIER Working Group Internet-Draft dated 1 October 2026 with intended Informational status. It is not an RFC. Its IANA section requests no allocation. Its Security Considerations section points to the underlying BIER, multipoint BFD, MVPN failover, Ping and BFD documents rather than supplying a complete threat model for the combined control loop.
Those facts do not diminish the draft's insight. They constrain the claim. The text is strong evidence that its authors and working group recognize observation-path mismatch, false results and unnecessary takeover as design concerns. It is not evidence that a named vendor implemented the safeguards, that an operator deployed them, or that any multicast service survived a real failure.
Lu Heng's Reality Layers offers the right editorial discipline: keep the detector's symbolic state attached to the physical path it sampled. The Policy Mirror adds the institutional question: who was allowed to turn that state into a topology change, for whom and with what consequence?
The failover receipt
For every automatic switch, retain: flow identity; selected and backup BFIR identities; affected BFER set; standby mode; overlay protocol and selection state; detector type and session identifiers; directed observation path; evidence that probe treatment matches the protected flow; timer values and missed-packet count; precise transition time; actor that authorized the switch; old and new UMH per BFER; forwarding start and stop receipts from both ingresses; duplicate, loss and reordering counters; receiver acceptance state; and application continuity evidence.
Record negative knowledge explicitly. If co-routing was not proved, say path_equivalence_unverified. If a backup cannot observe a receiver path, say receiver_path_unknown. If only a controller state changed, do not fill in a forwarding result. A failover system earns authority by preserving these gaps, not by painting over them.
Sources
- BIER source-protection Datatracker record
- Document history
- Revision 11 text
- Revision 11 HTML
- Revision 11 XML
- Multicast Redundant Ingress Router Failover, revision 10
- BIER BFD, revision 12
- BIER Ping and Trace, revision 29
- RFC 8279: BIER architecture
- RFC 8562: Multipoint BFD
- RFC 5880: BFD
- RFC 9026: MVPN Fast Upstream Failover
- RFC 8556: MVPN Using BIER
- RFC 8029: MPLS LSP Ping
- Lu Heng: Minimum Initial Specification
- Lu Heng: Running-Code Primacy
- Lu Heng: The Policy Mirror
- Lu Heng: Reality Layers
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
