Summary
- RFC 10018 can advertise an SR P2MP P-tunnel as
<Root, Tree-ID>and derive its leaf set from MVPN or EVPN Auto-Discovery routes. That proves bounded control-plane state, not that the controller installed one complete PTI or that every intended leaf received the right service payload. - A P2MP service should close only when the intended, advertised, controller-accepted, installed, forwarding, OAM-responsive and payload-confirmed leaf sets reconcile for the same candidate path and PTI Instance-ID; withdrawals need an equally explicit terminal record.
The healthiest dashboard can hide the simplest failure
Imagine a broadcast service with twelve egress PEs. Eleven receivers show clean traffic. Aggregate throughput is close to normal. The ingress PE advertises the expected PMSI Tunnel Attribute. The controller shows the tree as active. Yet the twelfth site receives nothing.
That is not an exotic contradiction. It is the natural failure shape of a point-to-multipoint system. A single root, several intermediate replication points and many leaves create a service in which partial success is numerically dominant. A green average can be accurate and still be operationally misleading.
RFC 10018 gives operators an important common contract for this environment. MVPN and EVPN can advertise P-tunnels realised by Segment Routing point-to-multipoint tree instances. The same architecture can use SR-MPLS or SRv6; the document also specifies a separate ingress-replication option. BGP Auto-Discovery routes connect service membership to an SR P2MP Policy's leaf set. New endpoint behaviours and tunnel types give independent implementations shared wire meanings.
The standard is strongest when its authority remains narrow. It describes how the systems can name and exchange state. It does not observe a controller's private transaction, a line card's forwarding entry, the loss on a remote branch or the payload accepted by an application. Those are later facts under different control.
The leadership problem is therefore not whether a tree object exists. It is whether everyone is closing the same tree.
One advertisement crosses several state machines
For an SR P2MP P-tunnel, RFC 10018 places a tunnel identifier in the PMSI Tunnel Attribute. It identifies the policy by <Root, Tree-ID>, encoded in the exact order Tree-ID then Root. The Tree-ID is a 32-bit unsigned value that is unique in the context of the root. IANA records tunnel type 0x0C for an SR-MPLS P2MP Tree and 0x0D for an SRv6 P2MP Tree.
That compact advertisement can be mistaken for a complete operational object. It is not. On the ingress PE, an MVPN or EVPN module creates a candidate path in the SR P2MP Policy module when the relevant A-D route is originated. The candidate can include constraints and an optimisation objective. The policy module then communicates the candidate to a controller using PCEP, BGP, NETCONF or another mechanism. RFC 10018 explicitly leaves those controller procedures outside its scope.
The separation matters. The service module can report that it originated a route. The policy module can report that it accepted a candidate. A transport session can report that it delivered a message. The controller can report that it computed a tree. These are four different success conditions.
RFC 10018 says an implementation should advertise an SR P2MP PTA only when the underlying data plane is supported, while leaving the provisioning and determination of that support out of scope. That is a reasonable interoperability boundary. It is also a warning against reading too much into the advertisement. “Supported” can mean a capability was recognized; it does not tell the reader which nodes were programmed, which release combinations were exercised or whether the current service payload reached all leaves.
The common record should therefore be treated as an invitation to reconcile state, not as a certificate that reconciliation has happened.
A leaf route is membership evidence, not a packet receipt
RFC 10018 makes leaf discovery concrete. In an MVPN case, the ingress PE adds an egress PE to the policy's leaf set when it imports the applicable Intra-AS I-PMSI or Leaf A-D route. It removes the egress PE when that route is withdrawn. An egress PE joins the SR P2MP Policy as a Leaf or Bud after importing the relevant root advertisement, and it originates a Leaf A-D route when the Leaf Information Required flag demands one.
EVPN follows the corresponding pattern with IMET, S-PMSI and Leaf A-D routes. An IMET or S-PMSI advertisement can create the candidate path. Imported IMET or Leaf A-D state adds an egress PE to the leaf set. Withdrawals reverse the operation. The details are valuable because they give service membership a reviewable lifecycle instead of leaving it implicit inside a controller.
But the lifecycle still belongs to the control plane. A Leaf A-D route can establish that the egress PE presented the required membership evidence. It cannot establish that the ingress policy module consumed the latest update, that the controller received the changed set, that the controller selected the same topology epoch, or that every replication segment toward that leaf exists in forwarding.
Nor does the reverse direction collapse into one event. When an ingress advertisement is withdrawn, an egress PE withdraws its prior Leaf A-D route and leaves the policy. The ingress side removes a leaf when it receives the corresponding withdrawal. The controller must then consume the reduced set, revise or replace the PTI, and remove obsolete per-node state. A withdrawn route is the beginning of data-plane retirement, not proof that retirement is complete.
The correct audit key includes the leaf and time. “PE-12 is a member” is weaker than “PE-12 entered leaf set version 184 at 10:02:11, controller revision 921 consumed it, PTI instance 37 included it, and the corresponding path was removed under version 185 at 10:17:44.” Without the epochs, a current snapshot cannot distinguish a late update from stale state.
A computed PTI can still be an incomplete tree
RFC 9960 supplies the general SR P2MP Policy architecture on which RFC 10018 relies. A policy can have multiple candidate paths. A candidate can have zero PTIs if the controller cannot compute a tree. During make-before-break it can temporarily have more than one PTI, but one and only one must be active. If two instances of the active candidate are active together, duplicate traffic can reach leaves.
Each PTI has an Instance-ID. Its replication segments are identified in the control plane by a tuple that includes Root, Tree-ID, Instance-ID and Node-ID. This is the level at which a statement such as “the tree was installed” becomes testable. The root segment, every intermediate replication segment and the leaf segments are separate local objects even when they share the same Tree-SID value.
RFC 9960 acknowledges partial failure directly. A node should report successful installation. Installation may fail—for example, because a Replication-SID conflicts with another SID—and the node should report the failure, preferably with a reason. The controller should retry with an upper bound and should alert after terminal failure. It may tear down a PTI when some segment installations fail; tearing it down is recommended when the root segment fails.
That “may” creates a local decision surface. An operator can choose to forbid activation until every required downstream segment is acknowledged. It can permit a named subset under a time-bounded exception. It can retain a partially instantiated PTI for diagnosis while ensuring it is inactive. The standard does not choose the service obligation.
One approach described by RFC 9960 makes the safety logic visible: install leaf and intermediate segments first, then install the root and activate the instance. This order reduces the chance that the root begins sending into a tree whose downstream state is unfinished. It is not proof that every controller uses that sequence. It is a useful control question to put in a procurement review and an incident timeline.
An aggregate controller status such as ACTIVE is therefore inadequate unless it has a precise meaning. Does it mean the PTI was computed? That every install request was sent? That all required nodes acknowledged? That the root has the active instance? That forwarding counters advance? A credible system exposes the per-node transaction and the rule that converted those facts into the aggregate state.
The Tree-SID names forwarding; it does not certify the service context
The Tree-SID is the unique data-plane identifier of a PTI. The root encapsulates a payload into it. Provider routers use replication segments to make copies toward leaves. A leaf disposes of the Tree-SID and delivers the payload. That chain is elegant because a single identifier can steer one ingress packet into a distributed forwarding structure.
The identifier still does not contain the entire service result. When a P-tunnel is dedicated to one MVPN, the Tree-SID can be sufficient to identify the MVPN instance. When several MVPNs share a P-tunnel, an upstream-assigned MPLS label or an SRv6 Multicast Service SID provides the additional context. For SRv6, RFC 10018 specifies transposition rules and the new End.DTMC4, End.DTMC6 and End.DTMC46 endpoint behaviours. IANA assigns them code points 76, 77 and 78.
A packet can therefore traverse the right tree and still fail at the last semantic boundary. The wrong service label can select the wrong MVPN. A malformed or mismatched SRv6 service SID can prevent correct multicast-table lookup. An egress can have the tree state yet lack the right local service binding.
EVPN adds a different correctness gate. Ethernet-segment multihoming requires split-horizon filtering to prevent duplicate BUM traffic. In SR-MPLS that involves the ESI label position; in SRv6 it uses the Arg.FE2 argument with End.DT2M. A leaf that receives a packet is not necessarily a healthy leaf if it accepts a copy that should have been filtered or delivers the same BUM frame twice.
This is why “reachability” is too broad as a closeout label. The evidence must say which PTI carried which payload into which MVPN or EVI context, with which split-horizon state, and what the receiver observed. A Tree-SID is a forwarding handle, not an application receipt or tenant-authorisation token.
Ingress replication is not a smaller version of the same controller chain
RFC 10018 also supports ingress replication over SR. In that model, the ingress PE makes one copy per egress PE and sends each copy through unicast forwarding. The SR P2MP Policy module and controller do not participate.
The distinction is operationally decisive. Ingress replication can provide per-egress treatment, including colour-based selection of an SR-TE policy in the cases specified by the RFC. A P2MP PTI replicates inside the tree; consequently, the ingress dictates one traffic-engineering treatment for the tree rather than a different treatment for every leaf.
A dashboard or runbook that merges these modes loses the actor responsible for failure. In ingress replication, the evidence focuses on the egress membership, per-egress service identifier, matching unicast policy and individual copy. In SR P2MP, it must include the controller's PTI, stitched replication segments, active instance and tree OAM. “Multicast over SR” is not a sufficient operational category.
Migration between the modes also needs an epoch boundary. If an operator falls back from a damaged P2MP tree to ingress replication, it must show when the old root stopped injecting, when per-egress copies began, whether both modes overlapped, and when stale PTI state was removed. Otherwise recovery can introduce duplication while making the service appear restored.
OAM must name the instance it tests
RFC 9961 defines ping and traceroute mechanisms for SR P2MP Policies. They are necessary because a tree can fail to deliver traffic even when it was centrally provisioned. The probes target a specific candidate path and PTI, traverse its replication segments and return responses toward the root.
The precision matters during make-before-break and repair. A response from the old PTI cannot validate the new PTI merely because both belong to the same <Root, Tree-ID> policy. Implementations should be able to test each candidate and each PTI, including an inactive instance. The test record therefore needs the Instance-ID, probe time, intended leaf set and responses or timeouts by leaf.
Non-adjacent replication segments expose another boundary. P2MP OAM can test the replication structure, while ordinary unicast OAM may be needed for the unicast path that connects two non-adjacent replication segments. RFC 9961 explicitly places that unicast-path failure detection outside its own mechanism. A successful tree probe at one layer should not erase a missing test at the other.
Even complete P2MP OAM is not the last word. A diagnostic reply can prove that a bounded probe reached and was processed by the targeted forwarding context. It does not prove that production packets carried the correct service SID, that loss and jitter met the service objective, that split horizon suppressed the right copy, or that the receiving process accepted the content.
The last step needs a payload canary. That may be a sequence-numbered test frame, a cryptographically attributable content marker, a receiver counter tied to the service context or an application-level acknowledgement. Its design is local, but its identity must be tied to the same PTI epoch as the preceding evidence.
Closeout is a set equation
The cleanest operational model is to keep several leaf populations rather than one ambiguous count.
| Set | Question |
|---|---|
E |
Which egress PEs are intended by the service contract? |
A |
Which egress PEs appear in the applicable A-D and Leaf A-D state? |
C |
Which leaves did the controller accept for the current candidate path? |
I |
Which leaves have a complete path of successful replication-segment installations for this PTI? |
F |
Which leaves are represented in current forwarding for the active PTI? |
O |
Which leaves respond to OAM for this exact candidate and Instance-ID? |
D |
Which leaves received the expected payload in the correct MVPN or EVI context? |
W |
Which removed leaves have proven withdrawal, deprogramming and quiescence? |
The service is not complete merely because these sets have the same count. Their members and epochs must match. Twelve intended leaves and twelve OAM responses are not equality if one response belongs to a removed PE and one intended PE is absent.
The desired relationship depends on policy. A premium broadcast service may require E = A = C = I = F = O = D before activation. A maintenance mode may allow a named subset, but the difference needs an owner, reason, expiry, affected service and recovery test. Withdrawal closes when the removed member is absent from active sets and appears in W under the same retirement event.
This set model catches the failures that component dashboards normalize. E − A is a service-provisioning gap. A − C is controller-consumption drift. C − I is incomplete installation. I − F challenges the controller's success claim. F − O is a forwarding or diagnostic failure. O − D exposes the last semantic or application boundary.
No central institution needs to own every decision for this to work. The shared identifiers make independent records comparable. Each operator can choose its activation and failure policy, but it cannot honestly call the tree complete without defining the set relationship it accepted.
What the public record does not prove
RFC 10018 and its supporting standards document architecture and interoperable procedures. They do not establish that any named operator deploys the mechanism, that a vendor supports every option, that a controller exposes per-node acknowledgements, or that a particular leaf has failed.
The hypothetical dark leaf in this article is an operating test, not a reported incident. The standards justify why such partial state is possible and which records exist to distinguish its causes. They do not supply prevalence, outage duration, customer harm or product attribution.
The IANA registries prove that code points and tunnel types have assigned meanings. They do not prove that a running binary implements them. A standards-track status does not make deployment mandatory. The IETF provides a common language; the operator decides whether and how to use it and remains accountable for the forwarding surface it controls.
That boundary is consistent with Lu Heng's argument for Running-Code Primacy: an institutional or documentary claim cannot overrule the effect shown by the running system. It is also consistent with Minimum Initial Specification, Localized Future Decision, and Voluntary Adoption: the common layer can stay precise while later deployment and risk decisions remain local.
Running state is not automatically correct. A line card can hold stale state, a controller can compute from old topology, and a payload can reach a leaf that should no longer receive it. Primacy means the claimed effect must remain answerable to observable operation—not that observation loses its normative or service context.
Sources
- RFC 10018 — Multicast and Ethernet VPN with Segment Routing P2MP and Ingress Replication
- RFC Editor publication record for RFC 10018
- RFC 9960 — Segment Routing Point-to-Multipoint Policy
- RFC 9961 — OAM for Segment Routing P2MP Policy
- RFC 9524 — Segment Routing Replication Segment
- RFC 6514 — BGP Encodings and Procedures for Multicast in MPLS/BGP IP VPNs
- RFC 7988 — Ingress Replication Tunnels in Multicast VPN
- RFC 7432 — BGP MPLS-Based Ethernet VPN
- RFC 9572 — BGP Route Types for Multicast in EVPN
- RFC 9252 — BGP Overlay Services Based on Segment Routing over IPv6
- RFC 8986 — Segment Routing over IPv6 Network Programming Behaviours
- IANA — BGP Parameters
- IANA — Segment Routing Parameters
- Heng Lu — Running-Code Primacy
- Heng Lu — Minimum Initial Specification, Localized Future Decision, and Voluntary Adoption
- Heng Lu — On Reality Layers
- Heng Lu — On Data Sovereignty
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
