Summary

  • Revision 16 deliberately kept the NMOP incident model independent of confidence; revision 17 preserves a structured Probable Root Cause without a score, producer identity, model version or calibration record.
  • Neighbouring NMOP anomaly drafts do carry confidence, concern and annotator provenance. The operational risk appears when that upstream envelope is discarded while its cause classification survives.
  • A typed cause is a bounded hypothesis, not remediation authority. Operators need an external join from evidence and analytical version to incident revision, decision threshold, authorized action and observed service result.

A precise label can still be an uncertain claim

Imagine a control-room card that says an incident's Probable Root Cause is interface-hardware-failure. It names the affected node, the port and a loss-of-signal event. The priority is critical. The category and domain are present. Nothing about the card looks provisional.

Now ask a different question: was the conclusion produced by an optical threshold rule, a human engineer, a topology model, or a learned detector? Was the detector's confidence 51 or 99? Which version of the detector ran? Did another hypothesis finish one point behind it? Which telemetry interval was available, and which evidence arrived too late?

The base incident record does not answer those questions. That is not an accidental omission hidden in a field name. The change log for draft-ietf-nmop-network-incident-yang says that revision 16 kept the incident model independent of confidence. Revision 15 had described the related anomaly-detection work as scoring a result with confidence and concern. Revision 16 removed that sentence while preserving the relationship to the anomaly architecture. Revision 17, dated 24 September 2026, retains the separation.

The design choice deserves to be read literally. The incident vocabulary standardizes an operational object. It does not standardize every method by which an operator becomes convinced that the object has a particular cause.

What the common incident core actually carries

Revision 17 can represent an incident before its source is known. Its source list may be empty and populated after diagnosis. Once a cause is available, probable-causes can identify a network and node, optionally a resource, and attach a cause-name identity plus descriptive detail. probable-events can refer back to related events. Domain, priority and category are mandatory in the incident grouping.

That is useful structure. A receiver no longer has to infer whether one free-text paragraph means loss of signal, routing misconfiguration, service misconfiguration or protection failure. The identity hierarchy supports common classification. The node and resource references locate the claim. The event references keep a path back to observations.

But none of those fields expresses epistemic weight. Priority answers how urgently the incident should be treated, not how likely the diagnosis is to be right. Category answers what family of incident is represented, not how the conclusion was calibrated. A related event shows association, not causal sufficiency. A formal identityref makes a label interoperable; it does not make the label true.

The draft's own definition sets a demanding semantic bar. A Probable Root Cause is the fault condition whose complete removal stops the incident and prevents recurrence; a contributing condition is not the same thing. Yet the cause object has no standardized field recording a removal experiment, recurrence window, counterfactual test or competing explanation. The definition tells implementers what the phrase means. It does not certify that every populated instance has already passed the test.

Confidence did not disappear from NMOP; it stayed in another record family

The neighbouring anomaly work makes the architectural split visible. draft-ietf-nmop-network-anomaly-lifecycle-07 defines a 0–100 confidence score for an anomaly or detection strategy and a separate 0–100 concern score for operator attention. Its model can carry the anomaly's version and stage, an annotator identity and name, whether the annotator is human or algorithmic, and the annotator's version. Detection, validation and refinement form a loop rather than a one-way declaration.

draft-ietf-nmop-network-anomaly-semantics-06 carries the same distinction into a serialized vocabulary: confidence, concern, strategy, annotator type and annotator version are explicit. Its examples also allow confidence to be null. The envelope can therefore be incomplete; merely retaining it does not manufacture certainty.

The important point is narrower. A network may have a rich upstream account of how an anomaly was produced and judged, while the normalized incident record contains only the selected cause and its related events. If the two remain joined, the separation is clean. If the join is lost, the downstream cause can look stronger precisely because the fields that qualified it have vanished.

This is a classic compression hazard. Aggregating thousands of alarms and metrics into one incident is valuable because it reduces noise. Compression, however, is a decision. It selects what survives. A standardized cause name may survive because downstream systems need it, while detector version, rejected hypotheses and calibration context are treated as analytics detail. Later, a remediation engine sees the durable label but not the uncertainty that accompanied its creation.

Confidence and concern are not interchangeable

The anomaly drafts distinguish two questions that dashboards often collapse. Confidence asks how strongly the detector regards an observation as anomalous. Concern asks how much attention or corrective action the condition deserves in light of likely service impact and operator knowledge.

A high-confidence anomaly can deserve low concern: a predictable traffic shift may be unmistakably unusual yet harmless. A low-confidence diagnosis can deserve high concern: a weak signal that might indicate an optical failure under a critical service can justify cautious investigation. Incident priority adds a third dimension. It can be critical because potential impact is severe even when the leading causal hypothesis is still uncertain.

Therefore, copying one number into all three places would be worse than leaving confidence out of the base model. The scores depend on their producer, population, training or rule set, observation window and calibration method. A “90” from one detector is not automatically comparable with a “90” from another. Independence protects the common incident object from pretending that a method-specific scale is universal.

The operational obligation follows from that virtue. Systems must preserve the distinction outside the common core instead of replacing it with a single green, amber or red badge.

Authorization protects action; it does not validate the diagnosis

Revision 17 expects YANG management access through secure transports and mutual authentication, and it points to NACM for restricting operations and content. Those controls matter. Incident records can expose customer-identifiable services and the broken shape of a network. Diagnose and resolve operations can consume resources or affect running services.

Authentication can establish which management peer connected. NACM can decide which principal may read a probable cause or invoke an operation. Neither control measures whether the cause was well calibrated. A principal may be fully authorized to act on a poorly supported hypothesis. Conversely, a high-confidence diagnosis does not itself grant permission to change the network.

The receipt chain must therefore retain both dimensions: epistemic support for the cause and institutional authority for the action. Treating one as the other gives an analytical score power it was never designed to hold.

Running code is evidence, with a boundary

The draft includes an implementation-status section stating that Huawei iMaster NCE implements the incident model with intent-management and AI tools, including RESTCONF, incident lifecycle management, notifications and incident-list query. That is relevant running-code evidence.

The same section says the information was supplied by contributors, was not verified by the IETF, does not imply endorsement and is not a catalogue of available implementations. The frozen record supplies no interoperability trace showing how confidence and annotator provenance survive incident creation, no accuracy measurement, no calibration report and no production incident outcome.

The honest conclusion is neither “there is no implementation” nor “the architecture is proven.” A named implementation exists in the document's disclosed status. The confidence-preservation question remains an empirical one.

This is not the earlier command-to-clear argument

BTW already publishes A Cleared Incident Does Not Prove Which Command Cleared It. That article examines revision 14's asynchronous boundary between an incident-resolve request and a later cleared notification. It asks which command produced a state change.

This briefing asks a different question introduced by later revisions: what epistemic evidence accompanied the cause before anyone chose a command? A perfect request identifier would not restore a discarded confidence record. A perfectly correlated repair could still be based on an overconfident diagnosis. Conversely, a well-calibrated cause would still need separate authorization and outcome evidence. Both articles remain necessary because causal belief and causal execution fail at different joins.

Sources