Summary

  • RFC 9940 distinguishes a measured Value, an Event, a Fault, a Problem, a Symptom, a Cause, an Alert and an Alarm; its workflow diagrams explain relationships rather than grant an automatic remedy.
  • An alarm can identify an undesirable state and an operator can close work on it, while the causal question, service outcome and authority for the next change remain separate records.

A dashboard changes from green to amber. The visual change is real: some monitored characteristic crossed a threshold, a system classified it as relevant, and an operator deserves to know. The dangerous shortcut begins when that sequence is retold as one sentence: “the fault caused the outage, so the automation fixed it.” Each clause may require a different source, owner and test.

RFC 9940, an IETF Informational document published in April 2026, supplies terms for network fault and problem management at the network layer and below. It is deliberately a common vocabulary for management models and protocols that report or manage faults and problems. It is not a protocol that discovers a cause, proves service recovery, or assigns a system permission to change another system.

The terms make the chain more exact. A Characteristic is an observable or measurable aspect of a Resource; a Value measures it. A Change is variation in the value over time, while an Event is such variation at a distinct moment. A Condition interprets one or more values, and a State is a condition a resource has at a particular time. That is already a progression from a reading to an interpretation. A single loss sample, for example, is not automatically a degraded service state.

The next transition is conditional, not mechanical. Relevance depends on policy, perspective, intent and other information. An Occurrence is a relevant Event or Change. A Fault is an unwanted Occurrence that may point to a current or future unwanted State. A Problem, by contrast, is an undesirable State that may need remedial action and cannot necessarily be tied to one Cause. A Symptom indicates a Problem; a Cause may be indicated or determined from multiple faults, problems and symptoms. An Alert indicates a Fault. An Alarm signifies an undesirable resource State requiring corrective attention.

The distinction matters most after the graph stops moving. RFC 9940's light-loss example allows services to return while the recent fault remains unexplained; a repaired microbend can settle one causal question but still leave the prevention-of-recurrence problem open. “Current service looks normal” is therefore not a universal closeout condition. It is evidence about one state at one time.

RFC 8632 makes an adjacent operational safeguard explicit. Its candidate root-cause resources are hints for the client, not settled attribution. It also separates an alarm's is-cleared state from an operator's closed state. A cleared signal can tell the operator that a condition no longer presents in the observed form; a closed work item says someone considers corrective action successful. Neither field, by itself, reproduces the measurement population, proves a cause or demonstrates that every dependent service is operating as intended.

The source documents also resist a second collapse: telemetry into intent. RFC 9940 describes telemetry, monitoring, analytics and observability as a chain, and says telemetry data does not contain service-definition intent. RFC 9315 defines intent as declarative goals and outcomes; the network does not automatically know a provider's particular goals. RFC 9417 separately names a metric, a symptom, a health score and the case in which a metric cannot be collected. An inferred health score may guide attention; it is not a standing authorization to reroute, withdraw, bill, declare an incident over, or make a customer-facing promise.

This is where Daniel Kade applies the reality-layer discipline in docs/heng-lu-note.md as an editorial lens, not as an IETF rule. Preserve the source reading, the threshold and policy that made it relevant, the classification, the causal hypothesis, the authorized decision, the actual change and the post-change observation as distinct facts. Then a system can be quick without asking a familiar label to do the work of an evidence chain.

Sources