Summary

  • RFC 9544 defines compact counts and ratios for intervals in which a service violates configured SLO thresholds; its “precision” refers to the promise being assessed, not to the accuracy of the measuring instrument.
  • The result inherits every prior choice: measurement points, traffic direction, interval length, packet population, thresholds, statistical window and the exact SLO version in force.
  • A PAM can support monitoring and accounting, but it does not supply a complete availability state machine, prove user-visible harm, identify cause, trigger a contractual remedy or show that remediation occurred.

The most important field in an availability report is often the one executives never see: who chose the denominator?

A dashboard may show a Violated Interval Ratio to four decimal places. The number can be calculated exactly. Yet its meaning changes if the observation points sit at the provider edge rather than the customer application, if intervals last one second rather than ten milliseconds, if maintenance is outside the governing denominator, or if the critical threshold was revised halfway through the quarter. Arithmetic does not remove those choices. It seals them into the result.

RFC 9544, published by the IETF as an Informational RFC in March 2024, is valuable because it defines a compact language for that result. It calls the language Precision Availability Metrics, or PAMs. The word “precision” is easy to overread. The RFC explicitly says it describes the precision of the service requirement being assessed, not the precision of the mechanism performing the measurement. Accurate measurement methods are a separate subject.

That boundary turns PAMs from a universal truth machine into something more useful: a disciplined answer to a bounded question. Given this SLO, these observation points, these thresholds and this measurement session, when did the service comply?

Availability can fail while connectivity survives

RFC 9544 uses “availability” in a specific way. A service governed by SLOs is unavailable whenever those SLOs are violated, even when basic connectivity still works. A reachable path can therefore be unavailable because latency, loss, throughput, ordering or another measured parameter crossed the agreed line.

This is not wordplay. It separates “the network carried packets” from “the network delivered the promised service.” It also creates a risk of false equivalence. A violated SLO is not automatically a total outage, and it is not necessarily visible to every user in the same way. Conversely, a green connectivity check does not prove compliance.

The RFC classifies each fixed time interval in a measurement session as a Violated Interval, a Severely Violated Interval or a Violation-Free Interval. A VI exists when at least one performance parameter misses its configurable optimal threshold. An SVI exists when at least one misses its configurable critical threshold. A VFI requires all parameters to remain at or better than their optimal levels.

Those labels carry two facts at once: what was observed and what rule was applied. “Severe” is not an intrinsic property floating in the packet. It is a relationship between an observation and a configured critical threshold. The RFC leaves the mechanism for setting that threshold outside scope.

Cadence changes the event

The interval is variable, although it must remain constant within a session. The RFC cites one second as reasonable and gives a decamillisecond as an example of finer granularity. That choice alters the event the counter can see.

A brief burst may contaminate one large interval, produce several small violated intervals, or disappear inside an aggregate, depending on the measurement and classification rules. Two services can report the same VIR while one suffered frequent short violations and the other a compact cluster. A mean time between violated intervals can summarize recurrence without preserving the longest disruption or its internal shape.

Packet counts add useful texture. Violated Packets Count and Severely Violated Packets Count can help distinguish an isolated outlier from a broad breach across many packets. They still do not identify cause. A queue spike, route convergence, insufficient resources, a customer-side bottleneck or a faulty instrument may produce similar surface evidence. Attribution needs its own receipt.

Statistical SLOs add another layer. A contract may allow a proportion of packets to exceed a target. In that case, one outlier is not necessarily a breach; the distribution and its window decide. RFC 9544 describes that problem but leaves SLO-aligned histogram metrics for future work. Tail harm to a particular user can coexist with aggregate compliance.

Compression is not replay

PAMs are designed to avoid retaining and post-processing an unlimited time series of raw measurements. The compact counters can be actionable at scale. That is an operational advantage, not proof that the summary preserves everything needed later.

A dispute asks different questions. Which packets were included? Were clocks aligned? Did the observation cover both directions and the relevant path? What happened during missing data? Which SLO version applied? Can the classification be replayed? What did the user experience? A compact ratio may answer none of them unless the organization preserved a linked audit record.

The RFC itself reinforces this discipline in its security section. Instrumentation and SLO configuration are attractive targets. A metric instance must be unambiguously bound to the SLO that was in force. If that binding cannot be made, the RFC says it is preferable to retain service-level statistics without claiming which observations were violations.

That is a stronger rule than “secure the dashboard.” It means the promise is part of the security perimeter.

A counter does not create a consequence

RFC 9544 discusses how PAMs might feed an available/unavailable state model. It offers possible transition settings based on successive violated or clean intervals. But the state model is outside scope. There is no universal outage trigger, restoration rule, hysteresis period, alarm route or service-credit decision hidden in the RFC.

Nor does the document supply the remaining implementation machinery. A YANG model, IPFIX Information Elements, longest-disruption measures and statistical histograms are listed as future work. The RFC has no IANA actions. The concepts are defined; several configuration and export surfaces are not.

The evidence chain therefore has to continue beyond the metric. A credible claim needs the agreement and the authority to amend it; the exact SLO, endpoints, direction and thresholds; the measurement topology, clocks, packet population and missing-data policy; the session interval and configuration hash; enough raw evidence to audit classification; any chosen state-transition rule; the alert or operational action; the clause mapping a breach to a consequence; and independent evidence of user impact and restoration.

Heng Lu's reality-layer distinction is helpful here. The SLO is a written rule. The PAM is a classified record. A reroute, repair, capacity change or credit is an executed act. Restored service and avoided loss are outcomes. Moving from one layer to the next requires a receipt, not a more emphatic label.

RFC 9544 makes the common layer thinner and clearer. That is its strength. It provides names that providers and customers can use without pretending to choose their commercial bargain. The mistake would be to inflate a shared metric vocabulary into authority over the promise itself.