Summary

  • draft-ietf-cats-oam-fw-01 is an active CATS Working Group Internet-Draft. It is a WG Document with an IESG state of I-D Exists, not an RFC, deployment record or operator policy.
  • The draft distinguishes link, path, instance and service OAM. It can collect telemetry, assess end-to-end conditions, and verify whether observed forwarding aligns with a CATS Path Selector decision.
  • Administrators define the steering policy, including the weight given to network and computing metrics. The draft requires detection of push/pull inconsistency but does not mandate a single resolution strategy.
  • A telemetry-to-remedy receipt should keep observation, authorized policy, conflict disposition, chosen action, execution and outcome review separate without exposing customer data or attack-enabling telemetry.

An observation layer is not an operating mandate

The CATS Working Group's charter starts with a familiar distributed-service problem. Multiple service instances may sit at different sites. A client experience can depend not only on bandwidth and latency, but on computing capacity, storage and processing conditions. The charter therefore contemplates a framework for distributing network and compute metrics so that an edge node may steer traffic toward a service instance. It confines the assumed model to a single domain, and describes that steering function as operating inside the service provider's network while logically separated from service operation.

That final distinction matters. A network-side steering component may have a legitimate, narrow job: interpret the signals made available to it and direct traffic under an existing local policy. It does not thereby become the owner of the service, the party entitled to alter the service's recovery plan, or the institution that may settle a commercial or customer consequence. A framework can make the interface between those roles more legible. It cannot merge them without a separate grant of authority.

The present OAM text is still a working document. The Datatracker calls it an active Internet-Draft, records it as a CATS WG Document, lists an IESG state of I-D Exists, and gives it an intended Informational status and a January 2027 expiry. Those are meaningful process facts. They show that a working group has taken the document forward. They do not show that any carrier, cloud provider, hospital, vehicle operator or enterprise has deployed it; nor do they establish a policy applied to a particular customer flow.

That restraint should not be confused with dismissal. A draft can clarify a design question before it becomes an RFC. But an IETF work item is evidence of a standards-process state. It is not a universal implementation command. The source material itself says Internet-Drafts may be updated, replaced or obsoleted and should be treated as work in progress. A reader who turns that state into a field report or operating authorization has skipped the very boundary the document needs operators to preserve.

Four layers make a better question, not a sovereign answer

The draft's contribution begins with differentiation. Link OAM concerns a physical link or single-hop interface. Path OAM concerns a path between CATS Forwarders. Instance OAM concerns a service instance's computing resources and operational health. Service OAM reaches from an ingress CATS Forwarder toward a target service instance and can measure end-to-end conditions such as liveness, transaction success rate and application-layer latency.

Those layers are valuable because a reachable IP interface is not proof of an available service. The draft identifies the familiar failure mode: a node can remain reachable while its application has crashed, deadlocked or exhausted resources. It also identifies another limit: the same client-visible degradation can reflect network congestion or resource exhaustion in a service instance. A combined OAM view can reduce the temptation to call every symptom a network failure.

But more layers do not eliminate judgment. An instance CPU figure is not the same object as an application transaction. An end-to-end latency measurement is not a contractual interpretation of an SLA. A path trace can help localise a fault; it does not establish the cause of a business interruption, the allocation of loss, or the person who must approve a change. The draft says that its specific Instance and Service OAM protocol implementations are out of scope. That is not an omission to hide. It is a public reminder that an architecture, a protocol choice and an operating procedure remain separate things.

The charter supplies a parallel limit. It seeks to study metrics, their distribution and the control/data-plane building blocks that use them. It does not place the service operator's internal governance inside the CATS WG. The draft may say what an OAM component needs to observe. It cannot decide which customer population warrants a failover, which maintenance window is permitted, what data may be discarded, or which local safety constraint outranks a lower latency value.

Verification proves alignment to a prior choice

The most easily overstated phrase in the draft is "policy verification." Service OAM is described as checking that the actual forwarding path from an ingress CATS-Router to the selected service instance aligns with the steering decision made by the CATS Path Selector, or C-PS. That is a strong and useful proposition: it asks whether execution matches a prior steering decision.

It is not a decision engine hiding inside an evidence engine. The question is not "what should we do?" It is "did the forwarding state correspond to what the named policy mechanism selected?" A result of alignment does not prove that the prior policy had the correct network-versus-compute weights. A result of misalignment does not select a new weight, direct the next remediation, or tell an operator whether to drain an instance, change a deployment, page a team, pause a release or leave a degraded service running for a safety reason.

The draft makes the policy owner legible: administration requirements say the system must allow administrators to define CATS-specific policies, including weighting factors for network and computing metrics. Those policies tell the C-PS how to interpret raw telemetry when calculating a "best" service instance. The telemetry pipeline can supply evidence to the calculation. The C-PS can apply the policy. The administrator or another accountable local authority remains responsible for the policy's content and for any later change to it.

This is where Heng Lu's distinction between evidence and mandate is most practical. A rich measurement may make inaction harder to defend. It does not make the sensor the principal. If a metric collector is treated as the actor that chose a remedy, a later review cannot tell whether the real issue was stale data, a policy version, an authorised override, an implementation mismatch, a safety constraint or a human decision that did not follow the recommendation. Responsibility becomes a property of a dashboard rather than a traceable act.

Detection of a conflict is not a rule for resolving it

The draft deliberately allows both push and pull modes for Instance OAM. In one, a computing node reports aggregated data periodically. In the other, a receiving entity retrieves data on demand. When values obtained in the two modes are inconsistent, the system must be able to detect the inconsistency. The document then says it does not mandate a single conflict-resolution strategy, because the optimal approach depends on deployment scenarios.

That sentence is a governance boundary, not a defect to be papered over with implied automation. It preserves questions that only a deployment can answer: Which source is authoritative for this metric? How much staleness is acceptable for a particular class of traffic? Is an outlier evidence of load, a clock problem, an integrity problem or a data-path failure? May a system fall back to a last known value? May it suppress steering changes during a maintenance window? Who can override a default when an automatic reaction would increase another risk?

The operation requirements reinforce this. They contemplate timestamps and an optionally configured maximum acceptable staleness threshold for each metric type. They ask for periodic and threshold-triggered reporting to balance precision and control-plane overhead. None of those settings is a neutral fact. Each turns a distributed measurement into a decision rule with costs: excessive change, delayed response, wasted capacity, unnecessary traffic movement or an overlooked outage. The appropriate setting may differ among a public web service, a time-sensitive industrial system and an internal batch workload.

The responsible answer is not to demand that an Internet-Draft dictate every local action. It is to make the local decision visible when a deployment uses the framework to act. An operator may choose a conservative hold, a short automatic drain, an escalation to an on-call authority, a verified failover, or an explicit decision not to move traffic. The draft supports the information needed for such choices. It does not make one of them universally right.

The missing public object is a telemetry-to-remedy receipt

For a material steering change or a material non-action, a deployment should be able to preserve a compact receipt without publishing sensitive raw data. The first field is the evidence boundary: the collection source, measurement time, subject and freshness rule. The second is integrity: whether the relevant channel, identity and validation checks passed. The third is policy: the policy version, its authorised owner, the relevant threshold or weighting class, and any approved exception.

The fourth field is conflict disposition. If push and pull differ, the record should say that they differed and name the rule or responsible role that selected the usable input. It need not publish capacity figures, customer identifiers or exploitable topology. A bounded state and a reason are enough to distinguish a measured discrepancy from a silent override.

The fifth is the action boundary: who selected a response, which response was selected, and what execution evidence later showed. The sixth is correction: a rollback condition, expiry or review date. The seventh is outcome: a separate observation of whether the service effect improved, rather than an assertion that moving traffic proved success. This is Daniel Kade's recommendation, not an IETF requirement or a claim that CATS deployments currently retain such records.

The receipt protects both sides of the operating relationship. It prevents a policy owner from saying "the telemetry decided" when a person or local rule made the decision. It prevents a responder from being blamed for a decision that was actually produced by stale input or a policy they did not control. And it gives a later auditor a way to distinguish a forwarding mismatch from a policy defect, an authorised exception, a remediated operational event or an unresolved outcome.

Sources

  1. draft-ietf-cats-oam-fw-01
  2. CATS Working Group charter
  3. RFC 7282 — On Consensus and Humming in the IETF
  4. Heng Lu, The Multi-Stakeholder Mirage