Summary

  • CATS metric-definition revision 11 followed a 4 September merge that added optional observation and validity times plus substantial operational guidance. It remains an active working-group Internet-Draft, not an RFC.
  • The working group made an explicit August choice not to place a normalization-function identifier in every metric. Vendors instead negotiate functions before deployment and synchronize a version-controlled configuration manifest offline.
  • Runtime components must then assume that scores use the agreed functions and are fully comparable. The metric field list does not identify the active manifest version, digest or epoch.
  • A partial configuration rollout can therefore leave two authentic, fresh scores on different semantic scales without giving the selector an in-message way to prove the difference.
  • CATS does not need runtime algorithm negotiation. It needs a smaller control: verify a shared manifest identity before admitting a score, declare the mismatch action and retain a decision receipt.

The update chose an operating model

On 4 September, maintainers of the CATS metric-definition repository merged pull request 42. The fixed merge commit became the basis of revision 11. The Datatracker record describes an active CATS working-group document whose IESG state is still “I-D Exists.” It is a proposal under development, not an RFC and not evidence that a network has deployed it.

The revision nevertheless makes an important operational commitment. CATS—Computing-Aware Traffic Steering—needs a selection function to compare service instances using network and computing information. Raw measures such as delay, utilization or available resources may have different units and ranges. The draft permits Level 1 category scores and a Level 2 global score, with normalization and aggregation turning unlike observations into comparable values.

That comparison is only meaningful if the functions behind it agree. Revision 11 says that in multi-vendor deployments all parties must settle the score range, normalization method and parameters, aggregation formula and weights, and whether higher or lower is better. The results should be compiled into a formal configuration manifest, sent offline during initialization to every component that needs metrics for decisions, and placed under version control.

Then comes the decisive sentence: at runtime, CATS components must assume received scores used the agreed functions and are fully comparable. Dynamic negotiation is unnecessary. This is not an accidental omission dressed up after the fact. It is the architecture the authors chose.

A rejected field explains the tradeoff

The decision has a public history. An August contributor proposed adding three per-instance fields: the applied normalization function, observation time and validity interval. The proposal argued that a selector holding two values otherwise could not tell whether the underlying conditions differed, the functions differed or one observation had expired. Pull request 41 supplied concrete text and a registry for parameter-bound function identities.

In a 25 August mailing-list response, a draft co-author accepted the two time fields as optional but rejected function identity in every report. Cross-vendor consistency, the reply said, should be arranged out of band at initialization. Making a selector parse function identifiers and adjust its algorithm while CATS operates would add excessive complexity.

The contributor then accepted that boundary: an offline agreement is stronger than reasserting the function on every message if the framework never contemplated per-message renegotiation. Revision 11 accordingly carries optional Observation_Time and Validity_Interval, but no function field. A comment on PR 41 says PR 42 incorporated those two optional fields; PR 41 itself remained open at the evidence cutoff.

That history matters. The governance question is not “why did the authors forget provenance?” They did not. It is whether the lower-complexity model proves the single fact on which it depends: all decision-making components are actually using the same agreed configuration now.

Version control is not shared state

A repository can hold a perfect sequence of manifests while a running fleet holds inconsistent copies. Revision 11 says the manifest should be version-controlled and synchronized offline. It does not identify a canonical digest, an activation epoch, a completion condition or the response when one component misses an update.

This exact weakness appeared before merge. A reviewer commented on PR 42 that synchronization cannot be assumed always to succeed in a distributed system. The proposed addition required updates to reach all related components and failed or incomplete synchronization to be detected and reported. The merged text kept the manifest and version-control sentence but not that proposed sentence. A review comment is not working-group consensus; it is evidence that the failure mode was visible before the new revision shipped.

Consider a hypothetical rollout. Vendor A's metric producer and the selector still use manifest M1. Vendor B's producer has activated M2, which changes a normalization bound or one aggregation weight. Both producers report a Level 2 value of 7. Both messages can have valid signatures, arrive within their validity intervals and be accepted from authorized publishers. Yet the two sevens are not necessarily the same proposition. The selector sees the result, not the configuration lineage that produced it.

The draft's field list makes the distinction concrete. Source says whether a value was directly measured, estimated, aggregated or normalized. Observation_Time can say when it was observed. Validity_Interval can limit its useful life. None names the manifest. Freshness answers “is this value old?” A signature answers “did an authorized key bind these bytes?” Neither answers “did producer and selector use the same rules?”

The new alarms stop one layer short

Revision 11 responds seriously to an early operational review. That review had marked revision 10 “Has issues” and asked for cross-vendor comparability guidance, policy identifiers, calibration, failure handling, management expectations and migration advice. The new Operational Considerations section is not decorative.

It discusses the visibility/overhead tradeoff between Level 1 and Level 2, recommends restraint in routine advertisement rates, and allows event-driven updates. When metrics go missing or fail freshness checks, an operator may use last-known-good values for two or three measurement windows, downgrade or exclude an instance, or fall back to network-only information. Operators should alarm on freshness failures, component failures and anomalous scores, and log fallback events.

Those are useful behaviours. None closes the configuration-state gap. A score produced under M2 is not necessarily stale, missing, malformed or anomalous. It may be completely plausible under M1. A component can be healthy while holding the wrong version. Last-known-good can even prolong the ambiguity if “good” was judged without recording the governing manifest.

The current CATS framework confines the system to one administrative domain and requires common functions across participating components. That reduces the number of authorities involved. It does not make partial rollout impossible. Within one operator's domain, configuration equality is easier to enforce—and therefore reasonable to prove.

Bind the decision, not the algorithm negotiation

The repair can preserve the authors' complexity budget. A selector does not need to receive a sigmoid definition, recompute every score or negotiate weights on the wire. Before admitting scores, the control plane can establish a small piece of shared state: canonical manifest identifier and digest, administrative-domain and component scope, version or activation epoch, effective time and synchronization-completion status.

Each producer and selector can attest which digest is active. The live message may carry that compact identity, or the transport/session can bind an authenticated identity to all messages it conveys. The important requirement is not placement in one particular packet. It is that a steering decision can be joined to the manifest under which the producer calculated and the selector interpreted the score.

Mismatch behaviour should be explicit. One deployment may reject the value. Another may use Level 1 metrics whose interpretation remains available, fall back to network-only information, or hold the decision until the rollout completes. These are local operational choices. What cannot remain implicit is that an unrecognized score was treated as comparable.

A compact receipt would retain the producer, selector, score identity, observation/validity clock, active manifest digests, synchronization result, chosen fallback, alarm acknowledgment, decision and restoration condition. The manifest's sensitive weights and thresholds need not be public. A digest proves identity without publishing the configuration itself.

RFC 8911 offers a useful precedent, not a ready-made CATS rule: its registry model binds a performance metric to fixed parameters left open by a definition. CATS can choose a different mechanism. The principle is the same—the name of a result should preserve enough context to know when two implementations measured the same thing.

Sources

  1. IETF Datatracker — CATS metric definition
  2. CATS metric definition, revision 11
  3. CATS metric definition, revision 10
  4. IETF Datatracker revision history
  5. OPSDIR early review of revision 10
  6. CATS list — co-author decision on offline function agreement
  7. CATS list — contributor response on that boundary
  8. Pull request 41 — proposed per-instance provenance
  9. Pull request 42 — operational considerations
  10. Merge commit for pull request 42
  11. Current CATS framework
  12. RFC 8911 — Performance Metrics registry
  13. Heng Lu on minimum shared rules and local choice