Summary
- Revision 12 of the CATS metric-definition Internet-Draft adds a dedicated test of whether normalized or aggregated steering values still reflect service-instance capability.
- It leaves validation intervals and divergence thresholds to operator policy; no common score or authentic message alone establishes application quality.
A traffic selector can receive an intact, fresh score from the correct publisher and still send work to the wrong service instance. Cryptographic origin answers who sent the number. A shared normalization function answers how different implementations made it comparable. Neither answers whether the number continues to predict the experience of a real application.
That missing comparison is the substantive addition in the 29 September revision of the CATS working group's metric-definition draft. Version 11 already described raw Level 0 measures, Level 1 category measures, a Level 2 global score, offline agreement on multi-vendor calculation functions, observation-time fields and security checks. Version 12 adds section 7.4: operators should periodically validate that normalized or aggregated values still represent actual instance capabilities.
The document proposes three kinds of evidence: trends against application QoE or SLA observations, distributions against known-good reference instances, and fixed workloads against expected baselines.
These are alternatives for validation, not a protocol-defined pass mark. The draft explicitly makes frequency and deviation thresholds matters of operator policy, while saying persistent divergence from application QoE should trigger recalibration. Its new section 7.5 also sorts management objects into configuration, operational state and statistics. A configured weight or normalization bound is not the live score; the live score is not its historical trend; none is the observed user outcome. The draft points to a separate CATS data-model document for YANG mapping, rather than claiming this metric definition supplies the management model.
The distinction has operational consequences. Version 12 further says cross-vendor Level 1 calibration can be negotiated per category, while a Level 2 composite should follow agreement on every category's aggregation and the global normalization function. Where those parameters cannot be agreed, parties may use a Level 0 measure with an explicit unit and source. That is a narrower, legible comparison—not permission to call any two vendor scores interchangeable.
The new text does not document a failed deployment or prove that any calibration method solves the problem. It is still an active working-group Internet-Draft in I-D Exists state, not an RFC. Its contribution is to make a previously implicit question explicit: when the service moves, loads change or software changes, who demonstrates that the steering score remains connected to the outcome it is supposed to represent?
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

