Summary
- Revision 07 proposes YANG structures for unitary OAM tests and user-ordered test sequences with periods, recurrence, status and schema-mounted device models.
- A planned or successful sequence can be a truthful orchestration record without proving which version ran, whether every step and packet population was comparable, which cause was unique, who could change the network or whether service actually recovered.
The schedule solves coordination, not epistemology
Network diagnosis often begins with an improvised sequence: run a continuity check, trace a path, test delay, inspect a device, repeat during the suspected window, then compare the results. The value of draft-ietf-opsawg-scheduling-oam-tests-07 is that it attempts to make this sequence declarative. Its proposed ietf-oam-unitary-test module identifies test definitions and target network elements. Its ietf-oam-test-sequence module places those tests in user-controlled order and adds time constraints, recurrence and execution count.
That is a useful common surface for controllers and orchestrators. It can make a diagnostic plan reviewable before it runs and give recurring work a consistent lifecycle. It can also create a dangerous visual shortcut: one row says success, a metric field is populated, and the surrounding organization begins to behave as if a root cause has been established.
The Datatracker record supplies the first restraint. Revision 07 is an active OPSAWG Working Group Internet-Draft intended for Standards Track. It has no RFC number and the IESG has not begun processing it. The document history shows two completed early reviews with issues and a YANG Doctors review still incomplete. The work is a proposal under development, not proof of implementation, deployment or conformance.
The deeper restraint is operational. Scheduling is one layer of evidence. Diagnosis, authority and recovery live in later layers. A defensible system needs at least eight receipts.
Receipt one: the plan that was meant to run
The draft reuses the state, version, local time, counters and occurrence fields of RFC 9922. A unitary test can reference one or more ne-config entries, each with a node identifier, a managed flag, a test-type identity and a schema-mount root. A sequence is an ordered list of such test references.
The first receipt must freeze that material as one plan identity: plan name, version, ordered test list, node set, mounted module and library identity, parameters, schedule rule, time zone and approving owner. A human-readable name is not enough. A version counter is useful only if the system states who increments it, what change it covers and which immutable digest corresponds to the configuration that reviewers approved.
RFC 9922 explicitly allows the version to be maintained by the embedding entity and notes that it may be unused. It is therefore not a universal content hash. If a sequence is edited between review and execution, the evidence must bind the later occurrence to the exact version and test definitions that fired.
Receipt two: acceptance is not application
Revision 07 models planned, configured, ready, on-going, stop, error and success. The sequence adds failure. These states make orchestration observable, but their meaning is local to the implementation that reports them.
The draft itself does not define an RPC or action for immediate execution. The OPSDIR early review calls out this mismatch between the abstract's on-demand ambition and a model centered on period or recurrence configuration. An accepted NETCONF or RESTCONF edit can prove that a server processed a request. It does not by itself prove that every target network element received the intended configuration or that a particular test became runnable.
RFC 8342 makes the distinction concrete: running, intended, applied and operational configuration are related but not identical. Configuration can be present in intended state while a resource prevents it from appearing in operational state. The application receipt must therefore compare the intended plan with applied/operational state at every target, record partial failures and preserve the schema and capability view used for that comparison.
Receipt three: which occurrence actually happened
A recurring schedule is a rule, not an execution instance. last-occurrence, upcoming-occurrence, a counter and a local clock help locate what the scheduler believes happened. They do not necessarily identify a single run across retries, clock changes, overlapping recurrences, controller failover or delayed work queues.
The occurrence receipt needs its own immutable identifier. It should bind the scheduled instant, actual start and finish, schedule version, orchestrator identity, time source and synchronization quality, retry lineage, skipped or deferred status, and the target-node set actually reached. A success reported after the planned window may be useful, but it is not proof that the test sampled the incident window the analyst intended to study.
This matters especially when results from several devices are compared. A timestamp in a YANG leaf is a representation. Trust in the measurement requires the clock source, offset bound and device identity that give the timestamp meaning.
Receipt four: whether the sequence is complete
Revision 07 makes ordering explicit with ordered-by user. It also says an error in one or more unitary tests does not prevent subsequent tests from executing. That behavior preserves useful evidence after a partial failure; it also means “the sequence finished” and “the diagnostic procedure was complete” are different claims.
The PERFMETRDIR early review notes an unresolved asymmetry: stop leads to success in the unitary-test figure but to failure in the sequence figure. OPSDIR separately asks for clearer failure-versus-error semantics, cross-node rollback, notifications and correlation. Until those questions are resolved, consumers should not infer completeness from one terminal identity.
A step-completeness receipt must list every required unitary test, its prerequisites, node, start and finish, result pointer, error, continuation decision and effect on later interpretation. If three of five tests ran, the record must say which causal hypotheses can no longer be excluded. Silence is not a passed step.
Receipt five: what the result actually measured
The scheduling draft deliberately does not redefine detailed inputs or outputs. It expects them to come from mounted device-level OAM modules such as the TWAMP model in RFC 8913. RFC 8528 explains why that indirection must be verified: schema mount defines how a model appears below another model, but it does not assume the source of instance data, and mount-point instantiation and control can remain outside its scope.
The measurement receipt must therefore cross the mount boundary. It should name the concrete test type and module revision, parameters, source and destination, direction, packet population, traffic class, size and rate, actual sampling interval, clock semantics, device identity, raw result or digest, and retention path. A scheduling status that points toward a mounted result is not the result itself. A result field without its population and units is not an operational conclusion.
RFC 7799 distinguishes active, passive and hybrid measurement methods. That distinction determines what traffic was observed or introduced. A diagnostic record must not silently move from “the probe measured this” to “the service experienced this.”
Receipt six: whether the probe shared the relevant reality
The closest existing evidence boundary is RFC 10014. It separates OAM method from topological path congruence and from equal forwarding treatment. A probe can traverse the same nodes and links as production traffic while entering another queue, using another QoS mark, selecting another ECMP member, avoiding the policer that dropped customer packets or missing the load that created the fault.
The scheduling draft can order an appropriate OAM test; it cannot make that test comparable by declaration. The comparability receipt must state what the probe shared with the affected service: topology, forwarding class, encapsulation, hashing inputs, maintenance domain, time window, load, failure mode and relevant device processing. If only topology is common, the conclusion must remain topological.
This is not an argument against active tests. It is an argument for assigning them the authority their design earns. A dedicated packet may isolate reachability quickly. A hybrid method may carry information closer to production traffic. Neither becomes a universal service witness merely because an orchestrator scheduled it.
Receipt seven: a likely cause is not a unique cause or a mandate
The draft's troubleshooting use case says OAM tests can narrow a problem and help locate candidate root causes. That wording is careful. RFC 9940 separates event, fault, problem, symptom, cause, alert, alarm and incident; causes may require several inputs. RFC 8632 likewise says raw alarms do not necessarily identify service status or root cause, and its root-cause-resource entries are hints for a client application.
The diagnostic receipt should record the inference, not just the winning label: observations used, competing explanations, excluded hypotheses, confidence, scope and the evidence that would falsify it. A sequence that detects loss after one hop and not before it may narrow the search. It does not automatically distinguish congestion, queue policy, optical impairment, software error, transient reroute or measurement artifact.
Even a strong diagnosis does not grant change authority. The account allowed to create OAM schedules may not be entitled to alter routing, QoS, interfaces or customer service. RFC 8341 makes access control explicit for YANG operations. The change receipt must separately identify the authorized owner, exact proposed change, policy basis, preconditions, blast radius, maintenance window, rollback and approval.
Receipt eight: the service after the change
A controller can accept a change, a device can expose new operational state and the original test can turn green while the customer-visible problem persists. The change may have shifted traffic to a worse path, hidden one symptom, created another failure class or recovered only the synthetic probe.
The outcome receipt must come from the surface capable of observing the promised result. It should identify the post-change service population, test window, customer or SLA criteria, independent signals, residual alarms, rollback status and duration of stability. Re-running the same probe is useful, but it is not independent when the probe's comparability was the disputed premise.
This is where scheduled OAM becomes genuinely powerful. The system does not need one omniscient verdict. It needs narrow records that can be joined without being confused: plan, application, occurrence, steps, measurement, comparability, diagnosis, authority and outcome. The draft gives the first five a common orchestration vocabulary. Operational credibility comes from refusing to pretend it has already supplied the last four.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
