Summary

  • draft-feng-netconf-naim-op-00 models compensation as one or more Operation IR objects intended to reverse or mitigate earlier changes. Purpose is not outcome: the draft separately anticipates compensation failure, timeout and loss of connectivity during rollback.
  • A shared transaction identifier correlates operations but does not specify atomicity, isolation, ordering, locking or a single restoration point. Compensation must face its own authorization, validation, logging and live-state checks.
  • A credible restoration claim needs the exact forward and compensating operations, before-state and concurrency evidence, acknowledgements, post-state observations, independent service validation and a ledger of effects that cannot be undone.

A recovery plan is executed in a changed world

Consider a routine request: move a group of edge interfaces to a new routing policy, verify reachability, and undo the move if the check fails. An AI agent can express the intent, a Handler can validate its fields, and a transaction identifier can keep the related operations together. The recovery object may say: restore the old policy reference.

The test fails after three of ten targets accept the change. During those seconds, another controller updates one of the same interfaces. A remote system consumes a notification. A route flap reaches a customer. The operator's credentials still authorize the original update, but the emergency role used for recovery has a narrower scope. Then the management path drops.

What does “rollback” mean now? It cannot mean time travel. The old policy reference may be written back on two devices and rejected on the third. The external notification cannot be unreceived. The concurrent edit may be legitimate and must not be erased. The service may stay impaired even if the modeled configuration again resembles its starting value.

This is the boundary exposed by draft-feng-netconf-naim-op-00. Its Compensation Operation is “intended to reverse or mitigate” an earlier operation. The verbs matter. An intention can be precise without being fulfilled; mitigation can reduce harm without recreating an earlier state.

What Operation IR adds—and what it does not

Operation IR is proposed as a protocol-neutral object between natural-language intent and NETCONF, RESTCONF or another management backend. It can represent writes, retrievals, RPCs or actions, filters, datastore selection, preconditions, expressions, transaction metadata and compensation. The proposed division of labour is sensible: the AI is an intent encoder, while a deterministic Handler validates the object, checks live preconditions, produces protocol messages and executes them.

The draft is also young. Datatracker lists revision 00, dated 18 July 2026, as an active individual Internet-Draft with no RFC stream and no formal Intended RFC status. It says the document is not endorsed by the IETF and has no formal standing. The submitted header says “NETCONF Working Group” and “Intended status: Standards Track”; those are author-supplied header claims, not evidence of working-group adoption or IETF consensus.

Reading the proposal at its actual maturity makes its honesty more valuable. The draft does not claim to standardize automatic compensation-derivation algorithms, scheduling internals or private execution logic. It provides a place to carry compensation intent without pretending that every system will infer the same inverse.

That restraint leaves an obligation for implementers. If the Handler generated the recovery steps, the execution record must preserve the exact objects it generated. “Automatic compensation triggered” is not enough to reconstruct which targets, values, datastores, credentials or assumptions were used.

A transaction identifier is not an atomicity guarantee

Section 12 says multiple Operation IR objects may be associated by a shared transaction identifier. Association is valuable. It allows a request, its constituent operations and its compensations to be correlated in logs.

But the draft does not say the identifier creates an all-or-nothing commit. It does not specify serializable isolation, a global lock, a durable transaction log, a mandated order, or one point to which all heterogeneous targets can return. The identifier names a relationship between operations; it does not supply the execution semantics that a database transaction would need.

This distinction becomes decisive when one Operation IR group crosses devices or protocols. One target may support candidate configuration and confirmed commit. Another may write directly to running configuration. A third action may invoke an RPC with an external side effect. A single label can correlate all three. It cannot make their clocks, locks, capabilities or failure modes identical.

The same reasoning applies to order. If an operation creates an object and the next operation assigns a reference to it, reversing the second before the first may be necessary. If another actor has since attached a legitimate reference, deleting the object may now cause fresh damage. A syntactically exact inverse can be operationally wrong because the world against which it was derived no longer exists.

Compensation is a new claim on authority

The draft says compensation operations should be subject to the same authorization, validation and logging expectations as normal operations. Its security section separately calls out authorization for compensation and audit logging of execution and rollback.

That is not ceremony. A forward operation and its compensation can touch different paths, objects or privileges. Creating a temporary policy may be permitted while deleting it is not. Changing one interface may be within scope while reverting a group is not. A principal authorized at 10:00 may have lost the role at 10:05. A dynamic reference may resolve to a different target when compensation runs.

Therefore, authority must be evaluated at the recovery boundary rather than copied from the original decision. The Handler needs the principal, current policy, resolved target and decision result for every compensation step. If authorization is denied, the correct outcome is not to bypass control in the name of safety. It is to record incomplete recovery, contain the blast radius and escalate to an authorized path.

Validation is equally temporal. The draft allows precondition_state to declare values that must hold before an operation executes. A compensation object can use preconditions to avoid overwriting a state that has legitimately changed. Yet a passed precondition is only a gate at a particular observation time. It does not prove that the subsequent write landed, remained in place or restored service.

The draft treats rollback failure as a real state

The most important lines may be the least glamorous. Implementations that support transactions are expected to define behaviour for precondition failure, failure in the middle of a transaction, compensation failure, timeout during grouped execution and loss of connectivity during rollback.

Those cases establish the right mental model. Recovery is not a magic exception to the failure model. It is another distributed operation that can be refused, interrupted or only partly observed. The operator must know whether silence means “not attempted,” “request sent,” “applied but acknowledgement lost,” “partly applied,” or “compensation itself failed.” Retrying each state blindly can duplicate side effects or erase later changes.

The execution report therefore needs a state machine richer than success/failure. At minimum, each step should distinguish planned, authorized, precondition-checked, dispatched, acknowledged, independently observed and service-validated. Recovery should add an explicit residual state: what remains different from the accepted before-state, and who owns the next decision.

NETCONF's narrower rollback proves the need for scope

RFC 6241 provides a useful comparator. When a server supports rollback-on-error, an edit-config can stop after an error and restore the specified configuration to its state at the start of that edit. That is a narrower, protocol-defined promise with a declared capability.

Even there, the RFC warns about shared configuration. Rollback can inadvertently alter or remove changes made by other NETCONF sessions unless configuration locking is used. NETCONF also has a rollback-failed error for a requested rollback or discard that could not be completed. Confirmed commit has its own capability and candidate-datastore dependency.

Operation IR compensation should not inherit these semantics by vocabulary. It may map one operation to NETCONF rollback-on-error, another to a compensating RESTCONF write, and a third to an application action. Whether the first target can restore its edit boundary says nothing about the other two. “Rollback” is a family of mechanisms, not a universal receipt.

Configuration restored is still not the past restored

RFC 8342 distinguishes intended configuration from operational state. That alone prevents a simple equality claim from doing too much work. A configuration subtree may match a prior snapshot while the device's applied state, learned state or dependent service remains different.

The gap is larger outside the datastore. Packets already forwarded cannot be recalled. A BGP withdrawal already observed by a peer is part of history. A notification already consumed may have triggered another controller. A credential exposed to a target is not made secret again by deleting its configuration. Timeouts, billing events, customer alarms and human decisions persist.

Compensation may be entirely successful and still be mitigation rather than restoration. That is not failure of the concept. It is why the draft's definition wisely includes both verbs. The error begins when an operations dashboard collapses “compensation completed” into “no impact” or “original state restored.”

Dry-run is a disclosure surface, not a prophecy

The draft's dry-run design reinforces the boundary. A Handler must not apply configuration in dry-run mode and should present targets, values, datastore, precondition checks, protocol summary, compensation plan, known side effects and limitations. It also says a preview may be incomplete when live-state dependencies, authorization decisions or external conditions cannot be fully verified without execution.

A good preview should therefore state its unknowns prominently. It should identify which compensation targets were resolved against a snapshot, which authorization checks are deferred, which external systems cannot be simulated and which side effects have no inverse. A green preview that hides those limitations is not safer because its formatting is structured.

The evidence chain leadership should demand

The common Operation IR object can remain small. The operational record around it cannot remain implicit. For each transaction, preserve the immutable forward intent and the semantic context version; the canonical model used by the Handler; authorization decisions; live precondition reads; exact generated protocol messages; an ordered account of acknowledgements and timeouts; the exact compensation objects; their independent authorization and validation; any locks or concurrent writes; post-compensation configuration and operational observations; external residual effects; and a service check from a vantage independent of the Handler.

The final status should not be rolled_back: true. It should say what was restored, what was only mitigated, what could not be observed, what changed concurrently and which consequence remains irreversible.

That is the practical meaning of keeping symbolic and operational reality separate. The compensation field is a valuable symbol. Running code, observed state and service outcomes determine what happened. A transaction identifier joins the record. It does not absolve anyone from proving the result.

Sources