Summary

  • draft-smith-opsawg-ai-network-governance-01 says constraints communicated to an AI service do not replace independent programmatic enforcement; the executing agent therefore needs a per-action receipt naming the exact policy and state it applied.
  • A compliant proposal, accepted management operation or attempted rollback is not proof of a safe outcome. Authority, mutation, readback and service effect remain separate claims.

The prompt says the interface is protected. The model proposes changing it anyway. The local agent rejects the proposal.

That is the reassuring demonstration. It is also the easy one.

The harder case begins when the model proposes an allowed action with an allowed parameter against an apparently ordinary target. Which Action Registry version classified the action? Which operator allow list was live? Had the protected-target set changed seconds earlier? Which rate window survived the last restart? Had a human already revoked autonomy? Did the executor validate the same limits summarized in the prompt, or an older local copy? And after the management RPC succeeded, did the service improve?

Those questions define the real control surface of AI-mediated network operations. A prompt can describe policy. Only the enforcement path can exercise authority.

draft-smith-opsawg-ai-network-governance-01, posted on 27 September 2026, recognizes that distinction explicitly. It is an active individual Internet-Draft, not an IETF-endorsed document, working-group consensus or RFC. Its page says it has no formal standing in the IETF standards process and no RFC stream. The header calls the intended status Informational. It should be read as a serious architecture proposal, not a deployed standard.

Its scope is deliberately narrow. All four conditions must hold: the system executes on, or has direct management-plane access to, a router or switch; it consults an external AI service; it can alter device configuration or operational state without per-action human approval; and it uses management interfaces such as NETCONF, RESTCONF or gNMI. Advisory systems requiring human execution are outside the scope.

The architecture separates the external AI service from the local autonomous agent. The service receives structured anomaly context and returns a proposal. The agent collects telemetry, detects the anomaly, consults the service, applies safety controls and reaches the device. The service must not receive direct device access.

This is good separation. It confines the probabilistic component to recommendation and keeps credentials and execution in a local enforcement boundary. But the local agent consequently becomes the privileged point where policy, state and device authority meet. Its decision deserves stronger evidence than the model's explanation.

The draft supplies two useful vocabularies. The Action Registry is a fixed, developer-defined catalogue whose entries define parameters, risk and reversibility. The Operator Allow List is the deployment-specific subset approved for autonomous execution. AI proposals outside the registry are discarded; the model cannot create new action types. Parameters outside validated ranges are rejected regardless of model confidence.

Section 12.5 then requires the prompt or context to summarize allowed and blocked actions, protected targets, rate limits and the risk ceiling. Section 16 explains why: the AI can avoid obviously inadmissible proposals and can explain which constraints led it to escalate.

Then comes the sentence on which the entire architecture depends: communicating constraints does not replace programmatic enforcement. Every proposal must still be independently checked against governance rules before execution.

That sentence prevents prompt engineering from becoming security theatre. It also creates an audit obligation. If prompt policy and enforcement policy are separate, a post-incident reviewer needs to know both—and must be able to prove which one controlled the device operation.

A hash of the prompt is insufficient. The prompt is a projection assembled for an external service. It may omit sensitive detail, simplify a target pattern, describe a rate budget at query time or lag a configuration reload. Conversely, a bare “guardrail passed” event is insufficient because it does not identify the inputs that made the decision true.

The minimum action receipt should bind the human-approved policy snapshot; Action Registry and allow/block list versions; protected-target rules; parameter bounds; risk ceiling; rate-counter state; degradation state; revocation state; proposal and parser result; guardrail verdict; point-of-execution validation; authenticated device operation; captured pre-state; device reply; fresh post-check; and rollback, hold or escalation outcome. Each value needs a stable identity and time relation. Otherwise the organization can show that rules existed without showing that these rules governed this action.

The draft itself makes state unavoidable. It proposes limits for actions per hour, actions per target per day, irreversible actions, AI queries and retries. Appendix B gives defaults and maxima—five and twenty remediation actions per hour, for example, and three and five actions per target per day. Those numbers are proposals, not universal safety constants. The frozen sources provide no empirical derivation across network sizes, vendors or failure domains.

More importantly, limits are historical claims. An action is “the fifth this hour” only if the system has a reliable clock, a defined window, complete prior events and continuity across failover. Revocation, cooldown, retry history and degradation state are historical too.

Section 17.3 says checks should be stateless where possible, drawing inputs from a persistent audit trail rather than volatile memory so restarts do not reset counters. That is a sound implementation direction. It does not eliminate state. It relocates state authority into the log and the projection that reads it. Completeness, ordering, retention, clock semantics, policy activation and replica convergence then become safety properties.

The proposed fail-safe examples reinforce the point. If rate data cannot be queried, assume the limit is exhausted. If protected-target matching fails, assume the target is protected. If the registry cannot be consulted, assume the action is unregistered. Unknown policy state should block execution, not silently collapse into permission.

Device mutation introduces a second receipt boundary. Before acting, the agent must capture enough state to restore the target. After acting, it must gather fresh telemetry, re-run detection and roll back if the condition fails to improve or worsens. Cached or duplicate-suppressed observations cannot establish recovery.

Yet pre-state is not a time machine. A NETCONF configuration snapshot can record the relevant tree before a change. It cannot by itself restore a withdrawn route learned elsewhere, an exhausted queue, a peer's timer, packets already dropped or application sessions already broken. Rollback is another forward operation through a changing system. It must earn its own device reply, readback and service observation.

This is also why “single target” must not be mistaken for single impact. One interface, metric or protocol instance may sit inside a shared failure domain. The proposal can be syntactically local while the consequence propagates through convergence, traffic placement and remote dependencies.

The prompt-injection section is appropriately cautious. Device-generated logs, interface descriptions or other strings may influence the external service. Sanitization helps, but the architectural defence is that the model remains advisory and every output crosses independent controls. Testing therefore cannot stop at a screenshot showing the model refused a malicious instruction. It must exercise parser, registry lookup, target matching, parameter checks, rate history, execution path and failover path.

The draft calls the audit trail the primary accountability mechanism and recommends integrity protection plus possible external retention. That is where the policy receipt belongs. Startup configuration logs and a known-good governance-document hash are valuable, but incident reconstruction needs the effective policy at the exact action, not merely the configuration seen at process start.

The acknowledgement says the concepts derive from operational experience deploying such systems on production infrastructure. That is an author statement. It is not an independently documented deployment study, benchmark or safety result. The proposed numbers and controls should therefore become testable hypotheses in each operating environment, not inherited assurances.

Heng Lu's minimum-initial-specification principle sharpens the design. The shared minimum should contain invariants that can be verified locally: an AI cannot enlarge the action vocabulary; unknown policy state fails closed; the agent cannot rewrite its safety rules; and every mutation identifies the exact authority that admitted it. Local operators may choose tighter bounds, but they should not weaken the receipt needed to prove which rules ran.

The reality-layer distinction supplies the second correction. “Allowed” is a symbolic judgment inside the policy layer. An interface, routing adjacency, queue and delivered service occupy the operating layer. A valid symbol does not become the physical outcome merely because the device accepted an RPC.

Running-code primacy supplies the last test. The governance document may be elegant. The prompt may reproduce it perfectly. The guardrail may evaluate the intended version. None of those facts proves service health. The action becomes operationally true only when the device state and intended service effect are independently observed.

Sources