Summary

  • draft-parsons-opsawg-security-operations-02 asks protocol designers to document retained, lost and new observable artefacts, as well as detectable error conditions. Revision 02 also expands its tooling account into collection, detection, investigation and response.
  • A structured event is only the first receipt. Safe action requires evidence of collection coverage, time and identity context, correlation, triage, investigation, decision authority, execution, service impact, recovery and learning.

The event that could not answer the next question

Imagine an implementation that does exactly what security guidance appears to request. It emits a machine-readable event when an abnormal state occurs. The event has a stable type, a timestamp, a severity and fields a parser understands. Transport succeeds. Storage succeeds. The SIEM shows a green ingestion indicator and a red alert.

The responder asks a different question: may I isolate the endpoint?

The event does not identify the current owner of the asset. It does not show whether an administrator made an approved change. The source clock is drifting, surrounding authentication logs are delayed and the topology record still points to yesterday's address. An alert rule has correlated the event with a threat-intelligence indicator, but the analyst cannot see which version of the indicator set or rule ran. The event is visible. The conclusion is not established, and the power to act has not followed automatically from the power to observe.

That gap is the useful centre of draft-parsons-opsawg-security-operations-02. Dated 9 September 2026, revision 02 remains an individual Internet-Draft. The Datatracker does not assign it an RFC stream or IETF consensus status; its header proposes Informational status. It should therefore be read as developing guidance, not as a deployed control framework or a verdict delivered by the IETF.

Its premise is nevertheless consequential. Protocol security is often assessed through cryptography, access control and resistance to obvious attack. Security operators inherit another problem after deployment: can they tell what happened, investigate it and respond without creating a second failure? The draft asks protocol designers to consider that operating need before observability has to be retrofitted.

Revision 02 draws a longer operational chain

The change from revision 01 to revision 02 matters because the tooling discussion becomes explicit. Collection, Detection, Investigation and Response appear as distinct subsections. The structure recognises that “the protocol produced a log” is not a complete security-operations story.

Collection gathers observations from endpoints, network infrastructure, applications, identity systems, asset inventories and other sources. Dynamic and ephemeral resources make the inventory problem harder: the object that produced an event may have disappeared or changed identity by the time an analyst sees it. Authentication and authorisation logs also require tamper protection, with privileged activity receiving particular scrutiny.

Detection is a later transformation. A SIEM can enrich records, compare them with baselines, apply rules and correlate multiple sources. That processing can turn observations into alerts, but it can also generate false positives and alert fatigue. A visible event does not certify that the right collection point saw all relevant traffic, that sampling preserved the decisive record, or that a detection rule interpreted it correctly.

Investigation adds an incident-management record, protocol dissectors, linked evidence and playbook steps. These tools give work a reproducible shape. They do not make the analyst's inference true merely because every required box was checked. A dissector has a version; a playbook embeds assumptions; a case can close while a compromised process continues.

Response moves from knowledge to intervention. The draft describes actions such as isolating an endpoint or blocking traffic and notes that SOAR systems may execute predefined playbooks without human intervention. Automation makes authority more important, not less. The ability to issue a command does not establish the right target, permitted scope, acceptable collateral effect or rollback condition.

The four headings are therefore not stages that collapse into one another. They are boundaries at which meaning, custody and authority change.

Protocol artefacts have bounded semantics

Revision 02 proposes a disciplined design question for new or changed protocols: which observable artefacts remain available, which familiar indicators disappear, and which new artefacts can support detection, investigation or response? It also asks designers to document errors and abnormal conditions that can be detected.

That is more useful than a generic instruction to “add logging.” It forces a comparison between old and new visibility. Encryption may hide an indicator that a middlebox once inspected. A new state machine may create a reliable error transition. An implementation may expose a counter that carries no stable identity across restarts. Each fact changes what defenders can reasonably infer.

But an artefact should not be asked to testify beyond its semantics. An error event may prove that one implementation entered a named state. It does not, by itself, prove attack, compromise, attribution, severity or the proper remedy. An Indicator of Compromise may be network-, endpoint- or behaviour-based and may help link incidents or drive blocking, yet RFC 9424's broader treatment reinforces why an indicator remains evidence to be contextualised rather than a universal verdict.

The first receipt should therefore name the exact protocol or draft revision, the implementation and build, the event definition, the producer, the deployment point and the observation time. If the artefact changed because the protocol changed, that change needs to remain visible downstream. Otherwise a parser can keep accepting a field whose operational meaning has moved.

Collection is a claim about coverage, not a pipe icon

A common architecture diagram draws a clean arrow from producer to SIEM. The arrow hides most of the collection problem.

Was the event enabled on every relevant instance? Did a restart reset the configuration? Was transport buffered, sampled or dropped? Which clock supplied the time, and what was its observed offset? Did a collector transform, redact or deduplicate the record? How long was it retained? Can the organisation prove that the queried interval was complete enough for the question being asked?

The draft stresses the need for a range of log sources because each offers a different view. That is not an argument for indiscriminate accumulation. It is an argument for preserving the source and coverage of each observation. Asset management may say what a device is. Authentication logs may say who established a session. Change management may say who was authorised to alter it. Network telemetry may show what traffic followed. None can silently substitute for the others.

Surrounding context is especially decisive when distinguishing an error or misconfiguration from hostile activity. “Configuration changed” is observable. “The authorised engineer changed the intended device within the approved window” depends on identity, asset, target, approval and time. “An intruder changed it” demands still more evidence. Joining those records without their clocks, provenance and uncertainty can make the resulting case look stronger than its parts.

A collection receipt should record source coverage, transport status, loss or sampling, transformation, clock condition, retention and access controls. Without it, “no matching event” can mean no activity, no sensor, no delivery, no retention or the wrong query. Silence is not one fact.

Standard form reduces friction, not uncertainty

The draft points to structured logging, including qlog, as a positive way to avoid fragmented private formats. The benefit is substantial. A common schema can reduce bespoke parsers, allow tools to preserve event identity and let investigators compare implementations without first reverse-engineering every log line.

Schema agreement does not create observation agreement. Two producers can emit valid qlog-shaped records with different clocks, retention policies, event coverage and implementation defects. A required field can be present but stale. A correlation identifier can be locally unique yet useless after aggregation. A secure channel can deliver a faithfully authenticated partial record.

Standardisation should therefore state its boundary honestly. It can define event names, field types, units and relationships. It may describe detectable errors. It does not establish a universal incident schema, an organisation's baseline, a retention period, a cross-system clock model, an incident verdict, response authority, rollback policy or recovery receipt. Those are later contracts.

The gain is still real: stable syntax frees operators to spend less effort decoding and more effort testing coverage and meaning. Problems begin when parser success is displayed as operational truth.

Correlation changes the claim

When a SIEM joins an event to asset inventory, identity logs, threat intelligence and network activity, it creates a new analytical object. That object needs its own receipt.

Which query ran? Which time window and join keys were used? Which baseline was active? What versions of the rule, enrichment feed and asset table were read? Which candidate records were excluded? If a threat-intelligence indicator had been withdrawn or its confidence reduced, did the correlation retain that state? If clocks disagreed, how was order reconstructed?

Those questions are not clerical. A rule can be syntactically valid and operationally inappropriate. A baseline can mistake a planned migration for anomalous behaviour. An alert can be important while still being inconclusive. False positives do more than waste time: repeated low-value alerts teach people to discount the channel through which a real compromise will later arrive.

The safe output of detection is not “attack confirmed.” It is a bounded claim such as: rule version R matched these records, under these assumptions, at this time, with this confidence and these unresolved alternatives. The alert becomes reviewable because its ingredients and transformation remain available.

Investigation needs interpreters and dissent

Protocol dissectors are often where a raw capture becomes intelligible. An incident-management system is where evidence, hypotheses, decisions and responsibilities can be assembled. A playbook can ensure that obvious checks are not skipped at three in the morning.

Each tool can also harden a mistaken interpretation. A dissector may not understand a new extension. A playbook may assume an asset class or organisational boundary that does not apply. A case-management field may force a binary severity before evidence warrants it. Automation may close a branch because an expected record was absent, even though collection was incomplete.

An investigation receipt should therefore preserve the dissector and tool versions, evidence links, queries, completed and skipped playbook steps, analyst or automation identity, competing explanations and confidence. It should distinguish observations from inferences and inferences from decisions. That separation makes dissent possible without erasing the common evidence base.

The draft distinguishes SOC security responsibilities from NOC performance responsibilities while treating security operations as coordinated whole-system work. The useful implication is not that the two teams must merge. They may read the same event differently because they hold different responsibilities. A NOC can establish a service-impact fact without deciding attribution. A SOC can identify a credible threat without possessing authority to alter a production route. Coordination works when those boundaries are explicit.

Authority begins where the alert ends

An incident verdict is still not a response command.

Isolation and blocking change a running system. They can stop an attacker, interrupt a customer, destroy volatile evidence, sever an investigator's access or move traffic onto a less observable path. The decision therefore needs a named authority, target, scope, duration, guardrails, dependencies and rollback path. Emergency authority may be designed for speed; it should not be invisible.

SOAR makes the distinction sharper. A platform may be technically able to disable an account or push a network block as soon as a rule fires. The relevant question is which principal delegated that capability, for which incidents, under which confidence threshold, with which exclusions and how revocation works. “The playbook ran” describes execution. It does not prove that the delegation was valid or that the operational effect matched intent.

The response receipt should bind the incident verdict to the approving human or policy, the command, exact target, expected effect, protected services, timeout and rollback condition. The executor then needs to acknowledge what it actually accepted. A controller's 200 response and an endpoint's isolation are different events.

Recovery outranks closure

Security operations often end in a case state: contained, remediated, closed. Running systems do not read case labels.

Observed recovery asks whether the malicious or abnormal behaviour ceased, whether the intended service returned, whether collateral damage remained, whether compensating controls took effect and whether the system stayed healthy after the rollback or restoration. It may require network observations, endpoint state, customer reports and fresh authentication evidence. A green ticket is useful workflow metadata; it is not the final witness.

Post-incident analysis completes the chain only when it can update the parts that failed: protocol guidance, event semantics, collection coverage, correlation rules, playbooks, delegations and recovery tests. If the lesson is compressed into “improve monitoring,” the organisation loses the location of the failure.

The hierarchy of evidence is simple. A documented intention is weaker than an accepted command. An accepted command is weaker than observed state. A closed case is weaker than a stable recovered service. Running evidence does not make governance irrelevant; it tells governance whether its decision entered reality.

Privacy is not an afterthought to visibility

Security-operations data can expose identities, behaviour, topology, privileged actions and content-adjacent metadata. Revision 02 consequently calls for segregation, secure storage, controlled access and auditing of tools and actions. RFC 6973 provides a wider vocabulary for assessing privacy threats, while operational guidance from NCSC and CISA illustrates why useful forensic and protective-monitoring evidence often spans multiple systems.

There is no universal maximisation rule. More data can improve reconstruction and increase harm if misused, breached or retained without purpose. Encryption can protect users while removing an observation previously used by defenders. Redaction can reduce exposure and also break correlation. The correct balance depends on the threat model, role, jurisdiction, system and decision being supported.

The defensible design records the choice. Which data is collected, for what purpose, at which granularity, for how long, accessible to whom and audited how? Which indicators disappeared under a privacy or encryption change, and what safer artefacts replaced them? A hidden trade-off produces both poor security and unaccountable surveillance.

The artefact-to-recovery receipt

A reviewable response chain should preserve at least ten separable records:

  1. the exact protocol, draft or RFC revision and the semantics of the artefact or detectable error;
  2. producer identity, implementation build, deployment point and observation time;
  3. collection coverage, transport, loss, sampling, clock condition, transformation and retention;
  4. asset, account, authorisation, topology and change-owner context;
  5. the correlation query, baseline, rule and enrichment versions that produced or suppressed an alert;
  6. triage priority, false-positive treatment and the analyst or automation identity;
  7. investigation record, dissector version, evidence links, alternative explanations and playbook steps;
  8. incident verdict, uncertainty and the authority permitted to decide or act;
  9. response command, target, scope, approval, guardrails, dependencies and rollback path; and
  10. execution acknowledgement, observed containment, service impact, recovery evidence and post-incident changes.

The chain is deliberately longer than a dashboard. It lets each participant make a bounded statement without pretending to own the next layer. Protocol designers can describe what becomes observable. Implementers can show what they emit. Operators can prove what they collected. Analysts can expose how they inferred. Governors can delegate action. Running systems can answer whether the intervention worked.

An event deserves to be called useful when it helps this chain advance. It does not deserve to rule the chain merely because it was the first thing the organisation could see.

Sources