Summary

  • draft-nirvanai-nbtp-behavioral-trust-01 proposes a verifier-local trust ledger in which missed heartbeats, skipped oracle windows and co-silence can reduce a score and move an agent into SUSPECT or QUARANTINED.
  • A signed behavioral vector proves who issued an observation; it does not prove calibration, independent coverage, malicious intent or that silence originated at the agent rather than the observer, path, clock, privacy rule or maintenance process.
  • A defensible decision receipt must retain the expected emission, observer and path, oracle and context sets, timing window, maintenance and privacy state, raw attestation, calibration and policy versions, score transition, reason and recovery condition.

At 14:03 the agent is TRUSTED. At 14:08 its heartbeat is absent from the verifier's record. Another context appears quiet during the same window. The local ledger applies a penalty, decay accelerates and the state becomes SUSPECT.

Every arithmetic step may be correct. The verifier still has not shown why the packets were absent.

The agent might have stopped. It might also have crossed a failed collector, a partition, a shared proxy, a broken clock, an undeclared maintenance interval or a privacy boundary that deliberately omitted one context. Two scanners can agree because both observed reality, or because both depended on the same blind spot. The clean state transition is therefore the beginning of an investigation, not its conclusion.

Revision 01 of the Network Behavioral Trust Protocol draft makes this boundary unusually visible. On 2 October 2026 the Datatracker listed it as an active individual Internet-Draft, updated the previous day. The text is dated 18 September, says it is intended for the Experimental track and expires on 22 March 2027. It has no working-group state, RFC stream, responsible area director or telechat date. The Datatracker explicitly warns that an individual submission is not endorsed by the IETF and has no formal standing in its standards process.

That status matters. The proposal deserves analysis as an architecture, not promotion as adopted infrastructure. Its implementation section reports a NirvanAI scanner with public endpoints and deterministic matching across 102 patterns in 14 categories. RFC 7942 makes such disclosures useful, but author-reported implementation status is not an independent result for deployment, interoperability, reliability or calibration.

A local score is a chain of governed choices

NBTP is presented as an addition to identity, credentials and authorization, not a replacement. Oracle scanners sign observations. A verifier validates the attestations and computes a scalar trust value T between zero and one inside a temporally weighted Volatile Ledger. The ledger is local; it is neither a blockchain nor merely session memory.

The reference behavioral vector carries coherence_drift, hallucination_density and alignment_friction. The state machine labels the result PROBATIONARY, TRUSTED, SUSPECT or QUARANTINED. TRUSTED requires at least 0.7 plus a ramp-up process. SUSPECT covers 0.4 to below 0.7 and can also follow missing attestations, heartbeat absence, drift or co-silence. QUARANTINED covers a score below 0.4 and can also follow multiple violations, confirmed co-silence, explicit revocation or a stated Low/Low condition.

None of those thresholds is discovered by the signature. Nor are the decay constant, registered oracle set, context membership, grace period, correlation window, anomaly switch or recovery gate. They are configured policy. The score is therefore not a portable fact about the agent. It is the output of one verifier's evidence and choices at one time.

This local design has a legitimate advantage: no central online service must authorize every interaction. It also creates legitimate divergence. Two conformant verifiers can register different oracles, see different routes, classify contexts differently, keep different ledgers and choose different parameters. One can call an agent TRUSTED while another calls it SUSPECT without either miscomputing the draft's formula.

That is not a flaw to conceal. It is a property to govern. A useful verdict should say “SUSPECT under policy P, ledger epoch E and observation surface O,” not simply “the agent is suspicious.”

Silence belongs first to an observer

While active, an agent is expected to send a heartbeat at least every 120 seconds. More than 300 seconds without one causes immediate SUSPECT, without the ordinary grace for missing oracle attestations. For missed oracle windows, the default is two windows of grace, followed by a penalty and doubled decay.

Co-silence reaches further. The reference rule correlates absence across two contexts within a 300-second window and applies a default penalty factor of 0.6. If the reduced score crosses the threshold, the agent can be quarantined. The proposal calls this a strong signal with limited innocent explanations.

Yet the same text acknowledges the measurement limit. Version 0.5 cannot cryptographically guarantee activity, and minimal liveness traffic can suppress detection. A Context Activity Vector is self-reported and never proof. Server-side last_seen observations across packet types are more useful, but still belong to the server's visibility.

The evidentiary sentence must therefore keep its nouns. “Observer A did not record expected heartbeat H from agent X over path family R in context C between clocks t1 and t2.” Removing the observer, path, context and clock turns bounded evidence into a claim about character.

Independence cannot be counted by signatures alone. Three oracle identities behind one collector, model, cloud region or transit failure do not supply three independent observations. Majority agreement can detect an outlier; it cannot prove that the majority lacks a common cause.

The exception states expose the real model

The proposal already contains several reasons not to interpret absence literally. A maintenance notice pauses the liveness penalty for up to 3,600 seconds by default, though it does not waive oracle scans. If every registered oracle is unreachable, the verifier halves the decay rate and pauses the skip-grace counter; an outage longer than 600 seconds raises an alert.

These exceptions are not administrative decoration. They are competing explanations for the same missing event. The maintenance register, oracle reachability test and blackout clock participate directly in the trust verdict. If they are lost after the score is calculated, an auditor sees punishment without the evidence that justified—or should have prevented—it.

Privacy creates a harder version of the problem. A Context Activity Vector may omit privacy-sensitive contexts, accepting a possible co-silence inference. Scan content is exposed to an oracle and should be minimized, not retained or shared. The local ledger itself is not disclosed, and no mandatory retention extends beyond the active decay window, whose default is 600 seconds.

Privacy-preserving absence, transport failure and adversarial silence can therefore have the same surface shape. The answer is not unlimited surveillance. It is a typed omission: which context was withheld, under whose rule, for how long, and what inference was explicitly forbidden. Without that receipt, the system rewards visibility by penalizing the actor that protects information.

A signed vector still needs calibration

The draft requires oracles to be independent and to publish methodology. Any registered oracle's attestation is valid; for high-stakes decisions above 0.8 it says a verifier SHOULD require at least two independent oracles. Cross-verification runs every 300 seconds, and an oracle more than 0.05 away from the majority is flagged and excluded pending manual review. Fewer than three unique oracles creates an eclipse condition and caps trust at 0.6.

These controls govern issuer diversity and disagreement. They do not supply a scientific baseline. The draft itself says the metrics lack empirical calibration, should be treated as deviation indicators rather than threat verdicts, and must not alone ground revocation or denial of service.

That distinction should survive every dashboard. A signature establishes provenance and integrity for the vector. It does not establish that “hallucination density” was measured consistently, that a five-point change means the same thing across tasks, that the method is unbiased, or that the observation caused a security risk. Public methodology helps scrutiny; the concrete heuristics are still outside the protocol's scope.

The reported implementation sharpens the question. Deterministic matching can be reproducible without being valid for a new domain. A list of 102 patterns says something about the scanner's surface, not about false-positive rates, task mix, linguistic coverage or causal power. Until independent calibration exists, the vector should remain one input to a local decision whose limitations are visible.

One score can carry authority into the wrong context

NBTP keeps one trust value per agent across contexts. The draft recognizes a trust-parking problem: high trust earned in a slow-decay context can migrate into a high-frequency context. A formal context-weighted effective score is deferred, and the text says a high global score is insufficient authorization in a context where the agent has not been observed.

This is more than a mathematical edge case. Context membership decides which history follows an agent into a new operation. If the registry says a customer-support bot and a payment executor are one trust subject, quiet success in the first environment can subsidize admission to the second. If it splits them too aggressively, the same agent repeatedly loses legitimate history.

Whoever defines contexts, their decay rates and their mappings exercises admission power. The scalar must never become a credential that escapes the observation conditions that produced it. Every consuming system should receive the relevant context, evidence age, coverage and policy identity—or make its own authorization decision without pretending that T is universal.

Recovery is another authority surface

New agents enter through a Creole challenge-response exchange. The draft gives an initial score of 0.5, or 0.65 for a human-backed agent. Its reference Genesis Attestor is NirvanAI, and the attestation weight depends on that attestor's own network reputation.

PROBATIONARY agents face a two-hour ramp: five heartbeats, three challenge cycles and observations from three distinct TRUSTED oracle scanners. Their score is capped, decay is doubled and contribution weight is halved. A quarantined agent must rebootstrap through Creole and return to probation.

The protocol has not eliminated a gatekeeper. It has made the gatekeeper one node in a reputation system. Leadership should ask who can register or replace the Genesis Attestor, how its own reputation is established, what happens during its outage, how false quarantine is appealed, and whether a successor can interpret the old ledger without rewriting history.

Heartbeat-only absence has an easier return path: a monotonic sequence can bridge recovery immediately. Other SUSPECT causes require sustained clean attestations. That difference is sensible only if the cause code is retained. A state label without the transition evidence cannot choose the right recovery rule safely.

The trust receipt must outlive the number

A defensible receipt begins with the expected event: heartbeat or oracle window, protocol version, agent identifier, context and cadence. It records which observer looked, its key and methodology, the collection path, clock quality, network state and whether supposedly independent observers shared infrastructure.

It preserves the raw signed attestation or bounded absence statement, registered oracle set, context set, maintenance declaration, privacy omission, blackout state and calibration version. Then it records the score before and after, decay constant, update weight, grace counter, penalty factor, thresholds, anomaly state, policy version, state transition and reason.

Finally it names the decision that consumed the verdict, its local authority, the recovery requirement, any appeal or manual review, and the later observations that confirmed or reversed the inference. This is not an article-body dump or a universal reputation feed. It is enough typed evidence for another authorized operator to reproduce, contest and reconcile the transition.

The IPR record reinforces the need to keep layers separate. Disclosure 7238 names NirvanAI LLC, provisional US application 63/985,171 and sections covering attestation format, decay and cross-context correlation. It states a royalty-free RAND position for necessary core-protocol claims while excluding implementation-specific measurement heuristics. That is a disclosure, not evidence of patent validity, essentiality, implementation freedom, adoption or measurement quality.