Summary

  • On 5 September, revision 01 of the individual Testimony Record Internet-Draft replaced an unsupported survey sentence, specified digest computation and added an explicit implementation-status boundary.
  • Its four cumulative levels move from a parseable record to explained beliefs, gated consequential actions and verifiable integrity, but a non-acting emitter can satisfy the gating level by declaring acts:false.
  • The revision distinguishes checks a reader can settle from the record from assertions the emitter merely makes, such as where a risk class or approver identity originated.
  • An RFC 3161 anchor can bind a digest to a time and outside signer. It does not establish that the testimony is accurate, complete or the only record the emitter produced.
  • Buyers, auditors and regulators should require a scope-aware proof matrix—per-level results, verified and attested checks, integrity scheme and validator identity—not a bare TR-4 badge.

The correction is part of the evidence

The IETF announcement records a modest institutional event: an 18-page Internet-Draft became available on 5 September. The Datatracker page draws the boundary more sharply. This is an active individual draft. Anyone may submit one. It has no RFC stream, responsible Area Director or formal standing, and it is not endorsed by the IETF.

Within that boundary, revision 01 deserves attention because it makes its own failures visible. The fixed -00 text had said that none of eight surveyed systems recorded who approved an action. The fixed -01 text says the census did not support that statement. Four of the five relevant systems not written by the draft’s author were assessed absent; the fifth was undetermined. The other three included the author’s own implementation. The stronger sentence was withdrawn.

That is not a cosmetic erratum. A format intended to preserve what a machine claimed to know is being tested by the same rule. “Not found,” “could not be established” and “present in the author’s own system” are different evidence states. Revision 01 stopped compressing them into a universal absence claim.

The official diff shows a second correction. Revision 00 named a digest but did not specify the algorithm, serialization or ordering needed to reproduce it. Two implementations could accept the same entries and calculate different values. A validator could also accept a declared digest without recomputing it. The level called Verifiable was therefore not reachable from the document alone in the ordinary meaning of that word.

Revision 01 now requires SHA-256 over a defined ordered set of entries. It specifies compact JSON serialization, member ordering by Unicode code point, omission of underscore-prefixed reader annotations, UTF-8 encoding, LF joins without a trailing LF and bounds for numeric values that otherwise serialize inconsistently. The reference validator recomputes the result. This is the kind of repair that makes a claim locally contestable.

Four levels answer more than one question

The format records scope, belief, evidence, conflict, decision, approval and integrity entries. TR-1 asks whether the sequence parses, uses known types, carries required members, orders write times and avoids duplicate identifiers. TR-2 asks whether beliefs name their evidence—even an empty array for an ungrounded belief—and whether both sides of a contradiction survive along with any later resolution.

TR-3 moves to consequential action. Risk classification must be attributed to a source outside the proposing model’s output. A refused action cannot also be recorded as executed. An executed high-risk action needs an approval naming a human, an identity source outside model text and an approver other than the proposer.

TR-4 then asks about record integrity. Covered entries must exist; their digest must recompute; and an external anchor must bind the declared imprint. These are valuable properties. They are not one property.

The draft says this itself: cumulative ordering conflates whether a system gates actions with whether its record can be shown unaltered. The problem becomes concrete in version 0.2. An emitter can declare acts:false. If it then contains no decisions, it satisfies TR-3 by having no action to gate and may reach TR-4. The result may properly be reported as “TR-4, record only.” Dropping the last two words turns a scope-qualified statement into a general assurance label.

This is not a defect that can be fixed by forbidding non-acting systems. A memory store or observer may need strong, tamper-evident records without possessing any actuation power. Its integrity evidence should be visible. The mistake would be to let that evidence imply that consequential actions were subjected to effective gates when no such actions were in scope.

Scope is still testimony

The scope entry resolves a real validator problem, but it is self-declared. A system that actually acts can say it does not and skip the decision requirements. If its record contains a decision, the contradiction is visible and TR-1 fails. If it simply omits the action path, the record alone cannot prove the declaration false.

That limitation matters in procurement. A vendor may expose a “recording” component for conformance while consequential calls occur in a connector, sidecar or host outside the emitter’s declared boundary. The record can be internally impeccable and operationally incomplete. A buyer must therefore identify the executable boundary separately: which process can propose, authorize, dispatch and observe an external effect; which paths bypass the recorder; and what independent evidence connects that boundary to the declared scope.

The author’s public repository makes the specification, validator, adapters and conformance corpus available. That is stronger than a prose-only promise because outsiders can inspect and run them. It remains author-controlled evidence. The draft’s new Implementation Status section says every implementation known at the cutoff was by the same author. Two separately written validators can expose ambiguity and test internal consistency; they do not demonstrate independent interpretation by another implementer.

The census directory is similarly useful because it pins assessed commits, requires evidence for absence and discloses conflicts. It also documents defects found in the author’s own work. It is not an independent audit merely because it examines third-party frameworks. Its author defines the rubric, writes the reference implementation and performs the assessment.

A level contains verified and attested checks

Revision 01 introduces a distinction more important than the ordinal. A verified check can be settled from the record: a cited evidence entry is present; both sides of a conflict remain; a refusal is not marked executed; a digest matches the exact entries it covers; an outside timestamp token signs that digest rather than another one.

An attested check is a statement the record makes that the reader cannot settle from the record. The named risk source may not exist. An approver name may not have come from the claimed authenticated session. A replay engine may not reproduce the claimed behavior. Requiring those fields is still useful: a specific assertion can later be contradicted, while an omitted field cannot. But an assertion does not become evidence because a schema requires it.

The same TR level can therefore rest on different mixtures. One record may pass mostly through arithmetic over retained bytes. Another may reach the same label with key facts supplied by the emitter. A headline number erases that composition.

Integrity schemes differ too. A digest computed by the emitter can detect a third party’s later change but cannot stop the emitter from creating a different record and recomputing. A signature places identity behind a statement, subject to key custody and trust policy. An external RFC 3161 timestamp places the imprint and time claim outside the emitter. RFC 3161 requires the requester to verify that the token carries the requested imprint and acceptable policy. The Testimony Record draft correctly limits the conclusion: the token fixes bytes and time, not accuracy, completeness or the absence of another discarded record.

RFC 8785 supplies useful JSON canonicalization context, while the draft makes its own covered-entry rule explicit. RFC 9162 and the IETF working draft on the SCITT architecture show how evidence can be placed in transparency infrastructure beyond one emitter. Neither can retroactively establish that a self-declared scope was complete or that a belief was true.

Replace the badge with a proof matrix

The revised draft already recommends reporting each level’s own result and the declared scope. Consequential users should make that the minimum, not optional fine print.

A compact proof matrix can name the exact specification version and hash; the emitter and executable boundary; declared scope and whether it is self-declared or externally evidenced; the pass, fail or scope-qualified result for every TR level; the identifiers and counts of mechanically verified checks; the identifiers and counts of attested checks; the integrity scheme and covered entries; the anchor authority, policy and validation outcome; the validator build; the conformance-corpus version; and the observation time.

Such a matrix preserves strong partial evidence. A record may have valid external integrity while its approval-origin claim remains unverified. Another may demonstrate a real action gate while lacking durable anchoring. The matrix shows both facts instead of forcing the weaker property to erase the stronger or the stronger to decorate the weaker.

The format’s privacy boundary remains unresolved. It recommends retaining source identifiers and digests rather than duplicating sensitive content, and marking redaction visibly. It also acknowledges that append-only history and erasure duties pull in opposite directions. Rewriting history destroys evidentiary continuity; keeping deleted content under a tombstone does not erase it. A conformance badge cannot decide that legal and architectural choice.

Nor does the draft implement the EU AI Act merely because it discusses record-keeping and human oversight. It expressly disclaims that conclusion. Legal sufficiency belongs to the parties under obligation and the competent regulator.

Revision 01 is therefore a credible act of self-correction and a better technical object than revision 00. The durable governance lesson is not to turn that progress into a larger mandate. A record can show exactly what it contains, expose contradictions and bind its bytes. The person relying on it must still ask which parts were observed, which were merely asserted, what the emitter left outside its scope and who is authorized to accept the remaining uncertainty.

Sources

  1. IETF announcement of Testimony Record revision 01
  2. IETF Datatracker status
  3. Testimony Record revision 01
  4. Testimony Record revision 00
  5. Official revision 00–01 diff
  6. Machine Testimony repository
  7. Conformance census directory
  8. RFC 3161 — Time-Stamp Protocol
  9. RFC 8785 — JSON Canonicalization Scheme
  10. RFC 9162 — Certificate Transparency Version 2.0
  11. SCITT architecture revision 22
  12. Heng Lu on running-code primacy