Summary
draft-abak-ai-evaluation-claim-preservation-00proposes requirements for carrying AI evaluation assertions across formats without silently changing their attribution, meaning, scope or material qualifications.- A successful run, a numerical score, an original criterion verdict and a consumer's later policy judgment are different records. Preserving one does not manufacture the others.
- A valid signature or transparent receipt can protect a lossy conversion. Later authentication identifies the protected view; it does not reconstruct context discarded at an earlier hop.
One number, four unanswered questions
An evaluation operator finishes a run. The log reports status: success and an accuracy value of 0.734. An exporter copies both into a compact evidence view. A reviewer later opens the view and sees an authenticated record that appears complete.
Revision 00 of Claim-Preserving Exchange of AI Evaluation Evidence asks the uncomfortable question: complete for which claim?
Process success can mean that the evaluation machinery finished without an operational error. The score can mean that one named metric produced 0.734 for one selected result. Neither fact says that the original evaluator applied a threshold. Neither establishes a pass or failure. If the source does not disclose whether a criterion existed, the honest result is not “no criterion.” It is “criterion not established by the available source.”
That distinction is easy to erase because dashboards prefer one status. A converter turns success into PASS; a consumer later applies a local threshold of 0.70; the exported record then reads as though the evaluator had declared a successful safety test before the run. Three legitimate facts have been collapsed into one false chronology.
The draft's proposed repair is not a universal AI assurance format. It is a set of preservation obligations for mappings that choose to adopt them. A mapping names its source and destination interpretations, converter version, profile revision and bounded claim set. It keeps source assertions, explicit derivations and later policy judgments separately attributed. If a required distinction cannot be represented, the stronger claim stays unsupported.
That is a modest design. It is also a direct challenge to the idea that an authenticated number can carry its meaning by itself.
A run identifier does not select a result
Evaluation outputs can contain several tasks, scorers and metrics under one run. The draft's synthetic LightEval-style example places two values under the same fictional task and run: exact match at 0.62 and majority-at-eight at 0.8. A record saying only “run-17 passed 0.75” has not selected either one. It has also failed to identify who supplied the threshold and what population the comparison covered.
A precise selector helps, but only within a precise source snapshot. JSON Pointer can identify one value in one object. It does not state the metric definition, the dataset split, the attempt, the aggregation rule or whether the referenced URL later changed. A pointer into mutable content is a coordinate without a preserved map.
Revision 00 therefore treats a result unit as more than a scalar. Task, scorer, metric, run attempt, aggregation, dataset or split and population context travel when they affect the claim. An ambiguous selector cannot be resolved by choosing the first, last, largest or most favorable value. If the profile cannot establish which result was meant, it has not preserved the requested claim.
This rule matters because apparently harmless export defaults carry authority. “Take the first metric” is not merely a parser choice when the first metric passes and the second fails. It is an undisclosed decision about which evidence gets to exist downstream.
The denominator is part of the evidence
The draft's clearest counterexample begins with a planned population of 100 samples. Eighty run. Seventy-six of those meet a sample criterion and four do not. The local fraction among completed samples is 76/80, or 0.95.
That is not evidence that 95 of the planned 100 samples passed. Twenty did not run. A summary that exports only 0.95 removes the difference between observed performance and campaign coverage. If another service later signs the summary and labels it “complete evaluation passed,” the signature can be perfectly valid while the assertion is stronger than the source.
Retries, duplicate records, invalidated samples and changing sample sets create the same problem. A complete population claim needs an identified population or reproducible inclusion rule and an accounting method. Unknown totals remain unknown. Selection and deduplication that change the denominator are transformations, not housekeeping.
Even preserving every record supplied by the operator does not prove that every real attempt was supplied. Claim preservation can carry a source assertion accurately; it cannot make an incomplete or dishonest source complete and honest.
A display value can reverse the result
The number itself can survive while its decision changes.
Revision 00 gives a synthetic decimal example: a source value of 0.94996 is displayed as 0.950, and the criterion is greater than or equal to 0.95. Comparing the original exact decimal produces failure. Comparing the rounded display produces pass.
Rounding is not forbidden. Concealing which value drove the comparison is. Units, scale, metric definition, comparator direction, precision and uncertainty are part of the claim when they affect interpretation. A consumer may deliberately adopt a rule that compares rounded values, but that is a newly attributed assessment. It cannot silently replace the source comparison.
The same discipline applies to normalization, recomputation and aggregation. Missing uncertainty is not zero uncertainty. A source estimate is not uncertainty introduced by conversion. An integer identifier is not expendable merely because a destination represents all numbers with a lossy type.
A reference is not the bytes behind it
Evidence services often carry links instead of logs. A publisher may receive a restricted URL for a run record but lack permission to retrieve it. The honest export can preserve the URL, the fact that the publisher did not retrieve it, the absence of locally recomputed digest knowledge and the source's assertion that an object exists.
It cannot insert an all-zero hash, claim that the bytes were checked, infer the contents, or promise that the source retained a complete log. Public location, permission, successful retrieval and future availability are four different properties.
The distinction persists when a digest exists. One actor may recompute a digest from bytes it received. Another may merely copy a digest reported by the source. Both can carry the same hexadecimal string, but they do not possess the same evidence. A content digest identifies selected bytes under a stated algorithm; it does not identify the semantic rules used to interpret those bytes.
This is also why an evidence reference is not permission to fetch or execute its target. Evaluation logs, prompts and tool outputs remain untrusted data even when embedded in a valid signed record. Retrieval destinations, redirects, schemes, decompression and resource use require independent policy.
Later authentication preserves the loss too
SCITT receipts, in-toto statements and other signed carriers can protect mapping records and their relationships. RATS offers a disciplined separation between evidence appraisal and relying-party policy. These are useful building blocks. None can infer a metric's missing meaning merely because the envelope verifies.
Revision 00 makes the multi-hop rule explicit. Every mapping hop used to support end-to-end preservation has to be accounted for, or a later actor must directly recheck an adequately identified earlier source. If the first converter dropped the population and exported only 0.95, a second converter cannot restore the old context by attaching a new statement. It can build a new derived view after checking the original, but it must disclose that recheck. It cannot pretend the context survived the earlier hop.
Likewise, a signed loss report is an assertion about loss, not proof that its inventory is complete. A valid carrier under an unsupported semantic profile may be stored or forwarded as opaque data. It should not be described as an established evaluation claim.
The verification vocabulary therefore has to stay typed. Structural validation, digest recomputation, signature verification, attestation appraisal, mapping verification and substantive evaluation judgment answer different questions. A single green badge that merges them is a semantic failure at the final interface, even if the raw attachment preserved every distinction.
Preservation is not truth, authority or effect
Accurately preserving a false, biased or selectively published assertion keeps it false, biased or selective. A compromised evaluator can sign false claims. A competent converter can faithfully carry them. Claim preservation does not certify benchmark quality, evaluator independence, model identity, scientific validity or deployment safety.
The decision boundary is separate again. A consumer may apply its own threshold and record a new assessment. A decision maker may rely on that assessment under a named policy. Neither record proves that the decision maker had authority, that a control reached its target, that enforcement occurred or that the intended external effect followed.
Revision 00 deliberately stops before those claims. It specifies requirements and synthetic test obligations, not a wire format or execution mechanism. Same-author tests are not independent interoperability. A schema check is not a semantic mapping test. A claim of compliance must identify the profile, implementation role and tested version; “claim-preserving” by itself is an empty adjective.
The reality-layer test
Heng Lu's reality-layer doctrine supplies the operational reading. A source byte sequence, a parsed result, a semantic claim, an authenticated carrier, a consumer assessment, an authorized decision, a delivered control and an observed outcome are related layers. No layer acquires the properties of the next merely because a dashboard aligns them in one row.
Minimum initial specification means defining the smallest bounded claim set that a mapping can preserve and leaving everything else explicitly unsupported. It does not mean reducing the record until only a score and a green label remain.
Running-code primacy then asks for observable behavior. Can a version-pinned producer and consumer distinguish missing criteria, two metrics under one run, incomplete populations, inaccessible evidence, conflicting revisions and threshold-changing rounding? Do negative cases remain visible? A schema-valid example and a documentation crosswalk are symbolic evidence. Interoperability begins when independent implementations produce and consume the distinctions under tested conditions.
The practical rule is severe but simple: never let the portability of a number exceed the portability of its meaning.
Sources and limits
- https://datatracker.ietf.org/doc/draft-abak-ai-evaluation-claim-preservation/
- https://datatracker.ietf.org/doc/draft-abak-ai-evaluation-claim-preservation/history/
- https://datatracker.ietf.org/doc/html/draft-abak-agent-control-delivery-evidence-01
- https://datatracker.ietf.org/doc/html/draft-nobuo-scitt-composite-evidence-verification-00
- https://datatracker.ietf.org/doc/html/draft-ozturk-scitt-prml-profile-00
- https://datatracker.ietf.org/doc/html/draft-watts-agent-evidence-boundary-00
- https://heng.lu/minimum-initial-specification-localized-future-decision-voluntary-adoption-internet-coordination-system/
- https://heng.lu/on-reality-layers-symbolic-power-and-why-clarity-feels-so-hostile/
- https://heng.lu/running-code-primary-the-patch-needed-to-preserve-the-internet-original-design/
- https://github.com/huggingface/lighteval/blob/main/docs/source/saving-and-reading-results.mdx
- https://inspect.aisi.org.uk/eval-logs.html
- https://github.com/in-toto/attestation/blob/main/spec/v1/statement.md
- https://slsa.dev/spec/v1.1/verification_summary
- https://www.ietf.org/archive/id/draft-abak-ai-evaluation-claim-preservation-00.txt
- https://www.rfc-editor.org/rfc/rfc2119.txt
- https://www.rfc-editor.org/rfc/rfc6901.txt
- https://www.rfc-editor.org/rfc/rfc8174.txt
- https://www.rfc-editor.org/rfc/rfc8259.txt
- https://www.rfc-editor.org/rfc/rfc9334.txt
- https://www.rfc-editor.org/rfc/rfc9943.txt
The evidence was frozen on 30 September 2026 Asia/Shanghai. Revision 00 is an active individual Internet-Draft with an Informational target. It is not an RFC, IETF consensus, Working Group product, deployed protocol, implementation report, benchmark result, safety certification or measured incident. The document claims no implementation or interoperability result, and its examples are synthetic. Interpretations applying Heng Lu's doctrine are Daniel Kade's analysis, not draft text.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

