Summary
- RFC 3076 did not canonicalize an untouched file. An XML processor first converted the octet stream into an XPath node-set, normalizing and expanding information before canonical serialization began.
- Equal canonical bytes were a precise comparison receipt for a named data model, comment choice and algorithm. They did not reconstruct the original XML, retain every DTD or base-URI fact, or decide application-specific meaning.
The stable stream began in the middle
In March 2001, RFC 3076 and the W3C Recommendation for Canonical XML Version 1.0 addressed a practical problem. XML allowed more than one physical spelling of information that many applications considered the same. Attribute order could vary. Character encoding could vary. An empty element could be written in short or expanded form. Entity references and literal replacement text could lead to the same parsed content. Digital signatures and other comparisons needed one reproducible representation.
The specification delivered that representation. It encoded output as UTF-8, removed the XML declaration and document type declaration, converted empty elements to start-and-end pairs, chose double quotes around attribute values, escaped prescribed characters, removed superfluous namespace declarations, and sorted namespace declarations and attributes. With or without comments was an explicit choice. When a well-formed canonical result passed through the same method again, the result stayed unchanged.
That sounds like a function from file to file. RFC 3076 described something more exact and more consequential: a function whose formal input was an XPath node-set. An implementation had to accept a well-formed XML octet stream, but it converted that stream into the data model first. The canonicalizer met the document after parsing.
The distinction determines what the output can prove.
The parser had already performed work
To construct the XPath nodes, the XML processor normalized line endings and attribute values. It replaced CDATA sections with their character content. It resolved character references and parsed entity references. A validating or appropriately configured processor could supply default attributes declared by a DTD. Namespace processing supplied namespace nodes. Consecutive characters became text nodes in the UCS character domain.
Those were not cosmetic printing choices made by the final serializer. They were earlier transformations that established the object to be serialized.
Imagine two source files. One writes a character literally; another uses a character reference. One contains an internal entity name; another contains the entity's replacement text. One uses an empty-element tag; another writes a start tag followed immediately by an end tag. Canonical XML intentionally lets these differences converge when the parsed information is the same under its rules. That convergence is the product.
It also means the canonical output is not a reversible archive of source syntax. The XML declaration is absent. The original character encoding is gone because the output is UTF-8. The DTD is absent. Entity boundaries have disappeared into replacement text. CDATA boundaries have disappeared into character content. Original attribute ordering has yielded to a defined sort. The output can be the correct canonical form and still be incapable of answering how the input was spelled.
The receipt therefore begins with a parser contract, not merely an algorithm URI. Which bytes arrived? From which base URI? Which XML version and declared encoding were used? Was validation enabled? Which external subset and parsed entities were resolved, from where, at what time, and with which bytes? Which defaults entered the node-set? A hash of the final canonical stream does not reconstruct these inputs.
Information could disappear while behavior changed
RFC 3076 did not hide this boundary. Its limitations section identified information that was unavailable in the XPath data model: base URI, particularly for replacement text from external parsed entities; notations and external unparsed-entity references; and DTD-declared attribute types.
The base-URI case is concrete. An external entity can contain a relative URI. When the parser replaces the entity reference with its content, that content moves into the host document's node-set. If the entity's original base location is not preserved, the same relative characters can resolve somewhere else. The canonical bytes may be internally stable while the resource reached by later application processing changes. RFC 3076 advised applications to use xml:base appropriately or resolve relative URIs before the canonical form was needed.
The DTD boundary cuts in two directions. A default attribute may already have been added to the parsed element and therefore appear in canonical output, even though the document type declaration that supplied it is removed. Meanwhile, declarations that gave values the types ID, IDREF, an enumeration, NOTATION or related constraints no longer accompany the canonical file. A later parser can see the characters without recovering the original type contract. Canonical form records the resulting node representation, not the full lineage that constructed it.
Unparsed entities and notations make the same point. The XPath model did not carry every binding that an application might use to interpret external non-XML data. The RFC called these circumstances unusual, not impossible. “Unusual” is a frequency judgment; it is not a warrant to erase the input receipt when the application actually relies on one of them.
A node-set was not just a subtree
RFC 3076 also supported canonicalizing document subsets. An XPath node-set is a mathematical set of individual nodes, not shorthand for “this element and everything below it.” An element may be selected while one of its attributes or text children is not. An excluded node is not rendered merely because its parent is present. An omitted ancestor can still contribute namespace context or inherited attributes to a selected descendant.
The specification therefore placed responsibility on the creator of the input node-set to preserve the information needed for the full semantics of its members. A subset's canonical output might not even be well-formed XML. It was still a defined octet sequence for the selected nodes, but an application could not infer completeness from the mere appearance of an element.
This is distinct from the signature question in RFC 3075. That document asked which referenced data and transforms a signature covered. RFC 3076 asks what object exists before a canonical byte is emitted. A perfect signature over perfect canonical output still inherits the parser and node-set boundary that produced those octets.
Canonical did not mean every kind of equivalent
The word “canonical” invites a larger claim than the RFC made. RFC 3076 said that identical canonical forms established equivalence within the defined scope, subject to its limitations. It explicitly rejected the converse as a goal. Two documents could remain equivalent for an application while producing different canonical forms.
An application might ignore some whitespace. It might treat black and rgb(0,0,0) as the same color. Another XML-related specification might define an equivalence that general canonical XML did not know. Unicode character-model normalization was not performed as a general step. Namespace prefixes were deliberately preserved because blindly rewriting them could break XPath or QName-like strings embedded in attribute values or text.
So canonical comparison was neither a universal semantic engine nor an arbitrary byte cleanup. It was a carefully bounded normalization of the XPath representation. Its value came from being executable and testable. Its authority stopped where application-specific meaning began.
Later revisions made the boundary visible again
Experience after 2001 exposed a problem in subset inheritance. A 2006 W3C note showed that Canonical XML 1.0 could copy a relative xml:base value onto a selected descendant while losing the ancestor path needed to resolve it correctly. The later xml:id Recommendation also made it clear that an identifier should not be inherited merely because Canonical XML 1.0's general rule copied attributes from the XML namespace.
Canonical XML 1.1, published as a W3C Recommendation in 2008, revised the subset rules: it did not inherit xml:id and performed special xml:base path fixup. That evolution does not prove every C14N 1.0 implementation failed. It shows that even a deterministic serializer depends on the adequacy of the data model and context rules supplied to it.
Exclusive XML Canonicalization addressed a related but different problem: how a subset's namespace context should behave when it is moved or embedded elsewhere. RFC 3275 made Canonical XML 1.0 a required method in the XML Signature environment. Neither development enlarged a canonical hash into proof of original source bytes, trusted external inputs or authorized application behavior.
The receipt chain that history should preserve
For an archival comparison, signature, audit or migration, preserve more than the canonical digest. Retain the received octet hash and retrieval URI; media type and declared encoding; parser identity and settings; validation and external-resource policy; every DTD, external subset and parsed entity actually consumed; the effective base URI; namespace processing; defaulted attributes; the selected XPath node-set; whether comments were included; the exact canonicalization method and version; and the canonical output hash.
Then continue the chain. Record the comparison or signature result, the application-specific equivalence rule, the decision taken, and the observed effect. A canonical byte match answers a real question. It becomes misleading only when it is made to answer all the others.
Lu Heng's later running-code and reality-layer arguments offer a useful disclosed lens here. The executable output is stronger than a slogan because another implementation can reproduce it. But an execution receipt does not absorb the source, authority, meaning and outcome layers around it. RFC 3076 is durable precisely because it made one layer sharply testable.
Sources
- RFC Editor record for RFC 3076
- RFC 3076: Canonical XML Version 1.0
- IETF Datatracker record for RFC 3076
- W3C Canonical XML Version 1.0 Recommendation
- XPath Version 1.0
- XML 1.0
- Known Issues with Canonical XML 1.0
- Canonical XML Version 1.1
- Exclusive XML Canonicalization Version 1.0
- RFC 3275: XML-Signature Syntax and Processing
- Lu Heng, “Running-Code Primacy”
- Lu Heng, “On Reality Layers, Symbolic Power, and Why Clarity Feels So Hostile”
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
