Summary
- The active CBOR working-group draft defines general, preferred-plus and deterministic serialization. Deterministic serialization adds bytewise lexicographic sorting of deterministically encoded map keys to preferred-plus, giving independent encoders one repeatable byte sequence without turning maps into ordered application objects.
- Determinism is needed when parties independently rebuild bytes for a signature, hash, content address, cache key or equality test. It is usually unnecessary when the exact protected bytes travel intact, and it never proves that the decoded value is complete, current, authorized or operationally effective.
The failed verification looked impossible in the application log. Both sides printed the same diagnostic object: the same integer, the same text string, the same two map members. Yet their hashes differed. One encoder chose a longer legal representation for an argument and emitted map entries in insertion order. The other used shortest forms and sorted keys. Both had produced valid CBOR. Only one had produced the bytes the protocol designer imagined.
That imagined rule had never been written down.
This is the governance problem addressed by draft-ietf-cbor-serialization-08, dated 29 July 2026. The document is an active CBOR working-group Internet-Draft in Working Group Last Call; the Datatracker still records the IESG state as I-D Exists. It is work in progress, not a final RFC and not evidence that any product implements the rules. Its practical value is sharper: it gives protocol authors names for three different serialization contracts and asks them to assign the right one at the point where bytes acquire consequence.
The distinction matters because CBOR is both a data model and a family of encodings. A decoded value can survive several valid byte representations. That flexibility is useful for streaming, constrained devices and existing implementations. It becomes dangerous only when a later system treats one representation as if it were inevitable.
Three contracts, not one ladder of virtue
General serialization is the theoretical default when a CBOR-based protocol says nothing else. For every type it supports, a conforming decoder is expected to accept all encodings CBOR permits, including definite and indefinite lengths. This is a broad interoperability promise. The draft also makes an awkward operational observation: such breadth is not widely implemented. A specification can therefore invoke the general set while deployed decoders quietly accept a smaller language.
Preferred-plus serialization narrows encoder freedom. Arguments use their shortest form; floating-point values use the shortest representation that preserves the value exactly; lengths are definite; NaN handling follows the stated trivial representation; and integers and bignums are normalized as specified. It aims to be a practical common choice for most protocols. Crucially, it does not require map sorting.
Deterministic serialization takes preferred-plus and adds the rule that map items are ordered by the bytewise lexicographic order of the deterministic encodings of their keys. Independent conforming encoders given the same model value can therefore converge on the same bytes, provided the surrounding protocol has removed the remaining ambiguities in what the value means.
These sets are nested in permitted output, but they are not a moral ranking. General serialization preserves capabilities that some designs need: indefinite-length construction, a meaningful distinction between integers and bignums, or non-trivial NaN payloads. Deterministic serialization deliberately removes those choices. Selecting it is an interoperability decision, not a declaration that every other legal encoding is defective.
The decoder boundary is equally important. Deterministic decoding imposes no requirements beyond preferred-plus decoding. A deterministic encoder produces a restricted form, but a receiver is not thereby instructed to reject every non-deterministic input unless the end-to-end protocol separately requires that rejection. “We emit deterministic CBOR” and “we accept only deterministic CBOR” are different controls with different compatibility and security consequences.
Ask whether the bytes travel or are rebuilt
The fastest way to choose a serialization contract is to follow the protected bytes.
If a sender signs a CBOR payload and transmits those exact payload bytes with the signature, the receiver can verify the bytes it received. It need not decode the value and recreate an identical serialization first. The draft uses COSE payload handling as the intuitive case. Hashing and signatures do not automatically require deterministic serialization; they require both sides to agree on the byte string that the cryptographic operation covers.
The situation changes when each party constructs that byte string independently. A COSE Sig_structure is assembled by the signer and again by the verifier from agreed fields. If their encoders can choose different legal representations, equal model values may yield unequal signature inputs. The same risk appears in content addressing, cache keys, deduplication, reproducible manifests and byte-level equality checks. Here determinism is not aesthetic tidiness. It is the missing agreement that makes independent reconstruction possible.
This test prevents two opposite mistakes. One team imposes map sorting and shortest forms on every payload, increasing migration cost without improving a design that already carries the exact protected bytes. Another team says “CBOR is interoperable” while independently rebuilding signature input with library defaults. The right question is neither “Is cryptography involved?” nor “Is CBOR valid?” It is “Which exact bytes must two independent processes obtain, and where is that obligation written?”
Determinism is local. A deterministic requirement for Sig_structure does not automatically govern an embedded payload, an outer envelope or every future field. Protocol text should name the byte-producing boundary precisely. Otherwise one small rule becomes an accidental platform policy and later implementations cannot tell which deviations actually break interoperability.
Sorting a map does not give the map an order
CBOR maps are unordered in the generic data model. Deterministic serialization sorts their encoded keys to select one byte layout. That does not promote the chosen layout into business sequence, display priority or processing order.
This is easy to forget because packet tools show keys in a line. A developer may iterate the decoded map and assume the first key should win. Another runtime may preserve insertion order, a third may retain the encoded order, and a fourth may use a hash table. All can decode the same CBOR value. If an application needs precedence, it must model precedence explicitly—through an array, a priority field or another specified structure—not borrow it from serialization.
Key sorting also does not complete the data model. The protocol must still define allowed key types, tags, duplicate-key behavior, schema evolution, equivalence rules and the treatment of unknown fields. Preferred-plus happens to yield a deterministic result when no maps are present, but that fact cannot resolve whether two tagged values are equivalent or whether a missing member changes authorization.
Byte equality can be stricter than semantic equality. Two encodings may represent values an application considers equivalent, while a byte comparison rejects them. Semantic equality can also be dangerously vague: Unicode normalization, numeric domains, tag interpretation and application defaults may differ. Deterministic CBOR chooses bytes after the model has been chosen. It cannot choose the model on the protocol’s behalf.
The word “canonical” should therefore be handled with care. Different profiles can deliberately choose different restrictions. The separate deterministic CBOR application profile in draft-mcnally-deterministic-cbor-17 narrows CBOR further; its rules must not be silently attributed to the working-group draft. Say which profile and revision applies. A generic claim of “canonical CBOR” conceals the exact interoperation contract that an audit needs.
A valid signature is still only a byte receipt
When deterministic reconstruction succeeds and a signature verifies, the receipt is meaningful: a named algorithm under a named key protected a defined byte sequence. That is stronger than a log message saying two decoded objects looked alike. It is still not an authorization decision.
The signature does not prove that the schema admitted every fact relevant to the decision. It does not prove a tag was permitted in this context, a timestamp was current, an issuer still held authority, an optional field was not omitted, or a duplicate key was resolved safely. It does not say that a verified request may spend funds, change routing policy or modify an account. It does not observe whether the approved operation later occurred.
The layers should remain separately receipted: the information model; the CBOR data model; the serialized byte string; cryptographic verification; schema and semantic validation; local policy; authorization; attempted execution; and observed effect. A green indicator at one layer cannot inherit evidence from the next.
CDDL helps describe CBOR data shape. The serialization-control mechanism discussed by the draft can state a serialization requirement alongside such a description. That is useful coordination, but a schema is not running code and not a grant of authority. A CDDL match cannot show which library version parsed the message, which duplicate-key policy ran, which trust anchor was selected or which state transition committed.
COSE and CWT illustrate the division of responsibility. A framework can support many end-to-end protocols and cannot know every application’s reconstruction or streaming needs. Where the framework does not force a set, the draft recommends preferred-plus, while the incorporating protocol states the actual serialization requirements. Library defaults cannot safely stand in for that decision. Defaults change; wrappers hide them; language bindings expose different switches.
The decoder gap is part of the deployment contract
The general set’s theoretical breadth creates a practical trap. Procurement language may say “supports CBOR” while one decoder rejects indefinite strings, another rejects legal alternate widths and a third accepts inputs the system never tests. A standards label then substitutes for an observed acceptance matrix.
Running-code evidence needs exact artifacts: library and version, build flags, encoder mode, decoder options, schema revision, input value, emitted bytes, accepted variants, rejected variants and downstream result. Golden vectors should cover both the happy deterministic output and the boundaries the deployed receiver claims to accept. Differential tests should run the same corpus through every implementation that participates in the protocol.
Negative tests matter. Feed an alternate but legal general encoding to a component that claims general decoding. Feed a non-shortest form to a preferred-plus gate. Exchange map insertion orders while preserving value. Exercise integer/bignum boundaries, definite and indefinite lengths, floating-point width and NaN behavior. Record not only parser acceptance but the value the application received and the policy path that followed.
An encoder conformance test cannot prove decoder breadth. A decoder acceptance test cannot prove the system emits deterministic bytes. A round trip through one library can conceal both failures because the same implementation repeats its own assumptions. Interoperability evidence begins when independently built paths meet.
NaNs and special values expose hidden model choices
Floating-point handling is where a “same value” claim can become slippery. Preferred-plus asks for the shortest exact representation and prescribes the trivial quiet NaN form. Yet NaNs can carry payload bits, and some applications distinguish or preserve those payloads. A protocol that needs non-trivial NaNs cannot simply claim deterministic serialization and hope the information survives.
Similarly, a data model that distinguishes a bignum from an ordinary integer may need general serialization rather than normalization. Streaming systems may depend on indefinite-length construction because the final size is unavailable when transmission begins. These are legitimate requirements. They should produce an explicit special serialization contract, not an undocumented encoder exception.
The smallest shared rule is often the most durable. Specify deterministic reconstruction only where independent byte equality is required. Use preferred-plus for ordinary exchange when its restrictions fit. Retain general or a named special profile where the application genuinely needs the extra representational choices. This is Minimum Initial Specification in practice: coordinate the minimum surface that must be common while leaving local future decisions outside it.
Fewer legal encodings reduce one covert channel, not every one
Multiple valid serializations can carry information not present in the decoded value. An encoder can vary argument widths, chunk indefinite-length items differently or choose other legal representations. A compromised component may use those choices as a covert channel while downstream software sees the expected model value.
Preferred-plus and deterministic rules reduce that freedom. Deterministic encoding can make unexpected byte variation conspicuous. But it does not prove the absence of exfiltration. The attacker may choose among semantically allowed values, reorder data before it enters the model, manipulate timestamps, select tags, exploit padding in another layer or leak through timing and traffic shape.
Therefore record both the normalized model and the actual bytes at the trust boundary where appropriate, with retention and privacy controls. Alert when supposedly deterministic producers emit divergent bytes for the same test corpus. Do not turn that test into the claim “no covert channel exists.” It proves only that this serialization degree of freedom was constrained in this execution.
Build a receipt that survives a library change
A useful serialization receipt names the protocol and draft or RFC revision; the required set at each byte boundary; the CBOR data-model profile; allowed tags and key types; duplicate-key policy; encoder library, version and options; decoder library, version and accepted set; schema revision; original input; emitted bytes or a durable digest plus retrievable evidence; cryptographic input; key provenance; verification result; semantic validation; policy decision; authorization; attempted operation; and observed effect.
It also records where exact bytes travel and where they are reconstructed. That one diagram frequently reveals the real defect. Teams discover that a proxy decodes and re-encodes a signed payload, that a cache hashes a parsed object, or that two services construct a structure from fields with different default handling.
Ownership should be distributed rather than blurred. Protocol designers own the byte-boundary rule. Implementers own encoder and decoder evidence. Security engineering owns cryptographic and parser assurance. Application owners define semantics and equivalence. Policy owners define authorization. Operations prove the effect. A transaction identifier and time join the receipts without pretending they are one assertion.
Lu Heng’s Running-Code Primacy asks what actually ran, not which label appeared in a design document. Reality Layers prevents a deterministic byte string from borrowing the authority of a schema or a signature from borrowing the authority of a business decision. The result is not scepticism about standards. It is a cleaner use of them.
Deterministic CBOR solves a precise problem: independent encoders need one byte sequence. It does not solve every problem located after decoding. The value can be the same while the bytes differ; the bytes can be identical while the meaning remains disputed; the meaning can be accepted while the action remains unauthorized. Good architecture keeps all three sentences available.
Sources
- CBOR Serialization Considerations, revision 08
- Datatracker record for CBOR Serialization Considerations
- Revision history for CBOR Serialization Considerations
- CBOR Serialization Considerations, revision 07
- RFC 8949: Concise Binary Object Representation
- RFC 9052: CBOR Object Signing and Encryption
- RFC 8392: CBOR Web Token
- RFC 8610: Concise Data Definition Language
- RFC 9413: Maintaining Robust Protocols
- Deterministic CBOR application profile, revision 17
- Lu Heng: Running-Code Primacy
- Lu Heng: Minimum Initial Specification, Localized Future Decision, and Voluntary Adoption
- Lu Heng: On Reality Layers, Symbolic Power, and Why Clarity Feels So Hostile
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
