Summary

  • RFC 5256 makes sorting and threading deterministic, but every result is a search-scoped projection shaped by the requested algorithm, collation, header quality and mailbox sequence state.
  • A thread edge can come from a declared Message-ID reference, a missing-message placeholder or a subject-based merge; none alone proves author intent, complete ancestry, custody or delivery.

The tree arrived before the evidence

The compliance reviewer saw six messages nested beneath one root. The product team saw a decision followed by five acknowledgements. The incident team saw a command chain. All three were reading the same screen, and all three were about to grant the screen more authority than the protocol did.

RFC 5256 defines two things that feel ordinary because mail software has made them invisible: ordered search results and server-side conversation trees. They are excellent interface machinery. They let a client ask for messages that match a predicate, sort them without downloading a mailbox, and receive a compact parenthesized thread structure. What they do not provide is an attestation that the visible tree is the historical conversation.

The distinction begins before the first edge. THREAD first runs a search. Messages outside that predicate are not ordinary members of the returned set. Change the date range, text query or mailbox state and the visible family can change while every stored message remains byte-for-byte identical. The output answers, “How does this algorithm arrange the matching records now?” It does not answer, “What was the complete human exchange?”

One command contains several authorities

The client chooses the search criteria, the character set and the threading algorithm. The server controls the selected mailbox state, parses the messages, applies the required normalization and emits sequence numbers or UIDs. The sender supplied the headers. The user interface decides how to draw the result. No one actor owns the whole claim.

That separation matters because each layer can be correct while the final interpretation is wrong. A server can implement REFERENCES exactly, a client can render the response faithfully, and a user can still infer a reply relationship that a false header invented. Conversely, a missing parent can be absent because the search excluded it, the mailbox never contained it, or it was removed before the current snapshot. The same visual gap can represent different histories.

Lu Heng's reality-layer discipline is useful here. The record layer contains messages and server metadata. The transformation layer contains search, normalization, collation and tree construction. The presentation layer contains indents, disclosure triangles and labels such as “conversation.” A clean interface may compress those layers; an evidentiary system must keep them separable.

ORDEREDSUBJECT is useful because it admits its poverty

RFC 5256 calls ORDEREDSUBJECT “poor man's threading.” It reduces a subject through a fixed base-subject procedure, groups equal results and orders the group by sent date. The first message becomes the root. Later messages become its children and siblings of one another. There are no grandchildren.

That is not a defective version of reply ancestry. It is a different product: subject grouping. Its honesty lies in the rule. A mailing list can use it when references are unavailable, and a user may find the result convenient. But the flat tree does not show who answered whom. Equal base subjects can join unrelated messages; different subjects can split an actual conversation.

The base-subject process itself makes editorial choices. Encoded words are decoded. Tabs, continuations and repeated spaces collapse. Recognized trailers, reply prefixes, forward wrappers and bracketed blobs can be removed. All compliant connected and disconnected implementations are told to use the same procedure so their views do not drift. Consistency is valuable. It does not turn removal into truth. The RFC explicitly warns that significant text can be mistaken for an artifact and produce incorrect collation.

REFERENCES builds ancestry by resolving imperfect claims

REFERENCES begins closer to the apparent meaning of a thread. It uses Message-IDs listed in References, or under defined conditions the first valid Message-ID in In-Reply-To. It normalizes quoted and unquoted forms that denote the same identifier. It then tries to construct parent-child links without loops.

But the algorithm must operate in the real world, where evidence is missing and contradictory. A message without a valid Message-ID receives a unique identifier invented for this computation. If several messages claim the same Message-ID, the first by lowest sequence number keeps it and later duplicates receive invented identifiers. If a referenced ancestor is missing, the server creates a dummy message so the remaining shape can be built.

Those dummy nodes are then pruned under explicit rules. A childless dummy disappears. A dummy with children may disappear while its children are promoted. A structural dummy can remain when promotion would change the root shape. Existing parent links can be kept or broken when truncated reference lists create conflict. Loops are refused.

This is disciplined repair, not forensic discovery. The resulting tree records how one standards-defined procedure resolved the available header claims. A dummy node is not a recovered email. A promoted child is not proof that no intermediary existed. An invented identifier is not a newly authenticated identity.

Subject fallback can join what headers did not

After the Message-ID tree is constructed, REFERENCES still compares base subjects among roots. Threads with the same non-empty base subject can be merged according to rules that prefer a non-reply root in some cases or create a dummy parent in others.

This is why a displayed edge needs provenance. A direct header reference, a chain inferred from several identifiers, and a subject-level root merge are not equivalent claims. They may look identical after the client draws connectors. Without the rule that created each edge, a reviewer cannot tell whether the tree follows declared ancestry or merely shares a normalized subject.

RFC 5256 says the danger plainly: false References: data can cause one thread to be incorporated into another. Normalizing a Message-ID prevents quoting differences from causing false non-matches. It does not authenticate the sender of the field or prove the field's narrative.

Order is also a policy result

SORT first searches, then applies the requested criteria. Strings follow a collation. Dates and sizes have defined ascending orders. Equal explicit criteria fall back to mailbox sequence number. That implicit final key means REVERSE SUBJECT is not simply the complete reverse of SUBJECT; the sequence tie-break does not reverse with it.

The DATE criterion also needs care. It starts with the message's Date: header and normalizes the time zone. Invalid or missing components receive prescribed substitutes. When no usable sent date exists, the server uses INTERNALDATE. One list can therefore contain author-declared time, corrected assumptions and server-held arrival metadata under one “date” column.

RFC 5957 later added display-name sorting for From and To. It decodes a display name when present, otherwise falls back toward mailbox and host material, and deliberately does not guess a language-dependent surname. That is the right restraint. A useful order should declare its key instead of pretending to be culturally universal.

Stable output is not a historical receipt

UID variants improve identification relative to volatile sequence numbers, but even a UID needs its mailbox validity context. RFC 5267 can keep a search or sort view updated as the mailbox changes. RFC 5182 can save a result set for reuse. These capabilities make the interface efficient and responsive. Neither turns the view into an immutable audit trail.

For decisions that depend on conversational ancestry, retain the messages, relevant raw headers, mailbox identity, UID validity, UID, search predicate, algorithm, collation, server revision and result hash. Record whether each edge came from a direct reference, fallback, dummy node, promotion or subject merge. Preserve the client rendering revision separately.

Then state the conclusion at the right layer. “The RFC 5256 REFERENCES view placed B beneath A in this snapshot” is auditable. “B was sent in response to A” is a stronger claim requiring independent support. “The recipient read and accepted A” is stronger still and is outside this protocol result.

The point is not to distrust email threads. It is to stop asking an interface projection to certify events it was designed only to arrange.

Sources