Summary
- RFC 1439 showed that name-derived account strings trade predictability for collision risk; the namespace, comparison rules and assignment event—not resemblance to a name—create uniqueness.
- Its most consequential advice was operational: always expose the disambiguating suffix, reject the ambiguous unsuffixed address and do not reuse the identifier while old mail can still act on it.
- Later RFCs separate the layers more explicitly: a display name, username, provider-issued stable ID, client-scoped external ID, mailbox, SMTP handoff, credential and authorization decision are different evidence.
The second Craig changes the system
Imagine an early campus network publishing one simple convention: take a first name, a middle initial and a family name, place punctuation between them, and the result is the account. The convention is useful precisely because a correspondent can infer the address without consulting a directory. Then another person arrives whose normalized name produces the same candidate.
The collision is not a rare defect added from outside. It is a property of the format. The rule has compressed a varied social object—a human name—into a smaller technical namespace. Someone must now decide whether the second account receives a number, whether the first account also displays a number, whether case and punctuation distinguish entries, and what happens when either person leaves. Every answer changes what a sender can safely infer from the visible string.
That was the practical subject of RFC 1439, published by Craig Finseth in March 1993. The memo was Informational, not an Internet standard. Its title, “The Uniqueness of Unique Identifiers,” sounds playful. Its core claim was severe: systems needed short account identifiers to be unique, but the formats that made them guessable from names did not contain enough information to make collisions exceptional.
The memo separated three assignment strategies. A system could issue an opaque unique string unrelated to the person. That made uniqueness manageable but made guessing impossible. It could derive most of the string from the name and vary the result—often with a numeric suffix—until it was unique. Or it could derive the string from the name, allow duplicates to arise, and handle them ad hoc. Electronic mail increased the value of the second and third models because a sender who knew the convention could predict an address.
That convenience created a false visual equivalence. Craig.A.Finseth looks like a proposition about Craig A. Finseth. Operationally, it is only a candidate key under one issuer's rules. The name supplies input. A normalization rule supplies a comparison form. An availability check observes a namespace at a moment. An allocation record binds a slot. None of those acts is identical to the person.
A birthday problem in the staff directory
RFC 1439 tried to quantify the compression. It estimated the typical and maximum information carried by initials, first names, family names and combinations. Most formats fell in a typical range of 8 to 20 bits. The richest typical combination in its table reached 26 bits, and none of its maximum estimates exceeded 40 bits. On those assumptions, duplicates were not merely possible; they became likely as the population grew.
The memo used the birthday-problem calculation: collision risk rises with pairs, not merely with the fraction of the namespace apparently occupied. Its worked example assigned a typical 17 bits to a first-name-plus-family-name format. Among 100 people, it estimated a duplicate probability between 2 and 5 percent, probably around 4 percent. Among 1,000, the probability was much greater than 20 percent.
Those figures should not be mistaken for a universal demographic law. The source books, cultural assumptions, transliteration practices and organizational populations were historically situated. Different scripts, naming customs and normalization rules produce different distributions. The durable insight is structural: a readable format's collision rate depends on the actual distribution of inputs, the comparison function and the number of assignments. Counting visible characters is not enough.
The same structural lesson applies whenever a system strips accents, folds case, drops punctuation, shortens long names, transliterates scripts or removes spaces. Each transformation can improve interoperability in one context while merging candidates that were distinct in another. A namespace cannot announce “unique” without also naming its scope and equality rules.
The suffix is part of the address, not an exception note
The appendix posed the decisive example. Suppose the format is First.M.Last-#. May the first holder use the pretty unsuffixed form while only the second receives -2? RFC 1439 answered no. If an outsider sends to the unsuffixed form, there is no reliable signal that the message may be going to the wrong person. If every account carries a suffix and the unsuffixed form is rejected, the failed attempt tells the sender that a discriminator is required.
This changes the purpose of an error. Rejection is not merely failure to serve a convenient alias. It preserves uncertainty that the system cannot safely resolve. Silently choosing the first claimant would convert a naming collision into confident misdelivery.
The memo tied that reasoning to the care owed to electronic mail and cited the 1987 United States Electronic Communications Privacy Act. That citation belongs to its 1993 rationale; it is not a statement of current law. The technical point survives without a legal extrapolation: when two plausible people share a name-derived candidate, successful routing to one mailbox is weaker evidence than an explicit refusal to guess.
RFC 1439 then advised that such identifiers should not be reused during the life of the mail system. Reuse creates a collision across time rather than across simultaneous staff. Old address books, archived messages, access-control entries, subscriptions, recovery contacts and human memory may continue to refer to the former holder. A namespace that reassigns the same visible slot has changed the person behind the string while preserving every stale reference that makes the string seem trustworthy.
The relevant lifetime is therefore not necessarily employment or account activity. It is the lifetime of dependent references. If an old message, ACL, audit record or recovery channel can still cause an action, the identifier remains operationally alive.
SMTP can accept responsibility without proving the person
Later SMTP language makes the evidence boundary precise. RFC 5321 says an address is a character string that identifies a user to whom mail will be sent or a location into which mail will be deposited; a mailbox is that depository. Contemporary addresses serve more purposes than simple usernames, so only the host named by the domain assigns meaning to the local-part.
This is a sharply bounded authority. A remote sender can preserve the string and route toward the domain. The receiving domain decides what its local-part denotes. The syntax does not reveal whether the target is a person, a shared queue, a forwarding alias, a program or an abandoned account kept alive for continuity.
SMTP itself contains another clean boundary. After a server sends a successful response at the end of the message data, responsibility formally passes: that server must deliver the message or properly report failure. The success proves a protocol handoff. It does not prove that the intended person owns the mailbox, that a human read the message, or that any downstream action occurred.
RFC 2142 makes the non-person case deliberate. Addresses such as postmaster, abuse, noc and security are well-known names for services, roles and functions. A domain supporting the function is expected to route the name to an appropriate recipient for that role. Continuity of the function is the goal; identity of a particular person is not.
Even case has no universal meaning detached from the namespace. RFC 5321 requires preservation of mailbox-local-part case and formally treats the local-part as case-sensitive, while discouraging deployments from exploiting that sensitivity because it impedes interoperability. Domain names are case-insensitive. A string comparison performed by an intermediate system cannot safely invent the destination host's account semantics.
Later identity schemas made the layers visible
RFC 7643, the SCIM core schema, does not solve the historical mail problem. It is useful here because its vocabulary refuses to pretend that one attractive field can serve every identity function.
A SCIM id is issued by the service provider, unique across that provider's resources, stable and non-reassignable. An externalId is issued by a provisioning client and scoped to that client's provisioning domain; the service provider does not enforce its uniqueness. A userName is a user-facing identifier, required and unique across the provider's Users. Human name components are separate attributes.
The differences identify four authorities. The provider controls its resource key. The client controls its external reference. The provider controls the login namespace. A person or organization supplies presentation data. Two strings can look identical while making claims in different scopes, and two different strings can legitimately refer to the same resource through different systems.
RFC 8265 adds the internationalized comparison layer. It defines a username as a string designating an account, often but not necessarily used by a person. It explicitly notes that a username does not necessarily map to a particular application identifier and encourages separation between restrictive account identifiers and expressive display names or nicknames.
It also defines two profiles: one maps case and one preserves it. The choice belongs to the application protocol, implementation or deployment. Case mapping is generally preferred where failure to map could cause false accepts or user confusion, but preservation may be required for compatibility. Mapping loses information, so the RFC advises delaying it until the last possible moment. The lesson is not “always lowercase.” It is “declare which comparison the verifier actually runs.”
An evidence ladder for a person-like string
The Internet-history mistake is to narrate all these operations as “identity.” A more useful ladder preserves what each step proves.
A displayed human name is presentation evidence. A normalization or transliteration output is a candidate comparison string. An availability response says that candidate is unused under one namespace's current rules. Assignment says an issuer bound a slot to a resource at a time. A mailbox address identifies a depository or user under the destination domain's semantics. SMTP success proves responsibility for further delivery or failure reporting. Authentication proves control of credentials under a verifier's rules. Authorization proves permission for one action under one policy.
A reply or independent channel may support a claim of human participation, but even that proof has a time and context boundary.
Every rung is useful. None automatically inherits the meaning of the rung above it.
That separation follows a broader Internet discipline: keep the common layer minimal and deterministic. Uniqueness needs a stated namespace, comparison rule, collision rule and assignment transition. Interoperability needs the visible identifier to be transported without intermediaries inventing semantics. Everything else—display preference, organizational naming custom, alias aesthetics—can remain local. A directory entry is a coordination record. Running comparison and routing code create the executable outcome.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
