Summary

  • RFC 3454 defined an ordered framework—map, normalize, prohibit, then check bidirectional rules—but every using protocol still had to name a profile, repertoire and policy.
  • Equal prepared output meant equality under that profile and Unicode version. It did not establish equal spelling, visual appearance, user intent, account control, authorization or human identity.

The Internet needed comparison before it could claim meaning

Unicode allowed protocols to carry far more than US-ASCII, but it also made simple comparison unreliable. The same user-perceived text could arrive through different code-point sequences, compatibility characters, case variants or invisible marks. Published in December 2002, RFC 3454 defined stringprep as a framework for preparing such strings before a protocol stored or compared them.

The ambition was practical: increase the chance that two people entering what they believed was the same string would produce character-for-character matching output. The wording mattered. Stringprep improved likelihood under specified rules; it did not define linguistic truth. It expressly could not cover every alternative spelling across languages and scripts.

The plain text, RFC Editor record, Datatracker file, history, references, later citations and errata search establish the standards record. They do not show what any particular user intended to type or which account controlled a resulting identifier.

The framework did nothing until a profile made choices

Stringprep was not a standalone preparation policy. A protocol had to define a profile stating its applicability, permitted repertoire, mapping tables, normalization choice, prohibited output, bidirectional test and any additions. Two protocols using different profiles could legally treat the same input differently.

That design preserved local protocol needs while preventing hidden improvisation. A profile could map some characters to nothing, replace one character with another or expand one into several. Mapped output was not rescanned during the same mapping stage, so the transformation path itself was part of the contract.

If a profile selected Unicode normalization, RFC 3454 required Normalization Form KC. The underlying normalization discipline is documented in Unicode Standard Annex #15. Compatibility normalization made more forms converge, but convergence was not a claim that they looked identical in every font or meant the same word in every language.

Order was an executable part of the rule

The four operations had to run in sequence: map, normalize, prohibit, then check bidirectional constraints. Reordering them could change the result. A character removed during mapping could never be rejected later; normalization could produce a sequence that the prohibition stage then had to test.

The prohibition stage returned either a prepared string or an error, never both. Controls, private-use code points, noncharacters, surrogates, tagging characters and other classes could be excluded by the profile. The bidi rule constrained strings containing right-to-left characters so their edges and directional classes would not create unstable protocol identifiers. These checks controlled syntax and display risk. They did not authenticate the typist.

Security guidance such as Unicode Technical Report #36 later made visual-confusability risks more explicit. Even perfect normalization cannot make every look-alike equal, and making two code-point sequences equal cannot prove that one person owns both names.

The version number lived inside the evidence

RFC 3454 fixed its tables to Unicode 3.2 and warned that the document did not automatically apply to later Unicode versions. That pinning made implementations reproducible, but it also meant that “the string was valid” was incomplete without the profile and table version.

Unassigned code points exposed the compatibility problem. Stored strings were not allowed to contain them, because a later Unicode version might assign them properties that changed processing. Queries could treat them more permissively so newer clients could search older stores. The same input could therefore be rejected for creation and tolerated for lookup without contradiction. Operation type was part of the result.

This asymmetry protected interoperability, but it ruled out a single timeless truth value. A log that retains only the final output loses whether the input was a stored identifier or a query, which tables were used, and which transformations occurred.

Nameprep showed both usefulness and scope

RFC 3491 made Nameprep a stringprep profile for the original IDNA system, while RFC 3490 placed it in the domain-name workflow. The later RFC 4690 review recorded difficulties, and the IDNA2008 vocabulary in RFC 5890 moved away from stringprep.

The migration was not a verdict that comparison was unnecessary. It showed that a framework tied to a fixed Unicode version and inherited profile choices was difficult to evolve. RFC 6885 described the PRECIS problem; RFC 7564 obsoleted RFC 3454; and RFC 8264 later replaced that first PRECIS framework. Standards succession proves continued repair, not that every older stored identifier changed meaning at once.

Equality needed a receipt with its scope

Heng Lu’s reality-layer discipline makes the record boundary precise. Original input, chosen profile, Unicode tables, mapping result, normalized result, prohibition or bidi decision, final comparison, account binding and authorization are different receipts. A boolean “match” cannot replace them.

Running-code primacy requires the claim to follow the operation actually executed. The code can show that two outputs matched; it cannot appoint their owner. The minimum-initial-specification lens explains why the generic framework left identity and authorization to using protocols. This is a later editorial interpretation, not a statement about the authors’ private intent.

RFC 3454’s lasting lesson is not that normalization is weak. It is that normalization is strong only inside its named boundary. It can make comparison deterministic. It cannot turn comparison into identity.

Sources