Summary

  • RFC 3987 made Internet identifiers readable in Unicode without declaring every readable, encoded or normalized form interchangeable.
  • It required operators to separate the original character string, transport bytes, URI mapping, local comparison and the resource reached.

In January 2005, RFC 3987 introduced the Internationalized Resource Identifier. Its status record, errata record and Datatracker history preserve a deliberately limited move. The RFC did not redefine the URI. It created a complementary protocol element—a sequence of Unicode characters—and defined a deterministic way to map that sequence into a URI.

That choice mattered because readability and identity are not the same operation. A person may enter an address in Arabic, Japanese or accented Latin script. A legacy protocol may carry a percent-encoded URI. A cache may compare a normalized form. A DNS implementation may convert one host label. The application may then retrieve a resource. These can be coordinated steps, but none is a substitute receipt for the others.

RFC 3986, with its status, errata and history, had already made URI comparison purpose-dependent. RFC 3987 inherited that ladder and added a character-versus-octet problem. Two strings stored as UTF-8 and UTF-16 can denote the same character sequence while failing a byte comparison. Simple IRI comparison therefore works code point by code point after bringing both strings into a common encoding form.

The RFC then draws its sharpest boundary: that character comparison must not first map the IRIs to URIs. The mapping can create additional, spurious equivalences. A conversion useful for transport is not automatically safe as an identity function. The RFC follows the rule to its operational consequence: an IRI should not be modified in transit when it may be used as an identifier.

Normalization does not disappear. It moves to named places. An IRI created from paper or a known non-Unicode encoding is prepared in NFC during IRI-to-URI conversion. An IRI already carried in UTF-8 or UTF-16 is not normalized by that step. Authors are advised to create IRIs in NFC, but third parties are warned not to normalize an existing identifier arbitrarily when they do not know how its owner treats the sequence.

RFC 5198 and its status, errata and Datatracker record later explained why NFC makes network text easier to compare and why assigned normalized strings were designed for stability. That network-text discipline supports creation and interchange. It does not cancel RFC 3987's identifier-custody rule. A canonical text form is still not permission for an intermediary to replace the evidence it received.

Percent encoding supplies a useful test. Different hex case or an encoded unreserved character can be aligned when an application needs a particular local equivalence. RFC 3987 permits conversion and escape alignment for that comparison, then requires the original form to be preserved if the identifier will be passed onward or used again. The normalized comparison key is an index; it is not the source record.

Nor can string comparison reveal the whole world. Two equivalent strings can be shown equivalent under a selected rule. Two different strings cannot thereby be proven to name different resources: one owner may serve the same object under multiple identifiers. Retrieval also does not solve comparison in the general case, because the comparer rarely has complete knowledge or control of both resources.

Domain names add another boundary. RFC 3987 was written with the earlier IDNA model. RFC 5890, its status, errata and history later distinguished Unicode U-labels, ASCII A-labels, registration and lookup. RFC 5891, with status, errata and Datatracker history, defines the validation protocol. A readable host label is therefore not yet a registered label, a successful lookup or an authorised resource.

RFC 5895, its status, errata and history, made the input boundary explicit. Locale-aware mapping belongs in the user interface before domain-name protocol processing. It is not a universal wire algorithm. Turkish case, Asian width forms, typing, voice input and copy-and-paste can require different user-facing decisions.

The historical achievement of RFC 3987 was therefore not a single magic conversion. It was an architecture of custody. Keep the original characters. Record the encoding. Name the mapping. State the comparison purpose. Apply scheme and domain rules at their own boundaries. Record resolution and application outcome later. The identifier became readable because more human language could enter the system; it remained reliable only if the system resisted collapsing every representation into one undocumented string.

Sources