Summary

  • RFC 3987 gave Unicode-capable Internationalized Resource Identifiers a defined place beside URI-centred software. Its conversion rules depend on the component: a domain-style host can require IDNA processing, while Unicode in a path or query is represented through UTF-8 and percent-encoding.
  • Martin J. Dürst is named with Michel Suignard as an RFC 3987 author; Dürst’s university CV describes him as the main author of the IRI specification. The standard joins reader-facing scripts to older software, but it does not register a name, establish control of a domain or prove that visually similar strings identify the same resource.

A URL looks singular; its boundaries are not

Consider an address containing Japanese characters in its path and a domain name written in another script. A person may see a single web address. A client must parse it into a scheme, authority, path, query and perhaps a fragment; different components then meet different protocol rules. The human-readable surface and the representation passed to a URI-only component can be related without being character-for-character identical.

That distinction was the point of Internationalized Resource Identifiers, or IRIs. RFC 3986 defines the generic syntax for Uniform Resource Identifiers (URIs), whose character repertoire is limited to a subset of US-ASCII. RFC 3987, published in January 2005, adds a Unicode-capable companion rather than silently changing the older URI definition. Its authors explain that a new protocol element preserved a clear boundary and avoided incompatibilities with existing software. An IRI can be used where a format or component accepts it; when a retrieval path accepts only a URI, the IRI has to be mapped into the corresponding URI form.

This was not a promise that every old network component would suddenly understand every script. It was an interoperability design: let capable software retain a broader character sequence, and define how to cross into a narrower interface when necessary. The distinction matters because an identifier may be stored, displayed, copied, or used to retrieve a resource. Those operations need not occur in the same component or at the same time.

The host is a special case

The most consequential split appears after //, in the authority portion of a web address. When the host is a DNS-style domain name, RFC 3987’s 2005 mapping calls for the IDNA ToASCII operation on each dot-separated label. The result is an ASCII-compatible form suitable for URI-era processing. The RFC’s example transforms the host résumé.example.org to xn--rsum-bpad.example.org.

That example is useful because it shows what Punycode is—and is not—doing. It encodes a domain label into an ASCII-compatible label. It does not translate an entire URL, nor should it be applied to every non-ASCII component. xn-- is also not a magic badge: RFC 5890 distinguishes an A-label that has passed IDNA validity requirements from a string that merely resembles an A-label. Validation is part of the protocol; visual appearance is not.

There is also a historical boundary in the standards. RFC 3987 cites RFC 3490, the 2003 IDNA protocol, for this host conversion. The later IDNA2008 framework, including RFCs 5890 and 5891, revised the terminology and protocol rules. RFC 5895, an Informational document, discusses mappings that an application may apply to user input before it enters the IDNA2008 protocol. It explicitly recognizes that a useful input mapping can depend on locale, application and input method.

So the RFC 3987 example should be read as the rule in that specification’s 2005 context—not as a guarantee that every present-day browser applies one universal transformation.

A path is not a domain label

Move the same kind of characters to the path and the treatment changes. A path segment such as /研究 can be represented in a URI as /%E7%A0%94%E7%A9%B6, with its UTF-8 octets percent-encoded. Punycode is not the path algorithm. The path may be interpreted by a web server, a framework, a file store or an application-specific router; DNS does not resolve each path segment.

Queries and fragments introduce their own semantics as well. A percent sign, slash, question mark or hash can be structural rather than ordinary data, so parsing and escaping order matter. RFC 3987 retains URI component syntax while extending which characters may appear directly in an IRI. Its mapping rules are therefore not “replace every Unicode character with an ASCII spelling.” The scheme and the component determine the right operation.

This is why the standard recommends delaying conversion until a component that cannot handle IRIs is reached. Mapping too early can discard a useful human-facing form before another IRI-capable application receives it. Mapping inconsistently, or decoding in a different order at the server, can make two systems disagree about what path was requested. Dürst and Suignard’s design is thus about a seam between systems, not simply about printing non-Latin text in an address bar.

A valid encoding does not confer a name

IDNA processing answers a narrow question: can a label be represented and validated under the applicable rules? Domain registration and lookup are separate processes in RFC 5891. The standard says registrar-facing intake and pre-registration handling are outside the IDNA protocol’s definition; registry or zone-manager processing validates the specific string submitted for registration. A syntactically valid A-label is not evidence that the label has been registered, delegated in the DNS, or placed under the control of the service a reader expects.

That separation is easy to miss when a familiar-looking Unicode name is copied into a document. There are at least four different receipts: the characters a person entered, the mapping a user interface applied, the host label presented to DNS, and the response returned by a web service. Registration records and proof of service control are further evidence, not alternate spellings of the same string. A successful conversion alone settles none of them.

The distinction is also a security issue. RFC 3987 warns about spoofing in both host and path components: visually similar characters, differing normalization expectations, or a mismatch between client and server processing can make two addresses look alike while selecting different resources. The standard does not say that Unicode itself is unsafe. It says that a system must understand the component and transformation it is relying on. A display string is evidence about presentation; it is not a certificate of identity.

Dürst’s role, without the lone-inventor story

The RFC’s byline names M. Dürst and M. Suignard. Dürst’s official Aoyama Gakuin University CV calls him the main author of the IRI specification and records his earlier work in web internationalization, Unicode use and composite-character normalization. It also places him in the W3C Internationalization Activity, which he led for much of the period when RFC 3987 took shape. The university identifies him as a professor in its College of Science and Engineering.

That record supports a substantial contribution, not solitary authorship. The significance of the work is visible in the design choice: instead of making URI readers guess what new characters mean, IRI gives Unicode-capable software its own character-level representation and specifies where conversion is needed. Dürst’s contribution belongs to a wider standards effort, and the precise work product is the RFC jointly credited to him and Suignard.

The article’s link to Heng Lu’s “Right to Accurate Records” is deliberately narrow. The note’s bill is written for Regional Internet Registries and Internet number resources; it is not DNS policy and does not govern domain names. Its useful question here is only whether a record describes the state it claims to describe. In this case, a Unicode spelling, an A-label, a DNS delegation and a web service’s resource key are related records at different layers. One cannot stand in for all the others.

Sources