Summary

  • RFC 1766 gave language tags visible internal structure, but instructed applications to handle each complete tag as one token; the divisions were an administrative mechanism, not a navigation aid.
  • ISO-derived values, IANA registrations and private-use tags occupied different parts of the namespace, while the standard using a tag defined how it related to its information object.
  • Later matching rules added explicit operations. They did not make every hyphen in the 1995 grammar a path, menu or promise about how an application behaved.

The dash looked like a breadcrumb

In March 1995, RFC 1766 offered a compact way to label the language used in an information object. The syntax looked as if it might invite a tree: one primary tag, then optional pieces separated by hyphens. A reader could see a familiar shape in en-US, or notice that az-arabic and az-cyrillic shared an opening segment. The RFC placed a firm limit on that intuition. Applications should treat the whole value as a single token. Its division into a main tag and subtags was an administrative mechanism, not a navigation aid.

That instruction matters because the hyphen made the string readable without assigning every piece a universal behavior. The primary tag could contain one to eight letters; each subtag had the same length limit. Whitespace was not allowed, and case did not change the value. Conventions might capitalize country codes or write language codes in lower case, but RFC 1766 said those conventions carried no meaning of their own.

The namespace was not a free-form list. Two-letter primary tags followed ISO 639. The primary value i was reserved for IANA-defined registrations; x opened private use, whose following subtags IANA would not register. Other primary values were held for a future revision. In the first subtag, two-letter codes followed ISO 3166 alpha-2, while three-to-eight-letter values could be registered with IANA. Later subtags could also be registered.

The structure therefore did two jobs that are easy to confuse. It let people read a tag as a sequence, and it gave administrators a way to partition a namespace. It did not say that each segment should be traversed like a directory. Nor did it establish that the same initial characters always had the same application-level meaning. The memo said the context-defining standard determined how a tag related to its information object.

Registration gave the pieces a public record

RFC 1766’s examples ranged across country identification (en-US), dialect or variant (no-nynorsk, en-cockney), an IANA-defined language (i-cherokee) and script variants (az-arabic, az-cyrillic). The memo added an important qualification: none of the example subtags had actually been assigned. They illustrated the syntax, not a list of available values.

For a value outside the predefined ISO assignments, a requester filled in a language-tag form with a name, native name, published description and other details. The form went to an open mailing list for a two-week review. A reviewer appointed by the IETF Applications Area Director could forward the request to IANA or reject it after significant objections; the decision could be appealed to the IESG. A shared spelling became legible to other participants through a public record, not merely because someone put another suffix after a hyphen.

The mechanism was deliberately narrow in another way: it registered identifiers and references, not an all-purpose user interface. A tag did not tell every consuming protocol whether it should choose a representation, filter a library, select a route or display a language menu. That behavior belonged to the context in which the tag was used.

A header could list languages; it still did not define the chooser

RFC 1766 also defined Content-Language, which could list one or more complete tags. For MIME multipart/alternative, it introduced a Differences parameter so a reader could distinguish alternatives that differed by content language. The memo described why a reader might use language information when choosing a body part, but left the mechanism for making that choice outside its scope.

That separation kept three questions apart: what string identifies the language of an object, which protocol element carries that string, and what a receiving application does with it. A language tag could appear in a content header; a container could say its alternatives differed by language; a particular reader still needed its own selection behavior. The RFC supplied the shared label and a point of attachment, not a universal navigation interface.

The later record makes the distinction clearer. RFC 3066, published in 2001, extended the syntax to digits in subtags and introduced language-range matching. A range could match a complete tag or a prefix ending at a hyphen boundary. That was a defined matching operation. RFC 3282 later specified Content-Language and Accept-Language headers, including preferred language ranges and optional quality values. These additions created explicit protocol behavior around tags; they do not turn the original hyphen into a general path.

RFC 4646 in 2006 and RFC 5646 in 2009 continued the BCP 47 revision trail. The changes matter as evidence that grammar and registry rules evolved. They do not prove that a particular client parsed every value correctly or that a tag automatically selected a page, translation or rendering. RFC 1766’s lasting design lesson is narrower: a structured identifier can remain one indivisible application value while its segments serve administrative conventions.

Sources and limits

The specification is RFC 1766. The bounded revision trail is RFC 3066, RFC 3282, RFC 4646 and RFC 5646. They establish grammar, registration and later protocol changes; they do not establish universal implementation or deployment.