Summary
- RFC 3066 standardized a case-insensitive language-tag grammar, public registration and a prefix-based language-range, while leaving each application to define what the label meant for its own information objects.
- The RFC explicitly warned that shared tag prefixes did not guarantee mutual intelligibility. A match was a selection result, not evidence that content was correctly labeled, rendered well or understood by a reader.
A small label for a large human fact
The Internet did not need a universal theory of language to route multilingual information. It needed a shared identifier that mail, HTTP, markup, libraries, speech tools and other systems could carry without each inventing a private vocabulary.
RFC 3066, published as BCP 47 in January 2001 and replacing RFC 1766, supplied that minimum. A language tag consisted of a primary subtag followed by zero or more hyphenated subtags. Comparison ignored case. Two-letter primary values came from ISO 639; three-letter values came from ISO 639-2. i opened an IANA-registered branch, while x opened private use. A two-letter second subtag could identify an ISO 3166 area.
That grammar created interoperable spelling. It did not create one owner of linguistic truth. ISO maintenance bodies assigned source codes. IANA administered the namespace and public registrations. The protocol carrying a tag defined its relationship to an object. The sender chose the label. The receiving application decided how to act on it. The reader still experienced the result.
The separation was not an omission. It was the design.
Context owned the assertion
RFC 3066 refused to give one fixed meaning to “this object has language tag X.” For a single document, a set of tags might name the languages required for complete comprehension. For a collection, it might merely list languages found among components. For a multipart set of alternatives, the labels were hints directing inspection of the individual choices. In HTML or XML, a tag could apply to one span inside a document and help a dictionary or speech synthesizer choose appropriate rules.
The characters of the tag were the same. The assertion was not.
This is why metadata cannot be audited in isolation from its binding. A database field called language might describe the prose, the user preference, the interface shell, the available alternatives or the requested output. Those are different claims even when their values all read fr. The standard using the field has to say which one it means.
RFC 3066 also advised senders to be precise only where precision was known and useful. When a two-letter ISO 639-1 code existed, it took precedence over a three-letter form. If the protocol did not force a value, omission was preferable to und for unknown language. If the protocol could carry several tags, several tags were preferable to flattening them into mul. Unknown and multiple were states to preserve, not details to hide behind a convenient token.
The prefix was a filter, not a family tree
The major addition over RFC 1766 was the named language-range. A range matched a tag when it was identical or when it formed a prefix ending at a hyphen boundary. * could match any tag, with protocols free to define additional wildcard semantics. That gave applications a simple selection tool: a request for a broader label could collect more specific labeled variants.
The RFC immediately bounded the inference. Languages whose tags began with the same sequence were not guaranteed to be mutually intelligible. The prefix rule worked where an application judged it useful; the string hierarchy did not prove a hierarchy of human understanding.
That distinction survives in RFC 4647, which names the old algorithm Basic Filtering and separates it from Extended Filtering and Lookup. A filter can return a set. Lookup can choose one tag. Neither operation interviews the reader. A successful comparison says that declared strings met an algorithm under a particular priority list. It says nothing by itself about literacy, dialect, script, accessibility, translation quality or whether the content was mislabeled.
The correct operational chain therefore has several receipts: the preference actually disclosed; the labels available on candidate content; the matching algorithm and version; the selected object; the object's real language and script; rendering or speech output; any user correction; and the eventual task outcome. Collapsing that chain into language_match=true grants a text comparison authority it never had.
Registration created memory, not comprehension
Tags outside the generative rules went through a public process. A requester supplied a description and reference, the open list reviewed it for two weeks, and an appointed reviewer forwarded or rejected it. Appeals followed the IETF process. Registrations were not deleted. When a tag became obsolete, its record was amended to recommend another value.
This was an institutional memory system. A published description helped operators learn what a tag referred to, and deprecation preserved the interpretation of old data. It did not certify every object bearing that tag. Registration answered “what identifier do we coordinate on?” It did not answer “is this document genuinely in that language?”
Private-use tags made the boundary even clearer. An x- tag could serve parties with a private agreement, but IANA neither registered its continuation nor supplied a global meaning. Private coordination was legitimate precisely because it did not masquerade as universal interoperability.
RFC 4646 later replaced the whole-tag registry with a Language Subtag Registry. It gave positions such as language, script, region and variant more explicit structure and strengthened stability: valid tags should not later become invalid. RFC 5646 refined that architecture and, with RFC 4647, remains BCP 47. IANA now keeps the RFC 3066 Language Tags registry closed and points new specifications to the subtag registry.
The successor architecture did not erase RFC 3066's central restraint. Better structure makes labels easier to parse and preserve. It does not turn them into observations of human understanding.
Preference disclosure had a human cost
RFC 1766 had treated security as irrelevant. RFC 3066 recorded a concrete concern: a language-range sent during content negotiation could help a receiver infer nationality and identify surveillance targets. The document did not claim that language equals nationality; the risk arose because observers might make that inference.
This is a subtle constitutional point. The user disclosed a preference to improve service. The server received a fingerprintable signal. The matching algorithm needed only enough information to choose among available representations, yet logging, correlation and retention could give the signal a second life.
The minimum useful label is therefore not automatically the minimum-risk disclosure. Systems should keep preference lists proportionate, avoid inferring identity from them, and separate the receipt needed for selection from profiles built for unrelated purposes.
Sources
- https://www.rfc-editor.org/rfc/rfc3066.html
- https://www.rfc-editor.org/info/rfc3066/
- https://datatracker.ietf.org/doc/rfc3066/
- https://www.rfc-editor.org/rfc/rfc1766.html
- https://www.rfc-editor.org/rfc/rfc2277.html
- https://www.rfc-editor.org/rfc/rfc4646.html
- https://www.rfc-editor.org/rfc/rfc4647.html
- https://www.rfc-editor.org/rfc/rfc5646.html
- https://www.iana.org/assignments/language-subtags-tags-extensions
- https://heng.lu/running-code-primary-the-patch-needed-to-preserve-the-internet-original-design/
- https://heng.lu/on-reality-layers-symbolic-power-and-why-clarity-feels-so-hostile/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
