Summary
- RFC 3536 separated the encoded octets carried by a protocol from abstract characters, scripts, languages and the glyphs selected by a renderer; none of those layers can stand as automatic proof of the next.
- The 2003 glossary was deliberately Informational and non-normative, and RFC 6365 obsoleted it in 2011. Its durable lesson is evidentiary: a plausible display is a rendering receipt, not proof of language or meaning.
Consider a support engineer looking at a customer name on a screen. The letters appear familiar, evenly spaced and perfectly legible. The engineer marks the text pipeline healthy. Yet the same picture could have been produced after the wrong charset label, an unintended normalization, a font fallback, a reordered bidirectional run or a language guess made by software. The screen has supplied one fact: a renderer found shapes to draw. It has not reconstructed the custody of the text.
RFC 3536 was published in May 2003 under the title Terminology Used in Internationalization in the IETF. It was not a character encoding or an interoperability profile. It was an Informational glossary assembled because engineers were using the same words for different layers. Its authors were unusually candid about the limits: the definitions were not normative, the list was not complete, and the community did not agree on every term—especially “language.” The RFC Editor’s frozen errata record lists no errata, but an empty errata list does not turn a provisional vocabulary into settled ontology.
The glossary begins where network protocols do: encoded data. Bits are grouped into octets; a charset tells a recipient how octets correspond to abstract characters. RFC 3536 describes that arrangement as a coded character set plus a character encoding scheme and points to IANA-registered names. UTF-8, standardized for Internet use in RFC 3629, is one such encoding scheme. A charset label therefore answers a decoding question. It does not say whether the decoded sentence is French, Arabic or Japanese, whether it is well written, or what it means.
The next boundary is easy to lose because screens conceal it. A character is an abstract data element, identified by a name and properties rather than by one fixed appearance. A glyph is a rendered form. A renderer selects glyphs using fonts, shaping rules and context. One character can have several glyphs; several character sequences can produce forms that viewers judge identical; and one visible form may be ambiguous without linguistic context. “It looks right” is an observation about output, not a reversible proof of input.
Script and language do not collapse into each other either. A script is a set of graphic characters used for written forms of languages. One script can serve many languages, while one language can be written in more than one script. Recognizing a script run narrows some rendering choices, but it does not identify the language. Language tags, developed in RFC 3066 and later comprehensively specified by RFC 5646, carry a separate declaration. Even a syntactically valid tag is supplied metadata, not proof that a passage is fluent, correctly classified or understood.
RFC 2277 had already expressed the human purpose of this machinery: protocols themselves do not have a natural language, but the text strings that pass through them are used by people. This matters because “internationalized” cannot mean merely that a parser accepted non-ASCII bytes. The chain continues through input, storage, comparison, search, sorting, rendering, speech or tactile output, and finally a person’s ability to act on the result.
Normalization reveals why appearance is an especially weak shortcut. Unicode may represent canonically equivalent text with different code-point sequences. Unicode Standard Annex #15 defines normalization forms, while RFC 5198 later specifies a preparation profile for network interchange. RFC 3536’s crucial procedural point is that a protocol must say where normalization happens. If a system silently normalizes on entry, another at comparison and a third not at all, identical-looking strings can behave as different keys—or distinct records can be merged without an auditable decision.
That topic stops before the existing RFC 5137 boundary. The question here is not whether a Unicode escape became a normalized identifier or whether two account handles collide. It is the wider representation chain: how received octets acquire character identity, script and language context before a renderer chooses a glyph. Identifier policy is one application of that chain, not the subject of this history.
Ordering supplies another test. Collation is not the same as sorting by numeric code point. Human readers expect ordering shaped by language and convention; two communities using the same script may place the same strings differently. A database can return a perfectly deterministic code-point order and still deliver the wrong index to its readers. The receipt “sort completed” says nothing about whether the chosen collation served the declared language.
Bidirectional text adds a spatial trap. Memory order and display order can differ. A screenshot cannot reliably reveal the underlying sequence, and cursor movement, copying or spoken output may expose an order that the static picture concealed. Rendering itself is not limited to pixels: RFC 3536 includes audible and tactile forms. A visual glyph that looks plausible may coexist with broken speech synthesis or unusable braille output.
The historical status deserves precision. RFC 3536 was never an Internet Standard. In September 2011, RFC 6365 was published as BCP 166 after IETF consensus and public review and explicitly obsoleted RFC 3536. Later, RFC 7997 changed the rules for non-ASCII characters in RFCs themselves. These developments show a vocabulary becoming more operational and publication practice becoming more inclusive; they do not retroactively make every 2003 definition normative.
RFC 3536’s security section says that security is not discussed. That sentence should remain intact. Modern operators can observe that encoding, directionality and confusable presentation affect risk, but they should not attribute a mature threat model to this glossary. The disciplined historical claim is narrower: systems become hard to audit when they erase the boundaries between encoded data, abstract text, linguistic context and rendered outcome.
The complete evidence chain is therefore longer than a screenshot: retain the received octets and charset label; decode them; validate the character sequence; apply a named normalization policy at a recorded stage; preserve language tags and their provenance; identify scripts and directional runs; shape with known engines and fonts; record fallback; render through the intended visual, audible or tactile channel; apply the selected collation; and test whether the intended audience understood the result. A luminous glyph closes only one link.
Sources
- RFC 3536 — HTML
- RFC 3536 — plain text
- RFC Editor information page
- IETF Datatracker record
- IETF Datatracker history
- RFC 3536 errata
- RFC 2277 — IETF policy on character sets and languages
- RFC 3066 — language tags
- RFC 5646 — current language-tag framework
- RFC 3629 — UTF-8
- RFC 5198 — Unicode format for network interchange
- RFC 6365 — terminology, BCP 166
- RFC 7997 — non-ASCII characters in RFCs
- IANA character-set registry
- Unicode Standard Annex #15
- W3C Character Model
- Heng Lu — Running Code Is Primary
- Heng Lu — On Reality Layers
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
