Summary
- RFC 2987 registered
charsetandlanguageas equality-tested media features, but called the first usually a device capability and the second usually a user preference. - Principal lowercase charset names and whole-token language matching reduced ambiguity without proving decoding, rendering, speech quality or comprehension.
- A quality weight ranked a predicate inside a local decision; it was not a receipt for the selected representation, the user's satisfaction or the service outcome.
Two neighbouring registrations
RFC 2987 is a short document. Most of its pages are registration forms. That apparent modesty makes its central decision easy to miss.
The first form registered charset, OID 1.3.6.1.8.1.31, as the ability to display particular registered character sets. The second registered language, OID 1.3.6.1.8.1.32, as the ability to display particular human languages. Both values were tokens. Both could be compared only for equality. Both comparisons ignored case. Both examples used an OR expression with q weights.
The forms looked parallel because RFC 2506 had created a transport- and protocol-neutral namespace for media features. A negotiation mechanism could carry assertions about presentation without inventing new names for every application. RFC 2533 supplied an algebra for combining feature predicates and attaching quality values. Registration made the vocabulary portable.
But portable vocabulary was not portable meaning by itself. RFC 2987 inserted a semantic asymmetry directly into the registrations.
For most devices, it said, charset was usually a capability. A device could not intelligently process text in a charset it did not know. language, by contrast, was usually a preference rather than a requirement. The document contemplated “display” of language most often as computer speech, but the allocation is broader: language choice normally expresses what a user wants, not a physical fact about what the machine can parse.
One grammar therefore carried two different kinds of constraint.
A decoder limit is not a vote
The charset example ranked utf-8 at 1.0, iso-8859-1 at 0.9 and utf-16 at 0.5. It is tempting to read those numbers as three tastes. The registration warns against that reading.
If a device lacks the selected decoder, there may be no meaningful output to rank. The bytes do not become text because a predicate won. A font may still be absent after successful decoding. A renderer may still fail. A valid glyph sequence may still be unintelligible to the user. But the first boundary is harder than preference: the decoder must exist and accept the received representation.
RFC 2913 had registered type as a separate media feature. Keeping charset outside the media-type token made that boundary visible. A system might support text/plain but not every encoding of text. Representation type and character encoding were related claims, not one receipt.
RFC 2987 also asked authors not to use charset aliases in feature expressions. RFC 2978 allowed a charset to have several registered names, but required one principal name and made each name identify only one charset. A manipulation tool could canonicalise an alias to that principal name. If one side ranked an alias while the other ranked the principal spelling, rewriting could unexpectedly change the expression it appeared to preserve.
Principal lowercase names were therefore an interoperability discipline. They did not certify the decoder. Canonical spelling reduced a naming disagreement; it did not prove that the running code implemented the named mapping correctly.
The registration added another guardrail: a charset assertion should accompany any capability to handle textual data. Claiming “text” without naming the encoding surface left the receiver unable to test the most basic precondition for turning bytes into characters.
A language preference is not a machine wall
The language example ranked no-nynorsk, no-bokmaal and i-sami-no. Here the same equality operator served a different purpose.
RFC 1766 built language tags from a primary tag and optional subtags, but told applications to treat the complete tag as one token. Subtags were administrative. They were not a general-purpose inference tree, and a shared prefix did not guarantee mutual comprehension. RFC 2987 repeated the operational rule: compare the whole token, case-insensitively, and do not use subtags in the comparison.
That rule prevented a negotiator from inventing linguistic authority from string structure. Matching no inside two longer tags did not prove that either choice satisfied a particular reader or listener. A language tag named an option under a registration system. It did not measure literacy, dialect familiarity, pronunciation, accessibility or cultural fit.
Yet language was usually a preference, not a requirement. Treating a failed equality test as a device incapacity could discard content the user would rather receive than nothing. Treating the preference as meaningless could repeatedly deliver a language the user could not use. The correct response depended on local policy, available variants and the user's stated priorities.
The feature expression did not settle that policy. It made the input inspectable.
What the quality value could not certify
RFC 2533 allowed quality values to qualify predicates and gave an unqualified predicate a default quality of 1. It deliberately did not define one universal way to combine every preference. Different applications could need different choice rules.
That limit matters. A q=1.0 attached to one language is evidence that a profile ranked that predicate highly in a particular expression. It is not evidence that the profile is current, that content exists in that language, that a speech engine sounds intelligible, that a reader understands the result or that the task succeeds.
The same is true for charset with a harder edge. A quality value can order supported alternatives only after support is real. It cannot install a decoder. It cannot prove that the selected payload's label matches its bytes. It cannot show that a gateway preserved those bytes or that a rendering path retained every character.
A complete operational record therefore needs more than the winning expression. It needs the received feature set, normalization result, evaluator policy, available candidates, selected representation, runtime decoder or speech resource, render result and—where the service matters—the user's outcome.
These are not redundant logs. They answer different questions.
A registry can name the question
The IANA media-feature registry still lists charset and language as entries 31 and 32 in the IETF tree. That continuity shows that RFC 2987 performed a durable coordination task. Independent systems could refer to the same names and equality semantics without one application owning the vocabulary.
Registration did not create a central evaluator. The sender or device profile controlled what it declared. The representation supplier controlled which variants existed. The receiving implementation controlled canonicalization, matching, weighting and fallback. The user bore the effect of the choice.
This division resembles Lu Heng's Minimum Initial Specification: standardize the minimum shared language needed for interoperability, then keep future decisions local unless a common rule is necessary. RFC 2987 fixed tag names, value domains and comparison semantics. It did not appoint a universal ranking authority.
Local freedom, however, did not eliminate responsibility. The implementation that treated one predicate as mandatory and another as advisory controlled the decisive interpretation. If it silently softened a charset limit, it risked unreadable output. If it silently hardened a language preference, it risked needless denial. Its policy needed a receipt.
Running-Code Primacy supplies the next test. A standards registration is a shared specification, not evidence that a deployed device has the decoder, content, font, voice, fallback path or user model it claims. Publication can make an expression valid. Only execution can show what was selected and what worked.
The reality layers remain separate:
- the feature tag and value are registered;
- an expression declares a predicate and weight;
- aliases and case are normalized;
- complete tokens are compared;
- local policy chooses an alternative;
- the required decoder or language resource exists;
- the representation is executed and presented;
- the user perceives and understands it;
- the intended task succeeds.
No earlier layer has authority to testify for all the later ones.
The small security disclosure
RFC 2987 included a narrow security warning for each tag. If a display bug was known for a particular charset or language environment, revealing that the device accepted it might slightly help an attacker. The document did not identify a product or incident, and neither should historical analysis invent one.
The warning nevertheless reinforces the main point. A capability declaration is operational data. It can shape selection, reveal attack surface and create expectations. The safer system does not collect or expose more than the decision needs, and it does not confuse disclosure with proof that the advertised path is safe.
RFC 2987's lasting contribution is not that it made charset and language equal. It is that it let them share machinery without erasing their difference.
The tags looked the same. The comparison looked the same. The weights looked the same.
One answer usually said, “the machine can or cannot do this.” The other usually said, “the person would rather have this.”
Interoperability depended on preserving both the common grammar and the different authority of those two statements.
Sources
- IETF Datatracker history for RFC 2987
- Lu Heng — Minimum Initial Specification, Localized Future Decision, and Voluntary Adoption
- Lu Heng — Reality layers and symbolic power
- Lu Heng — Running-Code Primacy
- IANA Media Feature Tags registry
- RFC Editor errata for RFC 2987
- RFC Editor information for RFC 2987
- RFC 1766 — Tags for the Identification of Languages
- RFC 2277 — IETF Policy on Character Sets and Languages
- RFC 2506 — Media Feature Tag Registration Procedure
- RFC 2533 — A Syntax for Describing Media Feature Sets
- RFC 2913 — MIME Content Types in Media Feature Expressions
- RFC 2978 — IANA Charset Registration Procedures
- RFC 2987 — Registration of Charset and Languages Media Features Tags
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
