Summary
- RFC 5147 defines URI fragment identifiers for character and line positions or ranges in
text/plainMIME entities. - Fragment evaluation happens on the client after retrieval; the fragment is not part of resource resolution.
- Positions are zero-based, and a position identifies a zero-length boundary rather than a piece of text.
- Character counts depend on decoding and are not byte offsets; a byte-order mark is not counted as a character.
- Every supported line-ending representation counts as one character, so protocol and storage context affect interpretation.
- An out-of-range position clamps to the entity's final position instead of proving that the intended passage still exists.
- Invalid syntax and reversed ranges must be ignored, not repaired or guessed by the client.
- Optional
lengthandmd5parameters compare the decoded MIME entity after transport encodings are removed. - Clients may ignore integrity checks, and a charset mismatch can make a check unusable or transcoding unreliable.
- MD5 can expose accidental drift in this narrow design but is not a signature, issuer credential or modern collision-resistant proof.
- Fragment-only display can hide context or small print and can support spoofing or phishing misunderstandings.
- Consequential citations need separate receipts for revision, bytes, decoding, context, issuer authority and the decision made.
The highlight survived the edit
The dangerous incident begins with a success signal. The document downloads. The fragment parses. The viewer scrolls. Three lines are highlighted. Every local metric says the reference worked.
But a line address is relative to one representation. Insert five lines at the top and #line=80,83 can select what used to be lines 75 through 78. Convert a file, change its line endings, declare a different charset or replace a combining sequence with a precomposed character, and a character position can move even when a reader considers the prose visually similar. A working highlight proves that one client found positions in the entity it decoded. It does not prove that those positions carry the proposition reviewed when the link was created.
RFC 5147 is unusually clear about this boundary. It says that modifications can break a fragment and offers optional integrity information so a client can warn about that possibility. The mechanism makes references more inspectable. It does not promise immutable meaning.
The fragment never reaches the authority
A generic URI contains the fragment after #, but the fragment's interpretation belongs to the media type. For text/plain, RFC 5147 defines that interpretation as a client-side operation after retrieval. The fragment is not used when resolving the URI and is not sent as an instruction that the origin must validate.
This divides the evidence chain. DNS and routing can bring the client to an endpoint. HTTP or another retrieval mechanism can return a representation. A media type tells the client which fragment grammar may apply. Decoding turns octets into characters. Only then can the client count positions and display a range.
No layer in that chain delegates legal or institutional authority merely because the next layer succeeds. A server log may never contain the fragment. An access-control decision at the origin therefore cannot be inferred from a client highlight. Conversely, a trusted origin does not guarantee that the client decoded, counted or displayed the intended material.
Unsupported clients expose another limit. RFC 5147 was designed for incremental deployment: a client that does not understand the new fragment can still retrieve the resource. The link appears alive while the sub-resource selection silently disappears. Availability of the whole file and support for the fragment are separate outcomes.
A position is not the text
The syntax combines characters or lines with positions or ranges. A position is a zero-length point between two characters or lines. Position zero comes before the first unit. A range spans from its starting position to its ending position. That distinction matters when an application turns a caret location into a quoted passage.
Character counting also begins only after decoding. UTF-8, UTF-16 and other encodings can use different numbers of octets for the same conceptual text. A leading byte-order mark does not count as a character. RFC 5147 uses code-point counting rather than bytes or user-perceived grapheme clusters. A glyph that looks like one letter may be represented by more than one code point.
Line handling adds a second normalization boundary. Internet text/plain conventionally uses CRLF, but HTTP, operating systems and local files can expose LF, CR and other conventions. The specification requires each recognized line ending to count as one character regardless of its representation. That rule improves interoperability only if the client recognizes the relevant convention.
These are not academic edge cases for evidence systems. A policy archive may normalize line endings. A gateway may transcode. A browser may infer a charset that the origin omitted. A litigation export may add a byte-order mark. If the assurance record retains only the displayed quote and the fragment string, it cannot reconstruct which entity and counting rules produced the selection.
Failure can collapse to the end of the file
RFC 5147 gives out-of-range positions a defined behavior: they identify the final character or line position of the entity. That makes processing deterministic, but it is not a positive verification of the reference.
Imagine that a document was truncated from 900 lines to 120. A link to line 760 can now land at the end. A viewer that reports only “fragment handled” may turn catastrophic content loss into a successful navigation event. The intended passage is gone; the syntax still has an answer.
Open-ended ranges have similar operational consequences. A missing first number extends from the beginning, and a missing second number extends to the end. These are useful forms, but a growth event can expand the selected material far beyond what a reviewer saw. The stable fragment does not freeze the selected bytes.
Some failures are intentionally stricter. If range positions are reversed, the fragment must be ignored. Syntax errors must also be ignored, and the client must not guess a correction. This is a valuable minimum rule: a convenient repair would create an unrecorded interpretation. Yet an ignored fragment can still leave the whole resource visible, which makes telemetry and user interfaces responsible for distinguishing “document opened” from “target selected”.
Integrity checks detect drift; they do not appoint a speaker
RFC 5147 permits length and md5 parameters after the positional scheme. They apply to the MIME entity after content encodings and content-transfer encodings have been removed. Length is expressly a weak check. MD5 produces a stronger accidental-change signal for the design's historical context.
Neither parameter is mandatory for clients. A client may ignore integrity information altogether. If it implements a check and detects a changed entity, it should not interpret the fragment and may warn the user. “Should” and “may” leave room for different product behavior. A document workflow cannot assume that every viewer enforces the same stop.
Charset qualification creates another fork. If the fragment names a charset, the client must compare it with the retrieved entity's charset and must not apply the check when they differ. It may transcode and then check, but RFC 5147 warns that this can be unreliable because characters or sequences may be lost or normalized.
MD5 must also be kept in its lane. RFC 6151 later concluded that MD5 is not prudent where collision resistance is required and is unacceptable for digital signatures. Even without a crafted collision, an unkeyed digest says nothing about who supplied the entity or whether that party could issue the policy. A matching RFC 5147 MD5 parameter can help answer “is this the same encoded entity?” It cannot answer “is this the authorized revision?”
Media type is the dispatch key
Fragment semantics are not universal. RFC 5147 applies when the retrieved entity is text/plain. Some URI schemes or retrieval environments do not provide an explicit media type, so clients use heuristics. The RFC calls such processing inherently unreliable.
This matters in archives and desktop workflows. A local file named .txt may be decoded under platform assumptions. An FTP transfer may not supply the same metadata as an HTTP response. A file URI can lead to a representation whose line-ending and charset conventions differ from the network copy. The fragment string is unchanged while the processing contract changes.
IANA records the RFC against the text/plain media type. That is coordination evidence: the syntax has a documented home. It is not telemetry about which user agents implement it, how they report mismatches or whether any particular product preserves context. A registry entry must not be promoted into an adoption metric.
Selection can manufacture confidence
The security section anticipates the social problem. Implementations can show different fragments, or one can support the mechanism while another ignores it. A product that displays only the selected region can hide surrounding legal language or present site-key-like text as if it came from a trusted context.
The exploit does not require breaking URI resolution. It relies on a reader confusing a bounded view with the whole statement. A genuine document can be selectively framed. A correct fragment can still be misleading. A valid retrieval can sit inside a deceptive inclusion surface.
For evidence, context is therefore an input, not decoration. Store the full entity or an immutable revision, the exact selected text, enough surrounding material to interpret it, the declared and effective charset, the line-ending policy, the fragment, the client version and the time of retrieval. If a decision depends on issuer authority, retain the signature or provenance chain separately.
Sources
- RFC 5147, HTML
- RFC 5147, text
- RFC Editor record for RFC 5147
- IETF Datatracker record for RFC 5147
- RFC 5147 document history
- RFC 5147 errata search
- RFC 6838: Media Type Specifications and Registration Procedures
- RFC 2046: MIME media types
- RFC 3986: URI generic syntax
- RFC 3987: Internationalized Resource Identifiers
- RFC 3629: UTF-8
- RFC 1321: MD5
- RFC 6151: MD5 security considerations
- RFC 9110: HTTP semantics
- RFC 8089: file URI scheme
- IANA media types registry
- W3C Media Fragments URI 1.0
- Minimum Initial Specification, Localized Future Decision, and Voluntary Adoption
- On Reality Layers, Symbolic Power, and Why Clarity Feels So Hostile
- Running-Code Primacy
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
