Summary

  • RFC 3548 addressed a deceptively small interoperability problem: specifications often said “base64” without naming the alphabet, line-break, padding or rejection rules they intended.
  • Its central lesson was that MIME’s mail-transfer behavior is one profile, not a universal decoder contract; a protocol should say which profile it uses.

Text editors can open a Base64 string. That convenience can make the conversion look self-explanatory: bytes go in, printable characters come out, and any decoder will reverse the process. The history recorded in RFC 3548 says otherwise. Implementations had accumulated small differences, while protocols invoked “base64” as though the name settled all of them.

The mismatch was not about the six-bit arithmetic. It was about the envelope around it. Should an encoder insert line breaks? Must the final = padding be present? Does a decoder reject a stray character, or ignore it? Which 64 symbols occupy the alphabet? A system that answers these questions by borrowing a nearby format may work in that format and still disagree with a peer expecting something else.

RFC 3548, published in July 2003 as an Informational RFC, set out to reduce that ambiguity. Its introduction singles out a recurring shortcut: protocol documents used “base64” without a precise description or reference, often treating MIME as the default without considering line wrapping or non-alphabet characters. The document therefore collected common Base16, Base32 and Base64 schemes and made the surrounding choices visible.

MIME was an important source of confusion precisely because it was specific. RFC 2045 defines Base64 as a Content-Transfer-Encoding for message bodies. Its 76-character line limit belongs to that mail context; PEM had used 64-character lines, also in a mail-era setting. RFC 3548’s rule for a generic referring specification was not to add line feeds unless that specification asks for them. A line break is not decoration when another parser might count it as data or reject it.

Padding and tolerance were also profile decisions. RFC 3548 says encoders include appropriate padding unless the referring specification says otherwise. Decoders reject characters outside the alphabet unless a referring specification explicitly chooses another policy. MIME may ignore such characters, including CRLF, but that exception belongs to its defined behavior. Copying it into a different protocol changes what inputs are accepted. The RFC discusses covert channels and implementation errors as reasons for care; it does not report a particular attack caused by a named implementation.

The alphabet itself also needed a name. The conventional Base64 alphabet uses + and / for values 62 and 63. RFC 3548 documents a URL- and filename-safe alternative using and _. It says this is not the same encoding and should not be called only “base64.” That distinction matters because a string moved into a path, filename or identifier field meets rules that a mail body does not.

In October 2006, RFC 4648 obsoleted RFC 3548 as a Standards Track document. It retained the profile-centered approach and added a canonical-encoding section: pad bits in Base64 and Base32 output must be zero, or multiple strings can decode to the same bytes. That successor point is a separate issue from MIME line wrapping. Together, the two RFCs show why “decode succeeds” is not a complete protocol specification: the sender and receiver still need agreement about spelling, padding, tolerance and the layer at which a value is interpreted.

Base encoding is not encryption and does not authenticate a message. Its value is a transport-friendly representation of octets. The historical work of RFC 3548 was more modest and more durable: it made protocols stop treating a familiar label as a complete set of rules.

Sources