Summary
- RFC 1952 defines a gzip file as consecutive members, each with its own header, compressed blocks and CRC32/ISIZE trailer; a later member can extend the decoded output without rewriting earlier members.
- Optional names, comments, timestamps and extra fields describe one member and can be skipped. They are not authenticated provenance, and the format remains stream-oriented rather than randomly indexed.
- CRC32 detects corruption in one member's recovered data and ISIZE records that member's length modulo 2^32. Neither field identifies an appender or states the absolute size of an arbitrary multi-member file.
The byte after the apparent ending
Suppose a program has written a valid gzip stream and closed it. Hours later, another program opens the same file for append and writes a second valid gzip stream. Nothing in the first stream changes. There is no new outer header, no rebuilt table of contents and no pointer inserted near byte zero. Yet a decoder reading from the beginning can recover the first input and then the second.
This is not a tolerant accident. RFC 1952 defines a gzip file as a series of members appearing one after another without additional information between them. A member is the unit that closes: header, compressed blocks, then a trailer. The file can continue because the decoder knows where one unit has finished and where the next recognizable unit begins.
The RFC Editor record dates P. Deutsch's document to May 1996 and classifies it as Informational. That status matters. The specification records a portable format and compliance rules; it is not itself a standards-track declaration or a census of every decoder that later encountered gzip.
The design goals were practical and narrow. The format was to move across processors, operating systems, file systems and character sets; work as a stream with bounded intermediate storage; remain implementable without patented techniques; and match the gzip format already in use. Random access was explicitly outside the design. A sequence can be scanned reliably without becoming an index.
One member, four different kinds of fact
The fixed header starts with two identification octets. A compression-method field follows; method 8 denotes DEFLATE. Flags announce optional material, while reserved flag bits must remain zero. These facts let a decoder recognize a member, choose the right decompressor and traverse the header.
The optional fields answer softer questions. MTIME can preserve a modification time, but zero means no timestamp is available. FNAME can carry an original name. FCOMMENT can carry a human-readable comment. FEXTRA can contain length-delimited subfields. FTEXT offers a tentative indication about text. A decoder must be able to skip the declared optional material even if it does not use it.
That distinction is architectural. A filename may be useful; it is not the member's identity. A comment may explain custody; it does not prove custody. An operating-system code can guide local conversion; a compliant decoder may ignore it. Because these fields are not authenticated, an investigator should treat them as claims carried by the object, not as facts independently established about its origin.
If FHCRC is present, it checks the preceding header using the low sixteen bits of a CRC32 calculation. The trailer's CRC32 is different: it covers the uncompressed data of that member. Beside it, ISIZE stores the member's original input size modulo 2^32. Header traversal, recovered-data integrity and descriptive metadata occupy separate surfaces even though they sit a few bytes apart.
The check ends where the member ends
The trailer makes corruption observable at a useful boundary. A decoder can inflate one member, compute its CRC32 and compare both checksum and size. If the comparison succeeds, it has evidence that the recovered bytes match those checks under the format's accidental-error model.
It does not have a signature. Anyone able to replace or append data can calculate a new CRC32. It does not have an unbounded length statement either: ISIZE deliberately wraps after 2^32. In a multi-member file, each trailer speaks only for its member. Summing recovered lengths is an aggregate operation performed by the reader; it is not a total written into the first or last header.
The compression layer is also separate. RFC 1951 specifies DEFLATE's blocks and coding. RFC 1950 places DEFLATE in the different zlib wrapper, with its own header and Adler-32 check. The later media-type registration in RFC 6713 states the distinction plainly: zlib is a stream format, while gzip adds a header and trailer suited to files.
That separation prevented the compression algorithm from having to carry every file-level decision. A gzip member can change optional metadata without inventing a new DEFLATE grammar. Another wrapper can use the same compressor with different validation and framing. Compatibility lives at named boundaries rather than in one undifferentiated “compressed file” concept.
Running code makes the continuation visible
The current GNU gzip manual documents concatenation as an operating technique. Several compressed files can be joined, and gunzip emits all members. Appending the compressed form of a second input therefore produces the same recovered byte sequence as concatenating the two original inputs. The manual also records two costs: compressing the inputs together usually achieves better compression, and gzip --list reports the uncompressed size and CRC of the last member rather than an aggregate for all members.
Those details are more than command-line trivia. They show how one correct container can present different truths to different tools. The decompressor follows the sequence. A listing command may summarize only its final unit. A scanner that stops after the first trailer sees less than a conforming stream consumer. “The file is valid” is incomplete unless the speaker names which members were examined and which decoder policy was used.
HTTP gives the format another setting but not another ontology. RFC 9110 registers gzip as a content coding by reference to RFC 1952. That permits gzip-coded representation data; it does not prove that every intermediary, server, client or security product preserves and reports member boundaries in the same way. The source set supports the format rule, not a universal HTTP-behaviour claim.
A minimum specification can leave meaning local
Running-Code Primacy offers a later discipline for reading this history: the published document is one reality layer, while implementation, validation and use add others. RFC 1952 defines the common syntax. GNU gzip supplies bounded evidence that concatenated members operate in one maintained implementation. Neither fact should be inflated into a claim about every system.
Minimum Initial Specification sharpens the virtue of the member boundary. The common layer is strict where independent programs must agree—recognition bytes, compression method, flags, optional-field lengths, reserved bits and trailer checks. It can remain permissive about whether a local program displays a comment, restores a timestamp or uses an OS hint.
Reality Layers helps separate the symbol from the operation. A single filename is a convenient symbol. Operationally, the bytes may contain several closure events, several checksums, several size residues and several metadata claims. The label is not false; it is simply coarser than the executable structure.
RFC 1952's durable achievement was not to make compression mysterious. It made the stopping point small and repeatable. Because one member could finish completely, the file could begin again.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
