Summary

  • RFC 2035 did not send a complete JPEG file in every video frame. It omitted table and marker structures that were stable or inferable, then repeated a compact eight-octet RTP/JPEG header carrying Type, Q, Width, Height and a 24-bit Fragment Offset with each packet.
  • Timestamp, sequence number, offset and marker answered different reconstruction questions: frame membership, stream order, byte placement and declared end. Together they exposed gaps and supported out-of-order placement; none alone proved a complete frame.
  • Restart markers could bound decoder state, and Types 4 and 5 aligned restart intervals with RTP packets to permit limited decoding after loss. That opportunity did not prove pixel correctness, deadline arrival, visual continuity or quality of experience.

A packet can arrive second and still know that it belongs first. That is the quiet operational power of RFC 2035’s Fragment Offset. The receiver need not confuse network arrival order with image order: the 24-bit field says where the packet’s scan bytes belong inside a JPEG frame. It is a coordinate, not a verdict.

That distinction captures the whole design. Published in October 1996, RFC 2035 carried JPEG-compressed video over RTP by refusing to treat a frame as a file that must be reproduced verbatim on the wire. Much of a conventional JPEG interchange stream describes tables, dimensions and coding choices. When those descriptions remain stable—or can be generated by an agreed rule—repeating them in every frame consumes bandwidth without adding new picture information. RFC 2035 removed that repetition and made the receiver reconstruct the necessary syntax from a small header attached to every RTP packet.

The saving was real, but it created a new evidence contract. Once the full JPEG wrapper disappears, the receiver must know exactly which declarations replace it, where scan bytes go, when a frame begins and ends, and where decoder state can safely restart. RFC 2035 supplies those coordinates. It does not promise that the picture behind them is intact.

The missing header was a protocol decision

The payload format targets a restricted JPEG surface: sequential DCT, a single interleaved scan and a deliberately small set of sampling arrangements. The restriction matters. It lets comparatively constrained hardware codecs meet on a common representation rather than forcing every endpoint to implement the full range of baseline JPEG possibilities.

An RTP/JPEG payload begins with entropy-coded scan data, not with the familiar sequence of JPEG marker segments and table definitions found in a complete interchange image. Restart markers are the exception. Frame and scan header information is represented instead by an eight-octet header in every packet. That header includes a type-specific byte, a three-byte Fragment Offset, Type, Q, Width and Height.

Type identifies the agreed JPEG interpretation. Q identifies the quantisation rule. Width and Height express dimensions in units of eight pixels. Their presence is not bureaucratic duplication. RFC 2035 anticipates that a source may adapt rate by changing quality or resolution, so a receiver cannot safely infer these values once for the lifetime of a session. For Q values from 1 through 99, the document defines how quantisation tables are derived. Values from 100 upward depend on dynamically supplied tables through a session mechanism outside RFC 2035; zero is reserved.

These fields are therefore reconstruction inputs. If a packet says Q=50 and a particular width and height, the receiver has locally observable declarations with which to build omitted JPEG syntax. It does not have proof that the encoder used the declared tables correctly, that the compressed scan is uncorrupted, or that the resulting pixels correspond to the scene a camera captured.

Four signals, four questions

Every fragment of one frame carries the same RTP timestamp. RTP sequence numbers expose the order in which packets were sent and make missing positions in that packet stream observable. Fragment Offset gives the byte position of this payload within the frame’s JPEG scan. The RTP marker bit identifies the packet declared to end the frame.

It is tempting to compress those signals into one dashboard flag called “frame received”. That destroys the useful distinctions. A shared timestamp groups packets but says nothing about completeness. A sequence gap identifies a missing RTP packet but does not itself say which scan-byte range is absent. A Fragment Offset gives placement but does not authenticate the bytes placed there. A marker supplies an end declaration but cannot prove that every earlier fragment arrived.

Used together, the fields support a much stronger local test. A receiver can allocate a frame buffer, copy each fragment directly to the offset it declares, compare the next offset with the preceding offset plus payload length, and flag gaps or overlaps. Arrival order stops being a reconstruction dependency. This is why the RFC prefers RTP-level fragmentation over leaving the work to IP. If one IP fragment disappears, the IP layer may withhold the entire datagram, including intact bytes. Separate RTP packets preserve independent delivery and give the sender finer control over packet size and rate.

The document also exposes an awkward boundary case. If the last packet of a frame is lost, its marker goes with it. The next frame’s zero-offset packet becomes essential evidence that a new frame has begun; continuity through to that frame’s marked final packet can then establish an intact successor. The lost marker and the lost scan bytes are different absences. A useful capture preserves both facts.

Restart points reduce dependency, not damage

JPEG entropy coding carries state. Huffman decoding and DC prediction normally make later data depend on earlier data. Restart markers reset that state, creating points at which a decoder can begin again. RFC 2035 assigns different Type behaviours around those markers.

Types 0 and 1 use no restart markers. Types 2 and 3 permit them without aligning RTP packet boundaries to restart intervals. Types 4 and 5 go further: restart intervals align with packets, and a type-specific restart count describes whether a packet is at the beginning, middle or end of the interval sequence. A receiver that has an intact aligned interval may decode that interval without possessing every earlier scan byte.

“May” carries the argument. The format creates a bounded recovery opportunity. It does not invent the missing image region, guarantee that an intact interval contains useful subject matter, or ensure the packet arrived before its playout deadline. Partial decode is a decoder capability conditioned on observed structure. It is not packet-loss concealment and not a service-level promise.

This is also where historical accuracy needs discipline. RFC 2435 later obsoleted RFC 2035 and changed the payload format. The current IANA registry cites that successor for static payload type 26. Those facts describe the present standards record. They do not let us import RFC 2435’s independent optional headers or revised restart mechanism into the older specification.

Reconstruction evidence stops at the decoder door

An operator can prove quite a lot from a packet capture: which packets declared the same timestamp, how sequence numbers progressed, which scan ranges their offsets covered, whether the marked end appeared, which Type and Q were declared, what dimensions were announced, and whether restart-aligned intervals were structurally intact. A decoder log can add whether reconstructed syntax was accepted and which regions were produced.

Each added observation is a new layer. A legal JPEG reconstruction does not prove lossless pixels; JPEG is lossy by design. A complete frame does not prove it met a playout deadline. A timely decode does not prove continuous motion. A rendered buffer does not prove that a viewer saw, understood or valued the image. Modern “quality” scores are further interpretations again and lie outside this RFC.

Heng Lu’s running-code discipline is useful here because the protocol’s best claims are local and falsifiable. Preserve the packet bytes. Preserve timestamp, sequence, marker and offset separately. Record the exact reconstruction rule and decoder result. Then refuse to make the symbolic statement—“the picture arrived”—carry more weight than those events support.

RFC 2035’s economy came from a thin common layer: omit what can be reconstructed, repeat the minimum context whose changes matter, and keep placement visible above lower-layer fragmentation. Its honesty comes from the same thinness. The packet can tell a receiver where it belongs. The evidence for a picture must still be assembled, layer by layer.

Sources