Summary

  • RFC 2343 made a bundled MPEG-2 programme legible to RTP with picture boundaries, a 90 kHz timestamp, audio length and a signed audio offset.
  • Those fields described the sender's packetization claim; receipt, reconstruction, decoding, lip synchronization, output quality and actual viewing remained separate events.

One packet could look like a whole programme

The seductive object in RFC 2343 was a packet that appeared self-sufficient. It could contain integral MPEG video slices followed by integral audio frames. Its RTP header named a payload type, a sequence position, a picture time and a picture endpoint. A short BMPEG header said what kind of picture was involved, whether header state had changed, how many audio bytes were present and how far their start lay from the picture timestamp.

For a video-on-demand server in 1998, that compression of responsibility was attractive. A programme could use one port rather than separate ports for sound and picture. Interleaved content on disk could stay interleaved on the wire. Shared packet overhead could fall. The receiver might need less buffering than two independently delayed streams, and the sender could govern the programme's aggregate bandwidth more directly.

The memo called the resulting relationship implicit synchronization. That was a property of the transport representation, not a receipt from a television set. The packet placed audio and video in one temporal account. It did not prove that every later component would honour that account.

Bundling was a deliberate trade

RFC 2343 was Experimental, not an Internet standard. Its abstract was candid: bundling could be useful when its advantages justified sacrificing the modularity of separate audio and video streams.

That sacrifice changed the failure surface. One port simplified session handling, but it also tied the fate of the two media more closely together. A lost bundled packet could damage both picture and sound. A receiver that wanted only one medium could not treat the programme like two independent RTP streams. A packetization choice that reduced headers did not reduce the importance of path MTU, decoder state or output clocks.

The format also removed MPEG systems-layer information it considered redundant with RTP. Efficiency came from assigning those duties elsewhere. When information disappears from one layer because another layer is expected to carry it, the operational question is not whether the bytes are fewer. It is whether every participating implementation shares the same division of responsibility.

The packetizer had to respect application boundaries

The format imposed structure. A sequence header, when present, began an RTP payload. A GOP header either began the payload or followed the sequence header. A picture header either began the payload or followed the GOP header. Each packet contained an integral number of video slices.

Those rules gave a receiver known places at which it might resume. They did not make packets immune to lower-layer fragmentation. RFC 2343 left the sender responsible for adjusting slice size and packet composition to stay under the path MTU. If it failed, lower layers could fragment the packet, with consequences for loss amplification and traffic classification.

The audio rule was equally precise and equally bounded. Video data was followed by enough complete audio frames to cover the duration of the video segment. An audio frame could cover longer than the video in one packet, so subsequent packets might legitimately contain no audio. A sender could repeat the latest audio frame to improve resilience.

A syntactic checker could confirm those boundaries. It could not tell whether the chosen MTU still matched the path, whether a fragment vanished, whether repeated audio was perceptually acceptable or whether the receiver had enough data to maintain a continuous clock.

Time was described, not delivered

The RTP timestamp was a 32-bit value on a 90 kHz clock. It represented the sampling time of the MPEG picture and was shared by all packets belonging to that picture. The Marker bit identified a packet containing the end of a picture.

Neither field was a progress bar. B pictures made timestamps legitimately non-monotonic because transmission order and presentation order could differ. A packet carrying only sequence, extension or GOP headers borrowed the time of the subsequent picture. Sequence numbers described transport order while timestamps described media time.

The BMPEG audio offset connected these two domains. It was a signed count of audio samples between the start of the audio frame and the RTP timestamp of the packet. At 44.1 kHz its range was roughly plus or minus 750 milliseconds. The memo even warned that this might be insufficient at a very low video frame rate such as one frame per second.

With B pictures, audio was not reordered alongside the pictures. It stayed in transmission order and relied on the offset to state when it belonged. A correct offset was therefore a coordinate. It was not evidence that the receiver decoded the right picture, scheduled the audio sample, kept the output clocks aligned or presented both at the intended instant.

Loss exposed the difference between coordinates and recovery

RFC 2343 gave receivers useful clues after loss. Sequence numbers and timestamps could reveal a gap. Slice number and the first macroblock position could help estimate the damaged region. When only slices from one picture were lost, a decoder might continue and repeat pixels from an earlier picture. Missing audio could be replaced with a suitable frame, such as background noise, to conceal damage and preserve lip synchronization.

These were recovery options, not restoration receipts. Repeated pixels were not the missing pixels. Inserted audio was not the missing sound. Concealment could make a defect less noticeable while leaving the underlying loss untouched.

The limits became sharper around header state. The N bit said whether sequence, extension, GOP or picture-header data had changed since previously sent headers. If state had changed, discarding data until a new picture start could be safer unless the receiver obtained headers by another channel. Heavy losses could require waiting for a new sequence header.

RFC 2343 omitted a dedicated temporal-reference picture counter. A lost GOP header could go undetected and cause incorrect decoding of following B pictures in some edited material. The payload could therefore remain parseable while its decoder context had already diverged.

Seven receipts lay beyond format conformance

A production system needed at least seven distinct observations after packet creation.

First, did the packet arrive intact and before its deadline? Second, did all lower-layer fragments and RTP packets needed for the picture arrive? Third, did the decoder have the correct sequence, GOP and picture state? Fourth, did audio and video clocks turn the timestamp and offset into aligned presentation times? Fifth, did the renderer produce valid frames and samples? Sixth, did the screen and audio path expose them without mute, occlusion, device failure or background suspension? Seventh, was a viewer present and able to perceive an acceptable programme?

No earlier receipt could silently stand in for a later one. RTP itself did not reserve resources or guarantee quality of service. A valid packet could be late. A complete packet could depend on missing earlier state. A decoded frame could be dropped. A synchronized render could be hidden. A lit screen could face an empty chair.

RFC 2343's historical value lies in making one layer unusually clear. It specified how a sender could express a bundled audio/video relation. The clarity of that expression is precisely why its evidentiary boundary can be drawn: the packet was a statement about packaging, not a witness to playback.