Summary

  • RFC 3952 carried one or more fixed-size iLBC frames without an in-band frame count; the receiver divided the RTP payload length by 38 or 50 according to the mode agreed through SDP.
  • A 2026 reported erratum changes one section's 32/50 divisor to 38/50; the conflict shows why packet bytes, active negotiation state, decoder acceptance and audible playout are separate receipts.

A count removed, not a fact eliminated

An RTP packet can arrive with an exact length, an orderly sequence number and a plausible timestamp while still leaving one essential question outside itself: how should its payload be divided? RFC 3952 answered that question for the internet Low Bit Rate Codec, iLBC, by avoiding a separate count field. Frames were concatenated. Their number was calculated.

The codec supplied two fixed sizes. A 20-millisecond frame occupied 38 octets; a 30-millisecond frame occupied 50. Once the receiver knew the mode, arithmetic recovered the boundary: payload length divided by the expected frame size yielded the number of frames. A 100-octet payload under mode 30 meant two frames. A 76-octet payload under mode 20 also meant two. The saved field did not make the count unknowable. It made the count dependent on prior agreement.

That distinction is the design's historical interest. The packet contained evidence of length. The session description contained evidence of mode. Only their join produced a frame count. Neither fact could substitute for the other.

The mode lived in the control plane

RFC 3952 mapped the codec to a dynamic RTP payload type with iLBC/8000. The mode travelled separately in SDP as an a=fmtp parameter. If it was absent, 30 milliseconds was the default. The document emphasized that the result was bidirectional: both ends had to use the same mode.

Its offer/answer rule is easy to misread if “20” and “30” are treated merely as durations. The offerer named a preference; the answerer could name either value; the common result was the lower-bandwidth mode. Because 30-millisecond frames use 50 octets rather than 38 every 20 milliseconds, both mixed examples resolve to mode 30. The signalling decision therefore selected not only a cadence but also the divisor by which every later payload would be interpreted.

The wire rules reinforced the dependency. A frame could not be split between RTP packets. Frames from the two modes could not be mixed inside one packet. Aggregation was bounded by the transport MTU, while the number of frames per packet balanced header overhead against delay and the amount of speech lost with one datagram. These constraints made length arithmetic workable; they did not make it self-describing.

Thirty-two appeared where thirty-eight belonged

Sections 2 and 3.1 of RFC 3952 identify the 20-millisecond representation as 38 octets. RFC 3951, which defines the codec, is consistent with that packetized size. Yet section 3.2 says that a receiver divides by the expected octets per frame, printed as 32/50.

The RFC Editor errata record now contains Errata ID 8866, submitted on 3 April 2026, proposing 38/50. Its status is Reported. That status matters. A reported erratum is evidence that a conflict has been identified and a correction proposed; it is not the same as a formally verified correction. The internally consistent reading is strong, but editorial confidence must not be mislabeled as process completion.

The typo is revealing precisely because the payload omitted the count. A receiver implementing the surrounding rules would know that 38 is the frame size. A tool copying only the isolated sentence could inherit 32 and reject sound packets, invent impossible boundaries or produce misleading diagnostics. The packet bytes would remain unchanged. The interpretation layer would be wrong.

Divisibility was necessary, never sufficient

Suppose the active mode is 20 milliseconds and a payload has 114 octets. Division by 38 yields three frames with no remainder. That is a useful framing receipt. It does not prove that each 38-octet block is valid iLBC, that the sender was authorized, that SRTP authentication succeeded, that the jitter buffer admitted the packet, that decoding produced acceptable speech or that anyone heard it.

The converse also needs care. A length inconsistent with the active divisor is strong evidence of malformed media, stale negotiation state, payload-type confusion or capture error. But some lengths can be divisible by both 38 and 50. Arithmetic alone then cannot reconstruct intent. The active offer/answer state, payload type and packet time must remain available.

RFC 3550 gives RTP sequence and timestamp semantics for ordering and timing. RFC 2327 supplies the session-description form, while RFC 3264 supplies the offer/answer model. Each layer contributes a different record. A sequence gap is not a mode change. A successful SDP exchange is not proof that a sender obeyed it. A clean frame count is not a playout receipt.

Security did not collapse into framing

RFC 3952 points confidentiality toward secure RTP and discusses denial-of-service exposure from malicious payloads. RFC 3711 defines SRTP protection separately. That separation is operationally important: cryptographic acceptance can establish that protected bytes arrived under an expected context, but it cannot make an incorrect mode correct. Conversely, a perfectly divisible unprotected payload proves no identity.

The RFC Editor metadata page records the document as Experimental, published in December 2004. Its lesson is not that fixed-size framing was a mistake. Removing redundant fields is often excellent engineering. The lesson is that compression of representation does not eliminate state; it relocates state. RFC 3952 placed the missing count in a three-part proof: packet length, agreed mode and sender compliance. Operations that retain only one of those parts cannot later explain what the receiver believed.

Sources