Summary

  • RFC 3389 let a sender replace ordinary speech packets during inactive intervals with a compact description of background noise: one required level byte and optional spectral coefficients.
  • The receiver generated the audible ambience locally, so a valid comfort-noise packet proved that a model crossed the network, not that the distant room had been recorded or faithfully reproduced.

Imagine an early packet-voice call after the person at the far end stops speaking. The stream of speech frames ceases. If the receiver substitutes mathematical zero, the line becomes unnaturally dead. Air conditioning, microphone electronics and the low grain of an occupied room vanish at once. A listener may think the call has failed and ask, “Are you still there?”

The economical answer was not to keep shipping full-rate samples of a room in which nobody was talking. It was to describe the room’s background briefly and let the other endpoint invent a plausible version of it.

RFC 3389, published on the Standards Track in September 2002, standardized the small object that crossed that divide. Its title was technical—Real-time Transport Protocol Payload for Comfort Noise—but its architectural decision was unusually clear. The network would carry a minimum set of comfort-noise parameters. Detection, analysis and sound production would remain local.

This format was intended primarily for audio codecs such as G.711, G.726, G.727, G.728 and G.722 that did not already provide an integrated comfort-noise mechanism. During active speech, the sender transmitted the normal codec’s RTP packets. During an inactive voice segment, a voice-activity detector could decide that ordinary speech encoding was no longer worth the rate. A discontinuous-transmission algorithm then decided when to send a comfort-noise update.

The RFC did not standardize either decision. It did not prescribe how sensitive a detector should be, how long it should wait, or how a generator should make artificial room tone. It standardized the handoff between independently built endpoints.

That handoff began with one byte. RFC 3389 called the comfort-noise object a Silence Insertion Descriptor, or SID frame. Its first octet carried the magnitude of the noise level in negative decibels relative to digital overload. Seven bits represented values from 0 to 127; the top bit was unused and had to be zero. Zero meant 0 dBov, while 127 meant −127 dBov.

This number was not a sample of the distant room. It was a bounded statement about level. The distinction matters because a statement can drive many possible outputs. Two receivers may accept the same byte and produce different waveforms, even while both conform to the wire format.

The sender could append spectral information. Each later byte represented a quantized reflection coefficient for an all-pole noise model. The number of coefficients—the model order—was inferred from packet length rather than sent as another field. The encoder could choose that order according to quality, complexity, expected noise and signal bandwidth. A decoder could even reduce its local complexity by setting higher-order coefficients to zero.

So interoperability had layers. Every conforming exchange had a level. Some had a richer description of spectral colour. An older implementation, built for the earlier one-byte form, could read the first octet and ignore the rest. It had not failed to parse the packet; it had accepted a thinner account of the source environment.

The RTP rules made the object behave like a small independent codec. The timestamp marked the beginning of the comfort-noise period. The marker bit should not be set on the CN packet. Exactly one CN payload per channel belonged in each RTP packet because the payload length varied; multichannel payloads had to use the same spectral model order for every channel.

Payload type also carried a boundary. Static RTP payload type 13 identified comfort noise at an 8,000 Hz RTP clock. It did not grant the same meaning at any arbitrary rate. A session using another clock rate had to assign a dynamic payload type and signal a mapping such as CN/16000. A familiar number was therefore not enough: number, clock and negotiated media description had to agree.

At the transition into an inactive segment, the sender issued a CN packet in the same RTP stream. What happened afterwards was deliberately not global policy. An implementation might send periodic updates. It might wait until the measured background changed significantly. Between updates, the receiver continued from the last model and its own synthesis algorithm.

This is where an apparently simple packet becomes an evidence problem. A capture of a CN packet can establish its payload type, timestamp, level byte and any coefficient bytes. It may establish that the sender advertised a noise model. It cannot establish that the local voice detector classified the scene correctly. It cannot show what analog sound entered the microphone before analysis. It cannot recover the exact waveform the remote listener heard, because that waveform was generated at the receiver.

Nor does an interval without packets prove comfort-noise operation. RFC 3389’s SDP guidance explicitly says that omitting CN from a media description does not imply that silence suppression will not occur. RTP permits discontinuous transmission for any audio payload format. After such a gap, a receiver can observe a timestamp discontinuity even when the sequence number advances by one. The missing packets may reflect intended suppression, but observed absence alone may also reflect loss, interruption or failure. Other receipts are necessary.

Later documents exposed the reach of the choice. RFC 6263 considered CN packets as a way to keep RTP paths through network address translators alive. It also warned that the remote side had to support and negotiate CN and might actually render the keepalive as noise. A supposedly operational packet could become audible. Level selection was no longer merely an aesthetic matter.

RFC 6465 reused the same negative-dBov convention for an RTP audio-level extension. That reuse made measurement vocabulary more consistent; it did not turn the audio-level header extension and the CN payload into the same assertion. One reported a level alongside media. The other supplied parameters from which a receiver could create sound during an inactive interval.

Lu Heng’s principle of Minimum Initial Specification offers a useful reading of this design. RFC 3389 specified what had to be common for independent endpoints to interoperate: level encoding, optional coefficient encoding, packing, timing, channels, payload naming and clock-rate rules. It did not centralize the future of noise detection or synthesis. A receiver designer could improve the generated result without changing the shared packet format, so long as the received model retained its agreed meaning.

Running-Code Primacy adds the necessary restraint. A document did not make the listener’s room sound natural. Implementations had to detect, send, decode and synthesize; operators had to observe the result. The standard’s authority ended at the conditions it actually defined. Perceptual fidelity remained a claim for measurements and listening tests, not for the existence of an RFC number.

That separation is the enduring history. RFC 3389 did not transport silence and did not transport a room. It transported a compact permission and set of parameters for a receiver to make a local substitute. The speech packets stopped. The physical room continued. The audible room at the other end became a new artifact, constructed from one small message and one endpoint’s running code.

Sources