Summary
- The new
draft-nygate-ippm-mrl-00defines mouth-to-ear response latency from the last transmitted speech sample to response onset, using one monotonic clock at the calling endpoint. - It requires both packet-arrival and buffered-playout results, onset-threshold uncertainty, continuity after the first sound, calibration, turn indices and the full attempted/reported/discarded denominator.
- The proposal is an active individual Internet-Draft, not adopted IPPM work or a registered metric; it measures time, not whether the answer is correct or useful.
A millisecond number without its interval
Imagine two voice agents. One returns a short “um” after 420 milliseconds, falls silent for half a second and begins its answer at 960 milliseconds. The other stays quiet until a continuous answer begins at 720 milliseconds. A detector that stops at the first sound declares the first system faster. A caller may reach the opposite conclusion.
That is not merely a dispute about taste. The two measurements end at different events. The first records acoustic onset; the second is closer to the start of sustained response. Neither number can explain itself once detached from the rule that produced it.
The new individual Internet-Draft Mouth-to-Ear Response Latency for Conversational Voice Systems tries to make that hidden rule travel with the result. Datatracker records revision 00 as uploaded and accepted on 7 September 2026. The author marks it Informational and discusses IPPM, DISPATCH and Independent Submission as possible homes. It is not an adopted IPPM document, an IETF consensus result or an IANA-registered metric. Its value today is as a testable specification under review.
The draft defines Mouth-to-Ear Response Latency, or MRL, as t1 - t0. The start, t0, is the instant when the final speech sample of a prerecorded caller stimulus is transmitted at the caller's RTP boundary. The sample is annotated offline, mapped into its RTP packet and located within the frame. Runtime voice-activity detection cannot define the boundary because the detector's own decision lag is part of the system interval being measured. Moving the start later would make every result look faster.
The end, t1, is the first response sample observed at that same endpoint. Both timestamps come from one monotonic clock on one host. That avoids the clock-offset problem of one-way measurements between machines. It does not remove network delay; it makes clear that the path and system under test are measured as a pairing.
MRL is signed. A system that begins a backchannel or predicts the end of speech aggressively can produce a negative result because response audio arrives before the stimulus is finished. The draft says to preserve that observation, not clamp it to zero. A negative value may describe real behaviour—or a greeting mistakenly detected as the response. The capture and detection boundary decide which.
Arrival and playout are different evidence
The proposal requires two versions of the end time. Ingress MRL observes the onset sample when its packet arrives, before buffering. It exposes the contribution of the system and path, including jitter. Playout MRL observes when that sample would leave a de-jitter buffer of a declared target depth. It is closer to the wait experienced by a caller, but it includes the endpoint's buffering policy.
Neither is allowed to stand alone. Ingress can understate the human wait. Playout can make a system look slower because one measuring client chose a deeper buffer. A headline that does not carry both cannot tell a buyer which difference belongs to the service and which belongs to the observer.
The reference point matters for the same reason. An RTP endpoint, host packet capture, bridge recording, carrier recording and endpoint audio device each put different stacks, relays, transcoding, buffering and hardware inside the interval. The draft permits a different observation point if it is declared. Unless its added terms have been characterised, however, the result is an upper bound and must not be pooled with RTP-endpoint figures.
This is where apparently exact rankings become governance claims. A vendor can publish three decimal places while leaving the interval undefined. Precision in arithmetic does not supply comparability in evidence.
First audio can be a measurement shortcut
Synthesised speech often ramps in rather than starting at full level. “First sample” therefore needs a threshold. The draft specifies sensitive, headline and strict onset variants and requires the spread among their MRL values to be published as onset-definition uncertainty. A gradual ramp produces a wider uncertainty band. The headline figure is incomplete without it.
Greetings need a separate boundary. If a detector scans the entire received stream, it can select audio spoken before the caller's stimulus. The draft describes a reference capture in which a genuine 900-millisecond response became an apparently valid negative result because an 800-millisecond greeting was selected. Unprompted audio must end before response detection begins.
Then comes filler. An earcon, breath or filled pause can lower first-audio MRL while conveying nothing. The draft deliberately avoids a semantic classifier: an external reviewer could not reliably re-derive “meaning” from a published capture. Instead it uses continuity. Within 2000 milliseconds after first onset, it looks for the longest below-threshold interval. If that silence exceeds 150 milliseconds, the response is marked discontiguous and a second MRL is reported at the start of the final continuous segment.
Those constants are not sacred. The author explicitly says that no real filler corpus informed them and asks reviewers whether they are right. That uncertainty is important. The proposal has supplied a falsifiable rule, not discovered a universal boundary between hesitation and answer.
The denominator belongs with the percentile
The measurement bundle is broader than two latency values. It includes codec, frame period, buffer target, threshold parameters, stimulus identifier and hash, calibration conditions, SUT identifier, recommended configuration hash, advisory flags and turn indices. Raw payloads and raw timestamps are kept so a reviewer can re-run a different onset rule without recollecting the call.
First and later turns cannot be silently merged. The first exchange may include session setup, cold model state, empty caches and unopened connections. Later turns may benefit from accumulated context. Both are real, but they answer different operational questions. A distribution must say which turn indices it contains.
It must also report how many trials were attempted, reported and discarded, with reasons that sum to the discard count. A run that removes every case with no detected onset can post an attractive percentile for its survivors while hiding a reliability failure. Packet loss and late discard remain visible advisory conditions; if the onset frame is lost, the number becomes an upper bound.
Calibration is another line between a number and evidence. A reference responder emits after a programmed delay. The instrument reports bias and error under stated conditions. A device that has not been calibrated does not get to omit that fact; its figures remain upper bounds. The draft even warns that calibration only at frame-aligned delays can conceal frame quantisation errors.
What the metric cannot decide
MRL contains no judgement about response correctness, recognition accuracy, relevance or usefulness. A system can answer quickly and wrongly. It can begin fluent speech promptly and still hallucinate. Those outcomes need different instruments and cannot be inferred from a latency field.
Nor is active measurement consequence-free. Generating calls against another operator's service without authorisation can breach terms and, at volume, resemble denial of service. Captures contain stimulus and response audio; recorded human speech can carry personal and biometric significance. Synthetic or consented material, access control and a retention limit belong in the measurement design.
The proposed report is therefore a minimum evidence package, not a universal ranking machine. It keeps observable boundaries separate: final caller speech, packet arrival, buffered playout, first sound, continuous speech, configuration identity and surviving sample set. That is the useful news. A buyer should not ask only which system posted the smallest number. The prior question is whether the numbers describe the same event.
Sources
- Current Datatracker record
- Datatracker revision history
- Recent Internet-Drafts listing
- Revision 00 text
- RFC 3550: RTP
- RFC 6076: SIP end-to-end performance metrics
- RFC 6390: new performance metric guidance
- RFC 7679: one-way delay
- RFC 7799: active and passive measurement
- RFC 8911: Performance Metrics Registry
- Heng Lu on a minimum initial specification
- Heng Lu on reality layers and symbolic power
- Heng Lu on running code as primary
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
