Summary
- RFC 5390 showed how a SIP request that would normally produce one downstream attempt could, in its bounded UDP example, produce as many as 18 requests and as many responses when 503 replies, alternate-server retries and retransmissions interacted.
- The failure was not simply “too much traffic.” The 503 mechanism lacked reliable scope, graded reduction and shared dependency knowledge, so it could amplify load, idle healthy capacity or oscillate traffic between saturated servers.
- Later work separated monitor, control decision, feedback and upstream actuator. A refusal, an overload signal, a throttle, a completed transaction and a successful call remain different receipts.
The negative answer still consumed capacity
RFC 3261 gave SIP a basic response for temporary overload or maintenance. A server could return 503, optionally say when to retry, and an upstream proxy could attempt the request at another server. The behavior sounds resilient: move work away from the node that says it is busy.
That logic assumes the problem is local. If one member of a farm is overloaded while another has spare capacity, retrying can recover useful work. RFC 5390 examined the opposite state. All three downstream servers in its first topology were already overloaded. The load-balancing proxy sent the request to the first server, received a 503, tried the second, then tried the third. Each node had to parse, schedule and reject work that could not complete.
Delay made the accounting worse. While an overloaded server struggled to produce its refusal, a UDP client could retransmit. The RFC's bounded example counted four SIP transactions including the client-to-proxy leg and up to seven retransmissions per UDP transaction. Before timeout, one client request could result in as many as 18 requests and as many responses. The number is a model inside the standard, not production telemetry, but its lesson is operational: rejection has a cost, and retry can multiply that cost.
Reliable transport did not erase the control error. In the same example, even with no TCP segment retransmission, one request could produce three downstream requests and four responses. Transport reliability reduced one multiplier. It did not tell the proxy that every alternative shared the same exhaustion.
This is the first evidence boundary. A 503 proves that one processing path chose not to complete one request. It does not prove that capacity was returned, that another destination is healthy, that offered load fell or that the user received a final outcome.
Hidden dependency state made the graph repeat itself
RFC 5390 then placed another proxy above the first pair of load balancers. Proxy P1 learned through failed attempts that servers S1, S2 and S3 were all overloaded. But that knowledge did not travel back as usable control state. The upstream proxy tried P2, which explored the same three servers again.
The topology exposes a failure of provenance. P1 possessed evidence about its downstream set, but the 503 sent upstream could be interpreted as P1's own condition rather than a statement about shared dependencies. RFC 3261 had discouraged forwarding 503 upstream precisely to avoid that confusion. The result was equally dangerous: the next branch lacked the information needed to avoid replaying the same search.
An operator looking only at response counts might diagnose “healthy failover.” The graph tells a different story. More nodes are participating, yet no new capacity or path diversity has entered the decision. The retry tree expands while the set of exhausted resources stays constant.
This distinction matters beyond SIP. Redundancy is not the number of endpoint labels. It is independence of the resources that determine success. Two servers backed by one failed database are one failure domain for the relevant request. A refusal without dependency scope can turn apparent diversity into repeated expenditure.
One 503 could also switch off healthy capacity
The same ambiguity caused the opposite error. RFC 5390 noted that RFC 3261 did not make the scope of 503 sufficiently clear. Did the signal apply to an IP address, a host name or a URI? Some implementations used the host name. If DNS SRV mapped that name to a cluster, one member's 503 could stop traffic to every member.
The system then underused capacity instead of amplifying load. Healthy servers became unreachable through policy because a narrow condition had been interpreted broadly. Both failures came from the same missing field: the receiver could not identify the exact object whose capacity the signal described.
Scope is therefore not metadata decoration. It is the boundary of delegated action. A server may have authority to report its local processor state. It cannot, merely by emitting the same code, establish that its siblings, database, route, host name or application service share that state.
REQ 18 in RFC 5390 made this explicit: overload indication had to be unambiguous about whether it applied to an IP address, host or URI. A useful dashboard should preserve that object, the observation time and the upstream decision. “503 received” is not enough to reconstruct what was disabled.
Binary backoff could create an oscillator
Retry-After gave the overloaded server time to clear its backlog, but the baseline behavior was all or nothing. RFC 5390's two-server example put both servers at full capacity. When S1 rejected a request and asked for quiet, the proxy shifted all traffic to S2. S2 then received twice its capacity, issued its own refusal and sent the traffic back when S1's timer expired.
Nothing in either local decision was irrational. Each node reported a real constraint. The instability emerged from the actuator: zero or all traffic, applied across a small number of upstream relationships, with no graded target. The signal did not carry how much load the server could still usefully process.
The RFC carefully bounded this observation. Where a server has many independent clients each contributing a small share, sending 503 to a subset can approximate a finer reduction. The article therefore does not claim that Retry-After always oscillates. It claims that the mechanism did not guarantee stable graded control across the topologies RFC 5390 needed to protect.
REQ 7 demanded a throttle that recognized degrees of overload. REQ 21 demanded that throughput settle when offered load fell below capacity. These were not cosmetic enhancements to an error response. They changed the system from a collection of refusals into a feedback loop.
An error code had acquired several incompatible meanings
RFC 5390 also found that implementations used 503 for conditions unrelated to a server's own resource exhaustion. A gateway might have ample SIP-processing capacity while a downstream PSTN path could not take a particular call. Another farm might share a failed database. In the first case an alternate route could help; in the second, an alternate front end could lead back to the same failed dependency.
The upstream proxy could not infer which world it inhabited. Retrying might recover the request or magnify useless work. Avoiding retry might protect the network or strand healthy capacity. The same visible token represented different causal states.
REQ 6 therefore required an explicit overload signal to say that overload, rather than another failure, caused it. REQ 14 required clear retry guidance, especially for connection establishment and registration after a restart. REQ 8 and REQ 9 held both sides of the boundary: do not retry into an overloaded or unknown target, but do not prevent a genuinely healthy target from doing useful work.
This is a useful governance pattern. A common label does not create a common authority. The decision-maker needs the cause, scope and requested action separately enough to accept or reject each one.
Later standards split observation from actuation
RFC 6357 gave the control loop explicit components. A monitor samples the protected SIP processor. A control function turns those samples into feedback. An actuator on the sending side enforces a throttle. The receiving node measures and reports; the upstream node actually reduces, delays, rejects or redirects traffic.
That separation makes failure visible. A monitor can be accurate while feedback is stale. Feedback can arrive while the actuator ignores it. The actuator can meet a rate target while choosing the wrong messages to discard. A lower offered rate can still fail to restore useful throughput if a hidden dependency remains unavailable.
RFC 6357 also stated why local rejection was not sufficient: rejecting requests consumes server resources. It can serve as a last layer of protection, but it cannot alone prevent congestion collapse. The actuation has to occur before excess work reaches the scarce processor.
RFC 7339 later carried overload information hop by hop in the topmost Via header. Because that Via entry is consumed by the adjacent client, the feedback belongs to one neighbor relationship. It can help incrementally where both adjacent elements support the mechanism and does not require a separate control message during overload.
The control still remained bounded. oc described a reduction under the default loss-based scheme; oc-validity limited its lifetime; oc-seq ordered updates; oc-algo identified the selected algorithm class. Publication of those fields did not prove that every intermediary implemented them or that the whole call path was protected.
A percentage and a rate were different promises
RFC 7339 required support for a loss-based algorithm. A server could ask an upstream client to reduce the share of requests it forwarded. This gives a lightweight feedback contract, but a percentage follows the offered load: if offered load rises, the accepted portion can rise too.
RFC 7415 added an optional rate-based method. The overloaded server supplies a maximum request rate for a client, constant until a later update. That establishes an upper bound between updates. It also requires more per-client control because one target rate may not suit every upstream source.
Neither value is a service promise. A maximum of 150 requests per second does not say the server will complete 150 transactions, that those messages have equal cost, or that 150 calls will succeed. RFC 7415 notes that message mix affects processor cost, and RFC 5390 left prioritization to local policy. The upstream node still decides which messages consume the allowance.
The receipt chain must therefore keep rate or loss feedback separate from the selected requests, completed transactions and user-visible calls. A compliant throttle can protect the server while refusing a commercially or operationally important class. Conversely, a prioritization policy can preserve emergency or teardown work while total request count falls.
Useful throughput was the governing metric
RFC 5390's first requirement did not optimize the number of responses emitted. It aimed to preserve useful throughput when offered load greatly exceeded capacity. That choice prevents an overloaded system from declaring success because it rejected every request quickly.
Useful throughput requires a joined view: work admitted, work completed, state released and desired sessions preserved. It is sensitive to message type. Completing an in-dialog teardown may free more resources than admitting a new call. Processing a registration refresh can have a different cost and consequence from processing an INVITE.
The IETF requirements deliberately stopped short of dictating one priority algorithm. That authority remained local. A provider could preserve emergency traffic or existing dialogs according to its obligations. The shared protocol supplied feedback; it did not become a global scheduler.
This is where Lu Heng's reality-layer discipline clarifies the design. The standard can define interoperable evidence and a narrow exchange. The running operator still measures its actual bottleneck, owns the admission decision and must prove what the implementation did. Consensus over a Via parameter does not create capacity, complete a call or transfer liability for a bad priority choice.
Evidence boundary
The frozen record establishes RFC text, publication status, stated deployment problems and later protocol architecture. It does not establish that any current carrier or product uses RFC 7339 or RFC 7415, that an outage followed the 18-request model, or that a specific 503 came from overload. Those claims require configuration, packet, resource and outcome evidence from the deployment itself.
The 18-request number must remain attached to its assumptions: the RFC's three-server load-balancer topology, UDP transactions, retransmission behavior and timeout. It is an explanatory bound, not a universal multiplier. Likewise, the oscillation example applies to a small number of upstream relationships, not every fan-out pattern.
What can be carried forward is the control invariant. Refusal is not reduction. A trustworthy system must show the monitored resource, causal scope, feedback version, upstream actuation, local priority choice and useful result. Without that chain, the network may be producing perfectly valid errors while operationally moving farther from recovery.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
