Summary
- RFC 3539 requires AAA systems to expect a failover duplicate on any connection. An alternate copy may arrive before packets on the original path have drained, so connection failure does not establish transaction extinction.
- Authentication is not always idempotent. The RFC gives the exact case in which an initial request receives Accept, a duplicate reaches another server after login state changes and receives Reject, and the client's outcome depends on which answer arrives first. Duplicate identity, answer selection, NAS enforcement and user outcome therefore need separate receipts.
Availability engineering often begins with an innocent sentence: if this peer stops answering, send the work somewhere else. The instruction sounds like movement. In a distributed system, it is usually copying. The sender cannot prove that every byte, request or application decision on the first route has disappeared before it activates the second.
RFC 3539, published in June 2003 as a Proposed Standard, made that ambiguity unusually concrete for Authentication, Authorization and Accounting. It did not merely warn that duplicate packets might occur. It described how failover, failback and load balancing can expose one logical AAA transaction to multiple agents or servers, and why a state-dependent authentication can return opposite answers.
That mechanism is the Article's bounded subject. It is not evidence that a named operator experienced a double login, a billing error or a Diameter outage. It is a protocol-design warning about what must be proved before “failover succeeded” can stand as an authorization result.
A connection can fail before its transaction is dead
A client that maintains more than one AAA route may use primary and secondary peers or distribute load. When it decides to fail over, it can send a pending transaction on an alternate connection before the old connection's traffic is guaranteed to have left the network. The alternate copy can arrive first. The original can arrive later. Both can be valid representations of the same logical attempt.
RFC 3539 therefore requires agents and servers to handle duplicates and to assume that a duplicate can arrive on any connection. Connection identity is not transaction identity. A new socket, new path or new server does not make the replay a new business request. Conversely, the failure of one connection does not prove that the request assigned to it was never processed.
This is the point at which an availability dashboard can mislead. “Primary down, secondary up” describes peer routing. It does not say whether the original request reached a proxy, whether a downstream server committed state, whether an answer is returning late, or whether the alternate is about to evaluate the same identity against a changed database.
The receipt chain needs two clocks immediately: the peer-failure clock and the transaction-liveness clock. They may overlap, but they are not interchangeable.
The watchdog's jurisdiction ends at the immediate peer
RFC 3539 requires an application-layer watchdog so a AAA implementation can detect both transport and application-function failures more quickly. The watchdog is intentionally local. It tests the immediate peer, not every server, database or policy engine behind that peer. It is not a cluster heartbeat.
That limit prevents a downstream problem from destabilising every upstream client. Any AAA response from the peer counts as evidence that the peer is up. Silence on an ordinary request is not enough to declare it down because the missing answer may be caused by a downstream proxy or server. Only failure to answer the watchdog supports the peer-down transition in the described algorithm.
The distinction is operationally valuable and evidentially narrow. A watchdog answer proves that one adjacent application endpoint answered one liveness exchange. It does not prove that the user's home realm is available, that its state is current, that an authorization was performed, or that the NAS applied a result.
The timer also reallocates risk rather than eliminating it. The profile gives an unjittered default of 30 seconds and permits a value as low as six seconds, excluding jitter, but warns that shorter timing increases duplicates and spurious failover and failback. A short interval reduces waiting for a genuinely failed peer. It also gives a merely delayed peer less time before the pending queue is copied elsewhere.
There is no green number that maximises both properties. The timer owner chooses where uncertainty is paid.
Failover replays a queue, not a philosophical intention
To perform failover, the client or agent maintains a pending-message queue for its peer. An answer removes its corresponding request. When failover begins, the queued messages are sent to an alternate agent if one is available. The mechanism works on records: still pending here means eligible for replay there.
The queue cannot know that an old request has just crossed a remote commit point. It cannot infer that a server decided Accept but its answer is delayed. Nor can it know that the first route stopped at a proxy while the second reached the final realm. Its local fact is smaller: no correlating answer has yet retired this entry.
That is why the replay receipt should preserve the old peer, alternate peer, trigger, queue snapshot, transaction identifier, request digest, retransmission marker and timestamps. “Retried” is too vague. It hides whether the same bytes were sent, whether identity was preserved, whether the target changed, and whether the first attempt was still eligible to finish.
Later Diameter practice makes the fields explicit. RFC 6733 retains RFC 3539's transport-failure algorithm, forwards pending requests to an alternate agent with the retransmitted-request flag, and warns that multiple identical requests or answers may result. It uses the End-to-End Identifier together with Origin-Host to detect duplicates. The Hop-by-Hop Identifier serves the adjacent request-answer exchange.
Those coordinates are not substitutes. The hop identifier helps retire a local queue entry. The end-to-end identity says that work observed across agents belongs to one logical request. Neither field is a cryptographic proof of freshness, a global lock or a guarantee that only one effect occurred.
Recognition is not reconciliation
The phrase “duplicate detection” can suggest that the hard problem is over once two messages share a key. In practice, classification begins the problem. A processor still needs a disposition.
It can return a cached answer, repeat a previously committed decision, wait for an authoritative state owner, reject a conflicting replay, or re-evaluate the request. Each choice has consequences. RFC 6733 says duplicate requests should cause the same answer to be transmitted, apart from hop-specific details. That is a stabilising rule because it prevents one logical request from becoming an uncontrolled second policy decision.
But a stable identifier cannot reconstruct a decision that was never durably recorded. It cannot make two servers share a state snapshot. It cannot compensate an access or charging effect already committed. It cannot tell the client which of two received answers it actually applied.
A defensible duplicate record therefore contains more than the key. It identifies the server, its state version, the applicable simultaneous-use or other constraint, whether it found a prior decision, whether it replayed or recomputed the answer, and the time at which that decision became durable.
Why one request can honestly receive opposite answers
RFC 3539 distinguishes accounting from state-dependent authentication. Duplicate accounting records can be weeded out using the accounting session ID, event timestamp and NAS identity. That is already a multi-field statement: a record is not the same event merely because it resembles another packet.
Authentication can be harder. The RFC's example applies a simultaneous-use restriction. The initial request reaches one server while the user is not yet presumed logged in and elicits Accept. A duplicate reaches another server after state has changed and elicits Reject because only one concurrent session is allowed. The client can receive both answers, and the result can depend on which arrives first.
Neither server needs to be malicious. Both can be internally consistent with the state they observed. The contradiction is created by time, replication and an operation that is not idempotent.
That makes “first response wins” a policy, even when no one wrote it down. Network latency, queueing and failover topology begin selecting the authorization outcome. A secondary server closer to the client may beat an older Accept returning through a congested proxy. The reverse ordering could admit the session before the Reject arrives.
The RFC describes a possible mechanism, not a production incident. The governance conclusion is an inference: if response order can determine the outcome, the operator should either make that rule explicit and auditable or design the server-side decision so duplicates cannot create a second answer.
Accept is still not access
Even after the client chooses an answer, the chain is unfinished. Accept is an authorization message under the protocol's policy context. The NAS must correlate it to the pending request, validate it, install the returned profile and change the access state. A later accounting start must identify the session. Observed traffic must show that the service actually became usable.
Reject is similarly bounded. It records a server decision, not proof that every path denied access. If another Accept had already been applied, the Reject might arrive too late, trigger cleanup, be discarded as a duplicate, or expose an unresolved inconsistency.
The final evidence chain therefore needs distinct receipts for request creation, immediate-peer liveness, failover replay, every server decision, client selection, NAS enforcement, accounting disposition and user-visible outcome. A transport owner can prove that the alternate connection worked. That owner cannot sign for the policy database or the user's access.
Older RADIUS identity and this failover boundary are not the same article
Classic RADIUS has its own transaction coordinates: the one-octet Identifier, Request Authenticator, transport context and shared secret matter to response matching, retransmission and duplicate caching. RFC 2865, RFC 2866 and RFC 5080 bound those mechanisms.
They must not be rewritten as Diameter fields. RFC 3539's cross-connection warning and RFC 6733's End-to-End Identifier plus Origin-Host solve a different coordination problem. One tells a system that messages seen through different agents belong to the same logical work. The other describes the older RADIUS request-response identity under its own protocol.
The operational lesson is shared but not interchangeable: preserve enough identity to distinguish a replay from new intent, then preserve enough decision state to avoid executing that intent twice.
The receipt that a failover review should demand
Start with the logical request: Origin-Host or its protocol-specific source identity, end-to-end key, immutable request digest, user/session scope, application and creation time. Record the pending-queue entry and the peer to which it was first sent.
For liveness, retain the immediate peer, connection generation, last qualifying response, watchdog request and answer, timer value including jitter, failure classification and the distinction between transport silence and downstream silence. Do not promote this record into a home-server health claim.
For failover, record the old and alternate peers, trigger, exact queue snapshot, retransmission marker and send times. At every decision server, preserve state epoch, relevant constraints, duplicate lookup, prior-decision reference, new or replayed answer and commit time.
At the client, retain all answers received, correlation results and the rule that selected one. “First” should mean an observed sequence with timestamps, not a retrospective story. At the NAS, record what was enforced. In accounting, link the session ID, event time and NAS identity to the dedupe disposition. Then observe the service or denial that the user actually experienced.
The executive sentence can be exact: “The watchdog declared this immediate peer unavailable; the client replayed these still-pending identities to this alternate; both server decisions were reconciled under this rule; the NAS applied this answer; the resulting session and accounting state were observed.” If one clause is missing, say so.
Evidence boundary
This Article identifies no operator, vendor, AAA realm, RADIUS or Diameter deployment, account, login, NAS, incident, outage, security event, duplicate charge or affected user. It reports no adoption rate, current configuration, measured latency or observed contradictory answer.
RFC 3539 is treated as a June 2003 Proposed Standard transport profile. RFC 3588 is historical Diameter context and was obsoleted by RFC 6733. RFC 6733 preserves the failover and duplicate-identification mechanism; it does not prove that any named system implements it correctly. RFC 8174 bounds the use of normative language, and RFC 6298 supplies retransmission-timer context without becoming AAA deployment evidence.
Heng Lu's Running-Code Primacy and Minimum Initial Specification notes are disclosed editorial lenses. They support the discipline of separating a published mechanism from implementation, local decision and observed result. They do not establish the RFC authors' intent or any operational fact.
The bounded conclusion is simple: failover can copy one logical AAA request before the original has disappeared. A common identifier makes the copies recognisable. Only controlled idempotency, answer selection, enforcement and outcome evidence can make the effect singular.
Sources
- https://www.rfc-editor.org/rfc/rfc3539.html
- https://www.rfc-editor.org/info/rfc3539
- https://datatracker.ietf.org/doc/rfc3539/
- https://www.rfc-editor.org/rfc/rfc6733.html
- https://www.rfc-editor.org/rfc/rfc3588.html
- https://www.rfc-editor.org/rfc/rfc2865.html
- https://www.rfc-editor.org/rfc/rfc2866.html
- https://www.rfc-editor.org/rfc/rfc5080.html
- https://www.rfc-editor.org/rfc/rfc8174.html
- https://www.rfc-editor.org/rfc/rfc6298.html
- https://heng.lu/running-code-primary-the-patch-needed-to-preserve-the-internet-original-design/
- https://heng.lu/minimum-initial-specification-localized-future-decision-voluntary-adoption-internet-coordination-system/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
