Summary

  • RFC 5351 described Reliable Server Pooling as a way to resolve a logical pool handle, choose one or more Pool Elements and obtain another candidate after failure. A resolution response is evidence about the handlespace view and selection policy, not proof that the candidate is reachable now or can continue the old application operation.
  • The specification makes the hard boundary explicit: failover that depends on application state or transaction status cannot generally be defined without application-specific knowledge. The optional last cookie is opaque to the Pool User and may assist recovery, but authenticity, freshness, completeness, compatibility and safe replay remain separate receipts.
  • High availability therefore needs an evidence chain beyond discovery: scoped name, registered member, selected candidate, transport reachability, authenticated application peer, known commit boundary, valid state, safe retry, durable execution and observed outcome. Collapsing those steps turns “another address” into a recovery claim the protocol never made.

The address arrived before the answer

RFC 5351 presents a clean abstraction. The application asks for a server pool by handle instead of carrying a list of primary, secondary and tertiary hostnames. An ENRP server resolves that handle according to its knowledge of registered Pool Elements and the pool's selection policy. In the simple example, a GETPRIMARYSERVER-like primitive returns the first address. After a failure, GETNEXTSERVER reports the failed element and returns another address according to the best information then available.

That is useful infrastructure. It moves candidate discovery and elementary selection out of each application. It also creates a tempting sentence that the RFC does not justify: “The session failed over.” The actual event is narrower. One component stopped using one candidate and learned another candidate.

Consider a payment instruction, a reservation or a configuration write. The client sent the request. The connection failed before a response arrived. The old Pool Element might have rejected the instruction, committed it but lost the reply, or committed it and failed before replicating application state. A second address cannot distinguish those histories. Retrying may be required, harmless, duplicative or destructive depending on application semantics.

The architecture is candid about this limit. It says mechanisms that rely on application state or transaction status cannot generally be defined without more specific knowledge of the application. RSerPool supplies simple hooks. It does not manufacture a transaction boundary where the application has not exposed one.

A pool handle was deliberately smaller than a global name

The pool handle is a unique byte string in a flat handlespace with limited operational scope. RFC 5351 says administration of pool handles was not addressed by the protocol documents. RFC 3237 similarly confines validity to the operational scope and leaves interoperability between namespaces to other mechanisms.

This limitation is not an embarrassment. It tells an operator what the receipt proves. A successful lookup shows that the queried RSerPool instance recognized the handle. It does not show that another scope uses the same bytes for the same service, that a cross-domain caller is entitled to use it, or that a selected address belongs to the expected trust zone.

The RSerPool protocols between Pool User, Pool Element and ENRP server are also separate from the application protocol between user and element. RFC 5351 assumes compatible application protocols have been configured independently. A valid Pool Element record can therefore coexist with a version mismatch, missing feature, incompatible security policy or application-level rejection.

Name resolution is strongest when treated as a minimal coordination layer. It answers where a candidate may be found under one scope. It should not be inflated into identity, health, authority, state continuity or completion.

The handlespace was synchronized, not omniscient

ENRP servers exchange pool membership and synchronize their handlespace. Presence messages include checksums over Pool Element identities owned by a peer. A mismatch can trigger a scoped audit and resynchronization. When an ENRP server fails, peers negotiate takeover of its Home ENRP responsibilities.

Those mechanisms matter because a registry with a single point of failure would undermine the service it supports. Yet synchronization has a temporal meaning. Each server acts on information received and reconciled over time. A registration may expire. A deregistration may still be propagating. A Pool Element can fail between a keep-alive acknowledgement and the next application connection.

ASAP adds a local cache. RFC 5352 associates entries with a stale timer. On a stale cache hit, an implementation may refresh in parallel and answer from the entry, or block until an update arrives. If the entry is not stale, it normally avoids a new resolution request. “Not stale” therefore means “within a configured protocol age,” not “independently proven reachable at this instant.”

This is a normal distributed-systems tradeoff. Blocking for every fresh observation increases latency and does not eliminate the race after observation. Reusing cached state improves speed and creates an explicit uncertainty window. The failure occurs when the window is hidden from the evidence model and a cached address is presented as live fact.

Selection policy ranked candidates, not futures

RFC 5356 defines round-robin, weighted, random, priority and load-sensitive policies. These policies give a pool a repeatable way to order eligible members. They do not predict that a chosen operation will finish.

The word “load” is especially important. RFC 5356 defines a numeric range but says what utilization means is application-dependent and out of RSerPool scope. Members of one pool using load information must share a definition. One service may report active users, another CPU utilization and another memory use. Even within a common definition, the value is sampled state that may change before the next connection.

“Least used” is thus a statement inside a measurement contract: among the elements known to the selection process, under a shared definition and at the represented time, one value was lower. It is not proof of spare database locks, warm caches, current replication, authorization capacity or future response time.

A leadership dashboard should retain the policy name, input age, candidate set and rejection history. Reporting only the winning address erases the conditions under which it won. Reporting only a load number erases the definition that gave the number meaning.

Failure detection found absence, not causality

Home ENRP servers can send Endpoint Keep-Alive messages, and Pool Elements acknowledge them. A Pool User can report an endpoint unreachable. SCTP contributes its own association and path reachability mechanisms. These observations help retire failed candidates and select alternatives.

They still do not answer every application question. A timeout does not prove that the remote process never handled the request. A transport failure does not prove that the server failed; a path, firewall or local resource may have failed. A keep-alive acknowledgement proves a bounded response on the control relationship, not readiness for the next expensive business operation.

The distinction becomes more important when SCTP multihoming is involved. SCTP can move traffic to another path for the same association and peer endpoint. RSerPool can move the application toward a different Pool Element. Those are different continuity problems. The first may preserve association state at one endpoint. The second crosses a server boundary and may require application state and security context to move or be rebuilt.

RFC 9260 later replaced RFC 4960 as the SCTP base specification. That standards succession does not alter the analytical boundary: transport-path recovery, transport-association recovery and application-session recovery require different receipts.

The cookie was intentionally opaque

RFC 5351 offers an optional cookie mechanism. A Pool Element may periodically send state information. The Pool User keeps only the last cookie and, after failover, sends it to the new element. For the Pool User, the cookie has no structure; it is stored and forwarded.

Opacity is useful. The client does not need to understand every application's checkpoint format. It also means the client cannot inspect whether the cookie includes the last accepted request, whether it is newer than replicated server state, whether the new version understands it or whether replay will repeat a side effect.

The old Pool Element should sign the cookie and the new element should verify it. RFC 5352 says verification details are out of scope. A valid signature can establish bounded provenance and integrity under the chosen key arrangement. It cannot prove semantic completeness. An authentic stale checkpoint remains stale. An authentic checkpoint from an incompatible software version remains incompatible. An authentic record of “request received” is not necessarily a record of “transaction committed.”

Support is itself conditional. Cookie and business-card handling is not a universal mandatory capability for every Pool User and Pool Element role. Pool membership must not be used as a proxy for state-transfer support. Capability discovery and recovery testing need explicit evidence.

A business card was advice from one participant

The business-card mechanism lets a Pool Element recommend another element for failover. The recommendation may incorporate load considerations or knowledge that another server holds more current state. That can be more useful than a generic pool ordering.

It is still guidance from a participant. The recommended element may fail after the card is issued. Its state may lag. The recommending element may have an incomplete view. A malicious or compromised element may misdirect the user. The new element must still be contacted, authenticated and evaluated under application policy.

The phrase “next best” hides a metric. Best for load is not necessarily best for state freshness. Best for data locality may not be best for jurisdiction or security context. Best for accepting a read may not be best for safely continuing a write. The recommendation becomes auditable only when the relevant criterion is declared.

An implementation can use the card to prioritize a candidate without calling it a recovery certificate. That small language discipline preserves both the optimization and the uncertainty.

Graceful deregistration exposed two timelines

RFC 3237 requires transparent registration and deregistration. After a Pool Element deregisters, it may continue serving Pool Users whose connections began earlier, while new connections go elsewhere. Membership and active work therefore follow different timelines by design.

This is operationally healthy. Draining a server should not destroy existing sessions. It also means a handlespace snapshot cannot be used to enumerate all work still executing. “Not a current member for new selection” is not “no longer processing requests.”

The reverse is also true. “Registered” means eligible under the protocol state, not necessarily warmed up for every application shard or authorized tenant. Rollouts, drains and incident response need an application-level readiness gate above registration.

Without that gate, orchestration can register too early, send traffic before migrations or caches are ready, deregister and assume work has ended, then terminate sessions that were explicitly allowed to remain. The pool protocol did not create the mistake; an operator collapsed its membership receipt into an application-lifecycle receipt.

Security protected the map, not every business decision

RFC 5355 catalogues attacks against registration, handlespace synchronization and resolution. A rogue Pool Element can poison the pool. A malicious ENRP server can return an attacker-controlled address. Replay and flooding can corrupt or exhaust the system. The response includes mutual authentication and authorization requirements, with TLS and pre-shared-key mechanisms described for the single-administrative-domain model.

Those controls protect an essential chain: which infrastructure node may join the registry, which registry server may answer, and whether protocol messages were altered. They do not automatically authorize an end user's payment, database write or administrative change at the selected Pool Element.

There are at least two principals. The infrastructure principal is permitted to participate in the pool. The application principal is permitted to perform an operation. A third policy may govern whether application state can cross to another element or jurisdiction. Treating the first decision as all three creates a privilege escalation through availability machinery.

Cookie security illustrates the same rule. Signing prevents undetected modification under the assumed key trust. It does not decide which fields the receiving application may act on, whether the old authorization is still valid, or whether a recovered session must be challenged again.

The evidence chain needed ten receipts

A defensible failover record separates at least ten events:

  1. The pool handle existed and was meaningful in the intended operational scope.
  2. A Pool Element was registered in the handlespace view actually consulted.
  3. Resolution or cache returned that candidate under a declared selection policy.
  4. The candidate was reachable at the required transport endpoint at the relevant time.
  5. It spoke the required application protocol and authenticated the expected principal.
  6. The old operation's state was known: not started, uncommitted, committed or indeterminate.
  7. Any cookie or transferred state was authentic, fresh, complete and compatible.
  8. Retry or continuation was safe under idempotency, ordering and concurrency rules.
  9. The new element executed the intended operation and durable state reflected it.
  10. Independent observation confirmed the user-visible or operator-visible outcome.

No earlier receipt implies the next. ENRP can prove the second and contribute to the third. SCTP can contribute to the fourth. TLS can contribute to peer authentication. The application must own the commit boundary, state semantics and replay rule. Outcome monitoring must remain independent enough to detect a false internal success.

This chain also improves diagnosis. If the user saw failure, the team can locate the first absent receipt instead of labeling everything “failover failed.” If the user saw a duplicate, the chain points toward commit ambiguity or replay control rather than pool resolution. If the wrong principal gained access, it points toward the gap between infrastructure trust and application authority.

The registry documented symbols, not deployment

IANA still publishes RSerPool message types, parameter types, error causes and pool-member selection policy values. The table makes ASAP_HANDLE_RESOLUTION, cookie parameters and policies interoperable references. It is not a census of implementations or active pools.

The existence of a codepoint does not show that current software supports it, that operators enable it, or that a pool uses the corresponding feature. The same caution applies to RFC status. RFC 5351 is Informational and describes a suite targeted for the Experimental Track. That is neither proof of operational adoption nor proof of abandonment.

Lu Heng's Minimum Initial Specification note offers a later analytical lens: keep the shared coordination object small and preserve localized decision authority. Here, the handle, message vocabulary and basic policy identifiers form the shared layer. State interpretation, safe replay and business authorization remain local to the application that can actually know them.

His Reality Layers note adds a related distinction. A handlespace record is symbolic protocol state. A live process, durable transaction and user-visible result are different operational facts. The notes are disclosed lenses, not historical sources for RSerPool. The RFCs and IANA registry carry the protocol claims.