Summary

  • draft-ietf-radext-radiusdtls-bis-18 makes RadSec health a per-connection, next-hop fact: a proxy hides the operational state of realms and credential backends behind it.
  • Status-Server, TLS/DTLS state and transport keepalives are valuable bounded receipts. Realm routing, backend processing, retry ownership, authorization and delivered access still need separate evidence.

A network access server sends an Access-Request for a subscriber in the affected customer realm. Its TLS session to the federation proxy is established. The certificate verifies. The socket is writable. When ordinary requests stop receiving answers, the application watchdog sends Status-Server; the proxy replies immediately. The dashboard therefore keeps the peer green.

The subscriber still cannot connect.

The easy explanation is that monitoring failed. The more accurate one is that monitoring answered the question it was designed to answer. RADIUS is hop-by-hop. To the access controller, the proxy is the server. To the home server, the same proxy is a client. The proxy may route dozens of realms to different home servers or credential stores. Its response proves that this adjacent process can receive, authenticate and answer a non-forwardable watchdog packet. It cannot prove that the route for example.net is correct, that the next link is up, that the selected home server is processing its queue or that its identity database is healthy.

Revision 18 states this boundary directly. A RadSec client combines TCP, TLS or DTLS connection state with the RFC 3539 application-layer watchdog. It may mark a connection down when the network stack says the path is no longer viable, when the protected transport offers no usable connection or when the watchdog has declared it down. The unit is the connection. If one of several connections to a server becomes unresponsive, the client must not mark all the others down.

That prevents one bad shard or forgotten DTLS association from poisoning every route to the same peer. It also blocks the opposite mistake: treating one responsive connection as proof that everything reachable through the proxy is healthy.

RFC 5997 supplies the reason. A Status-Server packet must not be forwarded by a RADIUS proxy or server. The reply describes the destination at the address and port being tested. Because the server does not use the User-Name realm to forward the query, the answer cannot establish reachability for a particular realm. Its non-forwardability is a feature. It separates first-hop failure from downstream silence. The same property limits the claim.

The draft requires RadSec clients to implement that watchdog and servers to answer it. It recommends reserving Identifier zero on each connection because RADIUS has only 256 simultaneous Identifier values. TCP keepalives or DTLS heartbeats may accelerate notice that the transport has failed. The text calls them essentially redundant in the presence of the application watchdog. None turns a connection probe into a realm probe.

This creates two symmetrical operational hazards.

The first is premature failover. A slow or failed home realm produces no Access-Accept or Access-Reject, so the client declares the whole proxy dead and shifts every realm to another proxy. Healthy traffic moves, connection pools churn and the secondary inherits load caused by one hidden dependency. RFC 3539 described this problem before RadSec: a client separated from a home server by an agent cannot estimate end-to-end transport parameters from its local transport alone.

The second is sticky failure. The proxy answers Status-Server, so automation keeps sending example.net requests to it even though that route is broken. The green signal is true and the operational conclusion is false. The missing evidence is not a stronger TLS cipher. It is a receipt tied to the requested realm and the decision behind it: route selection, downstream dispatch, home-server response, backend result and final NAS action.

Hop-local error handling follows the same architecture. The draft replaces several older unwanted-packet replies with Protocol-Error. That packet applies to one RADIUS hop and must not be forwarded. It can tell a client to stop retransmitting the original packet on the current connection and try another route. It cannot certify what a downstream server would have done. A useful error becomes dangerous only when its scope is silently enlarged.

Transport confidentiality is equally bounded. TLS or DTLS authenticates the adjacent RadSec peer and protects that segment. The draft warns that an intermediate proxy may then forward the packet over ordinary RADIUS/UDP. The original client has no protocol mechanism to enforce protected transport across the whole chain. An organization that requires end-to-end secure carriage must govern every intermediary. A padlock on the first edge is not a map of the path.

Mixed transports expose who owns retry state. RADIUS/TCP and RADIUS/TLS are reliable and do not retransmit RADIUS packets on the same connection. RADIUS/UDP and RADIUS/DTLS require application retransmission. If a proxy receives over an unreliable hop and forwards over a reliable one, it must absorb the duplicate rather than forward it again. If it receives over a reliable hop and forwards over an unreliable one, the proxy becomes responsible for retransmission. When the incoming connection closes or the client reuses an Identifier and abandons an old request, the proxy must stop the downstream retry chain.

One timeout therefore has several possible owners. The client may have given up while the proxy is still retrying. The proxy may be alive while one downstream server is overloaded. The home server may have completed the operation while its response is delayed. A single “RADIUS unavailable” alarm erases the distinction needed for safe remediation.

Accounting makes the cost visible. The draft recommends a stable Event-Timestamp for the time an event occurred and says it must not change on retransmission. Repeatedly updating Acct-Delay-Time creates new RADIUS packets for substantially the same event, often with new Identifiers. At a TLS-to-UDP boundary, the proxy may start a separate downstream retransmission sequence for each packet while the slow server is already processing the original. Appendix A.2 shows one fact multiplying into IDs 101, 102 and 103.

Stable event time helps the server recognize duplicates. It is not a global transaction key and does not promise exactly-once accounting. The operational record still needs to connect the event, packet identity, hop, retry owner and server decision.

Even a failed cryptographic validation needs scope. Appendix A.1 describes a race in which an Identifier is reused after timeout and a late response to the old request is validated against the new request. The Response Authenticator fails, but the peer need not be hostile. The correct action is to discard that response and keep the connection open. Packet failure, connection failure and peer compromise are three different conclusions.

The draft is an active Working Group Internet-Draft with an intended Proposed Standard status. Revision 18 was posted on 30 September 2026 and expires on 3 April 2027. It has no RFC number. Its proposed replacement of experimental RFC 6614 and RFC 7360 will matter only if the work is approved and published. Nothing in the record proves current deployment, product behavior or an outage at a named roaming federation.

The safe claim is narrower and more useful: this identified RadSec peer answered on this connection at this time. For a service claim, add the realm selected, downstream route, home-server or credential-backend response, authorization decision, NAS enforcement and observed user outcome. Running code belongs at the end of that chain, not inside the color of its first light.

Sources