Summary

  • RFC 5247 warns that either an EAP peer or authenticator can reboot or reclaim resources and clear cached key material even while the other side's agreed lifetime has not expired.
  • Cached reuse needs a live, protected proof of joint state. If no shared key remains to protect reconciliation, the safe recovery is a fresh authentication path—not an inference from a timer.

At 09:17, both dashboards were telling the truth

The peer showed a cached key with fourteen minutes remaining. The authenticator showed no matching entry. A maintenance restart had cleared volatile state while the peer was asleep. When the peer returned, it selected the key it had used successfully before and attempted the abbreviated path. The authenticator rejected the attempt because, from its perspective, that state no longer existed.

Neither clock had lied. The peer's record really was younger than its configured limit. The authenticator really had forgotten it. The contradiction appeared only because an operations screen had turned two local stores into one global statement: “key valid.”

RFC 5247 refuses that compression. In its requirements for a Secure Association Protocol, it notes that either party can reboot or reclaim resources, clearing part or all of a key cache. It then draws the operational conclusion: key-lifetime negotiation cannot guarantee that the caches remain synchronized, and a peer may not know whether the authenticator still holds a key until it tries to use it.

That sentence is more than an implementation warning. It identifies an authority boundary. A lifetime is a rule about how long state may be used if the state still exists. It is not a lease on another machine's memory.

EAP produces a chain, not one green light

The framework separates three linked phases. An EAP method runs between peer and EAP server and may export keying material such as the Master Session Key and Extended Master Session Key. A AAA exchange can transport relevant keying material and authorization information to an authenticator. A lower-layer Secure Association Protocol then proves possession, negotiates capabilities and creates the transient keys that protect traffic.

These stages share inputs but answer different questions. Method success can establish peer/server authentication and key derivation. AAA transport can deliver material to a named authenticator. The Secure Association Protocol can show that the live peer and authenticator both possess the selected material and can derive fresh transient session keys. Activation can put those keys into use. Protected packets can then cross the lower layer. Application traffic may still fail farther downstream.

A cache does not collapse the stages. It preserves selected state so a later session need not repeat all earlier work. The optimization is valuable precisely because the earlier authentication was expensive. Yet the reuse attempt becomes a new live claim: both parties still possess the right state, under the right name and scope, now.

The peer's local hit proves only half of that claim.

Expiry and existence are independent dimensions

Operators often model a cached credential or key with a single deadline. Before that time it is “valid”; after that time it is “expired.” The model omits existence. A record can be within lifetime but absent, or present but outside policy lifetime. Those are different failures and should not share one status.

RFC 5247 describes several reasons for divergence without assigning blame. A peer or authenticator may reboot. Either may reclaim resources. A cache can lose some entries rather than all of them. Server selection can also change: an authenticator may reach a different backend EAP server whose persistent state is not shared with the first. The framework does not authorize an observer to call every reuse failure revocation, credential compromise or malicious rejection.

An accurate state model therefore needs at least two local observations and one joint receipt. The peer can attest that it retains key name K under scope S until local time T. The authenticator can report the same facts for its own store. A recent protected exchange can prove that both held compatible material at that moment. None of those statements guarantees future retention.

This is the distributed-systems version of Heng Lu's reality-layer argument. The useful record is not the institution's or component's preferred label. It is the smallest fact the component actually observed.

The secure association must name what it is reusing

Caching introduces another ambiguity: the parties may share more than one key of the same type. RFC 5247 therefore requires the Secure Association Protocol to name the key used in its proof-of-possession exchange. A successful lookup by peer identity is not enough. Both sides must select the same key context.

The protocol also needs fresh transient session keys. Reusing exported EAP material directly as traffic keys would risk reusing the same session keys. RFC 5247 requires a post-EAP Secure Association Protocol where caching is supported, normally mixing nonces or counters into fresh unicast and, where applicable, multicast TSKs.

The evidence ladder is consequently longer than “cache hit.” It includes a selected key name, mutual possession proof, fresh inputs, derived TSK identifiers, activation and protected-packet evidence. If the abbreviated path fails before joint proof, the system knows only that the two current views did not compose. It does not yet know which store was wrong or why.

Reconciliation itself needs authority

The obvious response to divergence is to synchronize. But a synchronization message that changes security state must itself be protected. RFC 5247 recommends key-state resynchronization through the Secure Association Protocol or a lower-layer indication. That works when the parties still jointly hold a key suitable for protecting the exchange.

When they do not, secure resynchronization is not possible on the strength of the missing state. An unsigned “I forgot; please restore K” message cannot inherit the authority of K. The framework points to alternatives such as initiating EAP re-authentication after a timer.

That fallback matters because it preserves the direction of proof. Fresh authentication may produce new keying material and a new authorization decision. It does not ask an unauthenticated cache-repair request to recreate the authority it is supposed to prove.

The same discipline applies to scope. RFC 5247 recommends that the Secure Association Protocol let each party determine the scope of the other's cache and negotiate usage restrictions. A key retained for one authenticator, port, service set or traffic profile should not silently become authority for another merely because its bytes still exist.

A cache miss is not a credential verdict

The user-visible symptom may be a delay, a fresh login or failed connectivity. The tempting operational shortcut is to translate any abbreviated-path failure into “bad credentials.” That destroys useful evidence and can trigger the wrong response.

If the authenticator rebooted, rotating the user's long-term credential does not restore the lost cache entry. If the peer selected the wrong key name, increasing cache capacity does not fix selection. If the backend changed, blaming the access point misses server-state topology. If policy now requires full authentication, treating the fallback as an outage hides a deliberate control.

Record the earliest boundary that failed: no matching key name, possession proof failed, fresh TSK derivation failed, activation failed, access enforcement rejected, or protected traffic produced no service. Preserve “unknown” until the next receipt narrows the cause.

Resilience begins by testing asymmetric loss

Happy-path tests usually retain or clear both sides together. Real failures are asymmetric. A useful test matrix has four cases: both peer and authenticator retain state; only the peer retains it; only the authenticator retains it; neither retains it. Each case should have a deterministic result, bounded delay and an observable transition to re-authentication when reuse cannot be proved.

Restart and eviction storms deserve separate testing. Cached reuse can reduce authentication load in steady state, yet a fleet restart can erase many authenticator entries at once and redirect the whole population to the full EAP path. A system sized on cache-hit traffic can fail exactly when the cache is least available.

The correct capacity question is therefore not “how many cached sessions fit?” It is “how many safe cache misses can recover per second without turning control-plane repair into a wider access failure?”

Sources