Summary

  • RFC 5171 separates physical carrier, correct Layer 2 neighbour pairing and continuing UDLD evidence. Light on a fibre is not proof that transmit and receive strands lead to the intended peer.
  • Normal mode preserves silence as undetermined; aggressive mode may convert bounded silence into a protective shutdown only after earlier bidirectionality, expiry and repeated attempts to restore the relationship.

A fibre receiver can see light from the wrong place. The transmit strand can lead to one switch while the receive strand comes from another. Every interface may look electrically alive while the Layer 2 relationships form a dangerous loop. RFC 5171 begins with that very practical insult to a familiar dashboard: “up” can be true at one layer and misleading at the next.

UDLD adds a different receipt. A device advertises a Device-ID and Port-ID. Its neighbour returns the pairs it actually heard on the same interface in an Echo TLV. The question is no longer merely whether photons or link pulses exist. It is whether the expected port identity made the round trip through the intended relationship.

That receipt is still bounded. It is carried by one control protocol, processed by a device CPU and retained in a timed cache. It does not exercise every VLAN, frame size, queue, forwarding entry or application path. It cannot turn a healthy control exchange into a universal certificate for the link.

A cache entry is a dated observation

UDLD sends periodic hellos and retains learned neighbours only for a holdtime. A later hello replaces the previous entry and resets its timer. Disabling the interface or protocol, or resetting the device, clears relevant state and asks neighbours to flush theirs.

This matters because a command such as show udld is not a photograph of the cable. It is a view of a state machine assembled from recent messages, local timers and configuration. A neighbour row proves that an advertisement was accepted within the cache window. It does not prove that the state is current at the instant an operator reads it, or that data traffic follows the same fate.

When a new neighbour appears or asks to resynchronise, the device sends a train of echoes. The RFC explicitly assumes that N messages are enough for at least one to traverse the link despite possible drops. That is an engineering assumption, not an observation about the cause of every missing reply.

The implementation runs in the control plane. CPU load, task scheduling, event processing and software defects bound detection time. The protocol therefore used conservative timers to reduce false positives. The detector’s clock is part of the evidence.

Normal mode refuses to invent an event

Normal mode is event-based. It can act when received messages reveal proper pairing or an explicit mismatch. When no useful information arrives, including after bidirectional communication disappears, it does not manufacture a diagnosis. It marks the link undetermined.

That restraint is not weakness. Silence has several possible causes: a cut fibre, a stuck interface, loss in both directions, high error rate, a duplex problem, an overloaded control plane, disabled UDLD, an incapable neighbour or a configuration change whose flush message was lost. Missing messages collapse these causes into one observation.

Current Cisco guidance makes the same point operationally. Neighbour information can age out because of errors or a duplex mismatch; that packet loss does not by itself mean the link is unidirectional. Normal mode therefore does not disable merely because the cache expires.

Aggressive mode buys containment with availability

Aggressive mode uses a different policy. If a relationship was already established as bidirectional, the neighbour evidence expires and repeated last-resort probes still cannot restore it, the local device may disable the port. Current Cisco guidance describes eight one-second attempts on relevant implementations before errdisable. RFC 5171 deliberately does not make one missing message sufficient.

The conversion is defensible only because more context is present: the link had a known prior state; silence lasted beyond the declared window; active retries failed; the link remains physically up; and the operator selected a context where neighbour communication loss is itself unacceptable. The RFC recommends aggressive mode only for limited scenarios, typically point-to-point links.

The action is protective, not diagnostic. Disabling a port can prevent a forwarding loop or stop traffic entering a black hole. It still does not prove which strand, transceiver, ASIC, process or remote configuration failed. The errdisable event says why the state machine acted, not what a field engineer will find.

The race belongs in the record

UDLD does not act alone. STP or RSTP can change the forwarding topology while the UDLD timer is running. Cisco’s current guidance warns that a classic STP timing approximation does not guarantee UDLD will act before RSTP transitions. Platform, release, link type, timers, topology and fault timing all affect the race.

A responsible incident record therefore keeps separate timestamps for carrier state, last accepted hello, cache expiry, resynchronisation probes, port disablement, STP/RSTP changes, alternate-path forwarding and application recovery. “UDLD caught it” is too compressed. “The service recovered” is a later claim again.

The protocol’s enduring lesson is not that aggressive is better than normal. It is that an operator must declare when absence remains uncertainty and when a bounded absence is costly enough to justify fail-closed action.

Sources