Summary
- In EAPS, one master observes a special Control VLAN, blocks or unblocks a secondary ring port and orders bridge tables to be flushed. Those actions restore a loop-free forwarding opportunity; they do not directly observe every protected VLAN or user transaction.
- Recovery evidence should preserve the fault signal, master decision, port transition, distributed flush, relearning epoch, packet behavior and application outcome separately. “Ring complete” is a protocol state, not the last receipt in that chain.
The frame came home before the service did
A health-check frame leaves the master node's primary port, crosses the ring and returns on the secondary port. The fail timer resets. The state is normal. On a diagram, the circle is whole again.
That is a precise and valuable observation. It is also narrower than the conclusion an operations screen may invite. The frame travelled on the Control VLAN. It did not test every protected VLAN, every learned MAC destination, the paths before the ring or after it, or the application whose customers are waiting. It proved what it observed: continuity for one control mechanism at one moment.
RFC 3619 was published in October 2003 as an Informational description of EAPS Version 1, a technology invented by Extreme Networks. It explicitly does not specify an Internet Standard. The document says EAPS can converge in less than one second and often in less than 50 milliseconds. That is the RFC's description of the mechanism, not a measurement of a present network and not a promise about a current product.
The useful leadership question is therefore not whether the old mechanism was fast. It is who was allowed to translate a control observation into forwarding authority, what state that decision invalidated, and where independent evidence had to take over.
One master owns a narrow gate
An EAPS domain occupies one Ethernet ring. Protected VLANs are configured on the ring ports, one node is the master, and the others are transit nodes. The master designates a primary and secondary port. In normal operation it blocks non-control frames on the secondary port, making the physical ring appear loop-free to ordinary Ethernet switching and learning. The Control VLAN is allowed through.
When the master concludes that the ring has failed, it opens the secondary port. Frames can then use the surviving direction. This is a deliberately concentrated authority: one node's state machine changes whether a whole class of data frames may cross a particular gate.
Concentration is not automatically a defect. A ring needs one coherent answer to the loop question. The mistake is allowing the authority to become vague. A defensible configuration names the domain, current master, Control VLAN, exact protected-VLAN set, primary and secondary ports, timer values and configuration version. Without that scope, “the master opened the ring” can sound like a statement about the whole network when it is only a statement about one configured recovery domain.
This is Heng Lu's distinction between a record and reality in mechanical form. The master may author the domain's forwarding permission under the shared rules. It does not author the physical link, the learned destination, or the customer's outcome. Running code gives the decision effect; it does not enlarge the meaning of the observation that triggered it.
Alert and absence are different witnesses
EAPS has two fault paths. A transit node that detects link-down immediately sends a LINK-DOWN control frame to the master. Separately, the master sends periodic health-check frames. If a health frame does not return before the fail-period timer expires, the master enters ring-fault state. Polling backs up the alert path if an alert is lost.
These signals should not be collapsed. A link-down frame is a positive report from one transit node about one local port. A timer expiry is an inference from absence: the expected control frame did not return in time. Both can rationally trigger the same safety action while retaining different uncertainty.
The receipt should therefore say which witness acted, which sequence or port it concerned, when the observation was made, how the timer was configured and whether the other signal agreed. “Ring fault” is the result of a decision function, not a complete causal diagnosis.
That separation matters during partial or asymmetric failures. It also matters when the Control VLAN and protected traffic see different conditions. A health frame that returns does not certify capacity, loss, ordering or reachability for another VLAN. A missing health frame does not enumerate which services were harmed. The control mechanism should remain authoritative only for the decision it was designed to make.
A flush is controlled forgetting
On a fault, the master opens the secondary port, flushes its bridging table and tells the transit nodes to flush theirs. Each switch then starts learning the new topology. This is a powerful recovery design because it refuses to preserve locations learned on a path that may no longer exist.
But the flush is not the recovery outcome. It is the start of a new knowledge epoch. Immediately after forgetting, the switches possess less forwarding knowledge than before. Unknown destinations may be flooded until traffic teaches the bridges where MAC addresses now reside. Quiet destinations may remain unknown. The first learned location can arrive through traffic that is itself transient.
Operators need at least four timestamps: flush command issued, flush observed at each node, first correct relearn for each critical destination, and first successful user transaction. They should also observe loss, duplication, disorder and unexpected flooding. RFC 4427 later gave useful general vocabulary: detection, correlation, notification, recovery switching and total recovery time are different intervals. It defines hitless switching more strictly—no loss, duplication, disorder or bit errors.
The difference is not pedantry. A port can switch inside 50 milliseconds while a low-rate service takes longer to repopulate state. A bridge can relearn quickly while an application session remains broken. The number attached to convergence must say where its clock starts and which observable condition stops it.
Restoration is a second risky transition
When the physical ring returns, the master is still in ring-fault state and its secondary port remains open. The master continues sending health frames. Once a frame comes back on the secondary port, it returns to normal, blocks non-control traffic on that port and orders another flush and relearn.
RFC 3619 identifies a dangerous interval before that decision. A transit node may already see its local link restored while the master has not yet declared the ring complete. With the master's secondary port still open, the physical loop could return. The transit node therefore blocks the protected VLANs on the restored port and enters PRE-FORWARDING. Only after the flush instruction does it unblock and return to normal.
PRE-FORWARDING is an unusually honest state name. It says that link-up is not forwarding permission, and forwarding permission is not yet safe restoration. The protocol preserves a gap in which new physical reality must be reconciled with distributed control state.
Failure switching and restoration should consequently have separate event identifiers. Restoration can create a second outage or loop risk even when failure protection worked perfectly. Record the restored-link observation, entry into PRE-FORWARDING, returned health sequence, master reversion, block action, flush epoch, protected-VLAN unblock and user outcome. Do not overwrite the failure record with a final green state.
Several domains can share the steel and not the truth
RFC 3619 allows a switch to belong to several rings, requiring one EAPS instance per protected ring. It also allows several EAPS domains on the same ring, each with its own master and protected VLANs. This supports spatial reuse of bandwidth. It also makes the word “ring” ambiguous.
One domain can be normal while another is failed or pre-forwarding. Their Control VLANs, timers, master identities and protected-VLAN sets can differ. A physical link can be common evidence without making the state machines identical. Dashboards should therefore key every event by domain and VLAN scope, not merely by chassis or fibre.
Later recovery frameworks use the idea of a recovery domain for the same reason: recovery has boundaries. A mechanism inside the boundary does not automatically protect its edges, upstream paths, downstream paths or higher-layer service. If several layers react to the same event, hold-off and reversion policy must prevent one recovery action from fighting another.
The ordinary service receipt begins where ring state ends. It tests representative traffic at the ingress and egress that matter, then checks the application. A complete circle in a topology view is not evidence for dependencies that the circle does not contain.
A control frame is not self-authenticating authority
RFC 3619's security section is direct: someone with physical access to the Ethernet connections could forge bridge or EAPS frames and disrupt the network. It recommends link encryption or suitable upper-layer protection where active attacks are a significant risk.
The point is broader than malicious access. A frame type tells the receiver how to parse a message; it does not by itself prove that the sender should exercise the requested authority. Later survivability work similarly warns that bogus or captured fault indications can falsely trigger recovery and that management commands require authorization.
An operations design should bind the signal to an allowed physical interface, domain, source identity, configuration epoch and expected sequence. It should alarm on impossible state transitions and retain the triggering frame or counter evidence. Security cannot be reduced to “the packet arrived on the control VLAN,” because the packet's consequence is to change forwarding for other traffic.
This does not make automated protection undesirable. It makes automation accountable. The more quickly one signal can invalidate distributed state, the more precisely the signal and its scope must be recorded.
Measure the recovery the customer bought
The minimum useful receipt chain is physical observation, alert or polling result, master decision, port action, flush distribution, FDB epoch, relearn, packet behavior and application result. Each stage has a different owner and a different rollback.
The second-order effect of weak evidence is bad tuning. If only the master-state transition is measured, teams can make the timer ever shorter while hiding false switches, repeated flushing or slow relearning. If only application availability is measured, engineers lose the internal evidence needed to locate the delay. Both views are necessary; neither should impersonate the other.
The third-order effect is authority drift. A narrow health mechanism becomes “the network is healthy,” and the party operating it gains the power to close incidents that its sensors never observed. The remedy is not a larger dashboard label. It is a smaller claim and a longer evidence chain.
EAPS gives a clear example. The health frame may declare the configured ring path complete. The master may safely close the alternate gate. The bridges may begin learning again. The service still deserves its own witness.
Sources
- https://www.rfc-editor.org/rfc/rfc3619.html
- https://www.rfc-editor.org/rfc/rfc3619.txt
- https://www.rfc-editor.org/info/rfc3619/
- https://datatracker.ietf.org/doc/rfc3619/
- https://www.rfc-editor.org/errata_search.php?rec_status=0&rfc=3619
- https://www.rfc-editor.org/rfc/rfc3386.html
- https://www.rfc-editor.org/rfc/rfc3469.html
- https://www.rfc-editor.org/rfc/rfc4427.html
- https://www.rfc-editor.org/rfc/rfc5654.html
- https://www.rfc-editor.org/rfc/rfc5921.html
- https://www.rfc-editor.org/rfc/rfc6371.html
- https://www.rfc-editor.org/rfc/rfc6372.html
- https://www.rfc-editor.org/rfc/rfc6378.html
- https://www.rfc-editor.org/rfc/rfc6974.html
- https://www.rfc-editor.org/rfc/rfc8227.html
- https://www.rfc-editor.org/rfc/rfc9522.html
- https://heng.lu/running-code-primary-the-patch-needed-to-preserve-the-internet-original-design/
- https://heng.lu/minimum-initial-specification-localized-future-decision-voluntary-adoption-internet-coordination-system/
- https://heng.lu/on-reality-layers-symbolic-power-and-why-clarity-feels-so-hostile/
- https://heng.lu/on-authority-belief-and-the-internets-addressing-system/
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
