Summary

  • BFD and its BGP integration carry documented defects in the specification record, including an RFC 5880 timeout condition that fails to signal in a specific unidirectional-failure scenario, and an implementation divergence on the diagnostic field that Jeffrey Haas himself reported.
  • The strongest repair mechanisms — RFC 9384's named "BFD Down" subcode and the BGP BFD strict-mode draft's negotiated capability — are implemented in shipping software, but the deployment evidence is overwhelmingly vendor-authored and self-reported.
  • The evidence layers that would make a repair durably verifiable — independent interoperability testing, published test methodology, and post-revision field data — are the ones thinnest in the public record.

The mechanism, and where its text broke

BFD is a lightweight hello protocol: two systems exchange control packets at an agreed interval, and if the expected packets stop arriving within the Detection Time, the session goes Down and the routing protocols riding above it react. RFC 5880 defines the base protocol; RFC 5882 defines how BFD attaches to applications like BGP. The design intent is precise — a failure should be detected fast, attributed unambiguously, and acted on predictably.

The errata record shows how hard that precision is to hold. Erratum 5205 against RFC 5880, reported by Dave Katz — one of BFD's original authors — concerns the state machine's timeout condition. The original text requires that the session be in Init or Up state for a timeout to register. The consequence, documented in the erratum's notes: if system A signals AdminDown and the link then fails unidirectionally, system B gives no timeout indication in its outgoing control packets. A failure-detection protocol whose state machine can be maneuvered into a state where it stops signaling failures is, in the most literal sense, a mechanism whose repair mattered.

A second erratum touches the article's subject directly. Erratum 7240, reported by Jeffrey Haas, concerns the initialization of bfd.LocalDiag, the diagnostic field a system presents about its own state. The erratum's notes record that several implementations reset the value to zero when the session returns to Up, while also stating that the RFC text itself is correct and reflects the authors' intent. That is a specific and instructive category of defect: not a specification bug, but a documented divergence between what the specification says and what shipping code does — with the diagnostic field, the operator's primary window into why a session failed, as the casualty.

What each layer of evidence can and cannot prove

The working group's own activity acknowledges the repair is ongoing. The BFD working-group document list shows active revisions of all four base specifications — draft-ietf-bfd-rfc5880bis, 5881bis, 5882-bis and 5883-bis — with a December 2026 milestone to advance RFC 5880 and 5882 toward Internet Standard. A bis revision is the standards process's formal admission that the earlier text accumulated defects, ambiguities and dead references worth fixing in one pass. RFC 5882's own errata illustrate the smaller end of that spectrum: errata 8921 and 8922, reported by Xiao Min in May 2026, correct internal cross-references that drifted when a new section was inserted — and erratum 8921 was verified by the IESG, the closest thing the errata process offers to a verified repair. It verifies a citation, not a runtime behavior. The distinction matters.

On the implementation side, the record is richer but structurally different in kind. RFC 9384, authored by Haas and published as a Proposed Standard in March 2023, defines Cease NOTIFICATION subcode 10 so that a BGP speaker tearing down a session because BFD went Down tells its peer exactly that, rather than leaving the cause ambiguous. The RFC's own rationale states the receiving speaker "will then have an understanding that the connection is being terminated because of a BFD-detected issue and not an issue with the BGP speaker." The implementation report on the IETF community wiki lists two implementations — Juniper 22.3R1, with Haas as the named contact, and Arista 4.29.0, with Bill Fenner — reported conformant on sending the subcode, on recording the reason in operational state when total loss of connectivity prevents the NOTIFICATION itself, and on carrying the reason in the RFC 8538 Hard Reset data portion.

Read carefully, that report proves implementation and self-assessment. Both entries are vendor-authored and self-reported. No test methodology, no capture files, no third-party verification appear in the record. The same structure repeats for BFD strict-mode, the mechanism that prevents a BGP session from establishing until the BFD session beneath it is up. The IDR draft — revision 19 dated 26 August 2026, with Haas listed among its authors — defines a negotiated capability plus hold-down and dampening timers to keep the fix itself from introducing session flapping. The IETF 118 vendor implementation report slides describe successful interoperability between Junos and SRoS with the strict-mode capability exchanged, and simultaneously document a proprietary implementation that supports no dynamic signaling, requiring static configuration on both sides — the slides themselves call this deployment-complicating.

Shipping documentation closes part of the gap. Juniper's Junos documentation describes strict mode waiting for BFD Up before the BGP session establishes, with the wait bounded by the BGP hold-time or a configured interval, and states that on expiry the session resets with a BFD Down notification — tying the vendor's deployed behavior directly to RFC 9384's subcode. FRR's documentation describes an independently developed open-source implementation of the same strict-mode semantics, which is meaningful corroboration across implementation traditions, though still documentation rather than test evidence.

The rung that is missing

What the record does not contain is the layer above self-report: independent, methodologically documented testing of the repaired behavior under the failure conditions that exposed the original defects. An RFC 9978 comparison is instructive. BFD Stability, an experimental mechanism to detect packet loss within a BFD session, states its own evidentiary position plainly: it is on the Experimental track "because there are no known implementations or proof of concept." The RFC Editor's process here does something the implementation-report process does not — it labels the evidence gap in the document itself. For the repaired mechanisms this article follows, no equivalent label exists; the gap is filled by conformance tables and slide decks, which are the vendor's words about the vendor's code.

This is not an accusation of bad faith. Haas's own reported erratum 7240 demonstrates that implementers surface their own divergences into the standards record, which is the system working. The finding is structural: the standards ecosystem has strong mechanisms for recording what was specified, moderate mechanisms for recording what vendors say they implemented, and weak mechanisms for recording what was independently verified to still work after the fix.

Between the IESG-verified citation erratum and the shipping Junos and FRR manuals lie two rungs — self-reported conformance and vendor-narrated interop — and nothing above them for the specific failure-signaling paths at issue.

For operators, the practical translation is direct. A repair is durably verified when the failure that motivated it can be reproduced, the repaired behavior observed, and the observation reproduced again by someone other than the repairing party. For BFD strict-mode and the BFD Down subcode, the public record supports the first two steps for two major vendors and stops there. The December 2026 Internet Standard milestone for RFC 5880 and 5882 will certify the specification's maturity; it will not, by itself, certify the durability of any implementation's repair.

Sources