Summary
- RFC 7606 does not repair malformed routing information; it chooses how much valid state must be sacrificed to contain it.
- A peer that remains Established is not sufficient evidence of health: every NLRI in a treat-as-withdraw UPDATE may have disappeared from the receiving RIB and forwarding plane.
Imagine one BGP UPDATE carrying several ordinary destinations and one malformed path attribute. Under the base error model, the receiver sends a NOTIFICATION, terminates the session and loses every route learned over it. Under revised handling, the receiver may keep the session Established but process every destination in that particular UPDATE as withdrawn. The dashboard turns green. Those destinations do not.
That is the crucial change introduced by RFC 7606. The standard is often described as a way to stop malformed UPDATEs from flapping BGP sessions. It is, but that description is incomplete. The revision creates a hierarchy of fault-containment decisions. Each decision destroys or ignores a different unit of routing state, and each leaves a different evidentiary burden for the operator.
The base specification, RFC 4271, defines UPDATE error subcodes for malformed attribute lists, missing mandatory attributes, invalid origins, bad next hops and malformed AS paths. RFC 7606 characterises its original response as a session reset: send a NOTIFICATION and terminate the session. That action is simple and conservative, but it makes one bad attribute everybody's problem. Valid routes exchanged over the same session are removed with the offending information.
RFC 7606 orders four responses from strongest to weakest: session reset, disabling one AFI/SAFI, treat-as-withdraw, and attribute discard. These are not four names for graceful degradation. They are four different decisions about where uncertainty is allowed to stop.
Four blast radii, not one repair
A session reset sacrifices the whole adjacency. It remains necessary when the receiver cannot safely locate and parse the routing information or when the applicable specification still requires termination. The receiver cannot contain an error to routes it cannot identify.
An AFI/SAFI disable sacrifices an address-family boundary. RFC 4760 allows a receiver confronting an incorrect MP_REACH_NLRI or MP_UNREACH_NLRI attribute to delete routes for the affected family and ignore subsequent routes for that AFI/SAFI over the session, with session termination still available. The TCP and BGP session may remain, and other negotiated families may continue, while one family's routing authority is effectively suspended.
Treat-as-withdraw sacrifices every route carried in the offending UPDATE. The receiver acts as though all those routes had been explicitly withdrawn and removes them from Adj-RIB-In. It does not isolate the one prefix an operator happens to be investigating. UPDATE packing matters: several NLRIs sharing one attribute set share the same containment fate.
Attribute discard sacrifices only the attribute and processes the rest of the UPDATE. RFC 7606 permits this only when the attribute has no effect on route selection or installation. Even that boundary depends on real policy. An attribute that is normally informational can become decisive if local policy matches it. “The protocol does not use it by default” is not proof that the network does not use it.
This hierarchy explains why the word fault tolerance needs discipline. The receiver is not repairing malformed bytes or recovering the sender's intended meaning. It is choosing the least destructive action that still preserves a defensible routing state.
Parseability is the line that policy cannot negotiate away
Treat-as-withdraw works only when the receiver can find and parse the entire relevant NLRI field, including MP_REACH_NLRI or MP_UNREACH_NLRI where applicable. If corrupt lengths or structure prevent that, the receiver cannot know which routes to withdraw. RFC 7606 leaves the stronger RFC 4271 or RFC 4760 response in place: reset the session or disable the affected family.
This distinction is more important than a feature checkbox. A platform may advertise enhanced error handling and still reset on a class of errors that prevents safe parsing. Conversely, an operator-configured attribute filter may deliberately choose treat-as-withdraw for a parseable UPDATE that is syntactically acceptable but carries an attribute the receiver refuses to trust.
Multiple errors also resist optimistic interpretation. RFC 7606 says that when different errors prescribe different responses, the strongest applicable action wins. An attribute that could be discarded does not rescue the same UPDATE from another error that requires withdrawal or reset.
The attribute-specific rules show why a universal “malformed means withdraw” policy would be wrong. RFC 7606 assigns treat-as-withdraw to revised errors involving ORIGIN, AS_PATH, NEXT_HOP, MULTI_EXIT_DISC and LOCAL_PREF. It assigns attribute discard to specified ATOMIC_AGGREGATE and AGGREGATOR errors. Duplicate MP_REACH_NLRI or MP_UNREACH_NLRI attributes still cause a Malformed Attribute List NOTIFICATION, while later duplicates of most other attributes are discarded after the first occurrence.
Later specifications inherit the same decision framework but define their own syntax. RFC 8092 says a Large Communities attribute whose value length is not a non-zero multiple of twelve is malformed and triggers treat-as-withdraw; repeated community values alone are not malformed and are silently deduplicated. RFC 7607 proscribes AS 0 in specified BGP fields, but sends AS_PATH and AGGREGATOR cases to RFC 7606 and AS4_PATH and AS4_AGGREGATOR cases to RFC 6793. The correct action comes from the attribute's specification and context, not from a generic desire to keep the peer up.
The current IANA BGP Parameters registry provides the path-attribute type codes and UPDATE error subcodes. Registration identifies the vocabulary. It does not authorize an operator to discard a selection-relevant attribute or reinterpret an error contrary to its defining standard.
A green session can conceal a real outage
Treat-as-withdraw removes the broad collateral damage of a reset, but it can still make every destination in the affected UPDATE unreachable or suboptimal. It can also create inconsistent state inside one AS. RFC 7606 warns that applying the action on an iBGP session may produce long-lived forwarding loops or black holes if different routers retain different usable paths.
That is why dropping the malformed UPDATE is not an acceptable shortcut. BGP is incremental. Ignoring a new UPDATE without withdrawing its NLRIs can leave an older, now-invalid route installed. Treat-as-withdraw makes the loss explicit in routing state; silent message disposal can preserve a false past.
The operational evidence must therefore start before the RIB. RFC 7606 requires diagnostic facilities that identify the involved NLRI and retain the entire malformed UPDATE. Without the original PDU, the receiver's action may be visible but its justification is not. A counter labelled “malformed updates” cannot show which attribute was broken, which routes shared the message, whether the NLRI remained parseable or why one router reset while another withdrew.
RFC 7854 gives BMP Route Mirroring a useful forensic role. A monitored router may send verbatim received messages and mark an errored PDU that was treated as withdrawn. But mirroring is optional, may be sampled, may report messages lost and carries buffering and convergence trade-offs. A collector is evidence only when its capture policy, loss signals and timestamp relationship are known.
Follow the error from wire to packet
A defensible investigation begins with the exact peer, direction and negotiated family. Preserve the UPDATE, or an explicitly loss-bounded equivalent, with the attribute code, flags, declared and actual length, and every NLRI it contains. Record the receiver's software release and effective configuration after peer-group inheritance. The same command present on two routers does not prove the same parser or action.
Next, record the verdict. Did the receiver send a NOTIFICATION and reset? Did it disable one AFI/SAFI? Did it mark all routes in the UPDATE as withdrawn? Did it remove one attribute and accept the rest? A log line should be tied to the exact message, not merely to a neighbor that emitted many UPDATEs.
Then compare state at each layer. In Adj-RIB-In or a hidden/rejected view, prove which paths were removed or retained. In Loc-RIB, prove whether an alternative won. Across route reflectors, prove that the AS did not split into inconsistent beliefs. On every material egress, check whether the new state was re-advertised or contained. In the FIB or hardware table, prove what forwarding action was actually installed. Finally, send controlled packets in both directions and inspect the failure modes that a green session cannot reveal.
Current vendor documentation illustrates why this must be release-specific. Cisco IOS XR documents malformed-UPDATE logs that name the neighbor, message length, action, attribute details, family and NLRI, including TreatAsWithdraw and DiscardAttr examples in its BGP implementation guide.
Junos BGP error documentation describes reset, treat-as-withdraw, hidden-route treatment and diagnostic logging. Nokia SR OS documents an update-fault-tolerance control that chooses revised handling for specified non-critical errors instead of legacy reset behavior.
These sources prove implementation choices exist. They do not establish one cross-vendor default, one log format or one safe configuration.
Recovery requires a clean message, not a quiet alarm
The durable fix belongs at the source of the malformed information. RFC 7606 recommends tracing a malformed attribute received over iBGP back to the ingress router where it was created or received externally, then filtering or correcting it there to restore consistency.
Recovery should show a clean replacement UPDATE, not merely the disappearance of the error counter. The receiver should accept the corrected routes, remove hidden or withdrawn error state, converge consistently across reflectors, install the intended FIB entries and restore packet delivery. If an attribute filter was introduced as containment, its removal or continued scope needs an explicit decision; emergency filters easily become permanent policy whose original defect nobody remembers.
The leadership question is not whether session uptime is valuable. It is who may trade one kind of failure for another, under what evidence. Resetting every route may be too destructive. Discarding policy meaning may be too permissive.
Treating one UPDATE as withdrawn may be the correct middle action—but only when the receiver can name every route it is sacrificing and prove what happened next.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
