Summary
- An operator asked NANOG whether a flat IS-IS Level 2 domain with roughly 300 routers, including aggregation devices with about 100 downstream adjacencies, was practical.
- Saku Ytti reported experience with a few thousand IS-IS nodes on much older control planes, while Mark Tinka said his network had already operated at about 300 nodes in 2009.
- Tom Beecher said the scale should generally be workable, but linked the answer to geographic spread, link-state database size and SPF-delay tuning. Dan Snyder added hardware, fault-domain size and convergence targets.
- A newly indexed reply from Matthew Petach reframed the decision: protocol capacity does not determine how much of the network one automation error, fibre cut or long repair cycle should be allowed to disturb.
- IETF documents support the architectural distinction. IS-IS areas constrain flooding and shortest-path computation, while hierarchy introduces its own trade-off between scale and path optimality.
- The next design review should therefore test failure scope and degraded traffic flow, not merely prove that every router can hold the topology.
The NANOG thread began with a capacity question. An operator wanted real-world evidence for a flat Level 2 IS-IS network of about 300 routers and asked whether aggregation routers maintaining roughly 100 downstream adjacencies would be a concern.
The early replies made the numerical answer unusually clear. Saku Ytti described an IS-IS deployment with a few thousand nodes on control planes from two decades ago. Mark Tinka said a network he operated reached approximately 300 nodes in 2009 on Cisco CRS-1 and ME3600X equipment and Juniper M320, T320 and MX480 platforms. Tom Beecher likewise said 300 nodes in one flat Level 2 area generally should be feasible.
None of those messages is a vendor benchmark, and the questioner did not publish the proposed topology, router mix, LSDB size or convergence objective. The thread therefore cannot establish a universal ceiling. It does establish something more useful than a guessed threshold: experienced operators did not treat 300 as the limiting fact.
The qualification is where the design starts
Beecher attached two conditions to the reassuring answer. Geographic distribution can affect SPF timing, and LSDB size can require tuning. He also made a sharper architectural point: if the database is large enough to be the central concern, the decision to remain flat may deserve review.
Dan Snyder reduced that qualification to three variables: hardware, acceptable fault-domain size and required convergence time. These are not interchangeable. A fast control plane can calculate routes quickly while the failure domain remains too broad. A small LSDB can coexist with a topology whose remaining paths cannot carry the traffic after a major cut.
Matthew Petach's later reply made this separation the news event. He proposed beginning with the blast radius of recurring failures—automation defects, human input mistakes, hardware faults and fibre cuts—and asking where “blast doors” belong. He agreed with the technical scale judgment: 300 nodes, and considerably more, can be within IS-IS capability. His objection was to treating that capability as evidence of a good operating design.
The phrase “blast door” is a metaphor, not an instruction to add areas mechanically. It identifies the control objective: one class of error should have a defined boundary, and operators should retain a place to steer traffic when shortest-path behaviour alone no longer produces an acceptable outcome.
Failure duration changes the value of a boundary
A failed optic inside a staffed facility and a damaged long-haul cable do not impose the same operating problem. The first may be replaced quickly. The second may remain constrained long enough for traffic patterns, capacity reservations and customer commitments to matter more than nominal convergence.
Petach used transoceanic links to expose this difference. Spare subsea paths are costly, physical diversity can conceal shared fate, and the few links left after a regional disruption may be reachable but unable to carry every flow that a flat topology sends toward them. That is a capacity-allocation problem after convergence, not merely a convergence problem.
The article should not turn that example into a claim about a particular cable system. The thread describes no outage and supplies no route measurements. Its operational value lies in the test it proposes: during each credible failure, can the surviving paths carry the resulting demand, and does the design provide a higher-level control point before congestion becomes the only signal?
IS-IS hierarchy limits scope, but it is not free
The standards record explains why the operator distinction is durable. RFC 5302 says an IS-IS domain can be divided into Level 1 areas connected by a Level 2 topology. Containing LSPs inside an area limits the size of the link-state database and the complexity of shortest-path computation.
That same RFC rejects a simplistic “hierarchy always wins” rule. Summarisation and abstraction can remove information needed for optimal path selection. Distributing more detailed prefixes across the domain improves visibility but consumes memory, transmission capacity and computation. The engineering choice is a trade-off between scalability and optimality, not a free reduction in risk.
RFC 9377 describes processing and flooding overhead as the eventual constraint on a single IS-IS flooding domain and identifies multiple Level 1 domains plus a Level 2 backbone as the standard scaling approach. RFC 8405 adds an operational warning: SPF back-off parameters should be consistent within the same area or level, and the appropriate values can change over the life of the network.
Together, these documents support neither a fixed 300-node limit nor permission to ignore boundaries. They support measurement: topology size, link count, update rate, LSDB growth, flooding time, SPF behaviour and the capacity of alternate paths all belong in the decision.
A useful review has four tests
The first test is technical headroom. Operators should measure LSDB size, LSP churn, CPU and memory during topology changes, adjacency establishment time and full versus incremental SPF behaviour on the slowest deployed platform—not only on the newest router.
The second is failure containment. A design review should enumerate configuration automation errors, route-policy mistakes, device loss, site loss and simultaneous long-haul cuts. For each event, it should state which routers recompute, which traffic moves, which operational team owns the response and what must remain unaffected.
The third is degraded capacity. Successful convergence is not sufficient if surviving links saturate or if shortest-path routing sends traffic toward a path operators would prefer to protect. The test must apply demand matrices to failure states and show where traffic engineering, admission control or service prioritisation enters.
The fourth is future operability. Petach asked the questioner to consider three-, five- and ten-year growth. That should include not just more nodes but new geographies, mixed hardware generations, larger maintenance domains and the number of engineers who can safely reason about the topology during an incident.
The watchpoint is the first documented failure budget
No entity can select the right architecture without the questioner's network details. A geographically compact deployment with abundant diverse capacity may rationally remain flat. A dispersed network with slow repairs, concentrated long-haul paths or strict service obligations may need boundaries well before control-plane scale becomes uncomfortable.
The thread's contribution is to separate two approvals that are often collapsed. One approval says the protocol and hardware can carry the topology. The other says the organisation accepts the topology's failure radius and can control traffic under degraded conditions.
For a 300-router proposal, the next useful document is therefore not a larger node-count anecdote. It is a failure budget: the events the network is designed to absorb, the maximum affected scope and duration, the capacity left after each event, and the control point available when ordinary shortest-path routing is no longer enough.
Sources
- Original NANOG question on a 300-router flat IS-IS domain
- Saku Ytti's historical IS-IS scale experience
- Mark Tinka's 2009 deployment example
- Tom Beecher on LSDB size, geography and SPF delay
- Dan Snyder on hardware, fault-domain size and convergence
- Matthew Petach on failure radius and design boundaries
- RFC 5302 on two-level IS-IS and the scalability–optimality trade-off
- RFC 8405 on SPF back-off timing
- RFC 9377 on IS-IS flooding-domain scale

