Summary

  • A March 2017 customer case attributes to Erik Bais and the A2B Internet team the decision to move a virtual routing platform from lab testing into production after validation. The decision is person-level evidence; the implementation remained team and company work. [4]
  • The same case reports that a full Border Gateway Protocol table converged in three to four seconds, with faster recovery from routing flaps. The figure is a vendor-published A2B result, not an independently reproduced benchmark. [4]
  • The deployment account also reports IPv6 validation, multihoming and a basis for further automation. Those details describe bounded deployment outcomes and do not show universal performance or market-wide impact. [4]
  • Internet Society independently reported that Bais opened a RIPE 76 routing-security discussion in 2018 with a presentation on persistent distributed denial-of-service sources and operator responsibility. [3]
  • The practical lesson is not that one product or person solved routing resilience. It is that continuity improves when operators test running systems, preserve role boundaries, measure recovery and use evidence about network hygiene in interconnection decisions.

A production choice with a measurable operating consequence

The most useful starting point is a bounded production decision. The 2017 case says A2B Internet tested Juniper's virtual MX, or vMX, routing platform in a laboratory before moving it into the production network. It attributes the decision to Bais and the A2B team, rather than presenting the deployment as an anonymous corporate change. It also says the platform was used for A2B's Internet-facing connections. That combination matters: it links a named operator's decision to running infrastructure and to an observed result, while still leaving implementation as team work. [4]

The official A2B account provides the narrower identity and operating context. It says Bais founded A2B in 2010 and describes work involving Internet transit, management of full routing tables, fibre connectivity and Dutch data-centre services. That page is an organisation's account of itself, so it is useful for identity and context rather than independent proof of impact. The article does not need to turn it into a broad career profile. The relevant question is how an operator approached a routing constraint that could affect continuity. [1]

The problem described in the customer case was practical. Internet routing tables were growing, and convergence was an operating constraint. Border Gateway Protocol, usually called BGP, is the system through which independently operated networks exchange information about which address ranges they can reach. A router that carries a full table holds a large set of those reachability paths. When a connection fails or a path changes, the router must process new information and settle on usable alternatives. Until that process finishes, traffic may take a worse path, pause or fail.

The case reports that A2B wanted faster convergence when a link failed. It says the team validated the new platform in the lab and then made the production move. That sequence is more informative than a product announcement. Laboratory testing creates a controlled place to check whether the system can handle expected routes, protocols and failure conditions. Production use exposes the design to real traffic and operational dependencies. Neither step guarantees future reliability, but together they show a method: establish a requirement, test it, place it into service and observe the result.

The reported three-to-four-second figure gives the case its concrete edge. A number allows readers to see what the operator was trying to change. It is not merely a claim that routing became “better” or “more resilient.” At the same time, the evidence does not provide the full test design, independent replication, traffic conditions or every comparison variable. The number should therefore remain attached to its source: a vendor-published customer account describing A2B's experience. [4]

What full-table BGP convergence means in ordinary business terms

Convergence is the period during which routers work toward a consistent view after routing information changes. Imagine that a delivery company has several roads to every destination. If one road closes, dispatchers must learn about the closure, remove the unusable option and select another route. BGP performs an analogous function across networks, although the real system is far more complex. Networks announce address ranges, attach path information and apply local policies about which routes they prefer or accept.

A full routing table contains routes to the broad public Internet rather than only a small default or selected subset. Processing that table requires memory, computing capacity and software behaviour that remains predictable under change. When a link fails, a neighbour withdraws a route or policy changes, the router may need to evaluate many affected paths. Fast processing does not eliminate every interruption, but it can shorten the interval during which reachability is uncertain.

That interval matters commercially because customers experience routing as service availability. They do not normally see BGP messages. They see applications that stall, sessions that reset, calls that break or services that become unreachable. A long convergence period can amplify the effect of a local failure because traffic waits while the routing system settles. A shorter period can reduce exposure, provided that alternate paths exist and the surrounding network is designed to use them safely.

The case also refers to faster resolution of BGP flap incidents. A route flap occurs when a route or connection repeatedly appears and disappears. Frequent changes can force routers to recalculate paths and can spread instability to neighbours. A platform that processes changes quickly may help an operator respond, but speed alone is not the whole answer. Operators also need to understand why a link or route is unstable, whether policy is correct and whether repeated announcements should be damped, filtered or escalated.

The reported three-to-four-second result is therefore meaningful as an operational observation. It suggests that A2B reduced one part of the recovery window in the described environment. It does not tell readers the complete end-to-end outage experienced by every customer. Application recovery, transport sessions, name resolution, upstream behaviour and remote networks may add their own delays. The router's convergence time is one layer in a larger chain.

For non-specialist buyers, that distinction offers a practical procurement lesson. A provider can quote a fast routing metric without demonstrating the availability of alternate paths, the quality of failure tests or the impact at the application layer. A good review asks what was measured, where it was measured, which failure was introduced, what paths were available and what users experienced. The A2B account supplies a useful number and a named production decision. It also shows why a single number should lead to better questions rather than end the assessment.

Reading the three-to-four-second result without turning it into advertising

Vendor customer stories sit in an awkward but valuable evidence category. They often contain implementation detail and named customer statements that do not appear elsewhere. They are also designed to show the vendor's product in a favourable light. Responsible analysis does not discard the source, but it keeps the commercial context visible.

In this case, the reported outcome belongs to A2B's described deployment. The case attributes the operating priority and decision to Bais, and it presents the company result through Juniper's publication. The article can state those facts clearly: the team tested the system, put it into production and reported full-table convergence in three to four seconds. It should not upgrade the report into a neutral laboratory comparison or claim that the figure applies to unrelated networks. [4]

The report also does not support an invention claim. Bais did not invent BGP, virtual routing, IPv6, multihoming or automation in this evidence. His documented contribution was an operator choice about how to apply available technology in a production environment. That is a substantial form of infrastructure work. Operators create value not only by inventing protocols but also by deciding which systems to trust, how to test them and when evidence is strong enough for production use.

IPv6, multihoming and automation belong to the same continuity design

The customer case says the deployment included IPv6 validation, supported multihoming and created a foundation for automation. Each term points to a different continuity concern. IPv6 is the newer Internet Protocol address system designed to provide a vastly larger address space than IPv4. Validation matters because a network can carry IPv6 in theory while still failing in routing policy, filtering, monitoring or customer delivery.

Multihoming means connecting a network through more than one external path or provider. It can improve continuity by giving traffic another route when one connection is unavailable. The benefit is not automatic. The operator must announce and accept routes correctly, maintain coherent policy and ensure that alternative paths have enough capacity. Faster BGP convergence becomes more valuable when the network has a viable alternative to select.

Automation can help operators apply repeatable configuration and respond to change at scale. In routing, manual changes across many devices create opportunities for inconsistency. A system that supports programmatic control can make validation and deployment more repeatable, but automation also amplifies mistakes when controls are weak. The public case says the new environment provided a foundation for automation; it does not document every automated workflow or prove a later efficiency result. [4]

These elements fit together as design conditions rather than as a list of product features. IPv6 expands the protocol environment that must work. Multihoming supplies path diversity. Convergence determines how quickly the routing system can use changed information. Automation can improve repeatability. Operational continuity emerges only when the pieces are tested together and when people understand the failure modes between them.

The 2018 routing-security intervention adds a separate kind of discipline

Internet Society's account of RIPE 76 provides an independent, person-level source for a second episode. Published on 17 May 2018, it says Erik Bais of A2B Internet opened a routing-security discussion with a presentation about persistent distributed denial-of-service traffic. A distributed denial-of-service attack, or DDoS attack, overwhelms a target or supporting network with traffic from many systems, making legitimate access difficult or impossible. [3]

The report says the presentation highlighted networks that repeatedly appeared as traffic origins and called on operators to clean up their networks. It also connected the discussion to the Mutually Agreed Norms for Routing Security, known as MANRS. MANRS is a set of operator-oriented practices intended to reduce common routing and traffic-abuse problems through actions such as coordination, filtering and accurate information. The cited report confirms the presentation and its operator-responsibility message; it does not prove that the message caused measurable reductions in attacks. [3]

This independent account is valuable because it is not a customer case written by the routing-platform vendor. It shows Bais participating in an operator-community discussion and directing attention toward repeated network behaviour. The contribution is not a protocol invention or a claim of sole responsibility for routing security. It is a public technical intervention: identify a recurring operational problem, show where patterns appear and urge networks to act on the evidence available to them.

Routing security and DDoS response overlap without being identical. BGP determines reachability between networks. DDoS traffic can travel over correctly announced routes and still cause harm. Some routing failures involve hijacks, leaks or incorrect origin information; some attacks involve compromised systems sending unwanted traffic along otherwise valid paths. Operators need to distinguish these mechanisms. Network hygiene includes maintaining accurate routing, controlling spoofed or abusive traffic where possible, keeping contacts usable and responding when evidence points to persistent problems.

The intervention therefore broadens the continuity discussion. Recovery speed asks how quickly a network adapts when paths change. Routing hygiene asks whether operators reduce avoidable instability and abuse in the first place, and whether they cooperate when problems cross organisational boundaries. A network can converge quickly while accepting poor routes or ignoring abusive traffic. It can follow security guidance while still recovering slowly from a failure. Strong operations require both kinds of discipline.

The year between the cited production case and the RIPE meeting should not be presented as a causal sequence. The public sources do not say that the virtual routing deployment led to the DDoS presentation or that the presentation reflected results from that deployment. The connection belongs to analysis: both episodes reveal an operator's attention to observable behaviour and operational responsibility. Keeping that distinction explicit preserves the value of each source.

Network-hygiene evidence can inform peering without becoming a blacklist

An AMS-IX article provides additional industry context about A2B's approach to network hygiene. It describes the company analysing aggregated data about network misconfiguration and using a size-adjusted rating in peering decisions and traffic handling. Peering is the direct exchange of traffic between networks. It can improve performance, cost or control, but it also creates dependencies: each participant relies on the other to operate responsibly. [2]

The size adjustment is conceptually important. A large network may generate more observed incidents simply because it operates more systems and carries more traffic. Raw counts can therefore mislead. A rate or score that accounts for network scale may help distinguish widespread exposure from unusually poor hygiene. The public article supports the existence of the described method; it does not establish that the score is complete, unbiased or predictive of every future incident. [2]

Using such evidence in peering decisions can create an incentive for improvement. If repeated misconfiguration or abuse affects how another network treats traffic, an operator has a reason to investigate. The evidence can also help teams decide where closer monitoring or direct coordination is needed. Yet a score should not become an unexplained blacklist. Networks need to know which observations matter, how errors are corrected and how temporary incidents differ from persistent neglect.

The same caution applies to automated traffic handling. Automation can make a response faster and more consistent, but it can also entrench bad data. A false association, stale address mapping or poorly calibrated threshold could penalise legitimate traffic. Strong governance requires a path to review decisions, update evidence and reverse an action when the underlying facts change. The AMS-IX account is useful as an example of evidence entering interconnection practice, not as proof that every resulting decision was correct.

The approach also shows why contact data and operating evidence serve different functions. A registry or directory can tell an operator whom to reach. Network observations can show what needs discussion. Peering policy can define what action follows. Combining the three is useful; confusing them is dangerous. A registry contact is not proof of misconduct, and an observed incident does not establish permanent responsibility without careful attribution.

Recovery speed and routing hygiene are related, but not one causal story

It is tempting to tell a simple story: an operator made routing faster, then moved on to make routing safer. The sources do not justify that chronology as a cause-and-effect narrative. The 2017 account is a vendor-published deployment case. The 2018 account is independent reporting on a public presentation. The AMS-IX material describes a network-hygiene method. They concern related operating surfaces, but none says that one project produced the next.

The more defensible connection is a shared decision habit. In each episode, the operator's work is framed around observable behaviour. The deployment case begins with slow convergence and records a faster result after testing and production use. The RIPE discussion begins with recurring traffic sources and asks origin networks to act. The peering account uses aggregated observations to inform how networks interact. The common thread is not a single technology; it is the attempt to turn evidence into an operating decision.

That habit supports continuity because networks are dynamic. Routes change, links fail, software evolves, customers grow and abusive traffic shifts. Static declarations are not enough. An operator needs measurements that expose what the running system does, responsibility boundaries that identify who can act and decision processes that can change when evidence changes.

For infrastructure leadership, the distinction between correlation and causation is not academic. Investment decisions can go wrong when a visible improvement is credited to the wrong change. A new routing platform may coincide with better recovery while path diversity, configuration or operations also changed. A fall in abusive traffic may coincide with a policy intervention while external conditions shifted. Evidence should guide action, but the organisation must know which claims the evidence can carry.

What operators should ask before trusting a recovery claim

An operator evaluating a routing change can begin with the failure model. Which link, process or neighbour is expected to fail? Which alternate path should take over? What routing information must be processed, and what policy determines the replacement? A recovery number has little meaning without that context. Three seconds for one controlled withdrawal is not the same as three seconds during a wider control-plane event.

The next question is measurement. Teams should separate router convergence from customer-visible recovery. They can record when the failure begins, when the router selects a usable path, when forwarding resumes and when representative applications recover. Those stages may occur at different times. The public A2B account speaks to full-table convergence and faster handling of BGP flaps, not to every end-user application. [4]

Repeatability matters as well. A single successful test may reveal capability, but repeated tests across expected conditions reveal variation. Operators should include IPv4 and IPv6 where both are in scope, verify that multihomed paths carry the intended traffic and confirm that automation does not produce inconsistent policy. These are prudent operating questions derived from the case, not claims that the public source documents every such test.

Change control should preserve the human decision. Automation can execute configuration, but somebody must approve the target state, define success and decide what to do when evidence conflicts. The 2017 account is valuable because it connects lab work with a production decision. It reminds readers that infrastructure changes are not merely software events; they are accountable choices about acceptable risk.

Operators should also ask what happens after the platform performs well. Who monitors software updates, routing-table growth, compute capacity and new protocol demands? What evidence would trigger a retest? How are incidents classified? A result measured in 2017 cannot be treated as a permanent guarantee. Continuity requires a lifecycle, not a one-time acceptance ceremony.

Finally, teams should keep supplier evidence in proportion. A vendor case can support a shortlist or a technical hypothesis. It should be combined with local tests, contractual commitments, incident history and independent technical sources. The aim is not to distrust every vendor statement. It is to ensure that a claim used for a production decision can be reproduced or bounded in the buyer's own environment.

What customers and procurement teams can learn from the case

Customers often buy an outcome such as reliable connectivity rather than a routing platform. They may never specify BGP convergence in a contract. Still, the provider's ability to recover from path changes can affect the service they receive. Procurement teams can ask for evidence that connects underlying operations to customer experience without dictating every engineering choice.

Useful questions include whether the provider has diverse external paths, how it tests failover, how quickly it communicates incidents and whether IPv6 receives the same operational attention as IPv4. Buyers can ask for representative recovery evidence while recognising that network architecture and confidentiality may limit raw detail. A credible provider should be able to explain the method and boundary of its claims.

Routing hygiene belongs in the same conversation. A provider that exchanges traffic with many networks faces risks from misconfiguration, abuse and weak coordination. Buyers can ask how the provider maintains contact channels, evaluates recurring problems and avoids treating weak evidence as a permanent judgment. The goal is not to demand a universal blacklist. It is to understand whether the provider has a repeatable, reviewable process.

For boards and investors, the case offers a way to evaluate operational maturity. Look for evidence that management turns technical goals into tested decisions, that results are reported with boundaries and that the organisation participates in collective security practices. Avoid treating a single fast metric or conference appearance as proof of durable advantage. Maturity appears in the system that produces, checks and updates evidence.

Limits, uncertainty and the evidence worth watching next

The public evidence does not provide an independent replication of the three-to-four-second convergence result. It does not supply a full before-and-after configuration, a complete test methodology, application-level outage measurements or a long-term reliability series. Readers should not infer that every customer experienced the same recovery time or that the result remains unchanged today.

The sources also do not establish that Bais personally configured each system, built the automation or operated every incident response. They support a named operating priority, the lab-to-production decision and a later public routing-security intervention. The implementation outcomes belong to A2B and its team as reported in the customer case.

The 2018 presentation does not prove that originating networks changed their behaviour, that DDoS traffic fell or that MANRS adoption increased because of Bais's remarks. The AMS-IX account does not provide a complete validation of the hygiene rating or every traffic decision based on it. Those questions require later measurements, transparent methodology and evidence from affected parties.

Current organisational status is outside the article's claims. The analysis uses dated material and attributed page descriptions rather than assuming a present role. Future reporting could revisit the operating model if current, attributable sources document a new deployment, measured outcome or policy decision. It should not convert an unchanged biography or contact listing into a new contribution.

The most valuable next evidence would be specific and operational: repeated convergence distributions under defined failures; customer-visible recovery measurements; current IPv6 and multihoming validation; documented review of automated policy; transparent hygiene-score methodology; correction and appeal processes; and measured changes following operator outreach. Those records would allow readers to separate durable practice from a one-time case.

Image disclosure

Alt: AI-generated photorealistic editorial scene: an anonymous, fully concealed network operator viewed from behind.

Caption: Not a photograph or likeness of Erik Bais.

Sources

  1. A2B Internet, “About us,” official identity and operator context: https://www.a2b-internet.com/about-us/
  2. AMS-IX, “Predicting and mitigating DDoS attacks,” industry account of A2B's network-hygiene and peering method: https://www.ams-ix.net/ams/news/predicting-and-mitigating-ddos-attacks
  3. Internet Society, “RIPE 76 sees strong focus on routing security,” 17 May 2018: https://www.internetsociety.org/blog/2018/05/ripe-76-sees-strong-focus-on-routing-security/
  4. Juniper Networks, A2B customer case study, March 2017: https://www.juniper.net/us/en/customers/a2b-case-study.html