Executive summary
- ESnet is a US Department of Energy scientific user facility operated by Lawrence Berkeley National Laboratory, not a commercial carrier or standalone corporation.
- ESnet6 combines about 15,000 miles of dedicated fibre, optical and packet systems, private services, interconnection, measurement and programmable control; its 57 Tbps capacity figure is an aggregate, not a user speed.
- Science DMZ, perfSONAR, OSCARS, SENSE, High-Touch and EJFAT show that ESnet’s distinctive contribution is end-to-end engineering, from the instrument and campus edge to remote computing.
- ESnet7 and American Science Cloud planning point towards a network embedded more deeply in scientific workflows, but several components remain demonstrations, research systems or emerging services.
A scientific network enters the experiment itself
In April 2024, data from Jefferson Lab in Virginia crossed the United States to the Perlmutter supercomputer in California at 100 Gbps and was processed without first being written to disk. The demonstration used ESnet’s experimental EJFAT system to place the wide-area network inside the scientific workflow rather than after it. That shift captures the central question facing ESnet: can a network connect instruments, storage and remote computing closely enough that facilities separated by thousands of miles behave like parts of one machine?
A typical path begins with a detector, telescope, microscope or simulation at a laboratory or research facility. Local acquisition systems collect the data, storage and transfer servers prepare it, and a campus network carries it to a controlled perimeter. ESnet may then move the traffic across the United States, over the Atlantic, into a Department of Energy supercomputer, through a research-and-education partner or into a commercial cloud. The result can return as an analysed dataset, an alert, a model or a decision about what the experiment should do next.
No single organisation controls that complete path. Instrument teams, laboratory networks, regional research networks, cloud providers, international partners and supercomputing facilities can all answer to different institutions. ESnet operates the mission-oriented wide-area layer and works with those operators to make the path function as a system. A slow storage server, congested campus link or unsuitable firewall can waste the capacity of a fast national backbone; an ESnet fault can interrupt a well-designed local facility.
Institutionally, ESnet is a Department of Energy Office of Science user facility. Advanced Scientific Computing Research provides its principal stewardship, and Lawrence Berkeley National Laboratory’s Scientific Networking Division operates it. Berkeley Lab is itself managed for DOE by the University of California under contract. ESnet has no identified shareholders, stock, corporate valuation or standalone commercial balance sheet, so it is better understood as federal scientific infrastructure than as a network company with a public customer base.
Researchers usually encounter ESnet indirectly through an instrument, supercomputer, laboratory service, university connection or collaborating network. The facility distinguishes institutional site users from endpoint users and says it does not formally enrol or track each individual endpoint user. The 2024 Annual Report’s figure of more than 30,000 organisational users is therefore an institutional measure, not a subscriber count.
Ordinary carrier comparisons capture only part of the role. ESnet runs routing, optical transport and interconnection, but it also conducts requirements reviews, develops open-source tools, tests new services and helps sites diagnose end-to-end performance. Its public purpose is not to sell a connection. It is to make geographically distributed scientific facilities work together with enough speed, evidence and operational clarity to support the research itself. (ESnet governance; network services)
Remote computing created the need for a shared science network
ESnet’s official history reaches back to the mid-1970s, when computing and long-distance communications were both scarce. At Lawrence Livermore National Laboratory’s Controlled Thermonuclear Research Computer Center, staff connected a borrowed Control Data Corporation 6600 through four acoustic modems. The equipment and speeds belong to another era, but the operating problem is recognisable: a specialised scientific community needed remote access to expensive computing that could not be reproduced at every participating institution.
Separate Department of Energy research communities developed their own networks during the late 1970s and early 1980s. High-energy physics and magnetic-fusion research had different facilities, collaborators and data flows, and predecessor systems such as HEPnet and MFEnet reflected those boundaries. Purpose-built networking was rational when commercial services could not provide the required reach, performance or operational attention. It also created duplication. Each programme could negotiate circuits, maintain technology and build expertise independently, leaving collaboration and long-term upgrading fragmented.
ESnet’s formal formation in 1986 consolidated those functions into a broader science network. The change was organisational as much as technical. A shared facility could forecast demand across programmes, operate national links, coordinate international connections and retain specialist staff. It could also move capacity from a one-project concern into a Department-wide infrastructure plan.
The early model established several characteristics that remain visible. Users were defined by scientific mission rather than by a retail market. Central computers and instruments justified shared public investment. Network design followed research programmes with long timelines. Operations required people who understood both communications systems and the scientific consequences of failure. Demand arrived unevenly because one experiment could generate traffic far beyond the ordinary baseline.
Berkeley Lab had its own networking lineage. It connected a CDC 6600 to ARPANET in 1974, and researchers there later contributed foundational TCP congestion-control work. Van Jacobson and Mike Karels’ work belongs to the wider Berkeley Lab environment and should not be credited to ESnet alone. The institutional overlap nevertheless matters. It placed production scientific networking near researchers who treated protocol behaviour, measurement and performance as engineering problems that could be studied rather than accepted as fixed conditions.
Operations moved to Berkeley Lab in 1996. That move established the current home beside the National Energy Research Scientific Computing Center, networking researchers, software engineers and other Department of Energy programmes. It also reinforced a culture in which a production facility and an applied-research organisation share staff, laboratories and technical questions. (ESnet history)
Public funding lets ESnet build before demand arrives
A scientific user facility exists because some capabilities are too expensive, specialised or interconnected for each research team to build alone. A particle accelerator, light source or leadership supercomputer fits that logic. ESnet applies it to communications. Long-haul fibre rights, optical equipment, routers, international capacity, continuous operations, cybersecurity, performance measurement and engineering support are pooled into one facility serving many programmes.
The public-funding rationale follows from the shape of scientific demand. A commercial carrier can sell a high-capacity circuit, but it does not normally decide its national architecture around a detector scheduled to begin operation years later, a supernova burst that may never occur during the contract, or a research workflow that has little retail volume but high public value. ESnet can build ahead of measured utilisation because its mandate is scientific capability rather than near-term network revenue.
The facility model also changes how demand is discovered. ESnet conducts formal requirements reviews with Department of Energy science programmes. Researchers and facility teams describe instruments, data volumes, storage locations, computing destinations, timing needs, collaboration patterns and local bottlenecks over a five-to-ten-year horizon. Those findings become inputs to site-link size, route diversity, transatlantic procurement, orchestration research and service design.
This process is less glamorous than a new coherent optical link, but it may be more consequential. A submarine spectrum agreement, a router purchase or a second site entrance can take years to fund and deploy. Waiting until an experiment produces data would turn predictable infrastructure needs into emergencies. Requirements reviews convert scientific plans into engineering lead time.
They also expose uncertainty. Forecasts can change. Instrument schedules slip, data-reduction software improves, cloud use expands, and a facility may redesign its workflow. ESnet does not treat a case study as a guaranteed traffic order. It uses the evidence to construct a range of requirements and then decides where flexibility, reserve capacity or staged deployment is justified.
The user-facility model avoids per-gigabyte billing for individual scientists. Core operations are federally funded, while some site-level connection costs can be allocated according to programme sponsorship and the Site User Cost Policy. That arrangement is not free of incentives. A sponsoring programme may prefer a lower connection cost; ESnet may prefer capacity and resilience that serve several programmes; a site may delay local upgrades that are necessary to use the national network. The governance task is to align those decisions around scientific outcomes rather than around who can shift a cost to another budget line. (ESnet Site Coordinators Committee and cost policy; requirements-review reports)
DOE sets the mission; Berkeley Lab runs the network
ESnet does not have the corporate hierarchy of a carrier. Its authority is distributed through a federal programme chain. The Department of Energy Office of Science defines the mission. Advanced Scientific Computing Research supplies principal stewardship and budget oversight. Berkeley Lab houses the Scientific Networking Division that designs, operates and develops the facility. The University of California operates Berkeley Lab for the Department under contract. Connected institutions participate through site coordinators and the ESnet Site Coordinators Committee.
Inder Monga serves as ESnet Executive Director and directs Berkeley Lab’s Scientific Networking Division. Current public leadership also identifies Chin Guok as Chief Technology Officer and head of Planning and Innovation, Adam Slagell as Chief Security Officer, Jon-Paul Herron as Director of Network Engineering, and Susan Lucas as Deputy of Business Operations. Beneath those titles are functions covering optical engineering, routing, site engineering, network operations, software, measurement, orchestration, security, project management and science engagement.
The breadth of that structure corrects a common picture of a national network as a collection of fibre and routers. ESnet needs people who can procure spectrum, operate BGP, build telemetry systems, maintain open-source software, investigate packet behaviour, plan facilities, manage security evidence and translate scientific requirements into network designs. The organisation is both operational and developmental.
Public records contain one unresolved programme-management detail. ESnet’s current governance page names Benjamin Brown as its designated ASCR programme manager, while recent requirements-review material identifies Carol Hawk as ESnet programme manager. The evidence supports saying that public sources have not been reconciled or that responsibilities may have changed or been divided. It does not support selecting one exclusive current arrangement without further confirmation.
The ESnet Site Coordinators Committee gives connected institutions a formal operating channel. Each site designates a coordinator who can approve requests affecting its ESnet connection, communicate requirements and participate in policy or planning discussions. This is institutional user governance rather than direct voting by every scientist whose data crosses the network.
That structure creates a practical allocation of responsibility. ESnet controls its backbone and services. A connected site controls its campus equipment, data-transfer nodes, local security and power. ASCR controls programme-level funding. Scientific programmes define the consequences of delay or failure. International partners control their own networks. The complete workflow succeeds only when those separate decision surfaces remain aligned. (ESnet leadership; governance)
ESnet4 separated routine traffic from exceptional science flows
By the mid-2000s, distributed experiments were changing the traffic profile. The Large Hadron Collider would produce data at CERN and distribute it through a global hierarchy of laboratories and computing centres. General internet connectivity remained necessary, but a small number of scientific transfers could be large enough to dominate ordinary links. Building one undifferentiated network for both patterns made capacity management and service guarantees difficult.
ESnet4 addressed that problem with a hybrid architecture. An IP core carried general scientific communication. A separate Science Data Network used high-capacity optical circuits for large flows and supported dynamic path provisioning. Metropolitan rings connected major laboratories, while collaboration with Internet2 and international partners extended the system beyond Department of Energy sites.
The architecture separated traffic by operating requirement rather than by a simple “slow” and “fast” divide. General IP service needs broad reachability and resilient routing. A planned multi-terabyte transfer may benefit from a reserved Layer 2 path with a known start time, bandwidth allocation and endpoints. The science data network recognised that a small number of predictable, high-value flows can justify a different control model from ordinary packet forwarding.
OSCARS emerged from this environment. The On-demand Secure Circuits and Advance Reservation System allowed authorised users or applications to request network resources for a defined period. The system had to find a path, check topology and policy constraints, reserve bandwidth and VLAN resources, create device state, and remove that state when the reservation ended. It turned a task previously performed through manual coordination into a service that software could request.
ESnet4 also showed that raw speed was no longer the only bottleneck. A reserved circuit can protect capacity on the wide-area network, yet the application may still run slowly because a storage system cannot read fast enough or a local firewall drops packets. That lesson led directly towards Science DMZ architecture and end-to-end measurement.
The period established a recurring ESnet design pattern. A scientific problem appears first as a traffic requirement. The facility then creates physical capacity, a control mechanism and an operational method around it. When the method proves useful, software and architectural guidance spread beyond the original experiment. (ESnet4 and ESnet5 historical material; OSCARS)
Science DMZ showed that the bottleneck often sits at the edge
A national backbone can operate correctly while a scientist experiences poor transfer performance. The reason is straightforward: the backbone is only one segment of the path. Enterprise firewalls, shared campus cores, old routers, underpowered servers, packet loss and storage limitations can reduce a high-capacity route to a small fraction of its engineered rate.
Science DMZ is ESnet’s response to that end-to-end reality. The architecture places data-transfer nodes on a high-performance path near a controlled institutional perimeter. Security policy is tailored to the limited services exposed by those systems rather than forcing sustained scientific traffic through a general-purpose stateful firewall designed for many kinds of enterprise application.
The phrase can be misunderstood as a request to remove security. A Science DMZ still depends on hardened hosts, router access-control lists, monitoring, vulnerability management, limited application exposure and operational discipline. The design changes where and how controls are applied. It avoids putting a device with unsuitable throughput or session behaviour in the main data path simply because that appliance is part of the institution’s ordinary security pattern.
Performance measurement is part of the architecture rather than an optional dashboard. perfSONAR nodes run controlled tests for throughput, packet loss, latency and path changes. When two sites disagree about a transfer, measurements can help locate whether the problem begins at a host, a local link, a regional network, the ESnet backbone or an international partner. Without such evidence, every operator can report that its own segment looks healthy while the researcher still cannot move data.
Data-transfer nodes need their own engineering. Network interface cards, CPU placement, memory, disk or parallel file systems, TCP buffers and transfer software all determine achievable throughput. ESnet’s Fasterdata guidance and technical consulting turn those details into repeatable practice. The approach has spread well beyond Department of Energy laboratories, but ESnet does not operate every deployment that uses the name.
Science DMZ became influential because it reframed a procurement question as a system question. Buying a faster circuit does not solve a path that is constrained elsewhere. The useful unit of performance is the complete workflow from storage to storage or instrument to compute, with each administrative domain measured and assigned a clear responsibility. (Science DMZ guidance; network performance tools)
ESnet5 made continental 100G a production service
The next major capacity transition was supported by the Advanced Networking Initiative and US economic-stimulus funding. ESnet received $62 million through the American Recovery and Reinvestment Act to develop a 100 Gbps long-haul prototype and support the move towards ESnet5. The investment included more than new routers. ESnet obtained rights to spectral capacity across a national fibre footprint, giving the programme greater control over how wavelengths could be deployed and upgraded.
ESnet5 entered production in late 2012. At the time, DOE and ESnet described it as the world’s fastest science network. That historical claim was tied to a first-of-kind continental 100G deployment. It should not be repeated as an unqualified 2026 global ranking. Research networks now report capacity through different measures, including individual interface speeds, aggregate backbone capacity, optical spectrum and experimental trials. There is no neutral current comparison in the supplied evidence that makes one universal winner.
The economic effect of 100G was larger than a tenfold interface upgrade. Scientific projects could plan around routine long-distance transfers that would previously have occupied links for much longer. Laboratories could centralise some computing or storage rather than reproduce it locally. International collaborations gained a more capable US backbone. The new capacity also exposed local limitations more quickly, increasing the value of Science DMZ and perfSONAR.
ESnet5 retained the hybrid relationship between packet service and dedicated scientific paths. It also deepened the importance of automation. At 100G scale, reserving capacity, diagnosing faults and maintaining consistent service across a national footprint could not depend on ad hoc device configuration alone.
The project demonstrates how public infrastructure investment changes scientific options before any single experiment uses the full network. Capacity creates a reserve against future instruments, but reserve capacity is not waste in the same sense as an idle retail product. Its value includes avoiding delayed experiments, accommodating failures and allowing new workflows to be tested without waiting for another construction cycle.
The investment still has to be judged against scientific use, and the design must remain adaptable when forecasts change. ESnet’s later work on telemetry, caching and orchestration reflects an effort to extract more scientific value from each installed bit rather than treating continuous physical expansion as the only answer. (ESnet5 historical announcement)
ESnet6 combines optical capacity with programmable control
The ESnet6 project began in 2017 and reached public launch on 11 October 2022 after roughly six years of design and construction. Berkeley Lab reported more than 46 Tbps of aggregate capacity at launch across a dedicated footprint of about 15,000 miles, with backbone links from 400 Gbps to 1 Tbps. The 2024 Annual Report later gave 57 Tbps and a link range extending to 1.2 Tbps. These figures refer to different dates and a growing system rather than to conflicting definitions of one frozen network.
The physical layer includes long-haul fibre paths, optical line systems, amplifiers, spectrum rights, colocation facilities and local on-ramps. The 2024 report counted 278 optical-amplifier sites. Coherent optics carry several high-rate channels over fibre, while amplifiers and reconfigurable optical equipment maintain signal levels and route wavelengths. A path shown on a map may be controlled through dark fibre, spectrum, lit service or a partner arrangement. The 15,000-mile description does not establish that DOE holds legal title to every cable.
Above the optical layer, routers and service edges provide IP, private routing environments, Layer 2 services, peering and cloud connectivity. The annual report counted 78 router locations. Individual sites can connect at 10, 100 or 400 Gbps according to requirements, while backbone links can aggregate several interfaces or wavelengths into higher figures.
The software layer distinguishes ESnet6 from a capacity-only upgrade. Automation maintains device and service state. Programmable telemetry gives operators more detailed evidence. OSCARS provisions reserved paths. SENSE experiments with application-driven coordination across domains. High-Touch uses programmable hardware for richer packet and flow visibility. APIs and testbed hooks allow successful research to move towards operations.
The build also placed greater emphasis on resilience. Diverse routes, multiple site entrances, independent power, alternative routers and high-availability service locations reduce the chance that one failure removes a laboratory from the national system. These protections have different scopes. A redundant backbone does not help if both local circuits share one conduit, and a second router does not create resilience if both devices depend on the same power system.
ESnet6 should therefore be read as a layered facility. Fibre creates the physical reach. Optics create capacity. Routers create packet and private services. Software creates repeatability and programmability. Measurement establishes whether the promised path works. Organisational agreements make it possible to cross boundaries that ESnet does not own. (ESnet6 launch; 2024 Annual Report)
What 57 Tbps means—and what it does not mean
Network statistics often collapse different layers into one number. ESnet’s 57 Tbps figure for 2024 is aggregate engineered capacity across the system. It is not one physical link, the throughput available to every laboratory, or the traffic carried at every moment. The same report describes backbone links between 400 Gbps and 1.2 Tbps, seven US sites connected at 400 Gbps or above, and 2.7 Tbps of transatlantic capacity. Each figure answers a different question.
The same report described 78 router locations and 278 optical-amplifier sites across the approximately 15,000-mile footprint. It connected all 17 DOE national laboratories and reported 28 DOE user facilities, 277 research-and-education, commercial and other network relationships in five countries, more than 30,000 organisational users, and 137 staff members and contractors working across 24 states. These figures describe different populations and should not be added together as a customer total: routers, facilities, partner networks, institutional users and workforce are separate measures.
A site connection is the capacity at one institutional boundary. A backbone link is capacity between network locations. An optical system can carry several wavelengths, and a logical service may use only part of the physical layer. Aggregate capacity adds many resources that cannot all serve the same source and destination simultaneously. Traffic volume measures data moved over time rather than the maximum rate available at an instant.
ESnet carried 1.77 exabytes during calendar year 2024, up from 1.7 exabytes in 2023. The four per cent increase was low compared with the facility’s reported long-run average of roughly 55 per cent annual growth since 1989. That single year does not prove that scientific demand has stabilised. Large projects commission unevenly, software changes can reduce transfers, and one experiment or international event can alter the mix rapidly.
Traffic is also concentrated. The 2024 report said Large Hadron Collider activity accounted for roughly half of total traffic. Fermilab was the largest external sender at 136 petabytes, NERSC the largest Department of Energy recipient at 75.3 petabytes, and Oak Ridge the largest Department sender at 58.3 petabytes. Such figures reveal important workloads, but they are not a complete ranking of scientific value. A smaller transfer may be more time-sensitive or may support a unique instrument.
Capacity planning therefore needs several views at once: peak and average traffic, route-level utilisation, failure reserve, site link rates, scientific deadlines and future requirements. A network designed only around average use can fail during a supernova burst or a cable outage. A network designed only around the largest theoretical event can overbuild expensive assets that remain inflexible.
ESnet’s engineering problem is to place capacity where it can be reached by the workflows that need it, then verify that the end systems can use it. That is why the capacity number belongs beside Science DMZ, measurement, orchestration and requirements reviews rather than at the top of the article as a self-explanatory ranking. (ESnet by the Numbers)
AS293 routes for a mission rather than for a retail market
ESnet’s public autonomous-system number is AS293. The network exchanges routes with Department of Energy sites, research-and-education networks, commercial networks and paid upstream providers. Routing policy determines which paths can carry scientific traffic, how failures are handled and which external networks receive direct interconnection.
The peering policy is selective. Prospective peers must use BGP, maintain current Internet Routing Registry and PeeringDB information, avoid RPKI-invalid announcements and follow routing-security practices associated with MANRS. ESnet requires IPv6 or dual-stack peering and does not accept new IPv4-only relationships. Direct private interconnection starts at a 100G port, with 400G strongly preferred.
Those conditions reflect both mission and operating cost. A direct interconnection consumes ports, engineering time and monitoring capacity. ESnet does not need to peer with every network simply because public routes exist. It seeks relationships that improve Department of Energy science paths, reduce dependence on transit or connect major collaborators. Public-internet access is supplemented by paid transit from Hurricane Electric and Lumen for destinations that are not reached efficiently through research or direct peering.
The network publishes BGP communities that distinguish ESnet-connected sites, research-and-education networks and commercial networks. Connected institutions can use those tags to apply policy, although downstream interpretation remains outside ESnet’s control. A community is descriptive metadata, not a universal guarantee that every network will make the same routing choice.
Routing security also has limits. Rejecting RPKI-invalid announcements blocks one important class of bad route, but it does not prevent every leak, compromised peer, mistaken valid origin or incorrect route-origin authorisation. Route-hijack monitoring can identify unexpected origin changes, yet detection may occur after propagation. Blackhole routing can protect a site during attack by discarding traffic to a destination, but the protection works by sacrificing reachability for that prefix.
The mission-oriented design is visible in the service boundary. ESnet carries public IP traffic, but it is not attempting to win a consumer transit market. Its route policy is arranged around scientific reachability, resilience and trusted interconnection. The network’s value comes from the quality of the paths it creates for the community it serves, not from maximising the number of retail customers attached to AS293. (ESnet peering policy)
Private paths extend reach without eliminating external risk
ESnet’s service menu begins with physical connectivity. Eligible sites can connect locally or through a remote point of presence at 10, 100 or 400 Gbps. When a facility is distant from the dedicated footprint, a Service On-Ramp may use dark fibre or a lit carrier circuit. The access design is part of the scientific path, and a vendor-provided last segment can become the dominant failure or performance constraint.
Layer 3 IP service provides routed IPv4 and IPv6 connectivity. A Layer 3 virtual private network creates a logically separate IP and BGP environment for a multi-site programme while sharing physical infrastructure with other services. Layer 2 virtual private networks provide point-to-point or multipoint Ethernet connectivity, either configured statically or provisioned through OSCARS. These services let experiments connect facilities without exposing every relationship through the public internet.
Private service does not mean physically isolated service. Logical separation depends on router configuration, labels, routing instances, access controls and operational practice. A fault in the shared platform can affect several services, while a configuration error can defeat intended isolation. The value lies in controlled segmentation and predictable connectivity, not in a claim that no shared dependency exists.
Cloud Connect extends ESnet paths to major commercial cloud environments, including AWS, Microsoft Azure, Google Cloud and Oracle. A site can reach a provider on-ramp through dedicated Layer 2 or Layer 3 connectivity rather than ordinary public-internet routing. This can improve path control and reduce exposure to variable transit, but it does not remove cloud-side security, provider outages, identity design, egress charges or regional availability.
The cloud relationship illustrates ESnet’s changing role. Department of Energy science no longer lives only in laboratories and federal supercomputers. Researchers may use commercial object storage, specialised services or burst computing. ESnet must connect those resources without pretending that the cloud is part of the federal network or that a private circuit makes a cloud account safe by itself.
Other services include secondary DNS, network time, route-hijack monitoring and blackhole routing. Each addresses a supporting dependency that can interrupt research even when optical capacity remains available. A name-service failure, incorrect time, route leak or denial-of-service attack can make a scientific system unusable without damaging the fibre.
The service menu amounts to a set of operating contracts between network layers. Physical connectivity establishes the path. Routing creates reachability. Private services create logical boundaries. Cloud interconnection extends the system to commercial infrastructure. Security and supporting services reduce known failure modes. The user still has to build a functioning endpoint and local network at each end. (ESnet network services menu)
OSCARS makes capacity reservable
Most internet traffic uses whatever capacity is available when packets arrive. That model works well for general communication, but some scientific transfers are scheduled, large and expensive to delay. OSCARS gives authorised users and applications a way to reserve network resources for a specified period.
A reservation can identify endpoints, start and end times, bandwidth, VLAN information, route exclusions and other constraints. The system examines topology and existing commitments, finds an acceptable path, creates the required network state and removes it when the reservation expires. The service must avoid allocating the same scarce resource twice and must account for maintenance or failures that change the available topology.
The significance is operational rather than cosmetic. Manual circuit provisioning can take days of coordination among engineers. A programmatic reservation can be created in minutes and can become part of an experiment’s workflow. The network ceases to be a fixed background pipe and becomes a resource that software can request alongside storage or compute.
OSCARS is an open-source production system and has been adopted or evaluated by other networks and testbeds. Some public adoption figures appear on pages containing historical material, so the safest conclusion is that the system has travelled beyond ESnet without treating every listed deployment as current and identical.
A reservation does not guarantee application throughput. It can protect a defined amount of network capacity, but it cannot make a storage server read faster, repair packet loss in a campus network or tune an operating system. If the application achieves only a fraction of the reserved rate, the diagnosis must return to the end-to-end path.
Reservations also introduce allocation questions. A high-priority experiment can benefit from guaranteed capacity, while unused reservations can reduce flexibility for others. Policies must decide who can request service, how conflicts are resolved and whether a reservation should be pre-emptible. Those decisions are part of facility governance even when the provisioning code is automatic.
OSCARS remains one of ESnet’s clearest examples of applied software becoming infrastructure. It translated an operating practice into a repeatable service, preserved human policy around authorisation, and gave scientific applications a controlled way to express network requirements. (OSCARS)
SENSE tests whether applications can request an end-to-end outcome
A reserved network path solves only one part of a distributed workflow. Storage systems, data-transfer services and compute resources also have availability, capacity and policy. SENSE—the Software-defined network for End-to-end Networked Science at the Exascale—extends orchestration beyond a single network domain.
The project represents networks and end systems as discoverable resources. A scientific application can describe an intended result, such as moving a dataset between facilities at a required rate. SENSE gathers models from participating domains, negotiates resources, provisions network and transfer services, observes telemetry and releases resources after the workflow completes.
This is harder than configuring one circuit. Participating organisations retain their own authority and data models. One domain may expose bandwidth while another exposes storage endpoints or transfer services. Policies, credentials and maintenance windows differ. SENSE must coordinate without assuming that one central controller can command every system.
During the 2024 Large Hadron Collider Data Challenge, SENSE-enabled workflows sustained 330 Gbps between Caltech and the University of California, San Diego. The result shows that application-driven orchestration can operate at substantial rate in a specific live workflow. It does not show that SENSE is a universal production service across every ESnet site. The 2024 reporting still described the project as being in a research phase.
OSCARS and SENSE automate different units of work. OSCARS reserves network resources inside a known service domain. SENSE attempts to negotiate a complete path that includes systems owned by different organisations. The first has a clearer production boundary. The second offers broader scientific value but carries more governance and integration risk.
For operators, the practical issue is failure ownership. If an application declares intent and the workflow underperforms, the cause may lie in the resource model, the network, the storage service, a credential, the application or a cross-domain timing problem. Orchestration needs evidence precise enough to tell each operator which part of the negotiated state diverged.
SENSE points towards an infrastructure in which applications request outcomes rather than circuits. Its progress should be judged by repeatable multi-domain production use, clear support responsibility and the ability to recover from partial failure, rather than by the highest demonstrated transfer rate alone. (SENSE; 2024 applied research)
EJFAT moves scientific events directly to remote computing
Traditional scientific data movement often follows a file-based sequence. An experiment produces data, local systems process enough of it to write files, the files enter storage, and a later transfer sends them to another facility for analysis. That pattern is well understood and operationally dependable, but it delays feedback and requires local storage and computing sized for peak acquisition.
EJFAT—the ESnet–Jefferson Lab FPGA Accelerated Transport system—supports a different path. It distributes live, UDP-encapsulated event data from an instrument to available compute workers, potentially at a remote supercomputing centre. A programmable FPGA-based SmartNIC reads event identifiers and sends all fragments belonging to the same event to one worker. As compute nodes become available or busy, the control system can change the distribution.
The event grouping matters. A generic load balancer can spread packets across servers, but scientific processing may require every fragment of one detector event to arrive at the same worker. EJFAT separates the high-rate forwarding function from the scheduling of compute workers and allows each side to scale independently.
The April 2024 Jefferson Lab-to-Perlmutter demonstration described in the opening reached 100 Gbps without an intermediate disk stage. Later tests used computing resources across several facilities and approximately 20,000 cores, according to ESnet reporting. These results show the system’s potential, but they do not mean that every instrument can adopt it without engineering changes.
The Facility for Rare Isotope Beams provides a more recent operational example. ESnet reported in July 2026 that researchers streamed raw experiment data through EJFAT at roughly 5–6 Gbps to remote supercomputers. A run that generated 615 GB over 15 hours was processed in 20 minutes on eight Perlmutter nodes using machine-learning inference. The lower network rate compared with the 100G demonstration does not make the result less important; the scientific outcome was faster feedback during a real workflow.
Live streaming changes the economics of a facility. A detector can borrow shared computing instead of building enough local capacity for every peak. Researchers can inspect results while scarce beam time remains available. The system also creates new dependencies. A network interruption, compute-allocation problem or event-distribution error can affect the experiment in real time rather than delaying a later file transfer.
EJFAT remains an advanced prototype or emerging platform rather than a universally supported service. Production adoption will require instrument integration, security, FPGA availability, operational ownership, scheduling agreements and clear behaviour when UDP packets are lost. Its significance lies in showing that the wide-area network can sit inside the acquisition loop rather than after it. (EJFAT; FRIB science highlight)
High-Touch, perfSONAR and iperf3 make the path observable
A high-capacity network can fail in ways that aggregate utilisation charts do not reveal. Short bursts can overflow a queue. Packets can arrive out of order. Loss can affect one direction or one class of traffic. A route can change between tests. Security-relevant flows can disappear inside statistical sampling.
High-Touch uses programmable hardware and software to collect richer packet and flow telemetry than traditional sampled records. In 2024 it was used to investigate Large Hadron Collider throughput, packet reordering associated with Rubin Observatory traffic and security events. The purpose is diagnostic precision: identify behaviour that a coarse traffic total would hide.
Richer evidence has a cost. Packet and flow detail requires storage, processing, access control and careful interpretation. It can reveal communicating endpoints, collaboration patterns and the timing of scientific activity. A system designed for operational visibility therefore becomes a sensitive data asset.
perfSONAR addresses a different layer. It is a distributed measurement system supported by a multi-organisation partnership in which ESnet is a founding core participant. More than 2,000 locations were reported globally in 2024, and ESnet operates more than 30 sites. Controlled tests measure throughput, latency, loss and path behaviour between participating nodes.
The distributed design is central. A researcher can report a slow transfer, but a test between well-engineered perfSONAR nodes can show whether the underlying path supports the expected rate. Directional tests can isolate asymmetry. Repeated measurements can show when degradation began. Traceroute data can identify path change, although the observed route and the application’s actual forwarding behaviour may differ.
iperf3 provides active throughput testing over TCP, UDP or SCTP and supports IPv4 and IPv6. ESnet maintains the open-source tool and reports tests above 150 Gbps on 200G paths under specific conditions. A synthetic test is not an application benchmark. It helps determine whether the host and network can move traffic under controlled settings, after which storage and application behaviour can be examined separately.
Together, High-Touch, perfSONAR and iperf3 create different views of truth. High-Touch observes production flows with greater fidelity. perfSONAR supplies scheduled or on-demand path measurement across domains. iperf3 tests endpoint and network throughput directly. No single view is complete, but the combination reduces the temptation to diagnose a distributed failure from one operator’s local dashboard. (network performance tools; 2024 operational innovations)
Caching can remove traffic instead of carrying it faster
Adding capacity is not the only way to improve a scientific network. Some datasets are requested repeatedly by researchers in the same region. Moving every copy across the wide-area network wastes bandwidth and increases delay even when the backbone can absorb the load.
ESnet’s 2024 applied-research reporting described five regional cache nodes used with scientific data communities. Across the study, caches reduced wide-area traffic by an average of 33 terabytes per day. Two-thirds of requests were served locally, and the average cache hit rate was 94 per cent. The effect varied sharply by location: reported wide-area reduction reached 69.3 per cent in Southern California, 48.4 per cent in Chicago and 6.6 per cent in Boston.
Those differences are informative. Cache value depends on dataset popularity, local users, storage size, eviction policy and service configuration. A cache placed near a community repeatedly analysing the same data can remove large transfers. A cache serving diverse or rarely repeated datasets can consume storage without delivering the same benefit.
Caching also changes control. The network operator begins to participate in data placement, while scientific communities must decide which datasets can be copied, how freshness is maintained and who is authorised to access them. Storage faults and stale content become network-adjacent operational concerns.
The study shows why the most efficient bit can be the one that never enters the backbone. Optical upgrades increase supply. Caching changes demand. Workflow scheduling can move transfers away from busy periods. Compression or local filtering can reduce data before transmission. A mature infrastructure programme evaluates all of these mechanisms rather than measuring progress only through installed terabits.
The result should remain scoped to the studied deployment. It does not prove that every scientific dataset will achieve a 94 per cent hit rate or that caching can replace new capacity for unique live streams. It provides evidence that data architecture and network architecture can be designed together. (2024 applied research activities)
Transatlantic science requires physical diversity as well as capacity
ESnet’s domestic footprint cannot be separated from its international obligations. High-energy physics is organised around facilities and computing centres on both sides of the Atlantic, and the 2024 Annual Report attributed roughly half of ESnet’s traffic to Large Hadron Collider activity. A cable failure between Europe and North America can therefore affect a central scientific workload even when every domestic router is healthy.
During 2024 and early 2025, ESnet expanded reported transatlantic capacity from about 700 Gbps to 2.7 Tbps. A major part of the programme was a 15-year agreement with Aqua Comms for one quarter of a fibre pair on a route linking New York, Dublin and London. ESnet also shares capacity and costs with GÉANT across several cable systems. The engineering target is at least 3.2 Tbps over four physically diverse paths, with the underlying design offering growth potential beyond 10 Tbps.
Capacity and diversity solve different problems. Two logical services can share a cable, landing station or terrestrial duct and fail together. Subsea repair may require fault localisation, permits, a cable ship, suitable weather and the physical recovery of damaged fibre. Restoration can take weeks or months. ESnet therefore plans for scenarios involving three simultaneous cable failures rather than assuming that one backup path is enough.
Long-term spectrum rights give the facility more control over upgrades than repeated purchases of fixed finished circuits. They also create commitments to particular cable systems, landing infrastructure and operating partners. ESnet does not remove marine risk by leasing spectrum. It gains the ability to engineer capacity and resilience across several systems while remaining dependent on cable owners, repair processes and partner networks.
This is another place where the word “dedicated” needs care. ESnet controls dedicated rights and infrastructure over important routes, but the physical asset model varies. Domestic and international paths can involve dark fibre, spectrum, lit services, colocation and partner capacity. A map line does not disclose the same ownership or repair authority on every segment.
The GÉANT relationship shows why research networking is cooperative without being centrally controlled. The two organisations can share cost and coordinate science traffic, yet each remains accountable to its own institutions and members. End-to-end performance still depends on national research networks, campus systems and experiment facilities beyond either backbone. (transatlantic milestone; 2024 Annual Report)
Science arrives as steady scale, rare bursts and hard deadlines
The Large Hadron Collider supplies steady scale. Its distributed computing model moves experiment data among CERN, Fermilab, Brookhaven, universities and other centres, creating sustained international traffic. Planned high-luminosity operation is expected to increase detector output, simulation and replication. The network must carry routine enormous flows while retaining enough route diversity to survive cable or facility failures.
The Vera C. Rubin Observatory supplies a deadline. The telescope in Chile is designed to produce an image of roughly 13 GB every 30 seconds. The data path supports rapid processing at SLAC and the generation of alerts about transient objects so that other observatories can respond. ESnet reports a target of moving each image across approximately 12,000 miles in less than seven seconds. The route includes South American and international research-network partners, so the result depends on coordinated engineering beyond ESnet’s own domain.
DUNE supplies an exceptional burst. During a supernova event, the neutrino programme anticipates a requirement to move as much as 600 TB in 100 seconds. Its wider twenty-year programme is expected to generate about 900 PB. These are planning requirements for future operation, not a description of current routine traffic. They matter because the event cannot be scheduled after capacity becomes available. A network built only around average utilisation could fail at the moment of highest scientific value.
Fusion research combines international distance with a long development horizon. In May 2026, a test moved 176 TB of ITER data from Marseille to General Atomics’ DIII-D facility in San Diego at nearly 80 Gbps. ITER was still under construction, so the transfer demonstrated readiness rather than normal full-scale production. Requirements material anticipates some future fusion workflows around two petabytes per day and at least 200 Gbps.
FRIB and Jefferson Lab show a third pattern: the experiment may need remote analysis while data is still arriving. EJFAT allows event streams to reach computing at NERSC or Oak Ridge rather than waiting for files to be completed and copied. Light sources, electron microscopes and other instruments face related choices about local filtering, remote artificial-intelligence inference and the speed at which results must return to the operator.
Climate and Earth science broaden the demand again. Large simulations and observational datasets move among computing and storage facilities, while field sensors can sit where conventional fibre is unavailable. No single headline rate captures all of these needs. ESnet has to support persistent bulk transfer, rare bursts, low-latency feedback, international collaboration and remote acquisition under one facility model. (Rubin Observatory case study; DUNE case study; requirements reviews)
Requirements reviews turn scientific plans into network architecture
ESnet cannot wait for an experiment to produce its first dataset before ordering fibre, routers or transatlantic capacity. Long-haul agreements, equipment procurement, site construction and software development take years. The facility therefore conducts requirements reviews with Department of Energy science programmes over a five-to-ten-year horizon.
Researchers and facility teams describe instruments, expected data volumes, storage locations, computing destinations, timing constraints, international collaborators, cloud use and resilience needs. Network and computing specialists then examine the complete workflow. A request for a larger link may expose a more important problem in local storage, route diversity, security architecture or data-transfer software.
The review process is a form of co-design. Scientific programmes explain what they are trying to accomplish; ESnet translates that objective into infrastructure dependencies and tests whether the proposed path can work. The result can influence site-link size, new routes, subsea procurement, Science DMZ deployment, orchestration research, staffing and the schedule for a future network generation.
The major High Energy Physics review completed in 2025 and published in January 2026 involved 127 contributors and 14 case studies in a report of roughly 400 pages. Its significance is not that every forecast will be correct. Instruments slip, software becomes more efficient and new workloads appear. The value is a documented record in which researchers, network engineers, computing centres and funders can see the same assumptions and revise them deliberately.
Forecasting remains difficult because scientific demand is discontinuous. A supernova may never occur during an experiment’s lifetime, while an artificial-intelligence workflow can expand faster than a formal review cycle. Average traffic growth is useful for budgeting, but it cannot substitute for case-specific planning. ESnet’s unusually low four-per-cent traffic increase in 2024 does not show that long-term demand has stopped; it shows why one year should not determine a multi-decade infrastructure programme.
Requirements reviews also distribute accountability. A laboratory cannot assume that the national backbone will repair an inadequate campus path, and ESnet cannot assume that a science programme will adapt after capacity is built. Putting the dependencies into a shared plan makes later disagreements more concrete. (requirements-review reports)
Wireless and quantum projects test where ESnet’s remit should end
A national fibre network reaches fixed institutions well. Scientific instruments do not always operate in those places. Geothermal fields, environmental observatories and temporary campaigns can lack carrier coverage, reliable power or a practical fibre route. ESnet’s Wireless Edge programme investigates how private cellular, Wi-Fi, directional radio and satellite systems can extend a scientific workflow to those sites.
A Nevada geothermal deployment publicised in June 2026 combined private 4G using Citizens Broadband Radio Service spectrum, long-range Wi-Fi HaLow, ordinary Wi-Fi, Starlink backhaul, directional radio and a portable self-powered tower. The design answered the conditions of one field site. It was not a general ESnet wireless product, and it cannot deliver the sustained capacity of long-haul fibre. Terrain, weather, spectrum, power and the satellite provider all remain part of the service path.
The important organisational point is that ESnet did not stop at the nearest fibre point and call the remaining problem somebody else’s responsibility. Its science-engagement function treated acquisition, backhaul and wide-area transport as one workflow. That approach can be more valuable than a uniform product because remote scientific sites have different physical constraints.
QUANT-NET explores another frontier. The ASCR-funded project is building a three-node quantum-network testbed linking Berkeley Lab and two University of California, Berkeley locations over approximately five kilometres. Work includes ion traps, photon coupling, entanglement swapping, Bell-state measurements, time-critical control and a modular software framework.
The connection to ESnet is principally control and orchestration. Quantum experiments need classical communication, precise timing and coordination among devices that are still difficult to operate. Open-source control software can make experiments more repeatable and reduce manual adjustment. The testbed is not a production quantum internet, and it does not move ordinary bulk scientific data through entanglement.
Wireless Edge and QUANT-NET should therefore be measured by what they teach and what can be reused. Neither is evidence that every ESnet user will receive a wireless access service or quantum link. They extend the facility’s research function into areas where future scientific networking may require new physical media and control systems. (Wireless Edge deployment; quantum networking research)
American Science Cloud and ESnet7 shift planning from sites to workflows
The Integrated Research Infrastructure programme and the emerging American Science Cloud reflect a change in how the Department of Energy describes its facilities. Instead of treating an instrument, a network, a data store and a supercomputer as independent services, the programmes aim to make them operate as a federated environment for scientific data, artificial intelligence and high-performance computing.
ESnet’s role is the connective and orchestration layer. Data must move from instruments to storage and accelerators; users and services need controlled access; workflows must discover available resources; telemetry has to show whether the system is meeting the scientific objective. EJFAT, SENSE, OSCARS, Cloud Connect and High-Touch each address part of that problem, but none alone constitutes the American Science Cloud.
At Confab26, Inder Monga was identified as American Science Cloud Project Deputy, and programme sessions included demonstrations and planning for the next network generation. This is evidence of active institutional work, not proof that a universal scientific cloud was complete in August 2026. Access policy, connected-resource inventory, production support and governance remained developing questions.
ESnet7 appears in the same planning environment. The label is real, and innovation sessions have examined telemetry, packet inspection, data movement, artificial intelligence and quantum networking. ESnet6 nevertheless remains the current production generation. No final public architecture, project baseline, vendor selection, complete budget or launch date was identified in the supplied evidence.
Network-generation labels create expectations long before equipment is installed. A planning programme can shape research priorities, vendor engagement and capital requests. It should not be reported as a completed backbone. The most defensible interpretation is that ESnet is using the operational experience of ESnet6 and current workflow projects to decide what a seventh generation should be.
Speed will remain important, but the evidence suggests that programmability and intelligence may become equally central. A faster link cannot tell an application where compute is available, identify a microburst, schedule a detector stream or prove that a cross-domain service has been released correctly. ESnet7’s strategic question is therefore how much of the scientific workflow the network should understand and coordinate without becoming an unmanageable central controller. (Confab26 programme; 2024 applied research)
Federal funding supports the facility, but the budget line is not revenue
ESnet does not finance its backbone by selling bandwidth to the public. The Department of Energy funds it principally through ASCR’s activity called High Performance Network Facilities and Testbeds. The FY2027 request proposed $103 million for that programme line, compared with $97.261 million in the preceding enacted column.
Those figures provide the best public programme-scale indicator, but they are not an ESnet income statement. The line also supports network-related testbeds, software stewardship, upgrades and research activity. It does not disclose a separate annual operating cost for the production backbone, capital expenditure by network generation or the revenue associated with individual site connections.
Public funding allows the facility to invest before demand becomes commercial. It can buy long-term spectrum, maintain open-source tools, support experiments with low volume but high scientific importance, and build route diversity beyond what immediate utilisation might justify. The same model creates dependence on congressional appropriations, Department priorities and the Berkeley Lab operating framework.
Site connections are not always costless. ESnet’s Site User Cost Policy allows institutional users to bear connection-related costs depending on sponsorship, eligibility and incremental requirements. The unit of the relationship is the site, not the individual researcher. Endpoint users are not separately billed or enrolled through ESnet.
The absence of standalone accounts limits financial analysis. Public material does not reveal a complete supplier concentration, payroll, depreciation schedule, route-by-route asset register or annual capital budget. Calling $103 million “ESnet revenue” would therefore be incorrect. The more accurate conclusion is that a federal programme funds a production facility, its upgrade work, software stewardship and testbeds through a broader activity line.
The funding structure also affects strategic choice. A commercial network can withdraw from a route that does not cover its cost. A public scientific facility must weigh national mission, scientific opportunity and long-term capacity against budget constraints. That can justify investment with no immediate financial return, but it also makes transparent prioritisation and requirements evidence essential. (FY2027 ASCR budget request; ESCC and site policy)
Reliability and visibility create obligations beyond uptime
ESnet’s 2024 reporting recorded a notable availability result: all ten Office of Science sites included in the relevant measure achieved 100 per cent service availability excluding planned maintenance. It was the first year that every site exceeded the 99.9-per-cent requirement at the same time. The annual report also gave a wider uptime figure of 99.99 per cent. These are useful but differently scoped measures, not proof that every service, institution and international path experienced zero interruption.
The Site Resilience Program examines diverse entrances, routers, power and failure domains at connected facilities. This reflects an uncomfortable fact: a national backbone can be redundant while a laboratory remains dependent on one local conduit or device. Reliability has to reach the building and the data-transfer system, not stop at the backbone map.
Routing policy supplies another layer of control. ESnet rejects RPKI-invalid announcements in peering, requires current routing records and supports route-hijack monitoring and destination blackholing. These controls reduce known risks but cannot make interdomain routing infallible. A valid route origin can still be associated with a leak or policy error, and a blackhole protects other systems by making the targeted destination unreachable.
Automation creates a similar trade-off. Standard configuration, inventory and orchestration can reduce manual error and speed recovery. A wrong template, policy or data record can also propagate quickly across many devices. The appropriate response is staged deployment, independent validation, clear rollback or forward-recovery procedures, and limits on the authority of any single automation account.
Operational data deserves equal attention. ESnet’s facility data policy says router utilisation and NetFlow data can be retained indefinitely and replicated across east- and west-coast storage. Active perfSONAR measurements are retained for six months on a single disk without backup. Router utilisation and traceroute data are public, while flow and security records have restricted access.
ESnet can lack an individual subscriber relationship while still collecting detailed operational metadata; these are different systems. ESnet lacks an individual membership and billing relationship, yet network metadata can identify endpoints, communicating institutions, timing and traffic patterns. High-fidelity telemetry improves diagnosis and security while increasing the need for minimisation, access control, retention review and accountable use.
A credible resilience model rests on bounded mechanisms: physical diversity, site engineering, routing security, measurement, incident response, data governance and clearly scoped metrics. (availability milestone; facility data policy; peering policy)
The test is whether distributed science works in practice
ESnet’s history can be read as a sequence of faster networks: specialised predecessors, ESnet4, a continental 100G ESnet5 and the multi-terabit ESnet6. That chronology is real, but it misses the more important continuity. The facility repeatedly changed the architecture around scientific work when raw bandwidth alone could not solve the problem.
The Science Data Network separated exceptional flows from ordinary traffic. OSCARS made capacity reservable. Science DMZ changed the institutional perimeter. perfSONAR and iperf3 turned performance disputes into measurements. SENSE linked application intent to several administrative domains. EJFAT brought the wide-area path into the detector loop. Caching changed demand, and transatlantic spectrum changed ESnet’s control over international growth.
The organisation’s authority remains bounded. DOE and ASCR set mission and funding. Berkeley Lab operates the facility. The University of California manages the laboratory under contract. Site coordinators approve changes affecting their institutions. International partners control their own networks. Cloud providers, carriers, cable owners and equipment vendors retain power over components on which ESnet depends.
That distribution of control is not an inconvenience that software can erase. It is the condition under which a scientific network operates. The practical achievement is coordination strong enough to deliver a workflow without pretending that one institution owns the entire path.
The same discipline should govern claims about the future. ESnet7 is planning, SENSE remains research phase in the latest complete report, EJFAT is an advanced prototype or emerging platform, QUANT-NET is experimental, and the American Science Cloud was still being built. Their evidence is valuable precisely when maturity is described accurately.
ESnet’s longer-term significance lies in the user-facility model. The network is treated as a scientific instrument with requirements, operations, software, measurement, research and public funding. That model recognises that a detector and a supercomputer separated by thousands of miles create little combined value unless the systems between them are engineered as one path.
No neutral current measure in the supplied evidence can rank ESnet as the world’s fastest network, and that is not the most useful test. The stronger measure is whether an instrument can reach remote storage or computing at the required rate, recover when one part of the path fails and produce evidence that shows where responsibility lies. ESnet’s standing will be confirmed or weakened by those production outcomes as American Science Cloud work, live streaming and the eventual ESnet7 programme move from demonstrations and plans into supported scientific service. (ESnet history; ESnet6 launch; 2024 Annual Report)
Sources
- ESnet governance and policies
- ESnet history
- ESnet 2024 Annual Report
- ESnet by the Numbers: 2024
- 2024 operational innovations
- 2024 applied research activities
- Transatlantic spectrum milestone
- Availability milestone
- ESnet leadership
- ESnet Site Coordinators Committee
- Network services menu
- Peering connections and routing policy
- Facility data management policy
- DOE ASCR FY2027 budget request
- ESnet6 launch at Berkeley Lab
- ESnet4 architecture and Science Data Network
- ESnet5 continental 100G launch
- Science DMZ
- OSCARS
- SENSE
- EJFAT
- Network performance tools
- Science requirements reviews
- Wireless Edge geothermal deployment
- Quantum networking research
- FRIB direct data-streaming result
- Rubin Observatory case study
- DUNE case study
- Confab26 programme
Strategic Circle: indicators, triggers and scenarios
ESnet’s next stage should be judged through operating evidence. The first indicator is a formal ESnet7 baseline. A funded project plan, defined architecture, procurement schedule and migration method would move the programme from innovation work into an accountable network-generation project. Until those elements appear, ESnet6 remains the production system and “ESnet7” is a planning label.
The second indicator is how the network absorbs the next scientific demand cycle. High-Luminosity LHC operation, Rubin commissioning, DUNE development, ITER data production and the growth of artificial-intelligence workloads can change traffic shape as much as traffic volume. Useful evidence will include site-link upgrades, sustained production rates, burst tests, missed deadlines and the number of workflows that require real-time rather than deferred transfer.
International resilience should be tracked separately from aggregate bandwidth. Reaching the planned 3.2 Tbps across four genuinely diverse transatlantic paths would be a material milestone. A major cable failure would provide a harder test: how much traffic can be restored, how quickly, and which scientific workflows receive priority when several systems are unavailable at once.
The transition of research software into supported service is another trigger. EJFAT would cross an important boundary if ESnet publishes a production support model, integration requirements and service-level responsibilities used by several instruments. SENSE would change status when multi-domain orchestration becomes routine outside selected demonstrations. High-Touch should be assessed by the operational decisions it improves and the controls governing the resulting data, not by telemetry volume alone.
American Science Cloud work requires the same discipline. A production launch should define eligible users, connected compute and storage, identity and authorisation, service support, data responsibility and the role of commercial clouds. Demonstrations show technical direction. They do not yet answer who is accountable when a workflow spans a DOE facility, ESnet and a commercial provider.
Funding is a practical trigger. The final FY2027 appropriation should be compared with the $103 million request for High Performance Network Facilities and Testbeds, while remembering that the line is broader than ESnet. A sustained funding reduction would not necessarily interrupt current traffic immediately, but it could delay optics, site resilience, software maintenance, international spectrum and ESnet7. A large increase could accelerate the programme while also raising the risk of pursuing too many research fronts at once.
Operations teams should watch the number of 400G site connections, the introduction of 800G or higher packet services, workforce growth and the age of critical optical and routing platforms. They should also watch traffic after the unusual four-per-cent growth recorded in 2024. A return to high growth would validate aggressive capacity planning; continued slower growth would strengthen the case for selective upgrades, caching and workflow optimisation rather than uniform expansion.
Several scenarios follow. In a workflow-integration scenario, ESnet becomes the connective and orchestration layer of the American Science Cloud, with live instrument streams and shared computing becoming normal facility design. In a capacity-constrained scenario, international science grows faster than subsea diversity and forces priority decisions. In a site-bottleneck scenario, the national backbone scales successfully while local storage, firewalls and campus links remain the limiting factor. In a funding-constrained scenario, ESnet6 lives longer and software, coherent-optics upgrades and caching carry more of the burden.
Wireless and quantum work may generate reusable control systems without becoming mainstream production services during the next generation.
For laboratory leaders, the professional implication is straightforward: an ESnet upgrade does not substitute for local engineering. Site resilience, Science DMZ design, data-transfer nodes, storage and measurement should be funded as part of the same scientific capability. For programme managers, the relevant question is whether a planned instrument has an executable end-to-end workflow, not whether a national backbone slide shows enough capacity.
Leadership Alliance: control, incentives and irreversible decisions
ESnet’s control structure is distributed among institutions with different incentives. The Department of Energy and ASCR define mission and allocate federal funds. Berkeley Lab’s Scientific Networking Division designs and operates the facility. The University of California provides the laboratory operating framework. ESnet leadership allocates engineering attention among operations, security, software, research and science engagement. Site coordinators control changes that affect their own connections.
Scientific programmes articulate future requirements, while partner networks, cable providers, cloud companies and equipment vendors control parts of the end-to-end system.
That structure prevents one actor from commanding the whole scientific path, but it can also make failure responsibility ambiguous. A central programme can fund a backbone without fixing a local facility. A laboratory can request capacity without preparing storage. A cloud provider can deliver a private interconnect while retaining control over regional availability and pricing. Effective leadership therefore depends on making boundaries explicit and requiring evidence at each hand-off.
The first decision is how much authority the network should acquire over the workflow. More orchestration can allocate capacity, storage and compute efficiently. It can also concentrate credentials and create a large automation blast radius. A prudent design keeps policy ownership with the institutions that bear the consequences, exposes narrow capabilities through authenticated interfaces and requires independent verification before high-impact changes.
The second decision is where to commit capital irreversibly. Fibre and spectrum rights, landing arrangements, optical platforms and site construction can shape the network for fifteen years or more. Waiting preserves flexibility but risks being unable to obtain capacity when an experiment begins. Early commitment secures control but can strand investment if science schedules or technologies change. Requirements reviews are valuable because they connect these commitments to named workflows instead of treating growth as a generic curve.
The third decision concerns data. Indefinite flow retention and high-fidelity telemetry can improve incident response, capacity planning and security. They also create durable records of scientific relationships and endpoint behaviour. Leadership should distinguish data that is essential for continuity from data retained because storage is available, document access and use, and review whether indefinite retention remains proportionate as telemetry becomes more detailed.
The fourth decision is the boundary between public infrastructure and commercial services. Cloud Connect can give researchers private paths to powerful resources, but the cloud account, security configuration, egress pricing and service availability remain external dependencies. ESnet should make those dependencies legible rather than allowing a private path to be mistaken for public control of the full service.
Second-order effects matter. If live streaming becomes standard, facilities may build less local compute and become more dependent on the wide-area network and national schedulers. That can improve utilisation while turning a backbone incident into an experiment incident. If application-driven orchestration becomes normal, scientific software may begin to assume that capacity can be reserved dynamically, making manual fallback harder. If caches and shared storage expand, data-placement policy becomes an infrastructure governance question.
Third-order effects reach the organisation of science. A reliable instrument-to-supercomputer path allows smaller facilities to use capabilities they could not finance locally. It can distribute access more broadly, but it can also concentrate essential computing in a few national centres. Artificial-intelligence workflows may increase demand for shared accelerators and give networking decisions greater influence over which experiments receive timely analysis.
The most serious irreversible risk is not one failed upgrade. It is designing a scientific ecosystem whose components can no longer operate independently enough to recover. ESnet should deepen integration while preserving observable boundaries, local fallback, route diversity, open interfaces and the ability to replace a supplier or service. The facility’s history supports that approach: its strongest contributions did not centralise every decision. They made shared infrastructure programmable and measurable while keeping responsibility visible.
Leadership should judge ESnet by three outcomes: whether scientific programmes receive dependable end-to-end capability rather than nominal bandwidth; whether long-lived investment preserves options; and whether automation and telemetry improve evidence without placing unreviewable authority in one system. Those tests determine whether the facility remains recoverable and useful as scientific demand, technology and public priorities change.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
