Summary

  • ESnet is a user science facility of the U.S. Department of Energy operated by Lawrence Berkeley National Laboratory, not a commercial telecommunications company or an independent institution.
  • ESnet6 encompasses approximately 15,000 miles of dedicated fibre, optical and packet systems, private services, peering, measurement, and programmable control; the 57 terabit-per-second figure represents aggregate capacity, not user speed.
  • Science DMZ, perfSONAR, OSCARS, SENSE, High-Touch, and EJFAT demonstrate that ESnet's distinctive contribution is end-to-end path engineering, from the instrument and site network edge to remote computing.
  • Plans for ESnet7 and the American Science Cloud point to a network more integrated into scientific workflows, though several components remain experimental demonstrations, research systems, or nascent services.

A science network that enters the heart of the experiment itself

In April 2024, data travelled from Jefferson Lab in Virginia across the United States to the Perlmutter supercomputer in California at 100 gigabits per second, and was processed without first being written to disk. The demonstration used ESnet's experimental EJFAT system, placing the wide-area network inside the science workflow rather than running it after data collection was complete. This shift sums up the central question facing ESnet: can the network connect instruments, storage, and remote computing so tightly that facilities separated by thousands of kilometres operate as parts of a single machine?

The typical path begins with a detector, telescope, microscope, or simulation inside a laboratory or research facility. Local acquisition systems gather data, storage and transfer servers prepare it for movement, and the site network carries it to a controlled institutional perimeter. From there ESnet may move traffic across the United States or the Atlantic, to a DOE supercomputer, via a research and education network partner, or into a commercial cloud environment. The result may come back as an analysed dataset, alert, model, or decision that determines the next step of the experiment.

No single institution controls this entire path. Instrument teams, laboratory networks, regional research networks, cloud providers, international partners, and supercomputing facilities may fall under different organisations. ESnet operates the wide-area layer oriented to the science mission, working with these operators to make the path perform as one system. A slow storage server, congested site link, or inappropriate firewall can waste the capacity of a fast national backbone, just as a fault within ESnet can disable a well-designed local facility.

Institutionally, ESnet is a user facility of the DOE Office of Science. The Advanced Scientific Computing Research programme provides primary oversight, while the Scientific Networking Division at Lawrence Berkeley National Laboratory operates it. Berkeley Lab itself is managed for DOE by the University of California under contract. ESnet has no known shares, shareholders, enterprise valuation, or independent commercial budget, and its role is therefore better understood as a federal scientific infrastructure than as a network company with a general customer base.

Researchers usually encounter ESnet indirectly through an instrument, supercomputer, laboratory service, university connection, or partner network. The facility distinguishes between institutional site users and end users, and says it does not formally log or track each individual end user. The figure of more than 30,000 organisational users in the 2024 annual report is thus an institutional measure, not a subscriber count.

Ordinary comparisons with telecoms companies explain only part of the role. ESnet manages routing, optical transport, and peering, but it also conducts requirements reviews, develops open-source tools, trials new services, and helps sites diagnose end-to-end performance. Its overall purpose is not to sell connectivity, but to enable geographically dispersed science facilities to work together with the speed, evidence, and operational clarity needed to support the research itself. (ESnet governance;Network services)

Remote computing created the need for a shared science network

The formal history of ESnet reaches back to the mid-1970s, when long-distance computing and communications were rare. At the Controlled Thermonuclear Fusion Research Computing Center inside Lawrence Livermore National Laboratory, staff connected a borrowed Control Data Corporation 6600 computer via four voice-grade modems. The equipment and speeds belong to another era, but the operational problem remains familiar: a specialised scientific community needed remote access to expensive computing that could not be duplicated at every participating institution.

Separate research communities within DOE built their own networks during the late 1970s and early 1980s. High-energy physics and magnetic fusion research had different facilities, collaborators, and data flows, and earlier networks such as HEPnet and MFEnet reflected those boundaries. Building dedicated networks made sense when commercial services could not supply the bandwidth, performance, or operational attention required, but it also led to duplication. Each programme negotiated circuits, maintained technology, and built expertise independently, keeping cross-programme cooperation and long-term upgrade fragmented.

The formal creation of ESnet in 1986 unified these functions within a broader science network. The change was as much organisational as technical. A shared facility could forecast demand across programmes, operate national links, coordinate international connections, and retain specialist expertise. Capacity could also shift from being a single-project concern to a department-wide infrastructure plan.

The early model embedded several characteristics that are still visible. Users were defined by science mission, not by a retail market. Central computers and instruments justified shared public investment. Network design followed research programmes with long timelines. Operations needed people who understood both communications systems and the scientific consequences of failures. Demand also arrived unevenly because a single experiment could produce traffic far exceeding the normal level.

Berkeley Lab has its own networking history. A CDC 6600 was connected to the ARPANET in 1974, and researchers there later contributed to foundational work on TCP congestion control. The work of Van Jacobson and Mike Karels belongs to the broader Berkeley Lab environment and should not be attributed to ESnet alone. Still, the institutional connection matters because it placed the operation of a production science network close to researchers who treated protocol behaviour, measurement, and performance as study-worthy engineering problems, not fixed conditions to be accepted.

Operations moved to Berkeley Lab in 1996. That move established ESnet's current home alongside the National Energy Research Scientific Computing Center (NERSC), network researchers, software engineers, and other DOE programmes. It also reinforced a culture in which a production facility and an applied research organisation share staff, laboratories, and technical questions. (ESnet history)

Public funding lets ESnet build before demand arrives

A user science facility exists because some capabilities are too costly, too specialised, or too interdependent for every research team to build on its own. That logic applies to a particle accelerator, a light source, or a leadership-class supercomputer, and ESnet applies it to communications. Long-haul fibre rights, optical equipment, routers, international capacity, ongoing operations, cybersecurity, performance measurement, and engineering support are pooled inside one facility that serves multiple programmes.

The case for public funding flows from the shape of scientific demand. A telecoms company can sell a high-capacity circuit, but it does not normally design its national architecture around a detector that will start operating years later, a supernova data burst that may not occur within the contract term, or a research workflow of low commercial volume and high public value. ESnet can build before measured use appears because its mission is scientific capability, not near-term network revenue.

The facility model also changes how demand is discovered. ESnet conducts formal requirements reviews with DOE science programmes. Researchers and facility teams describe instruments, data volumes, storage locations, computing destinations, timing needs, collaboration patterns, and local bottlenecks over a five- to ten-year horizon. The results become inputs for sizing site connections, path diversity, transatlantic procurement, coordination research, and service design.

This process may seem less exciting than a new coherent optical link, but it may be more consequential. A submarine spectrum agreement, a router purchase, or a second site entry can require years of funding and execution. Waiting for an experiment to produce its data turns anticipated infrastructure needs into emergencies. Requirements reviews convert science plans into an engineering lead time within which to act.

The reviews also reveal uncertainty. Forecasts change, instrument schedules slip, data-reduction software improves, cloud use expands, or a facility redesigns its workflow. ESnet does not treat a case study as a guaranteed traffic contract; it uses the evidence to build a range of requirements and then decides where flexibility, reserve capacity, and phased deployment are justified.

The facility model avoids billing individual scientists per gigabyte. Core operations are federally funded, while some site connection costs may be allocated according to programme sponsorship and the Site User Cost policy. This arrangement is not free of incentive differences; a sponsoring programme may prefer a lower connection cost, ESnet may prefer capacity and resilience that serve multiple programmes, and a site may delay local upgrades needed to take advantage of the national network. Governance aims to align these decisions with scientific outcomes rather than with a competition to shift cost into another budget line. (ESnet Site Coordinators Committee and cost policy;Requirements review reports)

The DOE sets the mission and Berkeley Lab operates the network

ESnet does not own the corporate hierarchy of a telecoms carrier. Its authority is distributed across a chain of federal programmes. The DOE Office of Science defines the mission, the Advanced Scientific Computing Research programme provides primary oversight and budget control, and Berkeley Lab hosts the Scientific Networking Division that designs, operates, and evolves the facility. The University of California operates Berkeley Lab for the department under contract, and connected institutions participate through site coordinators and the ESnet Site Coordinators Committee (ESCC).

Inder Monga serves as ESnet executive director and director of the Scientific Networking Division at Berkeley Lab. Current public sources identify Chin Guok as chief technologist and planning and innovation lead, Adam Slagell as chief security officer, Jon-Paul Herron as network engineering manager, and Susan Lucas as deputy for business operations. Beneath these positions are functions covering optical engineering, routing, site engineering, network operations, software, measurement, coordination, security, project management, and science communication.

The breadth of this structure corrects the simplified image of a national network as a collection of fibre and routers. ESnet needs people who can purchase spectrum, run BGP, build measurement systems, maintain open-source software, investigate packet behaviour, plan facilities, manage security evidence, and translate scientific requirements into network designs. The organisation is operational and developmental at the same time.

Public records contain unresolved detail concerning programme management. The current governance page names Benjamin Brown as the designated programme manager from ASCR, while recent requirements-review materials identify Carol Hawk as the ESnet programme manager. The evidence allows us to say that public sources have not yet been reconciled or that responsibilities may have changed or been divided, but it does not allow us to select a current exclusive arrangement without additional confirmation.

The ESnet Site Coordinators Committee provides a formal operational channel for connected institutions. Each site appoints a coordinator who can approve requests affecting the institution's ESnet connection, convey requirements, and participate in policy and planning discussions. This is institutional user governance, not a direct vote by every scientist whose data traverses the network.

The result is a practical distribution of responsibility. ESnet controls its backbone and services, the connected site controls its campus network equipment, data transfer nodes, local security, and power, ASCR controls funding at the programme level, science programmes define the consequences of delay or failure, and international partners control their own networks. The full workflow succeeds only when these separate authorities remain aligned. (ESnet leadership;Governance)

ESnet4 separated ordinary traffic from exceptional science flows

By the mid-2000s, distributed experiments were beginning to change the traffic pattern. The Large Hadron Collider was designed to produce data at CERN and distribute it through a global hierarchy of laboratories and computing centres. General Internet connectivity remained necessary, but a small number of science transfers could become large enough to dominate ordinary links. Building one undifferentiated network for both patterns made capacity management and service assurance difficult.

ESnet4 addressed the problem through a hybrid architecture. An IP core carried general science communications, while a separate Science Data Network (SDN) used high-capacity optical circuits for large flows and supported dynamic path provisioning. Metropolitan-area rings linked major laboratories, and collaboration with Internet2 and international partners extended the system beyond DOE sites.

The architecture separated traffic according to operational requirements, not a simplistic 'fast' and 'slow' division. A general IP service needs broad reach and flexible routing, while a planned multi-terabyte transfer might benefit from a reserved Layer 2 path with known start time, dedicated bandwidth, and known endpoints. The SDN recognised that a small number of predictable, high-value flows could justify a different control model than ordinary packet forwarding.

OSCARS emerged from this environment. The On-demand Secure Circuits and Advance Reservation System allowed authorised users or applications to request network resources for a defined period. The system had to find a path, check topology and policy constraints, reserve bandwidth and VLAN resources, create state in devices, and tear it down when the reservation ended. It turned a task that once depended on manual coordination into a service that software could request.

ESnet4 also demonstrated that raw speed was no longer the only bottleneck. A reserved circuit could protect capacity inside the wide area while an application remained slow because a storage system could not read at the required rate or a local firewall dropped packets. This lesson led directly to the Science DMZ architecture and end-to-end measurement.

This phase established a pattern that recurs in ESnet design. A science problem first appears as a traffic requirement, then the facility builds physical capacity, a control mechanism, and an operating method around it. When the method proves useful, the software and engineering guidance spread beyond the original experiment. (ESnet4 and ESnet5 historical material;OSCARS)

Science DMZ showed the bottleneck often sits at the edge

A national backbone can operate correctly while a scientist experiences poor transfer performance. The reason is clear: the backbone is only one part of the path. Institutional firewalls, shared university network cores, old routers, weak servers, packet loss, and storage limits can reduce a high-speed route to a fraction of its engineered capability.

Science DMZ represents ESnet's response to this end-to-end reality. The architecture places data transfer nodes on a high-performance path near a controlled institutional perimeter. Security policy is tailored to the limited services those systems expose, rather than forcing sustained science traffic through a general stateful firewall designed for many types of enterprise applications.

The term can be misunderstood as a call to remove security. Science DMZ relies instead on hardened hosts, router access control lists, monitoring, vulnerability management, limited exposed services, and operational discipline. The design changes where and how some controls are applied, and avoids placing a device with inappropriate throughput or session behaviour in the primary data path simply because it is part of the institution's default security pattern.

Performance measurement is part of the architecture, not an optional dashboard. perfSONAR nodes run controlled tests of throughput, packet loss, latency, and path variation. When two sites disagree about why a transfer is slow, the measurements help determine whether the problem begins at the host, the local link, the regional network, the ESnet backbone, or the international partner. Without that evidence, every operator can claim their segment is fine while the researcher remains unable to move data.

Data transfer nodes need their own engineering. Network interface cards, processor location, memory, disk or parallel file systems, TCP buffers, and transfer software together set the achievable throughput level. ESnet's Fasterdata guidance and technical consultations turn these details into repeatable practice. The approach has spread beyond DOE laboratories, though ESnet does not operate every deployment that uses the name.

Science DMZ became influential because it reframed a procurement question as a systems question. A faster circuit does not fix a bottleneck elsewhere. The useful unit of performance is the full workflow from storage to storage, or from instrument to computing, with each administrative domain measured and its responsibility clear. (Science DMZ guidance;Network performance tools)

ESnet5 made a 100G continental network a production service

The next major capacity step was enabled by the Advanced Networking Initiative and US economic stimulus funding. ESnet received $62 million under the American Recovery and Reinvestment Act to develop a 100 Gbit/s long-haul prototype and support the transition to ESnet5. The investment went beyond new routers: ESnet also acquired spectrum capacity rights across a national fibre footprint, giving the programme more control over wavelength deployment and upgrade.

ESnet5 entered production in late 2012. The DOE and ESnet described it at the time as the world's fastest science network. That historical claim was tied to a pioneering 100G continental deployment and should not be repeated in 2026 as an absolute global ranking. Research networks today report capacity numbers through different metrics, including individual interface speeds, aggregate backbone capacity, optical spectrum, and test beds. The available evidence does not contain a recent neutral comparison that declares a single winner under a universal metric.

The economic effect of 100G was larger than simply raising the interface speed tenfold. Science projects could plan regular long-distance transfers that would previously have occupied links for much longer periods. Laboratories could consolidate some computing or storage centrally instead of duplicating them locally, and international collaborations gained a more capable US backbone. The new capacity also revealed local constraints more quickly, increasing the importance of Science DMZ and perfSONAR.

ESnet5 retained the hybrid relationship between packet service and dedicated science paths. It also deepened the importance of automation. At 100G scale, reserving capacity, diagnosing faults, and maintaining consistent service across a national footprint could no longer rely on device-specific configuration.

The project illustrates how public infrastructure investment changes scientific choices before any single experiment uses the full network. Capacity creates headroom for future instruments, but reserve capacity is not waste in the same sense that an idle commercial service would be. Its value includes avoiding experiment delays, absorbing faults, and allowing new workflows to be tested without waiting for another build cycle.

Nevertheless, the investment must be judged against scientific use, and the design must remain adaptable when forecasts change. ESnet's subsequent work on measurement, caching, and coordination reflects an attempt to extract more scientific value from every installed bit, rather than treating continuous physical expansion as the only solution. (ESnet5 historical announcement)

ESnet6 combines optical capacity with programmable control

The ESnet6 project began in 2017 and was publicly launched on 11 October 2022 after approximately six years of design and construction. Berkeley Lab announced at launch more than 46 Tbit/s of aggregate capacity across a dedicated footprint of roughly 15,000 miles, with backbone links ranging from 400 Gbit/s to 1 Tbit/s. The 2024 annual report later gave a figure of 57 Tbit/s and a link range extending to 1.2 Tbit/s. The numbers belong to different dates and an expanding system, not to conflicting definitions of a static network.

The physical layer includes long-haul fibre routes, optical line systems, amplifiers, spectrum rights, co-location facilities, and local service entry points. The 2024 report counted 278 optical amplifier sites. Coherent optics carry multiple high-rate channels over the fibre, while amplifiers and reconfigurable optical equipment maintain signal levels and steer wavelengths. A path shown on a map may be under-pinned by dark fibre, spectrum, lit service, or an arrangement with a partner. The 15,000-mile description does not prove that DOE legally owns every cable.

Above the optical layer, routers and service edges provide IP, private routing environments, Layer 2 services, peering, and cloud connectivity. The annual report counted 78 router sites. Individual sites can connect at 10, 100, or 400 Gbit/s according to their requirements, while backbone links can aggregate multiple interfaces or wavelengths to reach higher figures.

The software layer distinguishes ESnet6 from a capacity-only upgrade. Automation maintains device and service state, and programmable measurement gives operators richer evidence. OSCARS provisions reserved paths, SENSE tests application-directed coordination across multiple domains, and High-Touch uses programmable hardware for richer packet and flow visibility. Application programming interfaces and testbed hooks allow successful research to move towards operations.

The build also placed greater weight on resilience. Diverse paths, multiple site entries, independent power, alternate routers, and high-availability service locations reduce the chance that a single failure disconnects a laboratory from the national system. But the protections have different scopes; a diverse backbone is unhelpful if both local circuits share the same duct, and a second router does not create resilience if both devices depend on the same power system.

ESnet6 should therefore be understood as a multi-layer facility. Fibre gives physical reach, optics create capacity, routers provide packet and private services, software delivers repeatability and programmability, and measurement proves whether the promised path is working. Institutional agreements allow crossing of boundaries that ESnet does not own. (ESnet6 launch;2024 annual report)

What 57 Tbit/s means and does not mean

Network statistics often compress different layers into one number. The 57 Tbit/s figure for 2024 represents the aggregate engineered capacity across the system. It is not a single physical link, not the throughput available to every laboratory, and not the traffic carried at any moment. The same report describes backbone links between 400 Gbit/s and 1.2 Tbit/s, seven US sites connected at 400G or above, and transatlantic capacity of 2.7 Tbit/s. Each number answers a different question.

The same report described 78 router sites and 278 optical amplifier sites across a footprint of roughly 15,000 miles. It connected all 17 DOE national laboratories, recorded 28 DOE user facilities, 277 relationships with research and education networks, commercial networks, and others in five countries, more than 30,000 organisational users, and 137 staff and contractors working in 24 states. These figures describe different populations and should not be added together as a customer count; routers, facilities, partner networks, institutional users, and workforce are separate measures.

A site connection represents capacity at one institutional boundary, while a backbone link represents capacity between two points inside the network. An optical system can carry multiple wavelengths, and a logical service may use only part of the physical layer. Aggregate capacity sums many resources that cannot all serve the same source and destination simultaneously. Traffic volume measures data transferred over a period, not the highest rate available at an instant.

ESnet carried 1.77 exabytes during calendar year 2024, up from 1.7 exabytes in 2023. The 4% increase was low compared with the facility's reported historical annual growth average of about 55% since 1989. A single year does not prove that scientific demand has stabilised. Large projects enter service irregularly, software changes can reduce transfers, and one experiment or international event can quickly alter the mix.

Traffic is also concentrated. The 2024 report said Large Hadron Collider activity accounted for roughly half of all traffic. Fermilab was the largest external sender at 136 petabytes, NERSC the largest DOE receiver at 75.3 PB, and Oak Ridge the largest DOE sender at 58.3 PB. These figures reveal important workloads but do not form a complete ranking of scientific value; a smaller transfer may be more time-sensitive or support a unique instrument.

Capacity planning therefore needs several viewpoints at once: peak and average traffic, per-path utilisation, failure reserve, site link rates, scientific timelines, and future requirements. A network designed around average usage alone could fail during a supernova burst or a cable cut, while a network designed around the largest theoretical event might build costly and inflexible assets.

ESnet's engineering problem is to place capacity where the workflows that need it can reach it, and then verify that the endpoint systems can use it. That is why the capacity number should sit alongside Science DMZ, measurement, coordination, and requirements reviews, not at the top of a story as a self-interpreting rank. (ESnet by the numbers)

AS293 serves a mission, not a retail market

ESnet's public autonomous system number is AS293. The network exchanges routes with DOE sites, research and education networks, commercial networks, and paid transit providers. The routing policy determines which paths can carry science traffic, how failures are handled, and which external networks receive direct peering.

Peering policy is selective. Potential partners must use BGP, maintain current information in Internet Routing Registry and PeeringDB, avoid RPKI-invalid announcements, and follow MANRS-related routing security practices. ESnet requires IPv6 peering or dual-stack, and does not accept new IPv4-only relationships. Direct private peering starts at a 100G port, with a strong preference for 400G.

These conditions reflect both mission and operating cost. Direct peering consumes ports, engineering time, and monitoring capacity. ESnet does not need to peer with every network simply because public paths exist; it seeks relationships that improve DOE science paths, reduce reliance on transit, or reach key collaborators. General Internet access is complemented by paid transit from Hurricane Electric and Lumen for destinations not efficiently reached through research networks or direct peering.

The network publishes BGP communities that distinguish ESnet sites, research and education networks, and commercial networks. Connected institutions can use these tags to apply their own policy, but downstream interpretation remains outside ESnet's control. They are metadata, not a universal guarantee that every network will make the same routing decision.

Routing security also has limits. Rejecting RPKI-invalid announcements prevents an important class of misdirected routes, but it does not stop every leak, compromised partner, legitimate origin misused, or inaccurate route-origin authorisations. Route-hijack monitoring can identify unexpected origin changes, but detection may come after propagation. Blackhole routing can protect a site during an attack by dropping traffic to a destination, at the cost of reachability to that prefix.

The mission-oriented design shows at the service boundary. ESnet carries general IP traffic but does not try to win a consumer transit market. Its policy is arranged around scientific access, resilience, and trusted interconnection. The network's value comes from the quality of the paths it creates for the community it serves, not from maximising the number of retail customers connected to AS293. (ESnet peering policy)

Private paths extend reach without removing external risk

The ESnet service list begins with physical connectivity. Eligible sites can connect locally or via a remote point of presence at speeds of 10, 100, or 400 Gbit/s. When a facility is distant from the dedicated footprint, a Service On-Ramp may use dark fibre or a lit circuit from an operator. Access design becomes part of the science path, and the last-mile supplier can become the dominant performance constraint or point of failure.

The Layer 3 IP service provides IPv4 and IPv6 routing. A private Layer 3 VPN creates a logically separate IP and BGP environment for a multi-site programme, sharing physical infrastructure with other services. Layer 2 VPNs provide point-to-point or multipoint Ethernet connectivity, either through static configuration or via OSCARS. These services let experiments link facilities without exposing every relationship over the public Internet.

A private service does not mean complete physical isolation. Logical separation depends on router configuration, tags, routing instances, access lists, and operational practice. A fault in the shared platform can affect multiple services, and a configuration error can defeat the intended separation. The value lies in controlled segmentation and predictable connectivity, not in a claim that no shared dependency exists.

Cloud Connect extends ESnet paths to major commercial cloud environments including AWS, Microsoft Azure, Google Cloud, and Oracle. A site can reach the provider's peering point over a dedicated Layer 2 or Layer 3 connection instead of general Internet routing. This can improve path control and reduce exposure to variable transit, but it does not remove cloud security, provider outages, identity design, egress fees, or zone availability.

The cloud relationship illustrates ESnet's changing role. DOE science no longer exists only inside laboratories and federal supercomputers. Researchers may use commercial object storage, specialised services, or burst computing. ESnet must connect these resources without claiming that the cloud is part of the federal network, or that a private circuit makes a cloud account inherently secure.

Other services include secondary DNS, network time, route-hijack monitoring, and blackhole routing. Each addresses an auxiliary dependency that can stop research even when optical capacity is available. A name service failure, time error, route leak, or denial-of-service attack can make a science system unusable without damaging fibre.

The service list is a set of operational contracts between network layers. Physical connectivity defines the route, routing creates reachability, private services provide logical boundaries, cloud peering extends the system into commercial infrastructure, and security and support services reduce known failure modes. The user still must build an endpoint and a local network that work correctly at each end. (ESnet service list)

OSCARS makes capacity reservable

Most Internet traffic uses whatever capacity is available when packets arrive. That model works well for general communications, but some science transfers are scheduled, large, and costly when delayed. OSCARS gives authorised users and applications a way to reserve network resources during a defined window.

A reservation can specify endpoints, start and end times, bandwidth, VLAN information, path exclusions, and other constraints. The system checks the topology and existing commitments, finds an acceptable path, creates the required state in the network, and tears it down when the reservation ends. It must avoid allocating the same scarce resource twice and must account for maintenance or failures that change the available topology.

The significance is operational, not cosmetic. A manually configured circuit could need days of coordination between engineers, while a software reservation can be created in minutes and inserted into an experiment workflow. The network stops being a fixed background pipe and becomes a resource that software can request alongside storage or computing.

OSCARS is an open-source production system, and it has been adopted or tested by other networks and testbeds. Some public adoption figures appear on pages containing historical material, so the safest inference is that the system has spread beyond ESnet without assuming that every listed deployment is current and identical.

A reservation does not guarantee application throughput. It can protect a defined amount of network capacity, but it cannot make a storage server read faster, fix packet loss inside a site network, or tune the operating system. If an application achieves only a fraction of the reserved rate, diagnosis must return to the full path.

Reservations also raise allocation questions. A high-priority experiment can benefit from guaranteed capacity, while unused reservations reduce the flexibility available to others. Policy must define who may request the service, how conflicts are resolved, and whether a reservation is pre-emptible. These decisions remain part of facility governance even when the provisioning code is automated.

OSCARS remains one of ESnet's clearest examples of application software becoming infrastructure. It turned an operational practice into a repeatable service, kept human policy around authorisation, and gave science applications a controlled way to express network requirements. (OSCARS)

SENSE tests whether applications can request an end-to-end outcome

A reserved network path solves only one part of a distributed workflow. Storage systems, data transfer services, and computing resources have their own availability, capacity, and policies. SENSE, or Software-defined network for End-to-end Networked Science at the Exascale, extends coordination beyond a single network domain.

The project models networks and end systems as discoverable resources. A science application can describe a desired outcome, such as moving a dataset between two facilities at a specified rate. SENSE collects models from participating domains, negotiates resources, provisions the network and transfer services, monitors measurement, and releases resources when the workflow completes.

This task is harder than provisioning one circuit. Participating institutions retain their authority and their data models. One domain may offer network bandwidth while another offers storage endpoints or transfer services, and policies, credentials, and maintenance windows differ. SENSE must coordinate without assuming that a central controller can command every system.

During the 2024 Large Hadron Collider Data Challenge, SENSE-supported workflows sustained 330 Gbit/s between Caltech and the University of California, San Diego. The result shows that application-directed coordination can work at significant rate in a selected live path, but it does not prove that SENSE is a general production service at every ESnet site. The 2024 report continued to describe the project as in the research stage.

OSCARS and SENSE automate two different units of work. OSCARS reserves network resources inside a known service domain, while SENSE attempts to negotiate a full path that includes systems owned by different institutions. The former has clearer production boundaries; the latter offers broader scientific value with greater governance and integration risk.

The practical question for operators is failure ownership. If an application declares an intent and the workflow performs poorly, the cause may lie in the resource model, the network, the storage service, credentials, the application, or a timing problem between domains. Coordination needs evidence precise enough to tell each operator which part departed from the agreed state.

SENSE points towards an architecture in which applications can request outcomes rather than circuits. Its progress should be judged by repeated production use across multiple domains, clarity of support responsibility, and ability to recover from partial failure, not by the highest experimental transfer rate alone. (SENSE;2024 applied research)

EJFAT sends science events directly to remote computing

Traditional science data movement often follows a file-based sequence. The experiment produces data, local systems process enough of it to write files, files enter storage, and a later process moves them to another facility for analysis. This pattern is well understood and operationally reliable, but it delays feedback and requires local storage and computing that can handle the peak acquisition rate.

EJFAT, or ESnet–Jefferson Lab FPGA Accelerated Transport, supports a different path. The system distributes live, UDP-encapsulated event data from the instrument to available computing workers that may reside inside a remote supercomputing centre. A programmable FPGA-based SmartNIC reads event identifiers and sends all fragments belonging to the same event to a single worker. When computing nodes become available or busy, a control system can change the distribution.

Assembling event fragments matters. A generic load balancer can spread packets across servers, but scientific processing may require every fragment of one detector event to reach the same worker. EJFAT decouples the high-rate forwarding function from worker scheduling, letting each side scale independently.

The April 2024 demonstration between Jefferson Lab and Perlmutter cited in the opening achieved 100 Gbit/s without an intermediate disk stage. Later tests used computing resources distributed across several facilities and roughly 20,000 cores, according to ESnet reports. These results show the system's potential, but they do not mean that every instrument can adopt it without engineering modifications.

The Facility for Rare Isotope Beams (FRIB) provides a more recent operational example. ESnet reported in July 2026 that researchers streamed raw experiment data over EJFAT at about 5–6 Gbit/s to remote supercomputers. One run produced 615 GB over 15 hours, then processed it in 20 minutes on eight Perlmutter nodes using machine learning inference. The lower network speed does not make the result less important, because the scientific outcome was faster feedback during a live workflow.

Live streaming changes facility economics. A detector can borrow shared computing instead of building local capacity sufficient for every peak, and researchers can inspect results while scarce beam time remains available. But the system also creates new dependencies; a network outage, a computing allocation problem, or an event distribution error can affect the experiment moment by moment, rather than merely delaying a later file transfer.

EJFAT remains an advanced prototype or emerging platform, not a universally supported service. Moving to production requires instrument integration, security, FPGA availability, operational ownership, scheduling agreements, and clear behaviour when UDP packets are lost. Its significance lies in showing that a wide-area network can enter the acquisition loop instead of operating after it. (EJFAT;FRIB science report)

High-Touch, perfSONAR, and iperf3 make the path visible

A high-capacity network can fail in ways that aggregate usage graphs do not reveal. Micro-bursts can fill a queue, packets can arrive out of order, loss can affect one direction or one traffic class, the path can change between two tests, or security-relevant flows can disappear inside statistical sampling.

High-Touch uses programmable hardware and software to collect richer packet and flow measurement than traditional sample-based logs. It was used in 2024 to investigate Large Hadron Collider throughput, packet re-ordering linked to Rubin Observatory traffic, and security events. The purpose is diagnostic precision: identifying behaviour that aggregate traffic numbers hide.

Richer evidence has a cost. Packet and flow information requires storage, processing, access control, and careful interpretation, and it can reveal communicating endpoints, collaboration patterns, and the timing of scientific activity. A system designed for operational visibility therefore becomes a sensitive data asset.

perfSONAR addresses a different layer. It is a distributed measurement system supported by a multi-institutional partnership in which ESnet is a core founding entity. More than two thousand sites were reported globally in 2024, and ESnet operates more than 30 instances. Controlled tests measure throughput, latency, loss, and path behaviour between participating nodes.

The distributed design is essential. A researcher can report a slow transfer, but a test between properly engineered perfSONAR nodes shows whether the underlying path supports the expected rate. Directional tests can isolate asymmetry, repeated measurements can show when degradation began, and traceroute data can identify a path change, though the observed route and the actual application forwarding behaviour may differ.

iperf3 provides active throughput testing over TCP, UDP, or SCTP, supporting IPv4 and IPv6. ESnet maintains the open-source tool and reports tests exceeding 150 Gbit/s on 200G paths under specific conditions. A synthetic test is not an application benchmark, but it helps determine whether the host and network can move traffic under controlled conditions, after which storage and application behaviour can be examined separately.

High-Touch, perfSONAR, and iperf3 together offer different views of the truth. High-Touch watches production flows with greater precision, perfSONAR provides scheduled or on-demand path measurement across domains, and iperf3 tests endpoint and network throughput directly. No single view gives the whole picture, but combining them reduces the temptation to diagnose a distributed fault from a single operator's local dashboard. (Network performance tools;2024 operational innovations)

Caching can remove traffic instead of moving it faster

Adding capacity is not the only way to improve a science network. Some datasets are requested repeatedly by researchers in the same region, and sending every copy over the wide area wastes bandwidth and adds latency even when the backbone can accommodate the traffic.

ESnet's 2024 applied research reports described five regional caching nodes used with science data communities. The nodes reduced wide-area traffic by an average of 33 TB per day over the study. Two-thirds of requests were served locally, and the average hit rate was 94%. The impact varied sharply by location, however, with reported reductions of 69.3% in Southern California, 48.4% in Chicago, and 6.6% in Boston.

These differences reveal the nature of value. Caching benefit depends on dataset popularity, local users, storage size, eviction policy, and service configuration. A node placed near a community that repeatedly analyses the same data can remove large transfers, while a node serving diverse or rarely reused data may consume storage without delivering the same benefit.

Caching also changes the control question. A network operator begins to participate in data placement, while science communities must decide which datasets are cacheable, how freshness is maintained, and who may access them. Storage failures and stale content become operational problems adjacent to the network.

The study illustrates why the most efficient bit can be the one that never enters the backbone. Optical upgrades add supply, while caching alters demand. Workflow scheduling can move jobs away from congested periods, and local compression or filtering can reduce data before transfer. A mature infrastructure programme evaluates all of these mechanisms rather than measuring progress only by installed terabits.

The result must remain bounded to the studied deployment. It does not prove that every science dataset will achieve a 94% hit rate or that caching can replace new capacity for unique live streams. It does offer evidence that data architecture and network architecture can be designed together. (2024 applied research activities)

Transatlantic science needs physical diversity alongside capacity

ESnet's domestic footprint cannot be separated from its international commitments. High-energy physics is organised around facilities and computing centres on both sides of the Atlantic, and the 2024 annual report attributed roughly half of ESnet traffic to Large Hadron Collider activity. A cable fault between Europe and North America can therefore affect a central science workload even when every domestic router is healthy.

During 2024 and early 2025, ESnet expanded its reported transatlantic capacity from approximately 700 Gbit/s to 2.7 Tbit/s. A key element was a 15-year agreement with Aqua Comms for a quarter fibre pair on a path connecting New York with Dublin and London. ESnet also shares capacity and costs with GÉANT across several cable systems. The engineering goal is to reach at least 3.2 Tbit/s across four physically independent paths, with a design scalability to more than 10 Tbit/s.

Capacity and diversity solve different problems. Two logical services can share the same cable, landing station, or terrestrial backhaul and therefore fail together. A submarine cable repair may require fault location, permits, a cable ship, suitable weather, and physical retrieval of damaged fibre, and can take weeks or months. ESnet therefore plans for scenarios that include three simultaneous cable cuts, rather than assuming that one backup path is sufficient.

Long-term spectrum rights give the facility more upgrade control than repeatedly buying fixed complete circuits. They also create commitments to specific cable systems, landing infrastructure, and operating partners. ESnet does not remove marine risk by leasing spectrum; it gains the ability to engineer capacity and resilience across multiple systems while still depending on cable owners, repair procedures, and partner networks.

This is another place where the word 'dedicated' needs precision. ESnet controls dedicated rights and infrastructure on important paths, but the physical asset model differs. Domestic and international routes may involve dark fibre, spectrum, lit services, shared co-location, or partner capacity. A line on a map does not disclose the same level of ownership or repair authority at every segment.

The GÉANT relationship shows why research networks cooperate without being under central control. The two organisations can share cost and coordinate science traffic, but each remains accountable to its own institutions and members. Full performance still depends on national research networks, university systems, and experiment facilities that lie outside any single backbone. (Transatlantic capacity milestone;2024 annual report)

Scientific demand arrives as steady load, rare bursts, and hard deadlines

The Large Hadron Collider provides the steady load. Its distributed computing model moves experiment data between CERN, Fermilab, Brookhaven, universities, and other centres, creating sustained international traffic. The planned High-Luminosity LHC is expected to increase detector, simulation, and replication output. The network must carry massive flows under normal conditions while retaining enough path diversity for cable or facility failures.

The Vera C. Rubin Observatory provides the hard deadline. The telescope in Chile is designed to produce an image of roughly 13 GB every 30 seconds. A rapid processing data path at SLAC supports transient alerts so that other observatories can respond. ESnet reports a goal of moving each image across roughly 12,000 miles in less than seven seconds. The route involves South American research network partners and international collaborators, so the outcome depends on coordinated engineering beyond ESnet alone.

DUNE provides the exceptional burst. During a supernova event, the neutrino programme expects a need to move up to 600 TB within 100 seconds. The wider programme is forecast to produce roughly 900 PB over twenty years. These are planning requirements for future operations, not a description of current daily traffic. Their importance is that the event cannot be scheduled after capacity becomes available; a network designed around average usage can fail at the moment of highest scientific value.

Fusion research combines international distance with a long development horizon. In May 2026, a test moved 176 TB of ITER data from Marseille to the DIII-D facility of General Atomics in San Diego at a rate approaching 80 Gbit/s. ITER was still under construction, so the transfer proved readiness rather than steady full operation. Requirements materials forecast that some future fusion workflows will reach roughly two petabytes per day and at least 200 Gbit/s.

FRIB and Jefferson Lab show a third pattern: an experiment may need remote analysis while data are still being acquired. EJFAT lets event streams reach computing at NERSC or Oak Ridge instead of waiting for files to close and copy. Light sources, electron microscopes, and other instruments face similar choices about local filtering, remote AI inference, and how quickly results return to the operator.

Climate and earth science stretch the shape of demand again. Large simulations and observational datasets move between computing and storage facilities, while field sensors may sit in places where traditional fibre does not exist. No single headline rate can describe all of these needs. ESnet must support sustained bulk movement, rare bursts, low-latency feedback, international collaboration, and remote acquisition within one facility model. (Rubin Observatory case study;DUNE case study;Requirements reviews)

Requirements reviews turn science plans into network infrastructure

ESnet cannot wait for an experiment to produce its first dataset before ordering fibre, routers, or transatlantic capacity. Long-distance agreements, equipment purchases, site builds, and software development need years. The facility therefore conducts requirements reviews with DOE science programmes over a five- to ten-year horizon.

Researchers and facility teams describe instruments, expected data volumes, storage locations, computing destinations, timing constraints, international collaborators, cloud use, and resilience needs. Network and computing specialists then examine the full workflow. A request for a bigger link may reveal a more important problem in local storage, path diversity, security architecture, or data transfer software.

The review process is a form of co-design. Science programmes explain what they are trying to accomplish, and ESnet translates the goal into infrastructure dependencies and tests whether the proposed path is workable. The result can influence site connection sizing, new routes, submarine procurement, Science DMZ deployment, coordination research, staffing, or the schedule of the next network generation.

The major high-energy physics review completed in 2025 and published in January 2026 involved 127 entities and 14 case studies in a report of roughly 400 pages. Its value does not depend on every forecast proving correct; instruments may be delayed, software may become more efficient, and new workloads may appear. The value is a documented record in which researchers, network engineers, computing centres, and funders can see the same assumptions and adjust them consciously.

Forecasting remains difficult because scientific demand is lumpy. A supernova may not occur during an experiment's lifetime, while an AI workload may expand faster than a formal review cycle. An average traffic growth figure helps with budget preparation but cannot replace scenario planning. The unusually low 4% traffic increase in 2024 does not prove that long-term demand has stopped; it illustrates why a single year should not define a decadal infrastructure programme.

Requirements reviews also distribute responsibility. A laboratory cannot assume that the national backbone will fix an under-provisioned local path, and ESnet cannot assume that a science programme will adapt after capacity is built. Putting the dependencies inside a joint plan makes later disagreements more visible. (Requirements review reports)

Wireless and quantum projects test the edges of ESnet's mission

A national fibre network reaches fixed institutions well, but science instruments do not always operate in such places. Geothermal fields, environmental observatories, and temporary campaigns may lack operator coverage, reliable power, or a practical fibre route. ESnet's Wireless Edge programme is investigating whether private cellular, Wi-Fi, directional radio, and satellite can extend the science workflow to those locations.

A deployment at a geothermal site in Nevada, announced in June 2026, combined a private 4G network using Citizens Broadband Radio Service spectrum, long-range Wi-Fi HaLow, conventional Wi-Fi, a Starlink link, a directional radio, and a self-powered portable tower. The design responded to the conditions of one field site; it was not a general wireless access product from ESnet, nor can it deliver the sustained capacity of long-haul fibre. Terrain, weather, spectrum, power, and the satellite provider remain parts of the service path.

The organisational point is that ESnet did not stop at the nearest fibre point and treat the rest as someone else's responsibility. The science-engagement function treated acquisition, backhaul, and wide-area transport as one workflow. That approach may be more valuable than a standardised product because remote science sites face different physical constraints.

QUANT-NET explores a different edge. The ASCR-funded project is building a three-node quantum network testbed linking Berkeley Lab and two University of California, Berkeley, sites over roughly five kilometres. The work involves ion traps, photonic interconnects, entanglement swapping, Bell measurements, time-sensitive control, and a standardised software framework.

The connection to ESnet is primarily about control and coordination. Quantum experiments need classical communication, precise timing, and orchestration among devices that remain hard to operate. Open-source control software can make experiments more repeatable and reduce manual tuning. The testbed is not, however, a production quantum Internet, nor does it move conventional high-volume science data through entanglement.

Wireless Edge and QUANT-NET should therefore be measured by what they learn and what becomes reusable. Neither is evidence that every ESnet user will receive a wireless access service or a quantum link. They extend the facility's research function into domains where future science networks may need new physical media and control systems. (Wireless Edge deployment;Quantum networking research)

American Science Cloud and ESnet7 shift planning from sites to workflows

The Integrated Research Infrastructure programme and the emerging American Science Cloud project reflect a change in how the DOE describes its facilities. Instead of treating the instrument, network, data store, and supercomputer as independent services, the programmes seek to make them work as a federated environment for science data, AI, and high-performance computing.

ESnet's role is in the connectivity and coordination layer. Data must move from instruments to storage and accelerators, users and services need controlled access, workflows must discover available resources, and measurement must show whether the system is achieving the scientific goal. EJFAT, SENSE, OSCARS, Cloud Connect, and High-Touch each address part of the problem, but none of them alone constitutes the American Science Cloud.

At Confab26, Inder Monga was identified as the American Science Cloud deputy project lead, and programme sessions included demonstrations and next-generation network planning. That is evidence of active institutional work, not evidence of a completed, generally available science cloud in August 2026. Access policy, the inventory of connected resources, production support, and governance remained matters under development.

ESnet7 appears inside the same planning environment. The name is real, and innovation sessions studied measurement, deep packet inspection, data movement, AI, and quantum networking. ESnet6 remains the current production generation, however. The available evidence does not yet contain a final public architecture, project baseline, vendor selection, full budget, or launch date.

Network generation names create expectations long before equipment is installed. A planning programme can shape research priorities, vendor engagement, and capital requests, but it should not be presented as a completed backbone. The most accurate interpretation is that ESnet is using the operational experience of ESnet6 and current workflow projects to define what the seventh generation should look like.

Speed will remain important, but the evidence suggests that programmability and intelligence may become equally so. The fastest link cannot tell an application where computing is available, identify a microburst, schedule a detector stream, or prove that a service has been correctly released across multiple domains. The strategic question for ESnet7 is therefore how much of the science workflow the network should understand and coordinate without becoming an unmanageable central controller. (Confab26 programme;2024 applied research)

Federal funding supports the facility but the budget line is not revenue

ESnet does not fund its backbone by selling bandwidth to the public. It is funded primarily by DOE through the ASCR activity called High Performance Network Facilities and Testbeds. The FY 2027 request proposed $103 million for this line, compared with $97.261 million in the prior appropriation column.

These figures provide the best public indicator of programme level, but they are not an income statement for ESnet. The line also supports networking-related testbeds, software maintenance, upgrades, and research activity. It does not disclose a separate annual operating cost for the production backbone, capital expenditure per network generation, or revenue associated with individual site connections.

Public funding allows the facility to invest before demand becomes commercial. It can purchase long-term spectrum rights, maintain open-source tools, support low-volume, high-scientific-value experiments, and build path diversity beyond what immediate use would justify. The same model creates a dependency on congressional appropriations, department priorities, and the Berkeley Lab operating framework.

Site connections are not always free. The Site User Cost policy allows institutional users to bear connection-related costs according to sponsorship, eligibility, and incremental requirements. The unit of relationship is the site, not the individual researcher, and endpoints are not separately billed or logged through ESnet.

The absence of standalone accounts limits financial analysis. Public materials do not reveal a full concentration of suppliers, salaries, a depreciation schedule, an asset registry by path, or an annual capital budget. Describing $103 million as 'ESnet revenue' would therefore be inaccurate. It is more precise to say that a federal programme funds a production facility and its upgrade work, software maintenance, and testbeds through a broader activity line.

The funding structure also affects strategic choice. A commercial network can withdraw from a route that does not cover its cost, while a public science facility must balance national mission, scientific opportunity, and long-term capacity against budget constraints. This can justify an investment that delivers no immediate financial return, but it makes priority clarity and requirements evidence essential. (FY 2027 ASCR budget request;ESCC and site policy)

Reliability and visibility create obligations beyond uptime

ESnet's 2024 reports recorded a notable availability result: the ten measured Office of Science sites achieved 100% service availability after planned maintenance was excluded. That was the first year in which every site simultaneously exceeded the 99.9% requirement. The annual report also gave a broader uptime figure of 99.99%. These are useful but different-scope metrics, and they do not prove that every service, institution, and international path suffered no outage.

The Site Resilience Program examines diverse entries, routers, power, and failure domains inside connected facilities. This reflects an uncomfortable truth: a national backbone can be multi-path while a laboratory remains dependent on a single local conduit or a single device. Resilience must reach the building and the data transfer node, not stop at the backbone map.

Routing policy provides another control layer. ESnet rejects RPKI-invalid announcements at peering, requires current routing registry records, and supports route-hijack monitoring and destination blackhole routing. These controls reduce known risks but do not make inter-domain routing foolproof. A route with a valid origin can still be associated with a leak or a policy error, and blackhole protection works by making the targeted destination unreachable to others.

Automation creates a similar trade-off. Standardised configurations, inventory, and orchestration reduce manual error and speed recovery, but a wrong template, policy, or data record can propagate quickly across many devices. The appropriate response is phased deployment, independent verification, clear rollback or roll-forward procedures, and limits on the authority of any single automation account.

Operational data deserve comparable attention. The facility data policy states that router usage data and NetFlow may be retained indefinitely and replicated between East and West Coast stores. Active perfSONAR measurements are kept for six months on a single disk without backup. Router usage data and traceroute are public, while flow and security logs are subject to restricted access.

ESnet may have no individual subscriber relationship with a user and yet collect detailed operational data; the two regimes are different. There is no individual subscription and billing link, but network metadata can identify endpoints, communicating institutions, timing, and traffic patterns. Higher-precision measurement improves diagnosis and security, but it increases the need for data minimisation, access control, retention review, and responsible use.

A credible resilience model rests on scope-limited mechanisms: physical diversity, site engineering, routing security, measurement, incident response, data governance, and clearly defined metrics. (Availability milestone;Facility data policy;Peering policy)

The real test is whether distributed science works in practice

ESnet's history can be read as a sequence of faster networks: the predecessor specialist systems, ESnet4, the 100G continental ESnet5, and the multi-terabit ESnet6. That timeline is real, but it misses the more important continuity. The facility repeatedly changed the architecture around scientific work when more bandwidth alone could not solve the problem.

The Science Data Network separated exceptional flows from ordinary traffic, OSCARS made capacity reservable, Science DMZ redesigned the institutional perimeter, perfSONAR and iperf3 turned performance disagreements into measurements. SENSE linked application intent across multiple administrative domains, EJFAT brought the wide-area path into the detector loop, caching altered demand, and transatlantic spectrum changed ESnet's level of control over international growth.

The organisation's authority remains bounded. The DOE and ASCR set the mission and funding, Berkeley Lab operates the facility, and the University of California manages the lab under contract. Site coordinators approve changes that affect their institutions, international partners control their own networks, and cloud providers, operators, cable owners, and equipment suppliers hold authority over components that ESnet depends on.

Distributed control is not a problem that software can erase; it is the condition under which a science network operates. The practical achievement is coordination that is robust enough to deliver the workflow without claiming that a single institution owns the whole path.

The same discipline should apply to future claims. ESnet7 remains in planning, SENSE was in the research stage as of the latest full report, EJFAT is an advanced prototype or emerging platform, QUANT-NET is experimental, and the American Science Cloud was still under construction. The evidentiary value of each lies in describing its maturity stage accurately.

ESnet's long‑term significance lies in the user-facility model. The network is treated as a scientific instrument with requirements, operations, software, measurement, research, and public funding. That model recognises that a detector and a supercomputer separated by thousands of kilometres deliver limited joint value unless the systems between them are engineered as a single path.

No recent neutral metric in the available evidence can rank ESnet as the world's fastest network, and that is not the most useful test. The stronger test is whether an instrument can reach remote storage or computing at the required rate, recover when part of the path fails, and produce evidence that shows where responsibility sits. ESnet's standing will be confirmed or weakened by these production outcomes as the American Science Cloud, live streaming, and the future ESnet7 programme move from demonstrations and plans to supported science services. (ESnet history;ESnet6 launch;2024 annual report)

Sources