Summary
- A SoftBank NCCL all-link performance check found one link at an unspecified fraction of normal throughput even though its LED was green, administrative state was up and optical power appeared normal. Those signals were valid but limited public evidence for workload acceptance.
- The JANOG58 deck names switch, optics, NIC, fibre and connector contamination as fault-isolation surfaces, but publishes no root cause, incident BER/FEC values, exact throughput ratio or measured repair. Its BER/FEC numbers and optics-replacement diagrams are examples.
- Rack-scale acceptance should bind nominal state to a reproducible per-link workload baseline and Layer 1 evidence, then replay the same test after repair. External routes, ASN visibility and installed GPU counts cannot certify this internal usable capacity.
Four observations on one slide overturn a familiar acceptance shortcut. SoftBank says an NCCL benchmark measured every link in its rack-scale GPU environment. One link delivered throughput at only a fraction of the normal rate. Yet its LED remained green, Link Status remained UP and optical power remained inside the normal range.
The deck does not say one-third. It publishes no exact ratio. That omission should remain visible, because an invented precision would distract from the more important result: the controls that ordinarily establish presence and light did not establish the rate required by the collective workload.
Yasuhiro Uchida and Chaocheng Chang presented the case for SoftBank at JANOG58 on 16 July 2026. The official Day 2 session, Rack-Scale GPUサーバーのNW設計と運用までの苦悩, covers the network and facility changes required by SoftBank's first rack-scale GPU deployment. Its 56-page final deck moves from physical architecture and three network fabrics through service design, cabling, optics and liquid-cooling problems.
The separate nine-page preread is not a shorter incident report. The official link uses a j56-lt4.pdf filename, and the content is earlier material about liquid-cooled switches, proprietary cooling interfaces and standardisation. It is useful background on facility dependence. It cannot fill the missing incident values.
The final deck's architecture explains why a single weak link matters. One GB200 NVL72 rack is described as 18 compute trays, with four B200 GPUs per tray, and nine switch trays, each carrying two NVSwitch devices. SoftBank says the rack exceeds 100 kW. NVIDIA's reference documentation independently matches the 72-GPU, 36-CPU, 18-compute-tray and nine-switch-tray shape, with a passive cable backplane, power shelves, busbar and liquid-cooling manifolds. Those are architecture statements, not performance results for the affected link.
SoftBank separates the external network into three fabrics. Compute Fabric carries GPU scale-out traffic. Converged Fabric carries front-end, storage and NCCL bootstrap traffic. OOB Fabric carries switch and server management, facility monitoring and NVSwitch management. The division is logical and operational; it does not make the physical dependencies independent.
The OOB section is unusually explicit. Management-port communication to the leaf switches can be a single point of failure. The deck explores using loopback reachability through the Compute Fabric underlay to improve operational access at far lower cabling cost than connecting every low-speed bonus port. It also says this is not a replacement for OOB connectivity. A path that helps repair the production fabric cannot be allowed to become the only path from which that fabric is repaired.
The degraded-link episode exposes a related boundary. A green LED proves that one status circuit sees the condition it was designed to see. Link Status UP proves that the interface reached the relevant administrative and protocol state. Optical power in range proves that received power did not cross the chosen alarm boundary. None of those statements promises a clean enough signal, lane behaviour, error-correction margin or end-to-end rate for the workload.
They are not useless indicators. The public evidence does not support replacing them with one benchmark. It supports adding an acceptance control that asks the question they did not answer.
SoftBank's all-link NCCL check did that. A collective workload depends on the slow or degraded part of the communication graph rather than on the average port label. A link that passes presence checks but delivers a fraction of the expected rate can strand far more expensive compute behind it. The installed accelerators remain in inventory and powered; the capacity available to a tenant is lower.
The deck then lists the fault-isolation surfaces: switch, optics, NIC, fibre and a dirty connector. This is a candidate chain, not a published root cause. A link can experience errors at the transmitter, receiver, lane, module, connector or cable, while the symptom appears as lower delivered throughput. The operator has to isolate the dependency across ownership and support boundaries.
Pages that follow explain pre-FEC and post-FEC BER and FEC histograms. They show illustrative normal and abnormal patterns and optics replacement inside example diagrams. These examples must not be converted into the incident record. The deck does not say that its actual link carried the displayed BER values, reached the displayed FEC bins or recovered after a module swap. It publishes neither before/after throughput nor a component-specific conclusion.
That absence matters because diagnosis and repair are different evidence stages. BER can show errors before correction; FEC can show how much correction margin is being consumed. A component swap can test a hypothesis. Only a repeat of the original performance test can show that the usable rate returned. Without the replay, the module was replaced and the link was restored are different statements.
SoftBank's proposed operating change is nevertheless clear. The deck labels Link UP = OK limited public evidence. It calls for physical-layer understanding, greater attention to BER and FEC, collection and analysis of interface and optics logs on weekly and monthly horizons, and a move from monitoring only interface-down events toward looking for degradation signs.
This is a monitoring design, not a published predictive result. No source in the packet gives alert thresholds, retention, false positives, a degradation caught before failure or an availability gain. The strongest supported conclusion is that the operator identified a blind spot and named a closer set of physical indicators.
Acceptance also changes with the unit sold. The deck compares rack-, tray- and GPU-level service. Its summary says SoftBank currently offers rack and tray units after weighing the trade-offs. GPU-unit service remains a considered design, not the declared current offer.
Rack service is simpler to partition because one customer can receive the whole NVL72. It also exposes the customer to rack-level dependencies including Compute Fabric, CDU, NVSwitch, compute tray and busbar. Tray service supports a smaller customer but requires NVLink partitioning. The deck says the NVSwitch management process is common inside the rack and treats it as a single point of failure; higher service levels require rack- or scalable-unit-level redundancy.
These design choices distribute power. SoftBank controls acceptance, workload admission, monitoring and the customer restoration decision for the platform it operates. NVIDIA and the switch, NIC and optics vendors control reference designs, firmware, privileged diagnostics, qualification lists and support remedies. Facility teams control power, liquid cooling and physical access. JANOG controls the session record. It does not certify the link or command any production network.
The same distribution identifies the payer. SoftBank funds dense cabling, high-speed optics, switches, spare parts, expensive labs, rack power, cooling and the labour required to cross vendor boundaries. A tenant pays in delayed jobs and idle accelerator time when nominally installed capacity fails its workload. A smaller buyer may find the spare rack, lab and specialist evidence disproportionately costly, increasing dependence on the integrated supplier.
SoftBank's December 2025 platform notice says 1,224 Blackwell GPUs began operation and that expansion beyond 4,000 is planned. The first is an operator deployment statement; the second is a plan. Neither number says how many links meet an acceptance baseline at a given moment. Installed GPU count and saleable compute capacity are related, not interchangeable.
Nor should a separate public benchmark be imported. TOP500 identifies SoftBank's CHIE-4 as a DGX B200 system using InfiniBand NDR400. Its rank belongs to that system, not to the GB200 NVL72 platform discussed at JANOG58.
A stronger acceptance record would start with topology identity. For every required link, retain the expected NCCL distribution or other portable workload baseline alongside link state, optical power, lane counters, pre/post-FEC BER and the FEC histogram. Timestamp the firmware, module, fibre and switch endpoints. If a link deviates, swap one bounded component at a time, retain the counter change and replay the exact all-link test after repair.
The evidence should survive organisational handoff. A tenant need not receive every vendor counter, but the service owner should be able to prove that the accepted link met a declared rate and that the repair returned it to the same distribution. Monitoring should distinguish a trend, an alert, a ticket, a component change and a verified restoration.
External reachability remains a separate layer. A rack can have reachable management and customer prefixes while its internal collective path under-delivers. ASN and BGP visibility help prove routing identity and access to the service edge. They do not grant a registry, standards body or conference authority to declare the GPU fabric fit, and they do not certify throughput behind the prefix.
The JANOG58 lesson is therefore narrower than BER is better and more consequential than watch more counters. A rack-scale link becomes capacity only when its required rate is demonstrated, its physical evidence can be followed, and the same workload check closes the repair. Green and up described the link's presence. The all-link test revealed whether it was usable.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance

