In brief
Training a large language model is not simply a matter of stacking GPUs. These thousands of chips must constantly exchange billions of parameters — and the speed at which they do so directly determines the overall system performance. High-performance networking is the invisible connective tissue that transforms a warehouse of processors into a coherent machine. Without it, raw compute power remains largely untapped.
The GPU paradox
An A100 or H100 GPU posts impressive numbers on its spec sheet. Those figures fill announcements, feed comparisons, and structure purchasing decisions. Yet at the scale of a training cluster, an empirical rule holds: half of performance gains come not from computation itself, but from the ability to move data between chips.
The reason is mechanical. Distributed training splits a model and its data across hundreds or thousands of GPUs. At each step, locally computed gradients must be aggregated, redistributed, and synchronized. If this transfer is slow, GPUs wait. A processor rated at 1000 TFLOP/s that waits 40% of the time effectively delivers only 600 TFLOP/s. The network is the silent bottleneck.
Three interconnect layers
A datacenter’s network infrastructure for AI is organized into three distinct levels, each with its own technology, bandwidth, and constraints.
First layer: the intra-node highway
Inside a server, GPUs communicate via NVLink. Version 4.0, introduced with Hopper, reaches 900 GB/s total bandwidth between eight GPUs. Version 5.0, with the Blackwell architecture, doubles that to 1,800 GB/s.
The highway analogy holds: NVLink is a dedicated, high-speed lane that does not share its capacity with other traffic. The NVSwitch enables an all-to-all topology — each GPU can talk directly to each of the other seven without routing through the CPU. With the NVL72 in the Blackwell architecture, 72 GPUs share the same NVSwitch fabric for 1.8 TB/s total bandwidth. This is no longer a server — it is an architecture in its own right.
Second layer: the inter-node highway
Between servers, the dominant technology is InfiniBand. Successive generations — HDR at 200 Gb/s, NDR at 400 Gb/s, XDR at 800 Gb/s — have multiplied throughput fourfold in just a few years. The key characteristic of InfiniBand is not only speed: it is sub-microsecond latency, ten to one hundred times lower than standard Ethernet.
For 10,000 GPUs connected via an InfiniBand NDR fat-tree topology, the cost in switches and cables runs into tens of millions of euros. That represents 15 to 30% of total infrastructure budget for a cluster of that size. The network is not an accessory — it is a major budget line item.
Third layer: direct memory access
RDMA — Remote Direct Memory Access — is the mechanism that allows a GPU to read or write into the memory of another GPU on a remote node without going through the CPU. InfiniBand supports RDMA natively. RoCE (RDMA over Converged Ethernet) attempts to replicate this behavior on Ethernet infrastructure with additional configuration requirements to guarantee lossless packet delivery.
In short: three plumbing layers — NVLink (1,800 GB/s between 8 GPUs in the same enclosure), InfiniBand (800 Gb/s between enclosures), RDMA (remote memory read without waking the CPU). The further the data travels, the narrower the pipe. The whole art is to minimize byte travel and maximize compute-communication overlap.
What the data actually does
During training, several collective operations continuously flow across the network. All-reduce aggregates gradients from all GPUs and redistributes the sum. All-gather collects tensor fragments distributed across GPUs. Reduce-scatter combines both in a single optimized pass.
These operations are orchestrated by NCCL (NVIDIA Collective Communications Library), which detects the available topology and selects the most efficient algorithm. The library is designed to simultaneously exploit NVLink intra-node and InfiniBand inter-node, overlapping both flows.
What the real numbers show
Published results allow the stakes to be quantified. The 2021 Megatron-LM paper (Narayanan et al.) reports 502 petaflops measured across 3,072 GPUs — 52% of theoretical efficiency. This figure explicitly incorporates tensor parallelism intra-node via NVLink and pipeline parallelism inter-nodes via InfiniBand: the two network layers work together.
In 2024, the MegaScale paper (Jiang et al.) describes a cluster of 12,288 GPUs achieving 55.2% MFU (Model FLOP Utilization). The 1.34x gain over Megatron derives largely from network optimizations: better straggler management, compute-communication overlap, refined topology.
ZeRO++ (Wang et al., 2023) reduces communication volume by 4x by quantizing weights and restructuring exchanges, translating to 2.16x higher throughput on a 384-GPU cluster. FlashOverlap (Hong et al., 2025) goes further by explicitly overlapping compute and communication phases to achieve a 1.65x speedup.
These gains do not come from new chips. They come from better use of the existing network.
| Type of optimization | Typical gain | Lever |
|---|---|---|
| Megatron-LM (tensor + pipeline parallelism) | 52% MFU on 3,072 GPUs | Combines intra-node NVLink and inter-node InfiniBand. |
| MegaScale (straggler management + topology) | +34% over Megatron on 12,288 GPUs | Adaptive scheduling, ZeroMQ for events. |
| ZeRO++ (gradient compression) | × 2.16 on 384 GPUs | Quantize the exchanged weights, restructure the communications. |
| FlashOverlap (compute/comm overlap) | × 1.65 | Overlap kernel execution with NCCL transfers. |
In short: on a 10,000-GPU cluster, doubling performance does not require replacing chips — it requires optimizing the network. That is why hyperscalers invest so much in NCCL, RDMA, and soon UEC.
The straggler problem
An often underestimated point: in a cluster of thousands of GPUs, a single node responding more slowly — degraded cable, congested switch, misaligned process — is enough to block the entire synchronization. This is the straggler problem. Strategies to mitigate it (active detection, path redundancy, adaptive scheduling) themselves consume network time. The network engineering of a large cluster is as much a discipline of reliability as of performance.
The UEC vs. InfiniBand battle
Since 2023, an industry alliance has been attempting to dethrone InfiniBand in this market. The Ultra Ethernet Consortium brings together AMD, Intel, Microsoft, Meta, and Broadcom around an open standard for high-performance networking on Ethernet infrastructure.
The stakes are considerable. InfiniBand is supplied almost exclusively by NVIDIA since its acquisition of Mellanox in 2020. Large cloud operators and hyperscalers — who pay the tens of millions mentioned above — have every reason to want an alternative. UEC’s technical argument rests on SMaRTT (Segment-based Multi-pathing and Rapid Transport for Telemetry), a multipath routing mechanism that reduces tail latency. Preliminary results (Bonato et al., 2024) show 50% improvement over classic RoCE.
The battle is not only technical. It is industrial and geopolitical. An open standard for AI networking reduces dependence on a single vendor and opens the market to broader competition. For actors seeking to build sovereign infrastructure — in Europe and elsewhere — this is a structural issue. An AI cluster whose network depends on a single American vendor is vulnerable to export restrictions, pricing changes, or unilateral roadmap decisions.
Gradient compression is another active area of debate. Reducing the precision of exchanged data (quantization, sparsification) decreases communication volume but can introduce noise into gradients and degrade convergence. The balance between network efficiency and training quality has not yet been settled by consensus.
In short: the UEC vs InfiniBand battle is not only about protocols. It is also about sovereignty — an AI cluster whose network depends on a single American vendor is exposed to roadmap decisions, export restrictions, and unilateral pricing adjustments. For Europe, the issue is structural.
Key takeaways
- An underutilized GPU cluster is often a network problem, not a compute problem: GPUs are waiting for their data.
- Three interconnect levels coexist — NVLink intra-node (900–1,800 GB/s), InfiniBand inter-node (up to 800 Gb/s), RDMA for CPU-bypass memory access.
- Collective operations (All-reduce, All-gather) are the dominant traffic on the network during training; NCCL orchestrates them.
- Published network optimizations (ZeRO++, FlashOverlap, MegaScale) deliver gains of 1.3x to 2.16x without changing compute hardware.
- The UEC vs. InfiniBand battle goes beyond the technical: it is a sovereignty issue for anyone seeking to break free from dependence on a single AI infrastructure vendor.