TL;DR — NVIDIA NVLink provides the high-bandwidth, low-latency interconnect that makes multi‑GPU clusters viable for AI training and HPC. By combining NVLink with NVSwitch, engineers can build topologies that scale from 2‑GPU pairs to thousands of GPUs while keeping effective bandwidth high. This post walks through the architecture, design patterns, and production considerations you need to know to plan a modern GPU cluster.

In the last five years, the demand for compute has shifted from single‑core CPU performance to massive parallelism across GPUs. Models like GPT‑4, Stable Diffusion, and scientific simulations require not just raw FLOPs, but a fabric that can move tensors between devices faster than the compute units can consume them. NVIDIA NVLink is the technology that makes this possible, and understanding its design is essential for anyone building or operating GPU clusters.

NVLink is a point‑to‑point, high‑speed serial interconnect that directly connects NVIDIA GPUs without going through the PCIe bus. Each NVLink lane operates at 25 Gbps (Gen 3) or 50 Gbps (Gen 4), and a single link can aggregate multiple lanes. In the latest H100 GPUs, a single NVLink link provides up to 900 GB/s of bidirectional bandwidth—roughly 12 times the bandwidth of PCIe 5.0.

The key architectural benefits are:

  • Direct GPU‑to‑GPU communication – bypasses the CPU and host memory, reducing latency and freeing PCIe lanes for other I/O.
  • Cache coherence – NVLink maintains a coherent memory view across GPUs, enabling unified memory models that simplify programming.
  • Scalability – multiple links can be combined into a “NVLink bridge” or “NVLink fabric” to connect more than two GPUs.

NVSwitch and Topology Options

NVSwitch is a dedicated silicon that acts as a crossbar for NVLink. A single NVSwitch chip can connect up to 18 GPUs in a non‑blocking, full‑mesh topology. By cascading multiple NVSwitches, you can build larger topologies.

Common topology patterns include:

  • Dual‑GPU pair – two GPUs connected by a single NVLink bridge; simplest form for small workloads.
  • NVSwitch‑based “full clique” – every GPU in the switch is directly connected to every other GPU; ideal for communication‑heavy all‑reduce operations.
  • Fat‑tree / Clos network – multiple NVSwitches arranged in a tree structure to scale beyond 18 GPUs while preserving bisection bandwidth.
  • Hybrid PCIe‑NVLink – some GPUs communicate via NVLink, others via PCIe; useful when mixing node types or when NVLink ports are exhausted.

Topology Patterns: Tree, Fat Tree, Clique

When designing a cluster, the choice of topology directly impacts performance, cost, and complexity.

Full Clique (NVSwitch)
A full clique offers the lowest latency and highest bisection bandwidth because every GPU can talk to every other GPU at full link speed. The trade‑off is that the number of GPUs is limited by the number of NVLink ports on the switch (typically 18). This pattern is perfect for training large models where all‑reduce is a bottleneck, such as GPT‑style transformer training.

Fat Tree
To scale beyond 18 GPUs, you can build a fat tree using multiple NVSwitches. In a fat tree, the aggregate bandwidth between “spine” switches and “leaf” switches is provisioned to avoid oversubscription. For example, a two‑tier fat tree with 9 spine switches and 18 leaf switches can support up to 162 GPUs while maintaining a bisection bandwidth of 900 GB/s per GPU. The design requires careful port planning and cable management.

Tree (or “Star”)
A simple tree topology connects GPUs in a hierarchical fashion, with a central switch connecting to leaf switches. This is easier to cabling but can suffer from oversubscription at the root, leading to bandwidth contention. It’s often used for inference workloads where traffic is less bursty.

Bandwidth and Latency Considerations

NVLink’s latency is typically sub‑microsecond (≈ 0.5 µs) for a single hop, which is critical for collective operations like all‑reduce. The effective bandwidth you see in practice depends on:

  • Topology – a full clique avoids the “hop” penalty that tree topologies introduce.
  • Message size – small messages (< 4 KB) may be limited by the overhead of NVLink flow control; large messages (> 1 MB) approach the raw link speed.
  • Congestion – in oversubscribed topologies, micro‑bursts can cause queue buildup and increased latency.

A rule of thumb: for latency‑sensitive training, aim for a topology where the average hop count is ≤ 1.5. For HPC workloads that perform large‑scale FFT or linear algebra, a fat‑tree with bisection bandwidth close to the aggregate link bandwidth is ideal.

Patterns in Production

Large‑Scale AI Training Clusters

Meta’s AI Research SuperCluster (RSC) uses a combination of NVSwitch and InfiniBand to connect 16,000 A100 GPUs. The design employs a three‑tier fat tree: each GPU is connected to a local NVSwitch, which in turn connects to spine switches. This allows any GPU to communicate with any other GPU with at most three hops, keeping the all‑reduce time within 10 ms for a 1 GB tensor.

Key production patterns:

  • NVLink‑only islands – groups of 8 GPUs form a fully connected clique via NVSwitch, then these islands are linked with InfiniBand. This reduces the number of NVSwitches needed while preserving low‑latency communication within the island.
  • Dynamic routing – some frameworks (e.g., NCCL) can adaptively choose between NVLink and PCIe paths based on message size and congestion, improving effective bandwidth by 15–20 %.
  • Power and cooling – NVSwitches draw ~30 W each; a 16‑GPU cluster with 4 switches can consume an additional 120 W, requiring careful power budgeting and liquid cooling in dense racks.

HPC and Scientific Computing

In HPC, NVLink is often used to build “GPU‑accelerated nodes” where 4–8 GPUs share a single CPU socket. For example, the Summit supercomputer at Oak Ridge National Laboratory uses NVLink to connect 6 Power9 CPUs and 6 V100 GPUs per node, with each GPU linked to two others via NVLink bridges. This topology enables direct GPU‑to‑GPU communication for scientific kernels like lattice QCD, where data movement can dominate runtime.

Production tips:

  • NUMA awareness – when a GPU is attached to a specific NUMA node, pinning memory to that node can reduce PCIe latency by ~30 %.
  • NVLink vs. PCIe – always prefer NVLink for collective operations; fall back to PCIe only when NVLink ports are exhausted or when communicating with GPUs on different nodes.
  • Monitoring – tools like nvlink-topo and NVIDIA’s DCGM can report link utilization, error counts, and topology maps, helping to spot cabling mistakes or failing links.

Key Takeaways

  • NVLink provides up to 900 GB/s bidirectional bandwidth per link, bypassing PCIe for direct GPU‑to‑GPU communication.
  • NVSwitch enables full‑clique topologies for up to 18 GPUs; larger clusters require fat‑tree or hybrid designs.
  • Topology choice directly affects latency and effective bandwidth: full clique minimizes hops, fat tree scales while preserving bisection bandwidth.
  • Production clusters (e.g., Meta RSC, Summit) combine NVLink islands with InfiniBand for global communication.
  • Always monitor link health and align GPU‑NUMA placement to maximize performance.

Further Reading