A datacenter fabric diagram showing GPU racks connected via spine-leaf switches.

AI Networking: How GPUs, Fabrics, and Collective Operations Reshaped the Datacenter

AI workloads pushed the network to its limits. This post walks through the protocols, topologies, and failure modes that define modern AI networking — from RoCE and collective operations to fat-tree fabrics and packet-level telemetry.

September 5, 2026 · 9 min · 1915 words · martinuke0

Optimizing Distributed Model Training on Bare‑Metal Clusters with RDMA and Low‑Latency Interconnects

Introduction Training state‑of‑the‑art deep‑learning models now routinely requires hundreds of GPUs working in concert. While public cloud providers offer convenient, on‑demand clusters, many research labs and enterprises still prefer bare‑metal clusters for three core reasons: Predictable performance – no noisy neighbors, no hypervisor overhead. Cost efficiency at scale – amortized CAPEX and lower per‑GPU price. Full control over hardware and software – ability to fine‑tune network stacks, install custom drivers, and leverage specialized interconnects. When you combine bare‑metal hardware with RDMA (Remote Direct Memory Access) and low‑latency interconnects such as InfiniBand or RoCE (RDMA over Converged Ethernet), you can dramatically reduce the communication overhead that traditionally limits distributed training speed. This article walks through the entire optimization stack—from networking fundamentals to concrete PyTorch code—so you can extract the maximum throughput from your cluster. ...

April 3, 2026 · 11 min · 2238 words · martinuke0
Feedback