Essays··8 min read
The 18% You Cannot Name
A 504-GPU training cluster runs at 82% of baseline for six hours: zero errors, all ports active, every temperature in range. Silent throughput degradation — partial fabric failures, NUMA misconfigurations, HBM bottlenecks — sits below every alert threshold and looks identical from a Grafana dashboard full of green panels. NCCL Inspector, released December 2025, is the first collective-level telemetry shipped natively with NCCL; this is a field report from the last training run that did not have it.
distributed-traininggpu-infrastructurencclinfiniband
Read