Dispatches
Essays··5 min read

The Interconnect Decision That Moved Upstream

CoWoS packaging capacity — not wafer fabrication — has become the binding constraint on AI hardware, and the shortfall has reorganised how clusters are designed. Network architects who once specified fabrics after compute teams locked the BOM now sit in allocation meetings, because the interconnect bandwidth budget and the packaging timeline are the same constraint. Memory, topology, and fabric, previously sequential decisions, now collapse into one.

The senior network architect at a hyperscaler received the final CoWoS allocation from TSMC in late September 2026, and the number that stopped the planning cycle was 22% below the compute roadmap. TSMC's CoWoS lines are fully booked, with total 2026 demand estimated near 1.0 million wafers and NVIDIA holding roughly 60% of capacity, which means every other buyer is working from a constrained denominator. The roadmap assumed 1,024 accelerators per pod, each carrying eight stacks of HBM3E, wired through a NVLink-class scale-up fabric. The allocation delivers 790 accelerators. That is not a line item to negotiate. The package supply is the compute supply, and advanced packaging has become the binding constraint on AI hardware in 2026, no longer wafer fabrication.

The engineer's job changed shape in that moment. For fifteen years the sequence ran: model the training workload, size the compute, then design the fabric to match. The interconnect decision sat downstream of the silicon decision. In September 2026 it moved upstream. The fabric now determines which accelerators fit the allocation, and whether the cluster can run the target model at all.

This is not a vendor-selection problem. Open alternatives like UALink and Ethernet-based options are moving from specification to announced hardware, with production in late 2026 and into 2027, but designs requiring greater numbers of SoCs and HBM modules need to transition to CoWoS-L, which NVIDIA and AMD AI accelerators have already adopted. The bottleneck is not which standard wins. It is that memory bandwidth and packaging throughput have converged into a single constraint, and the network topology has to absorb the shortfall.

The trade is arithmetic. Compute capabilities advance at a rate of 3x every two years, while HBM bandwidth scales only at less than 2x every two years. During autoregressive decoding, each forward pass requires reading all model weights from VRAM, so for a 70B parameter model at FP16 that is approximately 140 GB of data transferred per token step, taking roughly 42 milliseconds per token at maximum theoretical bandwidth on an H100 with 3.35 TB/s. The 42 ms lower bound cannot be improved by adding FLOPS or changing the network between nodes. It is a per-accelerator memory wall, and the only way to scale throughput is to spread the inference load across more devices, which immediately turns the problem into an interconnect-sizing exercise.

The cluster shrinks, the model stays the same size, and the per-node memory pressure rises. Under the original roadmap, tensor parallelism would have run inside 128-accelerator islands connected by NVLink's 1.8 TB/s per GPU, with pipeline and data parallelism handled by InfiniBand, the preferred choice for GPU communication due to its extremely low and predictable latency. At 790 accelerators the topology cannot support the same island size without re-splitting the workload. The model either shrinks to fit tighter islands, or the inter-island fabric has to carry tensor-parallel traffic it was never sized for.

The RFQ that went out in the second week of October asked for 400 Gbps InfiniBand NDR uplinks and RDMA over Converged Ethernet with lossless configuration using Priority Flow Control and Explicit Congestion Notification as the fallback. The vendor shortlist includes Broadcom's scheduled Ethernet and the Quantum-X800 delivering 144 ports at 115 Tbps aggregate bisection bandwidth. The decision will turn on whether the KV cache handoff between prefill and decode can tolerate RoCE jitter, because for a trillion-parameter model serving thousands of concurrent requests, the aggregate KV cache transfer bandwidth requirement can reach hundreds of terabits per second. If it cannot, the cluster pays the InfiniBand premium. If it can, the bill of materials drops by 18% and the power budget improves by 11%, but the software stack has to implement retries the hardware would have hidden.

Nobody in the room thinks the 790 number will hold for Q1 2027. Leading-edge 2nm is booked well into 2028 and HBM is allocated through 2026, and HBM3E is effectively sold out for 2026 with prices up double digits year-over-year, while HBM4 is ramping into late 2026 with SK Hynix holding an estimated 62% share. The supplier that confirms HBM4 qualification earliest wins the next allocation round, which moves part of the interconnect decision into the packaging timeline. If HBM4 lands in Q1, the architect can assume 12 stacks per accelerator instead of eight, the per-device bandwidth rises from 3.35 TB/s to roughly 5 TB/s, and the 790-accelerator cluster can run models it cannot touch today. The network does not change. The memory does, and the viable cluster size follows.

The alternative is the one nobody wants to model: CoWoS capacity is sold out well into 2026, HBM is fully allocated, and advanced packaging is now described as the single tightest link in the entire AI semiconductor supply chain. If the allocation drops further, the training roadmap breaks. The fallback is inference-only deployment on the reduced cluster, which requires a different fabric entirely. Prefill and decode split into separate pods, the KV cache transfer between prefill and decode pools becomes the critical interconnect bottleneck, and the architect is back to the RFQ process with a new set of latency requirements and no additional budget.

Three decisions that used to happen in sequence now collapse into one. The memory allocation determines the feasible cluster size. The cluster size determines the partition strategy. The partition strategy sets the interconnect bandwidth and latency budget. The engineer who used to specify switches after the compute team locked the BOM now sits in the allocation meeting, because the fabric is part of the constraint set, not the solution set.

The strategic winners will be the companies that treat interconnect as a first-order architectural decision, not a late-stage engineering problem. That sentence, from a co-packaged optics brief published in July 2026, reads differently in October. It is not strategy advice. It is a description of the only operating mode left when packaging capacity runs out before compute demand does. The network engineer's scope expanded, the timeline compressed, and the approval chain now runs through the packaging team. The part that has not changed is the deadline. The cluster goes live in March 2027, whether the allocation holds or not.


Tarry Singh is the founder and CEO of Real AI (realai.eu), an enterprise AI advisory and deployment firm working with global enterprises on production agent systems, model risk, and AI sovereignty strategy. He also leads Earthscan (earthscan.io) for Energy AI, and is a founding contributor to the EU-funded HCAIM and PANORAIMA programmes for responsible AI education across European universities. He writes at tarrysingh.com.

Cartouche
The Interconnect Decision That Moved Upstream · Dispatches, 8 October 2026 · T. Singh