Trust is earned, not given

A different perspective

2024-05-02 · Projects

AI Frontiers, part 39: The GPU supply chain and the cost of a training cluster

Part 39from the AI Frontiers series · 65 parts in all

The scaling laws describe a relationship between compute and capability. They do not describe how you obtain compute, and for the last several years that has been the more binding question. Building a frontier training cluster requires accelerators that a single company manufactures, memory that comes from three suppliers, packaging capacity that is the scarcest step in the chain, an interconnect fabric that only a few vendors make, and a quantity of electrical power that requires negotiating with a utility and often a local government.

This entry is about that industrial layer, because it explains a great deal about the shape of the field: why the number of organizations capable of a frontier run is small, why governments began treating compute as a strategic resource, and why the cost of a model is only partly a research question.

The component chain, and where the bottlenecks actually are

An accelerator is not one product. Several independent supply chains have to converge, and each has its own capacity limits.

Compute dies. The logic is the part everyone thinks about, and it is manufactured on the most advanced process nodes available — a small number of fabs, with lead times measured in months and capacity allocated years in advance. This is expensive and it has not been the binding constraint.

High-bandwidth memory. Each accelerator carries a stack of memory mounted beside the compute die, and this is where the industry's deepest constraint lives. HBM is difficult to manufacture, yields are hard-won, and it is produced by a handful of firms. Because both training and inference are memory-bandwidth bound — the theme running through quantization, speculative decoding and cache engineering — the demand for memory bandwidth grows with every capability improvement, and it grows faster than the demand for arithmetic.

Advanced packaging. Attaching memory stacks to a compute die requires chip-on-wafer packaging at a scale and precision that only a few facilities can achieve, and this has repeatedly been the step that limited accelerator shipments. It is the least visible stage and it is the one that most often determines how many units exist in a given quarter, which is an unusually strong argument for paying attention to packaging lines when forecasting compute supply.

Interconnect and networking. Training a large model requires thousands of accelerators to communicate continuously, and the fabric is a product in its own right — switches, optical transceivers, cables and the topology that arranges them. The bandwidth of that fabric determines which parallelism strategies are affordable, as the previous entry described, and it is why a cluster's topology is a design decision rather than a detail.

Power and siting. A large facility draws power on the scale of a small city, and the constraints are now physical: generation capacity, transmission, grid interconnection queues that can take years to clear, cooling infrastructure, and local permitting (International Energy Agency). This shifted the conversation from a software-industry problem to an infrastructure one, and the consequences for siting decisions have been visible ever since.

The demand curve, and why prices did not fall

The usual pattern in computing is that a fixed capability gets cheaper every year. That did not happen to AI accelerators in this period, for a reason worth understanding.

Demand grew faster than manufacturing capacity could. The training compute used for frontier models increased by several orders of magnitude over five years (Amodei and Hernandez; Epoch AI), and every organization that decided to build a model needed the same components from the same small set of suppliers. When demand outruns supply by that margin, prices for existing generations stay high and the newest generation sells out on allocation rather than on price. There was no surplus to drive the cost curve down, which is unusual and explains a great deal of the anxiety in the period.

The result was a market with several distinctive features. Long-term contracts and reservations mattered more than spot purchases, because access was the constraint rather than price. Rental markets emerged and their rates became a de facto reference price for the cost of training. And every large buyer pursued alternatives — internal silicon programs, competing vendors, and region-specific accelerators designed to comply with export rules — which is a rational response to a supplier with pricing power and a political dimension.

Export controls, which turned a product into an instrument

The most consequential non-technical development in this period was the use of export controls on advanced semiconductors as a foreign policy tool. Successive rounds of United States rules restricted the sale of the most capable accelerators to certain destinations and defined thresholds of compute and interconnect bandwidth that determined eligibility, with revisions tightening those thresholds over time as newer chips arrived.

The effects rippled through the entire landscape. Vendors designed region-specific variants that sat just below the thresholds, which made the thresholds themselves a design input. Buyers in affected regions accelerated domestic alternatives, which is the outcome the policy was intended to prevent in the short term and to provoke in the long term. And the rules created a compliance function inside hardware companies, where a chip's specifications are now anticipated by a policy team before the product is announced.

The analytical literature on this is unusually good (Sastry et al.), and its central observation is that compute is governable in a way that data and algorithms are not: it is physical, it is concentrated in a few manufacturing locations, it is detectable in shipment records, and it is hard to substitute at scale. That is why compute became the go-to lever for policy, and why any analysis of the field's future that ignores it is incomplete.

What a cluster actually costs

Total cost of ownership is a different number from purchase price, and the difference is where most of the misconceptions live.

Purchase price for the accelerators is the visible part. Around it sit the servers, the networking, the storage for checkpoints and datasets, the facility and its cooling, and the power to run all of it for years. Then there is the cost that matters most and gets discussed least: utilization. A cluster at 40 percent utilization costs two and a half times per useful hour what the same cluster at full utilization costs, and utilization depends on having enough users, enough work, and enough scheduling discipline to keep the machines busy. Organizations that bought accelerators without an answer to that question discovered that the true cost of a GPU is dominated by how long it sits idle.

The second underappreciated line item is people. The number of engineers who can keep a thousand-accelerator cluster healthy is small, they are compensated accordingly, and the debugging tools are immature enough that a fault can consume weeks. A cluster is an organization as much as a capital asset, and the organizations that operate them well have learned that scheduling, quotas and reliability engineering are what turn hardware into throughput.

Which is why rental economics are so attractive and why the rental market became as important as it did. Paying a rate that includes someone else's low utilization is often cheaper than owning the hardware, because the provider has more users and better utilization than any single buyer would. The same argument that produced cloud computing produced GPU rental, with the twist that supply is constrained enough that rental prices stayed high rather than collapsing.

The alternatives, and the moat that is not the chip

Every large buyer pursued a second source, and the interesting part of that effort is where it succeeded and where it stalled.

The hardware side is genuinely competitive. Accelerators from multiple vendors offer comparable memory capacity and bandwidth, and hyperscalers built their own — tensor processing units optimized around a different memory architecture, training and inference chips designed in-house, and inference specialists whose chips trade generality for latency on specific model shapes. The consolidation of several well-funded accelerator startups into acquisitions rather than independent competitors is itself informative: building the silicon turned out to be the tractable half of the problem.

The intractable half is software. The value of an accelerator ecosystem is the accumulated pile of kernels, libraries, compilers and debugging tools built over a decade, and a new chip with better specifications and an immature toolchain is not competitive for a team under deadline. Porting a research codebase is one thing; porting a production training pipeline where a ten percent throughput loss costs millions is another. This is why the substitution that happened was mostly horizontal — different vendors within an ecosystem, or new projects started on a second stack — and rarely a migration of an existing run.

The strategic read is that the moat is a software ecosystem, not a fab. That is better news for competition than it first appears, because software moats erode with accumulated engineering, and worse news than it appears for anyone hoping a better chip ends the concentration.

Inference and training are different markets

Most public discussion of compute focuses on training, and most compute over the life of a deployed model is spent on inference. As models became products, that imbalance grew, and the two markets have different requirements.

Training wants the highest possible arithmetic throughput and interconnect bandwidth, because the job is a small number of enormous, tightly coupled computations. Inference wants memory bandwidth per dollar and low latency per token, because the job is a large number of small, loosely coupled decoding steps — the serving analysis in this series is the whole argument. A chip optimized for one is not optimized for the other, which is why inference-specific silicon appeared and why the previous generation of training accelerators becomes attractive for serving once newer ones arrive.

Two consequences follow. First, the second-hand market matters: a three-year-old accelerator that is obsolete for frontier training is perfectly good for serving a 7B model, and that market determines the cost floor for a whole class of applications. Second, the long-run cost of intelligence in a product is set by the inference market rather than the training market, and every optimization in the latter half of this series — quantization, batching, speculative decoding, cache management, routing to smaller models — is a way of buying capability from inference capacity instead of training capacity.

The strategic consequences, which are now the main story

Three structural changes followed from this period, and each is still shaping the field.

Compute concentration. Because the capital requirement for frontier training rose into the billions and the supply chain could not expand quickly, the set of organizations capable of training a top-tier model stayed small and arguably shrank. That is a departure from the software industry's usual dynamics and it is the backdrop for every argument about open weights, safety regulation and competitive policy in this series.

Sovereign compute. Governments concluded that domestic training capacity is strategically necessary, and the result was national programs, public compute allocations, and a stated desire for regional model capability in languages and domains that commercial models do not serve well. This added a class of buyer with objectives other than return on investment, which changes procurement dynamics and puts a floor under demand independent of commercial cycles.

The efficiency turn. When compute is expensive and scarce, efficiency becomes the highest value research program available, which is exactly the ethos of the open ecosystem described in part 35. Sparse architectures, aggressive quantization, better attention kernels, speculative decoding and cache management are all, viewed economically, ways to get more capability out of the same industrial capacity. The frontier labs spent money on compute and the rest of the field spent effort on efficiency, and both were rational responses to the same constraint.

The framing worth carrying forward is that this is an industrial technology now. The questions that determine what gets built — who can obtain accelerators, how much power a site can draw, which countries are permitted to buy what — are answered by manufacturing capacity and policy rather than by anything in a paper. The algorithms still matter, and the supply chain decides who gets to run them at scale.

Works Cited

Amodei, Dario, and Danny Hernandez. "AI and Compute." OpenAI, 16 May 2018, openai.com/research/ai-and-compute. Accessed 2 May 2024.

Bureau of Industry and Security. "Implementation of Additional Export Controls: Certain Advanced Computing and Semiconductor Manufacturing Items." Federal Register, 2023. Accessed 2 May 2024.

de Vries, Alex. "The Growing Energy Footprint of Artificial Intelligence." Joule, vol. 7, no. 10, 2023, pp. 2191–94. Accessed 2 May 2024.

Epoch AI. "Trends in Machine Learning Hardware and Compute." Epoch AI, epochai.org/trends. Accessed 2 May 2024.

International Energy Agency. "Electricity 2024: Analysis and Forecast to 2026." IEA, Jan. 2024, www.iea.org/reports/electricity-2024. Accessed 2 May 2024.

Lawrence Berkeley National Laboratory. "Queued Up: Characteristics of Power Plants Seeking Transmission Interconnection." LBNL, 2023, emp.lbl.gov/queues. Accessed 2 May 2024.

NVIDIA. "NVIDIA Announces NVIDIA Blackwell Platform." NVIDIA Newsroom, 18 Mar. 2024, nvidianews.nvidia.com. Accessed 2 May 2024.

Sastry, Girish, et al. "Computing Power and the Governance of Artificial Intelligence." arXiv, 2024, arxiv.org/abs/2402.08797. Accessed 2 May 2024.

Zhang, Susan, et al. "OPT: Open Pre-trained Transformer Language Models." arXiv, 2022, arxiv.org/abs/2205.01068. Accessed 2 May 2024.