Trust is earned, not given

A different perspective

2025-01-16 · Projects

AI Frontiers, part 13: DeepSeek-V3 and the cost-efficiency shock

Part 13from the AI Frontiers series · 65 parts in all

DeepSeek released V3 on 26 December 2024, with a technical report and open weights. The architecture was what the previous entries in this series would have predicted — a mixture-of-experts transformer, 671 billion total parameters of which 37 billion are active per token, trained on 14.8 trillion tokens (DeepSeek-AI, "DeepSeek-V3 Technical Report"). The number that travelled around the world was a different one. The report states that the full training run took 2.788 million H800 GPU-hours, and that at a rental price of two dollars per GPU-hour this amounts to roughly $5.6 million.

Set that against a widely repeated figure for the previous generation of frontier models — hundreds of millions of dollars, hundreds of thousands of expensive GPU-hours — and the reaction makes sense. Either the number was misleading, or the cost of training a frontier model had fallen by more than an order of magnitude in about a year. The honest answer, which I want to work through carefully, is that the number is real but measures less than people think, and that the underlying efficiency gains are real and larger than most commentary credits. Both halves matter.

What the report actually claims

Start with the architecture, because the cost claim is a consequence of it. V3 is 61 layers of DeepSeekMoE with 256 routed experts plus one shared expert per layer, routing each token to the top eight. It uses multi-head latent attention, inherited from V2, which compresses the key-value projections into a low-dimensional latent vector before reconstructing them. It is trained in FP8 mixed precision, with a small number of sensitivity-driven exceptions kept at higher precision. Load balancing across experts is handled without an auxiliary loss. And it is trained with multi-token prediction, a secondary objective that asks the model to predict several future tokens at once (DeepSeek- AI; Gloeckle et al.).

Several of those choices are direct answers to a hardware constraint. The report notes that the training cluster had limited cross-node interconnect bandwidth — a consequence of the export-controlled variants of the accelerators available in that environment — and that this shaped the engineering. Where a well-connected cluster encourages you to spread a problem widely and shuffle aggressively, a bandwidth-starved one rewards keeping communication rare and local. The fine-grained expert parallelism and the communication scheduling in the report read like someone designing around a narrow pipe, because that is what they are.

None of this is secret. All of it is hard.

That is the part of the story I find most interesting and most under-reported. Every ingredient in V3 was published before V3 existed. Sparsely-gated mixture-of-experts layers go back to Shazeer et al. in 2017; Switch Transformers (Fedus et al.) made routing practical at scale; the part 8 discussion of the memory-versus-compute tradeoff is exactly this design space. FP8 formats were specified in 2022 by Micikevicius et al. and were known to be viable for training. Chinchilla (Hoffmann et al.) settled the tokens-per-parameter question. Multi-token prediction was published in April 2024. Auxiliary-loss-free balancing was published by the DeepSeek team itself in August 2024 (Wang et al.) — a fix for a known pathology, where the auxiliary loss that keeps experts evenly loaded also degrades the model's quality by fighting the main objective.

So the contribution is not invention. It is the willingness to assemble the whole set of published techniques in one run, fix the parts that break at scale, and publish the engineering: the pipeline parallelism layout, the FP8 quantization scheme with its per-block scaling factors, the communication overlap strategy, the mixture-of-experts balancing algorithm. That last item is a good example of why this kind of work is hard even when it is not novel. A routing scheme that works at 8 experts does not obviously work at 256, and the failure mode is not a crash — it is a model that quietly routes everything through a handful of experts and trains poorly for reasons nobody can see from the loss curve.

Do the arithmetic yourself

The best way to judge an efficiency claim is to normalize it. Tokens trained divided by accelerator-hours gives a throughput number that lets you compare runs on different hardware and different cluster sizes, and it is the number that actually indicates how well the hardware was used.

V3 trained on 14.8 trillion tokens in 2.788 million H800 GPU-hours. That is roughly 5.3 million tokens per GPU-hour. For comparison, the Llama 3 herd report states that training the 405B model took about 30.84 million H100 GPU-hours for roughly 15 trillion tokens, which is on the order of 0.5 million tokens per GPU-hour (Meta AI). Different models, different lab reports, different hardware — but the order-of-magnitude gap is the thing to explain, and the explanation is architectural rather than mysterious. A densely activated 405B model uses every parameter for every token. A model that activates 37 billion of 671 billion uses about a tenth of the compute per token while having a comparable number of parameters to store. The MFU-style comparison is unfair to the dense model in one direction (H100s are faster per chip) and to V3 in another (effective throughput depends on batch sizes and communication overlap), but the direction of the result is not in doubt.

There is a second normalization worth doing, and it is the one that predicts whether a claim will survive scrutiny: what is the denominator? If someone quotes a training cost, ask whether it includes ablations, curriculum experiments, data cleaning, failed runs, and the salaries. If the answer is no, the figure describes the final sprint, not the project. V3's report is unusually explicit about this exclusion, which is why I am willing to take the number seriously at all — the authors made the boundary of the claim visible instead of letting readers assume it covered everything.

Serving economics, which changed more than training economics

The number that actually changes purchasing decisions for most organizations is not the training cost. It is the price per million tokens, and here the MoE-plus-MLA design pays off twice. Because only 37 billion of the 671 billion parameters are active per token, the compute per token is roughly that of a much smaller dense model. Because MLA compresses the KV cache — the per-request working memory that part 7 identified as the real constraint on long context — the memory cost of serving a request is a fraction of what an equivalent dense model would need. DeepSeek's published API price for V3 was in the region of $0.27 per million input tokens and $1.10 per million output tokens, with a further discount for cache hits (DeepSeek, "Models & Pricing").

At the time, that was roughly an order of magnitude below the comparable tier at the frontier labs. It is worth being careful here, because the comparison is not apples to apples: DeepSeek is selling capacity on its own hardware with no distribution business to subsidize, so it is not competing on the same terms as a company selling subscriptions. But the direction of the pressure is unambiguous. When a credible open-weights model sits there at a tenth of the price, every downstream product's unit economics get renegotiated whether or not anyone switches vendors.

Why $5.6 million is both true and misleading

Now the caveats, in the order they matter.

It excludes everything except the final run. The report says so explicitly: the figure covers the official training run and excludes prior research, ablation experiments, failed runs, data acquisition and curation, and the salaries of the people who did the work. Research organizations burn far more compute on experiments than on the run that gets published. Any comparison to another lab's headline architecture cost is comparing one company's marginal cost to another's average.

It excludes the hardware. A two-dollar-per-GPU-hour rental price is a market rate; owning or reserving 2,048 high-end accelerators costs money whether or not they are busy. The interesting economic question is not what the run cost but what the capability cost to acquire, and that number is much larger.

It says nothing about the data. V3 trained on 14.8 trillion tokens of curated text and code. The curation pipeline is undocumented in the detail that would let anyone reproduce it, and data quality is where most of the hard-won advantage in this field lives — the lesson of part 9 and of the chinchilla-era rewrites alike. Compute was the easy thing to measure, so it got measured.

And it does not generalize. One team achieving an efficiency result on one hardware configuration with one set of tricks does not establish that the whole field's costs fell by 20x. It establishes that the frontier of efficient training moved, and that a well-resourced group outside the historical center of gravity can reach it.

What it actually changes

Three things, in descending order of confidence.

First, the moat argument needs revising. The comfortable 2023 position was that frontier capability requires capital on a scale only a handful of firms can muster, so the model layer would concentrate. V3 does not refute that — DeepSeek is well funded and has spent years on this — but it does chip away at the assumption that the efficiency frontier belongs to whoever spends most. Algorithmic improvements diffuse. They diffused here within months, in public, with weights attached.

Second, the constraint is real and it shapes the answer. The interesting technical content of this report exists because of a bandwidth limitation. That is a recurring pattern in engineering: the best work in any generation comes from people who cannot afford the obvious approach. Designing for a narrow interconnect produced an architecture that is also cheap to serve, which is a more durable advantage than winning a benchmark.

Third, the competitive pressure lands on price, and price lands on products. If a competent open-weights model costs a tenth as much per token, then every application whose economics depended on expensive inference is now worth re-examining. That is the real news of December 2024, and it is not about a leaderboard.

There is a policy layer to this that I do not want to dodge, because it was the loudest part of the public reaction. A recurring argument for restricting the export of advanced accelerators was that compute scarcity would slow the development of frontier models in other countries. V3 is awkward evidence for that proposition. Not because the restrictions failed — DeepSeek plainly worked under constrained hardware, and the report's bandwidth-limited design is a direct artifact of those constraints — but because the constraint was absorbed by engineering rather than by deciding not to build the model. The counterargument is that the restriction bought time, and that a team with unrestricted hardware would have moved faster still. Both readings are consistent with the same report, which is why the report is worth reading rather than arguing about.

What I would watch next is the same team's work on reasoning and reinforcement learning. The V3 report describes a strong base model trained with an efficiency discipline that the rest of the field will be studying for a while. That discipline applied to long-chain-of-thought training — the part 11 problem, where inference cost grows with thinking time — is where the next surprise is likely to come from. If $5.6 million buys a frontier base model, the follow-up question is what it buys when the model has to think before it answers.

Works Cited

Brown, Tom B., et al. "Language Models Are Few-Shot Learners." arXiv, 2020, arxiv.org/abs/2005.14165. Accessed 16 Jan. 2025.

DeepSeek. "Models & Pricing." DeepSeek API Docs, api-docs.deepseek.com/quick_start/pricing. Accessed 16 Jan. 2025.

DeepSeek-AI. "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model." arXiv, 2024, arxiv.org/abs/2405.04434. Accessed 16 Jan. 2025.

---. "DeepSeek-V3 Technical Report." arXiv, 2024, arxiv.org/abs/2412.19437. Accessed 16 Jan. 2025.

Fedus, William, et al. "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity." arXiv, 2021, arxiv.org/abs/2101.03961. Accessed 16 Jan. 2025.

Gloeckle, Fabian, et al. "Better & Faster Large Language Models via Multi-token Prediction." arXiv, 2024, arxiv.org/abs/2404.19737. Accessed 16 Jan. 2025.

Hoffmann, Jordan, et al. "Training Compute-Optimal Large Language Models." arXiv, 2022, arxiv.org/abs/2203.15556. Accessed 16 Jan. 2025.

Meta AI. "The Llama 3 Herd of Models." arXiv, 2024, arxiv.org/abs/2407.21783. Accessed 16 Jan. 2025.

Micikevicius, Paulius, et al. "FP8 Formats for Deep Learning." arXiv, 2022, arxiv.org/abs/2209.05433. Accessed 16 Jan. 2025.

OpenAI. "GPT-4 Technical Report." arXiv, 2023, arxiv.org/abs/2303.08774. Accessed 16 Jan. 2025.

Shazeer, Noam, et al. "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer." arXiv, 2017, arxiv.org/abs/1701.06538. Accessed 16 Jan. 2025.

Wang, Lean, et al. "Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts." arXiv, 2024, arxiv.org/abs/2408.15664. Accessed 16 Jan. 2025.