AI Frontiers, part 8: Mixture of experts and the memory-feasible frontier
Part 8from the AI Frontiers series · 65 parts in all
In December 2023, Mistral AI dropped a torrent link on Twitter with no paper, no blog post, and a 26-word caption. The model inside — Mixtral 8x7B — turned out to be the most architecturally interesting release of the year: a mixture-of-experts language model that matches or beats GPT-3.5 on most benchmarks while activating only 13B of its 47B parameters per token (Jiang et al.). Mistral followed with a proper paper in January. By March 2024, Mistral Large had arrived proprietary, Groq was serving Mixtral at hundreds of tokens per second, and MoE had gone from specialist technique to the assumed shape of the next frontier generation — a bet GPT-4 was already widely believed to embody. This entry is about the idea, which is thirty years old, and the economics, which are new.
The oldest new idea in the field
Mixture-of-experts dates to 1991, when Jacobs and colleagues at UMass and MIT proposed dividing a problem among networks that specialize, supervised by a gating network that learns whom to ask (Jacobs et al.). The modern revival is Shazeer's 2017 Sparsely-Gated MoE layer: replace a transformer's feed-forward block with dozens of parallel expert FFNs plus a router that sends each token to a chosen few — capacity exploded to 137B parameters while compute stayed near a small dense model's (Shazeer et al.). GShard carried the idea to multilingual translation at 600B scale with the load-balancing machinery that makes it trainable (Lepikhin et al.), and Switch Transformer simplified the routing to one expert per token, showing stable training at trillion-parameter scale (Fedus et al.). The belief that GPT-4 is an 8-way MoE — first reported by semi-anonymous sources, echoed ever since — is unconfirmed but consistent with its reported serving economics; treat it as informed rumor, not fact. What is fact: by early 2024, every frontier lab was training MoE. The idea won.
What Mixtral actually is
Strip Mixtral to its frame and it is Llama-shaped — MHA-derived attention, RoPE, SwiGLU, RMSNorm — with one substitution: at each of its 32 layers, the single FFN is replaced by eight expert FFNs (about 7B parameters each) plus a router. For each token, the router picks the top-2 experts by softmax score; the token's hidden state is processed by those two, summed, and passed on. Consequences worth internalizing:
Parameters and compute decouple. Total weights: ~47B. Active per token: ~13B. Knowledge capacity scales with the 47B; per-token cost (time, FLOPs, energy) scales with the 13B. You get a model that "knows" like a 47B dense model but runs — per token — like a 13B one.
The router is the whole ballgame. Tokens are routed per-token, per-layer; there is no persistent "topic assignment." Routing failures poison training: if the router develops a taste for one expert, the others starve, the router's preference deepens, and the model collapses toward a small dense model. Both GShard and Switch handle this with auxiliary load-balancing losses; Mixtral's paper reports it trained with an auxiliary load-balancing loss of the same family — and one of the paper's quiet findings is that experts do not specialize by topic. They specialize syntactically (punctuation, numbers, whitespace) and by position — the interpretable "expert on biology" story the internet told in December was, per the paper's own analysis, mostly wrong. Emergent routing still yields the benefits; the folklore was just folklore.
The system dimension is where it bites. MoE shifts cost from GPU-hours to RAM. You still hold 47B weights in memory to serve 13B-active compute — so per-token inference is cheap but deployment needs HBM (or aggressive quantization: 4-bit Mixtral runs in ~26GB, on a pair of 24GB consumer cards). Serving benefits enormously from expert parallelism across devices and from batch effects: once many requests flow, all experts stay warm and utilization is high; single-stream latency pays the whole memory bill. This is why Groq's LPUs and the vLLM generation of serving stacks mattered so much to Mixtral's reception — MoE rewards infrastructure that batch and parallelizes well.
What an MoE deployment actually runs like
Because several of you will meet MoE as operators before you meet it as trainers, here is what serving Mixtral-class models actually demanded in the first quarter of 2024, learned from the deployments that worked and the ones that spent February learning. Memory first: 47B parameters in float16 is ~94GB before KV cache — two 80GB cards minimum, or a 4-bit build at ~26GB that fits a pair of 24GB consumer cards with headroom for a modest cache. Expert parallelism: shards experts across devices so each card holds one or two experts; the all-to-all traffic per layer is the communication pattern to watch, and on PCIe-boxed consumer hardware it is the reason 2-card rigs outperform 4-card rigs per dollar beyond a point. Batching, per the worked example above, is where MoE pays: the vLLM-generation servers with continuous batching turned a single A100 into a Mixtral endpoint that hundreds of concurrent users could share — the model's active-13B cost is why its throughput ceiling beat same-class dense models so visibly on the leaderboards that measured it. And quantization: the QLoRA line from part 3 applies unchanged — a 4-bit Mixtral with 64-rank LoRA adapters fine-tunes on a single 48GB card, which is the sentence that mattered to every team without a cluster.
The failure reports clustered, instructively, around the router's systems consequences rather than the model's quality: hot experts on cold cards, load imbalance under skewed multi-tenant traffic, cache thrash when expert placement ignored routing statistics. The model was fine; the placement was the product. It is an early preview of a theme this series will keep meeting — at frontier scale, model quality and systems engineering stop being separable disciplines, and the teams that win treat the weight file as a hardware layout decision.
Why sparse beats dense at equal compute: the scaling argument
It is worth stating precisely why the field converted to MoE, because the reason is not "experts" as folklore imagines them — it is arithmetic. Dense scaling laws (part 2's economics) say quality per FLOP improves smoothly with model size, but a dense model spends every FLOP on every token: a 70B dense model computes all 70B parameters for "the," for "quantization," for every word. Most of that compute is irrelevant to most tokens — the information a token actually needs to route through is a small fraction of what the model knows. Sparsity exploits that slack: capacity grows with total parameters while per-token cost tracks only the routed subset. In the frontier labs' terms, MoE buys a better quality-per-training-FLOP frontier and a better quality-per-inference-FLOP frontier at the price of memory — and memory, unlike FLOPs, got cheap through exactly the period (2021– 2023) when 80GB HBM cards and expert-parallel collectives matured.
The catch — the reason the conversion took three years despite Switch Transformer's 2021 results — is that the arithmetic only pays at scale. Below tens of billions of parameters, the router overhead, the load-balancing losses, the all-to-all communication, and the memory bill exceed the savings; MoE is a frontier-scale technology, which is why your favorite 7B models are all dense and will stay so. There is a neat symmetry with part 9's small models: at the bottom of the market, data quality (Phi) substitutes for parameters; at the top, sparsity (Mixtral-class and beyond) extracts more from parameters. The middle dense workhorse — the 70B — is squeezed from both directions, and its holders know it.
MoE in one worked example
Abstract descriptions of routing age badly, so let me make the mechanism concrete with Mixtral's actual numbers. The model has 32 layers; at each layer, eight expert FFNs of roughly 7B parameters sit in parallel where a dense model would have one. A token arrives at layer 17. The router — a single linear projection from the token's hidden state to eight logits — scores all eight experts, picks the top two, and the token's hidden state flows through only those two FFNs (each a SwiGLU block: down-project to the expert's hidden width, gate, up-project back). Their outputs are summed, scaled, and handed to layer 18, whose router starts fresh. Over a 4,096-token sequence, that means 4,096 independent top-2 decisions at every one of 32 layers — around 260K routing events for a single pass, each cheap (one 4,096×8 matmul) but each consequential: different tokens, even adjacent ones, take different computational paths through the same weights.
Two facts follow from the arithmetic. First, batching changes everything: with 32 concurrent requests, each of the eight experts at each layer likely has several tokens assigned per forward pass, so all experts stay busy and throughput approaches a dense model's — MoE loves serving load. With a single stream, four experts per layer sit idle every token, and you pay full memory for half the compute: MoE hates solo latency. Second, load balance is a live systems property, not a training afterthought: a deployment whose traffic skews — one very chatty tenant, one domain of prompts — can skew routing, which skews expert utilization, which skews hot-vs-cold memory across devices. The paper's load-balancing loss handles training; production needs its own monitoring. Both facts were learned, or re-learned, by every team that stood up Mixtral serving in January and February of this year.
Reading the Mixtral moment
The reception is as instructive as the architecture. Mixtral's December release was an unsigned torrent — no announcement, no paper — and the community reverse-engineered it within hours: architecture identified from weight shapes, benchmarks run overnight, GGUF quantizations on TheBloke's shelf by the next morning. When the paper arrived in January, it read as confirmation of what the community had already measured. Compare that cycle to a closed-lab release: zero architecture information, benchmark tables only. The open release turned the model into infrastructure before its own documentation existed — and by February, Mixtral-based fine-tunes were out-performing the base on domain tasks using the exact part-3 pipeline, at 13B-active cost.
The reception also exposed how much evaluation debt the field carried. Mixtral beat GPT-3.5 on most static benchmarks — and the Arena Elo told a more nuanced story: roughly parity on hard prompts, with GPT-3.5 retaining an edge on instruction-following nuance and English prose quality. Both were true. Static benchmarks measure what a model knows when asked well-formed questions; the Arena measures what users experience when they ask whatever they actually ask. The MoE architecture exposed the gap between those two measurements more sharply than any release before it, which is worth remembering the next time a launch blog post leads with MMLU.
One more consequence nobody priced: MoE changed the open-weights licensing conversation. Mixtral shipped under Apache 2.0 — no commercial-use carve-outs, no MAU clause, no attribution requirements — a deliberate contrast with Llama 2's community license (part 4) and, not coincidentally, an enterprise procurement unlock. Mistral raised at a reported $2B valuation within days of the paper. The lesson generalizes: in infrastructure markets, license terms are product features, and the vendor who removes friction at the right moment captures the ecosystem regardless of who has the absolute best weights.
The economics, stated plainly
MoE trades your GPU-hours bill for a memory bill. Training: more total parameters to hold and communicate (expert-parallel all-to-all traffic is the engineering battle), but fewer FLOPs per token than a dense model of equal total size — and, per the scaling-law logic of part 4's economics, a better quality-per-FLOP frontier. Inference: latency and energy near the active-parameter model; memory near the total-parameter model. For API providers running massive batch, that trade is a win; for on-device, it is a challenge — which is why the small-model movement of part 9 and the MoE movement are the same conversation happening at different scales.
Watch this space through 2024, because the conversion is not finished. Databricks has been openly building on the MoE line (their 2023 work on MoE scaling was explicit about where the economics point); xAI's November-found XAI corporation hired the GShard and Switch lineage; and the rumors about GPT-4's architecture have hardened from semi-anonymous leaks into near-consensus. The dense frontier model is not dead — Llama-class dense models still define the open-weights quality curve, and dense training remains the path of least surprise below frontier scale — but the frontier's economics increasingly run through sparsity. Mixtral's December torrent, in hindsight, was the public announcement of that shift, and by the time this series next touches frontier architecture, I expect MoE to be the assumed default at the top of the market, not the exception.
Works Cited
Fedus, William, Barret Zoph, and Noam Shazeer. "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity." Journal of Machine Learning Research, vol. 23, no. 120, 2022, jmlr.org/papers/v23/21-0998.html. Accessed 21 Mar. 2024.
"Mixtral of Experts." Mistral AI, 11 Dec. 2023, mistral.ai/news/mixtral-of-experts/. Accessed 21 Mar. 2024.
Jacobs, Robert A., et al. "Adaptive Mixtures of Local Experts." Neural Computation, vol. 3, no. 1, 1991, pp. 79–87, doi:10.1162/neco.1991.3.1.79. Accessed 21 Mar. 2024.
Jiang, Albert Q., et al. "Mixtral of Experts." arXiv.org, 2024, arxiv.org/abs/2401.04088. Accessed 21 Mar. 2024.
Lepikhin, Dmitry, et al. "GShard: Scaling Giant Models to Trillion Parameters Using Sequence-level Pipelining and Sharding." arXiv.org, 2020, arxiv.org/abs/2006.16668. Accessed 21 Mar. 2024.
Shazeer, Noam, et al. "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer." arXiv.org, 2017, arxiv.org/abs/1701.06538. Accessed 21 Mar. 2024.