What Is a Mixture-of-Experts (MoE) Model? Why Modern LLMs Only Use Part of Themselves
An earlier article in this series explained a large language model as a fixed set of parameters doing the same work on every token. That’s true for a dense model — but a growing share of today’s largest and most efficient LLMs, including DeepSeek-V3, Mixtral, and (by widespread but unconfirmed report) GPT-4, are built differently, using an architecture called Mixture of Experts (MoE). The idea is simple to state and genuinely clever in what it trades away.
The core idea: don’t use the whole model on every token
In a dense model, every single parameter participates in processing every token — a 70-billion-parameter dense model does 70 billion parameters’ worth of computation for each token it generates, no matter how simple or complex that token’s context is. An MoE model instead splits a large chunk of the network into many smaller sub-networks called experts — commonly somewhere between 8 and 256 of them, depending on the model — and adds a small trainable component called a router (or gating network) whose job is to look at each token and decide which handful of experts should actually process it.
A typical setup uses top-k routing: for each token, the router picks the top 2 (or some similarly small number) of experts out of the full set, and only those experts do computation on that token — the rest sit idle for that step. DeepSeek-V3, for example, has 671 billion total parameters but activates only around 37 billion of them per token — roughly 5.5% of the model does the work on any given token, while the rest stays dormant until a different token routes to a different combination of experts.
Why this trade is worth making
The practical payoff is real: a sparse MoE model can have a much larger total parameter count — and correspondingly larger total knowledge capacity — than a dense model would allow for the same computational cost per token. You get to build something with the raw capacity of a very large model while paying, per token, roughly the compute cost of a much smaller one. This is the central reason MoE has become popular for frontier and efficiency-focused models alike: it’s a way to keep scaling total capacity without scaling inference compute at the same rate.
It’s worth being precise about what doesn’t get cheaper, though: memory. Every expert still has to be loaded and ready in case the router sends a token its way — you can’t predict in advance which experts a given conversation will need, so you can’t just keep a few experts in memory and discard the rest. This means MoE models trade cheaper compute for a real, unavoidable memory cost: serving a 671-billion-parameter MoE model still requires enough combined GPU memory to hold all 671 billion parameters, even though any single token only exercises a small slice of them. This is why MoE architecture shows up overwhelmingly in models designed to be served at scale across clusters of GPUs, rather than something that makes a model dramatically easier to run on a single consumer laptop.
The hard part: keeping the router honest
The genuinely tricky engineering problem in MoE isn’t the concept — it’s training the router well. Left alone, a router has no inherent incentive to spread tokens evenly across all the available experts; it can collapse into a lazy pattern where a handful of experts get picked constantly and the rest barely get trained at all, since a sub-network that rarely gets routed to also rarely gets a training signal, which makes it even less likely to develop anything useful, which makes the router even less likely to pick it — a self-reinforcing feedback loop that wastes most of the model’s supposed capacity if left unchecked. Modern MoE training addresses this with explicit load-balancing techniques — auxiliary loss terms during training that specifically penalize uneven expert usage, nudging the router toward spreading work across the full set of experts rather than a favored few.
This is also, honestly, why MoE isn’t a strictly-better default choice for every model. Dense models remain simpler to train predictably, simpler to fine-tune, and simpler to deploy on constrained hardware. MoE earns its added complexity specifically at the scale where the compute savings outweigh the extra engineering and memory cost — which is exactly the regime frontier labs and cost-conscious model builders are both operating in right now, for related but slightly different reasons: frontier labs want more total capacity without a proportional compute bill, and efficiency-focused labs (a category where several prominent Chinese model releases have specifically competed) want strong benchmark performance without frontier-lab-scale compute budgets.
Where this fits in the larger picture
MoE doesn’t change what an LLM fundamentally is — it’s still trained on the same next-token prediction objective covered in the dedicated article on how LLMs work, still processes tokens through the same kind of transformer attention mechanism covered in the piece on AI’s history. What MoE changes is purely an efficiency decision inside the architecture: whether every parameter works on every token (dense), or whether a learned router activates only a relevant subset (sparse, Mixture-of-Experts). It’s an engineering answer to a scaling problem, not a different kind of intelligence — which is exactly the kind of distinction worth being precise about, rather than folding into vague talk about a model being “more advanced.”
Is a Mixture-of-Experts model the same as running several separate AI models and voting?
No, and this is a common mix-up. Ensembling runs multiple complete, independently trained models and combines their outputs. An MoE model is one single network, trained end to end, where a router picks a small subset of its internal expert sub-networks to process each individual token — the experts aren't separate models and were never trained independently.
Does MoE make a model cheaper to run?
It makes the compute per token cheaper, since only a fraction of the total parameters do work on any given token. It does not reduce the memory needed to hold the model, since every expert still has to be loaded and available in case the router picks it. This split — cheap compute, expensive memory — is exactly why MoE models are efficient to run at scale on GPU clusters but aren't necessarily easier to run on a single consumer device.
Why don't all LLMs just use MoE if it's more efficient?
MoE training is genuinely harder to get right — routers can collapse into always picking the same few experts (wasting the rest of the model's capacity), which requires extra balancing techniques during training to prevent. It also adds real engineering complexity for serving the model efficiently across hardware. Dense models remain simpler to train, tune, and deploy, which is still a real advantage for many use cases.