Mixture-of-Experts (MoE) is the architectural trick behind many of today's largest and most capable language models. Instead of pushing every token through one enormous feed-forward network, an MoE layer holds many smaller expert networks and routes each token to only a few of them. The result is a model with a very large total parameter count but a much smaller active compute footprint per token.
MoE decouples model capacity (total parameters) from per-token cost (active parameters and FLOPs) — you get more knowledge without paying more compute for every token.
Dense vs. sparse: the core tradeoff
A standard "dense" Transformer applies all of its weights to every token. If you want a smarter model, you make the layers wider or deeper, and the cost of every forward and backward pass grows in lockstep with the parameter count. That relationship is the fundamental economic constraint of dense scaling: capacity and compute are welded together.
A sparse model breaks that weld. In an MoE layer, only a small, input-dependent subset of parameters is activated for any given token. The model can therefore carry an order of magnitude more parameters — more places to store patterns, facts, and skills — while the arithmetic performed per token stays close to that of a much smaller dense model.
The intuition worth internalizing: dense models pay for capacity on every token; sparse models pay only for the capacity each token actually uses.
Anatomy of an MoE layer
In a conventional Transformer block, each token passes through attention and then a single feed-forward network (FFN), typically a two-layer MLP with a large hidden dimension. In an MoE Transformer, that single FFN is replaced by:
- N expert MLPs — each an independent FFN with its own parameters. Experts usually share the same shape as the FFN they replace (or a fraction of it, in fine-grained designs).
- A gating / router network — a small learned function (usually a single linear projection followed by a softmax) that scores the experts for the current token and decides where it goes.
Attention layers are typically left dense and shared; only the FFN sublayers are "expertized." A common recipe interleaves MoE and dense FFN layers rather than converting every layer, which stabilizes training and controls memory.
Routing and top-k gating
For each token, the router produces a score (logit) for every expert. A softmax over these logits gives gating weights. The layer then selects the top-k experts by score — k is usually small, such as 1 or 2 — and sends the token only to those.
The layer's output is the weighted sum of the chosen experts' outputs, using the (renormalized) gate weights:
logits = router(x) # shape: [num_experts]
probs = softmax(logits)
idx, gates = top_k(probs, k) # e.g. k = 2
y = sum( gates[i] * expert[idx[i]](x) for i in range(k) )
Switch Transformer popularized k = 1 (route each token to a single expert) for maximum simplicity and throughput. GShard and Mixtral use k = 2, which gives the router a smoother, more forgiving signal and lets two experts collaborate on hard tokens. Larger k increases quality and cost; the sweet spot for most production systems has been small.
Because top-k is a hard, discrete choice, the gate weights are what carry gradients back into the router: the selected experts' contributions are scaled by differentiable weights, so the router learns which experts help. Non-selected experts receive no gradient for that token, which is exactly what makes the layer sparse — and also what creates the balancing challenges discussed below.
Token-choice vs. expert-choice routing
The default scheme is token-choice: each token picks its top-k experts. An alternative, expert-choice routing, inverts the problem — each expert picks the top tokens it wants up to a fixed budget. Expert-choice guarantees perfect load balance by construction (every expert gets exactly its budget) at the cost of some tokens being processed by more experts and others by fewer.
Why MoE buys parameters at near-constant FLOPs
This is the crux of the whole idea. Suppose a dense FFN has P parameters and costs roughly 2P FLOPs per token (multiply-add for a matmul). Replace it with N experts of the same size and route with k:
- Total FFN parameters grow to
N x P— the model's capacity scales with the number of experts. - Active FFN parameters per token stay at
k x P, because onlykexperts fire. WithN = 8andk = 2, you hold 8x the FFN parameters but activate only 2x. - Compute per token is therefore governed by
k, notN. Adding experts increases memory and capacity but leaves the per-token matmul cost essentially unchanged (aside from the negligible router projection).
This is why a model like Mixtral 8x7B has far more total parameters than a 7B dense model yet performs inference with a per-token compute cost closer to a ~13B dense model (the two active experts plus the shared attention and embedding weights), not 8x7B. The "8x7B" name refers to the eight FFN experts, not eight full copies of the network — attention and embeddings are shared, so the total parameter count is well below 56B.
The catch is that all experts must be resident in memory even though only a few run per token. MoE trades cheap, abundant memory/bandwidth for scarce, expensive compute — a good trade when you are compute-bound, which large-scale training and serving usually are.
Load balancing and capacity
The auxiliary load-balancing loss
Left to its own devices, a router tends to fall in love with a handful of experts and send most tokens there. That wastes the other experts and creates severe compute hotspots. To counteract this, MoE training adds an auxiliary load-balancing loss alongside the main language-modeling loss.
The classic formulation (from Switch Transformer) encourages the fraction of tokens dispatched to each expert and the average routing probability for each expert to both be close to uniform. Concretely it penalizes the dot product of two per-expert vectors — the fraction of tokens routed to each expert and the mean gate probability assigned to each expert — scaled by the number of experts. Minimizing it pushes traffic toward an even split. The loss is weighted by a small coefficient so it nudges balance without overwhelming the primary objective.
Capacity factor
For efficient batched execution on accelerators, each expert is given a fixed capacity — a maximum number of tokens it will process in a batch. Capacity is set as:
capacity = capacity_factor * (tokens_per_batch * k / num_experts)
A capacity_factor slightly above 1.0 (commonly 1.0–1.25 in training, sometimes higher) provides slack so that mild imbalance does not immediately overflow. Tokens that arrive at an already-full expert are dropped — they skip the FFN and pass through via the residual connection unchanged. Dropping wastes those tokens' capacity for that layer, so the capacity factor is a direct dial on the quality/efficiency tradeoff: too low drops too many tokens, too high wastes padding compute and memory.
Failure modes and how they are tamed
- Router collapse. The router converges to using only a few experts, starving the rest. Mitigations: the load-balancing loss, adding noise to router logits (noisy top-k) early in training, and careful router initialization.
- Training instability. Sparse models are notoriously sensitive; router logits can grow large and destabilize the softmax. The router z-loss (a small penalty on the log-sum-exp of the router logits) keeps logits bounded and markedly improves stability. Lower-precision routers are especially fragile, so routing math is often kept in higher precision.
- Expert under-utilization / token dropping. Persistent imbalance plus finite capacity means some experts idle while others overflow and drop tokens. Balancing loss, capacity tuning, and expert-choice routing all address this.
- Memory and communication overhead. All experts occupy memory even when idle, and in distributed training experts are sharded across devices, so routing turns into all-to-all communication — tokens are shuffled to wherever their expert lives and back. This all-to-all is frequently the dominant cost and drives system design (expert parallelism, topology-aware placement).
- Poor fine-tuning generalization. Sparse models can overfit more readily than compute-matched dense models on small downstream datasets, which motivates careful regularization and sometimes selective (dense-only) fine-tuning.
Real systems done correctly
Google: GShard and Switch Transformer
GShard (2020) scaled a multilingual translation model into the hundreds of billions of parameters using top-2 MoE layers, and introduced much of the sharding and capacity machinery that later systems reused. Switch Transformer (2021) simplified routing to top-1, showed that this improves throughput while remaining stable with the right losses, and demonstrated favorable pretraining speedups over dense baselines at matched compute.
Mistral: Mixtral 8x7B and 8x22B
Mixtral 8x7B is an open-weights decoder-only MoE with eight experts per FFN layer and top-2 routing. It brought MoE into the mainstream of practical LLM deployment, delivering strong quality at an inference cost far below its total parameter count. Mixtral 8x22B scaled the same recipe to larger experts. Both are widely used precisely because the active-parameter economics make them cheaper to serve than dense models of comparable quality.
DeepSeek: MoE, V2, and V3
The DeepSeek-MoE line refined the architecture with two influential ideas:
- Fine-grained experts — using many smaller experts (and routing to more of them) rather than a few large ones, which increases the combinatorial specialization the router can express.
- Shared experts — a small number of always-on experts that every token passes through, capturing common knowledge so the routed experts are free to specialize. This reduces redundancy across experts.
DeepSeek-V2 and V3 combined this fine-grained + shared-expert design with large total parameter counts and comparatively small active footprints, and V3 notably pushed an auxiliary-loss-free load-balancing strategy that adjusts per-expert routing biases directly rather than relying solely on an auxiliary loss term — reducing the interference that balancing losses can impose on the main objective.
xAI: Grok
Grok-1 was released as a large open-weights MoE model with top-2 routing, another data point that frontier-scale systems increasingly favor sparse designs to reach high capacity within a serving budget.
Dense vs. MoE: side by side
The table below is qualitative and directional — it compares a dense FFN to an MoE FFN with N experts routed top-k, holding the per-expert size roughly equal to the dense FFN.
| Dimension | Dense model | MoE model (N experts, top-k) |
|---|---|---|
| Total parameters | Baseline P (per FFN) | ~N x P (per FFN) — much larger |
| Active params / token | All of them (P) | Only k x P — a fraction of total |
| FLOPs / token | Scales with total params | Scales with k, roughly constant in N |
| Memory footprint | Proportional to (smaller) param count | Large — all experts must be resident |
| Communication | Standard data/tensor parallel | Extra all-to-all for expert routing |
| Training stability | Well-understood, robust | More delicate — needs balancing + z-loss |
| Quality at fixed compute | Baseline | Typically higher (more capacity per FLOP) |
| Quality at fixed total params | Typically higher (all params work per token) | Lower per-param, but far cheaper to run |
The two quality rows capture the essential nuance: MoE wins when your bottleneck is compute, and dense wins when your bottleneck is memory or total parameter budget. Frontier training and serving are usually compute-bound, which is why MoE has become so prevalent.
Practical guidance for engineers
- Reach for MoE when you are compute-bound but have memory/bandwidth to spare. If you can afford to hold many experts in memory and your throughput ceiling is FLOPs, MoE is the natural lever.
- Report active parameters, not just total. When comparing an MoE against a dense baseline, the fair inference-cost comparison is active-parameter count (and FLOPs), not the headline total.
- Budget for the all-to-all. In distributed settings, routing communication can dominate. Expert parallelism, careful expert placement, and interleaving dense layers all help.
- Tune capacity factor deliberately. Watch your token-drop rate. A rising drop rate signals imbalance or too-tight capacity and directly costs quality.
- Use both stabilizers. The load-balancing loss addresses utilization; the router z-loss addresses numerical stability. They solve different problems — most robust recipes use both (or a loss-free balancing scheme plus z-loss).
- Consider shared + fine-grained experts. The DeepSeek-style design of a few always-on shared experts plus many small routed experts is a strong, well-validated default for new sparse architectures.
Mixture-of-Experts is not a free lunch — it trades systems complexity, memory, and training fragility for a dramatically better capacity-per-FLOP curve. But that trade has proven decisive at scale, which is why sparse routing now sits at the heart of so many frontier language models.