A sparse MoE layer replaces one dense feed-forward block with several expert blocks and a router. The router consumes the current token representation at that layer and selects a subset of experts. Their weighted outputs return an update to the residual stream. Routing can change at the next layer or position; it does not choose a whole model for the entire prompt.
Route a current representation
Four toy experts, two dimensions, one MoE layer. x is the hidden representation of an existing token. Fixed weights are illustrative and have no assigned subject areas. A future token is not passed into the router.
E0
logit = [1.000, 0.200] · x = 0.720
softmax: 0.395
Mixture weight: 0.500
[0.617, -0.508]
E1
logit = [-0.500, 1.000] · x = -0.800
softmax: 0.086
Mixture weight: 0.000
Not evaluated
E2
logit = [0.400, -1.000] · x = 0.720
softmax: 0.395
Mixture weight: 0.500
[0.565, 0.000]
E3
logit = [-0.700, -0.300] · x = -0.440
softmax: 0.124
Mixture weight: 0.000
Not evaluated
Here the selected top-k softmax scores are renormalized to sum to 1. Only selected experts calculate tanh(Mᵢx). Their weighted sum gives the FFN update for this position. Other MoE recipes use different router normalization.
Σᵢ gᵢ Eᵢ(x) = [0.591, -0.254]
Inspect fixed expert matrices Mᵢ
[1.000, 0.200]
[-0.300, 0.800]
[-0.500, 0.700]
[1.000, 0.200]
[0.400, -0.800]
[0.500, 1.000]
[0.800, 0.300]
[0.100, -0.700]
Routing and specialization
Top-k selection keeps only a few experts active for each token. A router can be trained with auxiliary balancing losses or other strategies, such as bias-based balancing in DeepSeek-V3. Experts may develop patterns of specialization, but they are not reliably labelled code, facts or grammar modules. Inspect measured routing before making such claims.
Available weights are not the same as GPU-resident weights
All experts must be available somewhere. They do not all have to reside in GPU memory: implementations can keep some expert weights on the CPU or move data between devices. Full GPU residency can avoid transfers, while offloading trades memory placement against bandwidth and latency. Active parameters alone therefore do not determine either total memory or end-to-end speed.
Top-k does not make every cost constant
With fixed expert size and k, the selected expert arithmetic can stay similar as the number of experts grows. Router work, weight storage, device communication and load imbalance can still grow. Batch size and token distribution matter because different tokens can activate different experts.
A concrete reference: Mixtral 8x7B
The Mixtral paper (2024) describes eight feed-forward experts per layer, with two selected per token. It reports 46.7B total parameters and 12.9B active parameters per token. These are properties of that architecture, not a general formula for every model with eight experts.