Mixture of Experts

Expert

Learned routing among multiple feed-forward networks: what is active per token, what must be stored, and where the costs move.

Last updated: Sep 13, 2026

A sparse MoE layer replaces one dense feed-forward block with several expert blocks and a router. The router consumes the current token representation at that layer and selects a subset of experts. Their weighted outputs return an update to the residual stream. Routing can change at the next layer or position; it does not choose a whole model for the entire prompt.

Route a current representation

Four toy experts, two dimensions, one MoE layer. x is the hidden representation of an existing token. Fixed weights are illustrative and have no assigned subject areas. A future token is not passed into the router.

E0

logit = [1.000, 0.200] · x = 0.720

softmax: 0.395

Mixture weight: 0.500

[0.617, -0.508]

E1

logit = [-0.500, 1.000] · x = -0.800

softmax: 0.086

Mixture weight: 0.000

Not evaluated

E2

logit = [0.400, -1.000] · x = 0.720

softmax: 0.395

Mixture weight: 0.500

[0.565, 0.000]

E3

logit = [-0.700, -0.300] · x = -0.440

softmax: 0.124

Mixture weight: 0.000

Not evaluated

Here the selected top-k softmax scores are renormalized to sum to 1. Only selected experts calculate tanh(Mᵢx). Their weighted sum gives the FFN update for this position. Other MoE recipes use different router normalization.

Σᵢ gᵢ Eᵢ(x) = [0.591, -0.254]

Inspect fixed expert matrices Mᵢ
E0

[1.000, 0.200]

[-0.300, 0.800]

E1

[-0.500, 0.700]

[1.000, 0.200]

E2

[0.400, -0.800]

[0.500, 1.000]

E3

[0.800, 0.300]

[0.100, -0.700]

Routing and specialization

Top-k selection keeps only a few experts active for each token. A router can be trained with auxiliary balancing losses or other strategies, such as bias-based balancing in DeepSeek-V3. Experts may develop patterns of specialization, but they are not reliably labelled code, facts or grammar modules. Inspect measured routing before making such claims.

Available weights are not the same as GPU-resident weights

All experts must be available somewhere. They do not all have to reside in GPU memory: implementations can keep some expert weights on the CPU or move data between devices. Full GPU residency can avoid transfers, while offloading trades memory placement against bandwidth and latency. Active parameters alone therefore do not determine either total memory or end-to-end speed.

Top-k does not make every cost constant

With fixed expert size and k, the selected expert arithmetic can stay similar as the number of experts grows. Router work, weight storage, device communication and load imbalance can still grow. Batch size and token distribution matter because different tokens can activate different experts.

A concrete reference: Mixtral 8x7B

The Mixtral paper (2024) describes eight feed-forward experts per layer, with two selected per token. It reports 46.7B total parameters and 12.9B active parameters per token. These are properties of that architecture, not a general formula for every model with eight experts.

Primary sources

Review the dense feed-forward calculation →