Feed-Forward Networks

Intermediate

How a token vector expands, passes through a nonlinearity and projects back to model width.

Last updated: Sep 13, 2026

Feed-Forward Networks

How a token vector expands, passes through a nonlinearity and projects back to model width.

A token through a feed-forward block

A real 2 → 4 → 2 calculation with fixed toy weights and no biases. The same function is applied independently to each token. This block does not mix tokens.

Expanded vector xWup

[0.650, -0.900, 0.840, -0.190]

After activation / gate

[0.482, -0.166, 0.672, -0.081]

Output update

[-0.160, -0.144]

Inspect the fixed matrices (row-vector convention)

Wup

[1.000, -0.500, 0.800, 0.200]

[0.300, 1.000, -0.400, 0.700]

Wgate

[0.500, 1.000, -1.000, 0.200]

[1.000, -0.200, 0.400, 0.800]

Wdown

[0.500, -0.200]

[0.300, 0.400]

[-0.500, 0.100]

[0.200, 0.600]

GELU uses the common tanh approximation. SwiGLU here computes SiLU(xWgate) ⊙ (xWup), followed by Wdown. Real models learn these matrices; wider activations do not imply named semantic channels.

The token-wise MLP

After attention mixes information across tokens, the feed-forward network processes each token independently. It expands the hidden vector, applies a nonlinearity, then compresses it back.

  • Expansion creates room for many features.
  • The nonlinearity prevents the block from collapsing into one linear map.
  • The compression writes a useful update back to the model width.
  • Many parameters live in these blocks; factual and semantic patterns are distributed across weights and activations, not stored in one neat lookup table.

GELU, SwiGLU, and gates

Modern LLMs often use gated activations such as SwiGLU. One learned pathway controls another, letting the block select which features should pass through.

  • Original transformers used simpler activations such as ReLU.
  • Many GPT-era models used GELU.
  • LLaMA-style models commonly use SwiGLU variants.
Continue: how MoE selects among multiple FFNs per layer →

Key Takeaways

  • FFNs process tokens independently after attention has mixed context.
  • SwiGLU-style gated MLPs are common in modern LLMs.
  • MoE is many MLP experts plus learned routing, not magic subject drawers.