Mega-Kernels

Expert

From fused kernels to mega-kernels: how moving more and more of GPU execution into a single kernel radically cuts launch overhead and HBM traffic.

Last updated: Sep 24, 2026

Fused kernels: gluing operations together

A fused kernel combines several consecutive operations into a single GPU kernel. The classic example: MatMul + Bias + GELU. Instead of three separate kernels exchanging their intermediate results through main memory (HBM), a single kernel computes everything in one pass.

Without fusion
MatMul
→HBM
Bias
→HBM
GELU

3 kernel launches · intermediate results travel through HBM

With fusion
MatMul
→Bias
→GELU
Register / Shared Memory

1 kernel launch · intermediate results stay in registers/shared memory

Less launch overhead

Every kernel launch costs time — CPU scheduling, grid setup, synchronization. Three launches become one. Across many small operations, these savings add up dramatically.

Intermediate results in registers & shared memory

Instead of writing the MatMul result to HBM and immediately reading it back in the next kernel, it stays in registers or shared memory. That saves the most expensive memory accesses — the arithmetic gets no faster, but the memory traffic almost disappears.

Mega-kernels: the whole step in one kernel

A mega-kernel takes this idea to the extreme: not just a few related operations, but a complete transformer or MoE compute step runs in ONE kernel. RMSNorm, QKV projection, RoPE, attention, output projection, MLP — and for MoE even the expert routing — are assembled into a single GPU program.

The mega-kernel path

Every step of a transformer block — inside a single kernel

RMSNorm
Normalize the residual stream
→
QKV
Query, key, and value projection
→
RoPE
Apply rotary positional embeddings
→
Attention
Compute scores, softmax, weighted sum
→
Projection
Output projection back into the residual stream
→
MLP
Up/gate/down projection with activation
→
Routing
Expert selection and dispatch for MoE

Fused kernel vs. mega-kernel

Both cut overhead — but on a completely different scale

Dimension
Fused kernel
Mega-kernel
Scope
A few related operations (e.g. MatMul + Bias + GELU)
A large algorithm section — a whole transformer or MoE step
Kernel launches
Reduced — many launches become one per fused group
Extremely reduced — the entire block needs a single launch
HBM traffic
Intermediate results stay in registers/shared memory instead of HBM
Intermediate stages never leave the kernel — minimal memory traffic
Complexity
Moderate — established optimization, easy to handle
Very high — register pressure, scheduling, and failure modes grow drastically
Register usage
Slightly higher due to additional live values
High — many simultaneously live values limit occupancy
Scheduling
The GPU runtime still schedules each kernel separately
Partly inside the kernel — the kernel itself takes over control flow and synchronization
Examples
Fused RMSNorm, fused Bias+GELU, FlashAttention
Whole transformer path, MoE path with routing in one kernel

Persistent kernels

The logical next step: a kernel that simply never stops.

Start once, run continuously

With a persistent kernel, the CPU launches the kernel exactly once — after that the GPU stays active continuously. Instead of launching a new kernel for every compute step, loops on the streaming multiprocessors keep pulling new work: attention → MLP → routing → attention → … The kernel handles synchronization itself; the CPU stays completely out of the picture.

The attention-MLP-routing loop

The GPU stays active and pulls work in a loop

Attention
→MLP
→Routing
⟲… and back to the start

Especially valuable for MoE

Mixture-of-Experts models produce many small, irregular workloads — individual experts often receive only a handful of tokens. In the classic launch model, every mini workload costs a full kernel launch; the GPU spends more time launching than computing. A persistent kernel picks up these small tasks in its loop and fills exactly the gaps where the GPU would otherwise wait.

The core idea

Kernel fusion = Glue operations together.

Mega-kernel = Move the bulk of GPU execution into a single kernel.

Key takeaways

  • 1Fused kernels glue a few consecutive operations together: fewer kernel launches and intermediate results in registers/shared memory instead of HBM
  • 2A mega-kernel moves a complete transformer or MoE step — RMSNorm, QKV, RoPE, attention, projection, MLP, routing — into a single kernel
  • 3The price: very high complexity, heavy register pressure, and scheduling that partially moves into the kernel
  • 4Persistent kernels start once and keep running in loops — especially valuable for MoE with many small workloads