Fused kernels: gluing operations together
A fused kernel combines several consecutive operations into a single GPU kernel. The classic example: MatMul + Bias + GELU. Instead of three separate kernels exchanging their intermediate results through main memory (HBM), a single kernel computes everything in one pass.
3 kernel launches · intermediate results travel through HBM
1 kernel launch · intermediate results stay in registers/shared memory
Less launch overhead
Every kernel launch costs time — CPU scheduling, grid setup, synchronization. Three launches become one. Across many small operations, these savings add up dramatically.
Intermediate results in registers & shared memory
Instead of writing the MatMul result to HBM and immediately reading it back in the next kernel, it stays in registers or shared memory. That saves the most expensive memory accesses — the arithmetic gets no faster, but the memory traffic almost disappears.
Mega-kernels: the whole step in one kernel
A mega-kernel takes this idea to the extreme: not just a few related operations, but a complete transformer or MoE compute step runs in ONE kernel. RMSNorm, QKV projection, RoPE, attention, output projection, MLP — and for MoE even the expert routing — are assembled into a single GPU program.
The mega-kernel path
Every step of a transformer block — inside a single kernel
Fused kernel vs. mega-kernel
Both cut overhead — but on a completely different scale
Persistent kernels
The logical next step: a kernel that simply never stops.
Start once, run continuously
With a persistent kernel, the CPU launches the kernel exactly once — after that the GPU stays active continuously. Instead of launching a new kernel for every compute step, loops on the streaming multiprocessors keep pulling new work: attention → MLP → routing → attention → … The kernel handles synchronization itself; the CPU stays completely out of the picture.
The attention-MLP-routing loop
The GPU stays active and pulls work in a loop
Especially valuable for MoE
Mixture-of-Experts models produce many small, irregular workloads — individual experts often receive only a handful of tokens. In the classic launch model, every mini workload costs a full kernel launch; the GPU spends more time launching than computing. A persistent kernel picks up these small tasks in its loop and fills exactly the gaps where the GPU would otherwise wait.
The core idea
Kernel fusion = Glue operations together.
Mega-kernel = Move the bulk of GPU execution into a single kernel.
Key takeaways
- 1Fused kernels glue a few consecutive operations together: fewer kernel launches and intermediate results in registers/shared memory instead of HBM
- 2A mega-kernel moves a complete transformer or MoE step — RMSNorm, QKV, RoPE, attention, projection, MLP, routing — into a single kernel
- 3The price: very high complexity, heavy register pressure, and scheduling that partially moves into the kernel
- 4Persistent kernels start once and keep running in loops — especially valuable for MoE with many small workloads