Forward Process (Corruption)
The forward process gradually injects noise into clean samples until structure is mostly destroyed. This creates training pairs that teach the model what noise was added at each step.
Start from clean data
Sample x0 from real data (image, text embedding, or latent).
Add small noise repeatedly
Each timestep adds controlled Gaussian or discrete corruption.
Reach near-random state
After many steps, xT approximates a simple prior distribution.
Reverse Process (Generation)
Generation starts from noise and applies learned denoising steps. The model predicts a cleaner state (or the noise residual) and iteratively moves toward a realistic sample.
Noise Prediction Intuition
Predicting noise is often easier than directly predicting the clean sample. Once noise is estimated, you can subtract it and recover structure progressively.
Schedulers: Linear vs Cosine
Schedulers define how much noise is added or removed per step. They strongly affect stability, sample quality, and speed.
Linear schedule
Noise changes at a constant rate per step. Simple, predictable, and easy to implement.
Cosine schedule
Allocates denoising effort non-linearly, often preserving useful signal longer and improving perceptual quality.
Score Matching Intuition
A score function estimates the gradient of log-density, pointing toward more likely data regions. Denoising updates follow this direction, nudging noisy states back to the data manifold.
Forward noising with Gaussian noise
Change signal power in the diffusion equation. The reverse direction is explained separately because a learned sampler does not have the original image.
Forward noising, computed
xₜ = √ᾱ·x₀ + √(1−ᾱ)·ε, with a fixed seed and Gaussian ε. The display clips the numeric values into the screen's color range. The same seed lets you compare noise levels.
x₀=1, ε=-1.0392 → xₜ=1.0000Moving the slider back reveals the stored original; this is not a learned reverse sampler. During generation the original is unknown. The model must predict a denoising direction from the noisy state and the conditioning.
DDPM (2020)Key Takeaways
- Diffusion training learns to invert a known corruption process.
- Reverse sampling uses a learned prediction. Its number of steps depends on the sampler and possible distillation.
- Schedulers control where model capacity is spent across timesteps.
- Score estimation provides a geometric view of denoising direction.