Image Diffusion

Intermediate

Understand the latent diffusion pipeline, architecture choices, conditioning, and sampling tradeoffs.

Last updated: Sep 13, 2026

Latent Diffusion Pipeline

Latent diffusion performs denoising in a compressed representation. Training, text-to-image sampling and image-to-image input use different entry points.

Text-to-image inference starts with sampled latent noise. A VAE encoder is used when preparing image latents for training or image-to-image input; the VAE decoder maps the final generated latent to pixels.

  1. 1. Text → conditioning
  2. 2. Seed → latent noise
  3. 3. Denoiser + scheduler → updated latent
  4. 4. VAE decoder → pixels

U-Net vs DiT Backbones

U-Nets dominate early diffusion models with strong inductive biases for spatial detail. DiTs replace convolutions with transformer blocks and scale efficiently with data and compute.

U-Net

Convolutional encoder-decoder with skip connections for multi-scale spatial reconstruction.

DiT (Diffusion Transformer)

Transformer backbone over patchified latents, often strong at scale with large training budgets.

Text Conditioning (CLIP/T5 + Cross-Attn)

Prompts are encoded (for example by CLIP or T5) and injected into denoising layers via cross-attention, aligning generated content with text semantics.

Classifier-Free Guidance (CFG)

CFG blends conditional and unconditional predictions. Higher guidance strengthens prompt fidelity, but too high can reduce diversity and introduce artifacts.

Sampling Steps vs Quality Tradeoff

More denoising steps often improve fidelity but increase latency. Practical deployments tune step count, scheduler, and guidance for target quality-per-second.

Compute classifier-free guidance

Combine two supplied noise predictions and inspect how the guidance scale changes the resulting vector.

This is a dependency diagram, not a generated image. Keep seed, prompt, resolution, model and scheduler fixed when comparing guidance or step counts. More steps need not monotonically improve quality.

εᵤ = [0.2, −0.4] · ε꜀ = [0.6, −0.1]
ε = εᵤ + s(ε꜀ − εᵤ)
ε = [0.60, -0.10]

Computed guidance example using supplied noise vectors. At s=0 the result is unconditional, at s=1 it is the conditional prediction; larger s extrapolates beyond it. The vectors are illustrative, not model outputs or quality scores. The scheduler must still turn a prediction into the next sample.

Classifier-Free Diffusion Guidance (2022)

Key Takeaways

  • Text-to-image sampling starts from latent noise; VAE encoding is used for image inputs and preparing training latents.
  • U-Net and DiT represent different inductive bias and scaling tradeoffs.
  • Text conditioning and CFG control prompt alignment strength.
  • Image quality depends on the joint tuning of steps, scheduler, and guidance.