Text Diffusion

Intermediate

Understand how diffusion ideas adapt to discrete language tokens and iterative mask refinement.

Last updated: Sep 13, 2026

Discrete vs Continuous Formulations

Language is fundamentally discrete, so text diffusion often uses token masking/replacement processes. Some approaches embed tokens into continuous spaces and diffuse there before projecting back.

Mask-and-Predict Paradigm

Text diffusion commonly begins with heavily masked sequences and repeatedly predicts missing tokens. Confidence-based remasking/refinement can improve coherence over multiple passes.

MDLM-style models

Masked diffusion language models denoise token grids through iterative unmasking rather than left-to-right decoding.

SEDD-style models

Score-entropy variants adapt score-based ideas to discrete vocabularies with principled probabilistic objectives.

Padding Tokens and Fixed Length

Fixed-length batches can use PAD positions and an attention mask to exclude unused positions. Exact masking rules depend on the architecture and implementation.

How It Differs from Autoregressive LMs

Autoregressive models predict the next token conditioned on previous tokens. Diffusion-style text models refine many positions in parallel over multiple denoising iterations.

Diffusion-style decoding

Parallel token refinement, repeated denoising steps, and optional remasking for error correction.

Autoregressive decoding

Strict left-to-right generation with causal dependency and one-pass token commitment.

Mask Refinement Demo

Compare supplied candidate probabilities, commit the highest-confidence masked positions and keep PAD positions excluded.

Confidence-based masking: a worked sampler

A supplied candidate table stands in for a denoiser. Each step recomputes the table from the visible context and commits the two most confident masked positions. Confidence drives the selection. PAD positions are outside the sequence and never participate.

[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[MASK]
Masked
[PAD]
Padding, excluded
[PAD]
Padding, excluded
PositionCandidateCandidate probability
1The92%
5the91%
4on88%
3sits62%
8quietly60%
2cat59%
6warm56%
7windowsill53%

This example uses irreversible unmasking. Other samplers may revise tokens. These probabilities are specified teaching inputs, not output from a language model.

Masked Diffusion Language Models (2024)

Key Takeaways

  • Text diffusion adapts denoising to discrete token spaces.
  • Mask-and-predict enables parallel refinement instead of strict left-to-right decoding.
  • [PAD] plus attention masks are essential for fixed-length batching.
  • Compared to autoregressive LMs, text diffusion trades extra steps for iterative correction.