Discrete vs Continuous Formulations
Language is fundamentally discrete, so text diffusion often uses token masking/replacement processes. Some approaches embed tokens into continuous spaces and diffuse there before projecting back.
Mask-and-Predict Paradigm
Text diffusion commonly begins with heavily masked sequences and repeatedly predicts missing tokens. Confidence-based remasking/refinement can improve coherence over multiple passes.
MDLM-style models
Masked diffusion language models denoise token grids through iterative unmasking rather than left-to-right decoding.
SEDD-style models
Score-entropy variants adapt score-based ideas to discrete vocabularies with principled probabilistic objectives.
Padding Tokens and Fixed Length
Fixed-length batches can use PAD positions and an attention mask to exclude unused positions. Exact masking rules depend on the architecture and implementation.
How It Differs from Autoregressive LMs
Autoregressive models predict the next token conditioned on previous tokens. Diffusion-style text models refine many positions in parallel over multiple denoising iterations.
Diffusion-style decoding
Parallel token refinement, repeated denoising steps, and optional remasking for error correction.
Autoregressive decoding
Strict left-to-right generation with causal dependency and one-pass token commitment.
Mask Refinement Demo
Compare supplied candidate probabilities, commit the highest-confidence masked positions and keep PAD positions excluded.
Confidence-based masking: a worked sampler
A supplied candidate table stands in for a denoiser. Each step recomputes the table from the visible context and commits the two most confident masked positions. Confidence drives the selection. PAD positions are outside the sequence and never participate.
| Position | Candidate | Candidate probability |
|---|---|---|
| 1 | The | 92% |
| 5 | the | 91% |
| 4 | on | 88% |
| 3 | sits | 62% |
| 8 | quietly | 60% |
| 2 | cat | 59% |
| 6 | warm | 56% |
| 7 | windowsill | 53% |
This example uses irreversible unmasking. Other samplers may revise tokens. These probabilities are specified teaching inputs, not output from a language model.
Masked Diffusion Language Models (2024)Key Takeaways
- Text diffusion adapts denoising to discrete token spaces.
- Mask-and-predict enables parallel refinement instead of strict left-to-right decoding.
- [PAD] plus attention masks are essential for fixed-length batching.
- Compared to autoregressive LMs, text diffusion trades extra steps for iterative correction.