How diffusion models generate images
Image generators like Stable Diffusion learn one skill: removing a little noise from an image. Run that skill many times, starting from pure noise, and a new image appears.
- 1
Add noise, step by step
During training, real images are corrupted with a little random noise at each step, until nothing of the image is left.
- 2
Learn to predict the noise
A neural network sees a noisy image and the step number, and predicts the noise that was added. The loss measures how far off the prediction is.
- 3
Start from pure noise
To generate, start from an image of random noise. There is no picture hidden in it yet.
- 4
Remove noise a little at a time
At each step the network predicts the noise and removes some of it. After a few dozen steps, a clean image is left.
- 5
Guide it with text
A text encoder turns the prompt into numbers the denoiser reads at every step, steering it toward an image that matches the words.
Add noise, step by step
During training, real images are corrupted with a little random noise at each step, until nothing of the image is left.
Learn to predict the noise
A neural network sees a noisy image and the step number, and predicts the noise that was added. The loss measures how far off the prediction is.
Start from pure noise
To generate, start from an image of random noise. There is no picture hidden in it yet.
Remove noise a little at a time
At each step the network predicts the noise and removes some of it. After a few dozen steps, a clean image is left.
Guide it with text
A text encoder turns the prompt into numbers the denoiser reads at every step, steering it toward an image that matches the words.
In short
- The model never stores images to copy. It learns how to turn noise into images that look like its training data.
- Most modern generators run the process in a compressed latent space, then decode to pixels at the end. That is far cheaper.
- Fewer denoising steps means faster images. Samplers trade steps against quality.
- The image and the noise in the animation are drawn by hand, not produced by a real model.