Juan Pablo García
All writing
Explainer5 steps · scroll to play

How diffusion models generate images

Image generators like Stable Diffusion learn one skill: removing a little noise from an image. Run that skill many times, starting from pure noise, and a new image appears.

  1. 1

    Add noise, step by step

    During training, real images are corrupted with a little random noise at each step, until nothing of the image is left.

  2. 2

    Learn to predict the noise

    A neural network sees a noisy image and the step number, and predicts the noise that was added. The loss measures how far off the prediction is.

  3. 3

    Start from pure noise

    To generate, start from an image of random noise. There is no picture hidden in it yet.

  4. 4

    Remove noise a little at a time

    At each step the network predicts the noise and removes some of it. After a few dozen steps, a clean image is left.

  5. 5

    Guide it with text

    A text encoder turns the prompt into numbers the denoiser reads at every step, steering it toward an image that matches the words.

Add noise, step by step

During training, real images are corrupted with a little random noise at each step, until nothing of the image is left.

Learn to predict the noise

A neural network sees a noisy image and the step number, and predicts the noise that was added. The loss measures how far off the prediction is.

Start from pure noise

To generate, start from an image of random noise. There is no picture hidden in it yet.

Remove noise a little at a time

At each step the network predicts the noise and removes some of it. After a few dozen steps, a clean image is left.

Guide it with text

A text encoder turns the prompt into numbers the denoiser reads at every step, steering it toward an image that matches the words.

In short

  • The model never stores images to copy. It learns how to turn noise into images that look like its training data.
  • Most modern generators run the process in a compressed latent space, then decode to pixels at the end. That is far cheaper.
  • Fewer denoising steps means faster images. Samplers trade steps against quality.
  • The image and the noise in the animation are drawn by hand, not produced by a real model.

Have an AI feature to build? Let's talk for 15 minutes.