On diffusion models
This is part #14 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on Generative Models, part 2 by Justin Johnson. Like the lecture, we start with rectified flow, then connect it to diffusion and score-based models.
In the previous notes, we gave a GAN random noise and asked it to produce an image in one pass. We needed another network to learn whether its output looked real.
What if we already knew the answer to the training problem? Take an image, mix it with noise, and compute the direction between them. We can train a network to predict that direction.
Then, starting from noise, we apply the prediction a little at a time. The same network runs at every step. This is how a regression problem becomes an image generator.
A path between data and noise
Let's call our clean image , written as a vector. We sample independent Gaussian noise with the same shape. Each noise coordinate has mean 0 and variance 1.
For a time , construct:
At , we have the image. At , we have only noise. Between them is a straight line in the space of pixel values.
Here time goes from data to noise. We will generate in the opposite direction. If a paper reverses these endpoints, the velocity sign changes too.
For a fixed image and noise pair, the derivative is:
We get one velocity value per pixel coordinate. It tells us how the pixel value changes as we add noise, rather than how an object moves in the scene.
Rectified flow uses these straight training paths to learn a vector field. Given a point and its time, the network predicts a velocity:
We give the network and . The clean image and the sampled noise stay on the target side of the training problem.
One coordinate, with numbers
Suppose , , and :
Increasing time by 0.1 moves the coordinate down by 0.3. Decreasing time by 0.1 moves it up by 0.3. The same velocity describes both directions; the sign of the time step tells us which way to go.
Learning the velocity
We draw an image, noise, and a random time independently. Then minimize squared error:
The expectation averages over these draws. In practice we use a mini-batch and update through backpropagation, as in optimization.
Notice that we can compute directly from the image and noise. We don't have to run through times 0.1, 0.2, and so on. Training needs one sampled time and one network prediction for this example.
Flow matching connects these per-example targets to a field that moves the whole distribution. We are using straight paths here, but the construction allows other paths too.
Why doesn't it just predict an average image?
Different image and noise pairs can give the same . How can the network choose a direction if it doesn't know which pair we used?
For squared error, the optimal prediction is the conditional mean velocity:
The noisy input and time tell us which pairs to average over. We keep the pairs that could explain this input.
At full noise, the image and input are independent. If the mean training image is , the optimal endpoint prediction is . At the clean endpoint, zero-mean noise gives .
At intermediate times, some of the scene is visible. The model has to use that evidence to infer which directions make sense.
During generation, each move changes the input to the next prediction. Different starting noise samples can therefore follow different paths, even though the prediction at each point is a conditional mean.
The straight training lines can cross. Our learned field gives one velocity at each position and time, so its sampling paths generally bend away from those original lines. This is part of the rectified flow construction.
Sampling: apply the field repeatedly
Once training is finished, freeze the weights. Start with and solve the ordinary differential equation:
We solve it backward, from time 1 to time 0. With a positive step size , the simplest numerical method is an Euler step:
For equal steps, . Evaluate at and finish at 0. There is no extra step after reaching 0.
For an arithmetic example, suppose the model predicts at each queried point. Starting at with :
| Time | Current value | Next value |
|---|---|---|
| 1 | ||
| 0.75 | ||
| 0.5 | ||
| 0.25 |
Here the velocity stayed constant, so Euler was exact. A learned field changes along the path. After a finite step, we have to recompute the velocity, and the numerical solution has some error. We also won't generally recover a particular training image from random noise.
Smaller steps reduce numerical error. A higher-order solver uses several velocity predictions to estimate a better step. We still depend on the learned field being correct. For speed comparisons, count network evaluations: one solver step may call the network more than once.
Once we fix the starting noise and settings, this flow sampler is deterministic. Below we will also see a diffusion sampler that adds fresh noise during generation.
Conditioning and classifier-free guidance
To ask for a particular image, let's add a condition to the network:
might be a class, text, or another image. Training uses matched image–condition pairs; the velocity target stays the same.
We can train the same network to work with and without . For some training examples, replace it with an empty condition . This gives us the two predictions used by classifier-free guidance, or CFG. The CFG paper combines diffusion predictions this way; here we combine velocities.
At sampling time, compute both at the same :
With this convention, is unconditional, is ordinary conditional sampling, and extends past the conditional prediction. The lecture writes , which is identical when .
For example, let , , and . Then . A backward step of size 0.1 now changes the current value by . This is extrapolation, not averaging two predictions.
Increasing pushes harder toward the condition. Too much guidance can reduce variety and produce artifacts. We also need two predictions per step, although we can batch them. The name comes from doing this without a separate classifier.
What does the schedule change?
There are three different choices here:
- The path sets the amount of signal and noise in .
- The training time distribution sets which times we train on most often.
- The sampling grid sets where we call the network during generation.
For rectified flow, our path remains . We can change the other two choices without changing that equation.
Uniform training gives equal weight to equal intervals of time. Esser et al. study alternatives that spend more training effort at intermediate noise levels. One is logit-normal sampling:
Here and control the time distribution, not the image noise itself. With , reducing concentrates times near 0.5. Increasing shifts them toward the noise endpoint.
If a time occurs more often, its errors contribute more often to training. This changes the loss weighting unless we correct for the sampling probabilities.
Work in a smaller space
Now consider the cost of running this network over every pixel, many times per image. Can we do the repeated work on a smaller representation?
Latent diffusion starts with an image encoder and decoder :
The latent is a smaller spatial array. We train the autoencoder first, freeze it, and then train the generative model on these latents. To generate, we start with latent noise, repeatedly update it, and decode the result once. The encoder is only needed to turn training images into latents. Rombach et al., 2022.
For a hypothetical encoder with eightfold spatial reduction and four latent channels:
We have reduced 196,608 scalar values to 4,096, or 48 times fewer. The actual speedup depends on the network we run over that array.
The compression must preserve useful visual detail. The original LDM work combines reconstruction and perceptual objectives with an adversarial loss, and studies both KL-regularized and vector-quantized autoencoders. The learned generative prior then handles the latent distribution.
The autoencoder now limits what we can generate. More sampling steps won't restore information its representation cannot retain.
What network predicts the update?
We have defined the target and sampler. What goes inside ? We need a network that takes the noisy array and time, and returns an update with the same spatial shape.
U-Net
A U-Net first reduces the spatial resolution, then brings it back up. The low-resolution layers can combine information from a larger part of the image. Skip connections carry features directly from each downward level to its matching upward level.
Diffusion U-Nets use residual blocks, add the time embedding, and often include attention. The DDPM paper describes one such network. Depending on the training objective, its output can predict noise, clean data, or velocity.
We call this U-Net at every sampling step. Its downward and upward paths belong to the update network. The image autoencoder above is a separate pair of networks, used before training the generative model and after sampling.
Diffusion Transformer
We can also use a transformer. A DiT splits the noisy latent into patches and projects each patch into a token. Transformer blocks mix the tokens, then an output projection restores the patch values. We arrange them back into the grid.
The DiT paper compares ways to pass time and class information into the blocks. Adaptive layer normalization, or adaLN, uses those embeddings to control shifts and scales. adaLN-Zero also learns gates on the residual branches, initialized at zero.
The original DiT also predicts reverse-process variance. The diagram leaves that extra output out to show the main prediction path.
For our latent, patches give tokens. Each patch starts with values, which a learned projection maps to the model's token width.
With patches we would have 1,024 tokens. Four times as many tokens gives sixteen times as many entries in the dense self-attention matrix. So the patch size changes both the spatial representation and its cost.
For text conditioning, we can use cross-attention, with image features querying text features. Other designs process image and text tokens together. In either case, the update depends on both the noisy input and what we asked to generate.
For video, the latent has a time axis as well. Patches and attention must account for relationships across frames. Here video time and diffusion time are different axes: one describes the clip, the other describes noise level during generation.
Fewer steps through distillation
So far, speeding up the sampler meant changing how we use the same network. With distillation, we train a new network to do more in each pass.
In progressive distillation, a student learns to reproduce two deterministic teacher steps with one student step. Repeating this procedure can reduce the required step count further. A 32-step teacher could become a 16-step student, then an 8-step student; each reduction requires training.
Consistency models take another view: points along the same sampling trajectory should map to the same clean endpoint. They can be trained through distillation or directly from data.
The student has to approximate a larger part of the path at once. Reducing an ordinary sampler to one step skips that training, so we should not expect the same result.
Connecting this to diffusion notation
Now let's connect the straight-path construction to the diffusion notation we see in other papers.
DDPM: a stochastic noising process
A Denoising Diffusion Probabilistic Model defines a sequence that adds fresh Gaussian noise at each forward step. Let the integer run from clean data at 0 to almost pure noise at :
sets the noise variance added at step . Combining the Gaussian transitions gives a direct sample at any chosen time:
Again, we can sample one time directly for training. A common objective asks to predict the noise we added, using squared error. The DDPM paper derives this loss from a reweighted variational bound.
For , our earlier values and give . If the predicted noise is exactly , we recover the clean estimate:
This gives us an estimate of the clean image. To take one reverse step, a stochastic sampler instead computes a less noisy mean and adds some fresh noise:
Here is the fresh noise, and is the reverse-process standard deviation. We omit that noise on the final step. This gives the ancestral sampler.
Noise, clean data, velocity, and score
More generally, write a noisy input as . We can choose to predict clean data, noise, or a linear combination of them. With nonzero , a noise estimate gives a clean estimate through .
For a variance-preserving path, . Its common v-prediction target is:
It gives . This is velocity with respect to an angular parameter, as derived in Salimans and Ho. With , , , and , this target is , while our rectified-flow target was . Changing targets without changing the conversion and sampler gives the wrong update.
The score is another representation:
It points in the local direction of increasing log density of the noisy distribution at time . Under Gaussian corruption, the optimal noise predictor and score satisfy:
So a noise prediction also gives us a score estimate. We can use it in the dynamics of score-based generation.
In continuous time, diffusion can be described by a stochastic differential equation. Its reverse process uses the score to undo the spread caused by noise. A related probability-flow ODE has the same time marginals when the score is exact, while its individual trajectories are deterministic. A probability distribution can thus be sampled through either stochastic or deterministic dynamics; the path equations and solver must agree.
When reading a model's equations, first locate the clean and noisy endpoints. Then check what the network predicts and how the sampler converts that prediction into an update. Rectified flow, DDPM noise prediction, and VP velocity prediction use different conversions, even when the network shapes look the same.
Next, 3D vision: how images relate to cameras, depth, and the structure of a scene.