On diffusion models

This is part #14 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes are based on Generative Models, part 2 by Justin Johnson. Like the lecture, we start with rectified flow, then connect it to diffusion and score-based models.


In the previous notes, we gave a GAN random noise and asked it to produce an image in one pass. We needed another network to learn whether its output looked real.

What if we already knew the answer to the training problem? Take an image, mix it with noise, and compute the direction between them. We can train a network to predict that direction.

Then, starting from noise, we apply the prediction a little at a time. The same network runs at every step. This is how a regression problem becomes an image generator.

A path between data and noise

Let's call our clean image x0∈RDx_0\in\mathbb{R}^D, written as a vector. We sample independent Gaussian noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I) with the same shape. Each noise coordinate has mean 0 and variance 1.

For a time t∈[0,1]t\in[0,1], construct:

xt=(1−t)x0+tϵx_t=(1-t)x_0+t\epsilon

At t=0t=0, we have the image. At t=1t=1, we have only noise. Between them is a straight line in the space of pixel values.

Here time goes from data to noise. We will generate in the opposite direction. If a paper reverses these endpoints, the velocity sign changes too.

For a fixed image and noise pair, the derivative is:

v=dxtdt=ϵ−x0v=\frac{dx_t}{dt}=\epsilon-x_0

We get one velocity value per pixel coordinate. It tells us how the pixel value changes as we add noise, rather than how an object moves in the scene.

Rectified flow uses these straight training paths to learn a vector field. Given a point and its time, the network predicts a velocity:

vθ(xt,t)∈RDv_\theta(x_t,t)\in\mathbb{R}^D

We give the network xtx_t and tt. The clean image and the sampled noise stay on the target side of the training problem.

One coordinate, with numbers

Suppose x0=2x_0=2, ϵ=−1\epsilon=-1, and t=0.25t=0.25:

x0.25=0.75(2)+0.25(−1)=1.25v=−1−2=−3\begin{aligned} x_{0.25}&=0.75(2)+0.25(-1)=1.25\\ v&=-1-2=-3 \end{aligned}

Increasing time by 0.1 moves the coordinate down by 0.3. Decreasing time by 0.1 moves it up by 0.3. The same velocity describes both directions; the sign of the time step tells us which way to go.

Forward noise path and backward sampling direction For the chosen scalar pair, clean data is 2 at time 0 and noise is minus 1 at time 1. The straight path decreases by 3 per unit time. At time 0.25 its value is 1.25. A forward arrow points down the line toward noise; a backward arrow points up toward data. Coordinate value 2 1.25 −1 0 0.25 1 Time t Clean data Noise Forward: Δt > 0 Velocity = −3 Sample backward Δt < 0, so value rises
One chosen image–noise pair. The straight path supplies training targets; a learned sampling path need not follow this same line.

Learning the velocity

We draw an image, noise, and a random time independently. Then minimize squared error:

t∼Uniform⁡(0,1)L(θ)=Ex0,ϵ,t[∥vθ(xt,t)−(ϵ−x0)∥22]\begin{aligned} t&\sim\operatorname{Uniform}(0,1)\\ \mathcal{L}(\theta) &=\mathbb{E}_{x_0,\epsilon,t} \left[\left\|v_\theta(x_t,t)-(\epsilon-x_0)\right\|_2^2\right] \end{aligned}

The expectation averages over these draws. In practice we use a mini-batch and update θ\theta through backpropagation, as in optimization.

Notice that we can compute x0.7x_{0.7} directly from the image and noise. We don't have to run through times 0.1, 0.2, and so on. Training needs one sampled time and one network prediction for this example.

Flow matching connects these per-example targets to a field that moves the whole distribution. We are using straight paths here, but the construction allows other paths too.

Why doesn't it just predict an average image?

Different image and noise pairs can give the same xtx_t. How can the network choose a direction if it doesn't know which pair we used?

For squared error, the optimal prediction is the conditional mean velocity:

v∗(xt,t)=E[ϵ−x0∣xt,t]v^*(x_t,t)=\mathbb{E}[\epsilon-x_0\mid x_t,t]

The noisy input and time tell us which pairs to average over. We keep the pairs that could explain this input.

At full noise, the image and input are independent. If the mean training image is mm, the optimal endpoint prediction is v∗(ϵ,1)=ϵ−mv^*(\epsilon,1)=\epsilon-m. At the clean endpoint, zero-mean noise gives v∗(x0,0)=−x0v^*(x_0,0)=-x_0.

At intermediate times, some of the scene is visible. The model has to use that evidence to infer which directions make sense.

During generation, each move changes the input to the next prediction. Different starting noise samples can therefore follow different paths, even though the prediction at each point is a conditional mean.

The straight training lines can cross. Our learned field gives one velocity at each position and time, so its sampling paths generally bend away from those original lines. This is part of the rectified flow construction.

Sampling: apply the field repeatedly

Once training is finished, freeze the weights. Start with x1∼N(0,I)x_1\sim\mathcal{N}(0,I) and solve the ordinary differential equation:

dxtdt=vθ(xt,t)\frac{dx_t}{dt}=v_\theta(x_t,t)

We solve it backward, from time 1 to time 0. With a positive step size hh, the simplest numerical method is an Euler step:

xt−h=xt−h vθ(xt,t)x_{t-h}=x_t-h\,v_\theta(x_t,t)

For NN equal steps, h=1/Nh=1/N. Evaluate at t=1,1−h,…,ht=1,1-h,\ldots,h and finish at 0. There is no extra step after reaching 0.

For an arithmetic example, suppose the model predicts −3-3 at each queried point. Starting at −1-1 with h=0.25h=0.25:

TimeCurrent valueNext value
1−1-1−1−0.25(−3)=−0.25-1-0.25(-3)=-0.25
0.75−0.25-0.250.50.5
0.50.50.51.251.25
0.251.251.2522

Here the velocity stayed constant, so Euler was exact. A learned field changes along the path. After a finite step, we have to recompute the velocity, and the numerical solution has some error. We also won't generally recover a particular training image from random noise.

Smaller steps reduce numerical error. A higher-order solver uses several velocity predictions to estimate a better step. We still depend on the learned field being correct. For speed comparisons, count network evaluations: one solver step may call the network more than once.

Once we fix the starting noise and settings, this flow sampler is deterministic. Below we will also see a diffusion sampler that adds fresh noise during generation.

Conditioning and classifier-free guidance

To ask for a particular image, let's add a condition yy to the network:

vθ(xt,t,y)v_\theta(x_t,t,y)

yy might be a class, text, or another image. Training uses matched image–condition pairs; the velocity target stays the same.

We can train the same network to work with and without yy. For some training examples, replace it with an empty condition ∅\varnothing. This gives us the two predictions used by classifier-free guidance, or CFG. The CFG paper combines diffusion predictions this way; here we combine velocities.

At sampling time, compute both at the same xt,tx_t,t:

vc=vθ(xt,t,y)vu=vθ(xt,t,∅)vcfg=vu+s(vc−vu)\begin{aligned} v_c&=v_\theta(x_t,t,y)\\ v_u&=v_\theta(x_t,t,\varnothing)\\ v_{\mathrm{cfg}}&=v_u+s(v_c-v_u) \end{aligned}

With this convention, s=0s=0 is unconditional, s=1s=1 is ordinary conditional sampling, and s>1s>1 extends past the conditional prediction. The lecture writes (1+w)vc−wvu(1+w)v_c-wv_u, which is identical when s=1+ws=1+w.

For example, let vu=2v_u=2, vc=1v_c=1, and s=3s=3. Then vcfg=2+3(1−2)=−1v_{\mathrm{cfg}}=2+3(1-2)=-1. A backward step of size 0.1 now changes the current value by +0.1+0.1. This is extrapolation, not averaging two predictions.

Increasing ss pushes harder toward the condition. Too much guidance can reduce variety and produce artifacts. We also need two predictions per step, although we can batch them. The name comes from doing this without a separate classifier.

What does the schedule change?

There are three different choices here:

  1. The path sets the amount of signal and noise in xtx_t.
  2. The training time distribution sets which times we train on most often.
  3. The sampling grid sets where we call the network during generation.

For rectified flow, our path remains (1−t)x0+tϵ(1-t)x_0+t\epsilon. We can change the other two choices without changing that equation.

Uniform training gives equal weight to equal intervals of time. Esser et al. study alternatives that spend more training effort at intermediate noise levels. One is logit-normal sampling:

u∼N(μ,σ2),t=11+e−uu\sim\mathcal{N}(\mu,\sigma^2), \qquad t=\frac{1}{1+e^{-u}}

Here μ\mu and σ\sigma control the time distribution, not the image noise itself. With μ=0\mu=0, reducing σ\sigma concentrates times near 0.5. Increasing μ\mu shifts them toward the noise endpoint.

If a time occurs more often, its errors contribute more often to training. This changes the loss weighting unless we correct for the sampling probabilities.

Work in a smaller space

Now consider the cost of running this network over every pixel, many times per image. Can we do the repeated work on a smaller representation?

Latent diffusion starts with an image encoder EE and decoder DD:

z0=E(x0),x^0=D(z0)z_0=E(x_0),\qquad \hat{x}_0=D(z_0)

The latent z0z_0 is a smaller spatial array. We train the autoencoder first, freeze it, and then train the generative model on these latents. To generate, we start with latent noise, repeatedly update it, and decode the result once. The encoder is only needed to turn training images into latents. Rombach et al., 2022.

For a hypothetical encoder with eightfold spatial reduction and four latent channels:

256×256×3⟶32×32×4256\times256\times3 \quad\longrightarrow\quad 32\times32\times4

We have reduced 196,608 scalar values to 4,096, or 48 times fewer. The actual speedup depends on the network we run over that array.

The compression must preserve useful visual detail. The original LDM work combines reconstruction and perceptual objectives with an adversarial loss, and studies both KL-regularized and vector-quantized autoencoders. The learned generative prior then handles the latent distribution.

The autoencoder now limits what we can generate. More sampling steps won't restore information its representation cannot retain.

What network predicts the update?

We have defined the target and sampler. What goes inside vθv_\theta? We need a network that takes the noisy array and time, and returns an update with the same spatial shape.

U-Net

A U-Net first reduces the spatial resolution, then brings it back up. The low-resolution layers can combine information from a larger part of the image. Skip connections carry features directly from each downward level to its matching upward level.

U-Net with time conditioning and skip connections A noisy image or latent enters residual blocks. The downward path reduces spatial resolution twice, then the upward path restores it. Skip connections carry features between matching resolutions. A time embedding conditions the residual blocks. The output has the input's spatial shape and predicts the quantity used by the sampler. Noisy input xₜPrediction H × WH × W Res. blocksH × W Res. blocksH × W Res. blocksH/2 × W/2 Res. blocksH/2 × W/2 Middle blocksH/4 × W/4 skipskip ↓ downup ↑ t → embedding → residual blocks
A schematic U-Net with two reductions in resolution. Skip connections preserve features from the downward path. The time embedding conditions the residual blocks.

Diffusion U-Nets use residual blocks, add the time embedding, and often include attention. The DDPM paper describes one such network. Depending on the training objective, its output can predict noise, clean data, or velocity.

We call this U-Net at every sampling step. Its downward and upward paths belong to the update network. The image autoencoder above is a separate pair of networks, used before training the generative model and after sampling.

Diffusion Transformer

We can also use a transformer. A DiT splits the noisy latent into patches and projects each patch into a token. Transformer blocks mix the tokens, then an output projection restores the patch values. We arrange them back into the grid.

Diffusion Transformer over noisy latent patches The noisy latent is split into patches, then each patch is projected into a token with a positional embedding. Repeated transformer blocks use self-attention and MLPs. Time and class embeddings condition their shifts, scales, and gates. Tokens are projected to patch values and arranged back into the spatial prediction. The diagram omits the optional variance output. Noisy latent Split into patchesProject each patch Tokens + positions Transformer blocks Self-attention + MLP Time + classembeddings adaLN-Zero:shift, scale, gate Project patch outputs Arrange into a grid Noise or velocity prediction
DiT replaces the U-Net with a transformer over patches. The same network is called at each noise level. This diagram shows time and class conditioning through adaLN-Zero.

The DiT paper compares ways to pass time and class information into the blocks. Adaptive layer normalization, or adaLN, uses those embeddings to control shifts and scales. adaLN-Zero also learns gates on the residual branches, initialized at zero.

The original DiT also predicts reverse-process variance. The diagram leaves that extra output out to show the main prediction path.

For our 32×32×432\times32\times4 latent, 2×22\times2 patches give 16×16=25616\times16=256 tokens. Each patch starts with 2×2×4=162\times2\times4=16 values, which a learned projection maps to the model's token width.

With 1×11\times1 patches we would have 1,024 tokens. Four times as many tokens gives sixteen times as many entries in the dense self-attention matrix. So the patch size changes both the spatial representation and its cost.

For text conditioning, we can use cross-attention, with image features querying text features. Other designs process image and text tokens together. In either case, the update depends on both the noisy input and what we asked to generate.

For video, the latent has a time axis as well. Patches and attention must account for relationships across frames. Here video time and diffusion time are different axes: one describes the clip, the other describes noise level during generation.

Fewer steps through distillation

So far, speeding up the sampler meant changing how we use the same network. With distillation, we train a new network to do more in each pass.

In progressive distillation, a student learns to reproduce two deterministic teacher steps with one student step. Repeating this procedure can reduce the required step count further. A 32-step teacher could become a 16-step student, then an 8-step student; each reduction requires training.

Consistency models take another view: points along the same sampling trajectory should map to the same clean endpoint. They can be trained through distillation or directly from data.

The student has to approximate a larger part of the path at once. Reducing an ordinary sampler to one step skips that training, so we should not expect the same result.

Connecting this to diffusion notation

Now let's connect the straight-path construction to the diffusion notation we see in other papers.

DDPM: a stochastic noising process

A Denoising Diffusion Probabilistic Model defines a sequence that adds fresh Gaussian noise at each forward step. Let the integer kk run from clean data at 0 to almost pure noise at KK:

q(xk∣xk−1)=N(αkxk−1,βkI)αk=1−βk,αˉk=∏j=1kαj\begin{aligned} q(x_k\mid x_{k-1}) &=\mathcal{N}(\sqrt{\alpha_k}x_{k-1},\beta_k I)\\ \alpha_k&=1-\beta_k,\qquad \bar{\alpha}_k=\prod_{j=1}^k\alpha_j \end{aligned}

0<βk<10<\beta_k<1 sets the noise variance added at step kk. Combining the Gaussian transitions gives a direct sample at any chosen time:

xk=αˉkx0+1−αˉk ϵx_k=\sqrt{\bar{\alpha}_k}x_0 +\sqrt{1-\bar{\alpha}_k}\,\epsilon

Again, we can sample one time directly for training. A common objective asks ϵθ(xk,k)\epsilon_\theta(x_k,k) to predict the noise ϵ\epsilon we added, using squared error. The DDPM paper derives this loss from a reweighted variational bound.

For αˉk=0.64\bar{\alpha}_k=0.64, our earlier values x0=2x_0=2 and ϵ=−1\epsilon=-1 give xk=0.8(2)+0.6(−1)=1x_k=0.8(2)+0.6(-1)=1. If the predicted noise is exactly −1-1, we recover the clean estimate:

x^0=xk−1−αˉk ϵ^αˉk=2\hat{x}_0 =\frac{x_k-\sqrt{1-\bar{\alpha}_k}\,\hat\epsilon} {\sqrt{\bar{\alpha}_k}} =2

This gives us an estimate of the clean image. To take one reverse step, a stochastic sampler instead computes a less noisy mean and adds some fresh noise:

μθ(xk,k)=1αk(xk−βk1−αˉkϵθ(xk,k))\begin{aligned} &\mu_\theta(x_k,k)\\ &=\frac{1}{\sqrt{\alpha_k}} \left(x_k-\frac{\beta_k}{\sqrt{1-\bar{\alpha}_k}} \epsilon_\theta(x_k,k)\right) \end{aligned} xk−1=μθ(xk,k)+σkηx_{k-1}=\mu_\theta(x_k,k)+\sigma_k\eta

Here η∼N(0,I)\eta\sim\mathcal{N}(0,I) is the fresh noise, and σk\sigma_k is the reverse-process standard deviation. We omit that noise on the final step. This gives the ancestral sampler.

Noise, clean data, velocity, and score

More generally, write a noisy input as xt=atx0+btϵx_t=a_t x_0+b_t\epsilon. We can choose to predict clean data, noise, or a linear combination of them. With nonzero at,bta_t,b_t, a noise estimate gives a clean estimate through x^0=(xt−btϵ^)/at\hat{x}_0=(x_t-b_t\hat\epsilon)/a_t.

For a variance-preserving path, at2+bt2=1a_t^2+b_t^2=1. Its common v-prediction target is:

vVP=atϵ−btx0v_{\mathrm{VP}}=a_t\epsilon-b_t x_0

It gives x^0=atxt−btv^VP\hat{x}_0=a_t x_t-b_t\hat v_{\mathrm{VP}}. This is velocity with respect to an angular parameter, as derived in Salimans and Ho. With at=0.8a_t=0.8, bt=0.6b_t=0.6, x0=2x_0=2, and ϵ=−1\epsilon=-1, this target is −2-2, while our rectified-flow target was −3-3. Changing targets without changing the conversion and sampler gives the wrong update.

The score is another representation:

st(x)=∇xlog⁡pt(x)s_t(x)=\nabla_x\log p_t(x)

It points in the local direction of increasing log density of the noisy distribution at time tt. Under Gaussian corruption, the optimal noise predictor and score satisfy:

st(xt)=−E[ϵ∣xt]bt(bt>0)s_t(x_t)=-\frac{\mathbb{E}[\epsilon\mid x_t]}{b_t} \qquad (b_t>0)

So a noise prediction also gives us a score estimate. We can use it in the dynamics of score-based generation.

In continuous time, diffusion can be described by a stochastic differential equation. Its reverse process uses the score to undo the spread caused by noise. A related probability-flow ODE has the same time marginals when the score is exact, while its individual trajectories are deterministic. A probability distribution can thus be sampled through either stochastic or deterministic dynamics; the path equations and solver must agree.

When reading a model's equations, first locate the clean and noisy endpoints. Then check what the network predicts and how the sampler converts that prediction into an update. Rectified flow, DDPM noise prediction, and VP velocity prediction use different conversions, even when the network shapes look the same.

Next, 3D vision: how images relate to cameras, depth, and the structure of a scene.

← Back to blog