On generative models
This is part #13 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on lecture 13: Generative Models by Justin Johnson, and the GAN section of lecture 14.
So far we have taken an image and computed something from it. A class, a box, or a useful representation, as in self-supervised learning.
Now we want to go the other way and produce an image. If we ask for a cup, there isn't one correct output. We could get a different shape, colour, or viewpoint each time. So instead of predicting one answer, we learn a distribution and sample from it.
What are we modelling?
Let be an image and a condition, such as a class or a text description:
- A discriminative model learns : given this image, which label fits?
- An unconditional generative model learns : how are images distributed?
- A conditional generative model learns : which images fit this condition?
As before, is our set of learnable parameters. We don't know the data distribution . We only have examples drawn from it.
For discrete images, the probabilities over all possible images sum to 1. For continuous values, we use a density, whose integral is 1. A density value can exceed 1; probability comes from integrating it over a region.
A taxonomy of generative models
Before looking at the models, let's put them on a map. The lecture starts by asking whether we can compute for a given image.
An explicit density model defines a likelihood. Sometimes we can evaluate it directly. Sometimes the definition contains an integral we cannot compute, so we optimize a bound instead.
An implicit model defines a way to generate samples. For example, take random noise and pass it through a neural network. This gives us a distribution over outputs, even if we cannot evaluate its density at an arbitrary image.
For autoregressive models, we write the likelihood as a product of conditional probabilities. We can evaluate that product exactly for the chosen model. Generating a new image still takes a sequence of predictions.
Normalizing flows also give a tractable likelihood. They use invertible transformations with a tractable Jacobian determinant to turn a simple density into a more complex one. The determinant accounts for how each transformation expands or compresses volume. Real NVP is one example.
For a VAE, the model defines a likelihood through a latent variable, but evaluating it requires an integral. We will derive the lower bound that makes training possible. A GAN instead learns a generator through a discriminator, without evaluating an image likelihood.
Where do diffusion models fit? The lecture places them under indirect sampling because generation uses repeated updates. That is a useful description of the sampler. A DDPM also defines a latent-variable likelihood and trains with a variational bound. So "iterative sampling" and "implicit density" are not the same property.
Keep those two questions separate as we go: what objective can we train, and what do we run to get a sample? Let's start with likelihood.
Maximum likelihood
Given independent training examples , maximum likelihood estimation chooses parameters that make the observed dataset likely:
The log is increasing, so it preserves the optimum while turning a product into a sum. We normally minimize the negative log likelihood, or NLL:
Now we have a loss to minimize through gradient descent. We use natural logarithms, so the loss is measured in nats. See Kingma and Welling for the connection between likelihood and minibatch training.
Let's give two observed outcomes probabilities and . Their joint likelihood is , and the average NLL is about . Change the probabilities to and : the likelihood rises to and the average NLL falls to .
Notice that the first outcome became less likely, yet the loss improved. We optimize the combined likelihood of the examples. We then use held-out data to check whether that improvement generalizes.
Autoregressive models
Write an image as a sequence . The chain rule of probability gives:
Here means all values before position . We haven't assumed that pixels are independent. Each prediction can depend on every earlier value. The chain rule is exact; our network approximates the conditional distributions.
PixelRNN applies this to images. Visit pixels row by row, and colour channels in a fixed order. Each 8-bit channel value becomes a classification problem with 256 possible values. The network predicts a distribution over the next value using only earlier ones.
We have already used this idea for next-token prediction in the RNN notes. We just changed what goes into the sequence.
Training and sampling use different inputs
During training, the whole image is available. To predict , give the model the real prefix and score its distribution against the real target. This is called teacher forcing.
The image loss becomes a sum of ordinary cross-entropies:
We block access to the target and future values with a causal mask. Otherwise, the network could copy the answer. Networks such as MADE, masked CNNs, and causal transformers can evaluate the training conditionals in parallel. An RNN still has to compute its hidden states in sequence.
During generation, there is no real prefix to supply:
- Sample from the first distribution.
- Predict and sample .
- Append the sample and repeat.
Let's use a three-pixel binary image :
Multiplying gives , with an NLL of about nats. Each factor is the probability assigned to the pixel value we observed.
When we generate, the sampled prefix determines what happens next. If the second pixel were 1, we would use for the third. Taking the highest-probability value every time would turn this into greedy decoding.
To condition on a label or an image representation, include it at every step: . Conditional PixelCNN demonstrates this. The model's basic likelihood calculation stays the same.
This becomes slow for images. A RGB image has channel values, and here each one needs the previous samples. We can shorten the sequence by generating compressed image tokens. We then model those tokens and use a decoder to turn them into pixels.
From autoencoders to a generative model
An ordinary autoencoder maps an image to a code and back:
The encoder has parameters and the decoder has parameters . We train them together to reconstruct the input, usually with a bottleneck that forces the code to keep useful information.
But where do we get a new after training? We know how to encode an image, but we wanted to generate one. Random coordinates may fall between the codes the decoder learned to use.
A variational autoencoder, or VAE, starts with that missing sampling step. We choose a simple prior over the latent variable, often:
Then we sample an image from the decoder distribution . is the identity matrix, so each prior coordinate is an independent Gaussian with variance 1. The decoder outputs distribution parameters, such as the mean image.
The resulting image density is:
To evaluate an image, we have to account for all the codes that could produce it. The nonlinear decoder makes this integral difficult to compute. This is the problem the VAE paper starts from.
Which codes could explain this image?
The posterior answers that question:
And the denominator is the same integral again. So we train an encoder to approximate the posterior. follows from the generative model we defined. is our learned approximation to it.
A common encoder outputs a mean vector and a standard-deviation vector :
Now an image gives us a distribution over codes. We reuse the same encoder for every image, so one forward pass estimates a new image's posterior. This is called amortized inference.
Deriving the VAE objective
We want an objective we can evaluate without knowing the true posterior. Let for this section. The KL divergence is:
It is nonnegative and is zero when the distributions agree, up to sets of probability zero. It is not a symmetric distance: reversing and changes what we measure.
Start from the divergence between our approximate and true posteriors. Substitute Bayes' rule into its denominator:
The last term does not depend on , so it comes outside the expectation. Rearrange:
where:
Because the remaining divergence is nonnegative, . This is the evidence lower bound, or ELBO. “Evidence” refers to the observed-data likelihood. We maximize a bound on its log, not the likelihood itself. This identity is the basis of variational inference.
The first term rewards reconstruction. We sample a code for an image and ask the decoder to assign that image high likelihood.
The second term keeps the code distribution close to the prior we will sample later. If every encoded distribution exactly matched the prior, the code would tell us nothing about its input. Reconstruction pushes back against that loss of information.
Reconstruction likelihood is a modelling choice
For an isotropic Gaussian decoder with mean and fixed covariance across data dimensions:
Only the first term depends on the predicted mean. This is why squared reconstruction error appears in this VAE. The second term is constant only when is fixed. The probabilistic decoder definitions also give a Bernoulli decoder for binary data, whose negative log likelihood is binary cross-entropy.
This gives the familiar MSE plus KL loss. The Gaussian decoder and fixed variance are what made it MSE. Averaging one term over dimensions while summing the other would also change their relative weight.
We can also see why a predicted mean image might look blurry. Suppose one code leaves a pixel equally likely to be 0 or 1. The expected squared error for a prediction is:
The minimum is at . If the uncertainty is about where an edge goes, averaging its possible positions softens it. A different decoder distribution can handle that uncertainty differently.
Sampling while keeping gradients
We need samples from the encoder to estimate the reconstruction expectation. For a diagonal Gaussian, write the sample as:
means elementwise multiplication. The randomness is now in , whose distribution does not depend on the encoder parameters. For a fixed sampled , is differentiable with respect to both and . This is the reparameterization trick. The stochastic backpropagation paper derives gradients through this transformation.
Notice the standard deviation multiplying the noise. If our encoder outputs log variance , we need here.
The prior KL has a closed form. For latent dimensions and the standard normal prior:
This follows by substituting the two Gaussian densities into the KL definition and using and .
For one coordinate with and , a noise sample gives . Its KL contribution is:
If and , this contribution is zero. If the sampled reconstruction NLL is 12, the negative-ELBO estimate for the first case is about . We minimize this estimate, averaging across the batch. One sampled value is a noisy estimate; it need not itself satisfy the exact expectation's bound.
At generation time, discard the encoder, sample from the prior, and use the decoder. Displaying its mean is common, but is different from sampling the observation distribution as well.
Generative adversarial networks
So far we have trained likelihoods or a bound on them. Let's return to the implicit branch of the taxonomy and train a sampling procedure directly.
A GAN samples from a simple prior and computes . This is the generator. A second network, the discriminator , learns to tell real images from generated ones. It outputs a score between 0 and 1 for the probability of "real" in that task. is its set of parameters.
The original GAN objective is a two-player game:
The discriminator maximizes the score by assigning real examples values near 1 and generated examples values near 0. The generator minimizes the second term by making its outputs receive larger discriminator scores. The first term contains no generator parameters.
Alternating updates
We alternate two steps:
- Update : keep fixed, generate a batch, and minimize binary cross-entropy with real targets 1 and generated targets 0.
- Update : keep 's parameters fixed, but backpropagate through into the generated images and then .
Freezing discriminator parameters in the second step must not cut the gradient path to the generator. The PyTorch DCGAN tutorial shows these two update paths.
Early on, the discriminator may reject generated images with high confidence. The original minimax generator loss, , then gives a weak gradient through a sigmoid discriminator. The usual non-saturating generator loss instead uses:
During this update, we ask the discriminator to give generated images target 1. We changed the generator's loss, while still trying to match the real data distribution.
We can see the gradient difference directly. Let be the discriminator's logit and . Then:
At , these gradients are and . Both push the logit upward under gradient descent, but the second gives a much stronger signal at this point. The gradient still has to pass through the rest of and .
Matching a distribution takes more than a good image
With a fixed generator and equal weighting of real and generated examples, the ideal discriminator is:
Here is the distribution induced by the generator. When , wherever these densities are nonzero. This is the ideal result proved in the GAN paper. A real discriminator stuck near could also be poorly trained; its score alone does not prove success.
This also explains why we need to look at more than one sample. If the data has ten kinds of cup and different noise vectors all produce the same convincing cup, the generator has missed most of the distribution. This is mode collapse. Improved Techniques for Training GANs uses relationships between samples to help the discriminator detect this failure.
Each network changes the problem the other one is solving. So a falling generator loss may mean better images, or it may mean the discriminator got worse. This makes the loss curves harder to read than in ordinary supervised training.
How do we evaluate them?
We can assign high likelihood to held-out images, generate convincing samples, or cover many kinds of image. A model can do well on one and badly on another. Theis, van den Oord, and Bethge work through this problem.
- Held-out NLL checks probability assigned to unseen data, where likelihood is available. A VAE's ELBO is a bound, not an exact NLL. Compare scores only with consistent data representation and preprocessing.
- Sample inspection can reveal broken shapes or repeated outputs. A few selected images cannot establish diversity or rule out memorization.
- Fréchet Inception Distance, or FID, compares the means and covariances of real and generated images in a pretrained network's feature space, using Gaussian approximations. Lower is better for this comparison. It does not inspect every property of the images, and its value depends on the features, samples, and preprocessing.
For a conditional model, I would also check whether the samples follow the condition. A realistic blue cup is still a wrong answer when we asked for a red one.
We can compare the models through their training objective and sampling process:
| Model | Training and sampling |
|---|---|
| Autoregressive | Fit conditional predictions with exact log likelihood. Sample each next value from the generated prefix. |
| VAE | Maximize a lower bound on log likelihood. Sample a latent code, then the decoder distribution. |
| GAN | Train against a learned discriminator. Sample noise, then run the generator. |
Next are diffusion models. Instead of asking the generator to produce the whole image at once, we learn an update that we can apply repeatedly.