On generative models

This is part #13 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes are based on lecture 13: Generative Models by Justin Johnson, and the GAN section of lecture 14.


So far we have taken an image and computed something from it. A class, a box, or a useful representation, as in self-supervised learning.

Now we want to go the other way and produce an image. If we ask for a cup, there isn't one correct output. We could get a different shape, colour, or viewpoint each time. So instead of predicting one answer, we learn a distribution and sample from it.

What are we modelling?

Let xx be an image and yy a condition, such as a class or a text description:

  • A discriminative model learns pθ(y∣x)p_\theta(y\mid x): given this image, which label fits?
  • An unconditional generative model learns pθ(x)p_\theta(x): how are images distributed?
  • A conditional generative model learns pθ(x∣y)p_\theta(x\mid y): which images fit this condition?

As before, θ\theta is our set of learnable parameters. We don't know the data distribution pdatap_{\text{data}}. We only have examples drawn from it.

For discrete images, the probabilities over all possible images sum to 1. For continuous values, we use a density, whose integral is 1. A density value can exceed 1; probability comes from integrating it over a region.

A taxonomy of generative models

Before looking at the models, let's put them on a map. The lecture starts by asking whether we can compute pθ(x)p_\theta(x) for a given image.

An explicit density model defines a likelihood. Sometimes we can evaluate it directly. Sometimes the definition contains an integral we cannot compute, so we optimize a bound instead.

An implicit model defines a way to generate samples. For example, take random noise and pass it through a neural network. This gives us a distribution over outputs, even if we cannot evaluate its density at an arbitrary image.

Generative models by likelihood and sampling Explicit density models include autoregressive models and tractable normalizing flows with exact likelihoods, and VAEs and DDPMs trained with likelihood bounds. A GAN defines an implicit distribution through its generator. A separate sampling axis groups GANs and VAEs with parallel, factorized decoders as one generator or decoder pass after drawing noise. Autoregressive models, DDPMs, and continuous flows need sequential or iterative steps. Generative models Explicit density Implicit Exact likelihood Autoregressive Normalizing flows Likelihood bound VAE · DDPM Sample without evaluating p(x) GAN How do we sample? One generator or decoder pass GAN · VAE with a parallel decoder Sequential or iterative steps Autoregressive · DDPM · continuous flows Likelihood and sampling describe different properties.
Two questions about the same models. A tractable likelihood does not imply one-step generation. The VAE sampling row assumes a feedforward, factorized decoder, as in the example below.

For autoregressive models, we write the likelihood as a product of conditional probabilities. We can evaluate that product exactly for the chosen model. Generating a new image still takes a sequence of predictions.

Normalizing flows also give a tractable likelihood. They use invertible transformations with a tractable Jacobian determinant to turn a simple density into a more complex one. The determinant accounts for how each transformation expands or compresses volume. Real NVP is one example.

For a VAE, the model defines a likelihood through a latent variable, but evaluating it requires an integral. We will derive the lower bound that makes training possible. A GAN instead learns a generator through a discriminator, without evaluating an image likelihood.

Where do diffusion models fit? The lecture places them under indirect sampling because generation uses repeated updates. That is a useful description of the sampler. A DDPM also defines a latent-variable likelihood and trains with a variational bound. So "iterative sampling" and "implicit density" are not the same property.

Keep those two questions separate as we go: what objective can we train, and what do we run to get a sample? Let's start with likelihood.

Maximum likelihood

Given NN independent training examples x(1),…,x(N)x^{(1)},\ldots,x^{(N)}, maximum likelihood estimation chooses parameters that make the observed dataset likely:

θ∗=arg⁡max⁡θ∏n=1Npθ(x(n))=arg⁡max⁡θ∑n=1Nlog⁡pθ(x(n))\begin{aligned} \theta^* &= \arg\max_\theta \prod_{n=1}^{N} p_\theta(x^{(n)}) \\ &= \arg\max_\theta \sum_{n=1}^{N}\log p_\theta(x^{(n)}) \end{aligned}

The log is increasing, so it preserves the optimum while turning a product into a sum. We normally minimize the negative log likelihood, or NLL:

LNLL(θ)=−1N∑n=1Nlog⁡pθ(x(n))L_{\text{NLL}}(\theta) =-\frac{1}{N}\sum_{n=1}^{N}\log p_\theta(x^{(n)})

Now we have a loss to minimize through gradient descent. We use natural logarithms, so the loss is measured in nats. See Kingma and Welling for the connection between likelihood and minibatch training.

Let's give two observed outcomes probabilities 0.60.6 and 0.30.3. Their joint likelihood is 0.180.18, and the average NLL is about 0.8570.857. Change the probabilities to 0.50.5 and 0.40.4: the likelihood rises to 0.200.20 and the average NLL falls to 0.8050.805.

Notice that the first outcome became less likely, yet the loss improved. We optimize the combined likelihood of the examples. We then use held-out data to check whether that improvement generalizes.

Autoregressive models

Write an image as a sequence x=(x1,…,xT)x=(x_1,\ldots,x_T). The chain rule of probability gives:

pθ(x)=∏t=1Tpθ(xt∣x<t)p_\theta(x)=\prod_{t=1}^{T} p_\theta(x_t\mid x_{<t})

Here x<tx_{<t} means all values before position tt. We haven't assumed that pixels are independent. Each prediction can depend on every earlier value. The chain rule is exact; our network approximates the conditional distributions.

PixelRNN applies this to images. Visit pixels row by row, and colour channels in a fixed order. Each 8-bit channel value becomes a classification problem with 256 possible values. The network predicts a distribution over the next value using only earlier ones.

We have already used this idea for next-token prediction in the RNN notes. We just changed what goes into the sequence.

Training and sampling use different inputs

During training, the whole image is available. To predict xtx_t, give the model the real prefix x<tx_{<t} and score its distribution against the real target. This is called teacher forcing.

The image loss becomes a sum of ordinary cross-entropies:

−log⁡pθ(x)=−∑t=1Tlog⁡pθ(xt∣x<t)-\log p_\theta(x) =-\sum_{t=1}^{T}\log p_\theta(x_t\mid x_{<t})

We block access to the target and future values with a causal mask. Otherwise, the network could copy the answer. Networks such as MADE, masked CNNs, and causal transformers can evaluate the training conditionals in parallel. An RNN still has to compute its hidden states in sequence.

During generation, there is no real prefix to supply:

  1. Sample x1x_1 from the first distribution.
  2. Predict pθ(x2∣x1)p_\theta(x_2\mid x_1) and sample x2x_2.
  3. Append the sample and repeat.

Let's use a three-pixel binary image x=(1,0,1)x=(1,0,1):

p(1)=0.6p(0∣1)=0.7p(1∣1,0)=0.8\begin{aligned} p(1)&=0.6 \\ p(0\mid1)&=0.7 \\ p(1\mid1,0)&=0.8 \end{aligned}

Multiplying gives p(1,0,1)=0.6⋅0.7⋅0.8=0.336p(1,0,1)=0.6\cdot0.7\cdot0.8=0.336, with an NLL of about 1.0911.091 nats. Each factor is the probability assigned to the pixel value we observed.

When we generate, the sampled prefix determines what happens next. If the second pixel were 1, we would use p(x3∣1,1)p(x_3\mid1,1) for the third. Taking the highest-probability value every time would turn this into greedy decoding.

To condition on a label or an image representation, include it at every step: pθ(xt∣x<t,y)p_\theta(x_t\mid x_{<t},y). Conditional PixelCNN demonstrates this. The model's basic likelihood calculation stays the same.

This becomes slow for images. A 256×256256\times256 RGB image has 196,608196{,}608 channel values, and here each one needs the previous samples. We can shorten the sequence by generating compressed image tokens. We then model those tokens and use a decoder to turn them into pixels.

From autoencoders to a generative model

An ordinary autoencoder maps an image to a code and back:

z=fϕ(x),x^=gθ(z)z=f_\phi(x),\qquad \hat{x}=g_\theta(z)

The encoder has parameters ϕ\phi and the decoder has parameters θ\theta. We train them together to reconstruct the input, usually with a bottleneck that forces the code to keep useful information.

But where do we get a new zz after training? We know how to encode an image, but we wanted to generate one. Random coordinates may fall between the codes the decoder learned to use.

A variational autoencoder, or VAE, starts with that missing sampling step. We choose a simple prior over the latent variable, often:

z∼p(z)=N(0,I)z\sim p(z)=\mathcal{N}(0,I)

Then we sample an image from the decoder distribution pθ(x∣z)p_\theta(x\mid z). II is the identity matrix, so each prior coordinate is an independent Gaussian with variance 1. The decoder outputs distribution parameters, such as the mean image.

The resulting image density is:

pθ(x)=∫pθ(x∣z)p(z) dzp_\theta(x)=\int p_\theta(x\mid z)p(z)\,dz

To evaluate an image, we have to account for all the codes that could produce it. The nonlinear decoder makes this integral difficult to compute. This is the problem the VAE paper starts from.

Which codes could explain this image?

The posterior answers that question:

pθ(z∣x)=pθ(x∣z)p(z)pθ(x)p_\theta(z\mid x) =\frac{p_\theta(x\mid z)p(z)}{p_\theta(x)}

And the denominator is the same integral again. So we train an encoder qϕ(z∣x)q_\phi(z\mid x) to approximate the posterior. pθ(z∣x)p_\theta(z\mid x) follows from the generative model we defined. qϕ(z∣x)q_\phi(z\mid x) is our learned approximation to it.

A common encoder outputs a mean vector μ=μϕ(x)\mu=\mu_\phi(x) and a standard-deviation vector σ=σϕ(x)\sigma=\sigma_\phi(x):

qϕ(z∣x)=N ⁣(μ,diag⁡(σ2))q_\phi(z\mid x)=\mathcal{N}\!\left( \mu,\operatorname{diag}(\sigma^2) \right)

Now an image gives us a distribution over codes. We reuse the same encoder for every image, so one forward pass estimates a new image's posterior. This is called amortized inference.

VAE reconstruction and generation paths The reconstruction path starts with an input image, uses an encoder to choose a latent code, then a decoder to reconstruct an image. The generation path draws a code from the standard normal prior and uses the same decoder to generate a new image. It does not use the encoder. Reconstruct Generate Input image x Encoder qφ(z | x) Prior p(z) = N(0, I) Sample code z using μ and σ Sample code z from the prior Decoder pθ(x | z) Decoder pθ(x | z) Reconstruct the input Generate a new image Both decoders have the same weights.
Reconstruction starts with an image and uses the encoder. Generation starts from the prior. Both paths use the same decoder.

Deriving the VAE objective

We want an objective we can evaluate without knowing the true posterior. Let q=qϕ(z∣x)q=q_\phi(z\mid x) for this section. The KL divergence is:

DKL(q∥p)=Ez∼q ⁣[log⁡q(z)p(z)]D_{\mathrm{KL}}(q\|p) =\mathbb{E}_{z\sim q}\!\left[\log\frac{q(z)}{p(z)}\right]

It is nonnegative and is zero when the distributions agree, up to sets of probability zero. It is not a symmetric distance: reversing qq and pp changes what we measure.

Start from the divergence between our approximate and true posteriors. Substitute Bayes' rule into its denominator:

DKL(q∥pθ(z∣x))=Eq[log⁡q]−Eq[log⁡pθ(x∣z)]−Eq[log⁡p(z)]+log⁡pθ(x)\begin{aligned} &D_{\mathrm{KL}}(q\|p_\theta(z\mid x))\\ &=\mathbb{E}_{q}[\log q] \\ &\quad-\mathbb{E}_{q}[\log p_\theta(x\mid z)] \\ &\quad-\mathbb{E}_{q}[\log p(z)] +\log p_\theta(x) \end{aligned}

The last term does not depend on zz, so it comes outside the expectation. Rearrange:

log⁡pθ(x)=LELBO(x)+DKL(q∥pθ(z∣x))\begin{aligned} \log p_\theta(x) ={}&\mathcal{L}_{\mathrm{ELBO}}(x) \\ &+D_{\mathrm{KL}}(q\|p_\theta(z\mid x)) \end{aligned}

where:

LELBO(x)=Ez∼qϕ(z∣x)[log⁡pθ(x∣z)]−DKL(qϕ(z∣x)∥p(z))\begin{aligned} \mathcal{L}_{\mathrm{ELBO}}(x) ={}&\mathbb{E}_{z\sim q_\phi(z\mid x)} [\log p_\theta(x\mid z)] \\ &-D_{\mathrm{KL}}(q_\phi(z\mid x)\|p(z)) \end{aligned}

Because the remaining divergence is nonnegative, LELBO(x)≤log⁡pθ(x)\mathcal{L}_{\mathrm{ELBO}}(x)\le\log p_\theta(x). This is the evidence lower bound, or ELBO. “Evidence” refers to the observed-data likelihood. We maximize a bound on its log, not the likelihood itself. This identity is the basis of variational inference.

The first term rewards reconstruction. We sample a code for an image and ask the decoder to assign that image high likelihood.

The second term keeps the code distribution close to the prior we will sample later. If every encoded distribution exactly matched the prior, the code would tell us nothing about its input. Reconstruction pushes back against that loss of information.

Reconstruction likelihood is a modelling choice

For an isotropic Gaussian decoder with mean gθ(z)g_\theta(z) and fixed covariance s2Is^2I across DD data dimensions:

−log⁡pθ(x∣z)=∥x−gθ(z)∥222s2+D2log⁡(2πs2)\begin{aligned} -\log p_\theta(x\mid z) ={}&\frac{\|x-g_\theta(z)\|_2^2}{2s^2} \\ &+\frac{D}{2}\log(2\pi s^2) \end{aligned}

Only the first term depends on the predicted mean. This is why squared reconstruction error appears in this VAE. The second term is constant only when ss is fixed. The probabilistic decoder definitions also give a Bernoulli decoder for binary data, whose negative log likelihood is binary cross-entropy.

This gives the familiar MSE plus KL loss. The Gaussian decoder and fixed variance are what made it MSE. Averaging one term over dimensions while summing the other would also change their relative weight.

We can also see why a predicted mean image might look blurry. Suppose one code leaves a pixel equally likely to be 0 or 1. The expected squared error for a prediction aa is:

12a2+12(1−a)2\tfrac12a^2+\tfrac12(1-a)^2

The minimum is at a=0.5a=0.5. If the uncertainty is about where an edge goes, averaging its possible positions softens it. A different decoder distribution can handle that uncertainty differently.

Sampling while keeping gradients

We need samples from the encoder to estimate the reconstruction expectation. For a diagonal Gaussian, write the sample as:

ϵ∼N(0,I),z=μϕ(x)+σϕ(x)⊙ϵ\epsilon\sim\mathcal{N}(0,I),\qquad z=\mu_\phi(x)+\sigma_\phi(x)\odot\epsilon

⊙\odot means elementwise multiplication. The randomness is now in ϵ\epsilon, whose distribution does not depend on the encoder parameters. For a fixed sampled ϵ\epsilon, zz is differentiable with respect to both μ\mu and σ\sigma. This is the reparameterization trick. The stochastic backpropagation paper derives gradients through this transformation.

Notice the standard deviation multiplying the noise. If our encoder outputs log variance ℓ=log⁡σ2\ell=\log\sigma^2, we need σ=exp⁡(ℓ/2)\sigma=\exp(\ell/2) here.

The prior KL has a closed form. For dd latent dimensions and the standard normal prior:

DKL(qϕ∥p)=12∑j=1d(μj2+σj2−1−log⁡σj2)\begin{aligned} &D_{\mathrm{KL}}(q_\phi\|p)\\ &=\frac12\sum_{j=1}^{d} \left(\mu_j^2+\sigma_j^2-1-\log\sigma_j^2\right) \end{aligned}

This follows by substituting the two Gaussian densities into the KL definition and using E[(zj−μj)2]=σj2\mathbb{E}[(z_j-\mu_j)^2]=\sigma_j^2 and E[zj2]=μj2+σj2\mathbb{E}[z_j^2]=\mu_j^2+\sigma_j^2.

For one coordinate with μ=1\mu=1 and σ=0.5\sigma=0.5, a noise sample ϵ=−0.4\epsilon=-0.4 gives z=1+0.5(−0.4)=0.8z=1+0.5(-0.4)=0.8. Its KL contribution is:

12(1+0.25−1−log⁡0.25)≈0.818\tfrac12(1+0.25-1-\log0.25)\approx0.818

If μ=0\mu=0 and σ=1\sigma=1, this contribution is zero. If the sampled reconstruction NLL is 12, the negative-ELBO estimate for the first case is about 12.81812.818. We minimize this estimate, averaging across the batch. One sampled value is a noisy estimate; it need not itself satisfy the exact expectation's bound.

At generation time, discard the encoder, sample zz from the prior, and use the decoder. Displaying its mean is common, but is different from sampling the observation distribution as well.

Generative adversarial networks

So far we have trained likelihoods or a bound on them. Let's return to the implicit branch of the taxonomy and train a sampling procedure directly.

A GAN samples zz from a simple prior and computes x^=Gθ(z)\hat{x}=G_\theta(z). This is the generator. A second network, the discriminator Dψ(x)D_\psi(x), learns to tell real images from generated ones. It outputs a score between 0 and 1 for the probability of "real" in that task. ψ\psi is its set of parameters.

Generator and discriminator in a GAN Noise enters the generator to produce a generated image. A real image and a generated image are separate inputs to the same discriminator. The discriminator predicts a real score and learns targets one for real and zero for generated. For a generator update, the discriminator weights stay fixed and gradients pass through it to the generator. Only the generator is needed after training. Noise z Generator G Real image x Generated G(z) Target 1 Target 0 Discriminator D Score: real? Update D: classify real and generated images. Update G: backpropagate through fixed D. After training, generation uses G alone.
The discriminator sees either a real image or a generated image. On the generator update, gradients pass through the fixed discriminator and back into the generator.

The original GAN objective is a two-player game:

min⁡θmax⁡ψV(θ,ψ)\min_\theta\max_\psi V(\theta,\psi) V(θ,ψ)=Ex∼pdata[log⁡Dψ(x)]+Ez∼p(z)[log⁡(1−Dψ(Gθ(z)))]\begin{aligned} &V(\theta,\psi)\\ &=\mathbb{E}_{x\sim p_{\mathrm{data}}} [\log D_\psi(x)] \\ &\quad+\mathbb{E}_{z\sim p(z)} [\log(1-D_\psi(G_\theta(z)))] \end{aligned}

The discriminator maximizes the score by assigning real examples values near 1 and generated examples values near 0. The generator minimizes the second term by making its outputs receive larger discriminator scores. The first term contains no generator parameters.

Alternating updates

We alternate two steps:

  1. Update DD: keep GG fixed, generate a batch, and minimize binary cross-entropy with real targets 1 and generated targets 0.
  2. Update GG: keep DD's parameters fixed, but backpropagate through DD into the generated images and then GG.

Freezing discriminator parameters in the second step must not cut the gradient path to the generator. The PyTorch DCGAN tutorial shows these two update paths.

Early on, the discriminator may reject generated images with high confidence. The original minimax generator loss, log⁡(1−D(G(z)))\log(1-D(G(z))), then gives a weak gradient through a sigmoid discriminator. The usual non-saturating generator loss instead uses:

LG=−Ez∼p(z)[log⁡Dψ(Gθ(z))]L_G=-\mathbb{E}_{z\sim p(z)}[\log D_\psi(G_\theta(z))]

During this update, we ask the discriminator to give generated images target 1. We changed the generator's loss, while still trying to match the real data distribution.

We can see the gradient difference directly. Let uu be the discriminator's logit and r=sigmoid⁡(u)r=\operatorname{sigmoid}(u). Then:

∂log⁡(1−r)∂u=−r∂(−log⁡r)∂u=r−1\begin{aligned} \frac{\partial\log(1-r)}{\partial u}&=-r \\ \frac{\partial(-\log r)}{\partial u}&=r-1 \end{aligned}

At r=0.01r=0.01, these gradients are −0.01-0.01 and −0.99-0.99. Both push the logit upward under gradient descent, but the second gives a much stronger signal at this point. The gradient still has to pass through the rest of DD and GG.

Matching a distribution takes more than a good image

With a fixed generator and equal weighting of real and generated examples, the ideal discriminator is:

D∗(x)=pdata(x)pdata(x)+pG(x)D^*(x)=\frac{p_{\mathrm{data}}(x)} {p_{\mathrm{data}}(x)+p_G(x)}

Here pGp_G is the distribution induced by the generator. When pG=pdatap_G=p_{\mathrm{data}}, D∗(x)=1/2D^*(x)=1/2 wherever these densities are nonzero. This is the ideal result proved in the GAN paper. A real discriminator stuck near 1/21/2 could also be poorly trained; its score alone does not prove success.

This also explains why we need to look at more than one sample. If the data has ten kinds of cup and different noise vectors all produce the same convincing cup, the generator has missed most of the distribution. This is mode collapse. Improved Techniques for Training GANs uses relationships between samples to help the discriminator detect this failure.

Each network changes the problem the other one is solving. So a falling generator loss may mean better images, or it may mean the discriminator got worse. This makes the loss curves harder to read than in ordinary supervised training.

How do we evaluate them?

We can assign high likelihood to held-out images, generate convincing samples, or cover many kinds of image. A model can do well on one and badly on another. Theis, van den Oord, and Bethge work through this problem.

  • Held-out NLL checks probability assigned to unseen data, where likelihood is available. A VAE's ELBO is a bound, not an exact NLL. Compare scores only with consistent data representation and preprocessing.
  • Sample inspection can reveal broken shapes or repeated outputs. A few selected images cannot establish diversity or rule out memorization.
  • Fréchet Inception Distance, or FID, compares the means and covariances of real and generated images in a pretrained network's feature space, using Gaussian approximations. Lower is better for this comparison. It does not inspect every property of the images, and its value depends on the features, samples, and preprocessing.

For a conditional model, I would also check whether the samples follow the condition. A realistic blue cup is still a wrong answer when we asked for a red one.

We can compare the models through their training objective and sampling process:

ModelTraining and sampling
AutoregressiveFit conditional predictions with exact log likelihood. Sample each next value from the generated prefix.
VAEMaximize a lower bound on log likelihood. Sample a latent code, then the decoder distribution.
GANTrain against a learned discriminator. Sample noise, then run the generator.

Next are diffusion models. Instead of asking the generator to produce the whole image at once, we learn an update that we can apply repeatedly.

← Back to blog