On self-supervised learning

This is part #12 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes are based on lecture 12: Self-Supervised Learning by Ehsan Adeli.


So far, we have trained our networks by giving them an image and telling them what is in it. Someone had to provide the class, draw the bounding box, or mark the pixels that belong to an object.

What if we could get the training target from the image itself?

We could rotate an image, then ask a network to predict which rotation we applied. Or hide some pixels and ask it to fill them in. We know the answer because we made the change.

This is self-supervised learning. We still compute a loss and use backpropagation to train the network. The supervision comes from the data, so we can train on images that nobody has labelled.

What we want to keep is the representation h=fθ(x)h=f_\theta(x) learned along the way. Here xx is an input, ff is an encoder, and θ\theta is its parameters. We can then reuse hh for classification, detection, or another task.

Give the network something to predict

Let's start with the rotation example. Apply one of four rotations and predict its index with the same classification loss we already know. We call this a pretext task, since predicting rotations is how we train the features we will use later.

We can construct several tasks this way:

TaskInput → target
Rotation predictionRotated image → which of four rotations was applied
Relative patch positionTwo image patches → where the second was relative to the first
Jigsaw puzzleShuffled patches → their arrangement
ColourizationGrayscale image → missing colour information

For colourization, we remove the colour channels and use the original colours as the target. For the jigsaw task, we know the original patch arrangement before shuffling it.

To place a wheel below a car window, the network has a reason to learn what those parts look like. But it could also match an edge continuing from one patch into the next. The context prediction paper leaves gaps and randomly shifts patches to make that shortcut harder.

If matching edges solves the task, the loss has no reason to ask for anything more. We need to choose the task so that solving it requires features we can reuse.

Colour can teach correspondence

Now take a video. Keep the colour in one reference frame and remove it from later frames. To restore the colour, let the network point to locations in the reference frame and copy from them.

To copy the colour of a moving shirt, it has to find where that shirt was before. This gives us a correspondence between locations in two frames. Once we have those matches, we can copy an object mask or keypoints through the video too. Tracking Emerges by Colorizing Videos learns these matches without tracking labels.

We asked it to colour a video, and the intermediate computation gives us a way to track parts of it. I find that pretty cool.

Reconstructing what we hide

We can also ask the network to reconstruct the image. An autoencoder first encodes the input, then passes that representation to a decoder that tries to recover the original pixels.

If the encoder sees everything, it may learn to copy the input. Hide part of it and the decoder has to use the remaining content to fill the gaps.

Masked Autoencoders, or MAE, do this with a Vision Transformer. We divide the image into patches, randomly remove most of them, and pass only the visible patch tokens through the encoder. Each token keeps its position information.

After the encoder, we put a learned mask token at each missing position. A smaller decoder now reads the full sequence and predicts pixels:

Masked autoencoder with visible-only encoding A toy image has sixteen patches. Four visible blue patches become tokens for the encoder; twelve hidden gray patches are removed. After encoding, learned mask tokens restore the twelve missing positions for a smaller decoder. The decoder predicts all sixteen patches, while reconstruction loss uses only the twelve hidden targets. Position information is retained. 16 patches · hide 12 blue = visible · gray = hidden 4 visible tokens + positions Encoder Insert 12 mask tokensrestore positions for all 16 tokens Smaller decoder Predict pixels at all 16 positions Reconstruction lossonly the 12 hidden patches original pixels as targets
The encoder sees only visible patches. Mask tokens enter at the decoder, and only the hidden patches contribute to the reconstruction loss.

Notice that the large encoder never processes the mask tokens. With 16×16=25616\times16=256 patches and 75% masking, it reads just 64 patch tokens. The smaller decoder handles all 256 positions.

We compare its output with the original pixels at the masked positions:

Lmask=1∣M∣D∑i∈M∥x^i−xi∥22L_{\text{mask}} = \frac{1}{|M|D} \sum_{i\in M}\|\hat{x}_i-x_i\|_2^2

MM is the set of masked patch indices and ∣M∣|M| is its size. xix_i is the original patch, x^i\hat{x}_i is its prediction, and DD is the number of pixel values per patch. For a 16×1616\times16 RGB patch, D=16⋅16⋅3=768D=16\cdot16\cdot3=768. MAE also studies targets normalized within each patch.

If we hide only a small patch, its neighbours may reveal almost everything. Removing 75% makes the network combine information from farther away, while also giving the encoder fewer tokens to process.

After training, we discard the decoder and give the encoder the full image. We trained it through reconstruction, but we use it to produce features.

The decoder still has to predict details we may not need for recognition. Several handles could fit a partly hidden cup, yet pixel error rewards the one in this photograph. That leads to another approach: compare images without reconstructing their pixels.

Learning by comparison

Let's take two views of an image and ask the network to recognize that they came from the same image. This is the starting point for contrastive learning.

SimCLR makes the views with random crops and colour transformations. We call the two views a positive pair, and pair each view with views from other images to get negatives.

Pass both through the same encoder fθf_\theta, then through a small projection network gϕg_\phi:

hi=fθ(vi),zi=gϕ(hi)h_i=f_\theta(v_i),\qquad z_i=g_\phi(h_i)

Here viv_i is a view, hih_i is its encoder representation, and ziz_i is the vector we compare. ϕ\phi is the projection network's parameters. Both branches reuse the same weights:

Two views, one shared encoder An image is transformed into two views. Each passes through the same encoder f and projection g. The contrastive loss brings their projected vectors together relative to negative examples. The encoder representations h are kept for later tasks. Image x View 1View 2 Encoder fEncoder f same weights h₁h₂ Projection gProjection g z₁z₂ Contrastive loss + negatives
Two views of one image form a positive pair. Both paths use the same encoder and projection weights.

The loss trains both networks. Later we discard gϕg_\phi and keep fθf_\theta. The projection gives the comparison its own feature space, where it can ignore colour changes while allowing hih_i to retain colour information.

The contrastive loss

With NN images in a batch, we have 2N2N views. Pick view ii as our anchor and let jj be its positive partner. The network has to select jj from the other 2N−12N-1 views. One is the positive; the remaining 2N−22N-2 are negatives.

First, compare ii with each candidate kk using cosine similarity:

sik=ziTzk∥zi∥2∥zk∥2s_{ik}=\frac{z_i^Tz_k}{\|z_i\|_2\|z_k\|_2}

For nonzero vectors this gives a score between −1-1 and 11. It compares their directions, so making a vector longer does not improve its score.

Apply softmax to the scores, after dividing by a temperature τ>0\tau>0:

pik=exp⁡(sik/τ)∑a≠iexp⁡(sia/τ)p_{ik}=\frac{\exp(s_{ik}/\tau)} {\sum_{a\ne i}\exp(s_{ia}/\tau)}

Both kk and aa index candidate views. Notice that the denominator includes the positive partner but excludes ii itself. We then take the negative log-probability of the correct match:

ℓi,j=−log⁡pij\ell_{i,j}=-\log p_{ij}

It is the same cross-entropy loss, with another view's index as the class. Repeat with each of the 2N2N views as anchor and average the losses. That includes both directions of each positive pair.

SimCLR calls this NT-Xent, short for normalized temperature-scaled cross-entropy. It belongs to the InfoNCE family introduced in Contrastive Predictive Coding.

A small example

Take a photograph of a cup and one of a bicycle. Two views of each give us four vectors. Use the first cup view as the anchor; it needs to find the other cup view among three candidates.

Let's give those candidates similarities of 0.80.8, 0.30.3, and 0.10.1. Set τ=0.5\tau=0.5, divide, then apply softmax:

Candidates/τs/\tauProbability
Cup, positive1.60.619
Bicycle view 10.60.228
Bicycle view 20.20.153

The positive probability and loss are:

pij=e1.6e1.6+e0.6+e0.2≈0.6194p_{ij}=\frac{e^{1.6}}{e^{1.6}+e^{0.6}+e^{0.2}} \approx0.6194 ℓi,j≈−log⁡(0.6194)≈0.479\ell_{i,j}\approx-\log(0.6194)\approx0.479

If all scores were equal, each probability would be 1/31/3 and the loss would be log⁡3≈1.099\log 3\approx1.099. Here the positive already has the highest score, and the loss asks us to separate it further.

To see how, take the derivative with respect to the positive similarity:

∂ℓ∂sij=pij−1τ≈−0.761\frac{\partial\ell}{\partial s_{ij}} =\frac{p_{ij}-1}{\tau}\approx-0.761

For a negative candidate kk:

∂ℓ∂sik=pikτ\frac{\partial\ell}{\partial s_{ik}} =\frac{p_{ik}}{\tau}

The two negative candidates get gradients of approximately 0.4560.456 and 0.3050.305. At the level of these scores, gradient descent pulls the positive up and pushes the negatives down. The bicycle view that looks more similar gets the larger correction. Backpropagation then carries these derivatives into the vectors and shared weights.

Lower τ\tau to 0.10.1 without changing the vectors, and the positive probability becomes about 0.9920.992. Temperature makes the distribution sharper. It would also make a wrong match more confident if that match had the highest score, so lowering the loss this way does not mean we learned better features.

Which differences should disappear?

By making two transformed views agree, we ask the comparison to ignore those transformations. This is the invariance we are trying to learn.

For digit recognition, asking a rotated 6 to match a 9 would erase a distinction we need. Cropping also has a limit: one view might contain a cup handle while the other contains only the table. We need views that still share useful content.

A second photograph of a cup also becomes a negative, even though our later classifier may want both cups in the same class. That is a false negative relative to the later task. During pretraining, we only know which original image each view came from.

More negatives without a huge batch

In SimCLR the negatives come from the current batch. To compare with more images, we need a larger batch and more memory for its activations.

MoCo keeps encoded views from recent batches in a queue. We compute a query for the current image and compare it with its positive key and the queued negative keys:

MoCo momentum encoder and queue Two views of the same image enter separate query and key encoders. The contrastive loss compares the query with the positive key and negative keys from a queue. Gradients train only the query encoder. An exponential moving average of query weights updates the key encoder. After the current comparison, new keys enter the queue and the oldest keys leave. Two views of the same image View 1View 2 Query encodergradient update Key encodermoving average EMA of weights Query qPositive k⁺ Contrastive loss add new keys after comparison Queue of negative keys from recent batches · no gradients oldest out No gradient through the key branch
The query encoder learns through gradients. A moving average updates the key encoder; keys enter a queue and do not receive gradients.

Unlike SimCLR, the two encoders have separate weights. We train the query encoder through gradients and update the key encoder with a moving average:

θk←mθk+(1−m)θq\theta_k\leftarrow m\theta_k+(1-m)\theta_q

θq\theta_q and θk\theta_k are their parameters, and 0≤m<10\le m<1. With mm close to 1, the key encoder changes slowly.

That slow change matters because the queue contains vectors computed with older weights. We want old and new keys to remain comparable. After using the queue, add the new keys and remove the oldest ones. No gradient passes through the keys.

We can now increase the number of negatives without increasing the current batch. MoCo v2 combines this queue with SimCLR's projection head and stronger transformations.

Predicting a sequence in feature space

We can use the same comparison to predict what comes next in a sequence. This is Contrastive Predictive Coding, or CPC.

Encode each observation xtx_t into ztz_t, then use an autoregressive network to collect everything up to time tt into context ctc_t. Given that context, we ask it to identify the real future observation among sampled alternatives.

For the observation kk steps ahead, compute a score with a learned matrix WkW_k:

s=zt+kTWkcts=z_{t+k}^TW_kc_t

Here tt is the current step and kk is how far ahead we predict. There is a separate WkW_k for each distance. Exponentiate and normalize the candidate scores as before, then apply the loss to the real future observation.

The prediction happens in feature space, so we do not have to reconstruct every pixel or audio sample. Keep future targets out of the context network or it can read the answer. CPC also applies this to images by treating rows of patches as a sequence, predicting lower rows from upper ones.

What if we remove the negatives?

So far we have pulled positive pairs together and pushed negatives apart. If we keep only the first part, the network can output one constant vector for every image. All pairs agree and the loss is solved. This is collapse.

To learn without negatives, we need to change how the two branches train.

BYOL adds a predictor to the online branch and trains it to match a target branch on another view. As in MoCo, a moving average updates the target weights. The target receives no gradient.

SimSiam uses shared weights and an extra predictor, but removes the moving-average encoder. It stops the gradient on the branch that supplies the target.

With stop-gradient, the forward pass still uses the target value, but backpropagation treats it as constant. The predictor and the stopped path make training asymmetric. Removing only the negatives would not give us either algorithm; stopping a gradient by itself does not guarantee that we avoid collapse.

DINO

DINO keeps a student and a teacher. We train the student to predict the teacher's output distribution on another view of the image. The teacher follows a moving average of the student, so neither starts with labels. This is self-distillation.

Give the teacher large global crops and the student both global crops and smaller local ones. The student then has to relate a local view to the content the teacher sees in a larger view.

Let asa_s and ata_t be their output vectors, each with KK learned dimensions. These are not named classes. Apply softmax to get the distributions:

ps=softmax⁡(as/τs)p_s=\operatorname{softmax}(a_s/\tau_s) pt=softmax⁡((at−c)/τt)p_t=\operatorname{softmax}((a_t-c)/\tau_t)

τs\tau_s and τt\tau_t are the two temperatures. Before the teacher's softmax, subtract a moving average cc of its outputs to centre them. For a pair of views, the loss is:

L=−∑d=1Ksg⁡(pt,d)log⁡ps,dL=-\sum_{d=1}^{K}\operatorname{sg}(p_{t,d})\log p_{s,d}

dd indexes the output dimensions and sg⁡\operatorname{sg} stops the gradient. Backpropagation updates the student. The moving average then updates the teacher.

We could still collapse by selecting the same output dimension for every image, or by predicting a uniform distribution for all images.

Centering reduces the tendency for one dimension to dominate. A low teacher temperature sharpens the targets, working against the uniform solution. DINO combines both to avoid collapse in its experiments, without a negative-pair term.

After training, take the backbone features before the output head. In DINO's Vision Transformers, some self-attention maps separate foreground objects from the background even though training never used segmentation labels. The maps show spatial structure the model learned; we still need a segmentation head to predict a full set of semantic labels.

DINOv2 builds on this with selected training data, objectives at both image and patch level, and larger-scale training.

Another source of supervision: sound

We can also take two observations of the same event. Look, Listen and Learn takes an image and an audio segment, then predicts whether they belong together. Sample matching pairs from a video and make mismatched pairs with audio from elsewhere.

The visual and audio encoders have to find related content without anyone naming the object or the sound.

A guitar recording might have speech over it, or the sound might come from outside the frame. The pairing gives us useful supervision, but it is noisier than a manually checked label.

Using the encoder

Now take the trained encoder and attach a classifier for the task we wanted to solve. This is our downstream task, and we evaluate it on held-out examples.

First, freeze the encoder and train just a linear classifier. This is a linear probe:

y^=softmax⁡(Wh+b),h=fθ(x)\hat{y}=\operatorname{softmax}(Wh+b),\qquad h=f_\theta(x)

Only WW and bb change. h∈RDhh\in\mathbb{R}^{D_h} is the feature vector, W∈RC×DhW\in\mathbb{R}^{C\times D_h}, b∈RCb\in\mathbb{R}^{C}, and CC is the number of labelled classes. If this works, a linear rule can separate the classes in the features we already learned.

Unfreeze some or all of the encoder and we get fine-tuning. The features can now change to fit the task. A model that does better when frozen may not do better after fine-tuning, as MAE's experiments show.

For detection or segmentation, attach the corresponding head and use spatial features. Separating image classes with one vector does not test whether the model preserved object boundaries.

To compare the training methods, use the same encoder, pretraining data, and number of labelled examples. Keep the test images separate from pretraining too. Otherwise a score may reflect the extra data instead of the learning method.

In the next notes on generative models, we return to reconstruction and generation. There, the generated examples are the output we care about.

← Back to blog