On self-supervised learning
This is part #12 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on lecture 12: Self-Supervised Learning by Ehsan Adeli.
So far, we have trained our networks by giving them an image and telling them what is in it. Someone had to provide the class, draw the bounding box, or mark the pixels that belong to an object.
What if we could get the training target from the image itself?
We could rotate an image, then ask a network to predict which rotation we applied. Or hide some pixels and ask it to fill them in. We know the answer because we made the change.
This is self-supervised learning. We still compute a loss and use backpropagation to train the network. The supervision comes from the data, so we can train on images that nobody has labelled.
What we want to keep is the representation learned along the way. Here is an input, is an encoder, and is its parameters. We can then reuse for classification, detection, or another task.
Give the network something to predict
Let's start with the rotation example. Apply one of four rotations and predict its index with the same classification loss we already know. We call this a pretext task, since predicting rotations is how we train the features we will use later.
We can construct several tasks this way:
| Task | Input → target |
|---|---|
| Rotation prediction | Rotated image → which of four rotations was applied |
| Relative patch position | Two image patches → where the second was relative to the first |
| Jigsaw puzzle | Shuffled patches → their arrangement |
| Colourization | Grayscale image → missing colour information |
For colourization, we remove the colour channels and use the original colours as the target. For the jigsaw task, we know the original patch arrangement before shuffling it.
To place a wheel below a car window, the network has a reason to learn what those parts look like. But it could also match an edge continuing from one patch into the next. The context prediction paper leaves gaps and randomly shifts patches to make that shortcut harder.
If matching edges solves the task, the loss has no reason to ask for anything more. We need to choose the task so that solving it requires features we can reuse.
Colour can teach correspondence
Now take a video. Keep the colour in one reference frame and remove it from later frames. To restore the colour, let the network point to locations in the reference frame and copy from them.
To copy the colour of a moving shirt, it has to find where that shirt was before. This gives us a correspondence between locations in two frames. Once we have those matches, we can copy an object mask or keypoints through the video too. Tracking Emerges by Colorizing Videos learns these matches without tracking labels.
We asked it to colour a video, and the intermediate computation gives us a way to track parts of it. I find that pretty cool.
Reconstructing what we hide
We can also ask the network to reconstruct the image. An autoencoder first encodes the input, then passes that representation to a decoder that tries to recover the original pixels.
If the encoder sees everything, it may learn to copy the input. Hide part of it and the decoder has to use the remaining content to fill the gaps.
Masked Autoencoders, or MAE, do this with a Vision Transformer. We divide the image into patches, randomly remove most of them, and pass only the visible patch tokens through the encoder. Each token keeps its position information.
After the encoder, we put a learned mask token at each missing position. A smaller decoder now reads the full sequence and predicts pixels:
Notice that the large encoder never processes the mask tokens. With patches and 75% masking, it reads just 64 patch tokens. The smaller decoder handles all 256 positions.
We compare its output with the original pixels at the masked positions:
is the set of masked patch indices and is its size. is the original patch, is its prediction, and is the number of pixel values per patch. For a RGB patch, . MAE also studies targets normalized within each patch.
If we hide only a small patch, its neighbours may reveal almost everything. Removing 75% makes the network combine information from farther away, while also giving the encoder fewer tokens to process.
After training, we discard the decoder and give the encoder the full image. We trained it through reconstruction, but we use it to produce features.
The decoder still has to predict details we may not need for recognition. Several handles could fit a partly hidden cup, yet pixel error rewards the one in this photograph. That leads to another approach: compare images without reconstructing their pixels.
Learning by comparison
Let's take two views of an image and ask the network to recognize that they came from the same image. This is the starting point for contrastive learning.
SimCLR makes the views with random crops and colour transformations. We call the two views a positive pair, and pair each view with views from other images to get negatives.
Pass both through the same encoder , then through a small projection network :
Here is a view, is its encoder representation, and is the vector we compare. is the projection network's parameters. Both branches reuse the same weights:
The loss trains both networks. Later we discard and keep . The projection gives the comparison its own feature space, where it can ignore colour changes while allowing to retain colour information.
The contrastive loss
With images in a batch, we have views. Pick view as our anchor and let be its positive partner. The network has to select from the other views. One is the positive; the remaining are negatives.
First, compare with each candidate using cosine similarity:
For nonzero vectors this gives a score between and . It compares their directions, so making a vector longer does not improve its score.
Apply softmax to the scores, after dividing by a temperature :
Both and index candidate views. Notice that the denominator includes the positive partner but excludes itself. We then take the negative log-probability of the correct match:
It is the same cross-entropy loss, with another view's index as the class. Repeat with each of the views as anchor and average the losses. That includes both directions of each positive pair.
SimCLR calls this NT-Xent, short for normalized temperature-scaled cross-entropy. It belongs to the InfoNCE family introduced in Contrastive Predictive Coding.
A small example
Take a photograph of a cup and one of a bicycle. Two views of each give us four vectors. Use the first cup view as the anchor; it needs to find the other cup view among three candidates.
Let's give those candidates similarities of , , and . Set , divide, then apply softmax:
| Candidate | Probability | |
|---|---|---|
| Cup, positive | 1.6 | 0.619 |
| Bicycle view 1 | 0.6 | 0.228 |
| Bicycle view 2 | 0.2 | 0.153 |
The positive probability and loss are:
If all scores were equal, each probability would be and the loss would be . Here the positive already has the highest score, and the loss asks us to separate it further.
To see how, take the derivative with respect to the positive similarity:
For a negative candidate :
The two negative candidates get gradients of approximately and . At the level of these scores, gradient descent pulls the positive up and pushes the negatives down. The bicycle view that looks more similar gets the larger correction. Backpropagation then carries these derivatives into the vectors and shared weights.
Lower to without changing the vectors, and the positive probability becomes about . Temperature makes the distribution sharper. It would also make a wrong match more confident if that match had the highest score, so lowering the loss this way does not mean we learned better features.
Which differences should disappear?
By making two transformed views agree, we ask the comparison to ignore those transformations. This is the invariance we are trying to learn.
For digit recognition, asking a rotated 6 to match a 9 would erase a distinction we need. Cropping also has a limit: one view might contain a cup handle while the other contains only the table. We need views that still share useful content.
A second photograph of a cup also becomes a negative, even though our later classifier may want both cups in the same class. That is a false negative relative to the later task. During pretraining, we only know which original image each view came from.
More negatives without a huge batch
In SimCLR the negatives come from the current batch. To compare with more images, we need a larger batch and more memory for its activations.
MoCo keeps encoded views from recent batches in a queue. We compute a query for the current image and compare it with its positive key and the queued negative keys:
Unlike SimCLR, the two encoders have separate weights. We train the query encoder through gradients and update the key encoder with a moving average:
and are their parameters, and . With close to 1, the key encoder changes slowly.
That slow change matters because the queue contains vectors computed with older weights. We want old and new keys to remain comparable. After using the queue, add the new keys and remove the oldest ones. No gradient passes through the keys.
We can now increase the number of negatives without increasing the current batch. MoCo v2 combines this queue with SimCLR's projection head and stronger transformations.
Predicting a sequence in feature space
We can use the same comparison to predict what comes next in a sequence. This is Contrastive Predictive Coding, or CPC.
Encode each observation into , then use an autoregressive network to collect everything up to time into context . Given that context, we ask it to identify the real future observation among sampled alternatives.
For the observation steps ahead, compute a score with a learned matrix :
Here is the current step and is how far ahead we predict. There is a separate for each distance. Exponentiate and normalize the candidate scores as before, then apply the loss to the real future observation.
The prediction happens in feature space, so we do not have to reconstruct every pixel or audio sample. Keep future targets out of the context network or it can read the answer. CPC also applies this to images by treating rows of patches as a sequence, predicting lower rows from upper ones.
What if we remove the negatives?
So far we have pulled positive pairs together and pushed negatives apart. If we keep only the first part, the network can output one constant vector for every image. All pairs agree and the loss is solved. This is collapse.
To learn without negatives, we need to change how the two branches train.
BYOL adds a predictor to the online branch and trains it to match a target branch on another view. As in MoCo, a moving average updates the target weights. The target receives no gradient.
SimSiam uses shared weights and an extra predictor, but removes the moving-average encoder. It stops the gradient on the branch that supplies the target.
With stop-gradient, the forward pass still uses the target value, but backpropagation treats it as constant. The predictor and the stopped path make training asymmetric. Removing only the negatives would not give us either algorithm; stopping a gradient by itself does not guarantee that we avoid collapse.
DINO
DINO keeps a student and a teacher. We train the student to predict the teacher's output distribution on another view of the image. The teacher follows a moving average of the student, so neither starts with labels. This is self-distillation.
Give the teacher large global crops and the student both global crops and smaller local ones. The student then has to relate a local view to the content the teacher sees in a larger view.
Let and be their output vectors, each with learned dimensions. These are not named classes. Apply softmax to get the distributions:
and are the two temperatures. Before the teacher's softmax, subtract a moving average of its outputs to centre them. For a pair of views, the loss is:
indexes the output dimensions and stops the gradient. Backpropagation updates the student. The moving average then updates the teacher.
We could still collapse by selecting the same output dimension for every image, or by predicting a uniform distribution for all images.
Centering reduces the tendency for one dimension to dominate. A low teacher temperature sharpens the targets, working against the uniform solution. DINO combines both to avoid collapse in its experiments, without a negative-pair term.
After training, take the backbone features before the output head. In DINO's Vision Transformers, some self-attention maps separate foreground objects from the background even though training never used segmentation labels. The maps show spatial structure the model learned; we still need a segmentation head to predict a full set of semantic labels.
DINOv2 builds on this with selected training data, objectives at both image and patch level, and larger-scale training.
Another source of supervision: sound
We can also take two observations of the same event. Look, Listen and Learn takes an image and an audio segment, then predicts whether they belong together. Sample matching pairs from a video and make mismatched pairs with audio from elsewhere.
The visual and audio encoders have to find related content without anyone naming the object or the sound.
A guitar recording might have speech over it, or the sound might come from outside the frame. The pairing gives us useful supervision, but it is noisier than a manually checked label.
Using the encoder
Now take the trained encoder and attach a classifier for the task we wanted to solve. This is our downstream task, and we evaluate it on held-out examples.
First, freeze the encoder and train just a linear classifier. This is a linear probe:
Only and change. is the feature vector, , , and is the number of labelled classes. If this works, a linear rule can separate the classes in the features we already learned.
Unfreeze some or all of the encoder and we get fine-tuning. The features can now change to fit the task. A model that does better when frozen may not do better after fine-tuning, as MAE's experiments show.
For detection or segmentation, attach the corresponding head and use spatial features. Separating image classes with one vector does not test whether the model preserved object boundaries.
To compare the training methods, use the same encoder, pretraining data, and number of labelled examples. Keep the test images separate from pretraining too. Otherwise a score may reflect the extra data instead of the learning method.
In the next notes on generative models, we return to reconstruction and generation. There, the generated examples are the output we care about.