On vision and language
This is part #16 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on lecture 16: Multi-Modal Foundation Models by Ranjay Krishna, with the original papers linked below.
These notes build on self-supervised learning and attention. We now want to connect what an image shows with what a sentence describes.
Our image classifiers have so far ended with a learned output layer, with one score for each class. To add a class, we need to change that layer and train it.
What if we could produce the classifier weights from the class names themselves? We could then describe a new set of categories and use the same model to compare them with an image.
This is one way to use a foundation model. We pretrain it on broad data, then reuse it for different tasks. Sometimes we update its weights. Here, we will start by keeping the trained model fixed and changing the text we give it.
CLIP
In SimCLR, we took two views of the same image and learned to bring their representations together. Let's replace one of those views with the text that accompanied the image.
That is the pairing used by CLIP, or Contrastive Language-Image Pretraining, by Radford and colleagues. We send the image through an image encoder and the text through a text encoder. Each encoder has its own weights. Their outputs have the same dimension, so we can compare them with a dot product.
During training, we want the image to match its own text more closely than the other texts in the batch. We also want the text to match its own image more closely than the other images.
Notice that the image encoder never reads the text, and the text encoder never reads the image. There is no cross-attention between them in CLIP. We combine their output vectors to compute a score:
The supervision comes from the image-text pairs. People wrote the captions, alt text, and descriptions, but did not have to assign each image a label from a fixed set of classes.
The matching problem
Let's put names on these operations. We have a batch of pairs , where is an image and is its text. Call the image encoder and its learned projection , and the text encoder and its learned projection . Normalize each output:
Both and are unit vectors with dimension . We now take the dot product for every image and text :
is our cosine-similarity matrix. Dividing by the temperature gives the logits . The CLIP implementation learns a positive scale equivalent to along with the encoder weights.
The correct pairs lie on the diagonal. There are positives and off-diagonal pairings treated as negatives.
First, apply softmax across each row. We get a distribution over texts for a given image. Then apply softmax down each column to get a distribution over images for a given text:
In the first denominator, runs over texts. In the second, it runs over images. Now we apply cross-entropy to the diagonal and average over the batch:
We average those two losses to get CLIP's symmetric contrastive objective. Backpropagation then updates both encoders, their projections, and the logit scale. The two encoders run separately, but they learn from the same comparisons.
A batch with two pairs
For example, take a cup photograph paired with "a ceramic cup", and a bicycle photograph paired with "a bicycle". Suppose our encoders give us:
| Image | Cup text | Bicycle text |
|---|---|---|
| Cup | 0.8 | 0.2 |
| Bicycle | 0.1 | 0.7 |
With , divide by 0.5 to get:
For the cup image, the probability assigned to its text is:
For the cup text, the probability assigned to its image is:
Notice that we use in the row denominator but in the column denominator. We kept the positive score fixed and changed what it competes against.
The other diagonal probabilities are for the bicycle row and for its column. Using natural logarithms:
If every score were equal, each candidate would get probability , and the loss would be . Increasing the diagonal scores relative to the other entries reduces the loss.
Using the text encoder as a classifier
Now we can return to the classifier we wanted. Take a class name, put it in a prompt such as "a photo of a cup", and run it through the text encoder. Repeat for every class.
We have produced one vector per class, which is exactly what we need to score an image. Compare the image vector with each text vector and take the largest dot product:
For categories with normalized text vectors :
This is zero-shot classification. We have not trained on labelled examples from this target task. The concepts can still have appeared in pretraining.
The text vectors are acting as our classifier weights. To change the output categories, we encode different descriptions. This is what makes the classifier open-vocabulary.
We can use several descriptions per class too. The CLIP paper averages their embeddings and normalizes the result. That reduces our dependence on one choice of words.
For retrieval, we keep the same comparison and reverse what we search over. Encode a collection of images once, then rank their vectors against the encoded text query.
A softmax still sums to one over whatever candidates we supply. If the image shows a cup but our only choices are "bicycle" and "dog", one must receive the higher score. These scores are not automatically calibrated probabilities that a description is true.
The original CLIP study trained on 400 million image-text pairs and scaled the model and training compute as well. The authors explain why this efficient comparison task helped them train at that scale. Entering a new class name is possible because we have a text encoder. Recognizing that class still depends on what the model learned.
Compositionality
Let's change the task slightly. We now want to distinguish:
- A blue cup beside a red plate.
- A red cup beside a blue plate.
Both descriptions contain a cup, a plate, and the same colours. We need to associate each colour with the correct object. If we change "beside" to "behind", we also need the relative positions.
We call this compositionality. Winoground, by Thrush and colleagues, tests it with paired images and captions that use the same words in different orders.
Look back at our two-pair example. Finding a cup is enough to separate it from a bicycle. The loss never asks whether that cup is red or blue. A larger batch gives us more comparisons, but we may still never compare the two descriptions above.
One approach is hard-negative training. Add a plausible but wrong description, such as the one with swapped colours, so the loss has to separate them.
But we have to check what we call a negative. "A red plate beside a blue cup" still describes the first scene. Two photographs in a batch can also both show a ceramic cup. The pairing tells us which text accompanied the image, not every sentence that could describe it.
And an image-level caption may leave out the details entirely. "A kitchen" tells us little about the spoon on the table. With region-level supervision, we pair a description with a location in the image. The model now gets a training signal for that particular region.
CoCa
So far we can compare an image with a sentence, but we cannot generate the sentence. CoCa, or Contrastive Captioner, adds a captioning objective.
Its text decoder first processes text without looking at the image. We use that representation for the contrastive loss. Later decoder layers add cross-attention to visual features, then predict the caption one token at a time.
For caption tokens and image , we sum the next-token losses:
is the token position, and contains the earlier tokens. Add the contrastive loss:
The nonnegative weights and control how much each loss contributes. We now train the model both to compare image-text pairs and to predict text conditioned on the image.
LLaVA
We could also start with a language model that already generates text. What do we need to change so that it can read an image?
LLaVA, by Liu and colleagues, connects a pretrained CLIP image encoder to a pretrained language model with a learned projection.
This time we keep the image's patch features. We want the language model to have access to spatial information, rather than only the pooled vector that CLIP uses for matching.
The original LLaVA uses features from the penultimate layer, the layer before the last. Those patch features feed CLIP's final pooling computation. The final patch outputs themselves have no separate loss that trains them to represent each region. Taking the earlier features worked better in LLaVA's comparison.
Call the extracted features , where is the number of patches and is the feature width. We multiply each patch vector by the same learned matrix :
The result is , where is the language model's input width. We call these rows visual tokens. They are continuous vectors, with the same width as the text embeddings, so we can put them in the input sequence:
For example, a image split into patches gives:
If the visual width is 1024 and the language width is 4096, we go from a array to . Add 40 text tokens and we have 296 input positions, excluding any extra markers.
Notice that changed the width, but we still have 256 visual tokens. We have not converted them into 256 words, nor compressed them into a single vector. More patches mean a longer sequence for the language model. With video, we have the patch features of many frames to deal with.
Training the projection and the language model
First, freeze both pretrained models and train only on image-caption pairs. This is feature alignment. The language model stays fixed, so the projection has to produce visual inputs that help it predict the caption.
Then train and the language model together on images, instructions, and answers. This is visual instruction tuning. In the original LLaVA, the image encoder stays frozen through both stages.
For an instruction , we predict each answer token using the visual tokens , the instruction, and the previous answer tokens:
contains the parameters we are training in that stage. We compute the loss on the assistant's answer tokens, including where to stop. The question and image tokens provide context.
During training, we feed in the previous tokens from the reference answer. This is teacher forcing, as in our RNNs. When we generate an answer, we instead feed each predicted token back into the sequence and predict the next one.
The original work used a language-only GPT-4 to help create instruction data from captions and bounding-box descriptions. GPT-4 did not see the images in that process. If those descriptions leave out a detail, the generated answer can fill it in incorrectly, and that error becomes part of the training data.
Flamingo
Flamingo, published in 2022 before LLaVA, gives us another way to connect vision and language.
Instead of putting every patch vector into the language model's input sequence, we first reduce the visual features to 64 tokens. A trainable Perceiver Resampler does this. Its learned queries attend to the visual features and collect them into a fixed number of outputs.
We then add gated cross-attention blocks between the language model's existing layers. The text states supply queries, while the resampled visual tokens supply keys and values. This is the cross-attention operation we already know:
The gate starts at zero, so the new visual path initially contributes nothing. Training opens that path while the pretrained image encoder and language layers stay frozen. We update the resampler and the added cross-attention blocks.
Training sequences mix text with images or video. At each text position, cross-attention reads the most recent preceding visual input. The language model's self-attention can still carry information from earlier inputs.
At inference, we can give it a few examples of the task, each containing an image, a question, and an answer. Then append a new image and question. This is in-context learning. The examples become context for generation, without another weight update.
Compare this with LLaVA's diagram. LLaVA adds projected visual tokens to the input sequence. Flamingo keeps them on a separate path and lets text states read them through cross-attention. Its 64-token representation limits the amount of visual data passed into those blocks, but also asks the resampler to decide which details to preserve.
Grounding
Suppose we ask the model how many cups are in an image and it answers "three". We have the count, but we cannot see which cups it counted.
With Molmo, by Deitke and colleagues, we can ask the model to point to the cups as it counts. This connects the answer to locations in the image, which we call grounding.
To train that behaviour, we need data that contains those locations. PixMo includes pointing annotations as well as detailed image descriptions. People describe images aloud, and those recordings become text. This gives the model descriptions of visible details that an incidental web caption may leave out.
Now we can inspect the points. Did the model miss a cup? Did it point to the same cup twice? Did it count a bowl? We can check the answer against the image and pass the locations to another model.
Molmo and PixMo also make the distinction between open weights and open training data concrete. With the weights we can run the model. With the data, code, and evaluations, we can inspect how it learned to produce these answers.
A point can become a mask
For example, we can give the point to Segment Anything, or SAM. Its image encoder computes features for the image. A prompt encoder takes the point, box, or input mask. Then a mask decoder uses both to predict the selected region.
But a point on a person's shirt might refer to the shirt or to the whole person. To handle this ambiguity, SAM predicts several candidate masks and a quality score for each one. During training, it backpropagates the smallest mask loss among the candidates.
The standard released SAM takes spatial prompts, so we still need a grounding model to turn "the cup behind the plate" into a point or box. The lecture shows this combination with Molmo and SAM 2.
We now have two predictions to check. The point must select the intended object, and the mask must follow its boundary. If the point selects a bowl, a perfect bowl mask does not fix the mistake.
Did the model use the image?
For CLIP retrieval, we can rank the results and measure Recall@K, the fraction of queries with a correct match in the first results. Checking a generated answer takes more work. Two correct descriptions can use different words, while a wrong description can sound perfectly plausible.
Let's keep the question fixed and change the image. Move the cup to the other side of the plate, or remove one of the cups. The answer should change with the image. If it does not, the model may be completing a likely sentence rather than using the visual information the question requires.
When the model includes objects or details that the input does not support, we call that a hallucination. Recall our answer loss: we train the model to predict the reference tokens given the image and question. That does not guarantee that each generated claim follows from the pixels. The LLaVA paper discusses these errors.
So we test object recognition, reading text, counting, spatial relations, and grounding separately. A high score on recognizing objects can hide poor performance on their relationships, as we saw with Winoground.
The setup matters too. A model that sees a larger image may be able to read text that disappeared at a lower resolution. Examples in the prompt can teach the expected answer format. And an image seen during training cannot establish generalization to a new image. Human or model judges help with open answers, but we still need to inspect which tasks failed.
Chaining models
We used a point from Molmo as an input to SAM. We can connect other models this way too.
In CuPL, we ask a language model to describe a category, then use those descriptions as CLIP prompts. This gives the text encoder information about an object's appearance instead of only its name. We do not train a new classifier. We generate new text vectors for the comparison, so wrong descriptions can also produce wrong matches.
With VisProg, the language model writes a program that calls existing visual modules. For a question about two images, the program can find objects in each image, count them, then compare the counts.
Each step leaves an output we can inspect. If the final count is wrong, we can look at the detected objects before checking the arithmetic. If the program chose the wrong region, we can trace that back to the instruction and the step that selected it. We can now ask a visual question, run the operations needed to answer it, and inspect where the answer came from.