On vision and language

This is part #16 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes are based on lecture 16: Multi-Modal Foundation Models by Ranjay Krishna, with the original papers linked below.


These notes build on self-supervised learning and attention. We now want to connect what an image shows with what a sentence describes.

Our image classifiers have so far ended with a learned output layer, with one score for each class. To add a class, we need to change that layer and train it.

What if we could produce the classifier weights from the class names themselves? We could then describe a new set of categories and use the same model to compare them with an image.

This is one way to use a foundation model. We pretrain it on broad data, then reuse it for different tasks. Sometimes we update its weights. Here, we will start by keeping the trained model fixed and changing the text we give it.

CLIP

In SimCLR, we took two views of the same image and learned to bring their representations together. Let's replace one of those views with the text that accompanied the image.

That is the pairing used by CLIP, or Contrastive Language-Image Pretraining, by Radford and colleagues. We send the image through an image encoder and the text through a text encoder. Each encoder has its own weights. Their outputs have the same dimension, so we can compare them with a dot product.

During training, we want the image to match its own text more closely than the other texts in the batch. We also want the text to match its own image more closely than the other images.

Notice that the image encoder never reads the text, and the text encoder never reads the image. There is no cross-attention between them in CLIP. We combine their output vectors to compute a score:

CLIP dual-encoder architecture and symmetric loss A batch of N paired images and texts passes through separate image and text encoders. Each branch applies a learned projection and L2 normalization. All image and text vectors are compared by a dot product divided by temperature to form an N by N logit matrix. Three pairs are shown. Green diagonal entries are the positive pair targets. Row softmax and cross-entropy select each image's paired text; column softmax and cross-entropy select each text's paired image. The two losses are averaged, and gradients update both encoders. There is no image-text cross-attention in CLIP. N imagesN paired texts Image encodertrainable Text encodertrainable Learned projection+ L2 normalization Learned projection+ L2 normalization u₁, …, uNv₁, …, vN All pairs of dot productszᵢⱼ = uᵢᵀvⱼ / τ N × N logits · 3 pairs shown t₁t₂t₃ z₁₁z₁₂z₁₃ z₂₁z₂₂z₂₃ z₃₁z₃₂z₃₃ x₁x₂x₃ Green diagonal = paired samples Image → textRow softmaxCross-entropy on pairs Text → imageColumn softmaxCross-entropy on pairs Average the two losses No image-text cross-attention
CLIP training. Each encoder runs independently. The green diagonal contains the paired images and texts; both loss directions use it as their target.

The supervision comes from the image-text pairs. People wrote the captions, alt text, and descriptions, but did not have to assign each image a label from a fixed set of classes.

The matching problem

Let's put names on these operations. We have a batch of NN pairs (xi,ti)(x_i,t_i), where xix_i is an image and tit_i is its text. Call the image encoder and its learned projection ff, and the text encoder and its learned projection gg. Normalize each output:

ui=f(xi)∥f(xi)∥2,vi=g(ti)∥g(ti)∥2u_i=\frac{f(x_i)}{\|f(x_i)\|_2},\qquad v_i=\frac{g(t_i)}{\|g(t_i)\|_2}

Both uiu_i and viv_i are unit vectors with dimension DD. We now take the dot product for every image ii and text jj:

Sij=uiTvj,Zij=Sij/τS_{ij}=u_i^Tv_j,\qquad Z_{ij}=S_{ij}/\tau

SS is our N×NN\times N cosine-similarity matrix. Dividing by the temperature τ>0\tau>0 gives the logits ZZ. The CLIP implementation learns a positive scale equivalent to 1/τ1/\tau along with the encoder weights.

The correct pairs lie on the diagonal. There are NN positives and N2−NN^2-N off-diagonal pairings treated as negatives.

First, apply softmax across each row. We get a distribution over texts for a given image. Then apply softmax down each column to get a distribution over images for a given text:

pijrow=eZij∑k=1NeZikp^{\mathrm{row}}_{ij} =\frac{e^{Z_{ij}}}{\sum_{k=1}^{N}e^{Z_{ik}}} pijcol=eZij∑k=1NeZkjp^{\mathrm{col}}_{ij} =\frac{e^{Z_{ij}}}{\sum_{k=1}^{N}e^{Z_{kj}}}

In the first denominator, kk runs over texts. In the second, it runs over images. Now we apply cross-entropy to the diagonal and average over the batch:

LI→T=−1N∑i=1Nlog⁡piirowL_{I\to T}=-\frac1N\sum_{i=1}^{N}\log p^{\mathrm{row}}_{ii} LT→I=−1N∑i=1Nlog⁡piicolL_{T\to I}=-\frac1N\sum_{i=1}^{N}\log p^{\mathrm{col}}_{ii} Lcontrast=12(LI→T+LT→I)L_{\mathrm{contrast}}=\frac12(L_{I\to T}+L_{T\to I})

We average those two losses to get CLIP's symmetric contrastive objective. Backpropagation then updates both encoders, their projections, and the logit scale. The two encoders run separately, but they learn from the same comparisons.

A batch with two pairs

For example, take a cup photograph paired with "a ceramic cup", and a bicycle photograph paired with "a bicycle". Suppose our encoders give us:

ImageCup textBicycle text
Cup0.80.2
Bicycle0.10.7

With τ=0.5\tau=0.5, divide by 0.5 to get:

Z=[1.60.40.21.4]Z=\begin{bmatrix}1.6&0.4\\0.2&1.4\end{bmatrix}

For the cup image, the probability assigned to its text is:

p11row=e1.6e1.6+e0.4≈0.7685p^{\mathrm{row}}_{11} =\frac{e^{1.6}}{e^{1.6}+e^{0.4}}\approx0.7685

For the cup text, the probability assigned to its image is:

p11col=e1.6e1.6+e0.2≈0.8022p^{\mathrm{col}}_{11} =\frac{e^{1.6}}{e^{1.6}+e^{0.2}}\approx0.8022

Notice that we use e0.4e^{0.4} in the row denominator but e0.2e^{0.2} in the column denominator. We kept the positive score fixed and changed what it competes against.

The other diagonal probabilities are 0.76850.7685 for the bicycle row and 0.73110.7311 for its column. Using natural logarithms:

LI→T≈0.2633,LT→I≈0.2668L_{I\to T}\approx0.2633,\qquad L_{T\to I}\approx0.2668 Lcontrast≈0.2651L_{\mathrm{contrast}}\approx0.2651

If every score were equal, each candidate would get probability 1/21/2, and the loss would be log⁡2≈0.6931\log 2\approx0.6931. Increasing the diagonal scores relative to the other entries reduces the loss.

Using the text encoder as a classifier

Now we can return to the classifier we wanted. Take a class name, put it in a prompt such as "a photo of a cup", and run it through the text encoder. Repeat for every class.

We have produced one vector per class, which is exactly what we need to score an image. Compare the image vector with each text vector and take the largest dot product:

For CC categories with normalized text vectors vcv_c:

c^=arg⁡max⁡c∈{1,…,C} uTvc\hat{c}=\underset{c\in\{1,\ldots,C\}}{\arg\max}\ u^Tv_c

This is zero-shot classification. We have not trained on labelled examples from this target task. The concepts can still have appeared in pretraining.

The text vectors are acting as our classifier weights. To change the output categories, we encode different descriptions. This is what makes the classifier open-vocabulary.

We can use several descriptions per class too. The CLIP paper averages their embeddings and normalizes the result. That reduces our dependence on one choice of words.

For retrieval, we keep the same comparison and reverse what we search over. Encode a collection of images once, then rank their vectors against the encoded text query.

A softmax still sums to one over whatever candidates we supply. If the image shows a cup but our only choices are "bicycle" and "dog", one must receive the higher score. These scores are not automatically calibrated probabilities that a description is true.

The original CLIP study trained on 400 million image-text pairs and scaled the model and training compute as well. The authors explain why this efficient comparison task helped them train at that scale. Entering a new class name is possible because we have a text encoder. Recognizing that class still depends on what the model learned.

Compositionality

Let's change the task slightly. We now want to distinguish:

  • A blue cup beside a red plate.
  • A red cup beside a blue plate.

Both descriptions contain a cup, a plate, and the same colours. We need to associate each colour with the correct object. If we change "beside" to "behind", we also need the relative positions.

We call this compositionality. Winoground, by Thrush and colleagues, tests it with paired images and captions that use the same words in different orders.

Look back at our two-pair example. Finding a cup is enough to separate it from a bicycle. The loss never asks whether that cup is red or blue. A larger batch gives us more comparisons, but we may still never compare the two descriptions above.

One approach is hard-negative training. Add a plausible but wrong description, such as the one with swapped colours, so the loss has to separate them.

But we have to check what we call a negative. "A red plate beside a blue cup" still describes the first scene. Two photographs in a batch can also both show a ceramic cup. The pairing tells us which text accompanied the image, not every sentence that could describe it.

And an image-level caption may leave out the details entirely. "A kitchen" tells us little about the spoon on the table. With region-level supervision, we pair a description with a location in the image. The model now gets a training signal for that particular region.

CoCa

So far we can compare an image with a sentence, but we cannot generate the sentence. CoCa, or Contrastive Captioner, adds a captioning objective.

Its text decoder first processes text without looking at the image. We use that representation for the contrastive loss. Later decoder layers add cross-attention to visual features, then predict the caption one token at a time.

For caption tokens y1,…,yTy_1,\ldots,y_T and image xx, we sum the next-token losses:

Lcaption=−∑r=1Tlog⁡p(yr∣y<r,x)L_{\mathrm{caption}}=-\sum_{r=1}^{T} \log p(y_r\mid y_{<r},x)

rr is the token position, and y<ry_{<r} contains the earlier tokens. Add the contrastive loss:

L=λcLcontrast+λgLcaptionL=\lambda_cL_{\mathrm{contrast}} +\lambda_gL_{\mathrm{caption}}

The nonnegative weights λc\lambda_c and λg\lambda_g control how much each loss contributes. We now train the model both to compare image-text pairs and to predict text conditioned on the image.

LLaVA

We could also start with a language model that already generates text. What do we need to change so that it can read an image?

LLaVA, by Liu and colleagues, connects a pretrained CLIP image encoder to a pretrained language model with a learned projection.

This time we keep the image's patch features. We want the language model to have access to spatial information, rather than only the pooled vector that CLIP uses for matching.

The original LLaVA uses features from the penultimate layer, the layer before the last. Those patch features feed CLIP's final pooling computation. The final patch outputs themselves have no separate loss that trains them to represent each region. Taking the earlier features worked better in LLaVA's comparison.

Call the extracted features H∈RM×DvH\in\mathbb{R}^{M\times D_v}, where MM is the number of patches and DvD_v is the feature width. We multiply each patch vector by the same learned matrix WW:

V=HW,W∈RDv×DℓV=HW,\qquad W\in\mathbb{R}^{D_v\times D_\ell}

The result is V∈RM×DℓV\in\mathbb{R}^{M\times D_\ell}, where DℓD_\ell is the language model's input width. We call these rows visual tokens. They are continuous vectors, with the same width as the text embeddings, so we can put them in the input sequence:

Visual tokens enter a language model An image passes through a vision encoder to produce 256 patch features. A learned projection maps each feature from width 1024 to width 4096. Forty text tokens join the 256 visual tokens, giving 296 input positions. The language model uses this context to predict answer tokens. The dimensions are an example, and additional markers are omitted. 224 × 224 image Vision encoder14 × 14 patches 256 × 1024 features Projection W1024 → 4096 Text embeddingswidth 4096 256 visual tokens40 text tokens 296 positions · same width Language modelattends to image + text context Predict the next answer token
LLaVA's connection. The projection changes the width of each patch vector. The language model receives the visual tokens alongside the text embeddings.

For example, a 224×224224\times224 image split into 14×1414\times14 patches gives:

M=(224/14)2=256M=(224/14)^2=256

If the visual width is 1024 and the language width is 4096, we go from a 256×1024256\times1024 array to 256×4096256\times4096. Add 40 text tokens and we have 296 input positions, excluding any extra markers.

Notice that WW changed the width, but we still have 256 visual tokens. We have not converted them into 256 words, nor compressed them into a single vector. More patches mean a longer sequence for the language model. With video, we have the patch features of many frames to deal with.

Training the projection and the language model

First, freeze both pretrained models and train only WW on image-caption pairs. This is feature alignment. The language model stays fixed, so the projection has to produce visual inputs that help it predict the caption.

Then train WW and the language model together on images, instructions, and answers. This is visual instruction tuning. In the original LLaVA, the image encoder stays frozen through both stages.

For an instruction qq, we predict each answer token using the visual tokens VV, the instruction, and the previous answer tokens:

Lanswer=−∑r=1Tlog⁡pθ(yr∣V,q,y<r)L_{\mathrm{answer}}=-\sum_{r=1}^{T} \log p_\theta(y_r\mid V,q,y_{<r})

θ\theta contains the parameters we are training in that stage. We compute the loss on the assistant's answer tokens, including where to stop. The question and image tokens provide context.

During training, we feed in the previous tokens from the reference answer. This is teacher forcing, as in our RNNs. When we generate an answer, we instead feed each predicted token back into the sequence and predict the next one.

The original work used a language-only GPT-4 to help create instruction data from captions and bounding-box descriptions. GPT-4 did not see the images in that process. If those descriptions leave out a detail, the generated answer can fill it in incorrectly, and that error becomes part of the training data.

Flamingo

Flamingo, published in 2022 before LLaVA, gives us another way to connect vision and language.

Instead of putting every patch vector into the language model's input sequence, we first reduce the visual features to 64 tokens. A trainable Perceiver Resampler does this. Its learned queries attend to the visual features and collect them into a fixed number of outputs.

We then add gated cross-attention blocks between the language model's existing layers. The text states supply queries, while the resampled visual tokens supply keys and values. This is the cross-attention operation we already know:

Flamingo connects frozen models through trained cross-attention An image or video passes through a frozen vision encoder. A trainable Perceiver Resampler reduces the resulting visual features to 64 visual tokens. Text passes through frozen language-model blocks. A new trainable gated cross-attention and feed-forward block between language blocks reads the visual tokens as keys and values, using the text states as queries. The language model predicts the next answer token. One inserted block is shown; the model repeats this pattern. The original vision encoder and language layers remain frozen. Image or videoText embeddings Vision encoderfrozen Language blockfrozen PerceiverResamplertrainable Q 64 visual tokenskeys and values Gated cross-attn+ feed-forwardtrainable Language blockfrozen Next answer token One inserted blockis shown.
Flamingo keeps the pretrained image encoder and language layers frozen. The resampler and the new cross-attention blocks learn. One inserted block is shown; the full model repeats this pattern.

The gate starts at zero, so the new visual path initially contributes nothing. Training opens that path while the pretrained image encoder and language layers stay frozen. We update the resampler and the added cross-attention blocks.

Training sequences mix text with images or video. At each text position, cross-attention reads the most recent preceding visual input. The language model's self-attention can still carry information from earlier inputs.

At inference, we can give it a few examples of the task, each containing an image, a question, and an answer. Then append a new image and question. This is in-context learning. The examples become context for generation, without another weight update.

Compare this with LLaVA's diagram. LLaVA adds projected visual tokens to the input sequence. Flamingo keeps them on a separate path and lets text states read them through cross-attention. Its 64-token representation limits the amount of visual data passed into those blocks, but also asks the resampler to decide which details to preserve.

Grounding

Suppose we ask the model how many cups are in an image and it answers "three". We have the count, but we cannot see which cups it counted.

With Molmo, by Deitke and colleagues, we can ask the model to point to the cups as it counts. This connects the answer to locations in the image, which we call grounding.

To train that behaviour, we need data that contains those locations. PixMo includes pointing annotations as well as detailed image descriptions. People describe images aloud, and those recordings become text. This gives the model descriptions of visible details that an incidental web caption may leave out.

Now we can inspect the points. Did the model miss a cup? Did it point to the same cup twice? Did it count a bowl? We can check the answer against the image and pass the locations to another model.

Molmo and PixMo also make the distinction between open weights and open training data concrete. With the weights we can run the model. With the data, code, and evaluations, we can inspect how it learned to produce these answers.

A point can become a mask

For example, we can give the point to Segment Anything, or SAM. Its image encoder computes features for the image. A prompt encoder takes the point, box, or input mask. Then a mask decoder uses both to predict the selected region.

But a point on a person's shirt might refer to the shirt or to the whole person. To handle this ambiguity, SAM predicts several candidate masks and a quality score for each one. During training, it backpropagates the smallest mask loss among the candidates.

The standard released SAM takes spatial prompts, so we still need a grounding model to turn "the cup behind the plate" into a point or box. The lecture shows this combination with Molmo and SAM 2.

We now have two predictions to check. The point must select the intended object, and the mask must follow its boundary. If the point selects a bowl, a perfect bowl mask does not fix the mistake.

Did the model use the image?

For CLIP retrieval, we can rank the results and measure Recall@K, the fraction of queries with a correct match in the first KK results. Checking a generated answer takes more work. Two correct descriptions can use different words, while a wrong description can sound perfectly plausible.

Let's keep the question fixed and change the image. Move the cup to the other side of the plate, or remove one of the cups. The answer should change with the image. If it does not, the model may be completing a likely sentence rather than using the visual information the question requires.

When the model includes objects or details that the input does not support, we call that a hallucination. Recall our answer loss: we train the model to predict the reference tokens given the image and question. That does not guarantee that each generated claim follows from the pixels. The LLaVA paper discusses these errors.

So we test object recognition, reading text, counting, spatial relations, and grounding separately. A high score on recognizing objects can hide poor performance on their relationships, as we saw with Winoground.

The setup matters too. A model that sees a larger image may be able to read text that disappeared at a lower resolution. Examples in the prompt can teach the expected answer format. And an image seen during training cannot establish generalization to a new image. Human or model judges help with open answers, but we still need to inspect which tasks failed.

Chaining models

We used a point from Molmo as an input to SAM. We can connect other models this way too.

In CuPL, we ask a language model to describe a category, then use those descriptions as CLIP prompts. This gives the text encoder information about an object's appearance instead of only its name. We do not train a new classifier. We generate new text vectors for the comparison, so wrong descriptions can also produce wrong matches.

With VisProg, the language model writes a program that calls existing visual modules. For a question about two images, the program can find objects in each image, count them, then compare the counts.

Each step leaves an output we can inspect. If the final count is wrong, we can look at the detected objects before checking the arithmetic. If the program chose the wrong region, we can trace that back to the instruction and the step that selected it. We can now ask a visual question, run the operations needed to answer it, and inspect where the answer came from.

← Back to blog