On video understanding
This is part #11 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on lecture 10: Video Understanding by Ruohan Gao.
With a CNN, we can recognize what is in a frame. To recognize someone picking up a cup, we also need to relate that frame to the ones before and after it.
We already have most of the pieces: convolutions for image features, RNNs to carry a state through a sequence, and attention to look back at other positions. Let's see how we can put them together for video.
Video = images + time
I'll use a channels-first shape for one clip:
is the channel count, is the number of frames, and are their height and width. For RGB, . A batch of clips adds a leading dimension .
Notice that time has its own axis. A temporal filter will move along , just as a spatial filter moves across height and width. The channels hold RGB values or learned features at each position.
For now, let's classify the whole clip. We want class scores, with classes such as "opening a door" and "closing a door".
Which frames do we keep?
We usually cannot pass an entire long video through the network at once. During training, take a short clip. At test time, we can sample several clips and combine their predictions.
Suppose the video has frames per second. Take frames, each source frames after the previous one. The first and last samples are this far apart:
For , , and , we take frames . They span two seconds. With , the same number of frames covers only half a second.
Increasing the spacing lets us see further in time, but we may skip a fast movement. Taking more frames preserves the spacing and costs more computation.
Those 16 RGB frames at already contain 2,408,448 values, about 9.6 MB at four bytes each. During training we also keep activations and gradients, so the input is only part of the memory cost.
Start with an image model
Start by applying the same 2D CNN to every frame independently. A tennis court, an instrument, or a body pose already tells us something about the action. Karpathy et al. found that single-frame models could work surprisingly well.
Pool over the spatial positions and frame gives us a vector:
is the frame and is the feature dimension. Average the vectors, then apply a linear classifier:
and are the classifier's parameters. has one score per class. For a correct class , apply softmax and cross-entropy as before:
is the probability of class . When we backpropagate, the gradient passes through the average into every frame's features. All frames contribute to updating the shared CNN weights.
We combined the frames after computing their features, so this is late fusion. We could also classify each frame and average its probabilities. That gives a different result, since softmax is not linear.
Losing the order
Reverse the frames and the average stays the same. Both and give us the same :
So this network cannot tell picking up a cup from putting it down if the two clips contain exactly the same frames in reverse order.
Concatenating the vectors in order would let an MLP distinguish their positions. But we still computed each frame's features separately. Let's combine the frames earlier, while the network can compare their pixels.
Early fusion: combine the frames first
Stack the RGB channels of all frames. We turn into and pass that through a 2D convolution. Its first layer now reads all frames at once. This is early fusion, one of the approaches in the video CNN paper.
Eight RGB frames give us 24 input channels. A filter can now compare a patch in frame 1 with the same patch in frame 2.
But after that layer we only have a spatial feature map. The time axis disappeared into the channels, so later filters cannot slide through it. To learn a pattern and reuse it at different times, we need to keep time as an axis.
3D convolutions: combine frames gradually
A 3D convolution does that by giving the filter a temporal size as well as its spatial size . At each position, it reads a small volume across all input channels.
For input channels and output channels, the weights have shape:
This is an ordinary convolution without groups. We slide the same weights across time, height, and width, reusing a pattern at different times as well as different image locations. The axes follow Conv3d's convention.
The third dimension here is time. We are still processing 2D images, without a measurement of the scene's depth.
A small network
Let's use eight RGB frames. Apply a convolution with eight output channels, stride 1, and padding 1 along time, height, and width.
The parameter count, including biases, is:
The equivalent 2D layer has parameters. Extending the filter through time adds weights, and applying it at every time position adds computation.
Stack two of these convolutions, with a ReLU after each:
| Layer | Output shape | Temporal receptive field |
|---|---|---|
| Input | 1 frame | |
| 3D conv, 8 channels | 3 frames | |
| 3D conv, 8 channels | 5 frames | |
| Average over time and space | features | Entire clip |
After the first layer, a feature can depend on three frames. The next layer combines three of those features, extending its receptive field to five frames and a image region. At the boundaries, part of that region is padding.
We still average at the end, but this time the features already include local changes across frames.
We can shorten time with stride or pooling too. A pool with stride 2 halves time, height, and width. Later layers have fewer positions to process, and fewer time positions at which to describe the action.
Optical flow: give the model motion directly
Instead of asking the classifier to learn motion from RGB values, we can estimate it first. Optical flow gives each pixel an apparent displacement between two frames:
Here are image coordinates. and are the horizontal and vertical displacements from frame to frame . A pixel moving three places right has flow .
The two-stream network sends RGB frames through one CNN and stacked flow fields through another. Each produces class probabilities, which we can average:
For eight frames, we have seven adjacent frame pairs. Each flow field has two channels, so the motion input has shape .
This removes much of the texture and colour from the motion branch. The RGB branch can still recognize the cup, while the flow branch sees how its image moves.
Flow also includes camera motion. If the camera pans, even a stationary cup moves in the image. It costs time to compute and is harder to estimate when a region disappears behind something else or has little texture.
What about a longer sequence?
A short clip may show a hand opening a cupboard. To recognize someone making coffee, we may need to connect that action with events much farther apart.
Take a feature from each frame or short clip, then pass the features in order through an RNN:
is the hidden state and is a recurrent unit, such as an LSTM. Each update combines the new visual features with the previous state. Read the final state for a video label, or classify each state for a sequence of predictions. Long-term Recurrent Convolutional Networks combine CNN features and recurrence this way.
We are back to the fixed-size state from the RNN notes. Earlier frames can affect it, but details can be lost as we keep updating it. Each step also waits for the state from the previous step.
We do not have to flatten the spatial features into a vector first. Replace the recurrent unit's dense transformations with convolutions, and its hidden state can remain a feature map. Ballas et al. use these convolutional recurrent units to carry spatial features through time.
Attention across space and time
As with attention in an RNN, we can let the network look back at other positions instead of carrying everything through one state. Compute query-key similarities, apply softmax, and take the weighted sum of the values.
A non-local block does this inside a CNN. Take a feature volume and rearrange it into vectors. The primes are the reduced sizes at this layer. Each output can now combine information from other times and image locations, as in Wang et al..
We can also start with a transformer. Split each frame into patches, project each patch into a -dimensional token, and add information about where and when it came from. That gives us:
This count assumes and divide evenly by and excludes any extra classification token.
Counting the attention cost
For 16 frames at , using patches:
Full attention compares every token with every token: query-key pairs per head. Double the frames and we get four times as many pairs.
TimeSformer divides this into two operations. First attend through time at each patch location. Then attend across the image within each frame.
For our example, the pair counts become:
Together that is 664,832 pairs, about 15 times fewer. The projections, MLP, and any classification tokens add their own work, so this ratio only describes the attention pairs.
After both operations, information has moved through time and space, using fewer comparisons than joint attention. Notice that the temporal step stays at one patch location. An object can move into another patch, so that step alone is not tracking it.
Reusing a 2D network
We do not have to start a 3D CNN from random weights. I3D takes a pretrained 2D kernel and copies it into each of temporal slices, dividing each copy by .
If all frames are identical, those slices add back to the original spatial response wherever the temporal window contains real frames. We start with the image features and let training learn how to use time. This is the filter inflation in Carreira and Zisserman.
R(2+1)D splits a 3D convolution into a spatial operation followed by a temporal one, with a non-linearity between them:
The intermediate channel count controls the parameter count. We can choose it to match a 3D layer's count; splitting the operations does not automatically save parameters. Tran et al. study this separation.
SlowFast changes how we sample the video. Give one RGB pathway fewer frames and more channels to describe appearance. Give a second pathway more frames and fewer channels to follow rapid changes. Add connections between them so they can exchange features. Both read RGB, unlike the RGB and flow pair above. Feichtenhofer et al..
We can pretrain without action labels too. VideoMAE hides the same spatial locations across a clip and reconstructs the missing content. Hiding tubes through time stops the model from copying the same patch from an adjacent frame. We then fine-tune the encoder for the action task. Tong et al..
From a label to when and where
So far, one clip gave us one class. Let's put that clip back into a longer video. We now need to say when the action starts and ends.
In temporal action localization, we predict a class and an interval. We can propose candidate intervals, then refine and classify them, much as an object detector works with boxes. The predicted interval has to match the action's timing as well as its class. Chao et al..
Add a box for the person or object and we get spatiotemporal detection. The AVA dataset labels people with boxes and actions at annotated times. One person can be standing, talking, and holding something at once.
Those labels should not compete for one unit of probability as they do with softmax. Give each class a sigmoid and train it as a separate binary prediction:
is the class score, is 1 when that action is present and 0 otherwise, and is its loss. The probabilities can all be high at once and do not need to sum to 1.
A video can have sound too
So far we have used only the frames. A video of a guitar also has sound, which can help us tell whether someone is playing it.
Encode the frames and audio separately, then combine their features or predictions. For the audio, we can use a spectrogram, a grid of frequency content over time. Its steps may differ from the video frame rate; we align both inputs to the interval we want to classify.
We can combine them inside the network too. Attention Bottlenecks give the audio and visual streams a small set of shared latent tokens. Each stream updates a copy of those tokens, then we average the updates. In the next layer, both streams can read what the other contributed. Communication passes through these few tokens rather than every possible audio-visual pair.
Instead of a class, we could predict one speaker's voice from a mixed recording. VisualVoice uses facial appearance and lip motion along with the audio. The face helps identify whose voice to extract, and the lips help align speech with time. This is audio-visual source separation.
Sound can also help us choose frames. Listen to Look uses inexpensive image-audio observations to select moments worth processing with a more expensive visual model. We no longer have to spend the same amount of work on every part of the video.
Finally, connect the visual features to a language model. Video-LLaVA uses this to answer questions and produce descriptions, instead of choosing from a fixed set of action classes.
The language model still only sees what the video encoder retained. If we skipped the moment the cup moved, or averaged away its direction, a more flexible decoder does not restore that evidence.
Next are self-supervised learning notes, where we train visual features by hiding, transforming, and comparing the data itself.