On video understanding

This is part #11 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes are based on lecture 10: Video Understanding by Ruohan Gao.


With a CNN, we can recognize what is in a frame. To recognize someone picking up a cup, we also need to relate that frame to the ones before and after it.

We already have most of the pieces: convolutions for image features, RNNs to carry a state through a sequence, and attention to look back at other positions. Let's see how we can put them together for video.

Video = images + time

I'll use a channels-first shape for one clip:

X∈RC×T×H×WX\in\mathbb{R}^{C\times T\times H\times W}

CC is the channel count, TT is the number of frames, and H,WH,W are their height and width. For RGB, C=3C=3. A batch of BB clips adds a leading dimension BB.

Notice that time has its own axis. A temporal filter will move along TT, just as a spatial filter moves across height and width. The channels hold RGB values or learned features at each position.

For now, let's classify the whole clip. We want KK class scores, with classes such as "opening a door" and "closing a door".

Which frames do we keep?

We usually cannot pass an entire long video through the network at once. During training, take a short clip. At test time, we can sample several clips and combine their predictions.

Suppose the video has ff frames per second. Take TT frames, each ss source frames after the previous one. The first and last samples are this far apart:

Δt=(T−1)sf\Delta t=\frac{(T-1)s}{f}

For T=16T=16, s=4s=4, and f=30f=30, we take frames 0,4,8,…,600,4,8,\ldots,60. They span two seconds. With s=1s=1, the same number of frames covers only half a second.

Increasing the spacing lets us see further in time, but we may skip a fast movement. Taking more frames preserves the spacing and costs more computation.

Those 16 RGB frames at 224×224224\times224 already contain 2,408,448 values, about 9.6 MB at four bytes each. During training we also keep activations and gradients, so the input is only part of the memory cost.

Start with an image model

Start by applying the same 2D CNN to every frame independently. A tennis court, an instrument, or a body pose already tells us something about the action. Karpathy et al. found that single-frame models could work surprisingly well.

Pool over the spatial positions and frame tt gives us a vector:

ft=CNN⁡(Xt)∈RDf_t=\operatorname{CNN}(X_t)\in\mathbb{R}^{D}

XtX_t is the frame and DD is the feature dimension. Average the vectors, then apply a linear classifier:

fˉ=1T∑t=1Tftz=Wfˉ+b\bar f=\frac{1}{T}\sum_{t=1}^{T}f_t \qquad z=W\bar f+b

W∈RK×DW\in\mathbb{R}^{K\times D} and b∈RKb\in\mathbb{R}^{K} are the classifier's parameters. zz has one score per class. For a correct class yy, apply softmax and cross-entropy as before:

pk=ezk∑j=1KezjL=−log⁡pyp_k=\frac{e^{z_k}}{\sum_{j=1}^{K}e^{z_j}} \qquad L=-\log p_y

pkp_k is the probability of class kk. When we backpropagate, the gradient passes through the average into every frame's features. All frames contribute to updating the shared CNN weights.

We combined the frames after computing their features, so this is late fusion. We could also classify each frame and average its probabilities. That gives a different result, since softmax is not linear.

Losing the order

Reverse the frames and the average stays the same. Both (f1,f2,f3)(f_1,f_2,f_3) and (f3,f2,f1)(f_3,f_2,f_1) give us the same fˉ\bar f:

Frame order changes the motion The first row shows a dot moving from the left side to the middle and then to the right side of a frame. The second row reverses the same three frames, so the dot moves left. Averaging independent frame features gives the same representation for both sequences. Moving right Moving left frame 1 frame 2 frame 3 frame 1 frame 2 frame 3
The same three frames, in opposite orders. Independent frame features followed by average pooling cannot distinguish them.

So this network cannot tell picking up a cup from putting it down if the two clips contain exactly the same frames in reverse order.

Concatenating the vectors in order would let an MLP distinguish their positions. But we still computed each frame's features separately. Let's combine the frames earlier, while the network can compare their pixels.

Early fusion: combine the frames first

Stack the RGB channels of all frames. We turn 3×T×H×W3\times T\times H\times W into 3T×H×W3T\times H\times W and pass that through a 2D convolution. Its first layer now reads all frames at once. This is early fusion, one of the approaches in the video CNN paper.

Eight RGB frames give us 24 input channels. A filter can now compare a patch in frame 1 with the same patch in frame 2.

But after that layer we only have a spatial feature map. The time axis disappeared into the channels, so later filters cannot slide through it. To learn a pattern and reuse it at different times, we need to keep time as an axis.

3D convolutions: combine frames gradually

A 3D convolution does that by giving the filter a temporal size ktk_t as well as its spatial size kh×kwk_h\times k_w. At each position, it reads a small volume across all input channels.

For CinC_{\mathrm{in}} input channels and CoutC_{\mathrm{out}} output channels, the weights have shape:

Cout×Cin×kt×kh×kwC_{\mathrm{out}}\times C_{\mathrm{in}} \times k_t\times k_h\times k_w

This is an ordinary convolution without groups. We slide the same weights across time, height, and width, reusing a pattern at different times as well as different image locations. The axes follow Conv3d's convention.

The third dimension here is time. We are still processing 2D images, without a measurement of the scene's depth.

A small network

Let's use eight 32×3232\times32 RGB frames. Apply a 3×3×33\times3\times3 convolution with eight output channels, stride 1, and padding 1 along time, height, and width.

The parameter count, including biases, is:

(3⋅3⋅3⋅3+1)⋅8=656(3\cdot3\cdot3\cdot3+1)\cdot8=656

The equivalent 2D layer has (3⋅3⋅3+1)⋅8=224(3\cdot3\cdot3+1)\cdot8=224 parameters. Extending the filter through time adds weights, and applying it at every time position adds computation.

Stack two of these convolutions, with a ReLU after each:

LayerOutput shapeTemporal receptive field
Input3×8×32×323\times8\times32\times321 frame
3D conv, 8 channels8×8×32×328\times8\times32\times323 frames
3D conv, 8 channels8×8×32×328\times8\times32\times325 frames
Average over time and space88 featuresEntire clip

After the first layer, a feature can depend on three frames. The next layer combines three of those features, extending its receptive field to five frames and a 5×55\times5 image region. At the boundaries, part of that region is padding.

We still average at the end, but this time the features already include local changes across frames.

We can shorten time with stride or pooling too. A 2×2×22\times2\times2 pool with stride 2 halves time, height, and width. Later layers have fewer positions to process, and fewer time positions at which to describe the action.

Optical flow: give the model motion directly

Instead of asking the classifier to learn motion from RGB values, we can estimate it first. Optical flow gives each pixel an apparent displacement between two frames:

Ft(x,y)=(ut(x,y),vt(x,y))F_t(x,y)=(u_t(x,y),v_t(x,y))

Here x,yx,y are image coordinates. utu_t and vtv_t are the horizontal and vertical displacements from frame tt to frame t+1t+1. A pixel moving three places right has flow (3,0)(3,0).

The two-stream network sends RGB frames through one CNN and stacked flow fields through another. Each produces class probabilities, which we can average:

RGB and optical-flow streams A video clip supplies an RGB frame to an appearance CNN and adjacent frame pairs to an optical-flow estimator. Stacked flow fields enter a separate motion CNN. The two CNNs produce class probabilities, which are averaged for the prediction. Video clip sample a frameestimate flow RGB frame3 × H × W Stacked flow2(T − 1) × H × W AppearanceCNN MotionCNN Classprobabilities Classprobabilities Average predictions
One stream reads appearance from RGB. The other reads motion from stacked horizontal and vertical flow fields. Their predictions meet at the end.

For eight frames, we have seven adjacent frame pairs. Each flow field has two channels, so the motion input has shape 14×H×W14\times H\times W.

This removes much of the texture and colour from the motion branch. The RGB branch can still recognize the cup, while the flow branch sees how its image moves.

Flow also includes camera motion. If the camera pans, even a stationary cup moves in the image. It costs time to compute and is harder to estimate when a region disappears behind something else or has little texture.

What about a longer sequence?

A short clip may show a hand opening a cupboard. To recognize someone making coffee, we may need to connect that action with events much farther apart.

Take a feature ftf_t from each frame or short clip, then pass the features in order through an RNN:

ht=g(ft,ht−1)h_t=g(f_t,h_{t-1})

hth_t is the hidden state and gg is a recurrent unit, such as an LSTM. Each update combines the new visual features with the previous state. Read the final state for a video label, or classify each state for a sequence of predictions. Long-term Recurrent Convolutional Networks combine CNN features and recurrence this way.

We are back to the fixed-size state from the RNN notes. Earlier frames can affect it, but details can be lost as we keep updating it. Each step also waits for the state from the previous step.

We do not have to flatten the spatial features into a vector first. Replace the recurrent unit's dense transformations with convolutions, and its hidden state can remain a feature map. Ballas et al. use these convolutional recurrent units to carry spatial features through time.

Attention across space and time

As with attention in an RNN, we can let the network look back at other positions instead of carrying everything through one state. Compute query-key similarities, apply softmax, and take the weighted sum of the values.

A non-local block does this inside a CNN. Take a C×T′×H′×W′C\times T'\times H'\times W' feature volume and rearrange it into N=T′H′W′N=T'H'W' vectors. The primes are the reduced sizes at this layer. Each output can now combine information from other times and image locations, as in Wang et al..

We can also start with a transformer. Split each frame into P×PP\times P patches, project each patch into a DD-dimensional token, and add information about where and when it came from. That gives us:

N=THPWPN=T\frac{H}{P}\frac{W}{P}

This count assumes HH and WW divide evenly by PP and excludes any extra classification token.

Counting the attention cost

For 16 frames at 224×224224\times224, using 16×1616\times16 patches:

S=14⋅14=196patches per frameS=14\cdot14=196\quad\text{patches per frame} N=16⋅196=3136tokensN=16\cdot196=3136\quad\text{tokens}

Full attention compares every token with every token: N2=9,834,496N^2=9,834,496 query-key pairs per head. Double the frames and we get four times as many pairs.

TimeSformer divides this into two operations. First attend through time at each patch location. Then attend across the image within each frame.

For our example, the pair counts become:

temporal: ST2=50,176spatial: TS2=614,656\begin{aligned} \text{temporal: }&ST^2=50,176\\ \text{spatial: }&TS^2=614,656 \end{aligned}

Together that is 664,832 pairs, about 15 times fewer. The projections, MLP, and any classification tokens add their own work, so this ratio only describes the attention pairs.

After both operations, information has moved through time and space, using fewer comparisons than joint attention. Notice that the temporal step stays at one patch location. An object can move into another patch, so that step alone is not tracking it.

Reusing a 2D network

We do not have to start a 3D CNN from random weights. I3D takes a pretrained 2D kernel and copies it into each of ktk_t temporal slices, dividing each copy by ktk_t.

If all frames are identical, those slices add back to the original spatial response wherever the temporal window contains real frames. We start with the image features and let training learn how to use time. This is the filter inflation in Carreira and Zisserman.

R(2+1)D splits a 3D convolution into a spatial operation followed by a temporal one, with a non-linearity between them:

Joint and factorized video convolutions On the left, a 3D convolution has a kernel extending through time, height and width. On the right, R(2+1)D replaces that operation with a spatial convolution of temporal width one, a non-linearity, and a temporal convolution with spatial size one by one. Intermediate channels control the factorized layer's parameter count. Input feature volume 3D convolutionR(2+1)D Space + timek_t × k_h × k_w mixed in one filter Spatial filter1 × k_h × k_w Non-linearity Temporal filterk_t × 1 × 1 Output featuresOutput features Kernel sizes shown as time × height × width
A 3D filter mixes space and time together. R(2+1)D first mixes within each frame, then across frames, with a non-linearity between the operations.

The intermediate channel count controls the parameter count. We can choose it to match a 3D layer's count; splitting the operations does not automatically save parameters. Tran et al. study this separation.

SlowFast changes how we sample the video. Give one RGB pathway fewer frames and more channels to describe appearance. Give a second pathway more frames and fewer channels to follow rapid changes. Add connections between them so they can exchange features. Both read RGB, unlike the RGB and flow pair above. Feichtenhofer et al..

We can pretrain without action labels too. VideoMAE hides the same spatial locations across a clip and reconstructs the missing content. Hiding tubes through time stops the model from copying the same patch from an adjacent frame. We then fine-tune the encoder for the action task. Tong et al..

From a label to when and where

So far, one clip gave us one class. Let's put that clip back into a longer video. We now need to say when the action starts and ends.

In temporal action localization, we predict a class and an interval. We can propose candidate intervals, then refine and classify them, much as an object detector works with boxes. The predicted interval has to match the action's timing as well as its class. Chao et al..

Add a box for the person or object and we get spatiotemporal detection. The AVA dataset labels people with boxes and actions at annotated times. One person can be standing, talking, and holding something at once.

Those labels should not compete for one unit of probability as they do with softmax. Give each class a sigmoid and train it as a separate binary prediction:

qk=sigmoid⁡(zk)=11+e−zkq_k=\operatorname{sigmoid}(z_k)=\frac{1}{1+e^{-z_k}} ℓk=−yklog⁡qk−(1−yk)log⁡(1−qk)\begin{aligned} \ell_k={}&-y_k\log q_k\\ &-(1-y_k)\log(1-q_k) \end{aligned} L=∑k=1KℓkL=\sum_{k=1}^{K}\ell_k

zkz_k is the class score, yky_k is 1 when that action is present and 0 otherwise, and ℓk\ell_k is its loss. The probabilities qkq_k can all be high at once and do not need to sum to 1.

A video can have sound too

So far we have used only the frames. A video of a guitar also has sound, which can help us tell whether someone is playing it.

Encode the frames and audio separately, then combine their features or predictions. For the audio, we can use a spectrogram, a grid of frequency content over time. Its steps may differ from the video frame rate; we align both inputs to the interval we want to classify.

We can combine them inside the network too. Attention Bottlenecks give the audio and visual streams a small set of shared latent tokens. Each stream updates a copy of those tokens, then we average the updates. In the next layer, both streams can read what the other contributed. Communication passes through these few tokens rather than every possible audio-visual pair.

Instead of a class, we could predict one speaker's voice from a mixed recording. VisualVoice uses facial appearance and lip motion along with the audio. The face helps identify whose voice to extract, and the lips help align speech with time. This is audio-visual source separation.

Sound can also help us choose frames. Listen to Look uses inexpensive image-audio observations to select moments worth processing with a more expensive visual model. We no longer have to spend the same amount of work on every part of the video.

Finally, connect the visual features to a language model. Video-LLaVA uses this to answer questions and produce descriptions, instead of choosing from a fixed set of action classes.

The language model still only sees what the video encoder retained. If we skipped the moment the cup moved, or averaged away its direction, a more flexible decoder does not restore that evidence.

Next are self-supervised learning notes, where we train visual features by hiding, transforming, and comparing the data itself.

← Back to blog