On CNNs
This is part #7 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes follow Image Classification with CNNs by Justin Johnson, the convolutional network notes, and the CNN architectures lecture.
In backpropagation, we built a network from fully connected layers. For an image, that meant flattening all the pixels into one long vector.
That works, but we know something about images: nearby pixels are related, and a useful pattern can appear in different places. We can build those assumptions into the network.
Convolution
A filter, also called a kernel, is a small array of learnable weights. We move it across the image. At each position, multiply the weights by the corresponding input values, sum the products, and add a bias.
It's the same dot product we already know, applied to a small patch:
Here and are the filter and patch written as column vectors. Each position produces one number. Collect those numbers into a grid and we have a feature map.
The useful bit is that we reuse the same weights and bias at every position. We don't need to learn a separate edge detector for each corner of the image.
A small example
Let's use a single-channel image, a filter, and zero bias:
At the top-left position:
Move one pixel right, then repeat on the next row:
That's the entire operation. I picked these weights to make the arithmetic easy; in a CNN, training learns them.
Technically this is cross-correlation, since we don't flip the filter. Deep learning libraries still call the layer a convolution. PyTorch uses this convention too.
Channels and parameters
For these notes, image shapes are : height, width, channels. An RGB image has three channels.
In an ordinary convolution, a filter covers all input channels. On an RGB image, it has weights. Products from all three channels are added together to produce one number, then we add one bias.
Use 16 different filters and we get 16 output channels. Each filter has its own weights, shared across positions.
where is the filter width and height. For our 16 filters:
Notice that the image width and height don't appear in this count. A larger image needs more computation and produces more activations, but the filter weights stay the same.
Stride and padding
Two settings control where we apply the filter:
- Stride : how many pixels we move at each step.
- Padding : how many rows or columns of zeros we add on each side.
With a square filter of size , no dilation, and equal padding on both sides:
The same formula applies to width. means round down: count only positions where the whole filter fits.
For a input and a filter:
| Stride | Padding | Output |
|---|---|---|
| 1 | 0 | |
| 1 | 1 | |
| 2 | 1 |
Padding lets us keep the spatial size. A larger stride reduces it. Neither setting changes the number of filters.
Seeing more of the image
One convolution only sees a small patch. Stack another one and it combines information from neighbouring patches.
With stride 1 and no dilation, two layers cover a region of the original image; three cover . This region is the receptive field. Near the boundary, part of it may be padding. The lecture diagrams show how it grows.
We still need non-linearities between layers. As before, ReLU applies to each activation.
Pooling
Pooling reduces spatial resolution within each channel. Max pooling keeps the largest value in each window:
A window with stride 2 turns into . The channel count stays the same, and there are no weights to learn.
We save computation in later layers, but discard some position information. A convolution with stride 2 can also reduce resolution while learning its weights.
Putting it together
For a small ten-class classifier, we could use the layers below. Both convolutions use filters, stride 1, and padding 1. Max pooling uses a window and stride 2.
| Layer | Shape |
|---|---|
| RGB image | |
| Convolution, 16 filters + ReLU | |
| Max pooling | |
| Convolution, 32 filters + ReLU | |
| Average each channel over its spatial positions | |
| Fully connected layer |
Those scores go into the same softmax and cross-entropy loss as before. Backpropagation computes gradients for the filters, biases, and final classifier. Because a filter is reused, its gradient adds the contributions from every position where it was applied.
We choose the architecture. The loss tells it which patterns are useful.
Deeper networks
The next lecture builds on these pieces. Three ideas to keep:
- Normalization layers normalize groups of activations, then apply a learned scale and offset. The lecture uses layer normalization, which computes its statistics within each example.
- Residual connections, used in ResNet, add a block's input to its learned output: . The two shapes must match, or the shortcut needs a projection. This gives information and gradients a direct path around the block, making deeper networks easier to train.
- Transfer learning starts with a pretrained network. Replace its classifier for your classes, then train that layer or fine-tune more of the network. We can reuse features learned from another dataset.
Next are RNNs, where we reuse weights across a sequence instead of across image positions.