On CNNs

This is part #7 of my notes on CS231n. The course is openly available, including the video lectures and assignments.

These notes follow Image Classification with CNNs by Justin Johnson, the convolutional network notes, and the CNN architectures lecture.


In backpropagation, we built a network from fully connected layers. For an image, that meant flattening all the pixels into one long vector.

That works, but we know something about images: nearby pixels are related, and a useful pattern can appear in different places. We can build those assumptions into the network.

Convolution

A filter, also called a kernel, is a small array of learnable weights. We move it across the image. At each position, multiply the weights by the corresponding input values, sum the products, and add a bias.

It's the same dot product we already know, applied to a small patch:

y=wTxpatch+by = w^T x_{\text{patch}} + b

Here ww and xpatchx_{\text{patch}} are the filter and patch written as column vectors. Each position produces one number. Collect those numbers into a grid and we have a feature map.

The useful bit is that we reuse the same weights and bias at every position. We don't need to learn a separate edge detector for each corner of the image.

A small example

Let's use a single-channel image, a 2×22\times2 filter, and zero bias:

X=[120031210]K=[100−1]X = \begin{bmatrix} 1 & 2 & 0 \\ 0 & 3 & 1 \\ 2 & 1 & 0 \end{bmatrix} \qquad K = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}

At the top-left position:

y0,0=1(1)+2(0)+0(0)+3(−1)=−2y_{0,0} = 1(1) + 2(0) + 0(0) + 3(-1) = -2

Move one pixel right, then repeat on the next row:

Y=[−21−13]Y = \begin{bmatrix} -2 & 1 \\ -1 & 3 \end{bmatrix}

That's the entire operation. I picked these weights to make the arithmetic easy; in a CNN, training learns them.

Technically this is cross-correlation, since we don't flip the filter. Deep learning libraries still call the layer a convolution. PyTorch uses this convention too.

Channels and parameters

For these notes, image shapes are H×W×CH\times W\times C: height, width, channels. An RGB image has three channels.

In an ordinary convolution, a 3×33\times3 filter covers all input channels. On an RGB image, it has 3×3×3=273\times3\times3=27 weights. Products from all three channels are added together to produce one number, then we add one bias.

Use 16 different filters and we get 16 output channels. Each filter has its own weights, shared across positions.

parameters=(F2Cin+1)Cout\text{parameters} = (F^2 C_{\text{in}} + 1)C_{\text{out}}

where FF is the filter width and height. For our 16 filters:

(32⋅3+1)⋅16=448(3^2\cdot3+1)\cdot16 = 448

Notice that the image width and height don't appear in this count. A larger image needs more computation and produces more activations, but the filter weights stay the same.

Stride and padding

Two settings control where we apply the filter:

  • Stride SS: how many pixels we move at each step.
  • Padding PP: how many rows or columns of zeros we add on each side.

With a square filter of size FF, no dilation, and equal padding on both sides:

Hout=⌊H+2P−FS⌋+1H_{\text{out}} = \left\lfloor\frac{H+2P-F}{S}\right\rfloor+1

The same formula applies to width. ⌊⋅⌋\lfloor\cdot\rfloor means round down: count only positions where the whole filter fits.

For a 32×3232\times32 input and a 3×33\times3 filter:

StridePaddingOutput
1030×3030\times30
1132×3232\times32
2116×1616\times16

Padding lets us keep the spatial size. A larger stride reduces it. Neither setting changes the number of filters.

Seeing more of the image

One 3×33\times3 convolution only sees a small patch. Stack another one and it combines information from neighbouring patches.

With stride 1 and no dilation, two 3×33\times3 layers cover a 5×55\times5 region of the original image; three cover 7×77\times7. This region is the receptive field. Near the boundary, part of it may be padding. The lecture diagrams show how it grows.

We still need non-linearities between layers. As before, ReLU applies max⁡(0,x)\max(0,x) to each activation.

Pooling

Pooling reduces spatial resolution within each channel. Max pooling keeps the largest value in each window:

[1423]⟶4\begin{bmatrix}1 & 4\\2 & 3\end{bmatrix} \quad\longrightarrow\quad 4

A 2×22\times2 window with stride 2 turns 32×32×1632\times32\times16 into 16×16×1616\times16\times16. The channel count stays the same, and there are no weights to learn.

We save computation in later layers, but discard some position information. A convolution with stride 2 can also reduce resolution while learning its weights.

Putting it together

For a small ten-class classifier, we could use the layers below. Both convolutions use 3×33\times3 filters, stride 1, and padding 1. Max pooling uses a 2×22\times2 window and stride 2.

LayerShape
RGB image32×32×332\times32\times3
Convolution, 16 filters + ReLU32×32×1632\times32\times16
Max pooling16×16×1616\times16\times16
Convolution, 32 filters + ReLU16×16×3216\times16\times32
Average each channel over its spatial positions3232
Fully connected layer1010

Those scores go into the same softmax and cross-entropy loss as before. Backpropagation computes gradients for the filters, biases, and final classifier. Because a filter is reused, its gradient adds the contributions from every position where it was applied.

We choose the architecture. The loss tells it which patterns are useful.

Deeper networks

The next lecture builds on these pieces. Three ideas to keep:

  • Normalization layers normalize groups of activations, then apply a learned scale and offset. The lecture uses layer normalization, which computes its statistics within each example.
  • Residual connections, used in ResNet, add a block's input to its learned output: z=F(x)+xz=F(x)+x. The two shapes must match, or the shortcut needs a projection. This gives information and gradients a direct path around the block, making deeper networks easier to train.
  • Transfer learning starts with a pretrained network. Replace its classifier for your classes, then train that layer or fine-tune more of the network. We can reuse features learned from another dataset.

Next are RNNs, where we reuse weights across a sequence instead of across image positions.

← Back to blog