On object detection, segmentation, and model interpretation
This is part #10 of my notes on CS231n. The course is openly available, including the video lectures and assignments.
These notes are based on lecture 9: Detection, Segmentation, Visualization, and Understanding by Ehsan Adeli. I also cover feature inversion, adversarial examples, and style transfer from the course schedule, with their original papers linked below.
So far, we have used CNNs to predict one class for an image. Let's put two people and a bicycle in that image. Predicting “person” tells us something, but we still don't know where either person is, or that there is a bicycle.
We can ask for more structure in the output:
| Task | What the output tells us |
|---|---|
| Classification | A class for the image |
| Classification + localization | A class and a box for one object |
| Object detection | A class and a box for each object |
| Semantic segmentation | A class for each pixel |
| Instance segmentation | A separate mask for each object |
| Panoptic segmentation | Pixel classes across the scene, with separate object identities |
We can keep the feature extractor. What we need to change is the output and the loss that trains it.
Semantic segmentation: classify every pixel
Let's start with an image of shape . For possible classes, we want a score tensor of shape .
One idea is to classify a crop around every pixel. A crop gives us context: a grey pixel alone could be a road, a wall, or a bicycle. But adjacent crops contain almost the same pixels, so we repeat a lot of work.
Instead, run the convolutions once over the image and keep the spatial feature map. Apply a convolution to turn each feature vector into class scores. This gives us a fully convolutional network, or FCN. Long et al., 2015
There is a problem: the pooling and strides that helped classification have reduced spatial resolution. We need to bring it back up.
We can split the network into an encoder that reduces resolution and a decoder that increases it again. In the decoder, interpolate then convolve, or use a learned transposed convolution. The latter applies the transpose of a convolution's linear operator; the lost detail does not come back automatically.
To preserve that detail, pass earlier feature maps into the decoder through skip connections. The decoder can now combine fine edges from early layers with object features from deeper layers.
After upsampling, apply softmax across classes at each pixel. We now have a classifier at every pixel, so we can average the same cross-entropy over the labelled pixels:
Here is the true class of pixel , and is the probability assigned to that class. Pixels marked “ignore” do not enter the sum.
If two labelled pixels receive correct-class probabilities and , their average loss is:
The second pixel contributes much more. Notice that none of this distinguishes one person from another: both receive the class “person”.
Object detection: predict a set of boxes
For one object, we can attach two heads to the same features: one predicts the class, the other predicts four box coordinates. A head is just a branch of the network with its own output and loss.
Now add a second object. We need another box, and perhaps a third for the bicycle. The number of outputs depends on the image. Their order should not matter either: listing the left person first or the right person first describes the same scene.
Before looking at architectures, we need a way to compare boxes.
Intersection over Union
For boxes and , Intersection over Union is:
The vertical bars mean area. Identical boxes have IoU 1; boxes with no shared area have IoU 0.
Let's use corner coordinates . Let and . Each area is 16, and their overlap is :
Half of each box overlaps, but IoU is only one third. The denominator counts the entire union, not just one box.
We can apply the same operation to sets of mask pixels. For semantic segmentation, average the IoU of each class to get mean IoU. If most pixels are background, predicting background everywhere can give high pixel accuracy. The object classes still have zero IoU.
When is a detection correct?
A prediction needs the correct class and enough overlap with an annotated object. We choose an IoU threshold, then match detections to ground truth in descending confidence order. A normal ground-truth object can be matched once; another detection of that same object is a false positive.
Suppose a dataset contains three annotated bicycles. At IoU threshold 0.5, four predictions arrive in this order:
| Score | Match | Precision | Recall |
|---|---|---|---|
| 0.95 | Bicycle A | ||
| 0.90 | A again: duplicate | ||
| 0.80 | Bicycle B | ||
| 0.70 | Bicycle C |
Precision is the fraction of retained predictions that are correct. Recall is the fraction of ground-truth objects found. Notice what the duplicate does: it lowers precision without finding another bicycle.
Lowering the confidence threshold includes more predictions. Plot precision against recall as we do this, then summarize that curve with Average Precision, or AP. AP uses interpolated precision at recall levels; we cannot just average the table's precision column. Average AP across classes to get mAP.
The COCO evaluator uses 101 recall levels and averages AP over ten IoU thresholds, from 0.50 to 0.95 in steps of 0.05. AP50 and AP75 fix the IoU threshold at the named value. The full evaluator also handles crowds, ignored regions, and detection limits.
From crops to shared features
Running a classifier on every possible crop is expensive. There are many positions, sizes, and aspect ratios to try.
R-CNN reduces the search to region proposals: boxes that might contain objects. Resize each proposed crop, run it through a CNN, then classify its features and adjust the box.
But overlapping crops still repeat the CNN's work. Fast R-CNN moves that computation before the crop: run the image through the CNN once, then extract each region from the shared feature map. We train its class and box heads together. Girshick, 2015
The proposed regions still have different sizes. Divide each one into a fixed grid and max-pool within each bin. Now every region has the same feature size and can feed the same head. This operation is RoI pooling, where RoI means region of interest.
We can also learn where to place the proposals. Faster R-CNN adds a Region Proposal Network, or RPN, to the shared features. At each feature-map location, place reference boxes of several sizes and aspect ratios. These are the anchors. The RPN predicts an objectness score and four box corrections for each anchor. Ren et al., 2015
Notice the two paths from the shared feature map. One proposes boxes. The other supplies the features inside those boxes. Objectness first asks whether there is an object; the class head then asks which class it belongs to. The box head makes a second correction to its location.
These are the two stages: propose regions, then classify and refine them. We only ran the image backbone once.
What does a box head learn?
For a reference box with centre and dimensions , we can express the target box as relative changes:
We divide position changes by the reference size and take the log of size ratios. The same target can now mean “move right by half a box width” for a small or a large object. This is R-CNN's box regression, in Appendix C.
If the reference is centred at with size , and the target is centred at with size , the targets are:
To decode, and , with the same operations for and .
To train the region heads, add the class loss and the box loss:
Here is the true-class probability and balances the terms. The indicator turns the box loss off for background: there is no target object box to regress. Fast R-CNN sums smooth L1 over the four box-offset errors. For each error :
It is quadratic near zero and linear for larger errors.
One-stage detectors
What if we predict the final classes and boxes at each image location, and remove the second stage?
The original YOLO divides the image into a grid. Assign each object to the cell containing its centre, then train that cell to predict boxes, confidence values, and class probabilities. The grid gives us a fixed output tensor for images with different numbers of objects. In this version of YOLO, confidence combines object presence with box overlap. Redmon et al., 2015
Dense prediction creates a training problem: most candidates are easy background. Their combined loss can overwhelm the smaller number of useful examples.
We can reduce the easy examples' contribution by multiplying cross-entropy by a factor that shrinks as the prediction improves. This is focal loss, introduced with RetinaNet. Omitting its optional class-balancing weight:
Here is the probability of the correct binary label and controls the effect. At , this is cross-entropy. At , a correct-label probability of 0.8 multiplies cross-entropy by 0.04; a probability of 0.2 multiplies it by 0.64. The hard example keeps much more of its contribution. Lin et al., 2017
Removing duplicates
Many nearby candidates can predict the same object. To remove duplicates, keep the highest-scoring box, remove lower-scoring boxes that overlap it above an IoU threshold, then repeat with the remaining boxes. This is non-maximum suppression, or NMS, commonly applied separately for each class. The Torchvision operation follows this rule.
For three person detections, suppose the scores are 0.9, 0.8, and 0.7. The first two overlap at IoU 0.85, while the third is elsewhere. With NMS threshold 0.5, keep the first and third.
Two real people can overlap too, so NMS can remove a correct detection. Could we train the network to produce one prediction per object in the first place?
DETR: prediction as a set
DETR passes image features and positional information through a Transformer encoder. Its decoder starts with a fixed set of learned object queries. Each query produces a class and a box; unused queries predict “no object”. Carion et al., 2020
The queries attend to image features through cross-attention and interact through self-attention. We can think of each query as a slot available to predict an object. Training decides how those slots are used.
For each training image, compute a cost for every predicted/true pair. Then find the lowest-cost one-to-one assignment with Hungarian matching.
For two objects and two predictions, imagine these combined class-and-box costs:
| Prediction 1 | Prediction 2 | |
|---|---|---|
| Person | 1 | 4 |
| Bicycle | 3 | 2 |
Matching person to prediction 1 and bicycle to prediction 2 costs . Swapping them costs . The first assignment wins. If there were extra slots, they would learn “no object”.
Once we have the assignment, train every slot with classification loss and the matched objects with box loss. The original box loss combines L1 coordinate error with generalized IoU. Notice the order: matching chooses which target belongs to each prediction, then backpropagation improves those predictions. Training this way lets the original DETR produce its set without NMS. Matching and training losses
Why generalized IoU? Ordinary IoU stays zero when two boxes are disjoint. GIoU also measures the unused area in the smallest enclosing box, so separated boxes can still receive a localization signal.
Instance and panoptic segmentation
Return to our two people. Semantic segmentation gives both the same pixel label. Detection separates them, but a rectangle includes background. We want a separate pixel mask for each person.
Add a third branch to Faster R-CNN: for each region, predict a small binary mask for each class. We now have Mask R-CNN. For foreground regions, train only the mask for the true class with binary cross-entropy. The class head handles which class it is; the mask branch learns which pixels belong to it. He et al., 2017
For a mask, rounding a region's coordinates can shift the boundary by several image pixels. Replace RoI pooling with RoIAlign: keep the fractional coordinates and sample features with bilinear interpolation. Resize the predicted mask into the output box.
That separates countable things such as people and cars. To label the rest of the scene as well, include stuff classes such as road and sky. Each pixel gets one class and, for a thing, an instance identity. This is panoptic segmentation, and its final regions do not overlap. Kirillov et al., 2018
In our scene, this means person 1, person 2, bicycle 1, road, and sky. Mask AP evaluates scored instance masks. Panoptic Quality combines their overlap quality with missed and extra segments. For one class:
contains matched predicted/true segment pairs with IoU above 0.5. counts unmatched predictions; counts missed true segments. The bars count elements. Dataset rules handle void regions separately.
What has the model learned?
We have been changing the output to ask more of the network. Let's now keep a trained network and look at what its features respond to.
We can view first-layer filters directly, or collect image patches that strongly activate a channel. We can also make an input that activates it: start with an image, compute the activation's gradient with respect to the pixels, then change the pixels to increase it. Keep the network weights fixed.
This is activation maximization:
Here is the activation we want to increase. We subtract to penalize extreme pixels or rapid changes between neighbours. Change this regularizer and we change which activating image the optimization finds. Olah, Mordvintsev, and Schubert
Saliency: sensitivity to the input
For an existing image , we can ask a more local question: which pixels can change the class score most if we move them a little? Compute the gradient with respect to each pixel and take its largest magnitude across colour channels:
Here are pixel coordinates and indexes colour channels. Displaying gives us a basic saliency map.
For a small input change , the score changes approximately by:
This measures local sensitivity. A region can matter even if its gradient is small, for example when the model's response is flat around the current input. Simonyan et al., 2014
CAM and Grad-CAM
Suppose a classifier averages each final feature map over space, then uses a linear classifier with weights for class .
Take those same classifier weights and use them to sum the feature maps before spatial averaging:
Averaging over space recovers the class score, apart from its bias. But now we can see where each contribution came from. This is Class Activation Mapping, or CAM. Zhou et al., 2016
For a network without that pooling-and-linear-classifier structure, we need another way to weight its feature maps. Grad-CAM computes the score gradient at a chosen convolutional layer and averages it over space:
Here are the feature-map dimensions. Each becomes the weight for channel . Sum the weighted maps, keep positive contributions with ReLU, then upsample for display. Selvaraju et al., 2017
A heatmap can look right without explaining the learned model. One useful check is to randomize the weights or training labels and compare the maps. Some methods still produce similar-looking results, as Sanity Checks for Saliency Maps shows.
Feature inversion
Now take the features of a particular image and try to reconstruct that image from them. This is feature inversion.
Let be the features at layer , and let be our reference image. Start from noise and optimize:
At every step, run the reconstruction through the same fixed network and compare its features with the reference. Backpropagate that difference to the reconstructed pixels. If several images fit, the representation does not tell us which one was the original; the regularizer helps choose one. Mahendran and Vedaldi, 2015
We can repeat this at different layers to see how much image detail each representation preserves.
Adversarial examples
We can also optimize the input to make the model wrong, while restricting how much we change it.
For fixed parameters , true label , and classification loss , the Fast Gradient Sign Method, or FGSM, makes one step that increases the loss:
The clip operation keeps valid pixel values. limits the change per input coordinate, in the same units as the input. Each pixel moves in the direction that increases the linearized loss. Goodfellow et al., 2015
For a two-value input , gradient , and , we get . The first-order loss change is:
Notice how both coordinates increase the loss even though one moves up and the other down. These contributions can add across many dimensions. We chose the perturbation using the model's gradient; small random noise would not necessarily do the same. Whether it changes the predicted class, or is visible to a person, depends on the image, model, and .
DeepDream and style transfer
Let's change what we optimize the image for.
In DeepDream, start with an image and increase selected internal activations. A weak pattern becomes stronger, activates the feature more, and gets amplified again. Choosing a different layer changes which structures emerge. The network stays fixed throughout. Mordvintsev, Olah, and Tyka, 2015
For neural style transfer, give the optimization two images: one whose content we want to keep and another whose style we want to use. Compare content through feature activations. For style, compare how feature channels activate together. Gatys et al., 2015
Flatten a layer's spatial positions into columns, giving a feature matrix . Its Gram matrix is:
An entry measures how two channels activate together across positions. Summing over positions removes their exact arrangement from this comparison.
For two channels and two positions:
Swap the two columns of : the Gram matrix stays the same. We kept the feature correlations but changed where they occurred. This is why a Gram comparison can describe texture without preserving the whole scene layout.
For generated image , content image , and style image , the objective has the form:
selects a content layer. weights style layers, with normalization factors absorbed into it. balance content and style; sums squared matrix entries. Gradients change while the feature network stays fixed.
We are still using backpropagation, just with a different variable to update. The same fixed network can reconstruct an image, make an adversarial input, or transfer a texture depending on the loss we differentiate through.
Next, video understanding adds time: recognizing what changes across a sequence, rather than only what appears in one image.