15

Convolutional Neural Networks

Cantonese podcast title: 卷積神經網絡

Learning Objectives

  1. Compute the output spatial size of a convolution given input size, kernel size, stride, and padding, and explain why convolution is dramatically more parameter-efficient than a fully-connected layer for image data.
  2. Explain what a convolution filter does, how local connectivity and weight sharing differ from the dense layer pattern, and what the resulting activation map means.
  3. Distinguish "same" padding from "valid" padding and stride 1 from stride 2, and pick the right combination for a downsampling layer versus a same-resolution feature extractor.
  4. Describe max pooling and average pooling, what they contribute to translation invariance and to spatial downsampling, and when to prefer one over the other.
  5. Compute the receptive field of a unit at a given depth and explain why the receptive field must grow with depth for a CNN to recognise whole objects.
  6. Sketch the feature hierarchy (edges, textures, parts, objects) and explain why the early layers of trained CNNs look like Gabor filters and colour blobs regardless of the task they are trained on.
  7. Compare LeNet, AlexNet, VGG, and ResNet on depth, parameter count, and core innovation, and articulate the degradation problem that motivated the residual connection.
  8. Distinguish feature extraction from fine-tuning in transfer learning, and pick between them based on the size and domain similarity of the new dataset relative to the pretraining corpus.
Convolutional Neural Networks — visual guide
CNN convolution and pooling Convolution, pooling, and the receptive field input 32×32×3 3×3 kernel stride 1 feature maps 30×30×32 one filter -> one channel pool 2×2 pooled 15×15×32 what each layer learns edges layer 1 textures layer 2 parts layer 3 objects layer 4 Padding preserves spatial size; stride trades it away. Together they set how much context one layer can see. Pooling halves width and height but leaves the channel count alone, which is why the maps get deeper rather than smaller as you go down. Transfer learning reuses the early layers of a network trained on millions of labelled images and retrains only the head — usually all the data you need. The receptive field grows roughly linearly with depth, which is why a CNN needs many layers to reach the whole image.

A 224 by 224 colour image is a tensor of 150,528 numbers; a fully-connected layer that maps it to a thousand outputs owns 150 million weights before it sees a single training example. A convolutional layer doing the same job owns dozens. The trick — local connectivity plus weight sharing — is the defining idea of convolutional neural networks and the reason computer vision made its modern leap. This lesson walks through that trick: what a filter is and how it slides, how stride and padding change the output shape, what pooling buys you, why the receptive field grows with depth, the feature hierarchy that makes the layers stack into a recogniser, the canonical architectures from LeNet to ResNet, and the transfer-learning workflow that lets you reuse ImageNet backbones on a few hundred custom-class images.

Learning Objectives

  1. Compute the output spatial size of a convolution given input size, kernel size, stride, and padding, and explain why convolution is dramatically more parameter-efficient than a fully-connected layer for image data.
  2. Explain what a convolution filter does, how local connectivity and weight sharing differ from the dense layer pattern, and what the resulting activation map means.
  3. Distinguish "same" padding from "valid" padding and stride 1 from stride 2, and pick the right combination for a downsampling layer versus a same-resolution feature extractor.
  4. Describe max pooling and average pooling, what they contribute to translation invariance and to spatial downsampling, and when to prefer one over the other.
  5. Compute the receptive field of a unit at a given depth and explain why the receptive field must grow with depth for a CNN to recognise whole objects.
  6. Sketch the feature hierarchy (edges, textures, parts, objects) and explain why the early layers of trained CNNs look like Gabor filters and colour blobs regardless of the task they are trained on.
  7. Compare LeNet, AlexNet, VGG, and ResNet on depth, parameter count, and core innovation, and articulate the degradation problem that motivated the residual connection.
  8. Distinguish feature extraction from fine-tuning in transfer learning, and pick between them based on the size and domain similarity of the new dataset relative to the pretraining corpus.

1. From Dense Layers to Convolutions

Consider a 224 by 224 RGB image, three channels stacked into a tensor of shape (3,224,224)(3, 224, 224). A fully-connected ("dense") layer with 1000 output units owns a weight matrix of shape

(3⋅224⋅224)×1000=150,528×1000=150,528,000(3 \cdot 224 \cdot 224) \times 1000 = 150{,}528 \times 1000 = 150{,}528{,}000

parameters. That is 150 million weights before any optimisation begins, and it treats every pixel as an unrelated scalar. The network has no notion that the pixel at (x,y)(x, y) is adjacent to the pixel at (x+1,y)(x+1, y); the ordering of the input vector is arbitrary but the dense layer treats it as meaningful. It also learns a separate weight for the same input location in two different images, which is wasteful — edges, corners, and colour gradients look the same regardless of where they appear.

1.1 Two assumptions images actually satisfy

Convolutional layers exploit two facts that dense layers ignore:

  1. Locality. A pixel's meaning is determined mostly by its neighbours. A cat's ear is a cat's ear because of the local pattern, not because of the average colour of the whole image. Useful features can be extracted from small patches.
  2. Stationarity. The same local pattern (an edge, a corner, a blob) is informative anywhere in the image. There is no reason to learn a separate "horizontal edge detector" for the top-left corner and another for the bottom-right.

1.2 The parameter savings

A convolutional layer with CoutC_{\text{out}} filters of spatial size k×kk \times k over a CinC_{\text{in}}-channel input owns

Cout×(Cin⋅k⋅k+1)C_{\text{out}} \times (C_{\text{in}} \cdot k \cdot k + 1)

parameters (the trailing +1+1 is the bias per filter). With Cin=3C_{\text{in}} = 3, k=3k = 3, Cout=64C_{\text{out}} = 64, that is 64×(27+1)=1,79264 \times (27 + 1) = 1{,}792 parameters — five orders of magnitude fewer than the dense layer above, and the convolutional layer also produces a 64-channel feature map at every spatial position. The saving is not a constant factor; it is the difference between a model that fits in GPU memory and one that does not.

2. The Convolution Operation

A convolution filter (also called a kernel) is a small tensor of learnable weights, typically 3×33 \times 3 or 5×55 \times 5 in spatial extent, with the same depth as the input volume. The filter slides over the input one position at a time (stride 1; stride is treated properly in Section 3). At every spatial position it computes the dot product of its weights with the input patch underneath, adds a bias, and passes the result through a non-linearity. The collection of those scalar outputs across all spatial positions is the activation map (or feature map) for that filter.

2.1 A concrete 3×33 \times 3 example

Suppose we have a single-channel 5×55 \times 5 input and a 3×33 \times 3 filter with weights

K=(10−110−110−1).K = \begin{pmatrix} 1 & 0 & -1 \\ 1 & 0 & -1 \\ 1 & 0 & -1 \end{pmatrix}.

This is a vertical-edge detector: the dot product is large when the left column is bright and the right column is dark, and zero on a uniform patch. Sliding it over a 5×55 \times 5 input and computing the dot product at every legal position yields a 3×33 \times 3 activation map. At the top-left position, the dot product is

(1⋅a1,1+0⋅a1,2−1⋅a1,3)+(1⋅a2,1+0⋅a2,2−1⋅a2,3)+(1⋅a3,1+0⋅a3,2−1⋅a3,3).(1\cdot a_{1,1} + 0\cdot a_{1,2} - 1\cdot a_{1,3}) + (1\cdot a_{2,1} + 0\cdot a_{2,2} - 1\cdot a_{2,3}) + (1\cdot a_{3,1} + 0\cdot a_{3,2} - 1\cdot a_{3,3}).

The same operation repeats at nine positions and produces nine scalars. That is the entire convolution in the discrete case; backpropagation through it is the transpose of the same operation.

2.2 Local connectivity and weight sharing

Two phrases hide the two design choices:

  • Local connectivity. Each output unit looks at only a k×kk \times k patch of the input, not the whole image. This is what cuts the parameter count by orders of magnitude and what makes the operation cheap enough to stack.
  • Weight sharing. The same filter is used at every spatial position, so the layer learns one pattern, not H⋅WH \cdot W copies of it. This is what gives a CNN translation equivariance: if the input shifts by one pixel, the activation map shifts by one pixel.

In a deep stack, equivariance at one layer composes with equivariance at the next, so the whole network is equivariant to translations of the input. True invariance — the prediction is unchanged when the input shifts — comes from a global pooling layer at the head (Section 4) or from data augmentation.

2.3 Output volume size

For an input of spatial size H×WH \times W, a filter of size k×kk \times k applied with stride ss and zero-padding pp on each side produces an activation map of spatial size

H′=⌊H+2p−ks⌋+1,W′=⌊W+2p−ks⌋+1.H' = \left\lfloor \frac{H + 2p - k}{s} \right\rfloor + 1, \qquad W' = \left\lfloor \frac{W + 2p - k}{s} \right\rfloor + 1.

The volume stacks CoutC_{\text{out}} such maps, so the full output volume is shape (Cout,H′,W′)(C_{\text{out}}, H', W'). Worked example: a 32×3232 \times 32 input padded by p=1p = 1, convolved with a 3×33 \times 3 filter at stride 1, gives

H′=⌊32+2−31⌋+1=32,H' = \left\lfloor \frac{32 + 2 - 3}{1} \right\rfloor + 1 = 32,

so the spatial size is preserved and the output volume is (Cout,32,32)(C_{\text{out}}, 32, 32). With stride 2 instead, H′=⌊32/2⌋+1=16H' = \lfloor 32/2 \rfloor + 1 = 16 — exactly the downsampling factor we want in the encoder half of a typical classifier.

import numpy as np

#A from-scratch 2D convolution on a single-channel image.
#Real frameworks use im2col + GEMM, but the operation is identical.
def conv2d(x, kernel, stride=1, padding=0):
    """x: (H, W). kernel: (kH, kW). Returns (H', W')."""
    if padding > 0:
        x = np.pad(x, padding, mode="constant", constant_values=0)
    kH, kW = kernel.shape
    H, W = x.shape
    oH = (H - kH) // stride + 1
    oW = (W - kW) // stride + 1
    out = np.zeros((oH, oW), dtype=np.float32)
    for i in range(oH):
        for j in range(oW):
            patch = x[i*stride:i*stride + kH,
                       j*stride:j*stride + kW]
            out[i, j] = np.sum(patch * kernel)
    return out

#Vertical-edge kernel.
k = np.array([[1, 0, -1],
              [1, 0, -1],
              [1, 0, -1]], dtype=np.float32)
img = np.random.default_rng(0).normal(size=(8, 8)).astype(np.float32)
print(conv2d(img, k, stride=1, padding=0).shape)  # (6, 6)
import torch
import torch.nn as nn

#The real thing: a 2D convolution that preserves spatial size ("same").
#padding = k // 2 keeps H' = H for stride 1 and odd kernel size.
conv_same = nn.Conv2d(
    in_channels=3, out_channels=64,
    kernel_size=3, stride=1, padding=1, bias=True,
)
x = torch.randn(8, 3, 32, 32)   # batch of 8 RGB images
print(conv_same(x).shape)       # torch.Size([8, 64, 32, 32])

#Strided convolution: downsamples by 2 in each spatial dim.
conv_down = nn.Conv2d(3, 64, kernel_size=3, stride=2, padding=1)
print(conv_down(x).shape)       # torch.Size([8, 64, 16, 16])

3. Stride and Padding

Stride and padding are the two knobs that control the spatial geometry of a convolutional layer.

3.1 Stride

Stride is the number of input pixels the filter shifts between consecutive positions. Stride 1 (the default) is the dense sliding we described above. Stride 2 skips every other position and halves the output spatial size (modulo rounding). The full formula from Section 2.3 is

H′=⌊H+2p−ks⌋+1.H' = \left\lfloor \frac{H + 2p - k}{s} \right\rfloor + 1.

Stride 2 is the canonical way to halve spatial resolution in modern architectures. It replaces the older "conv then pool" pattern and lets the network learn its own downsampling filter.

3.2 Padding: "valid" vs "same"

There are two conventions.

  • Valid padding means no padding. The filter is only ever placed where it fits entirely inside the input. The output is strictly smaller than the input; for stride 1 and a 3×33 \times 3 filter, every convolution shrinks the spatial map by one pixel per side.
  • Same padding pads the input with zeros around the border so that stride 1 with a k×kk \times k filter produces an output of the same spatial size as the input. The required padding is p=(k−1)/2p = (k - 1)/2 pixels on each side, which is an integer when kk is odd. This is why every convolutional layer in practice uses odd kernel sizes — 3×33 \times 3, 5×55 \times 5 — even sizes require asymmetric padding or a one-pixel mismatch.

The arithmetic for a 32×3232 \times 32 input with a 5×55 \times 5 filter at stride 1:

  • Valid: H′=⌊32−5⌋+1=28H' = \lfloor 32 - 5 \rfloor + 1 = 28. Spatial size 32 → 28.
  • Same: pad by 2 on each side; H′=⌊32+4−5⌋+1=32H' = \lfloor 32 + 4 - 5 \rfloor + 1 = 32. Spatial size 32 → 32.
import torch.nn as nn

#"Valid" convolution: no padding, output shrinks.
conv_valid = nn.Conv2d(in_channels=64, out_channels=64,
                       kernel_size=5, stride=1, padding=0)

#"Same" convolution: pad so output equals input at stride 1.
conv_same = nn.Conv2d(in_channels=64, out_channels=64,
                      kernel_size=5, stride=1, padding=2)

3.3 Worked example: the VGG block

A VGG block applies two or three 3×33 \times 3 convolutions with same padding (so spatial size is preserved), then a stride-2 pooling or strided convolution that halves the resolution. For a 224×224224 \times 224 input the cascade is

224→conv 3×3,  s=1,  p=1224→conv 3×3,  s=1,  p=1224→pool 2×2,  s=2112.224 \xrightarrow{\text{conv } 3\times 3,\; s=1,\; p=1} 224 \xrightarrow{\text{conv } 3\times 3,\; s=1,\; p=1} 224 \xrightarrow{\text{pool } 2\times 2,\; s=2} 112.

Two stacked 3×33 \times 3 "same" convolutions preserve the spatial size and give the layer an effective receptive field of 5×55 \times 5 (Section 5), without paying the parameter cost of a single 5×55 \times 5 filter (2×32=182 \times 3^2 = 18 weights per channel versus 52=255^2 = 25).

4. Pooling

A pooling layer downsamples the activation map by summarising each non-overlapping (or slightly overlapping) window with a single number.

4.1 Max pooling

The most common variant. With a 2×22 \times 2 window and stride 2, each 2×22 \times 2 patch collapses to its maximum value. The intuition: "was this feature detected anywhere in the patch?" If a vertical-edge filter fires anywhere in the patch, the patch's max is high; if it fires nowhere, the max is low. The exact position of the firing within the patch is discarded. This is exactly the invariance property we want — the network can find the edge anywhere inside the patch and still detect it.

4.2 Average pooling

Each window collapses to its mean. Average pooling keeps more information about the distribution of activations (it preserves energy) but is less aggressive at discarding noisy single-pixel spikes. In modern architectures average pooling is used at the head of the network (global average pooling replaces the flatten-plus-dense pattern), and max pooling is used in the body for downsampling.

4.3 What pooling does

  • Spatial downsampling. A 2×22 \times 2 pool with stride 2 halves each spatial dimension; the volume goes from (C,H,W)(C, H, W) to (C,H/2,W/2)(C, H/2, W/2).
  • Invariance. A small translation of the input — moving the edge detector's response by one pixel — does not change the max-pooled output as long as the response stays inside the same pooling window. Stacking pooling layers gives a growing tolerance to translation.
  • Receptive-field growth. Each pooling layer doubles the effective receptive field of every later unit (Section 5).
  • Reduced compute. Smaller feature maps mean cheaper convolutions downstream.
import torch
import torch.nn as nn

x = torch.randn(1, 64, 32, 32)   # one image, 64 channels, 32x32 map
print(x.shape)

#Max pool 2x2, stride 2: spatial size halves, channels unchanged.
max_pool = nn.MaxPool2d(kernel_size=2, stride=2)
print(max_pool(x).shape)         # torch.Size([1, 64, 16, 16])

#Average pool over the whole map: collapses H and W to 1.
#This is "global average pooling", used at the network head.
gap = nn.AdaptiveAvgPool2d(output_size=1)
print(gap(x).shape)              # torch.Size([1, 64, 1, 1])

5. Receptive Field

The receptive field of a unit in layer LL is the region of the input image that can influence that unit's value. A unit in the first convolutional layer, with a 3×33 \times 3 filter, has a 3×33 \times 3 receptive field: only those nine input pixels can affect its output. A unit two layers deep, after two 3×33 \times 3 convolutions, has a 5×55 \times 5 receptive field. After three, it is 7×77 \times 7. In general, LL stacked 3×33 \times 3 convolutions give a receptive field of (2L+1)×(2L+1)(2L + 1) \times (2L + 1), growing linearly with depth.

5.1 The formula

For a stack of layers each with kernel size kik_i and stride sis_i, the receptive field of a unit at depth LL is

ri(L)=ri(L−1)+(kL−1)⋅∏j=1L−1sj,r_i(L) = r_i(L - 1) + (k_L - 1) \cdot \prod_{j=1}^{L - 1} s_j,

with ri(0)=1r_i(0) = 1. The product ∏sj\prod s_j is the cumulative stride — how many input pixels each step in layer LL's output corresponds to. Stride matters more than kernel size: doubling the stride at any layer roughly doubles the receptive-field growth of every layer that follows.

5.2 Why it has to grow

A classifier that says "this is a cat" must look at the cat, not at one whisker. The decision requires evidence from a region large enough to contain the whole object, which on ImageNet is up to most of the image. If the receptive field at the final layer is only 20×2020 \times 20 pixels, the network literally cannot see the whole cat. Modern classifiers are designed so the receptive field at the head covers the full input — for a 224×224224 \times 224 input going through five stride-2 downsamples (each halving the spatial map), the receptive field of any unit at the head is at least 32×3232 \times 32 in input pixels and typically larger because the strided convolutions and pooling windows accumulate.

5.3 Why two 3×33 \times 3 beats one 5×55 \times 5

The receptive field of two stacked 3×33 \times 3 same-padded convolutions is 5×55 \times 5. The receptive field of one 5×55 \times 5 convolution is also 5×55 \times 5. They are not equivalent: the stacked version has two ReLU non-linearities between input and output instead of one, giving the network more expressive power for fewer parameters (2⋅32⋅C2=18C22 \cdot 3^2 \cdot C^2 = 18C^2 versus 52⋅C2=25C25^2 \cdot C^2 = 25C^2 per channel pair). VGG's contribution was to recognise that this trade is almost always favourable, and to chain many such stacks together.

6. Feature Hierarchy

A trained CNN learns a hierarchy of features: low-level patterns in the early layers and high-level concepts in the late layers. The same hierarchy appears regardless of the task the network was trained on, because the early layers are solving the same problem — detect local patterns that are useful regardless of whether the final answer is "cat", "car", or "chest X-ray".

6.1 Edges, textures, parts, objects

The hierarchy, in order:

  1. Edges and colour blobs (layers 1–2). Oriented gradients, colour transitions, simple texture patches.
  2. Textures and simple shapes (middle layers). Grids, parallel lines, corners, circles, repeating motifs.
  3. Object parts (deeper layers). Eyes, wheels, wings — parts that are recognisable on their own but do not yet pin down a specific object.
  4. Whole objects (final layers). The activation at the head is driven by global evidence; a unit fires when it sees the whole object regardless of pose.

This is the answer to "why does the network work?" The early layers are general-purpose feature extractors that any image task benefits from, and the later layers are increasingly specialised to the classes the network was trained on.

6.2 Why early filters look like Gabor filters

A striking empirical finding: the first-layer filters of nearly any trained CNN on natural images, when visualised, look like a mix of oriented Gabor-like edge detectors and colour blobs. They were not designed that way. They were learned. The reason is that oriented edges are the most informative local patterns in natural images — any useful feature extractor has to detect them — and gradient descent discovers the same solution humans discovered by hand in classical vision.

This finding is the empirical justification for transfer learning (Section 8): if the early filters are nearly the same across tasks, we can train them on one large dataset and reuse them on another.

import torch
import torchvision

#Inspect what a trained AlexNet learned at layer 1.
model = torchvision.models.alexnet(weights=torchvision.models.AlexNet_Weights.DEFAULT)
first_conv = model.features[0]            # Conv2d(3, 64, kernel_size=11, stride=4)
weights = first_conv.weight.data.clone()  # shape (64, 3, 11, 11)

#Normalise each filter to [0, 1] for visualisation.
weights = weights - weights.amin(dim=(1, 2, 3), keepdim=True)
weights = weights / weights.amax(dim=(1, 2, 3), keepdim=True)
print(weights.shape)   # torch.Size([64, 3, 11, 11])
#Most filters, plotted, look like oriented Gabor wavelets
#in grayscale or colour-opponent patterns.

7. Canonical CNN Architectures

The history of convolutional architectures is the history of three ideas: going deeper (LeNet → AlexNet → VGG), going wider with smaller filters (VGG's contribution), and going deeper still without losing signal (ResNet's contribution).

7.1 LeNet-5 (1998)

LeNet-5 is the original convolutional architecture for digit recognition on MNIST. Two convolutional-pooling stages followed by two fully-connected layers, processing 32×3232 \times 32 grayscale digits into ten class scores. The architecture that introduced the modern convolutional-pooling-dense pattern; the network has roughly 60k parameters.

7.2 AlexNet (2012)

AlexNet won ImageNet 2012 by a large margin and ignited the deep learning era. Five convolutional layers and three fully-connected layers, trained on GPUs with ReLU activations, dropout, and data augmentation. The first layer uses 11×1111 \times 11 filters at stride 4 — large by modern standards but reasonable for 224×224224 \times 224 images. AlexNet has roughly 60 million parameters, almost all of them in the dense layers at the head.

7.3 VGG (2014)

VGG demonstrated that depth with small filters beats shallower networks with large ones. The standard VGG-16 has thirteen convolutional layers (all 3×33 \times 3, same-padded) interleaved with five max-pooling layers, followed by three fully-connected layers. VGG-16 has about 138 million parameters, and the bulk of them are still in the dense layers at the head. The lesson learned from VGG: stacking many small filters is more expressive than using a few large ones, and the parameter savings let the network go deeper.

7.4 ResNet (2015) and the degradation problem

Going deeper stopped working at some point. A 56-layer plain network on ImageNet had higher training and test error than a 20-layer plain network. This is the degradation problem: optimisation, not generalisation, is the bottleneck. Adding layers should at worst leave training error unchanged (the new layers could learn the identity), but in practice plain networks cannot even learn the identity cleanly once they are deep enough.

ResNet fixes this with the residual connection (or skip connection):

y=F(x,{Wi})+x,\mathbf{y} = \mathcal{F}(\mathbf{x}, \{W_i\}) + \mathbf{x},

where F\mathcal{F} is the residual branch — two or three convolutions and a non-linearity — and x\mathbf{x} is the identity shortcut from the input. If the residual branch learns nothing, the output equals the input and the layer is effectively an identity. This makes identity "free": the network can choose to pass information through unchanged by zeroing the residual branch, which is much easier than learning an identity from scratch.

A ResNet is a stack of residual blocks of this form, with the spatial size halved every few blocks by a strided convolution in the residual branch (and a 1×11 \times 1 convolution on the shortcut to match dimensions). ResNet-50, the most common variant, has 50 such blocks and roughly 25 million parameters. Crucially, the bulk of the parameters are gone from the head: the head is a global average pool plus a single dense layer, not three dense layers. The residual connection is what makes it feasible to optimise networks of that depth.

import torch
import torch.nn as nn

class ResidualBlock(nn.Module):
    """The core of ResNet: F(x) + x with a projection shortcut."""

    def __init__(self, in_ch, out_ch, stride=1):
        super().__init__()
        # Residual branch: conv -> bn -> relu -> conv -> bn.
        self.conv1 = nn.Conv2d(in_ch, out_ch, kernel_size=3,
                               stride=stride, padding=1, bias=False)
        self.bn1 = nn.BatchNorm2d(out_ch)
        self.conv2 = nn.Conv2d(out_ch, out_ch, kernel_size=3,
                               stride=1, padding=1, bias=False)
        self.bn2 = nn.BatchNorm2d(out_ch)
        self.relu = nn.ReLU(inplace=True)
        # Shortcut: 1x1 conv when stride or channel count changes.
        self.shortcut = (
            nn.Conv2d(in_ch, out_ch, kernel_size=1,
                      stride=stride, bias=False)
            if stride != 1 or in_ch != out_ch else nn.Identity()
        )

    def forward(self, x):
        out = self.relu(self.bn1(self.conv1(x)))
        out = self.bn2(self.conv2(out))
        out = out + self.shortcut(x)   # the residual connection
        return self.relu(out)

7.5 Shape comparison

ArchitectureYearConv layersParams (M)Core innovation
LeNet-5199820.06Convolutional-pooling-dense pattern for digits
AlexNet2012560ReLU + dropout + GPU training, 11×1111 \times 11 first-stage filters
VGG-16201413138Many 3×33 \times 3 filters stacked; depth via small filters
ResNet-5020154925Residual connections; identity is "free"; global avg pool head

The parameter column is instructive. ResNet-50 has roughly a fifth of VGG-16's parameters despite being three times as deep. The residual connection lets the network go deeper without going wider, and the global average pooling head removes the parameter-heavy dense layers that dominate VGG.

8. Transfer Learning and Pretrained Backbones

Training a deep CNN from scratch on a small dataset is a recipe for overfitting. The standard remedy is transfer learning: take a CNN that was pretrained on a large dataset (ImageNet, COCO, JFT) and reuse it on the new task, with one of two strategies.

8.1 Feature extraction

Freeze the convolutional backbone and train only a new classification head on top. The early and middle layers of the pretrained network already extract edges, textures, and parts (Section 6); for many tasks those features are exactly what is needed. This works best when:

  • The new dataset is small (a few hundred to a few thousand images).
  • The new task is similar in domain to pretraining (natural images on ImageNet weights; medical images on weights pretrained on radiology, or after a domain-adaptation step).
  • The classes are different from ImageNet — otherwise the head could just be the ImageNet classifier.

8.2 Fine-tuning

Unfreeze some or all of the backbone and continue training with a small learning rate. The backbone's weights, which were good for ImageNet, drift toward being good for the new dataset. This works best when:

  • The new dataset is medium-to-large (tens of thousands of images or more).
  • The new domain is different enough from pretraining that the pretrained features are not quite right.
  • The compute budget allows the extra training.

A common recipe is to fine-tune the later layers of the backbone at a higher learning rate and the early layers at a lower learning rate (or not at all). The intuition: the early filters are universal edge detectors and rarely need to change; the later filters are task-specific and benefit from adapting.

8.3 Dataset-size thresholds

Empirical rules of thumb for choosing the strategy on a natural-image task:

New dataset sizeRecommended strategy
Very small (< 1k images)Feature extraction: freeze the backbone, train only a linear head
Small (1k–10k)Feature extraction as a baseline; fine-tune the last block or two if accuracy is insufficient
Medium (10k–100k)Fine-tune the last several blocks; consider unfreezing the whole backbone with a low learning rate
Large (> 100k)Train from scratch or fine-tune the entire backbone with a moderate learning rate

These are not hard boundaries — a domain shift large enough that "cat" features look nothing like "satellite imagery" features can require full fine-tuning even on a medium dataset — but they are a defensible starting point.

import torchvision

#Pretrained ResNet-50 backbone, ready for either strategy.
weights = torchvision.models.ResNet50_Weights.DEFAULT
backbone = torchvision.models.resnet50(weights=weights)

#Strategy 1: feature extraction. Freeze every parameter.
for param in backbone.parameters():
    param.requires_grad = False

#Replace the head (the final fully-connected layer) with a new
#classifier for, say, 5 custom classes.
backbone.fc = torch.nn.Linear(backbone.fc.in_features, num_classes=5)

#backbone is now ready to train. Only backbone.fc will update;
#the backbone acts as a fixed feature extractor.

#Strategy 2: fine-tuning. Unfreeze selectively — say, layer4 only.
for name, param in backbone.named_parameters():
    if "layer4" not in name:
        param.requires_grad = False
    else:
        param.requires_grad = True
#Train with a low learning rate (e.g. 1e-4) on the new dataset.

8.4 Why this works

Three reasons, in order of robustness:

  1. The features are reusable. Section 6: early CNN layers learn the same Gabor-like edge detectors regardless of the supervised task. Pretrained on ImageNet, the backbone already extracts useful features for any natural-image problem.
  2. The optimisation is friendlier. Starting from pretrained weights puts the optimiser in a basin of the loss landscape that is much closer to a good solution than random initialisation. Fine-tuning needs fewer epochs to converge and is less likely to land in a bad local minimum.
  3. Implicit regularisation. Frozen early layers cap the hypothesis space, which is exactly what a small dataset needs.

9. Comparison and Practitioner Checklist

9.1 Convolution vs pooling

PropertyConvolution (strided)Max poolingAverage pooling
Has learnable parametersYes (filter weights, bias)NoNo
Effect on spatial sizeDownsamples by strideDownsamples by pool sizeDownsamples by pool size
Information keptWhatever the filter prefersMaximum activation in windowMean activation in window
Invariance introducedWeak (depends on weights)Strong (positional)Mild (averaging)
Use caseMost downsampling in modern netsOlder architectures (AlexNet, VGG)Head of network (global avg pool)
CostCout⋅Cin⋅k2C_{\text{out}} \cdot C_{\text{in}} \cdot k^2 per layerFreeFree

Modern networks prefer strided convolution for downsampling because the network learns the downsampling filter, and global average pooling at the head because it removes the parameter-heavy dense layers that VGG and AlexNet suffered from.

9.2 A short practitioner checklist

  1. Start with a pretrained backbone. ResNet-50 or a more recent EfficientNet/ConvNeXt, depending on the accuracy-versus-latency trade-off. Training from scratch is rarely the right default unless the new dataset is very large and very different from ImageNet.
  2. Pick the strategy by dataset size. Feature extraction for small data, fine-tuning for medium, fine-tune-the-whole-thing or from-scratch for large.
  3. Use odd kernel sizes and same padding. 3×33 \times 3 filters at stride 1 with padding 1 are the default building block; they keep spatial size and give a growing receptive field.
  4. Downsample with strided convolution. Pooling still has its uses (especially global average pooling at the head), but a 3×33 \times 3 strided convolution learns a better downsampler than a fixed max-pool.
  5. Watch the receptive field. Make sure the receptive field at the head covers the whole object. Five stride-2 downsamples is the typical minimum for ImageNet-sized inputs.
  6. Residual connections for anything deep. Past ~10 layers, plain convolutional stacks start to suffer from the degradation problem. Use ResNet-style residual blocks; identity shortcuts are free.

The shape of modern computer vision is roughly: a pretrained convolutional backbone, a small task-specific head, fine-tuned with a modest learning rate on the new data. That pattern is the lesson.

Key Takeaways

  • A convolutional layer is a dense layer with two restrictions baked in: local connectivity (each output looks at only a k×kk \times k patch) and weight sharing (the same filter is reused at every spatial position). For a 3×33 \times 3 filter over RGB input that replaces 150 million parameters with a few thousand, and it gives translation equivariance for free.
  • The output spatial size is H′=⌊(H+2p−k)/s⌋+1H' = \lfloor (H + 2p - k)/s \rfloor + 1. "Same" padding (p=(k−1)/2p = (k-1)/2 at stride 1) preserves spatial size; stride 2 halves it; "valid" padding (no padding) shrinks it by k−1k - 1 pixels per side per convolution.
  • Max pooling downsamples by taking the maximum of each window, which gives a small-translation invariance and reduces compute. Average pooling at the head (global average pooling) replaces the parameter-heavy dense layers of older architectures with a parameter-free summarisation.
  • The receptive field of a unit grows with depth and is multiplied by stride at each layer. Five stride-2 downsamples is the canonical way to grow it from 3 pixels to 32+ pixels for ImageNet-sized inputs. The receptive field must cover the whole object or the network literally cannot recognise it.
  • Trained CNNs learn a hierarchy of features, edges and colour blobs in early layers, textures and parts in middle layers, whole objects in the late layers. The early layers look like Gabor filters across tasks, which is the empirical reason transfer learning works.
  • ResNet introduced the residual connection y=F(x)+x\mathbf{y} = \mathcal{F}(\mathbf{x}) + \mathbf{x} to fix the degradation problem — the observation that adding layers to a plain network makes training error worse. The skip connection makes identity "free", so the network can choose to pass information through unchanged by zeroing the residual branch.
  • Transfer learning is the default workflow for new image tasks: feature extraction (freeze the backbone, train only a new head) for small datasets, fine-tuning (unfreeze some or all of the backbone with a small learning rate) for medium datasets, and full training from scratch only for large datasets with a domain far from the pretraining corpus.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

A 224x224 RGB image is flattened and fed to a fully-connected layer with 1000 output units. Roughly how many parameters does that layer have before any training?

0 of 8 answered

Pick a lesson to start the audio.