Convolutional Networks for Vision
Convolution defined properly, a Sobel edge detector worked by hand on a 5x5 image, why sliding one small kernel over an image beats a dense layer by five orders of magnitude in parameters, and what changed when kernels stopped being designed and started being learned.
Prerequisites: What Is a Neural Network?
Feed a photograph to the dense layer from What Is a Neural Network? and something goes wrong immediately: a modest colour image has 150,528 numbers in it, and connecting every one of them to every unit of a 1,000-unit layer needs over 150 million weights - for one layer. Worse, the model would have to learn what an edge looks like separately for every position in the image, because nothing ties the weight at one pixel to the weight at its neighbour.
Convolution fixes both problems with a single idea: slide one small set of weights over the whole image. This article defines the operation precisely, works an edge detector by hand, and then covers the shift that made it central to vision - the kernels stopped being designed by people and started being learned.
A. Images, and what "low-level" vision means
An image is an array of intensities. Russell and Norvig characterise the first operations in a vision pipeline by two properties: they are local, meaning they can be carried out in one part of the image without regard for anything more than a few pixels away, and they involve no knowledge, meaning they run without any consideration of what objects might be in the scene. They note this makes such operations good candidates for parallel hardware - a GPU, or an eye.
Locality is the property convolution is built to exploit. A detector for a vertical edge needs to see a handful of neighbouring pixels; it does not need the other 150,000.
B. Convolution, defined
Russell and Norvig give the definition directly. The function is the convolution of and , written , when
In practice - the kernel or filter - is zero outside a small window, so the sums are over a or patch rather than the whole plane. Each output pixel is a weighted sum of a small neighbourhood of input pixels, with the same weights used at every position.
A notational honesty note. The definition above has : the kernel is flipped before being applied. Deep-learning frameworks skip the flip and compute - strictly cross-correlation
- while still calling the layer a convolution. For the Sobel kernel below the two differ by a sign: our patch gives by cross-correlation and by true convolution. It makes no practical difference in a network, because the kernel is learned: whatever the framework's convention, training simply finds the kernel that works under it. It matters only when you are comparing formulas across a textbook and a library.
C. A worked edge detector
Edges, in Russell and Norvig's definition, are straight lines or curves in the image plane across which there is a significant change in brightness. The point of finding them is compression of a sort: to abstract away from the messy, multimegabyte image toward a more compact representation. They are careful to distinguish four physical causes that all produce the same kind of image edge - depth discontinuities, surface-orientation discontinuities, reflectance discontinuities, and illumination discontinuities such as shadows. Edge detection works on the image alone and cannot tell them apart; later processing must.
Take a image with a single vertical edge - dark on the left, bright on the right:
is the horizontal Sobel kernel: it subtracts what is on the left from what is on the right, so it responds to vertical edges and ignores flat regions.
A flat patch. The top-left window is entirely s:
The negative and positive weights cancel exactly. A kernel whose weights sum to zero returns zero on any constant patch - it measures difference, not brightness.
A patch straddling the edge. The window starting at column 2:
Sliding the unflipped kernel over all nine valid positions - cross-correlation, written , as frameworks compute it; the true convolution negates every entry:
The output - a feature map - is a picture of where the edge is. Flat regions are zero; the two columns spanning the intensity jump light up.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Interactive: nine weights, used everywhere
Valid padding, so the output is two smaller in each direction.
Image
Kernel
=
- This window
- 320
- Kernel weights sum to
- zero
320 here, because the window straddles a change. Move to a flat part of the image and the output drops to exactly zero. That contrast is the whole of edge detection, and it comes from nine numbers that sum to zero. Note also that those same nine numbers are used at every stop - the kernel is not re-learned per position, which is a claim that an edge looks the same wherever it appears.
D. Smoothing first, and a theorem that saves a pass
Real images are noisy, and differencing amplifies noise. The standard remedy is to smooth first with a Gaussian,
convolving the image with it, . Russell and Norvig note that of one pixel smooths a small amount of noise while two pixels smooths more but costs detail, and that since the Gaussian's influence fades quickly the infinite sums can be truncated at .
There is a nice economy available here. It is a theorem that for any and ,
the derivative of a convolution equals the convolution with the derivative. So instead of smoothing the image and then differentiating it, you can convolve the image once with the derivative of the smoothing function, , and mark as edges those peaks in the response that exceed a threshold. Two passes become one.
E. Why not just use a dense layer
Two properties make convolution the right tool, and both are countable.
Parameter sharing. One kernel over 3 colour channels is weights, and a layer of 64 such filters is . The dense layer from the opening - inputs to 1,000 units - is weights. That is a factor of about 87,000. The dense layer also has to learn an edge detector independently at every location; the convolutional layer learns one and applies it everywhere.
Equivariance. Widen our image to six columns, so that the shifted edge stays inside the valid region, and move the edge one column to the right: the response moves with it, unchanged in value:
(rows omitted; every row is identical). The detector does not care where the edge is, which is exactly right - an edge is an edge wherever it appears.
Equivariance is not invariance. The response translated with the input rather than staying fixed. Convolution gives translation equivariance; a classifier that must output the same label regardless of position needs something further - typically downsampling or pooling stages that progressively discard spatial resolution, so that by the final layer a shifted input maps to the same answer. The two terms are routinely conflated, and they are not the same property.
F. When kernels stopped being designed
Everything above uses a kernel somebody chose. Sobel weights are a human's encoding of "look for a horizontal brightness gradient". This was how vision worked for decades, and Chollet's example of the era is exact: before convolutional networks succeeded on MNIST digit classification, solutions were typically based on hardcoded features such as the number of loops in a digit image, the height of the digit, or a histogram of pixel values.
The change is that the kernel entries became weights, learned by the same gradient descent that trains any other layer. Chollet's framing of deep learning is that it automates feature engineering entirely - you learn all the features in one pass instead of engineering them yourself, which often replaces a sophisticated multistage pipeline with a single end-to-end model. Stack such layers and you get what section A promised: successive layers of increasingly useful representations, early ones responding to edges and later ones to combinations of them.
The empirical verdict was decisive. Chollet records that since 2012 convnets have been the go-to algorithm for essentially all computer-vision tasks, that by 2015 the winning ImageNet top-five accuracy reached 96.4% and the task was considered solved, and that after 2015 it was nearly impossible to find a presentation at a major vision conference that did not involve them.
Learned does not mean unconstrained. A convolutional layer still bakes in strong assumptions - that useful features are local, and that a feature worth detecting in one place is worth detecting everywhere. Those assumptions are what make it efficient, and they are also why it is the wrong architecture when they do not hold. Chollet lists the choice of architecture as one of the modelling priors you are responsible for, alongside the loss function and training configuration.
Key takeaways
- Convolution is - every output pixel is a weighted sum of a small neighbourhood, with the same weights everywhere.
- Frameworks compute cross-correlation (no kernel flip) and call it convolution; with learned kernels it makes no practical difference.
- A kernel whose weights sum to zero measures difference, not brightness: our Sobel filter returns on flat patches and across the edge (unflipped; as a true convolution).
- Smoothing before differencing controls noise, and collapses the two passes into one.
- Convolution beats a dense layer by roughly 87,000× in parameters on a image, and shares one detector across all positions.
- Convolution is translation equivariant, not invariant - the response moves with the feature.
- The decisive change was learning the kernels rather than designing them, replacing hardcoded features like "number of loops in a digit".
What's next
Images are grids with a natural notion of locality. Text is not - the relevant context for a word can sit anywhere in the sequence, so a fixed local window is the wrong prior entirely. Handling that requires a different mechanism, developed in Attention and Self-Attention, and before it, a way to turn text into vectors at all: Tokenization and Embeddings.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
- François Chollet, Deep Learning with Python, Manning (2nd edition, MEAP), 2020· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.