Skip to content
Kudos AI

What a Network Computes

Layers as parameterised functions, the affine-then-nonlinearity pattern, and why stacking linear maps alone buys you nothing.

IntermediateModule 130 min · 120 XP
Three stacked linear layers collapsing on screen into a single one, and then failing to collapse once a nonlinearity is put between them.

A neural network is a composition of simple functions with adjustable constants. Nothing about that sentence is metaphorical, and this lesson makes it concrete before any training is involved.

A layer is a function with parameters

The standard fully connected - dense - layer computes

h=σ(Wx+b),\mathbf{h} = \sigma(W\mathbf{x} + \mathbf{b}),

where WW is a weight matrix, b\mathbf{b} a bias vector, and σ\sigma an elementwise nonlinearity. Two distinct things happen: an affine map Wx+bW\mathbf{x} + \mathbf{b}, then a nonlinearity applied to each coordinate independently.

Chollet's framing is useful: layers are the building bricks of deep learning, and the shape of the data selects the brick. Vectors go to dense layers; sequences in 3D tensors go to recurrent or 1D convolution layers; images in 4D tensors go to 2D convolution layers.

Counting parameters

A dense layer from nn inputs to mm outputs holds an m×nm \times n matrix plus a bias per output unit:

parameters=mn+m.\text{parameters} = mn + m .

From 4 inputs to 3 outputs that is 3×4+3=153 \times 4 + 3 = 15. Forgetting the bias vector is the usual slip and gives 12.

Why the nonlinearity is not optional

Suppose we drop σ\sigma and stack three dense layers:

W3(W2(W1x+b1)+b2)+b3.W_3\big(W_2(W_1\mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2\big) + \mathbf{b}_3 .

Multiply it out. The weight matrices collapse into a single product W=W3W2W1W = W_3W_2W_1 and the biases collapse into a single vector, leaving Wx+bW\mathbf{x} + \mathbf{b} - an affine map. Three layers, and the family of functions is exactly what one layer could already express.

Depth alone buys nothing. A composition of affine maps is affine. Whatever expressive power depth provides arrives only once a nonlinearity sits between the layers, which is why σ\sigma is structural rather than decorative.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

ReLU, precisely

The standard choice is the rectified linear unit:

ReLU(z)=max⁡(0,z).\text{ReLU}(z) = \max(0, z).

Raschka describes it plainly: it thresholds negative inputs to 0. That is the whole definition. It does not squash values into (0,1)(0,1) - that is the logistic sigmoid - and it does not normalise anything.

Its effect on the backward pass matters as much as on the forward one: the derivative is 11 where z>0z > 0 and 00 where z<0z < 0, so a unit that is inactive for an example passes no gradient back through it for that example.

A worked forward pass

Two inputs, two hidden units, one output:

x=[12],W(1)=[0.10.30.20.4],W(2)=[0.5−0.5]\mathbf{x} = \begin{bmatrix}1\\2\end{bmatrix}, \quad W^{(1)} = \begin{bmatrix}0.1&0.3\\0.2&0.4\end{bmatrix}, \quad W^{(2)} = \begin{bmatrix}0.5&-0.5\end{bmatrix} z(1)=[0.1(1)+0.3(2)0.2(1)+0.4(2)]=[0.71.0]\mathbf{z}^{(1)} = \begin{bmatrix}0.1(1)+0.3(2)\\0.2(1)+0.4(2)\end{bmatrix} = \begin{bmatrix}0.7\\1.0\end{bmatrix}

Both entries are positive, so ReLU leaves them unchanged, and

y^=0.5(0.7)−0.5(1.0)=−0.15.\hat y = 0.5(0.7) - 0.5(1.0) = -0.15 .

That single number is what the network currently believes. The next lesson asks how wrong it is and pushes that error backwards to every weight.

The figure below is that same network, sliced: the output against the first input, with the second on a slider. Take the nonlinearity away and the slice is a straight line, because the two layers multiply out to the single row [-0.05, -0.05]. The distance from affine drops to zero at machine precision, and the collapse is something you can check rather than take on faith.

Put ReLU back and notice where you are standing. At the worked input (1, 2) both units are on, so the rectifier does nothing at all, and the network and its collapsed form both return -0.15. The nonlinearity earns its keep only across a boundary: at (2, -1.5) both units are off and the answers are 0 and -0.025.

So depth does not buy a curve. It buys a set of regions, each of which is still an affine map, and the kinks in the slice are exactly where a pre-activation crosses zero. That is worth carrying into the next lesson, because it is also why a unit that is off passes no gradient back.

Interactive: the collapse, and what stops it

Take the nonlinearity away and the slice becomes a straight line.

Pre-activations
0.700, 1.000
Output
-0.150
Collapsed layer says
-0.150
Units switched off
0
Distance from affine
0.150

The two layers multiply out to the single row [-0.050, -0.050], so without a nonlinearity three layers or thirty express exactly what one expresses: the distance from affine is 2e-16, which is zero to machine precision. Put ReLU back and it is 0.150. But look at where you are. Both units are on, so the rectifier is doing nothing here, and the network returns -0.150 while its collapsed form returns -0.150: the same number. That is also true at the lesson’s own input of (1, 2). Depth does not buy a curve. It buys a set of REGIONS, each of which is still an affine map, and the kinks in the slice are where a pre-activation crosses zero. Move the inputs across one and the two answers part company. One counting note while the layer is in view: a dense layer from four inputs to three outputs holds 15 parameters, not 12. The bias vector is the part that gets forgotten.

Continue with backpropagation and gradient descent for the general treatment.

References & further reading

  • François Chollet, Deep Learning with Python, Manning (2nd edition, MEAP), 2020· Kudos AI reference library
  • Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.