Understanding Neural Network
A neural network is a stack of layers, each performing two operations. First a linear step multiplies the incoming values by a weight matrix and adds a bias. Then a non-linear activation function is applied element-wise. The output of one layer is the input to the next, so the network as a whole is a deeply nested composition of these simple pieces.
The non-linearity is not decorative. Composing linear maps yields another linear map, so a network of any depth with no activation function is exactly equivalent to a single linear layer and can represent nothing more. It is the non-linearity between layers that lets the composition express functions no single layer could.
What depth buys is compositional representation. Early layers learn simple, local structure, and later layers combine those into progressively more abstract features. This is learned rather than designed: the network is told only what to predict and how wrong it currently is, and the intermediate representations emerge from optimizing that objective.
Training is ordinary supervised learning machinery. A loss function measures the gap between predictions and targets, backpropagation computes the gradient of that loss with respect to every weight, and an optimizer steps the weights downhill. Everything else, the activation functions, normalization layers, residual connections, regularizers, exists to make this optimization work reliably at depth.
Example of Neural Network
A small network for classifying 28×28 grayscale digit images flattens each image into 784 inputs, passes them through a couple of hidden layers of a few dozen units each with ReLU activations, and ends in a 10-unit output layer with a softmax, giving a probability for each digit.
Chollet illustrates precisely this shape: dense layers of sizes 784 → 32 → 64 → 32 → 10, with ReLU on the hidden layers and softmax on the output. The softmax turns the final scores into a probability distribution over the ten classes, which pairs with cross-entropy loss.
The parameter count grows quickly. The first layer alone holds 784 × 32 weights plus 32 biases, over 25,000 parameters, which is why even modest networks need substantial data or regularization, and why architectures that share weights, such as convolutional layers, are preferred for images.
Frequently Asked Questions
Why is a non-linear activation function required?
Because a composition of linear functions is itself linear. Without a non-linearity, a hundred-layer network would have exactly the representational power of a single linear layer, and depth would buy nothing at all.
How many layers and units are needed?
There is no formula. The choice is empirical, guided by the size and complexity of the data and constrained by overfitting. Practice is usually to start from an architecture known to work on similar problems and adjust from there.
What makes a network "deep"?
Having many layers, and therefore many successive stages of learned representation. The contrast is with shallow models that map inputs to outputs in one step; depth is what enables the hierarchical composition of features.
The Bottom Line
A neural network alternates linear transformations with non-linearities, and depth lets it compose simple learned features into complex ones. Backpropagation supplies the gradients and gradient descent does the fitting; nearly everything else in deep learning exists to keep that process stable.