Deep Learning
Neural networks from first principles: backpropagation, architectures, representation learning, and why depth buys something width does not.
Learning paths (2)
Deep Learning Foundations
What a neural network actually computes, how the chain rule delivers every gradient in one backward sweep, and why convolution is the right prior for an image.
Language Models and Generative AI
How text becomes numbers, how attention lets a token gather context from the whole sequence, and what pretraining then fine-tuning actually do to the weights.
Encyclopedia (8)
Backpropagation
The algorithm that computes the gradient of a neural network’s loss with respect to every weight, by applying the chain rule backwards through the network.
Cross-Entropy
A measure of the difference between two probability distributions, used as the standard loss function for classification.
Neural Network
A model composed of layers of simple units, each computing a weighted sum followed by a non-linear function, fitted by gradient descent using backpropagation.
Activation Function
The non-linear function applied to a layer’s output, without which a network of any depth would collapse to a single linear transformation.
Convolutional Neural Network
A neural network that applies learned filters across an input’s spatial extent, sharing weights so the same pattern is detected wherever it occurs.
Self-Attention
A mechanism that lets every position in a sequence attend to every other, computing each output as a weighted sum of values whose weights come from query-key similarity.
Transformer
A neural architecture built on stacked self-attention and feed-forward layers, which replaced recurrence as the standard for sequence modelling.
Pretraining and Fine-Tuning
The two-stage recipe of first training a model on a large generic corpus, then adapting it to a specific task with a much smaller labelled dataset.
Articles (8)
What Actually Makes Training Converge
A two per cent change in the learning rate separates a converged run from one five orders of magnitude away, a condition number predicts the convergence rate to six decimal places, and stochastic gradient descent with a fixed step never converges at all - it settles into a ball whose radius grows as the square root of the step. Every figure here was computed on a problem whose exact optimum is known.
What Is a Neural Network?
Layers as parameterised transformations, the forward pass, and why depth and non-linearity are not optional: a proof that no single linear layer can compute XOR, and a two-layer network that does, worked entirely by hand.
Tokenization and Embeddings
How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.
The Transformer Architecture
Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.
Pretraining and Fine-Tuning
How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.
Convolutional Networks for Vision
Convolution defined properly, a Sobel edge detector worked by hand on a 5x5 image, why sliding one small kernel over an image beats a dense layer by five orders of magnitude in parameters, and what changed when kernels stopped being designed and started being learned.
Backpropagation and Gradient Descent
How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.
Attention and Self-Attention
Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.
Tools (1)
Datasets (2)
Research (7)
The Perceptron: A Perceiving and Recognizing Automaton
Introduces the perceptron, a trainable unit that computes a weighted sum of its inputs and fires if the sum exceeds a threshold, with a rule for adjusting weights from labelled examples.
Learning Internal Representations by Error Propagation
Presents backpropagation as a general method for training multilayer networks, showing that hidden layers can learn useful internal representations rather than needing to be designed by hand.
Neural Machine Translation of Rare Words with Subword Units
Adapts byte pair encoding to text segmentation, building a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair, so that rare words decompose into known fragments.
Attention Is All You Need
Introduces the transformer, an architecture built entirely from attention and feed-forward layers with no recurrence, originally developed for machine translation.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Pretrains a transformer encoder to predict masked tokens using context from both directions, then fine-tunes the same model on downstream tasks with a small task-specific head.
Language Models are Few-Shot Learners
Describes GPT-3 and shows that a sufficiently large decoder-only language model can perform new tasks from a handful of examples supplied in its prompt, with no gradient updates.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Applies a standard transformer directly to images by cutting each image into fixed-size patches and treating the sequence of patches as tokens, with no convolutions.
Projects (2)
GPT From Scratch
A decoder-only transformer built up one component at a time - tokenizer, embeddings, attention, blocks, pretraining loop - with no modelling code imported from a library.
Neural Network From Scratch
A feed-forward network in NumPy with hand-derived backpropagation, validated against numerical gradients so the calculus is proven rather than trusted.