Generative AI
Models that produce rather than classify: likelihood and latent-variable models, diffusion, and the generative side of the transformer.
Learning paths (1)
Encyclopedia (4)
Tokenization
The process of splitting text into the discrete units a language model actually operates on, typically subword fragments rather than whole words.
Self-Attention
A mechanism that lets every position in a sequence attend to every other, computing each output as a weighted sum of values whose weights come from query-key similarity.
Transformer
A neural architecture built on stacked self-attention and feed-forward layers, which replaced recurrence as the standard for sequence modelling.
Pretraining and Fine-Tuning
The two-stage recipe of first training a model on a large generic corpus, then adapting it to a specific task with a much smaller labelled dataset.
Articles (4)
Tokenization and Embeddings
How text becomes numbers a model can train on: building a vocabulary, why byte pair encoding never needs an unknown token, the embedding layer as a lookup that is provably one-hot times a matrix, and why position has to be added back in by hand.
The Transformer Architecture
Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.
Pretraining and Fine-Tuning
How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.
Attention and Self-Attention
Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.
Tools (1)
Research (4)
Attention Is All You Need
Introduces the transformer, an architecture built entirely from attention and feed-forward layers with no recurrence, originally developed for machine translation.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Pretrains a transformer encoder to predict masked tokens using context from both directions, then fine-tunes the same model on downstream tasks with a small task-specific head.
Language Models are Few-Shot Learners
Describes GPT-3 and shows that a sufficiently large decoder-only language model can perform new tasks from a handful of examples supplied in its prompt, with no gradient updates.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Applies a standard transformer directly to images by cutting each image into fixed-size patches and treating the sequence of patches as tokens, with no convolutions.