8 min readBuilding a Language Model
The Transformer Architecture
Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.
Generative AIDeep LearningNatural Language Processing