Skip to content
Kudos AI

Tagged “transformers”

2 articles.

8 min readBuilding a Language Model

The Transformer Architecture

Assembling a GPT from attention: multi-head projections, layer normalization worked by hand, why shortcut connections rescue the gradient, the 4x feed-forward expansion, and a parameter count that reproduces GPT-2 small at 124 million exactly.

Generative AIDeep LearningNatural Language Processing
7 min readBuilding a Language Model

Attention and Self-Attention

Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.

Generative AIDeep LearningNatural Language Processing