7 min readBuilding a Language Model
Pretraining and Fine-Tuning
How next-word prediction turns unlabelled text into supervision, why cross entropy is just negative average log probability, what perplexity really measures, and why a model that completes text fluently still cannot follow an instruction.
Generative AIDeep Learning