Skip to content
Kudos AI

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin et al. · 2018 · arXiv:1810.04805

Natural Language ProcessingDeep LearningGenerative AIView source ↗

Summary

Pretrains a transformer encoder to predict masked tokens using context from both directions, then fine-tunes the same model on downstream tasks with a small task-specific head.

Why it matters

It made pretrain-then-fine-tune the standard recipe in natural language processing. Because masking lets the model condition on both left and right context, the learned representations proved far stronger for understanding tasks than left-to-right pretraining, and the two-stage workflow became the field’s default.