Language Models are Few-Shot Learners
Tom B. Brown et al. · 2020 · arXiv:2005.14165
Summary
Describes GPT-3 and shows that a sufficiently large decoder-only language model can perform new tasks from a handful of examples supplied in its prompt, with no gradient updates.
Why it matters
It demonstrated that scale alone changes what a model can do, and introduced in-context learning as an alternative to fine-tuning. Raschka identifies this as the paper describing the decoder-only GPT-3 model that inspired modern large language models and serves as the template for building one from scratch.