Skip to content
Kudos AI

Attention Is All You Need

Ashish Vaswani et al. · 2017 · arXiv:1706.03762

Generative AIDeep LearningNatural Language ProcessingView source ↗

Summary

Introduces the transformer, an architecture built entirely from attention and feed-forward layers with no recurrence, originally developed for machine translation.

Why it matters

Removing recurrence let every position in a sequence be processed in parallel during training and put any two positions one operation apart, which together made training at previously impractical scale feasible. Raschka records that this paper proposed the original transformer architecture; essentially every modern large language model descends from it.