An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy et al. · 2020 · arXiv:2010.11929
Summary
Applies a standard transformer directly to images by cutting each image into fixed-size patches and treating the sequence of patches as tokens, with no convolutions.
Why it matters
It showed the transformer is not specific to language: given enough data, a general sequence architecture matches or beats convolutional networks on image classification. Raschka cites it as illustrating that transformer architectures are not restricted to text inputs.