Understanding Transformer
Recurrent networks process a sequence one step at a time, each step depending on the previous one’s output. That serial dependency prevents parallelism across the sequence during training and makes long-range dependencies traverse many intermediate steps. The transformer’s contribution was showing that attention alone, with no recurrence, was sufficient, which removed both problems at once. Raschka records that the architecture was introduced in the 2017 paper "Attention Is All You Need" and was originally developed for machine translation.
Discarding recurrence costs the model its sense of order, because attention treats its input as a set: permuting the tokens would permute the outputs identically. Position information must therefore be added explicitly, by adding a position-dependent vector to each token embedding before the first layer. The original design used fixed sinusoids; learned and rotary position embeddings are common alternatives, but the requirement itself is unavoidable.
Each transformer block contains two sublayers doing complementary work. Multi-head self-attention mixes information across positions, deciding what each token should read from the rest of the sequence. A position-wise feed-forward network then transforms each position independently, using the same weights everywhere. Both sublayers are wrapped in residual connections and layer normalization, which is what makes stacking dozens of them trainable.
Three arrangements dominate. Encoder-only models such as BERT see the entire input at once and are suited to understanding tasks like classification. Decoder-only models, the GPT family, use causal masking so each position sees only what precedes it, and are trained to predict the next token, which makes them naturally generative. Encoder-decoder models keep both halves and are used where an input sequence must be transformed into an output sequence, as in translation. Modern large language models are overwhelmingly decoder-only.
Example of Transformer
A decoder-only language model embeds each token, adds a positional encoding, then passes the sequence through many identical blocks. Within each block, masked self-attention lets each position gather context from earlier positions, and the feed-forward sublayer processes each position independently.
At the output, a final linear layer maps each position’s representation to a score for every token in the vocabulary, and a softmax converts these into a probability distribution over what comes next. Training minimizes cross-entropy against the actual next token, over enormous quantities of text.
The same skeleton transfers beyond language. Raschka notes the vision transformer, introduced in "An Image is Worth 16x16 Words" (2020), which cuts an image into patches and treats each patch as a token, showing the architecture is not tied to text.
Frequently Asked Questions
Why is positional encoding necessary?
Unmasked self-attention is permutation-equivariant: it has no built-in notion of order, so "dog bites man" and "man bites dog" would be represented identically. Position information has to be supplied explicitly for the model to distinguish them (a causal mask carries some order of its own, but only in one direction).
What does the feed-forward sublayer contribute?
Attention moves information between positions but applies no substantial per-position transformation. The feed-forward network supplies that processing capacity, applied identically and independently at every position. The two sublayers do genuinely different jobs.
Why are modern language models decoder-only?
Next-token prediction is a single objective requiring no labelled data, applies to any text, and scales cleanly. At sufficient scale it produces models capable across tasks that once needed separate architectures, so the simpler design won on scalability.
The Bottom Line
The transformer replaced recurrence with attention, gaining parallel training and direct long-range connections at the cost of needing explicit position information. Its encoder-only, decoder-only, and encoder-decoder variants now underpin essentially all modern language modelling.