Understanding Pretraining and Fine-Tuning
Training a capable model from scratch for each task would demand a large labelled dataset every time, which is rarely available. The two-stage approach separates learning about the domain in general from learning the specific task. The first stage is expensive and done once; the second is cheap and repeated per application.
Pretraining works because its objective is self-supervised. Predicting the next token requires no annotation: the text itself supplies the answer, so effectively unlimited raw text can be used. To predict well the model must acquire a great deal of incidental competence, in syntax, factual association, and discourse structure, and that competence is what fine-tuning later exploits.
Fine-tuning continues training on task-specific data, usually at a much smaller learning rate so the pretrained knowledge is adjusted rather than overwritten. Because the representations are already good, the amount of labelled data needed is small, often a few thousand examples where training from scratch would need millions. Raschka frames the distinction as instruction and classification fine-tuning applied after pretraining a base model.
The economics of this split explain much of the current landscape. Pretraining costs are enormous and concentrated among a few organizations, while fine-tuning is within reach of almost anyone, so a single base model spawns many specialized descendants. Parameter-efficient methods, which update only a small set of added parameters and leave the base weights frozen, push the cost down further still.
Example of Pretraining and Fine-Tuning
A base language model is pretrained on a very large text corpus with a single objective: predict the next token. It is never told about sentiment, summarization, or question answering, yet it acquires the linguistic competence those tasks depend on.
To build a sentiment classifier from it, a small classification head is attached and the model is fine-tuned on a few thousand labelled reviews. It already understands negation, sarcasm cues, and intensifiers from pretraining, so it only has to learn the mapping onto sentiment labels.
Instruction tuning is the same mechanism applied differently: fine-tuning on examples of instructions paired with good responses turns a raw next-token predictor into something that follows requests. The underlying capability came from pretraining; fine-tuning shaped how it is expressed.
Frequently Asked Questions
Why does pretraining need no labels?
The objective is constructed from the data itself. Given a passage of text, the next token is already known, so the supervision signal is free. This self-supervision is what allows training on corpora far larger than any labelled dataset could ever be.
What is catastrophic forgetting?
When fine-tuning overwrites pretrained knowledge, so the model gains the narrow task but loses general capability. It is mitigated by using small learning rates, fine-tuning for few steps, or freezing most parameters and training only a small added set.
How does this differ from few-shot prompting?
Fine-tuning changes the model’s weights and persists. Few-shot prompting changes nothing: examples are placed in the context window and influence only that one request. Prompting is immediate and free; fine-tuning is more durable and usually more accurate for a fixed narrow task.
The Bottom Line
Pretraining buys broad capability from unlabelled data at great expense once; fine-tuning converts it into a specific competence cheaply and repeatedly. That division is why a handful of base models now sit beneath a very large number of applications.