Skip to content
Kudos AI

Overfitting

When a model learns the noise and idiosyncrasies of its training data rather than the underlying pattern, so it performs well in training and poorly on new data.

Understanding Overfitting

A model is fitted on a sample, but judged on data it has never seen. Overfitting is what happens when it treats accidental features of that particular sample, measurement noise, quirks of which observations happened to be drawn, as if they were real structure. The training error keeps falling while the ability to generalize quietly degrades.

Flexibility is the mechanism. A sufficiently flexible model can pass exactly through every training point; a curve threaded through all of them will have zero training error and no predictive value. The more capacity a model has relative to the amount of data constraining it, the more of the sample’s noise it is able to absorb.

The failure mode has a mirror image. A model that is too rigid, a straight line fitted to a genuinely curved relationship, cannot represent the real pattern regardless of how much data it sees. That is underfitting, and no amount of extra data fixes it. Choosing model complexity is the act of navigating between these two failures, which is precisely the bias-variance trade-off.

Because training error is a biased and optimistic estimate of future performance, detecting overfitting requires data the model has not been fitted on. That is the purpose of a held-out test set, and of cross-validation when data is too scarce to give a large sample away.

Example of Overfitting

Suppose ten points are generated from a gentle quadratic relationship, plus a little random measurement noise. Fit a straight line and it will miss the curvature: both training and test error stay high. That is underfitting.

Now fit a ninth-degree polynomial. It has enough freedom to pass exactly through all ten points, driving training error to zero. But between and beyond those points it swings wildly, because those swings are fitted to noise rather than signal. Its test error is far worse than the straight line’s.

A quadratic fit, matching the process that actually generated the data, has small training error and comparably small test error. The lesson is that the best model is the one whose complexity matches the structure in the data, not the one that fits the training sample most closely.

Frequently Asked Questions

How can I tell whether my model is overfitting?

Compare performance on data the model was fitted on against data it has never seen. A small training error alongside a substantially larger validation or test error is the signature of overfitting. If both errors are large and similar, the problem is underfitting instead.

Does more training data fix overfitting?

It generally helps, because a fixed amount of model flexibility is spread over more constraints, leaving less freedom to fit noise. It does not help underfitting at all, where the limitation is the model’s form rather than the quantity of data.

Is a model with zero training error always overfitted?

Not necessarily, but it is a strong warning sign and demands checking against held-out data. On genuinely easy, low-noise problems a good model can legitimately fit the training data almost perfectly while still generalizing.

The Bottom Line

Overfitting is the failure to distinguish signal from noise in a finite sample. It is controlled by matching model flexibility to the evidence available, and it can only be detected on data the model has not seen, which is why disciplined validation is not optional.