Skip to content
Kudos AI

Regularization

Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.

Also known as: Ridge regression, Lasso, Weight decay

Ridge asymptoting toward zero while lasso crosses it and is held there, and the same fact drawn as a disc without corners against a diamond with them.

Understanding Regularization

Fitting a model by minimizing training error alone gives it every incentive to contort itself toward the training sample. Regularization changes the objective: instead of minimizing error alone, the model minimizes error plus a penalty that grows with the size of its parameters. Large, elaborate fits must now justify themselves against that cost, so the model settles for a simpler one unless the data really supports more.

The two classical penalties behave differently. The L2 penalty, the sum of squared coefficients, gives ridge regression. Its gradient is proportional to the coefficient itself, so it shrinks every coefficient smoothly toward zero without reaching it; correlated predictors tend to share the influence between them. The L1 penalty, the sum of absolute values, gives the lasso. Its gradient has constant magnitude, which is enough to drive coefficients exactly to zero, so the fitted model performs variable selection and yields a sparse, more interpretable result.

Both are expressions of the bias-variance trade-off. Shrinking coefficients away from their least-squares values introduces bias, since the fitted model is no longer the best possible fit to the training data. In exchange, the fit becomes far less sensitive to which observations happened to be sampled. When predictors are numerous or collinear, that trade is usually strongly favourable.

The idea generalizes well beyond linear models. In neural networks the same L2 penalty appears as weight decay, and other devices, dropout, early stopping, data augmentation, act as regularizers even though they add no explicit penalty term. All of them limit how precisely the model can conform to its training sample.

How to Calculate

minimize L(θ) + λ · P(θ), where P(θ) = Σⱼ θⱼ² (L2) or Σⱼ |θⱼ| (L1)

where

L(θ)
the ordinary loss measured on the training data
P(θ)
the penalty on parameter magnitude
λ
the regularization strength, controlling the trade-off

Example of Regularization

Consider a regression with many predictors, several of them strongly correlated. Ordinary least squares can produce large coefficients of opposing sign that cancel one another out: the fit is good on the training data but wildly unstable, since a slightly different sample would produce very different coefficients.

Applying an L2 penalty suppresses that behaviour. Large opposing coefficients are expensive, so the fit distributes the influence more evenly across the correlated predictors and the coefficients become far more stable across samples.

Applying an L1 penalty of comparable strength does something qualitatively different: it drives most of the correlated predictors’ coefficients to exactly zero and keeps one. The result is sparser and easier to interpret, but which predictor survives can be somewhat arbitrary when the correlated group is nearly interchangeable.

Frequently Asked Questions

How is the penalty strength chosen?

By cross-validation. The strength cannot be fitted from the training data, because larger penalties always increase training error by construction. A grid of candidate values is evaluated by cross-validation and the one minimizing estimated test error is selected.

Should the intercept be penalized?

Normally no. The intercept represents the baseline level of the response rather than the influence of any predictor, and penalizing it would make the fit depend on where the response happens to be centred.

Why must features be standardized before regularizing?

The penalty is applied to coefficient magnitudes, and a coefficient’s magnitude depends on the units of its feature. Without standardization, a feature measured in small units gets a large coefficient and is penalized more heavily purely because of its scale.

The Bottom Line

Regularization trades a controlled amount of bias for a larger reduction in variance by making complexity costly. L2 shrinks coefficients smoothly and stabilizes correlated predictors; L1 zeroes them and selects features. In both cases the strength is a hyperparameter that only held-out data can set.