Skip to content
Kudos AI
Lire en français
Supervised Learning

Regularization: Ridge and Lasso

Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.

7 min readKudos AI

Prerequisites: Linear Regression from First Principles, The Bias-Variance Tradeoff

Ridge asymptoting toward zero while lasso crosses it and is held there, and then the same fact drawn as a disc without corners against a diamond with them.

Least squares has one objective: make RSS as small as possible on the training data. The Bias-Variance Tradeoff showed why that is not the same as being right, and Logistic Regression showed coefficients running away to ±∞\pm\infty when nothing restrains them.

Regularization adds a second term to the objective that makes large coefficients expensive. The result is deliberately worse on the training data and often much better on new data.

A. Ridge regression

Ridge minimises RSS plus a penalty proportional to the sum of squared coefficients:

∑i=1n(yi−β0−∑j=1pβjxij)2+λ∑j=1pβj2  =  RSS+λ∑j=1pβj2.\sum_{i=1}^{n}\Big(y_i - \beta_0 - \sum_{j=1}^{p}\beta_j x_{ij}\Big)^2 + \lambda\sum_{j=1}^{p}\beta_j^2 \;=\; \mathrm{RSS} + \lambda\sum_{j=1}^{p}\beta_j^2 .

The tuning parameter λ≥0\lambda \ge 0 controls the trade. At λ=0\lambda = 0 the penalty vanishes and we recover least squares. As λ→∞\lambda \to \infty the penalty dominates and every coefficient is driven towards zero.

Two details matter. The intercept β0\beta_0 is not penalised - it only sets the overall level and shrinking it would be meaningless. And because the penalty acts on coefficient magnitudes, predictors must be standardised first: otherwise measuring a variable in grams rather than kilograms would change how hard it is penalised, which is clearly wrong.

B. Lasso

The lasso replaces the squared penalty with an absolute-value penalty:

RSS+λ∑j=1p∣βj∣.\mathrm{RSS} + \lambda\sum_{j=1}^{p}\big|\beta_j\big| .

Ridge uses an ℓ2\ell_2 penalty, the lasso an ℓ1\ell_1 penalty. That single change has a consequence out of proportion to its size: the ℓ1\ell_1 penalty forces some coefficients to be exactly zero once λ\lambda is large enough. Ridge shrinks coefficients towards zero but, except in the limit, never to zero.

A model with exact zeros ignores those predictors entirely, so the lasso performs variable selection and yields sparse models - generally far easier to interpret than a ridge fit that keeps all pp predictors with small coefficients.

C. Watching them shrink

Using the five-observation dataset from Linear Regression from First Principles (x=1..5x = 1..5, y=2,4,5,4,5y = 2,4,5,4,5), whose least-squares fit was y^=2.2+0.6x\hat y = 2.2 + 0.6x:

λ\lambdaridge sloperidge interceptlasso slopelasso intercept
00.60002.20000.60002.2000
0.10.59412.21780.59502.2150
10.54552.36360.55002.3500
50.40002.80000.35002.9500
100.30003.10000.10003.7000
120.27273.18180.00004.0000
200.20003.40000.00004.0000

Both methods pull the slope down as λ\lambda grows, and the intercept rises to compensate - it is unpenalised, so it absorbs the level.

The difference in how they approach zero is the whole point. Ridge divides: its slope is 6/(10+λ)6/(10 + \lambda), so at λ=20\lambda = 20 it is 0.20.2, a third of the original, but alive. The lasso subtracts: its slope is (6−λ/2)/10(6 - \lambda/2)/10, which hits exactly 0.00.0 at λ=12\lambda = 12 and stays there. From then on the model is just y^=4.0=yˉ\hat y = 4.0 = \bar y - the predictor has been switched off completely.

These numbers fit the raw xx, not a standardised one. With a single predictor that only changes which λ\lambda produces each row, not the shape of either path.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it prints every row of the table, including lambda=12 ridge 0.2727 3.1818 lasso 0.0000 4.0000. Mind the rescaling: scikit-learn's Lasso divides RSS by 2n2n, so its alpha is λ/(2n)\lambda/(2n). Passing λ\lambda straight in would fit each lasso row at a penalty ten times larger than the ridge row beside it.

A caveat on this demonstration. With one predictor and five observations there is nothing to select - sparsity is only interesting when some predictors are genuinely irrelevant and the lasso can switch those off while keeping the rest. The table shows the mechanism honestly, not the setting where it pays.

Interactive: what the two penalties do to a coefficient

Eight predictors, three of them genuinely zero.

Ridge (L2)
0λ →
Lasso (L1)
0λ →
Lasso: switched off
3 of 8
of which genuine noise
3 / 3
Ridge: switched off
0 of 8

At this penalty the lasso has set 3 of 8 coefficients to exactly zero, and 3 of the 3 predictors that are genuinely noise are among them. Ridge has set none to zero, and never will at any penalty: its paths approach the axis asymptotically. That difference is the whole of variable selection. Note also which order the lasso switches things off in - the weakest predictors go first, which is the useful behaviour and also the dangerous one when two predictors are correlated and it keeps an arbitrary one of them.

D. Why the penalties behave differently

The geometric account is the clearest. Both problems can be written as: minimise RSS subject to a budget on coefficient size -

∑jβj2≤s(ridge),∑j∣βj∣≤s(lasso).\sum_j \beta_j^2 \le s \quad\text{(ridge)}, \qquad \sum_j |\beta_j| \le s \quad\text{(lasso)} .

In two dimensions the ridge budget is a circle and the lasso budget a diamond with corners on the axes. The solution is where the elliptical contours of RSS first touch the budget region.

A circle has no corners, so the touch point almost never lands exactly on an axis - both coefficients stay non-zero. The diamond's corners sit on the axes, and expanding contours hit a corner readily. A corner means one coefficient is exactly zero. In higher dimensions the diamond has ever more corners, edges, and faces, and hitting one zeroes out several coefficients at once.

E. Why this reduces error

Regularization is the bias-variance tradeoff applied deliberately. Shrinking coefficients away from their least-squares values introduces bias: the estimator is no longer centred on the truth. But it also makes the fit far less sensitive to which particular sample you drew, which reduces variance.

When variance falls faster than squared bias rises, total expected error falls. That is most likely when least squares is unstable: many predictors relative to observations, or strongly correlated predictors. In the extreme case p>np > n, least squares has no unique solution at all. Ridge still has exactly one; the lasso still has solutions, though they need not be unique (two identical columns can split their weight in many ways at the same cost).

The bias grows monotonically with λ\lambda and the variance falls monotonically, so there is an interior optimum - and it cannot be found from training error, which always prefers λ=0\lambda = 0. It is chosen by cross-validation over a grid of λ\lambda values.

F. Choosing between them

Ridge (ℓ2\ell_2)Lasso (ℓ1\ell_1)
CoefficientsShrunk, all retainedSome exactly zero
Variable selectionNoYes
InterpretabilityAll pp predictorsSparse subset
Best whenMany predictors matter a littleFew predictors matter a lot
Correlated predictorsShares weight among themTends to pick one arbitrarily

Neither dominates. It is an empirical question about the problem, settled by cross-validating both. The last row is a practical gotcha: with a group of correlated predictors the lasso's arbitrary choice can be unstable across resamples, which is why elastic net - combining both penalties - exists.

Key takeaways

  • Regularization adds a penalty on coefficient size to RSS, controlled by λ\lambda, trading training fit for stability.
  • Ridge (ℓ2\ell_2) shrinks coefficients towards zero; lasso (ℓ1\ell_1) sets some exactly to zero, performing variable selection.
  • Predictors must be standardised, and the intercept is never penalised.
  • The difference is geometric: a circular budget has no corners, a diamond has corners on the axes.
  • It works by trading a little bias for a large variance reduction; λ\lambda must be chosen by cross-validation, never by training error.

What's next

Every model so far has been a single global formula. A different strategy is to carve the predictor space into regions and predict a constant in each - which handles interactions and non-linearity without anyone specifying them in advance. That is Decision Trees and Ensembles.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatisticsMachine LearningMathematics
7 min readSupervised Learning

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

StatisticsMachine LearningOptimization
6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
← Back to all articles