Skip to content
Kudos AI
Lire en français
Statistical Learning Foundations

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

7 min readKudos AI

Prerequisites: What Is Statistical Learning?

Squared bias falling and variance rising drawn separately, then added into the familiar U - and the LOOCV training sets shown overlapping almost completely.

What Is Statistical Learning? showed a degree-15 polynomial fitting its training sample better than any competitor while being catastrophically wrong about the underlying function. That was a demonstration. This article gives the account: the expected test error splits into exactly three pieces, two of which we control and one of which we do not, and the two we control move in opposite directions.

A. The decomposition

Take a test point x0x_0, and let f^\hat f be a model fitted to a random training sample. The expected squared error at x0x_0, averaged over all the training samples we might have drawn, decomposes as

E[(y0−f^(x0))2]=[Bias⁡(f^(x0))]2⏟wrong on average+Var⁡(f^(x0))⏟unstable+Var⁡(ε)⏟irreducible,\mathbb{E}\big[(y_0 - \hat f(x_0))^2\big] = \underbrace{\big[\operatorname{Bias}(\hat f(x_0))\big]^2}_{\text{wrong on average}} + \underbrace{\operatorname{Var}(\hat f(x_0))}_{\text{unstable}} + \underbrace{\operatorname{Var}(\varepsilon)}_{\text{irreducible}} ,

where the bias is Bias⁡(f^(x0))=E[f^(x0)]−f(x0)\operatorname{Bias}(\hat f(x_0)) = \mathbb{E}[\hat f(x_0)] - f(x_0).

This is an equality, not an approximation or a heuristic. Each term means something concrete:

  • Bias is the error from approximating a complicated real problem with a simpler model. A straight line fitted to a curve is biased everywhere, and no amount of extra data removes that.
  • Variance is how much f^\hat f would change if you refitted it on a different training sample of the same size. A method that swings wildly from sample to sample has high variance, and any single fit is unreliable.
  • Irreducible error is Var⁡(ε)\operatorname{Var}(\varepsilon), the floor from the previous article.

Since squared bias and variance are both non-negative, their sum is non-negative, so expected test error can never fall below Var⁡(ε)\operatorname{Var}(\varepsilon). The floor is real.

B. Why the two terms fight

The general pattern: as a method becomes more flexible, bias falls and variance rises.

More flexibility means the model can bend towards the true ff, so it is less systematically wrong - bias down. But it also means the model can bend towards the particular noise in this sample, so a different sample would produce a noticeably different fit - variance up.

Test error is the sum. Early on, flexibility buys a large bias reduction for a small variance increase and total error falls. Past some point the trade reverses and total error climbs. The best model sits where the two rates of change cancel.

The figure draws schematic curves in arbitrary units, not the sin⁡(2πx)\sin(2\pi x) simulation below; its noise floor sits at 0.16, not the 0.09 used there.

Interactive: the bias-variance tradeoff

Training vs test error as complexity grows.

0.00.51.0irreducible noisemodel complexity →
Test errorTraining erroroptimal complexity
Regime
Well-fit
Training error
0.25
Test error
0.55

Near the sweet spot: the test error is close to its minimum.

Training error always falls as the model grows more flexible, so it is a misleading guide. Test error is bias² + variance + irreducible noise: it bottoms out where the two forces balance, then rises as the model fits noise. Add data (raise the sample size) and the variance term shrinks, pushing the sweet spot to higher complexity and lowering the whole test curve.

C. Measuring all three terms

The decomposition is usually presented as theory because in practice you cannot compute it - you do not know ff and you have only one training sample. But in a simulation we control both, so we can measure each term separately and check they really do add up.

The setup: f(x)=sin⁡(2πx)f(x) = \sin(2\pi x), noise ε∼N(0,0.32)\varepsilon \sim N(0, 0.3^2) so Var⁡(ε)=0.09\operatorname{Var}(\varepsilon) = 0.09, training samples of 30 points, and a single test point x0=0.35x_0 = 0.35. We refit 2,000 times on fresh samples and watch what the predictions at x0x_0 do.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it prints:

true f(x0) = 0.8090,  Var(eps) = 0.0900

degree 1:  bias^2 0.2684   variance 0.0130   + noise 0.0900   = 0.3713
degree 3:  bias^2 0.0059   variance 0.0105   + noise 0.0900   = 0.1064
degree 9:  bias^2 0.0000   variance 0.0295   + noise 0.0900   = 0.1195

D. Reading the table

Every claim in section B is visible here as a number.

Bias collapses as flexibility grows. Degree 1 has squared bias 0.26840.2684 - a straight line simply cannot pass near sin⁡(2πx)\sin(2\pi x) at x0=0.35x_0 = 0.35, and its average prediction over 2,000 fits was 0.2910.291 against a true value of 0.8090.809. By degree 3 the squared bias is 0.00590.0059, and by degree 9 it rounds to zero: on average, the flexible model is exactly right.

Variance grows as flexibility grows. It more than doubles from 0.01300.0130 at degree 1 to 0.02950.0295 at degree 9. The degree-9 fit is right on average but any individual fit is noticeably off, in a direction that depends on which 30 points it happened to see.

The sum is minimised in the middle. Total error runs 0.3713→0.1064→0.11950.3713 \to 0.1064 \to 0.1195. Degree 3 wins - not because it is best on either term individually (degree 9 has lower bias) but because it is the best compromise.

The floor holds. No configuration gets below 0.090.09. The best total, 0.10640.1064, is only 0.01640.0164 above the irreducible noise, and that remaining gap is the reducible error still on the table.

A caution about the degree-9 row. Squared bias printing as 0.00000.0000 does not mean the model is unbiased everywhere - only at this particular x0x_0, and only to four decimal places after averaging 2,000 fits. Bias is a function of xx; at a point near the boundary of the interval the same model would show substantial bias, for exactly the reason the degree-15 polynomial exploded in the previous article.

E. Does the decomposition actually hold?

The three terms should sum to the expected test error, which we can measure directly by generating fresh noisy observations at x0x_0 and averaging the squared prediction errors. Doing that alongside the decomposition gives:

degreebias² + var + noisemeasured test MSE
10.37130.3764
30.10640.1072
90.11950.1171

The two columns agree to within a few thousandths - the residual difference is Monte Carlo error from using 2,000 simulations rather than infinitely many. The identity is exact; our estimate of it is merely very good.

F. What this means in practice

You cannot compute bias and variance on real data, because you have one sample and no access to ff. What you can do is recognise their signatures:

SymptomLikely causeResponse
High training error, similar test errorHigh bias (underfitting)More flexible model, better features
Very low training error, much higher test errorHigh variance (overfitting)Regularise, simplify, more data
Both errors near the noise floorNear the optimumStop

Note the asymmetry in the fixes: more data reduces variance but does nothing for bias. Doubling the sample will not help a straight line fit a sine wave. If your model is biased, you need a different model, not a bigger dataset.

Key takeaways

  • Expected test error decomposes exactly into squared bias, variance, and irreducible noise.
  • Bias falls and variance rises with flexibility; the sum is U-shaped and the best model is the compromise, not the extreme.
  • In simulation the terms are measurable: degree 3 won with total 0.10640.1064 against 0.37130.3713 (biased) and 0.11950.1195 (high-variance).
  • Test error can never go below Var⁡(ε)\operatorname{Var}(\varepsilon).
  • More data cures variance, not bias. Diagnose which you have before choosing a remedy.

What's next

The diagnostic table above needs an honest estimate of test error, and the training error emphatically is not one. Getting that estimate from the data you already have - without a separate test set you may not be able to afford - is the job of Cross-Validation and Resampling.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

4 min readStatistical Learning Theory

One Parameter, Infinite Capacity

A classifier with exactly one real parameter fits all 1,048,576 labellings of twenty points, every time, and predicts a twenty-first at 0.5038 accuracy over twenty thousand trials. Counting parameters measures neither an upper nor a lower bound on what a model class can fit, which is why capacity has to be measured some other way.

Machine LearningMathematics
7 min readStatistical Learning Foundations

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

StatisticsMachine Learning
7 min readSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatisticsMachine LearningMathematics
← Back to all articles