Skip to content
Kudos AI
Lire en français
Statistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

7 min readKudos AI

Prerequisites: Probability from Zero

The error bar descending onto a noise floor of 0.25 and stopping dead there, however far the fit is improved.

Every supervised model on this site - linear regression, trees, neural networks - is an answer to the same question, posed in the same notation. Before comparing methods it is worth stating that question exactly, because it also tells you which part of your error you can hope to remove and which part you are stuck with no matter how good your method is.

A. The setup

We observe a response YY and pp different predictors X=(X1,X2,…,Xp)X = (X_1, X_2, \dots, X_p). We assume there is some relationship between them, written in the general form

Y=f(X)+ε.Y = f(X) + \varepsilon .

Here ff is a fixed but unknown function representing the systematic information that XX provides about YY, and ε\varepsilon is a random error term, independent of XX, with mean zero.

Statistical learning is the set of approaches for estimating ff.

The estimate is written f^\hat f, and the prediction it produces is Y^=f^(X)\hat Y = \hat f(X).

B. Reducible and irreducible error

Suppose for a moment that f^\hat f and XX are fixed. How wrong is Y^\hat Y as a prediction of YY? Following James et al., the expected squared error decomposes into two conceptually distinct pieces:

E[(Y−Y^)2]=[f(X)−f^(X)]2⏟reducible+Var⁡(ε)⏟irreducible.\mathbb{E}\big[(Y - \hat Y)^2\big] = \underbrace{\big[f(X) - \hat f(X)\big]^2}_{\text{reducible}} + \underbrace{\operatorname{Var}(\varepsilon)}_{\text{irreducible}} .

The first term is reducible: f^\hat f is not a perfect estimate of ff, and we can shrink that gap by choosing a better method or gathering more data.

The second term is irreducible, and this is the important one. Even with a perfect estimate - even if f^=f\hat f = f exactly - the prediction would still be wrong by ε\varepsilon, because YY depends on ε\varepsilon and ε\varepsilon is by definition not predictable from XX.

Why is the irreducible error larger than zero? Two reasons, both worth internalising. First, ε\varepsilon may contain unmeasured variables that would help predict YY - since we did not measure them, no ff built on XX can use them. Second, it may contain genuinely unmeasurable variation: the same inputs on two different days simply do not produce identical outputs.

The practical consequence is a hard ceiling. Var⁡(ε)\operatorname{Var}(\varepsilon) is an upper bound on how good any model can get, and it is almost always unknown in practice. A model that appears to beat it is not a triumph; it is a sign you are measuring test error on data the model has already seen.

C. Prediction versus inference

There are two reasons to estimate ff, and they pull in opposite directions.

For prediction, you only care that Y^\hat Y is close to YY. The estimate f^\hat f can be a total black box - you will never inspect it - provided it is accurate.

For inference, you want to understand the relationship: which predictors actually matter, whether each one raises or lowers the response, whether a simple linear summary is adequate. Now f^\hat f cannot be a black box, because the shape of the model is the answer.

This tension recurs constantly:

PredictionInference
GoalAccurate Y^\hat YUnderstand ff
Model preferenceFlexible, possibly opaqueRestrictive, interpretable
Typical choiceEnsembles, neural networksLinear and generalised linear models

A model chosen purely for accuracy will often be one you cannot explain, and a model chosen for explanation will often leave accuracy on the table. Knowing which of the two you are doing is a prerequisite to choosing sensibly.

D. Parametric and non-parametric estimation

Broadly there are two strategies for producing f^\hat f.

Parametric methods reduce the problem to estimating a fixed set of numbers. You first assume a functional form - most simply, that ff is linear:

f(X)=β0+β1X1+β2X2+⋯+βpXp,f(X) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \dots + \beta_p X_p ,

and then only need to estimate the p+1p+1 coefficients rather than an arbitrary pp-dimensional function. That is an enormous simplification. The risk is equally clear: if the true ff is far from linear, no choice of coefficients will fit it, and the model is wrong in a way more data cannot fix.

Non-parametric methods make no such assumption and let the data determine the shape. They can fit a much wider range of true functions, but because they do not reduce the problem to a few parameters, they need substantially more observations to pin the shape down reliably.

E. Why flexibility is not free

It is tempting to conclude that flexible non-parametric methods are simply better. They are not, for two reasons.

First, as noted, they are hungrier for data. Second - and less obviously - a very flexible method can follow the noise ε\varepsilon in the training sample as if it were signal. The fit to the data you have looks superb while the performance on data you have not seen degrades.

The figure draws schematic curves, not the polynomial fits below: as you raise Model complexity, training error falls steadily while test error turns back up.

Interactive: the bias-variance tradeoff

Training vs test error as complexity grows.

0.00.51.0irreducible noisemodel complexity →
Test errorTraining erroroptimal complexity
Regime
Well-fit
Training error
0.25
Test error
0.55

Near the sweet spot: the test error is close to its minimum.

Training error always falls as the model grows more flexible, so it is a misleading guide. Test error is bias² + variance + irreducible noise: it bottoms out where the two forces balance, then rises as the model fits noise. Add data (raise the sample size) and the variance term shrinks, pushing the sweet spot to higher complexity and lowering the whole test curve.

The following demonstrates the point on data where the truth is known, so we can compare the fit against ff itself rather than against noisy observations:

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it prints:

degree  1:  training MSE 0.1985   error vs true f 0.2465
degree  3:  training MSE 0.0542   error vs true f 0.0235
degree 15:  training MSE 0.0258   error vs true f 730.5159

Read the two columns against each other. Training MSE falls monotonically with flexibility - degree 15 fits the observed points best of all three, at 0.02580.0258. Yet its error against the true ff is 730.5730.5: not marginally worse than degree 3's 0.02350.0235, but roughly thirty thousand times worse.

The blow-up is worth understanding rather than just noting. A degree-15 polynomial fitted to 25 points has enough freedom to weave through the noise, and the wiggles it needs in order to do so become violent near the edges of the interval, where there are fewest points to constrain it. Evaluated on the dense grid, those boundary excursions dominate the average. This is a real and well-known failure mode of high-degree polynomial fits, not an artefact of the seed.

Meanwhile degree 1 shows the opposite failure: its training error is much worse than degree 15's, but it is honest about it - the error against the truth, 0.24650.2465, is about the same as its training error, because a straight line is too rigid to chase noise. It is simply the wrong shape for a sine wave.

Degree 3 wins on the only column that matters. This gap between "fits the sample" and "captures the truth" is the central phenomenon of the subject, and it has a precise decomposition.

Key takeaways

  • Statistical learning estimates an unknown ff in Y=f(X)+εY = f(X) + \varepsilon.
  • Prediction error splits into a reducible part, which better methods can shrink, and an irreducible part Var⁡(ε)\operatorname{Var}(\varepsilon), which nothing can.
  • Irreducible error exists because of unmeasured and unmeasurable variation; it is a hard ceiling on any model's accuracy.
  • Prediction favours flexible, opaque models; inference favours restrictive, interpretable ones. Decide which you are doing first.
  • Parametric methods assume a form and estimate few parameters; non-parametric methods assume less but need far more data.
  • More flexibility always improves training error, and beyond a point worsens accuracy on new data.

What's next

That last observation deserves an exact account rather than an anecdote. The reducible error itself splits into two competing pieces - one that shrinks as models get more flexible and one that grows - and their sum is what you actually pay. That is The Bias–Variance Tradeoff.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatisticsMachine LearningMathematics
7 min readProbability Foundations

Probability from Zero: The Language of Uncertainty

Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.

ProbabilityMathematicsArtificial Intelligence
6 min readProbability Foundations

Bayes' Theorem and Belief Updating

Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.

ProbabilityMathematicsArtificial Intelligence
← Back to all articles