Skip to content
Kudos AI
Lire en français
Supervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

7 min readKudos AI

Prerequisites: What Is Statistical Learning?

The residual squares drawn literally, shrinking as the line is tilted into place, then the five residuals summed on screen to exactly zero.

Linear regression is old, simple, and still the right first thing to try. It is also the cleanest place to see the pattern every other supervised method follows: write down a measure of how badly the model fits, then choose the parameters that minimise it. Here that minimisation can be done exactly, in closed form, with calculus you already have.

A. The model

With a single predictor, we assume

Y≈β0+β1X,Y \approx \beta_0 + \beta_1 X ,

where β0\beta_0 is the intercept and β1\beta_1 the slope. Together they are the model's coefficients. Estimating them from data gives β^0\hat\beta_0 and β^1\hat\beta_1, and predictions

y^i=β^0+β^1xi.\hat y_i = \hat\beta_0 + \hat\beta_1 x_i .

The ii-th residual is ei=yi−y^ie_i = y_i - \hat y_i, the gap between what we saw and what we predicted.

B. What "best fit" means

We need a single number for how badly a candidate line fits. The standard choice is the residual sum of squares:

RSS=∑i=1nei2=∑i=1n(yi−β0−β1xi)2.\mathrm{RSS} = \sum_{i=1}^{n} e_i^2 = \sum_{i=1}^{n}\big(y_i - \beta_0 - \beta_1 x_i\big)^2 .

Squaring does two things: it makes positive and negative misses count the same, and it penalises one large miss more heavily than several small ones. It is also differentiable everywhere, which is what makes the closed-form solution below possible.

Least squares chooses β^0,β^1\hat\beta_0, \hat\beta_1 to minimise RSS.

C. Deriving the coefficients

RSS is a smooth function of two variables, so its minimum is where both partial derivatives vanish.

With respect to the intercept:

∂ RSS∂β0=∑i=1n2(yi−β0−β1xi)(−1)=0  ⟹  ∑i=1n(yi−β0−β1xi)=0.\frac{\partial\,\mathrm{RSS}}{\partial \beta_0} = \sum_{i=1}^n 2\big(y_i - \beta_0 - \beta_1 x_i\big)(-1) = 0 \;\Longrightarrow\; \sum_{i=1}^n \big(y_i - \beta_0 - \beta_1 x_i\big) = 0 .

Dividing by nn gives yˉ−β0−β1xˉ=0\bar y - \beta_0 - \beta_1 \bar x = 0, so

β^0=yˉ−β^1xˉ.\hat\beta_0 = \bar y - \hat\beta_1 \bar x .

Two things follow immediately: the residuals sum to zero, and the fitted line always passes through (xˉ,yˉ)(\bar x, \bar y).

With respect to the slope:

∂ RSS∂β1=∑i=1n2(yi−β0−β1xi)(−xi)=0.\frac{\partial\,\mathrm{RSS}}{\partial \beta_1} = \sum_{i=1}^n 2\big(y_i - \beta_0 - \beta_1 x_i\big)(-x_i) = 0 .

Substituting β0=yˉ−β1xˉ\beta_0 = \bar y - \beta_1\bar x and rearranging yields

β^1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2.\hat\beta_1 = \frac{\sum_{i=1}^n (x_i - \bar x)(y_i - \bar y)}{\sum_{i=1}^n (x_i - \bar x)^2} .

The numerator is the sample covariance of xx and yy (up to a constant) and the denominator the sample variance of xx. So the slope is how much xx and yy move together, scaled by how much xx moves on its own - which is exactly what a slope ought to be.

D. A complete worked fit

Five observations:

iixix_iyiy_i
112
224
335
444
555

Step 1 - the means.

xˉ=1+2+3+4+55=3,yˉ=2+4+5+4+55=4.\bar x = \frac{1+2+3+4+5}{5} = 3, \qquad \bar y = \frac{2+4+5+4+5}{5} = 4 .

Step 2 - the deviation products.

xi−xˉx_i - \bar xyi−yˉy_i - \bar yproduct(xi−xˉ)2(x_i - \bar x)^2
−2-2−2-24444
−1-1000011
00110000
11000011
22112244
sum6\mathbf{6}10\mathbf{10}

Step 3 - the coefficients.

β^1=610=0.6=35,β^0=4−0.6×3=4−1.8=2.2=115.\hat\beta_1 = \frac{6}{10} = 0.6 = \tfrac{3}{5}, \qquad \hat\beta_0 = 4 - 0.6 \times 3 = 4 - 1.8 = 2.2 = \tfrac{11}{5} .

The fitted line is

y^=2.2+0.6 x.\hat y = 2.2 + 0.6\,x .

Step 4 - fitted values and residuals.

xix_iy^i=2.2+0.6xi\hat y_i = 2.2 + 0.6x_iyiy_ieie_iei2e_i^2
12.82−0.8-0.80.64
23.440.60.60.36
34.051.01.01.00
44.64−0.6-0.60.36
55.25−0.2-0.20.04
0.02.40

The residuals sum to exactly zero, as the derivation promised, and the line passes through (3,4)=(xˉ,yˉ)(3, 4) = (\bar x, \bar y) - check row 3.

RSS=2.40.\mathrm{RSS} = 2.40 .

E. How good is the fit?

RSS alone is uninterpretable: it depends on the units and the sample size. Compare it instead against the total sum of squares, the error of the best possible constant prediction yˉ\bar y:

TSS=∑(yi−yˉ)2=4+0+1+0+1=6.\mathrm{TSS} = \sum (y_i - \bar y)^2 = 4 + 0 + 1 + 0 + 1 = 6 .

Then

R2=1−RSSTSS=1−2.46=0.6.R^2 = 1 - \frac{\mathrm{RSS}}{\mathrm{TSS}} = 1 - \frac{2.4}{6} = 0.6 .

The predictor explains 60% of the variability in yy; the remaining 40% it does not. R2R^2 lies in [0,1][0,1] for a model fitted with an intercept, and is unitless, which is what makes it comparable across problems.

R2R^2 never decreases when you add a predictor, even a column of pure noise, because the extra freedom can only reduce RSS. It therefore cannot be used to choose between models of different sizes - that is what cross-validation is for.

Interactive: the squares that least squares minimises

RSS is the total area of the squares.

012345670123456(x̄, ȳ)this lineflat mean4.056
RSS
4.05
TSS
6
R²
0.325
Sum of residuals
-0.5

The squares cover RSS = 4.05, so R² is 0.325. Tilt and lift the line to shrink the total area, and watch the stacked bar fall toward its floor of 2.40, or snap straight to it.

F. Verifying every number

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Running it prints the table exactly - slope 0.60.6, intercept 2.22.2, residuals summing to 00 (to machine precision), RSS=2.4\mathrm{RSS} = 2.4, TSS=6.0\mathrm{TSS} = 6.0, R2=0.6R^2 = 0.6 - and sklearn: 2.2 0.6 0.6, an independent confirmation.

G. What the model assumes

Least squares always returns an answer. Whether it is meaningful depends on assumptions worth stating plainly:

  • Linearity. The relationship really is a straight line. If it curves, the fit is biased no matter how much data you collect - the failure mode from The Bias-Variance Tradeoff.
  • Independent errors. Correlated residuals (time series, clustered data) leave coefficients usable but make standard errors far too small.
  • Constant error variance. If spread grows with xx, the fit over-weights the noisy region.
  • Influential points. With n=5n = 5, a single outlier moves the line substantially - as LOOCV fold 1 showed, dropping (1,2)(1,2) moved the prediction there from 2.82.8 to 4.04.0, a shift of 1.21.2.

Plotting residuals against fitted values checks the middle two faster than any formal test.

Key takeaways

  • Least squares minimises RSS, and setting both partial derivatives to zero gives closed-form coefficients - no iteration required.
  • β^1\hat\beta_1 is the covariance of xx and yy over the variance of xx; β^0\hat\beta_0 forces the line through (xˉ,yˉ)(\bar x, \bar y).
  • The derivation guarantees the residuals sum to zero.
  • Our worked fit: y^=2.2+0.6x\hat y = 2.2 + 0.6x, RSS=2.4\mathrm{RSS} = 2.4, TSS=6\mathrm{TSS} = 6, R2=0.6R^2 = 0.6.
  • R2R^2 never decreases when predictors are added, so it cannot compare models of different sizes.

What's next

Squared error is the wrong target when the response is a category rather than a number - a fitted line will happily predict a probability of 1.41.4. Adapting the linear model to classification is Logistic Regression and Classification.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readStatistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

StatisticsMachine LearningMathematics
7 min readStatistical Learning Foundations

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

StatisticsMachine LearningMathematics
7 min readSupervised Learning

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

StatisticsMachine LearningOptimization
← Back to all articles