Statistics
Estimation, inference, and the discipline of saying how far a number can be trusted. The foundations under every model that claims to have learned something.
Learning paths (12)
Probability and Statistical Foundations
Reason about uncertainty precisely, then meet the central problem of learning from data: separating the error you can remove from the error you cannot.
Supervised Machine Learning
Derive the workhorse supervised methods rather than merely calling them: least squares, logistic regression, shrinkage penalties, and tree ensembles.
Unsupervised Learning
Find structure in data that has no response to predict, and face the consequence squarely: with no y there is no held-out error, so every choice you make has to be defended some other way.
Moving Beyond Linearity
Keep least squares and change what you regress on: fixed basis functions buy curvature, constraints buy smoothness, and a penalty buys a curve that chooses its own flexibility.
Learning Probabilistic Models
When the data are complete, learning a probability model is counting - the derivative of the log likelihood does the rest. When variables are hidden there is nothing to count, and the repair is to guess the counts, refit, and repeat until the likelihood stops rising.
Classification Methods Compared
There is one classifier no method can beat, and it needs the answer to build. Everything else - nearest neighbours, discriminant analysis, logistic regression - is a different guess at what it would have done, and the guesses fail in different directions.
Statistical Inference
What a sample can and cannot tell you about the population behind it: how an estimator misses, what a confidence interval actually promises, and what a p-value is - together with the three places where each of those is routinely read as something stronger than it is.
Causal Inference
Why a comparison between the treated and the untreated can carry the wrong sign, what randomisation actually buys, and the rule that says which variables to adjust for - including the ones that make the answer worse.
Time Series
What breaks when observations are not independent: a regression that finds a relationship between two series that have nothing to do with each other, standard errors that are wrong by a known factor, and a validation split that reports a model more than five times better than it is.
Experimentation and A/B Testing
What an online experiment reports when it is too small, watched too often, or read across too many metrics: an effect inflated 2.4 times, a false positive rate of 19% instead of 5%, and a winning segment in almost half of all experiments where nothing happened.
Recommender Systems
Two fitted offsets that deliver two thirds of the accuracy gain before any latent factor is learned, a model 1.28 times worse for the users who have told it least, and the blind spot that opens when a system only ever sees ratings for what it chose to show.
Anomaly Detection
A detector that never fires scores 99.5% accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when the anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.
Encyclopedia (28)
Overfitting
When a model learns the noise and idiosyncrasies of its training data rather than the underlying pattern, so it performs well in training and poorly on new data.
Bias-Variance Trade-off
The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.
Cross-Validation
A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.
Regularization
Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.
Bayes’ Theorem
A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.
Maximum Likelihood Estimation
A method of fitting a model by choosing the parameter values that make the observed data most probable.
Linear Regression
A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.
Logistic Regression
A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.
Bagging and Random Forests
Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.
Spline
A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.
k-Means Clustering
An unsupervised algorithm that partitions observations into k groups by alternately assigning points to the nearest centroid and recomputing the centroids.
Expectation–Maximization
An iterative method for maximum-likelihood estimation when some variables are unobserved: it computes the posterior distribution over the hidden variables under the current parameters, then refits the parameters as though those expected counts had been observed.
k-Nearest Neighbours
A nonparametric classifier that predicts the class of a point by taking a majority vote among the k training observations closest to it.
Linear Discriminant Analysis
A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.
ROC Curve
A plot of a classifier’s true positive rate against its false positive rate as the decision threshold is swept across its whole range, summarising every available trade between the two kinds of error.
Hierarchical Clustering
An unsupervised method that builds a tree of nested clusters by repeatedly fusing the two least dissimilar groups, so that cutting the tree at any height yields a clustering.
Principal Component Analysis
A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.
p-value
The probability of observing data at least as extreme as the data in hand, computed under the assumption that the null hypothesis is true. It measures how unusual the sample would be in a world where the effect is absent, and nothing else.
Confidence Interval
A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.
Statistical Power
The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.
Confounding
A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.
Causal Graph
A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.
Stationarity
A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.
Autocorrelation
The correlation of a series with a lagged copy of itself, measuring how long the influence of an observation persists. It is the structure that makes time-series data informative and the reason ordinary standard errors do not apply to it.
Multiple Comparisons
The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.
Matrix Factorisation
A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.
Precision and Recall
Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.
Anomaly Detection
Finding the few observations that were not produced by the process that produced the rest. The defining difficulty is not the algorithm but the base rate: at 0.5% anomalies, a detector that never fires is 99.5% accurate, and most standard metrics inherit that number rather than measuring skill.
Articles (21)
The Direction That Changes When You Change Units
Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.
The Control Variable That Invents a Relationship
Two independent causes and one common effect. Adjust for the effect and the causes acquire a correlation of exactly -1: a regression of A on B recovers a coefficient of +0.0030, and adding the common effect as a control turns it into -1.0000. Selecting a sample does the same thing invisibly, which is why "control for everything you measured" is not a defensible rule.
The 95% Interval That Covers 81% of the Time
The textbook confidence interval for a proportion has exact coverage you can compute by summing over the n+1 possible samples, and at n = 30 with p = 0.10 it is 0.8085 rather than 0.95. Coverage does not improve monotonically with n, and in a rare-event setting it can fall to 0.0392. Two one-line alternatives fix it.
A Score That Loses to Doing Nothing
A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.
The Detector That Never Fires Is 99.5% Accurate
At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.
The Treatment That Helps Everyone and Harms the Average
A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.
The Regression That Finds a Relationship That Is Not There
Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.
The Model That Picks Its Own Training Data
Two fitted offsets deliver 66% of a recommender’s gain in accuracy before any latent factor is learned, the error is 1.28 times worse for the users who have said least, only 30% of the catalogue reaches anyone’s top ten with no explicit popularity term in the model, and after six rounds of self-selected data the system is 1.14 times worse exactly where it stopped looking.
The Experiment That Was Going to Win Anyway
A test with 2,000 users per arm reports effects 2.4 times too large. An A/A test checked ten times comes out significant 19% of the time. Twenty independent null metrics produce a winner 64% of the time and twelve null segments 46%. Four numbers, one cause, and the decisions that have to be made before the data arrives.
What a Sample Can and Cannot Tell You
Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.
Learning the Numbers in a Probability Model
Where the numbers in a Bayesian network or a Gaussian actually come from: the three-step maximum-likelihood recipe worked through on discrete and continuous parameters, the Beta prior that repairs what it does to an unseen event, naive Bayes and the single zero count that destroys it, and the EM algorithm for the case where the counts cannot be taken at all - with every figure computed rather than asserted.
Comparing Classifiers, and What Accuracy Hides
The Bayes classifier nothing can beat and the error floor it leaves behind, k-nearest-neighbours as a nonparametric imitation with k as the flexibility dial, discriminant analysis and why a shared covariance forces a straight line, and the confusion matrix, thresholds and ROC curve that a single accuracy figure conceals - every number computed on simulated data where the optimum is known.
Moving Beyond Linearity: Splines and Additive Models
How to fit curved relationships without leaving least squares: basis functions, the constraints that turn a broken piecewise polynomial into a spline, the single extra column per knot that enforces them for free, and the roughness penalty that lets a curve choose its own flexibility.
Unsupervised Learning: Structure Without Labels
What changes when there is no response to predict: principal components as the direction of maximum variance, K-means and the local optima it settles into, hierarchical clustering and the linkage that decides the answer - and why none of the required choices can be validated the way a classifier can.
What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
The Bias-Variance Tradeoff
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.
Cross-Validation and Resampling
Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.
Linear Regression from First Principles
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Logistic Regression and Classification
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Regularization: Ridge and Lasso
Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.
Decision Trees and Ensembles
How recursive binary splitting builds a tree, why the Gini index beats accuracy as a splitting criterion, and how bagging and random forests turn a high-variance learner into a strong one, with the split arithmetic worked out.
Tools (7)
Base-rate calculator
Enter a prevalence, a sensitivity, and a false-positive rate to see what a positive test result is actually worth - as a probability and as whole people out of ten thousand.
Classifier metrics explorer
Move a decision threshold across an imbalanced population and watch precision, recall, F1, and the ROC point move with it - including the regime where accuracy looks excellent and the model is useless.
Confidence-interval simulator
Draw thirty samples, form a confidence interval from each, and count how many cover the truth - the clearest way to see that the confidence level describes the procedure, not any single interval.
Bias-variance explorer
Sweep model complexity and sample size to watch training error fall monotonically while test error turns upward, decomposed into bias, variance, and irreducible noise.
Distribution explorer
Change the parameters of the distributions that keep appearing in the material and watch the shape, mean, and spread respond.
Maximum-likelihood explorer
Move a parameter along the log-likelihood curve of a fixed sample and watch the fitted distribution track it, with the peak sitting exactly at the maximum-likelihood estimate.
Study planner
Pick a goal - foundations, supervised methods, or modern AI - and get this site’s own ordered plan for reaching it: which learning path to take at each stage, what to read alongside it, and how long the lessons run.
Datasets (3)
Auto fuel economy
Fuel consumption and engine characteristics for a few hundred cars - the worked example behind simple and polynomial regression.
Credit card default (simulated)
A simulated set of cardholders used to introduce classification - and a clean illustration of why accuracy misleads on rare events.
S&P 500 daily movements
Daily percentage changes in the S&P 500 from 2001 to 2005 - a deliberately hard classification problem where the honest answer is "barely better than chance".
Research (2)
Classification and Regression Trees
Establishes the CART methodology: growing decision trees by recursively choosing the split that most improves node purity, then pruning back the fully grown tree using held-out data.
Bagging Predictors
Introduces bootstrap aggregating: fitting a model to many bootstrap resamples of the training data and averaging the predictions, which reduces variance without increasing bias.
Projects (2)
Statistical Learning Toolkit
Least squares, logistic regression, ridge and lasso, and k-fold cross-validation implemented from their estimating equations and checked against scikit-learn.
Bayesian Inference Playground
Prior to posterior updating with conjugate families, alongside an entropy and KL-divergence explorer that measures what each observation actually told you.