Skip to content
Kudos AI

Machine Learning

Learning a function from data. Supervised and unsupervised methods, generalisation and overfitting, and the theory that says when learning is possible at all.

96 items

Learning paths (14)

Supervised Machine Learning

Derive the workhorse supervised methods rather than merely calling them: least squares, logistic regression, shrinkage penalties, and tree ensembles.

Unsupervised Learning

Find structure in data that has no response to predict, and face the consequence squarely: with no y there is no held-out error, so every choice you make has to be defended some other way.

Support Vector Machines

Classify by choosing the widest slab that separates two classes, then relax it so a few points may sit inside, and finally bend it without ever building the space it is bent in.

Moving Beyond Linearity

Keep least squares and change what you regress on: fixed basis functions buy curvature, constraints buy smoothness, and a penalty buys a curve that chooses its own flexibility.

Learning Probabilistic Models

When the data are complete, learning a probability model is counting - the derivative of the log likelihood does the rest. When variables are hidden there is nothing to count, and the repair is to guess the counts, refit, and repeat until the likelihood stops rising.

Classification Methods Compared

There is one classifier no method can beat, and it needs the answer to build. Everything else - nearest neighbours, discriminant analysis, logistic regression - is a different guess at what it would have done, and the guesses fail in different directions.

Optimization for Learning

Every model on this site is fitted by the same loop: look at the slope, take a step. What decides whether that loop converges in forty steps or diverges in three is not the model - it is curvature, noise, and the size of the step. All three are measurable before the first epoch runs.

Statistical Learning Theory

Why fitting a sample tells you anything about the world it was drawn from, what the capacity of a model class really measures, and the theorem that says no method is best everywhere - together with what that theorem does not say.

Causal Inference

Why a comparison between the treated and the untreated can carry the wrong sign, what randomisation actually buys, and the rule that says which variables to adjust for - including the ones that make the answer worse.

Time Series

What breaks when observations are not independent: a regression that finds a relationship between two series that have nothing to do with each other, standard errors that are wrong by a known factor, and a validation split that reports a model more than five times better than it is.

Information Theory

The one place in this subject where a bound is met exactly: entropy is the shortest any code can be, the best code reaches it, and the surcharge for using the wrong distribution is the loss function you already train with.

Experimentation and A/B Testing

What an online experiment reports when it is too small, watched too often, or read across too many metrics: an effect inflated 2.4 times, a false positive rate of 19% instead of 5%, and a winning segment in almost half of all experiments where nothing happened.

Recommender Systems

Two fitted offsets that deliver two thirds of the accuracy gain before any latent factor is learned, a model 1.28 times worse for the users who have told it least, and the blind spot that opens when a system only ever sees ratings for what it chose to show.

Anomaly Detection

A detector that never fires scores 99.5% accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when the anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Encyclopedia (41)

Gradient Descent

An iterative optimization algorithm that minimizes a function by repeatedly stepping in the direction opposite its gradient.

Stochastic Gradient Descent

Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.

Condition Number

The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.

Learning Rate Schedule

A rule that changes the step size over the course of training, large early so the run can travel and small late so it can settle.

Overfitting

When a model learns the noise and idiosyncrasies of its training data rather than the underlying pattern, so it performs well in training and poorly on new data.

Bias-Variance Trade-off

The decomposition of a model’s expected prediction error into bias, variance, and irreducible noise, and the tension that reducing one of the first two typically increases the other.

Cross-Validation

A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.

Regularization

Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.

Naive Bayes

A classifier that applies Bayes’ theorem while assuming all features are conditionally independent given the class.

Cross-Entropy

A measure of the difference between two probability distributions, used as the standard loss function for classification.

Linear Regression

A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.

Logistic Regression

A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.

Decision Tree

A model that predicts by applying a sequence of threshold tests on individual features, splitting the data into increasingly homogeneous groups.

Bagging and Random Forests

Ensemble methods that reduce variance by averaging many models fitted to bootstrap resamples, with random forests additionally decorrelating the trees by restricting the features available at each split.

Spline

A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.

Support Vector Machine

A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.

k-Means Clustering

An unsupervised algorithm that partitions observations into k groups by alternately assigning points to the nearest centroid and recomputing the centroids.

Expectation–Maximization

An iterative method for maximum-likelihood estimation when some variables are unobserved: it computes the posterior distribution over the hidden variables under the current parameters, then refits the parameters as though those expected counts had been observed.

k-Nearest Neighbours

A nonparametric classifier that predicts the class of a point by taking a majority vote among the k training observations closest to it.

Linear Discriminant Analysis

A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.

ROC Curve

A plot of a classifier’s true positive rate against its false positive rate as the decision threshold is swept across its whole range, summarising every available trade between the two kinds of error.

Hierarchical Clustering

An unsupervised method that builds a tree of nested clusters by repeatedly fusing the two least dissimilar groups, so that cutting the tree at any height yields a clustering.

Principal Component Analysis

A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.

Neural Network

A model composed of layers of simple units, each computing a weighted sum followed by a non-linear function, fitted by gradient descent using backpropagation.

Pretraining and Fine-Tuning

The two-stage recipe of first training a model on a large generic corpus, then adapting it to a specific task with a much smaller labelled dataset.

Q-Learning

A reinforcement learning algorithm that learns the value of taking each action in each state directly from experience, without a model of the environment.

VC Dimension

The size of the largest set of points a family of classifiers can label in every possible way. It measures capacity by what a class can do rather than by how many members it has, which is what makes it usable for infinite families.

PAC Learning

A definition of learnability in which an algorithm must return, with high probability, a hypothesis whose true error is within a chosen tolerance - using a number of samples that is bounded in advance rather than discovered afterwards.

Confounding

A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.

Causal Graph

A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.

Stationarity

A property of a series whose statistical behaviour does not depend on when you look at it: the mean, the variance and the correlation structure are the same in every window. Almost every classical method assumes it, and most real series lack it.

Autocorrelation

The correlation of a series with a lagged copy of itself, measuring how long the influence of an observation persists. It is the structure that makes time-series data informative and the reason ordinary standard errors do not apply to it.

Kullback-Leibler Divergence

The number of extra bits per symbol paid for describing one distribution with a code built for another. It is zero only when the two agree, it is never negative, and it is not symmetric, so it is a cost rather than a distance.

Mutual Information

How many bits observing one variable tells you about another. It is zero exactly when the two are independent, it catches dependence of any shape rather than linear dependence only, and nothing computed downstream can increase it.

Multiple Comparisons

The inflation of false positives that occurs whenever more than one test, metric, segment or stopping point is allowed to produce the headline. Each additional chance raises the probability that something crosses the threshold by luck alone.

Matrix Factorisation

A model that explains a sparse table of interactions as the product of two small matrices, giving every user and every item a short vector of learned traits whose dot product predicts the missing entries.

Softmax

A function that turns a vector of real scores into a probability distribution by exponentiating each score and dividing by the total, preserving their order while making them positive and summing to one.

Perplexity

The exponential of a model’s average cross entropy, read as the number of equally likely options it is effectively choosing between at each step.

Conjugate Prior

A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.

Precision and Recall

Two rates that split what accuracy hides: precision is the share of predicted positives that are real, and recall is the share of real positives that were found.

Anomaly Detection

Finding the few observations that were not produced by the process that produced the rest. The defining difficulty is not the algorithm but the base rate: at 0.5% anomalies, a detector that never fires is 99.5% accurate, and most standard metrics inherit that number rather than measuring skill.

Articles (27)

Which Wrong Distribution Do You Want?

One bimodal target, one Gaussian, and two directions of the same divergence. Minimising KL(P||Q) puts the Gaussian across both modes with almost no mass where the target actually lives; minimising KL(Q||P) puts it on one mode at a value of 0.6931 nats, which is ln 2 to four decimals and not a coincidence. Each fit is judged catastrophic by the other objective, 2.0976 against 15.2799.

The Two Features That Look Like Noise

A variable that determines another with a correlation of exactly 0.0000000000, and a pair of features whose every pairwise mutual information with the target is exactly zero while the two together determine it completely. Univariate screening discards both, and the second case is the one that matters: the features it removes are removed because they matter.

The Theorem That Says Nothing About Your Problem

Averaged over all 256 functions from three bits to one, a nearest-neighbour learner and a learner built to be wrong on purpose both score exactly 0.500000 off the training set. That is the no free lunch theorem, it is exactly true, and the moment the average is restricted to the six functions that depend on a single bit the two separate to 0.333333 and 0.666667.

The Slowest Direction Sets the Pace

The step size you are allowed is fixed by the steepest direction and the number of steps you need is fixed by the flattest, so the cost of gradient descent is their ratio. The same least-squares fit, to the same ten decimal places, takes 1742 steps in one basis, 147 in a rescaled one and exactly 1 in an orthonormal one, and momentum buys back the square root of the ratio rather than the ratio.

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

One Parameter, Infinite Capacity

A classifier with exactly one real parameter fits all 1,048,576 labellings of twenty points, every time, and predicts a twenty-first at 0.5038 accuracy over twenty thousand trials. Counting parameters measures neither an upper nor a lower bound on what a model class can fit, which is why capacity has to be measured some other way.

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

What Actually Makes Training Converge

A two per cent change in the learning rate separates a converged run from one five orders of magnitude away, a condition number predicts the convergence rate to six decimal places, and stochastic gradient descent with a fixed step never converges at all - it settles into a ball whose radius grows as the square root of the step. Every figure here was computed on a problem whose exact optimum is known.

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Why Learning From Data Works At All

The gap between the error you measure and the error you will suffer, why picking the best of a thousand identical hypotheses makes it look 0.1149 better than chance, how capacity is counted for infinite model classes, and the theorem that equalises every learner - with the assumption that makes it true.

The Treatment That Helps Everyone and Harms the Average

A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.

The Regression That Finds a Relationship That Is Not There

Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.

The Model That Picks Its Own Training Data

Two fitted offsets deliver 66% of a recommender’s gain in accuracy before any latent factor is learned, the error is 1.28 times worse for the users who have said least, only 30% of the catalogue reaches anyone’s top ten with no explicit popularity term in the model, and after six rounds of self-selected data the system is 1.14 times worse exactly where it stopped looking.

The Experiment That Was Going to Win Anyway

A test with 2,000 users per arm reports effects 2.4 times too large. An A/A test checked ten times comes out significant 19% of the time. Twenty independent null metrics produce a winner 64% of the time and twelve null segments 46%. Four numbers, one cause, and the decisions that have to be made before the data arrives.

The Bound That Is Actually Reached

Entropy is not a summary of a distribution but a floor that the best code meets to the last decimal, the surcharge for using the wrong distribution is exactly the loss every classifier already minimises, and mutual information puts a hard ceiling on everything downstream of a sensor. Three results, each unusually sharp.

Comparing Classifiers, and What Accuracy Hides

The Bayes classifier nothing can beat and the error floor it leaves behind, k-nearest-neighbours as a nonparametric imitation with k as the flexibility dial, discriminant analysis and why a shared covariance forces a straight line, and the confusion matrix, thresholds and ROC curve that a single accuracy figure conceals - every number computed on simulated data where the optimum is known.

Support Vector Machines: Margins and Kernels

Why the widest slab between two classes is a good boundary, why insisting on a perfect one is self-defeating, how a budget for violations buys back stability, and how a kernel bends the boundary by working in a space it never has to build.

Moving Beyond Linearity: Splines and Additive Models

How to fit curved relationships without leaving least squares: basis functions, the constraints that turn a broken piecewise polynomial into a spline, the single extra column per knot that enforces them for free, and the roughness penalty that lets a curve choose its own flexibility.

Unsupervised Learning: Structure Without Labels

What changes when there is no response to predict: principal components as the direction of maximum variance, K-means and the local optima it settles into, hierarchical clustering and the linkage that decides the answer - and why none of the required choices can be validated the way a classifier can.

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

Reinforcement Learning and Q-Learning

Learning to act well without a model of the world: temporal-difference updates, the Q-learning rule, exploration versus exploitation, and a run that recovers the planned optimum from experience alone.

The Bias-Variance Tradeoff

The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

Regularization: Ridge and Lasso

Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.

Decision Trees and Ensembles

How recursive binary splitting builds a tree, why the Gini index beats accuracy as a splitting criterion, and how bagging and random forests turn a high-variance learner into a strong one, with the split arithmetic worked out.

Tools (3)

Datasets (4)

Research (6)

Projects (1)

Related topics