Mathematics
The mathematical machinery every learning algorithm runs on: linear algebra, calculus, and the geometry that makes high-dimensional data behave in ways intuition does not predict.
Learning paths (9)
Probability and Statistical Foundations
Reason about uncertainty precisely, then meet the central problem of learning from data: separating the error you can remove from the error you cannot.
Logic and Knowledge Representation
The other tradition in artificial intelligence: representing what a system knows as sentences that are true or false, and deriving what must follow - with guarantees a learned model cannot offer.
Search and Heuristics
The oldest working idea in artificial intelligence: describe a problem as states and actions, then let a systematic exploration find the path. Which strategy you pick decides whether the answer is optimal, and whether you run out of memory before you find it.
Unsupervised Learning
Find structure in data that has no response to predict, and face the consequence squarely: with no y there is no held-out error, so every choice you make has to be defended some other way.
Support Vector Machines
Classify by choosing the widest slab that separates two classes, then relax it so a few points may sit inside, and finally bend it without ever building the space it is bent in.
Moving Beyond Linearity
Keep least squares and change what you regress on: fixed basis functions buy curvature, constraints buy smoothness, and a penalty buys a curve that chooses its own flexibility.
Optimization for Learning
Every model on this site is fitted by the same loop: look at the slope, take a step. What decides whether that loop converges in forty steps or diverges in three is not the model - it is curvature, noise, and the size of the step. All three are measurable before the first epoch runs.
Statistical Learning Theory
Why fitting a sample tells you anything about the world it was drawn from, what the capacity of a model class really measures, and the theorem that says no method is best everywhere - together with what that theorem does not say.
Information Theory
The one place in this subject where a bound is met exactly: entropy is the shortest any code can be, the best code reaches it, and the surcharge for using the wrong distribution is the loss function you already train with.
Encyclopedia (16)
Gradient Descent
An iterative optimization algorithm that minimizes a function by repeatedly stepping in the direction opposite its gradient.
Stochastic Gradient Descent
Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.
Condition Number
The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.
Backpropagation
The algorithm that computes the gradient of a neural network’s loss with respect to every weight, by applying the chain rule backwards through the network.
Spline
A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.
Support Vector Machine
A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.
Principal Component Analysis
A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.
VC Dimension
The size of the largest set of points a family of classifiers can label in every possible way. It measures capacity by what a class can do rather than by how many members it has, which is what makes it usable for infinite families.
PAC Learning
A definition of learnability in which an algorithm must return, with high probability, a hypothesis whose true error is within a chosen tolerance - using a number of samples that is bounded in advance rather than discovered afterwards.
Kullback-Leibler Divergence
The number of extra bits per symbol paid for describing one distribution with a code built for another. It is zero only when the two agree, it is never negative, and it is not symmetric, so it is a cost rather than a distance.
Mutual Information
How many bits observing one variable tells you about another. It is zero exactly when the two are independent, it catches dependence of any shape rather than linear dependence only, and nothing computed downstream can increase it.
Kalman Filter
The exact filtering algorithm for a continuous state that moves linearly with Gaussian noise and is measured linearly with Gaussian noise, carrying the whole belief as a mean and a variance.
Softmax
A function that turns a vector of real scores into a probability distribution by exponentiating each score and dividing by the total, preserving their order while making them positive and summing to one.
Conjugate Prior
A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.
Bellman Equation
The self-consistency condition that the utility of a state equals its immediate reward plus the discounted value of the best action available from it, averaged over the outcomes that action cannot control.
Arc Consistency
A property of a constraint problem in which every value in every variable’s domain has at least one supporting value in each neighbouring domain, and the algorithm that enforces it by deleting the values that do not.
Articles (17)
Which Wrong Distribution Do You Want?
One bimodal target, one Gaussian, and two directions of the same divergence. Minimising KL(P||Q) puts the Gaussian across both modes with almost no mass where the target actually lives; minimising KL(Q||P) puts it on one mode at a value of 0.6931 nats, which is ln 2 to four decimals and not a coincidence. Each fit is judged catastrophic by the other objective, 2.0976 against 15.2799.
The Two Features That Look Like Noise
A variable that determines another with a correlation of exactly 0.0000000000, and a pair of features whose every pairwise mutual information with the target is exactly zero while the two together determine it completely. Univariate screening discards both, and the second case is the one that matters: the features it removes are removed because they matter.
The Theorem That Says Nothing About Your Problem
Averaged over all 256 functions from three bits to one, a nearest-neighbour learner and a learner built to be wrong on purpose both score exactly 0.500000 off the training set. That is the no free lunch theorem, it is exactly true, and the moment the average is restricted to the six functions that depend on a single bit the two separate to 0.333333 and 0.666667.
The Slowest Direction Sets the Pace
The step size you are allowed is fixed by the steepest direction and the number of steps you need is fixed by the flattest, so the cost of gradient descent is their ratio. The same least-squares fit, to the same ten decimal places, takes 1742 steps in one basis, 147 in a rescaled one and exactly 1 in an orthonormal one, and momentum buys back the square root of the ratio rather than the ratio.
One Parameter, Infinite Capacity
A classifier with exactly one real parameter fits all 1,048,576 labellings of twenty points, every time, and predicts a twenty-first at 0.5038 accuracy over twenty thousand trials. Counting parameters measures neither an upper nor a lower bound on what a model class can fit, which is why capacity has to be measured some other way.
A Million Clauses, or Sixty-One
Converting one short formula to conjunctive normal form by distributing gives 1,048,576 clauses and 20,971,520 literals; naming the subformulas gives 61 clauses and 160 literals, a factor of 131,072 in literals, and loses nothing at all - the two have the same number of models, checked exhaustively. The encoding is where a satisfiability problem is won or lost, not the solver.
Why Learning From Data Works At All
The gap between the error you measure and the error you will suffer, why picking the best of a thousand identical hypotheses makes it look 0.1149 better than chance, how capacity is counted for infinite model classes, and the theorem that equalises every learner - with the assumption that makes it true.
The Bound That Is Actually Reached
Entropy is not a summary of a distribution but a floor that the best code meets to the last decimal, the surcharge for using the wrong distribution is exactly the loss every classifier already minimises, and mutual information puts a hard ceiling on everything downstream of a sensor. Three results, each unusually sharp.
What Is a Neural Network?
Layers as parameterised transformations, the forward pass, and why depth and non-linearity are not optional: a proof that no single linear layer can compute XOR, and a two-layer network that does, worked entirely by hand.
Probability from Zero: The Language of Uncertainty
Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.
Entropy and Information
Measuring uncertainty in bits: Shannon entropy and why the logarithm is base 2, information gain worked on a split, and how cross-entropy and KL divergence relate to entropy and to the loss functions used to train classifiers.
Bayes' Theorem and Belief Updating
Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.
What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
The Bias-Variance Tradeoff
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.
Linear Regression from First Principles
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Backpropagation and Gradient Descent
How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.
Game Theory and Nash Equilibrium
Strategic reasoning when players are not strictly opposed: dominant strategies, the prisoner's dilemma worked from its payoff matrix, Nash equilibrium, Pareto optimality, and why equilibrium and efficiency can conflict.
Tools (1)
Datasets (1)
Research (4)
On Computable Numbers, with an Application to the Entscheidungsproblem
Introduces an abstract machine that reads and writes symbols on a tape according to a finite table of rules, and uses it to show that no general procedure can decide whether an arbitrary program halts.
A Mathematical Theory of Communication
Defines information quantitatively, introduces entropy as the measure of a source’s uncertainty, and proves limits on lossless compression and on reliable transmission over a noisy channel.
Equilibrium Points in N-Person Games
Proves that every finite game with any number of players has at least one equilibrium point, provided players may use mixed strategies.
Support-Vector Networks
Introduces the support vector machine with a soft margin, separating classes by the widest possible margin while permitting bounded violations, and using kernels to obtain non-linear boundaries.
Projects (2)
Neural Network From Scratch
A feed-forward network in NumPy with hand-derived backpropagation, validated against numerical gradients so the calculus is proven rather than trusted.
Statistical Learning Toolkit
Least squares, logistic regression, ridge and lasso, and k-fold cross-validation implemented from their estimating equations and checked against scikit-learn.