Probability
The language for reasoning under uncertainty: random variables, distributions, conditioning, and the rules that turn partial information into a defensible belief.
Learning paths (7)
Probability and Statistical Foundations
Reason about uncertainty precisely, then meet the central problem of learning from data: separating the error you can remove from the error you cannot.
Reinforcement Learning
Acting well when outcomes are uncertain: the Bellman equation and how to solve it, then what changes when the environment is unknown and the agent has to learn from experience alone.
Probabilistic Reasoning with Bayesian Networks
Represent a joint distribution over many variables with a graph and a handful of small tables, then answer queries against it exactly when the structure allows and by sampling when it does not.
Probabilistic Reasoning over Time
Track a world that changes while you watch it through a noisy sensor: the two assumptions that make it tractable, the forward and backward recursions that answer every query about the past and present, and the separate algorithm needed for the most likely history.
Making Decisions under Uncertainty
Combine what you believe with what you want: expected utility as the criterion, the curve that explains why sensible people refuse favourable bets, and a price for information that is zero unless it changes your mind.
Decisions Under Partial Observability
An agent that cannot see which state it is in has to act on a distribution instead. That distribution is itself always observable, which turns the problem back into an MDP - over a continuous space, on which the exact algorithms do not close.
Statistical Inference
What a sample can and cannot tell you about the population behind it: how an estimator misses, what a confidence interval actually promises, and what a p-value is - together with the three places where each of those is routinely read as something stronger than it is.
Encyclopedia (16)
Bayes’ Theorem
A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.
Naive Bayes
A classifier that applies Bayes’ theorem while assuming all features are conditionally independent given the class.
Bayesian Network
A directed acyclic graph whose nodes are random variables and whose edges express direct influence, with a conditional probability table at each node, that together define a full joint distribution as a product of local factors.
Entropy
A measure of the uncertainty in a random variable, equal to the average number of bits needed to encode its outcome.
Maximum Likelihood Estimation
A method of fitting a model by choosing the parameter values that make the observed data most probable.
Hidden Markov Model
A temporal model in which a single discrete state variable evolves as a Markov chain and emits one observation per time step, so the state must be inferred from a noisy proxy rather than seen directly.
Belief State
The probability distribution an agent holds over the states it might be in, given everything it has done and perceived - the thing it can act on when the state itself is hidden.
Partially Observable MDP
A Markov decision process in which the agent cannot observe its state directly, only noisy percepts of it - solved in principle by treating the distribution over states as the state of an ordinary, fully observable MDP.
Markov Decision Process
A formal model of sequential decision-making in which outcomes are partly random, defined by states, actions, transition probabilities, and rewards.
Expected Utility
The probability-weighted average utility of an action’s possible outcomes, and the quantity a rational agent maximises when choosing what to do under uncertainty.
p-value
The probability of observing data at least as extreme as the data in hand, computed under the assumption that the null hypothesis is true. It measures how unusual the sample would be in a world where the effect is absent, and nothing else.
Confidence Interval
A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.
Statistical Power
The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.
Kalman Filter
The exact filtering algorithm for a continuous state that moves linearly with Gaussian noise and is measured linearly with Gaussian noise, carrying the whole belief as a mean and a variance.
Viterbi Algorithm
A dynamic programming algorithm that finds the single most likely sequence of hidden states given a sequence of observations, by carrying forward the best path to each state rather than the total probability of reaching it.
Conjugate Prior
A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.
Articles (12)
The Week That Cannot Have Happened
Take the most likely state on each day and write them down in order, and you have a report the model assigns probability exactly zero: on a four-day machine-monitoring example the day-by-day answer is healthy, healthy, failed, failed, and healthy to failed is a transition that cannot occur. What the two questions actually are, why smoothing and Viterbi answer different ones, and what the 0.411 posterior on the best path means for anyone who has to act on it.
A Hundred Thousand Samples, Four Hundred of Them Real
On the burglary network with both neighbours calling, rejection sampling keeps 183 of 100,000 draws and likelihood weighting keeps all of them at an effective sample size of 396. Both estimates are about 10% off a posterior of 0.284172, and the reason is exactly computable: 252 samples carry 76% of the weight and 99.975% of the squared weight.
What a Sample Can and Cannot Tell You
Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.
Learning the Numbers in a Probability Model
Where the numbers in a Bayesian network or a Gaussian actually come from: the three-step maximum-likelihood recipe worked through on discrete and continuous parameters, the Beta prior that repairs what it does to an unseen event, naive Bayes and the single zero count that destroys it, and the EM algorithm for the case where the counts cannot be taken at all - with every figure computed rather than asserted.
Acting When You Cannot See the State
What changes when an agent gets noisy percepts instead of its state: the belief state that replaces it and the filtering update that maintains it, the exact reduction of a POMDP to an MDP over beliefs, the piecewise-linear convex value function that makes the reduction computable in principle, and the measured reasons it is not computable in practice - with the fixed point of belief, the alpha vectors and the value function all computed rather than asserted.
Decisions under Uncertainty: Utility and Information
Why a bet with positive expected monetary value can be rational to refuse, what the curvature of a utility function measures, and how to price an observation before buying it - including the common case where the honest price is zero.
Reasoning About a Changing World
How two Markov assumptions turn an unbounded history into two small tables, the forward and backward recursions that answer every question about the present and the past, why the most likely sequence needs an algorithm of its own, and what changes when the state is a real number rather than a list.
Bayesian Networks and Probabilistic Inference
How a graph and a few small tables stand in for a joint distribution with thousands of entries, how to answer a query against it exactly by enumeration and variable elimination, and what to do when exact inference is out of reach: rejection sampling, likelihood weighting, and Gibbs sampling, each worked through on the same two networks.
Probability from Zero: The Language of Uncertainty
Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.
Entropy and Information
Measuring uncertainty in bits: Shannon entropy and why the logarithm is base 2, information gain worked on a split, and how cross-entropy and KL divergence relate to entropy and to the loss functions used to train classifiers.
Bayes' Theorem and Belief Updating
Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.
Markov Decision Processes
How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.
Tools (5)
Base-rate calculator
Enter a prevalence, a sensitivity, and a false-positive rate to see what a positive test result is actually worth - as a probability and as whole people out of ten thousand.
Entropy and information calculator
Edit two distributions side by side and read off entropy, cross entropy, and KL divergence in bits, with the identity H(p, q) = H(p) + KL(p ‖ q) held on screen.
Confidence-interval simulator
Draw thirty samples, form a confidence interval from each, and count how many cover the truth - the clearest way to see that the confidence level describes the procedure, not any single interval.
Distribution explorer
Change the parameters of the distributions that keep appearing in the material and watch the shape, mean, and spread respond.
Maximum-likelihood explorer
Move a parameter along the log-likelihood curve of a fixed sample and watch the fitted distribution track it, with the peak sitting exactly at the maximum-likelihood estimate.
Datasets (2)
Credit card default (simulated)
A simulated set of cardholders used to introduce classification - and a clean illustration of why accuracy misleads on rare events.
S&P 500 daily movements
Daily percentage changes in the S&P 500 from 2001 to 2005 - a deliberately hard classification problem where the honest answer is "barely better than chance".
Research (2)
A Mathematical Theory of Communication
Defines information quantitatively, introduces entropy as the measure of a source’s uncertainty, and proves limits on lossless compression and on reliable transmission over a noisy channel.
Games with Incomplete Information Played by Bayesian Players
Shows how games in which players are uncertain about one another’s payoffs can be transformed into games of complete but imperfect information, by treating each player as having a randomly assigned "type".
Projects (2)
Gridworld RL Lab
Value iteration, policy iteration, and Q-learning on the same gridworld, so a planner that knows the model can be compared directly against a learner that does not.
Bayesian Inference Playground
Prior to posterior updating with conjugate families, alongside an entropy and KL-divergence explorer that measures what each observation actually told you.