Optimization
How a model actually learns: objective functions, gradients, convexity, and the algorithms that search a parameter space for the bottom of a loss surface.
Learning paths (3)
Supervised Machine Learning
Derive the workhorse supervised methods rather than merely calling them: least squares, logistic regression, shrinkage penalties, and tree ensembles.
Deep Learning Foundations
What a neural network actually computes, how the chain rule delivers every gradient in one backward sweep, and why convolution is the right prior for an image.
Support Vector Machines
Classify by choosing the widest slab that separates two classes, then relax it so a few points may sit inside, and finally bend it without ever building the space it is bent in.
Encyclopedia (8)
Gradient Descent
An iterative optimization algorithm that minimizes a function by repeatedly stepping in the direction opposite its gradient.
Stochastic Gradient Descent
Gradient descent in which each step uses the gradient of a small random sample of the data rather than all of it, trading an exact direction for far more steps per unit of compute.
Condition Number
The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.
Learning Rate Schedule
A rule that changes the step size over the course of training, large early so the run can travel and small late so it can settle.
Backpropagation
The algorithm that computes the gradient of a neural network’s loss with respect to every weight, by applying the chain rule backwards through the network.
Regularization
Any technique that constrains a model’s effective complexity in order to reduce variance and improve generalization, typically by penalizing large parameter values.
Support Vector Machine
A classifier that separates classes with the boundary leaving the widest possible margin, determined only by the closest training points.
Bellman Equation
The self-consistency condition that the utility of a state equals its immediate reward plus the discounted value of the best action available from it, averaged over the outcomes that action cannot control.
Articles (6)
What Actually Makes Training Converge
A two per cent change in the learning rate separates a converged run from one five orders of magnitude away, a condition number predicts the convergence rate to six decimal places, and stochastic gradient descent with a fixed step never converges at all - it settles into a ball whose radius grows as the square root of the step. Every figure here was computed on a problem whose exact optimum is known.
Support Vector Machines: Margins and Kernels
Why the widest slab between two classes is a good boundary, why insisting on a perfect one is self-defeating, how a budget for violations buys back stability, and how a kernel bends the boundary by working in a space it never has to build.
Linear Regression from First Principles
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Logistic Regression and Classification
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Regularization: Ridge and Lasso
Adding a penalty on coefficient size to trade a little bias for a large reduction in variance, and why the L1 penalty sets coefficients exactly to zero while L2 only shrinks them, with both fitted numerically.
Backpropagation and Gradient Descent
How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.
Tools (1)
Research (2)
Learning Internal Representations by Error Propagation
Presents backpropagation as a general method for training multilayer networks, showing that hidden layers can learn useful internal representations rather than needing to be designed by hand.
Support-Vector Networks
Introduces the support vector machine with a soft margin, separating classes by the widest possible margin while permitting bounded violations, and using kernels to obtain non-linear boundaries.
Projects (2)
Neural Network From Scratch
A feed-forward network in NumPy with hand-derived backpropagation, validated against numerical gradients so the calculus is proven rather than trusted.
Statistical Learning Toolkit
Least squares, logistic regression, ridge and lasso, and k-fold cross-validation implemented from their estimating equations and checked against scikit-learn.