Skip to content
Kudos AI

Tagged “gradient-descent”

3 articles.

5 min readNeural Networks

The Slowest Direction Sets the Pace

The step size you are allowed is fixed by the steepest direction and the number of steps you need is fixed by the flattest, so the cost of gradient descent is their ratio. The same least-squares fit, to the same ten decimal places, takes 1742 steps in one basis, 147 in a rescaled one and exactly 1 in an orthonormal one, and momentum buys back the square root of the ratio rather than the ratio.

Machine LearningMathematics
10 min readNeural Networks

What Actually Makes Training Converge

A two per cent change in the learning rate separates a converged run from one five orders of magnitude away, a condition number predicts the convergence rate to six decimal places, and stochastic gradient descent with a fixed step never converges at all - it settles into a ball whose radius grows as the square root of the step. Every figure here was computed on a problem whose exact optimum is known.

OptimizationDeep LearningMachine Learning
8 min readNeural Networks

Backpropagation and Gradient Descent

How a neural network learns: the loss as a function of weights, gradient descent, and backpropagation as the chain rule applied backwards, with every partial derivative of a small network computed by hand and checked against autograd.

Deep LearningOptimizationMathematics