Skip to content
Kudos AI

Condition Number

The ratio of the largest to the smallest curvature of a loss surface, which alone determines how fast gradient descent can converge on it.

Also known as: Conditioning, κ

Understanding Condition Number

Near a minimum, a loss surface looks like a quadratic bowl, and the shape of that bowl is its Hessian. The condition number is the ratio of that matrix’s largest eigenvalue L to its smallest μ. At κ = 1 the level sets are circles and the gradient points straight at the minimum from anywhere. As κ grows they stretch into ellipses, and the gradient - always perpendicular to the level set - points more and more across the valley rather than down it.

This is not a metaphor for slowness but a calculation of it. A gradient step multiplies the error along each eigendirection of curvature λ by 1 − ηλ, independently of the other directions. Every direction must contract, so the step size is capped by the sharpest one at η < 2/L; and the slowest direction to contract is the flattest one. The two constraints together are what make κ, and nothing else about the problem, set the rate.

Balancing them gives the optimal step η = 2/(L + μ) and a contraction of (κ − 1)/(κ + 1) per step. The number is worth converting into a step count before dismissing it: even a mild κ of 24 contracts at 0.919488 per step, so shrinking the error a millionfold takes 165 steps - on a two-parameter problem. Real networks have condition numbers in the thousands.

Because κ is a property of the parameterisation and not of the underlying problem, it can be changed without changing what is being learned. Rescaling an input column moves the Hessian’s eigenvalues; so does standardising features, and so do batch and layer normalisation. This is the reason those techniques accelerate training so reliably: they are not statistical refinements, they are conditioning, and conditioning is the rate.

How to Calculate

κ = L/μ = λ_max(H) / λ_min(H), ρ_GD = (κ − 1)/(κ + 1)

where

H
the Hessian of the loss at the minimum
L
the largest eigenvalue: the sharpest curvature
μ
the smallest eigenvalue: the flattest curvature
ρ_GD
the factor the error shrinks by each step, at the optimal step size

Example of Condition Number

A least-squares problem is made ill-conditioned on purpose by shrinking one column of its design matrix by a factor of five. Its Hessian has L = 1.760627 and μ = 0.073849, so κ = 23.8410 and the predicted contraction is (κ − 1)/(κ + 1) = 0.919488 per step.

Measuring the run - taking the distance to the optimum at steps 20 and 100 and extracting the per-step ratio - gives 0.919488. Nothing was fitted: the prediction uses two eigenvalues and reproduces a run that knows nothing about them. Converted to a step count, 0.919488 predicts 164.6 steps to shrink the error a millionfold, and the run takes 165.

Momentum on the same problem, with β = 0.435629, reaches the same accuracy in 42 steps. The asymptotic rate (√κ − 1)/(√κ + 1) = 0.660022 predicts a fourfold-to-fivefold saving and the measured saving is 165/42 = 3.9; the shortfall is real and expected, because Polyak’s rate is asymptotic and this run reaches machine precision before the asymptotics arrive.

Frequently Asked Questions

How would I know my problem’s condition number without computing a Hessian?

For large models you do not compute it exactly; you infer it. A run that oscillates when the learning rate is raised slightly, but crawls when it is lowered, is reporting a large κ. The practical response is the same either way: normalise the inputs, add normalisation layers, and use momentum or an adaptive method.

Does a large condition number mean the model is bad?

No. It is a property of how the problem is parameterised, not of what the model can represent. The same function class in different coordinates can have a κ of 1 or of 10,000, with identical solutions - which is exactly why rescaling can transform training time without changing the model at all.

Why does momentum get √κ rather than κ?

Its velocity is a decaying average of past gradients. Along the flat direction successive gradients agree, so the average builds up; along the sharp direction they alternate in sign and largely cancel. One blind mechanism amplifies exactly the slow direction and damps exactly the oscillating one, and the resulting rate depends on √κ.

The Bottom Line

The condition number is the most useful thing to know about an optimisation problem before starting it: it predicts the convergence rate to several decimal places, it explains why feature scaling and normalisation work, and its square root is what momentum and adaptive methods are buying.