Skip to content
Kudos AI

Maximum Likelihood Estimation

A method of fitting a model by choosing the parameter values that make the observed data most probable.

Also known as: MLE

Understanding Maximum Likelihood Estimation

Maximum likelihood inverts the usual direction of probabilistic reasoning. Ordinarily a fixed model assigns probabilities to possible datasets. Here the dataset is fixed, because it has been observed, and the parameters vary: the likelihood function reports, for each candidate parameter value, how probable the observed data would have been under it. The estimate is the value that maximizes this.

In practice the log-likelihood is maximized rather than the likelihood. The likelihood of independent observations is a product of many small numbers, which underflows numerically and is awkward to differentiate. Taking logarithms converts the product into a sum without moving the maximum, since the logarithm is monotonic.

James and colleagues make the connection to familiar methods explicit: maximum likelihood is the general approach used to fit logistic regression and many other non-linear models, and in the linear regression setting least squares is itself a special case of it. Minimizing squared error is what maximum likelihood reduces to when the errors are assumed normal with constant variance.

The same identity connects it to deep learning. Maximizing the log-likelihood of observed labels is the same optimization as minimizing cross-entropy, so a network trained with cross-entropy loss is performing maximum likelihood estimation. This is why the loss functions used across statistics and deep learning look so similar once written out.

How to Calculate

θ̂ = argmax_θ Πᵢ P(xᵢ | θ) = argmax_θ Σᵢ log P(xᵢ | θ)

where

θ
the parameters being estimated
xᵢ
the i-th observed data point
P(xᵢ | θ)
the probability (or density) of that observation under θ
θ̂
the maximum likelihood estimate

Example of Maximum Likelihood Estimation

A coin is flipped 10 times and lands heads 7 times. Treating the flips as independent with unknown head probability p, the likelihood is proportional to p⁷(1 − p)³.

Maximizing the log-likelihood 7 log p + 3 log(1 − p) by setting its derivative 7/p − 3/(1 − p) to zero gives 7(1 − p) = 3p, hence p̂ = 0.7. The estimate is simply the observed proportion, which is reassuring.

The example also exposes the method’s main weakness. Had the coin landed heads all 10 times, maximum likelihood would report p̂ = 1.0, asserting that tails is impossible on the strength of ten flips. Regularization or a Bayesian prior is what prevents this kind of overconfident conclusion from small samples.

Frequently Asked Questions

What is the difference between likelihood and probability?

They are the same function read in opposite directions. Probability fixes the parameters and asks how likely various datasets are; likelihood fixes the observed data and asks how well various parameter values explain it. Likelihood is not a probability distribution over parameters and does not integrate to one.

Why maximize the logarithm instead of the likelihood itself?

Because the logarithm is strictly increasing, it has its maximum at the same place, while converting a product of many small probabilities into a numerically stable sum that is far easier to differentiate.

How does maximum likelihood relate to Bayesian estimation?

Maximum likelihood uses only the likelihood; Bayesian estimation multiplies it by a prior and works with the resulting posterior. Maximum likelihood coincides with the Bayesian maximum a posteriori estimate when the prior is uniform.

The Bottom Line

Maximum likelihood picks the parameters under which the observed data would have been most probable. It unifies least squares, logistic regression, and cross-entropy training, and its tendency toward overconfidence on small samples is exactly what regularization and priors exist to temper.