Skip to content
Kudos AI

Conjugate Prior

A prior chosen so that the posterior belongs to the same family, which turns Bayesian updating into arithmetic on the parameters and makes the prior readable as a number of imagined observations.

Also known as: Beta-binomial prior, Pseudo-counts

Understanding Conjugate Prior

Bayes rule multiplies a prior by a likelihood and renormalises. In general the result has no closed form and needs numerical integration or sampling. A conjugate prior is one chosen so that the multiplication keeps the answer inside a family the prior already belongs to, in which case updating reduces to changing the family parameters.

The best-known pair is Beta and Bernoulli. A Beta prior with parameters a and b, updated on c successes and l failures, becomes a Beta with parameters a + c and b + l. Nothing is integrated and nothing is approximated. The same pattern repeats elsewhere: Dirichlet with categorical, Gamma with Poisson, Gaussian with a Gaussian of known variance.

Reading the parameters as pseudo-counts is what makes the prior honest. A Beta(1, 1) prior is uniform and is worth two imagined observations, one of each kind; a Beta(50, 50) prior is worth a hundred, and will dominate a small sample. The strength of a belief and its direction are separate knobs, and the parameters name both.

Conjugacy is chosen for tractability, and that is a real cost as well as a convenience. If the honest prior is not in the conjugate family - bimodal, say, or with a hard cutoff - forcing it there to keep the arithmetic pretty is a modelling error. With modern sampling the arithmetic is no longer the binding constraint it once was.

How to Calculate

\text{Beta}(a, b) \times \text{Binomial}(c, l) \;\propto\; \text{Beta}(a + c,\; b + l), \qquad \mathbb{E}[\theta \mid \mathbf{d}] = \frac{a + c}{a + b + c + l}

where

a, b
prior pseudo-counts of successes and failures
c, l
observed successes and failures

Example of Conjugate Prior

Unwrap 25 candies and find 18 cherry. Maximum likelihood reports 18/25 = 0.72. With a uniform Beta(1, 1) prior the posterior is Beta(19, 8), whose mean is 19/27 = 0.703704 and whose mode is 18/25 = 0.72, the same as the maximum-likelihood answer. The mean is pulled toward one half; the mode is not.

Strengthen the prior and the pull grows: Beta(2, 2) gives a posterior mean of 20/29 = 0.689655, Beta(5, 5) gives 23/35 = 0.657143, and Beta(50, 50) gives 68/125 = 0.544. The data have not changed; only the number of imagined observations has.

The failure case is where this earns its keep. Unwrap one candy, find it cherry, and maximum likelihood reports 1: lime is not unlikely but impossible, and no later evidence can be multiplied back in. The Beta(1, 1) posterior mean is 2/3 instead, which is exactly add-one smoothing. The standard hack and the uniform prior are the same calculation.

Frequently Asked Questions

Does a conjugate prior make the answer more correct?

No. It makes it easier to compute. A conjugate prior that misrepresents your belief gives a tidy posterior that is wrong, and the tidiness is no evidence either way.

How much data does it take to overwhelm the prior?

Roughly as much as the prior is worth in pseudo-counts. A Beta(1, 1) prior is worth two observations and vanishes almost at once; a Beta(50, 50) prior is worth a hundred and still moves the answer noticeably after 25.

The Bottom Line

A conjugate prior keeps the posterior in the prior’s family, so updating is arithmetic on pseudo-counts. It is chosen for tractability rather than truth, and it explains why add-one smoothing, which looks like a hack, is the posterior mean under a uniform prior.