Skip to content
Kudos AI

Confidence Interval

A range computed from data by a procedure that, repeated over many samples, contains the true value a stated proportion of the time. The stated proportion is a property of the procedure, not of any particular interval it produces.

Also known as: Interval estimate, Coverage interval

Understanding Confidence Interval

An estimate on its own says nothing about its own precision. A confidence interval attaches that missing information by reporting a range instead of a point, together with a promise about how the range was built: apply this procedure to sample after sample, and the stated fraction of the resulting intervals will contain the true value.

The promise is about the procedure, which is the part most easily misread. Once the data are in hand and the numbers are fixed, the interval either covers the truth or it does not, and no probability remains to be assigned. Saying there is a 95% chance the parameter lies inside a computed interval treats a fixed unknown as if it were random. The usable reading is the frequency one: intervals built this way are right 95% of the time, and this is one of them.

Two sources of uncertainty enter when the population spread is unknown, and using the normal quantile accounts for only the first. At a sample of 10 the normal quantile 1.959964 gives an exact coverage of 0.918351 rather than 0.95, because the estimate of the spread is itself uncertain; the t quantile 2.262157 is the correction, and for normally distributed observations it is exact rather than conservative. The gap narrows as data accumulate: at 29 degrees of freedom the quantile is already 2.045230.

For a proportion, the textbook interval fails in a way that is worth seeing in full, because the same unlucky sample both moves the centre and mis-sizes the width. At n = 50, exact enumeration of all 51 outcomes gives coverage 0.878917 at p = 0.1 and 0.394848 at p = 0.01 - the latter because 0.605006 of such samples contain no successes at all, making the estimated standard error exactly zero and collapsing the interval to a single point. The Wilson interval, which does not estimate the width from the same slip, gives 0.970308 where the standard one gives 0.878917.

How to Calculate

x̄ ± t_{α/2,n−1}·s/√n, coverage = P(L(X) ≤ θ ≤ U(X))

where

x̄, s
the sample mean and sample standard deviation, both computed from the data
t_{α/2,n−1}
the Student t quantile; 2.262157 for a 95% interval at n = 10
θ
the fixed, unknown population value - not random, which is why the probability sits on L and U
L(X), U(X)
the endpoints, which are random because they are computed from the sample

Example of Confidence Interval

Twenty thousand simulated samples of size 10 from a normal population, each giving a 95% t-interval for the mean, cover the true mean 0.95 of the time. Replacing the t quantile with the normal one drops the exact coverage to 0.918351.

The standard interval for a proportion at n = 50, evaluated exactly over all 51 possible outcomes: 0.919867 at p = 0.05, 0.878917 at p = 0.1, 0.937531 at p = 0.2 and 0.935091 at p = 0.5. The Wilson interval on the same points gives 0.962224, 0.970308, 0.950701 and 0.935091 - identical at p = 0.5, materially better where the standard one fails.

Coverage is not monotone in the sample size. At p = 0.2 the standard interval covers 0.951214 at n = 23 and 0.872862 at n = 24: one additional observation costs nearly eight points. Coverage falls from one size to the next at 29 of the 180 steps between 20 and 200, though only three of those falls exceed five points.

Frequently Asked Questions

Is there a 95% chance the true value is in my interval?

No. The true value is fixed and your interval is fixed, so it is in or it is not. What is 95% is the long-run success rate of the procedure that produced it.

Does a wider interval mean a worse study?

It means a less precise one, which is honest rather than bad. Width falls as the square root of the sample size: with a population standard deviation of 2, the standard error is 0.4 at n = 25 and 0.2 at n = 100, so each halving costs four times the data.

When should the standard interval for a proportion be avoided?

Whenever the proportion may be near 0 or 1, and in practice more often than that. Its coverage is erratic in the sample size rather than converging smoothly, so checks like np > 5 do not make it dependable. The Wilson interval is a direct replacement.

The Bottom Line

A confidence interval reports precision honestly, provided its promise is read as a property of the procedure rather than of the one interval in front of you. And the promise is only as good as the construction: the standard interval for a proportion covers 0.878917 where it claims 0.95, so which interval you choose is a real decision.