Skip to content
Kudos AI
Lire en français
Statistical Inference

The 95% Interval That Covers 81% of the Time

The textbook confidence interval for a proportion has exact coverage you can compute by summing over the n+1 possible samples, and at n = 30 with p = 0.10 it is 0.8085 rather than 0.95. Coverage does not improve monotonically with n, and in a rare-event setting it can fall to 0.0392. Two one-line alternatives fix it.

4 min readKudos AI

Prerequisites: What a Sample Can and Cannot Tell You

Many intervals drawn one above another against a fixed true value, the ones that miss it highlighted, and the running tally of how many covered it settling below the level it promised.

A 95% confidence interval promises one thing: over repeated samples, the interval contains the true parameter 95% of the time. For a proportion this is checkable exactly. There are only n+1n + 1 possible samples, so you can build the interval for each, ask whether it contains pp, and add up the binomial probabilities of the ones that do. No simulation, no approximation.

Do that for the standard interval, p^±1.96p^(1−p^)/n\hat{p} \pm 1.96\sqrt{\hat{p}(1-\hat{p})/n}, at n=30n = 30 and p=0.10p = 0.10:

coverage=0.8085.\text{coverage} = 0.8085.

One in five samples produces an interval that does not contain the truth. The interval is labelled 95%.

A. Not a small-sample caveat

The usual defence is that 30 is a small sample. Here is the same calculation across a range of settings, with two alternatives alongside:

nnppWaldWilsonAgresti-Coull
300.100.80850.97420.9742
400.050.86810.95200.9861
1000.050.87750.96590.9659
200.200.92080.95630.9563
500.300.93470.95670.9567
1000.500.94310.94310.9431

A hundred observations of a 5% event still gives 0.8775. And the last row is worth noting on its own: at the most favourable proportion there is, with a hundred observations, the interval still covers 94.31% rather than 95%. It is never quite right, and in the corner of the parameter space where most real questions live - rare events, moderate samples - it is not close.

B. More data does not fix it monotonically

Coverage does not climb as nn grows. At p=0.10p = 0.10:

nn2530354045505560
coverage0.91870.80850.87130.91450.93560.87890.91530.9413

Going from 25 observations to 30 takes the coverage from 0.9187 down to 0.8085. Going from 45 to 50 takes it from 0.9356 down to 0.8789.

In the figure, pick 0.10 beside true p and drag sample size n to 30: Coverage at n reads 80.9%, the 0.8085 above.

Interactive: the coverage a 95% interval actually delivers

Exact, by enumeration. A binomial has only n + 1 outcomes.

95%80%n = 20200
Coverage at n
95.1%
Shortfall
0.0%
Sizes that lose ground
29
Samples with no successes
0.6%
true p

At n = 23 the interval advertising 95% delivers 95.1%. Note the shape: coverage does not creep up toward the line, it jumps across it 29 times in this range alone, each tooth being one more attainable value of the estimate. Drag n from 23 to 24 and watch eight points disappear with a single extra observation. No rule of thumb about large enough n protects against that, because the thing being counted is discrete.

The oscillation is not noise, because there is no noise here: every entry is an exact sum. It happens because the set of achievable p^\hat{p} values is discrete, and as nn changes, one of the n+1n + 1 possible outcomes moves across the boundary of containing pp and takes its whole binomial probability with it. Any statement of the form "the interval is fine once nn is large enough" has to survive this table, and the usual rules of thumb (np^>5n\hat{p} > 5, and so on) do not.

C. The degenerate case

Take n=40n = 40 and p=0.001p = 0.001. The coverage of the standard interval is

0.0392.0.0392.

A 95% interval that contains the truth in 4% of samples. The mechanism is immediate: P(k=0)=0.9608P(k = 0) = 0.9608, and when k=0k = 0 the estimate is p^=0\hat{p} = 0, the standard error is 0⋅1/n=0\sqrt{0 \cdot 1 / n} = 0, and the interval is the single point [0,0][0, 0], which does not contain 0.0010.001. The formula reports perfect certainty in exactly the situation where it has seen nothing.

This is the case that turns up in production: the rare failure, the rare click, the rare adverse event.

D. What to use instead

Both alternatives in the table are one line and neither needs a new idea.

  • Wilson. Invert the test rather than the estimate: keep the values of pp the data would not reject. The interval is no longer centred on p^\hat{p} and it never collapses to a point.
  • Agresti-Coull. Add two successes and two failures, then use the ordinary formula on the adjusted counts. It is the Wald interval with the estimate pulled off the boundary, and in the table above it is the more conservative of the two.

The general lesson is the one the exact calculation makes available: coverage is computable, so compute it. For a proportion it is a finite sum. For anything else, simulating the sampling distribution and counting how often the interval contains the value you generated from takes a few lines, and it is the only way to know whether the number on the label is the number you get.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

10 min readStatistical Inference

What a Sample Can and Cannot Tell You

Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.

StatisticsProbability
6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
3 min readCausal Inference

The Control Variable That Invents a Relationship

Two independent causes and one common effect. Adjust for the effect and the causes acquire a correlation of exactly -1: a regression of A on B recovers a coefficient of +0.0030, and adding the common effect as a control turns it into -1.0000. Selecting a sample does the same thing invisibly, which is why "control for everything you measured" is not a defensible rule.

Statistics
← Back to all articles