Understanding Multiple Comparisons
A 5% significance level promises that a single pre-specified test on data with no effect will mislead you one time in twenty. Nothing in that promise survives being applied repeatedly. With twenty independent null metrics the chance that at least one is significant is one minus 0.95 to the twentieth, which is 64.2%; at fifty metrics it is 92.3%. The tests are behaving exactly as designed - the reporting is not.
The same arithmetic applies to time rather than to metrics. An A/A experiment, where both arms are identical by construction, is significant 5.0% of the time when read once at its planned sample size. Check it ten times as data accumulates and stop at the first crossing, and it is significant 19.3% of the time. The estimate wanders as the sample grows, and a wandering quantity crosses a fixed line eventually; stopping there keeps the excursion and discards the return.
And to subgroups. Slicing a null experiment into twelve segments produces at least one significant segment in 46.4% of runs. More data does not help: with a larger sample a null test still manufactures spurious segments at the same rate, because the rate depends on the number of chances rather than on their precision.
Corrections exist and they work. Bonferroni divides the threshold by the number of tests, which brings twenty metrics back to a 4.9% family-wise rate, at the cost of each individual test running at 0.25% and needing far more data for the same power. False-discovery-rate procedures are less severe and control a different quantity. But every correction only counts the comparisons you declare, and the analyst who tried twenty metrics, three segments and four windows before reporting one has made comparisons no correction can see.
How to Calculate
P(at least one false positive) = 1 − (1 − α)^m; Bonferroni: test each at α/m
where
- α
- the significance level of a single test, usually 0.05
- m
- the number of independent chances: metrics, segments, variants (repeated looks at the same data are correlated and give less)
- family-wise error rate
- the chance of at least one false positive across all m
- false discovery rate
- the expected share of the claimed findings that are false
Example of Multiple Comparisons
At the 5% level, the chance that at least one independent null metric out of 1, 5, 20 and 50 comes out significant is 5.0%, 22.6%, 64.2% and 92.3%.
An A/A test read once: significant 5.0% of the time. The same test checked ten times and stopped early: 19.3%, an inflation of about four times.
A null experiment sliced into twelve segments yields at least one significant segment in 46.4% of runs, and a larger sample does not lower that rate.
Frequently Asked Questions
Does a correction fix a result I found by exploring?
Not really. It can only account for the comparisons you can count, and exploration generates comparisons you cannot. The honest treatment of an exploratory finding is as a hypothesis, confirmed or dropped by a fresh test aimed at it specifically.
Is it wrong to look at an experiment while it runs?
It is wrong to stop on what you see, under a test that assumed one look. Sequential designs and alpha-spending schedules exist precisely so that monitoring is allowed; they spend the error budget deliberately rather than accidentally.
Which correction should I use?
It depends on what you are protecting. Family-wise procedures such as Bonferroni or Holm are right when a single false claim is costly. False-discovery-rate procedures are right when you are screening many candidates and can tolerate a known share of false leads.
The Bottom Line
Multiplicity is the reason a study that looked at everything can find something in almost any dataset, including one where nothing is happening. The arithmetic is simple and the discipline is structural: declare the metric, the population and the stopping rule before the data arrives, because no correction applied afterwards can see the choices you did not record.