The Direction That Changes When You Change Units
Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.
Prerequisites: Unsupervised Learning: Structure Without Labels
Here are twelve people, each measured twice.
| Height | 168 | 172 | 175 | 177 | 180 | 181 | 183 | 185 | 186 | 188 | 191 | 193 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Weight | 58 | 63 | 61 | 68 | 70 | 66 | 74 | 72 | 78 | 75 | 82 | 80 |
Height in centimetres, weight in kilograms. Ask principal component analysis for the single direction along which these people vary most, and it answers
an almost exactly even blend of the two measurements, carrying 97.500% of the total variance.
Now record the heights in millimetres instead. Nobody has grown. The answer becomes : almost pure height, carrying 99.904%.
Record them in metres. The answer becomes : almost pure weight, carrying 99.999%.
Three different directions, three different stories about what drives variation in this group, and not one number in the underlying data has changed. The correlation between height and weight is 0.9500 in all three, because correlation is unitless and the first principal component is not.
A. What the method actually maximises
The first principal component is the direction , of unit length, that maximises the variance of the projected data:
The figure uses six unitless points rather than the twelve people, to show the objective itself: turn the Direction of the line and the direction that captures the most variance is the one that leaves the least squared distance to the line.
Interactive: what a projection keeps, and what it loses
Six points, one line, every angle available.
- Variance kept
- 2.0000
- Squared distance lost
- 6.1667
- Their sum
- 8.1667
Keeping 2.0000 of the variance. Watch the third figure as you turn the line: the first two move in opposite directions and their sum never changes, whatever angle you choose. Maximising the variance you keep and minimising the perpendicular distance to the line are not two criteria that happen to agree - they are one criterion written two ways. No angle beats the principal component; try to find one.
The constraint is what makes the problem well posed, and it is also where the units enter. Setting the length of to one treats a step of size one in the first coordinate as the same size as a step of one in the second. In centimetres and kilograms that is a defensible claim: one centimetre and one kilogram are comparable fractions of the spread in this group. In millimetres and kilograms it says one millimetre of height counts for as much as one kilogram of weight, which is not a claim anyone would make out loud. The method makes it silently, and then answers the question it was actually asked.
You can see it in the covariances. Height varies by
depending only on how it was written down, while weight sits at throughout. Variance is in squared units, so changing a unit by a factor of ten changes a variance by a factor of a hundred. The direction of greatest variance follows whichever variable was recorded in the smallest unit.
B. The proportion that explains nothing
The metres answer is the one worth dwelling on. Its first component carries 99.999% of the variance, which is the kind of figure that ends an analysis: one number replaces two with essentially no loss, and the scree plot has a cliff after the first component that could not be sharper.
It is a statement about metres. Height in metres varies by 0.0058, weight in kilograms by 58.45, and a ratio of ten thousand to one means the second direction has almost nothing left to carry. The component "explains" 99.999% of a total that weight alone very nearly is. Had the heights been in metres and the weights in grams, the same calculation would have declared height the whole story instead.
A proportion of variance explained is only a measure of concentration once the variables are on a footing you are willing to defend. Before that it is a measure of your spreadsheet.
C. What standardising chooses
The usual fix is to standardise: divide each variable by its own standard deviation before running PCA, which is the same as running PCA on the correlation matrix. Every variable then varies by exactly one, and no unit survives to tilt the answer.
For these twelve people that gives
Both numbers deserve a second look. is , and it is not a coincidence: with two standardised variables the correlation matrix has ones on the diagonal and off it, and its eigenvectors are always and whatever happens to be. The direction was fixed before anyone looked at the data. Only the sign carries information, and only about the sign of .
The proportion is similarly determined: the eigenvalues are and , so
which is the correlation, rearranged. On two standardised variables, PCA recovers what a correlation already told you and adds nothing. That is not an argument against standardising. It is a reminder that the method has no private source of information: it reports the covariance structure it is handed, and standardising is a decision about which covariance structure to hand it.
D. So when should you not standardise?
Standardising is right when the variables measure different things in different units, which is most of the time. It is wrong when the differences in scale are themselves the signal:
- Pixel intensities. All channels are in the same units on the same range. A channel that barely varies across the images is genuinely uninformative, and standardising promotes its noise to the same footing as real structure.
- Returns across assets. All in the same units, and a volatile asset is volatile. Standardising deletes exactly the thing a risk decomposition is looking for.
- Repeated measurements of one quantity, such as a spectrum or a time series sampled at many points. The relative sizes across the axis are part of the shape.
The test is not statistical. Ask whether a step of one unit in variable A is comparable to a step of one unit in variable B. If yes, leave the data alone. If the question is not even meaningful, standardising is how you decline to answer it, and the answer you get afterwards is about correlations rather than about quantities.
E. What to take from twelve people
PCA is often introduced as though it finds structure that is simply there, the way a regression finds a slope that is simply there. The three answers above came from one dataset and one method, and they disagree completely.
What PCA finds is the direction of greatest variance in the coordinates you supplied. Choosing those coordinates - millimetres or metres, raw or standardised - is a modelling decision of the same weight as choosing a predictor, and it is made whether or not anyone notices making it. The honest version of a PCA report says which choice was made and why, and the fastest way to find out whether it mattered is to make the other choice and look.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.