Understanding Matrix Factorisation
A recommender starts from a table with users down one side, items across the other, and almost nothing in the middle: in the simulation used throughout this path, 17.7% of the cells are filled. Matrix factorisation assumes that table is close to the product of two much smaller matrices, so that each user and each item is described by a handful of numbers and a prediction is their dot product. The factors are not named in advance; they are whatever directions turn out to explain the ratings.
What is worth internalising first is how much comes before the factors. Predicting the global mean for everything gives an RMSE of 0.9368. Adding a single offset per user and per item - capturing who rates generously and which items most people like - gives 0.6761. Eight latent factors on top of that give 0.5393. The offsets therefore carry about two thirds of the total gain, and they are cheap, robust on sparse rows and easy to explain.
Both parts need regularisation, and for the same reason. Counts vary by orders of magnitude, so an item with three ratings would otherwise receive an offset fitted to three numbers and be trusted like one fitted to three hundred. Shrinking each parameter towards zero in proportion to how little data supports it is what stops a handful of enthusiastic ratings from promoting an obscure item, which makes the shrinkage constant a real hyperparameter.
The error is not distributed evenly across users. The same fitted model scores 0.5210 for users with more than thirty ratings, 0.6018 between ten and thirty, and 0.6693 below ten - about 1.28 times worse for the sparsest users. Since the heavy raters supply most of the evaluation rows, the headline number is set by the people the system already knows, while a new user experiences the worst version of it. Any evaluation worth trusting is broken down by how much history each user has.
How to Calculate
r̂(u,i) = μ + b_u + b_i + p_u · q_i; minimise Σ (r − r̂)² + λ(‖p_u‖² + ‖q_i‖² + b_u² + b_i²)
where
- μ
- the global mean rating
- b_u, b_i
- per-user and per-item offsets, fitted before any interaction
- p_u, q_i
- the latent vectors, typically 8 to 200 numbers each
- λ
- the shrinkage that protects parameters estimated from few ratings
Example of Matrix Factorisation
On a simulated 800 by 300 catalogue that is 17.7% observed: RMSE 0.9368 for the global mean, 0.6761 after adding user and item offsets, 0.5393 after eight latent factors.
By user activity, the same model scores 0.6693 below ten ratings, 0.6018 between ten and thirty and 0.5210 above thirty: a cold-start penalty of 1.28 times.
Ranking every unrated item and taking each user’s top ten, only 91 of 300 items appear anywhere, even though the model has no explicit popularity term.
Frequently Asked Questions
How many factors should I use?
Enough that the validation error stops improving, and no more. More factors always fit the training ratings better and start to memorise the sparse rows, which is why the shrinkage constant and the factor count must be chosen together rather than one at a time.
Can the learned factors be interpreted?
Occasionally a direction lines up with something recognisable, but there is no guarantee: the factorisation is only identified up to rotation, so any interpretation you place on an individual axis is a story about one arbitrary basis among many.
Is RMSE the right thing to optimise?
It is the easy thing to optimise, and it measures the wrong task. Users see a ranked list, not a predicted number, so a model can improve its RMSE on ratings people would never have seen while leaving the top of every list unchanged. Ranking metrics and online tests measure what is actually delivered.
The Bottom Line
Matrix factorisation is a compact way to fill in a table nobody could fill by hand, and most of its accuracy arrives before the interesting part: two offsets carry two thirds of the gain. Its failure modes are concentrated exactly where the product needs it most, on the users who have told it least, so evaluate it in slices rather than by a single average.