The Model That Picks Its Own Training Data
Two fitted offsets deliver 66% of a recommender’s gain in accuracy before any latent factor is learned, the error is 1.28 times worse for the users who have said least, only 30% of the catalogue reaches anyone’s top ten with no explicit popularity term in the model, and after six rounds of self-selected data the system is 1.14 times worse exactly where it stopped looking.
Prerequisites: Supervised Machine Learning
Almost every model you train is a function applied to a fixed dataset. A recommender is not. It decides what gets shown, what gets shown decides what gets rated, and what gets rated is its next training set.
That one property changes what accuracy means, and it is worth putting numbers on. Everything below is measured on a simulated 800 by 300 catalogue where the true preferences are known - which real data never gives you.
Most of the accuracy arrives before the model does
The matrix is 17.7% observed. Three models, in order of ambition:
| model | RMSE | share of the total gain |
|---|---|---|
| predict the global mean | 0.9368 | - |
| + one offset per user and per item | 0.6761 | 66% |
| + 8 latent factors | 0.5393 | 100% |
Two numbers per user and per item - who rates generously, and which items most people like - deliver two thirds of everything the full model achieves.
That is the useful shape of the problem. A large part of any rating is not about the match between a person and an item at all, and fitting that part first is cheap, robust on sparse rows, and easy to explain to anyone who asks why an item was recommended.
The factors earn the remaining third by capturing which kinds of people like which kinds of items. Nobody names those directions in advance, and the factorisation is only identified up to rotation, so any interpretation of an individual axis is a story about one arbitrary basis among many.
Why the shrinkage is not a detail
Without the , an item rated three times gets an offset fitted to three numbers and is trusted exactly as much as one fitted to three hundred. At , an item whose three ratings average 1.2 above the mean keeps 0.327 of that - 27% - while an item with three hundred such ratings keeps 1.169, or 97%.
That asymmetry is the whole point, and it is what stops four enthusiastic ratings from putting an obscure item at the top of every list.
The figure below turns that denominator into a dial, and shows what it decides: not the offsets but the order. Take the constant to zero and a short film four people rated tops the list, because with no shrinkage the count does not enter the calculation at all. The arithmetic is not wrong - that really is the average of what those four said. It is just not an estimate of what the next person will think, and only the constant knows the difference.
Interactive: the constant that decides the leaderboard
Drag the shrinkage constant and watch the top of the list change hands.
- λ
- 8
- Top of the list
- a well-loved classic
- The 4-rating item keeps
- 33%
- The 300-rating item keeps
- 97%
At λ = 8 the four-rating item keeps only 33% of its apparent quality while the three-hundred-rating one keeps 97%, and the top of the list is a well-loved classic, on 300 ratings. That asymmetry is the whole mechanism: each offset is pulled towards zero in proportion to how little data supports it, so a thin row barely moves from the global mean and a thick one is left almost alone. Push λ further and the catalogue collapses towards the mean - which is the trade the constant controls.
The error is not spread evenly
The full model's RMSE is 0.5393. Split the same test set by how much history each user has:
| user's training ratings | RMSE | test rows |
|---|---|---|
| over 30 | 0.5210 | 6,928 |
| 10 to 30 | 0.6018 | 1,319 |
| under 10 | 0.6693 | 254 |
A 1.28× penalty for the sparsest users - and look at the third column. Heavy raters supply 6,928 of the 8,501 test rows, so they set the headline figure almost single-handedly, and they are exactly the people the model already knows.
A new user does not experience the 0.5210 model. They experience the 0.6693 one. The average runs over test rows rather than over users, and that row weighting hands it to the group who need help least, which makes it the wrong number to optimise and the wrong number to report.
For a user with nothing at all, the personalised term carries no information: initialise the unfitted vector at zero, or drop the term, and what remains is : quality and popularity, the same list for everybody. That is a reasonable default worth naming honestly, and it sets the bar for anything you build for new users.
Ranking concentrates on its own
Score every unrated item for every user, take each user's top ten, and count the distinct items across 8,000 slots: 91 out of 300, or 30% of the catalogue.
The model has no explicit popularity term. The rating count enters only through the shrinkage, which pulls rarely rated items toward the average. Beyond that, the concentration comes from the item offsets capturing quality, and quality being shared: an item most people like ranks near the top for most people, and personalisation reshuffles the order rather than replacing the pool.
This matters because it changes the remedy. Concentration is not a bug introduced by a popularity feature you can remove. It is what happens when a ranking system meets correlated tastes, so if catalogue coverage matters to you it has to be an explicit objective - no amount of improving the accuracy metric will produce it.
And then the loop closes
Let the system run. Each round it shows every user its five best unshown items, users rate only what they were shown, the model refits. Six rounds. To keep the loop simple, the model here is the offsets alone, , so every user is ranked by : the same list for everyone, minus what each has already been shown.
Afterwards, 28% of the matrix has been shown at some point. Comparing the refitted model against the known truth:
| region | RMSE against the truth |
|---|---|
| items it showed | 0.5441 |
| items it never showed | 0.6213 |
A 1.14× blind spot, in exactly the region it chose not to look at. Every round, the system's current beliefs decide what gets rated next, so its future training data is a sample of its present opinions. Where it was confident and right it collected confirmation; where it was confident and wrong it collected nothing, and nothing arrived to correct it.
Why offline evaluation cannot see it
Your logs contain ratings for the items you showed. Your offline test set is a held-out slice of that same log. So the evaluation runs in the region where the model is accurate. Against the truth the model is 0.5441 off there and 0.6213 off everywhere else, but the log holds only the first region, and noisy ratings of it: scored against those, on the shown cells it was fitted on, the same model reads 0.611, a number that carries no trace of the region it never showed. Holding out a slice of the log instead barely moves it, because a held-out slice of the log is still inside the shown region.
A model that has quietly stopped understanding the 72% of the matrix it never showed will look excellent by every offline metric you have, and will keep looking excellent as the blind spot grows.
What helps, honestly measured
Reserve one slot in five for a random item and run the same six rounds. The error on unshown items falls from 0.6213 to 0.6162, and the blind spot narrows from 1.14× to 1.12× - partly because the error on shown items rose, from 0.5441 to 0.5497.
That is a small effect, and the honest reading of a small measured effect is that it is small. A fifth of your recommendation slots bought a correction of under one percent in the region you were worried about. Exploration is insurance, not a fix: it keeps the tail of the catalogue from disappearing entirely, and it does not undo six rounds of self-selected data.
Two measures cost more and do more:
- log the propensity - the probability the system had of showing each item at the moment it showed it. With those recorded, offline estimates can be reweighted to correct for the selection, the same inverse-weighting idea used for observational data in causal inference. It has to be decided before you need it, because propensities cannot be reconstructed afterwards.
- experiment on the policy, not the model - randomise which ranking system a user gets and measure the outcome you care about. It is the only method that measures the system as deployed rather than a model in isolation.
The property underneath all of it
Two offsets carry two thirds of the gain in accuracy, so most of what a recommender knows is "who rates generously" and "what is good", not "who likes what". The error is worst for the users whose experience decides whether they stay. Ranking concentrates without being told to. And a system trained on its own output goes blind in the region it stopped showing while every metric it has continues to look fine.
All four follow from the same thing: the model's output determines its next training set. Once that is true, accuracy on logged data stops being a measurement of quality and becomes a measurement of habit - and the only way back out is to deliberately break the loop, by exploring, by recording why each choice was made, or by testing the policy itself.
References & further reading
- Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
- Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.