Skip to content
Kudos AI
Lire en français
Recommender Systems

The Model That Picks Its Own Training Data

Two fitted offsets deliver 66% of a recommender’s gain in accuracy before any latent factor is learned, the error is 1.28 times worse for the users who have said least, only 30% of the catalogue reaches anyone’s top ten with no explicit popularity term in the model, and after six rounds of self-selected data the system is 1.14 times worse exactly where it stopped looking.

8 min readKudos AI

Prerequisites: Supervised Machine Learning

A rating matrix mostly empty with two rows of offsets absorbing most of the error, an error bar splitting by how much each user has rated, and hundreds of ranked lists collapsing onto the same small set of items.

Almost every model you train is a function applied to a fixed dataset. A recommender is not. It decides what gets shown, what gets shown decides what gets rated, and what gets rated is its next training set.

That one property changes what accuracy means, and it is worth putting numbers on. Everything below is measured on a simulated 800 by 300 catalogue where the true preferences are known - which real data never gives you.

Most of the accuracy arrives before the model does

The matrix is 17.7% observed. Three models, in order of ambition:

modelRMSEshare of the total gain
predict the global mean0.9368-
+ one offset per user and per item0.676166%
+ 8 latent factors0.5393100%

Two numbers per user and per item - who rates generously, and which items most people like - deliver two thirds of everything the full model achieves.

That is the useful shape of the problem. A large part of any rating is not about the match between a person and an item at all, and fitting that part first is cheap, robust on sparse rows, and easy to explain to anyone who asks why an item was recommended.

The factors earn the remaining third by capturing which kinds of people like which kinds of items. Nobody names those directions in advance, and the factorisation is only identified up to rotation, so any interpretation of an individual axis is a story about one arbitrary basis among many.

Why the shrinkage is not a detail

bi=∑u(rui−μ−bu)ni+λb_i = \frac{\sum_{u} (r_{ui} - \mu - b_u)}{n_i + \lambda}

Without the λ\lambda, an item rated three times gets an offset fitted to three numbers and is trusted exactly as much as one fitted to three hundred. At λ=8\lambda = 8, an item whose three ratings average 1.2 above the mean keeps 0.327 of that - 27% - while an item with three hundred such ratings keeps 1.169, or 97%.

That asymmetry is the whole point, and it is what stops four enthusiastic ratings from putting an obscure item at the top of every list.

The figure below turns that denominator into a dial, and shows what it decides: not the offsets but the order. Take the constant to zero and a short film four people rated tops the list, because with no shrinkage the count does not enter the calculation at all. The arithmetic is not wrong - that really is the average of what those four said. It is just not an estimate of what the next person will think, and only the constant knows the difference.

Interactive: the constant that decides the leaderboard

Drag the shrinkage constant and watch the top of the list change hands.

a well-loved classic300 ratings+3a niche documentary11 ratings+1a steady favourite96 ratings+2a popular sequel180 ratings+2a cult short film4 ratings-4a broad crowd-pleaser420 ratings+1an out-of-print album3 ratings-5a divisive experiment27 ratings
raw averagefitted offset
λ
8
Top of the list
a well-loved classic
The 4-rating item keeps
33%
The 300-rating item keeps
97%

At λ = 8 the four-rating item keeps only 33% of its apparent quality while the three-hundred-rating one keeps 97%, and the top of the list is a well-loved classic, on 300 ratings. That asymmetry is the whole mechanism: each offset is pulled towards zero in proportion to how little data supports it, so a thin row barely moves from the global mean and a thick one is left almost alone. Push λ further and the catalogue collapses towards the mean - which is the trade the constant controls.

The error is not spread evenly

The full model's RMSE is 0.5393. Split the same test set by how much history each user has:

user's training ratingsRMSEtest rows
over 300.52106,928
10 to 300.60181,319
under 100.6693254

A 1.28× penalty for the sparsest users - and look at the third column. Heavy raters supply 6,928 of the 8,501 test rows, so they set the headline figure almost single-handedly, and they are exactly the people the model already knows.

A new user does not experience the 0.5210 model. They experience the 0.6693 one. The average runs over test rows rather than over users, and that row weighting hands it to the group who need help least, which makes it the wrong number to optimise and the wrong number to report.

For a user with nothing at all, the personalised term carries no information: initialise the unfitted vector at zero, or drop the term, and what remains is μ+bi\mu + b_i: quality and popularity, the same list for everybody. That is a reasonable default worth naming honestly, and it sets the bar for anything you build for new users.

Ranking concentrates on its own

Score every unrated item for every user, take each user's top ten, and count the distinct items across 8,000 slots: 91 out of 300, or 30% of the catalogue.

The model has no explicit popularity term. The rating count enters only through the shrinkage, which pulls rarely rated items toward the average. Beyond that, the concentration comes from the item offsets capturing quality, and quality being shared: an item most people like ranks near the top for most people, and personalisation reshuffles the order rather than replacing the pool.

This matters because it changes the remedy. Concentration is not a bug introduced by a popularity feature you can remove. It is what happens when a ranking system meets correlated tastes, so if catalogue coverage matters to you it has to be an explicit objective - no amount of improving the accuracy metric will produce it.

And then the loop closes

Let the system run. Each round it shows every user its five best unshown items, users rate only what they were shown, the model refits. Six rounds. To keep the loop simple, the model here is the offsets alone, μ+bu+bi\mu + b_u + b_i, so every user is ranked by bib_i: the same list for everyone, minus what each has already been shown.

Afterwards, 28% of the matrix has been shown at some point. Comparing the refitted model against the known truth:

regionRMSE against the truth
items it showed0.5441
items it never showed0.6213

A 1.14× blind spot, in exactly the region it chose not to look at. Every round, the system's current beliefs decide what gets rated next, so its future training data is a sample of its present opinions. Where it was confident and right it collected confirmation; where it was confident and wrong it collected nothing, and nothing arrived to correct it.

Why offline evaluation cannot see it

Your logs contain ratings for the items you showed. Your offline test set is a held-out slice of that same log. So the evaluation runs in the region where the model is accurate. Against the truth the model is 0.5441 off there and 0.6213 off everywhere else, but the log holds only the first region, and noisy ratings of it: scored against those, on the shown cells it was fitted on, the same model reads 0.611, a number that carries no trace of the region it never showed. Holding out a slice of the log instead barely moves it, because a held-out slice of the log is still inside the shown region.

A model that has quietly stopped understanding the 72% of the matrix it never showed will look excellent by every offline metric you have, and will keep looking excellent as the blind spot grows.

What helps, honestly measured

Reserve one slot in five for a random item and run the same six rounds. The error on unshown items falls from 0.6213 to 0.6162, and the blind spot narrows from 1.14× to 1.12× - partly because the error on shown items rose, from 0.5441 to 0.5497.

That is a small effect, and the honest reading of a small measured effect is that it is small. A fifth of your recommendation slots bought a correction of under one percent in the region you were worried about. Exploration is insurance, not a fix: it keeps the tail of the catalogue from disappearing entirely, and it does not undo six rounds of self-selected data.

Two measures cost more and do more:

  • log the propensity - the probability the system had of showing each item at the moment it showed it. With those recorded, offline estimates can be reweighted to correct for the selection, the same inverse-weighting idea used for observational data in causal inference. It has to be decided before you need it, because propensities cannot be reconstructed afterwards.
  • experiment on the policy, not the model - randomise which ranking system a user gets and measure the outcome you care about. It is the only method that measures the system as deployed rather than a model in isolation.

The property underneath all of it

Two offsets carry two thirds of the gain in accuracy, so most of what a recommender knows is "who rates generously" and "what is good", not "who likes what". The error is worst for the users whose experience decides whether they stay. Ranking concentrates without being told to. And a system trained on its own output goes blind in the region it stopped showing while every metric it has continues to look fine.

All four follow from the same thing: the model's output determines its next training set. Once that is true, accuracy on logged data stops being a measurement of quality and becomes a measurement of habit - and the only way back out is to deliberately break the loop, by exploring, by recording why each choice was made, or by testing the policy itself.

References & further reading

  • Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
  • Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
4 min readTime Series

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

StatisticsMachine Learning
7 min readAnomaly Detection

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Machine LearningStatistics
← Back to all articles