The holdout, and what it reveals a year later
Every experiment that ships is a small win, and small wins are supposed to add up. A holdout is the group deliberately kept on the old version, and its only purpose is to check whether they did. On the samples 120 shipped changes summed to 22.4% and moved the holdout gap by 3.6%.
- holdout
- 5%
- users a month
- 1,000,000
- experiments a year
- 120
- realised effect
- 3.6%
A holdout is a slice of visitors kept on the old version while every improvement is shipped to the rest. It exists to measure the accumulated effect rather than the individual one. On the samples 120 shipped experiments had individual lifts summing to 22.4%, while the holdout, a year later, sat 3.6% behind - the difference between an estimate and an outcome.
Why individual wins do not add up
Each lift is measured against the version that ran beside it, which is a moving baseline. A change can win against last month's screen and lose against the screen that will exist next month, and a change can win on the metric and be undone by the next one. The holdout is the only instrument that sees the whole series at once.
| Measure | Value | What it is |
|---|---|---|
| experiments shipped | 120 | each with its own measured lift at the time |
| sum of the individual lifts | 22.4% | what a scoreboard of wins shows |
| holdout users | 50,000 | 5% of 1,000,000, kept on the old version |
| realised difference after a year | 3.6% | treated 42.76 per user against holdout 41.20 |
| experiments that did not reproduce | 41 of 120 | 34.2% failed a later check |
| how much the sum overstated the outcome | 6.2 times | 22.4 / 3.6 |
What makes a result fail to reproduce
Four of the reasons are ordinary and one is uncomfortable. The ordinary ones are a novelty effect that wears off, a change that helped one segment and hurt another, a metric that drifted, and a measure taken in a different season. The uncomfortable one is that a result found by looking often enough is a result about the looking.
| Reason | Experiments | Share of the 41 |
|---|---|---|
| the effect faded after a few weeks | 16 | 39.0% |
| it helped one segment and hurt another | 11 | 26.8% |
| another change later removed its benefit | 7 | 17.1% |
| the season was different on the second look | 4 | 9.8% |
| it was found by looking repeatedly | 3 | 7.3% |
| failed to reproduce | 41 | 100% |
- Ask whether the site keeps a group that never receives changes, and at what share.
- Ask what the accumulated effect is, not how many experiments shipped.
- Ask how many past results were re-checked, and how many survived.
- Ask whether a lift is still present a month after the test ended.
- Ask whether the metric that improved is the one the reader would have chosen.