▤The Test Group experiment running Open the partner account
paid
Affiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / The holdout
The group that never receives the improvement

The holdout, and what it reveals a year later

Every experiment that ships is a small win, and small wins are supposed to add up. A holdout is the group deliberately kept on the old version, and its only purpose is to check whether they did. On the samples 120 shipped changes summed to 22.4% and moved the holdout gap by 3.6%.

Desk spec
holdout
5%
users a month
1,000,000
experiments a year
120
realised effect
3.6%
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answer

A holdout is a slice of visitors kept on the old version while every improvement is shipped to the rest. It exists to measure the accumulated effect rather than the individual one. On the samples 120 shipped experiments had individual lifts summing to 22.4%, while the holdout, a year later, sat 3.6% behind - the difference between an estimate and an outcome.

Why individual wins do not add up

Each lift is measured against the version that ran beside it, which is a moving baseline. A change can win against last month's screen and lose against the screen that will exist next month, and a change can win on the metric and be undone by the next one. The holdout is the only instrument that sees the whole series at once.

Sample E - twelve months of shipping, measured two ways
MeasureValueWhat it is
experiments shipped120each with its own measured lift at the time
sum of the individual lifts22.4%what a scoreboard of wins shows
holdout users50,0005% of 1,000,000, kept on the old version
realised difference after a year3.6%treated 42.76 per user against holdout 41.20
experiments that did not reproduce41 of 12034.2% failed a later check
how much the sum overstated the outcome6.2 times22.4 / 3.6
sample E - the scoreboard against the holdout individual lifts, added up = 22.4% realised effect, measured once = 3.6% overstatement = 22.4 / 3.6 = 6.2x per user per month: treated users spend = 42.76 holdout users spend = 41.20 difference = 1.56, or 3.6% experiments that failed to reproduce = 41 / 120 = 34.2% so one shipped change in three is later found not to be doing what it was measured to do, and the scoreboard cannot see it.

What makes a result fail to reproduce

Four of the reasons are ordinary and one is uncomfortable. The ordinary ones are a novelty effect that wears off, a change that helped one segment and hurt another, a metric that drifted, and a measure taken in a different season. The uncomfortable one is that a result found by looking often enough is a result about the looking.

Sample E - where the 41 failures came from, on the samples
ReasonExperimentsShare of the 41
the effect faded after a few weeks1639.0%
it helped one segment and hurt another1126.8%
another change later removed its benefit717.1%
the season was different on the second look49.8%
it was found by looking repeatedly37.3%
failed to reproduce41100%
sample E - the holdout as a cost, not a courtesy holdout share = 5% users a month = 1,000,000 users deliberately held back = 50,000 revenue per user a month = 42.00 revenue given up on the holdout = 50,000 x 42.00 x 0.036 = 75,600.00 a month that is the price of being able to tell an estimate from an outcome, and it is why so few sites keep a holdout at all.
The figures are invented. Whether a named operator runs a holdout, for how long and at what share is not something a reader can see from outside; the desk's point is structural, that a series of individually measured wins is not the same thing as a measured improvement, and only a group left alone can show the difference.
Questions that separate a measured improvement from a scoreboard
  • Ask whether the site keeps a group that never receives changes, and at what share.
  • Ask what the accumulated effect is, not how many experiments shipped.
  • Ask how many past results were re-checked, and how many survived.
  • Ask whether a lift is still present a month after the test ended.
  • Ask whether the metric that improved is the one the reader would have chosen.

Read next