▤The Test Group experiment running Open the partner account
paid
Affiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / The sample
How many visitors an answer needs, and how looking ruins it

How big the sample has to be

A result is not a result until it is bigger than the noise, and the size of the noise is set by the size of the sample. The samples need 9,550 visitors an arm to see a 0.86-point lift; the test ran 12,000. The harder lesson is that checking the result every morning changes what "significant" means.

Desk spec
lift to detect
0.86 point
required per arm
9,550
run per arm
12,000
days at this traffic
6.4
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answer

Sample size is the number of visitors an answer needs before it can be told apart from noise. On the samples a 0.86-point lift on a 4.30% base needs 9,550 visitors an arm at 95% confidence and 80% power; the test ran 12,000, which is 1.26 times the requirement. Checking the result 20 times during the test raises the chance of a false positive to 64.2%.

Significance and power, in plain terms

Two risks are being balanced. The first is calling a change a win when the difference was chance - the significance level, usually 5%. The second is missing a real change because the sample was too small - the power, usually set at 80%. A small sample fails the second test long before it fails the first, which is why tiny tests usually produce "no result" rather than a wrong one.

Sample G - the sample needed, at four different lifts on a 4.30% base
Lift to detectVisitors an armDays at 3,750 a dayPractical?
2.00 points (46.5% relative)1,7600.9yes, and usually too big a claim to be real
0.86 point (20.0% relative)9,5505.1yes, at this traffic
0.40 point (9.3% relative)44,00023.5a month of traffic for one answer
0.20 point (4.7% relative)176,00093.9not on a single site's own traffic
why small changes never resolven/an/athe sample grows with the square of the inverse of the lift
sample G - the size, the time and the run base rate = 4.30% lift to detect = 0.86 point, or 20.0% relative at 95% confidence and 80% power = 9,550 visitors per arm required daily visitors = 3,750 per arm per day = 3,750 / 2 = 1,875 days required = 9,550 / 1,875 = 5.1 the test actually ran 12,000 an arm: overshoot = 12,000 / 9,550 = 1.26x days run = 12,000 / 1,875 = 6.4 so the operator bought 26% more sample than the question needed, and a smaller lift would have been unaffordable instead.

Why looking every day is a defect

If a test is checked repeatedly and stopped the first time it looks significant, the 5% risk no longer applies once; it applies once per look. Twenty looks at a 5% level is the multiple comparison problem in its most common form.

Sample G - what repeated looks do to the risk of a false positive
Looks during the testChance of at least one false positiveWhat it means
1, decided in advance5.0%the level the method was built for
522.6%1 - 0.95 to the fifth
2064.2%1 - 0.95 to the twentieth
the arithmetic of peeking-the looks are not independent, so the real figure is lower than this simple product - but it is above 5% either way
sample G - the cost of a false positive a false positive ships a change that does nothing the change is then kept, and the next experiment is measured against it, so a wasted ship also moves the baseline if 20% of shipped changes are false positives and each costs 1,612.00 of engineering and 34 days of a slot: wasted slots = 120 x 0.60 x 0.20 = 14.4 a year wasted engineering = 14.4 x 1,612.00 = 23,212.80 and the direction of the ship is not neutral either: a change that looks good by chance is as likely to raise a guardrail as to lower it, and one of the two will be kept.
The sample sizes are computed from the invented rates and standard two-proportion arithmetic, and the peeking figures are the simple independent-look product, labelled as such. A real programme uses a sequential method or a fixed horizon to control this; a reader cannot see which from outside, but the effects of getting it wrong are visible as an interface that keeps changing without improving.
What a reader can infer about the quality of a loop
  • A site whose layout changes constantly is either testing a lot, or stopping tests early.
  • A claim of a small improvement on a small site is usually a claim about noise.
  • A change that is never revisited is a change whose result was never re-checked.
  • A guardrail that is never mentioned is one that probably has no threshold.
  • None of this is verifiable from outside, which is why the controls a reader sets in the account remain the only reliable limit.

Read next