How big the sample has to be
A result is not a result until it is bigger than the noise, and the size of the noise is set by the size of the sample. The samples need 9,550 visitors an arm to see a 0.86-point lift; the test ran 12,000. The harder lesson is that checking the result every morning changes what "significant" means.
- lift to detect
- 0.86 point
- required per arm
- 9,550
- run per arm
- 12,000
- days at this traffic
- 6.4
Sample size is the number of visitors an answer needs before it can be told apart from noise. On the samples a 0.86-point lift on a 4.30% base needs 9,550 visitors an arm at 95% confidence and 80% power; the test ran 12,000, which is 1.26 times the requirement. Checking the result 20 times during the test raises the chance of a false positive to 64.2%.
Significance and power, in plain terms
Two risks are being balanced. The first is calling a change a win when the difference was chance - the significance level, usually 5%. The second is missing a real change because the sample was too small - the power, usually set at 80%. A small sample fails the second test long before it fails the first, which is why tiny tests usually produce "no result" rather than a wrong one.
| Lift to detect | Visitors an arm | Days at 3,750 a day | Practical? |
|---|---|---|---|
| 2.00 points (46.5% relative) | 1,760 | 0.9 | yes, and usually too big a claim to be real |
| 0.86 point (20.0% relative) | 9,550 | 5.1 | yes, at this traffic |
| 0.40 point (9.3% relative) | 44,000 | 23.5 | a month of traffic for one answer |
| 0.20 point (4.7% relative) | 176,000 | 93.9 | not on a single site's own traffic |
| why small changes never resolve | n/a | n/a | the sample grows with the square of the inverse of the lift |
Why looking every day is a defect
If a test is checked repeatedly and stopped the first time it looks significant, the 5% risk no longer applies once; it applies once per look. Twenty looks at a 5% level is the multiple comparison problem in its most common form.
| Looks during the test | Chance of at least one false positive | What it means |
|---|---|---|
| 1, decided in advance | 5.0% | the level the method was built for |
| 5 | 22.6% | 1 - 0.95 to the fifth |
| 20 | 64.2% | 1 - 0.95 to the twentieth |
| the arithmetic of peeking | - | the looks are not independent, so the real figure is lower than this simple product - but it is above 5% either way |
- A site whose layout changes constantly is either testing a lot, or stopping tests early.
- A claim of a small improvement on a small site is usually a claim about noise.
- A change that is never revisited is a change whose result was never re-checked.
- A guardrail that is never mentioned is one that probably has no threshold.
- None of this is verifiable from outside, which is why the controls a reader sets in the account remain the only reliable limit.