The Test Group / The numbers
Every figure on the desk, in one place
Every figure in one place
This page is the desk's own index: the same ten samples, the same derived figures, with the arithmetic for each one shown so it can be re-derived rather than believed. Nothing on it is observed from a real operator, and every number is labelled illustrative.
Desk spec
- samples
- 10 invented
- pages indexed
- 18
- figures listed
- 11 groups
- source
- samples A to J
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answerThe desk's figures come from ten invented samples. The two arms convert 4.30% and 5.16%, a lift of 0.86 point or 20.0% relative. The funnel turns 100,000 visits into 2,074 first bets, 2.1% end to end. A year of 120 shipped changes sums to 22.4% and moves the holdout by 3.6%. One change is worth 860 extra first deposits, 29,412.00 of revenue and 6,192.00 of reader loss.
The eleven groups of figures
4.30%control first-deposit rate (sample A)
5.16%variant rate, a 0.86-point lift (sample A)
1,240events in one 14-minute session (sample B)
535.7 GBevents stored in a month (sample B)
2.1%end to end through the funnel (sample C)
145.9 GBreplays held at a 30-day retention (sample D)
3.6%realised effect against the holdout (sample E)
11.8%guardrail movement that shipped anyway (sample F)
9,550visitors an arm the test needed (sample G)
20.9%how far the dashboard overstates the real rate (sample H)
29,412.00revenue one experiment returns in 90 days (sample J)
The arms
The experiment's own numbers, both arms and the interval, with the lift worked out.
Sample A - the two arms and the lift
| Figure | Value | Derivation |
| control rate | 4.30% | 516 / 12,000 |
| variant rate | 5.16% | 619 / 12,000 |
| lift in points | 0.86 | 5.16 - 4.30 |
| lift in proportion | 20.0% | 0.86 / 4.30 |
| interval | +/-0.34 point | at 95% on 12,000 an arm |
| what the interval allows | 0.52 to 1.20 | a lower bound of 12.1% relative |
sample A - the lift, re-derived
control = 516 / 12,000 = 4.30%
variant = 619 / 12,000 = 5.16%
lift = 0.86 point = 20.0% relative
interval = +/-0.34 point, so the lift is between 0.52 and 1.20
and the lower end is 12.1% relative, still positive.
The event log and the funnel
The data underneath everything else: what one session writes down, and what happens to 100,000 of them.
Samples B and C - the log and the funnel
| Figure | Value | Derivation |
| events in one session | 1,240 | 640 scroll + 388 pointer + 186 click + 22 view + 4 fields |
| one session | 290.6 KB | 1,240 x 240 bytes |
| a month | 535.7 GB | 290.6 KB x 1,800,000 sessions |
| deliberate events | 17.1% | 212 of 1,240 |
| visits to first bets | 100,000 to 2,074 | six steps |
| end to end | 2.1% | 2,074 / 100,000 |
| lost before the deposit step | 76.0% | 76,000 of 100,000 never created an account |
samples B and C - the log and the funnel, re-derived
one session = 1,240 x 240 = 297,600 bytes = 290.6 KB
a month = 297,600 x 1,800,000 = 535,680,000,000 = 535.7 GB
deliberate = (186 + 22 + 4) / 1,240 = 212 / 1,240 = 17.1%
funnel = 24,000 / 100,000 = 24.0%; 9,600 / 24,000 = 40.0%
4,320 / 9,600 = 45.0%; 3,456 / 4,320 = 80.0%
2,074 / 3,456 = 60.0%; 2,074 / 100,000 = 2.1%
The replay and the holdout
The two instruments that are meant to check the loop: one to see the session, one to see the year.
Samples D and E - the replay and the holdout
| Figure | Value | Derivation |
| a week of recordings | 34.0 GB | 10,000 x 3.4 MB |
| held at 30 days | 145.9 GB | 34.0 x 4.29 weeks |
| ever watched | 2.1% | 210 of 10,000 |
| mean watched | 12.9% | 1.8 of 14 minutes |
| holdout | 50,000 | 5% of 1,000,000 |
| realised effect | 3.6% | 42.76 against 41.20 |
| gap to the scoreboard | 6.2 times | 22.4% against 3.6% |
| failed to reproduce | 34.2% | 41 of 120 |
| the check almost nobody runs | - | 0 of 40 records published |
samples D and E - the two checks, re-derived
a week = 10,000 x 3.4 MB = 34.0 GB
at 30 days = 34.0 x (30 / 7) = 34.0 x 4.29 = 145.9 GB
watched = 210 / 10,000 = 2.1%
holdout realised effect = (42.76 - 41.20) / 41.20 = 1.56 / 41.20
= 3.6%
overstatement = 22.4 / 3.6 = 6.2x
reproduce rate = 41 / 120 = 34.2% failed
The metric, the sample and the consent line
The three things that decide what a result means, and who it is a result about.
Samples F, G and H - the metric, the sample and the consent gap
| Figure | Value | Derivation |
| primary metric | 4.30% to 5.16% | the fleet-deposit rate, both arms |
| guardrail | 61.20 to 68.40 | +7.20 = +11.8% |
| rule set in advance | 5% | and it shipped at 11.8% |
| sample required | 9,550 | 0.86 point on a 4.30% base, 95% / 80% |
| sample run | 12,000 | 1.26 times the requirement, 6.4 days |
| peeking risk, 20 looks | 64.2% | 1 - 0.95 to the twentieth |
| accept analytics | 68.4% | and behave at 5.10% |
| unmeasured | 31.6% | assumed at 2.30% |
| population rate | 4.22% | 0.684 x 5.10 + 0.316 x 2.30 |
| overstatement | 20.9% | 5.10 / 4.22 |
| the rate a dashboard shows | 5.10% | against a population rate of 4.22% |
samples F, G and H - the three, re-derived
guardrail = (68.40 - 61.20) / 61.20 = 7.20 / 61.20 = 11.8%
sample = 0.86 point on 4.30% at 95% / 80% = 9,550 an arm
days = 12,000 / 1,875 = 6.4
peeking = 1 - 0.95 to the 20th = 1 - 0.3585 = 64.2%
population = 0.684 x 5.10 + 0.316 x 2.30 = 4.2152 = 4.22%
overstated = 5.10 / 4.2152 = 1.2099 = 20.9%
sample J - the money behind the figures
lift = 0.86 point on 100,000 monthly visitors
= 860 extra first deposits
operator, over 90 days = 860 x 34.20 = 29,412.00
reader, in the first week= 860 x 7.20 = 6,192.00
programme cost a year = 120 x 1,612.00 + 57,600.00 = 251,040.00
return on that cost = 1,411,776.00 / 251,040.00 = 5.6x
and the two figures stay the same order of magnitude:
29,412.00 of revenue against 6,192.00 of reader loss,
both produced by one 0.86-point move on one screen.
Every figure on this page is invented and derived from samples A to J, exactly as the arithmetic above shows. No real operator, vendor, product or experiment is described, and no live rate is reproduced. The page exists so that a reader can check any other page on the desk against the same ten samples rather than take a number on trust.
Read next