The Test Group: the interface that is tested on you
The screen you use has two versions. One is the version you were shown, the other is the version the next reader will get, and a count decides which one wins. This desk is about that loop rather than about the screen: what each interaction writes down, how the steps from a visit to a deposit are counted, what an experiment needs before its result means anything, and what a change costs the people it is tested on.
- samples
- 10 invented
- control rate
- 4.30%
- variant rate
- 5.16%
- lift
- 0.86 point
Of 100,000 visits, 2,074 end in a first bet - 2.1% end to end. The largest single loss by count is the 76,000 who never create an account; the largest loss by proportion is the 60.0% who create an account and never open the deposit page, which is the step this desk's whole loop lives on.
six steps; end to end 2,074 / 100,000 = 2.1%; one test can only move one of the six.
The two bars are drawn against the same axis and the top bar is 20.0% longer than the bottom one. That is the whole claim of an experiment, and it is only a claim once the interval is narrower than the gap: at 12,000 visitors an arm and a 4.30% base, a 0.34-point wobble either way would be ordinary noise, and the 0.86-point gap is comfortably outside it.
one test, two arms, 24,000 visitors in total; the gap is real at this sample size and says nothing about why.
A gambling interface is improved by measurement, not by taste: every interaction is written down as an event, the steps from a landing visit to a first deposit are counted as a funnel, and two versions of one screen are shown at the same time so a rate can be compared. On the samples the control arm converts 4.30% of visitors and the variant 5.16% - a 0.86-point lift, or 20.0% relative.
What the samples show
The desk's central instrument is one number produced by two arms of one experiment. The control converts 4.30% of its 12,000 visitors and the variant 5.16% of its 12,000, so the change is worth 0.86 of a point, or 20.0% relative. Read against the interval, the gap is wider than the noise: the 95% interval is +/-0.34 point, which is less than half the measured lift.
Two findings do most of the work. First, the loop optimises the operator's number and not the reader's: the metric on the samples is the first-deposit rate, and the change that lifted it by 20.0% also raised the mean first-week net loss of the readers it worked on by 11.8% (61.20 to 68.40). Second, the loop is blind in a way that is easy to miss: the 31.6% of visitors who decline analytics behave at 2.30% against 5.10% for those who accept, so every dashboard figure is a rate on the measured population and not on the population.
None of the samples describes a real operator, vendor, product or experiment. They are ten invented sets of counts and rates, defined on this page, and every other figure on the site is derived from them.
Ten samples
One experiment, two rates.
- control / variant
- 4.30 / 5.16%
- lift
- 0.86 point
- interval
- +/-0.34
What one session writes down.
- events
- 1,240
- bytes each
- 240
- per session
- 290.6 KB
Landing visit to first bet.
- visits
- 100,000
- first bets
- 2,074
- end to end
- 2.1%
One week of recordings.
- recordings
- 10,000
- each
- 3.4 MB
- watched
- 2.1%
The group that never changes.
- holdout
- 5%
- over a year
- 50,000
- realised effect
- 3.6%
What is optimised, and the guardrail.
- primary
- 4.30 -> 5.16%
- guardrail
- +11.8%
- rule
- 5%
How many visitors an answer needs.
- required
- 9,550
- run
- 12,000
- days
- 6.4
Who is measured and who is not.
- accept
- 68.4%
- decline
- 31.6%
- gap
- 20.9%
Two further samples are defined on the pages that use them: sample I on what the loop cannot fix, and sample J on one experiment carried through to money.
The loop in one table
The clearest place to start is the one thing every question about this subject comes back to: what does a change have to beat before it is shipped?
| Arm | Visitors | First deposits | Rate |
|---|---|---|---|
| Control, the screen as it was | 12,000 | 516 | 4.30% |
| Variant, the screen being tried | 12,000 | 619 | 5.16% |
| Lift of the variant over the control | - | +103 | +0.86 point |