paidAffiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / The two arms
Two versions at once, and the visit that decides between them
How the two arms of an experiment are built
An experiment is not a launch. It is one screen shown in two versions, at the same time, to two groups chosen at random, so the only difference between the groups is the screen. Everything the method has to defend is in that sentence: one change, a random split, and a sample big enough that the two rates can be told apart.
Desk spec
- arms
- 2
- visitors per arm
- 12,000
- control rate
- 4.30%
- interval
- +/-0.34 point
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answerA gambling interface is tested by showing two versions at once: the control arm keeps the screen as it was and the variant arm changes one thing, and visitors are assigned to an arm at random so the two groups differ only by the screen. On the samples the control converts 4.30% and the variant 5.16%, a lift of 0.86 point, or 20.0% relative, against an interval of +/-0.34 of a point.
One change, two arms, one random split
The method only works if one thing changes. If the variant moves the deposit button, shortens the form and adds a countdown, a better rate cannot be attributed to any of the three, and the result is a feeling with a number attached. The samples' variant moves one thing: where the deposit step sits in the flow.
Sample A - what each arm is, and what is held constant
| Element | Control arm | Variant arm |
| the screen | the version already running | the version being tried |
| what changed | nothing | the position of one step |
| who is in it | 50% of arriving visitors, at random | the other 50%, at random |
| the rate measured | 4.30% | 5.16% |
| the claim the method allows | the difference between the two rates, at this sample size, and nothing about why |
sample A - why the split must be random
control conversions = 516 of 12,000 = 4.30%
variant conversions = 619 of 12,000 = 5.16%
if the two groups were not chosen at random, an unequal mix of
new and returning visitors could produce the same 0.86 point
on its own:
a group that is 60% returning, converting at 2.00% and 40% new
converting at 7.75%, gives 0.6 x 2.00 + 0.4 x 7.75 = 4.30%
the same group the other way round, 40% returning and 60% new,
gives 0.4 x 2.00 + 0.6 x 7.75 = 5.45%
so the random split is what makes the 0.86 point a fact about
the screen rather than a fact about who happened to arrive.
The interval is the part that is usually missing
A rate measured on 12,000 people is not the true rate; it is an estimate with a range. At a 4.30% base and this sample, the range is about +/-0.34 point, so the honest reading of the control arm is "between 3.96% and 4.64%" and of the variant "between 4.82% and 5.50%". The two ranges do not overlap, which is what the interval is for.
Sample A - the two arms with their intervals, which is the only honest way to read them
| Arm | Point estimate | Interval at 95% | What it means |
| control | 4.30% | 3.96% to 4.64% | the range the true control rate most likely sits in |
| variant | 5.16% | 4.82% to 5.50% | the same range for the variant |
| does the gap survive? | 0.86 | two ranges, no overlap | a gap larger than the noise, at this sample size |
sample A - the lift against its own interval
lift = 0.86 point
interval on the lift = +/-0.34 point
so the lift is somewhere between 0.52 and 1.20 point
and the lower end of that range is
0.52 / 4.30 = 12.1% relative, still positive
a lift of 0.20 point would have been inside the interval
and would have meant "no result yet", not "a small win"
so the size of the interval decides what counts as an answer.
The method described here is generic and its numbers are invented. A real operator's split, its eligibility rules and the length of its tests are its own, and a reader cannot see them from outside. What a reader can see is the consequence: two visitors on the same day can be shown two different pages.
Before believing any before-and-after claim about a site
- Ask whether two versions ran at the same time, or whether the site was simply changed on a date.
- Ask how many people were in each group, because a rate on 200 people is not a rate.
- Ask what exactly changed, because three changes cannot be attributed to any one of them.
- Ask for the interval, not just the two rates, and check that it is narrower than the gap.
- Ask how long it ran, because a test stopped on the first good day measures a day.
Read next