▤The Test Group experiment running Open the partner account
paid
Affiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / The two arms
Two versions at once, and the visit that decides between them

How the two arms of an experiment are built

An experiment is not a launch. It is one screen shown in two versions, at the same time, to two groups chosen at random, so the only difference between the groups is the screen. Everything the method has to defend is in that sentence: one change, a random split, and a sample big enough that the two rates can be told apart.

Desk spec
arms
2
visitors per arm
12,000
control rate
4.30%
interval
+/-0.34 point
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answer

A gambling interface is tested by showing two versions at once: the control arm keeps the screen as it was and the variant arm changes one thing, and visitors are assigned to an arm at random so the two groups differ only by the screen. On the samples the control converts 4.30% and the variant 5.16%, a lift of 0.86 point, or 20.0% relative, against an interval of +/-0.34 of a point.

One change, two arms, one random split

The method only works if one thing changes. If the variant moves the deposit button, shortens the form and adds a countdown, a better rate cannot be attributed to any of the three, and the result is a feeling with a number attached. The samples' variant moves one thing: where the deposit step sits in the flow.

Sample A - what each arm is, and what is held constant
ElementControl armVariant arm
the screenthe version already runningthe version being tried
what changednothingthe position of one step
who is in it50% of arriving visitors, at randomthe other 50%, at random
the rate measured4.30%5.16%
the claim the method allowsthe difference between the two rates, at this sample size, and nothing about why
sample A - why the split must be random control conversions = 516 of 12,000 = 4.30% variant conversions = 619 of 12,000 = 5.16% if the two groups were not chosen at random, an unequal mix of new and returning visitors could produce the same 0.86 point on its own: a group that is 60% returning, converting at 2.00% and 40% new converting at 7.75%, gives 0.6 x 2.00 + 0.4 x 7.75 = 4.30% the same group the other way round, 40% returning and 60% new, gives 0.4 x 2.00 + 0.6 x 7.75 = 5.45% so the random split is what makes the 0.86 point a fact about the screen rather than a fact about who happened to arrive.

The interval is the part that is usually missing

A rate measured on 12,000 people is not the true rate; it is an estimate with a range. At a 4.30% base and this sample, the range is about +/-0.34 point, so the honest reading of the control arm is "between 3.96% and 4.64%" and of the variant "between 4.82% and 5.50%". The two ranges do not overlap, which is what the interval is for.

Sample A - the two arms with their intervals, which is the only honest way to read them
ArmPoint estimateInterval at 95%What it means
control4.30%3.96% to 4.64%the range the true control rate most likely sits in
variant5.16%4.82% to 5.50%the same range for the variant
does the gap survive?0.86two ranges, no overlapa gap larger than the noise, at this sample size
sample A - the lift against its own interval lift = 0.86 point interval on the lift = +/-0.34 point so the lift is somewhere between 0.52 and 1.20 point and the lower end of that range is 0.52 / 4.30 = 12.1% relative, still positive a lift of 0.20 point would have been inside the interval and would have meant "no result yet", not "a small win" so the size of the interval decides what counts as an answer.
The method described here is generic and its numbers are invented. A real operator's split, its eligibility rules and the length of its tests are its own, and a reader cannot see them from outside. What a reader can see is the consequence: two visitors on the same day can be shown two different pages.
Before believing any before-and-after claim about a site
  • Ask whether two versions ran at the same time, or whether the site was simply changed on a date.
  • Ask how many people were in each group, because a rate on 200 people is not a rate.
  • Ask what exactly changed, because three changes cannot be attributed to any one of them.
  • Ask for the interval, not just the two rates, and check that it is narrower than the gap.
  • Ask how long it ran, because a test stopped on the first good day measures a day.

Read next