▤The Test Group experiment running Open the partner account
paid
Affiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / Myths
Six beliefs about the loop, checked against the samples

Six beliefs, checked

The loop is understood badly in both directions: by readers who think the changes are arbitrary, and by readers who think a change that helps the reader is the point. Six beliefs, each checked against the ten samples, with four false, one partly true and one that holds.

Desk spec
beliefs checked
6
false
4
partly true
1
true
1
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answer

Four of the six common beliefs about interface testing are false, one is partly true and one holds. Layouts are not random - they are two versions shown at once, a control at 4.30% and a variant at 5.16%. The data is not anonymous once an account id joins the events. And a change is shipped on the operator's metric, not on the reader's outcome.

The six beliefs

false

"The layout changes at random." It changes between two versions shown at the same time. The control converts 4.30% of its visitors and the variant 5.16%, and the change is kept if the lift survives its interval - 0.86 point against +/-0.34. Randomness is in the assignment, not in the design.

false

"The data collected about me is anonymous." It is anonymous until an account id joins it, and on the samples the account field is attached after sign-in to 1,240 events of one session. A session replay of the same events shown to a human is not an anonymous record; it is a recording of a person's screen.

false

"If a change makes things easier for me it helped me." The same change lifted first deposits by 20.0% relative and raised the mean first-week net loss of the readers it worked on by 11.8%, from 61.20 to 68.40. Easier for the reader and better for the reader are two different claims.

false

"A small improvement is a small change." A 0.86-point lift on 100,000 monthly visitors is 860 extra first deposits, 34,400.00 of deposits and 29,412.00 of revenue over 90 days. A rate move that reads as trivial is a large number of people as soon as it meets real traffic.

partly true

"They test everything, all the time." On the samples 120 changes shipped in a year, but 41 of 120 failed to reproduce, 11 of 24 addressed a cause outside the page and only 3 of 40 records named a guardrail. Something is running constantly; whether it is measuring anything is a separate question.

true

"The metric is the operator's, not mine." This one holds. The primary metric is the first-deposit rate, and the reader's later outcome is a guardrail at best: the guardrail moved 11.8% against a rule of 5% and the change shipped. The direction of a screen is set by the number it is measured on.

samples A, F and J - the two changes in one arithmetic the reader is told: 5.16% is better than 4.30% the operator measures: 0.86 point = 20.0% relative the reader is not told: the same group's mean first-week net loss rose 61.20 -> 68.40 = +11.8%, against a 5% rule turned into money for one month: operator = 860 x 34.20 = 29,412.00 over 90 days readers = 860 x 7.20 = 6,192.00 in the first week so the belief that "easier means better" is not a small misunderstanding: it is the whole difference between the two columns, and only one of them is on the dashboard.

How the six were chosen

Each belief was taken from the way the subject is usually described in public - by readers, by the trade press and by the operators themselves - and each was then tested against one of the desk's ten samples. A belief was marked false when the samples contradicted it outright, partly true when the direction was right and the scale was wrong, and true when the samples supported it as stated.

The beliefs are the ordinary ones heard about how a site is run; the answers are checked only against the desk's invented samples. Whether a named operator does any of this is a fact about that operator, and the desk names none.
What follows if any of this is true of a site you use
  • Assume a change you notice was measured, and that the measurement was of the operator's number.
  • Assume a recording may exist if the privacy notice names analytics at all.
  • Use the account-level controls for a limit, because they are the part no experiment is allowed to optimise away.
  • Keep your own dates and figures, since they are the only comparable you can hold.
  • Ask, in writing, which metric the deposit page is tested against, and keep the reply.

Read next