▤The Test Group experiment running Open the partner account
paid
Affiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / Questions
The twelve questions readers ask most

Questions, answered

The questions are the ones a reader actually asks when a page looks different on Tuesday, when a deposit step moves, or when a privacy notice mentions analytics. Each answer is short, and each is drawn from the desk’s ten samples.

Desk spec
questions
12
derived from
10 samples
short answers
6 in the strip
status
explanatory only
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answer

Two versions of a page are shown at once and a rate decides which one stays; the events behind it are written down per interaction and can be replayed; a second number called a guardrail is supposed to catch the harm; and the metric being improved is the operator's, not the reader's. The twelve answers below are drawn from the desk's ten invented samples.

Short answers

Am I in a test? Probably, at 50% of arrivals, and invisibly.
Can they see my screen? By rebuild, not by camera - 2.1% of replays are watched.
Is a small lift small? 0.86 point is 860 extra deposits on 100,000 visitors.
Is the data anonymous? Until an account id joins the events; then it is not.
What stops a bad ship? A guardrail with a threshold - 3 of 40 records had one.
What can I do? Use the account controls and keep your own dated record.
samples A and H - the one calculation behind the commonest question "am I being measured?" depends on two things at once: the arm you are in: 50% of arriving visitors are in the variant the consent you gave: 68.4% of visitors are measured at all so a variant-served reader who accepted analytics is both in the test and in the log: 0.50 x 0.684 = 34.2% of all visitors a variant-served reader who declined is in the test and not in the log: 0.50 x 0.316 = 15.8% and their behaviour is invisible to the dashboard reading 5.10% against a population rate of 4.22%, a 20.9% overstatement.
The answers are general and the figures are invented. A specific operator's metrics, guardrails, vendors and retention periods are facts about that operator, answerable in writing and in its own documents; this page describes the machinery and names nobody.
Three questions worth putting in writing to an operator
  • Which metric is the deposit page tested against, and what is the guardrail beside it?
  • Is a session replay of my account recorded, what is masked in it, and for how long is it kept?
  • Which processors receive my events, and does a deletion request reach each of them?

Read next

The questions, in full

q01

Why does the deposit page look different on different days?

Because two versions are being shown at the same time and you have been assigned to one of them. On the samples the control converts 4.30% of visitors and the variant 5.16%, a lift of 0.86 of a point, and the version with the better rate is kept.

q02

Am I in a test group?

If a test is running and you are eligible, yes - one of two arms, chosen at random when you arrived. The assignment is normally invisible: no notice is shown, and the record of which arm you were in sits in the operator's own experiment log rather than on your account.

q03

Can the site see my screen?

Partially, and by rebuild rather than by camera. A session replay is redrawn from the event stream, so it contains the pointer path, the scrolls, the clicks, the pauses and whichever form fields are not masked. On the samples 2.1% of recordings are ever watched.

q04

What is an event, in one sentence?

One interaction plus its context - a name, a time, a session, an account, a device and the state of the screen - written down and kept. On the samples one 14-minute session produces 1,240 of them, of which 212 are deliberate actions.

q05

Why does the site record so much and watch so little?

Storage is cheap and attention is not. On the samples 10,000 replays a week weigh 34.0 GB and about 378 minutes of them, 0.27% of what was kept, are ever looked at - but a recording kept in full can be searched later, and that is the reason to keep it.

q06

What is a guardrail metric?

A second number a change is not allowed to move, meant to catch the harm the primary metric cannot see. On the samples the guardrail was the mean first-week net loss per depositing reader, and it rose 11.8% against a rule that said 5%, and the change shipped anyway.

q07

Is a small improvement worth anything?

It is worth a lot of people. A 0.86-point lift on 100,000 monthly visitors is 860 extra first deposits, 34,400.00 of deposits and about 29,412.00 of net revenue over 90 days - and about 6,192.00 of extra first-week loss for the readers it worked on.

q08

Why did a test that looked good turn out not to work?

Because a result measured inside a test is an estimate with a horizon. On the samples 41 of 120 shipped changes failed to reproduce, 16 of them because the effect faded, 11 because it helped one group and hurt another, and 3 because it was found by looking often enough.

q09

How big does a test have to be to mean anything?

Big enough that the interval is narrower than the effect. On the samples a 0.86-point lift on a 4.30% base needs 9,550 visitors an arm at 95% confidence and 80% power, which is 6.4 days at 3,750 visitors a day; the test ran 12,000 an arm.

q10

Does checking the result every day break it?

It changes what the result means. Twenty looks at a 5% level would give a 64.2% chance of at least one look falling the wrong way if the looks were independent, which they are not - but the risk is above 5% either way, which is why a horizon or a sequential method is used.

q11

If I decline analytics, am I missing from the numbers?

Yes. Declining removes you from the events and therefore from every rate computed from them. On the samples the consenting 68.4% behave at 5.10% and the unmeasured 31.6% at 2.30%, so the population rate is 4.22% and the dashboard overstates it by 20.9%.

q12

What can a reader actually do about any of this?

Four things that work: use the account-level controls, because a limit is the part no experiment optimises away; keep your own dated figures, because they are the comparable you own; ask in writing which metric the deposit page is tested against; and ask whether a replay of your own session exists and for how long.