Questions, answered
The questions are the ones a reader actually asks when a page looks different on Tuesday, when a deposit step moves, or when a privacy notice mentions analytics. Each answer is short, and each is drawn from the desk’s ten samples.
- questions
- 12
- derived from
- 10 samples
- short answers
- 6 in the strip
- status
- explanatory only
Two versions of a page are shown at once and a rate decides which one stays; the events behind it are written down per interaction and can be replayed; a second number called a guardrail is supposed to catch the harm; and the metric being improved is the operator's, not the reader's. The twelve answers below are drawn from the desk's ten invented samples.
Short answers
- Which metric is the deposit page tested against, and what is the guardrail beside it?
- Is a session replay of my account recorded, what is masked in it, and for how long is it kept?
- Which processors receive my events, and does a deletion request reach each of them?
Read next
The questions, in full
Why does the deposit page look different on different days?
Because two versions are being shown at the same time and you have been assigned to one of them. On the samples the control converts 4.30% of visitors and the variant 5.16%, a lift of 0.86 of a point, and the version with the better rate is kept.
Am I in a test group?
If a test is running and you are eligible, yes - one of two arms, chosen at random when you arrived. The assignment is normally invisible: no notice is shown, and the record of which arm you were in sits in the operator's own experiment log rather than on your account.
Can the site see my screen?
Partially, and by rebuild rather than by camera. A session replay is redrawn from the event stream, so it contains the pointer path, the scrolls, the clicks, the pauses and whichever form fields are not masked. On the samples 2.1% of recordings are ever watched.
What is an event, in one sentence?
One interaction plus its context - a name, a time, a session, an account, a device and the state of the screen - written down and kept. On the samples one 14-minute session produces 1,240 of them, of which 212 are deliberate actions.
Why does the site record so much and watch so little?
Storage is cheap and attention is not. On the samples 10,000 replays a week weigh 34.0 GB and about 378 minutes of them, 0.27% of what was kept, are ever looked at - but a recording kept in full can be searched later, and that is the reason to keep it.
What is a guardrail metric?
A second number a change is not allowed to move, meant to catch the harm the primary metric cannot see. On the samples the guardrail was the mean first-week net loss per depositing reader, and it rose 11.8% against a rule that said 5%, and the change shipped anyway.
Is a small improvement worth anything?
It is worth a lot of people. A 0.86-point lift on 100,000 monthly visitors is 860 extra first deposits, 34,400.00 of deposits and about 29,412.00 of net revenue over 90 days - and about 6,192.00 of extra first-week loss for the readers it worked on.
Why did a test that looked good turn out not to work?
Because a result measured inside a test is an estimate with a horizon. On the samples 41 of 120 shipped changes failed to reproduce, 16 of them because the effect faded, 11 because it helped one group and hurt another, and 3 because it was found by looking often enough.
How big does a test have to be to mean anything?
Big enough that the interval is narrower than the effect. On the samples a 0.86-point lift on a 4.30% base needs 9,550 visitors an arm at 95% confidence and 80% power, which is 6.4 days at 3,750 visitors a day; the test ran 12,000 an arm.
Does checking the result every day break it?
It changes what the result means. Twenty looks at a 5% level would give a 64.2% chance of at least one look falling the wrong way if the looks were independent, which they are not - but the risk is above 5% either way, which is why a horizon or a sequential method is used.
If I decline analytics, am I missing from the numbers?
Yes. Declining removes you from the events and therefore from every rate computed from them. On the samples the consenting 68.4% behave at 5.10% and the unmeasured 31.6% at 2.30%, so the population rate is 4.22% and the dashboard overstates it by 20.9%.
What can a reader actually do about any of this?
Four things that work: use the account-level controls, because a limit is the part no experiment optimises away; keep your own dated figures, because they are the comparable you own; ask in writing which metric the deposit page is tested against; and ask whether a replay of your own session exists and for how long.