paidAffiliate disclosure. The partner link in the masthead and in the band beside the copy on this page is a sponsored link to a partner operator, and this site may be paid if you open an account through it, at no extra cost to you. It carries rel="sponsored noopener" and opens in a new tab. A desk about how a screen is measured and tested should not leave its own funding unsaid: one link funds the site, no operator and no product is named, rated or recommended anywhere on it, and this site runs no analytics of its own on its readers.
The Test Group / The metric
The number that is optimised, and the one that must not fall
The metric and the guardrail
An experiment is not decided by whether the screen is better. It is decided by whether one chosen number went up. A guardrail is the second number that is supposed to stop the first one, and the whole ethical weight of the loop sits on whether the guardrail has the power to stop anything.
Desk spec
- primary metric
- first-deposit rate
- control to variant
- 4.30% to 5.16%
- guardrail moved
- +11.8%
- rule set in advance
- 5%
the eventOne interaction, written down. A click carries a name, a time, an account, a session and what was on the screen, and it is kept whether or not the reader chose to be measured.
the funnelThe order the steps happen in. A landing visit becomes an account, an account becomes a deposit page, and 100,000 visits end in 2,074 first bets - a 2.1% path the whole loop is aimed at.
the armsTwo versions of one screen shown at the same time, and a rate for each. 4.30% against 5.16% is a 0.86-point lift, and the interval decides whether it is a result or a coincidence.
Direct answerThe metric is the one number a test is judged on - on the samples, the first-deposit rate, which rose from 4.30% to 5.16%. A guardrail is a second number the change must not move, here the mean first-week net loss per depositing reader and the complaint rate. The guardrail rose 11.8%, against a rule of 5%, and the change shipped anyway.
Choosing a number chooses a behaviour
The metric is not neutral. A first-deposit rate rewards a screen that makes the step easier, and it is indifferent to what happens afterwards. Every metric has a shadow: the thing it will trade away because it cannot see it.
Sample F - a candidate metric and the behaviour it rewards
| Candidate metric | What it rewards | On the samples |
| first-deposit rate | moving more readers through the deposit step now | 4.30% -> 5.16% |
| deposits per user | repeat deposits by the same reader | 1.72 -> 2.06 |
| first-week net loss per depositing reader | nothing the desk would ship on its own | 61.20 -> 68.40 |
| complaint rate | nothing either | 0.42% -> 0.61% |
| the primary and the two guardrails | the first moved up, and so did the second and the fourth | - |
sample F - the change, measured on four numbers
primary: first-deposit rate 4.30% -> 5.16% = +20.0% relative
secondary: deposits per user 1.72 -> 2.06 = +19.8%
guardrail 1: mean first-week
net loss per depositing reader 61.20 -> 68.40 = +7.20 = +11.8%
guardrail 2: complaint rate 0.42% -> 0.61% = +0.19 point
the decision rule, fixed before the test, was:
ship if the primary rises and no guardrail moves more than 5%
the primary rose and the first guardrail moved 11.8%
so the rule said no and the change was shipped, which is the
outcome the loop is most often criticised for and the one
it records least often.
What a guardrail needs to work
A guardrail only stops a change if it has a threshold fixed in advance, a size big enough to see the movement, and a consequence attached. A guardrail that is reported after the decision, or measured on a sample too small to move, is decoration.
Sample F - the same guardrail measured on three sample sizes
| Sample per arm | Smallest detectable move | Would the 11.8% movement be visible? |
| 2,000 | +/-24.0% | no, the movement is inside the noise |
| 12,000 | +/-9.8% | yes, 11.8% is outside it |
| 40,000 | +/-5.4% | yes, and a 5% rule becomes checkable |
| the sample the rule needs | +/-5% | roughly 40,000 an arm to police a 5% threshold |
samples F and J - who gains and who pays
extra first deposits from the change:
0.86 point x 100,000 monthly visitors = 860
operator side, at 34.20 of net revenue per depositor over 90 days:
860 x 34.20 = 29,412.00
reader side, at 7.20 more net loss in the first week:
860 x 7.20 = 6,192.00
so the change moves 6,192.00 a month from the readers it worked
on to the operator, and the guardrail is the number that was
supposed to make that visible before the ship.
measured on 40,000 an arm it would have been; measured on 2,000
it would have looked like noise.
The rates are invented. What a named operator optimises, what it calls a guardrail and whether its guardrail can stop a ship are internal facts; the desk's purpose is to show that the choice of metric is a decision with a direction, and that a guardrail is only a limit if it is set in advance, sized to see, and allowed to say no.
What a reader can take from a metric, without ever seeing one
- Assume the interface is optimised for the operator's next number, not for the reader's outcome.
- Treat a step that becomes faster or easier as a decision that has been measured, not an accident.
- Use the controls that are meant to slow the decision down, because they are the ones the loop cannot remove.
- Read a change in your own behaviour as information: a page that makes you faster is working as designed.
- Ask a regulator or the operator in writing which metric its interface is tested against; the answer is a fact you are entitled to ask for.
Read next