Conversion Rate Optimisation
A/B testing is CRO's engine — the method that turns "I think this button would convert better" into "this button converts 12% better, proven" — and it's where most CRO efforts go wrong, not from not testing but from testing badly: calling winners too early, testing trivia, or misreading noise as signal. Done right, A/B testing is how you know rather than guess. Here are the fundamentals: what it is, how to run one properly, and the discipline that separates real results from wishful thinking.
What A/B testing is
The mechanic: show version A (the control) to half your visitors and version B (the variant, with one change) to the other half, measure which converts better, and — the crucial part — determine whether the difference is real (statistically likely to hold) or noise (random variation that would vanish on re-test). The "one change" matters: A/B tests isolate a single variable (this headline vs that headline) so you know what caused the difference — testing multiple changes at once (multiple variables) is multivariate testing, a different tool for a different question. A/B testing's power is causal clarity: change one thing, measure the effect, know what worked.
Running one properly
- Test a real hypothesis, not trivia: the change should come from research (a diagnosed leak, a hypothesised fix) and matter enough to move the needle — testing button-shades on a page whose real problem is the value proposition is optimising the wrong thing (the famous button-colour tests are mostly noise on trivial changes; test what your research says matters).
- One variable, clean: change one thing so the result is attributable; the multi-change "test" tells you something changed but not what.
- Run to significance, not to impatience: the single biggest A/B mistake — calling a winner after a few days because B is "ahead," when the sample is too small to be real (early leads reverse constantly). The test runs until it reaches statistical significance and a full business cycle (a week+ to capture day-of-week variation) — the discipline that separates real winners from noise-declared-victory.
- Measure the right metric: the actual conversion/revenue outcome, not a proxy — a variant that wins clicks but loses sales isn't a winner (the revenue-not-vanity rule applied to tests).
The discipline that makes results real
The mindset separating rigorous A/B testing from theatre: expect most tests to lose or draw (the honest hit rate — most changes don't move conversion, and a testing program that "always wins" is either testing trivia or misreading noise); respect statistical reality (small samples can't prove anything — below the traffic for valid tests, A/B testing isn't the right tool, and forcing it produces confident-wrong conclusions, per the significance rules); don't p-hack (stopping when you like the number, testing until something "wins" by chance — the statistical sins that manufacture false results); and document and iterate (each test, win or lose, informs the next — the roadmap of hypotheses tested in priority order). Real A/B testing is patient, rigorous and often humbling — which is exactly why its results are trustworthy where opinion isn't.
Frequently asked questions
How long should an A/B test run?
Until it reaches statistical significance and covers a full business cycle (at least a week, to capture day-of-week patterns) — not until it "looks like" a winner (early leads are usually noise that reverses). The exact duration depends on your traffic and effect size; the rule is significance-and-a-cycle, never impatience.
What should I test first?
What your research says matters most — the diagnosed leak with the biggest potential impact (the value prop, the checkout friction, the landing page), prioritised by the roadmap. Not button colours or trivia — test the changes big enough to move the needle, sourced from real research, not the "10 things to test" listicles.
Can I A/B test with low traffic?
Valid A/B testing needs enough traffic to reach significance in reasonable time — below that, tests take forever or never conclude, and forcing conclusions produces false results. Low-traffic sites shift to qualitative research and evidence-backed best practices instead (ship what's proven, measure before/after) — A/B testing is one CRO tool, right at sufficient traffic, wrong below it, on traffic content and authority provide (our lane).