Statistical Significance in A/B Testing, Explained
A plain-English deep dive into statistical significance: what it means, p-values and confidence, frequentist vs Bayesian, and the mistakes that make most "winning" tests fake.


Most "winning" A/B tests are not real. A button turns blue, conversions tick up 8%, someone screenshots it, the variant ships, and three months later revenue has not moved. The number was noise wearing the costume of a result.
Statistical significance is the tool that separates a real effect from random luck. This guide explains it properly - what it means, how to read it, and the traps that fool even experienced teams.
What an A/B test actually does
An A/B test splits your visitors randomly into two groups. One sees the current version (the control), the other sees your change (the variant). Because the split is random, the two groups are alike in every way except the change - so any difference in conversion is either caused by the change or by chance.
The entire job of statistics here is to estimate how much of the observed gap is "the change" and how much is "chance."
The coin-flip problem
Flip a fair coin 10 times and you might get 7 heads. The coin is not biased - a small sample is just noisy. A/B tests have the same problem. Two identical pages, shown to two random groups, will almost never convert at exactly the same rate. The difference you see is part signal, part noise.
Statistical significance answers one precise question: if the variant were actually no better than the control, how likely is it that you'd see a difference this large by chance alone?
The cleanest way to picture it: each version's true conversion rate is not a single point but a distribution of plausible values. When those distributions overlap a lot, the gap you measured could easily be noise. When they barely overlap, the gap is probably real.
p-value and confidence
That "likelihood it's chance" is the p-value. A p-value of 0.05 means there is a 5% chance you'd see a difference this big from noise alone. Confidence is the flip side: 95% confidence equals a p-value of 0.05.
The 95% threshold is the long-standing convention in conversion optimization. It is not sacred - it is a balance. Demand 99% and you'll miss real wins by being too cautious; accept 90% and you'll ship more false positives. For most ecommerce testing, 95% is the right default; reserve 99% for changes that are risky or expensive to roll back.
Significance tells you a difference is probably real. It does not tell you the difference is large, or that it will last. Those are separate questions - read the confidence interval for the size, and re-validate big wins.
Frequentist vs Bayesian, in one minute
You'll see tools advertise one or the other. Both are legitimate.
- Frequentist (p-values, the classic z-test) asks: "if there were no real difference, how surprising is this data?" Good frequentist tools offer sequential testing, which is designed to stay valid even if you check often.
- Bayesian asks a more intuitive question: "what's the probability the variant beats the control?" - and reports something like "94% to be best."
Pick whichever your tool does well and whichever your team will read correctly. What matters far more than the school of statistics is not abusing it - which brings us to the two big traps.
The two ways significance lies to you
1. Peeking. This is the big one. Significance assumes you look once, at a sample size you set in advance. If you check daily and stop the instant it crosses 95%, you've quietly broken the math - you've given random noise dozens of chances to wander across the line. A test with no real effect will hit "significant" at some point in its run far more often than you'd think.
2. Tiny samples. With a few hundred visitors per variant, even a real 10% uplift won't reach significance, and a fake one will swing wildly. You need enough traffic for the signal to rise above the noise. A sample size calculator tells you how many visitors each variant needs before the result means anything.
A worked example
Control: 20,000 visitors, 1,000 conversions (5.0%). Variant: 20,000 visitors, 1,100 conversions (5.5%). That's a 10% relative lift. Run it through a two-proportion z-test and you get roughly z = 2.24, about 97.5% confidence (p ≈ 0.025) - significant at the 95% level.
Now halve the traffic to 10,000 per side with the same rates. The lift is identical, but confidence drops below 95%, because the smaller sample leaves more room for noise. Same result, different certainty - which is exactly why sample size is destiny.
A pre-launch checklist
- Decide the sample size before you launch, from your baseline rate and the smallest uplift worth detecting.
- Run for full business cycles (one to two weeks) so weekday/weekend and payday effects average out.
- Read the confidence interval, not just the point estimate. "+10%, 95% CI +3% to +17%" is a real finding; "+10%" alone is not.
- Judge significance on the metric that pays you - revenue per visitor, not just conversion rate, so a CR win that drops AOV doesn't fool you.
- Don't peek-and-stop. Set the sample size, run to it, read once.
You can pressure-test any result with our free A/B test significance calculator - it runs the z-test and shows confidence, p-value and the uplift interval.
The honest summary
Significance is not a trophy you unlock; it's a guardrail that stops you betting the roadmap on noise. Size the test up front, run it to completion, read the interval, and judge it on revenue. Do that and your winners start surviving contact with reality - the only test that matters.
Want a team that runs experiments this way on your store? See how our CRO program works.
Related research
Always-Valid Inference in Conversion Testing
A practitioner review of sequential testing methods and the statistical cost of peeking at A/B test results.
Read the research ResearchBayesian vs. Frequentist Inference in Conversion Testing
A practitioner's guide to Bayesian and frequentist A/B testing: what p-values and posteriors really mean, the role of priors, and when each fits CRO.
Read the researchWant a team to run this for you?
See how we help