A/B Testing for Ecommerce: How to Run Tests That Reach Significance
A "winner" that never reached significance costs you more than running no test at all. Here is how to run ecommerce A/B tests you can trust, and the mistakes that fake a victory.


A/B testing is how you replace opinions with evidence. Done well, it tells you which changes genuinely grow revenue. Done badly, it produces confident-sounding conclusions that are simply wrong. Here is how to run ecommerce tests you can actually trust.
What an A/B test really is
You split your traffic between two versions: the control (what you have now) and a variant (your proposed change). Visitors are randomly assigned, and you compare a metric, usually conversion rate or revenue per visitor. Because assignment is random and simultaneous, any meaningful difference can be attributed to the change rather than to seasonality or traffic source.
Test one clear hypothesis
A good test starts with a hypothesis: "If we add reviews to the product page, then conversion rate will rise, because social proof reduces purchase anxiety." One change, one expected outcome, one mechanism. If you change five things at once and conversion improves, you have learned nothing about why.
Reaching statistical significance
This is where most ecommerce tests go wrong. Significance is the probability that the result is not just random noise. A few rules keep you honest:
- Decide your sample size up front. Use your baseline conversion rate and the smallest lift worth detecting to estimate how much traffic you need.
- Do not peek and stop early. Calling a winner the moment it looks good inflates false positives badly1. Let the test run to its planned sample or significance threshold.
- Run full weeks. Behavior differs by day of week. Always test in whole-week increments to avoid skew.
- Watch the right metric. Optimize for revenue per visitor or conversion rate, not vanity metrics like time on page.
Why traffic volume matters
Significance depends on sample size. A store with high traffic can read a test in a week; a smaller store may need a month or more, or a bigger expected effect. If you do not have the volume to test small tweaks, focus on bigger, bolder changes (or qualitative research) where the effect is large enough to detect.
Common mistakes that create false winners
- Stopping the test as soon as it looks like it is winning.
- Running tests for only a few days.
- Testing during an unusual period (a sale, a press spike) and generalizing.
- Ignoring the revenue impact and celebrating a click-rate lift that does not convert.
Tools and reporting
Modern ecommerce testing tools, such as Intelligems on Shopify, handle traffic splitting and give z-score-aware, revenue-based significance reporting. The output should tell you the lift, the probability to beat the baseline, and whether it is safe to roll out.
The payoff
Each test does two things: it either ships a real revenue gain, or it teaches you something about your customers. Over time, a disciplined testing program compounds into a store that converts far better than guesswork ever could. That program is the heart of what we do as a CRO agency.
References
- 1.Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2017). Peeking at A/B Tests: Why It Matters, and What To Do About It. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. https://doi.org/10.1145/3097983.3097992
Related research
Want a team to run this for you?
See how we help