CRO toolkit
Free A/B Test Significance Calculator
Check statistical significance, plan your sample size and how long to run, and simulate outcomes before you launch. Free, instant, no signup.
Significant
You can be 95% confident this result is real.
- Observed uplift
- 10.00%
- p-value
- 0.0250
- z-score
- 2.242
- Uplift CI
- 1.3% to 18.7%
Want results like these on your store?
These tools tell you if a test won. We design, build, and run the experiments that move revenue.
Get a free CRO auditSignificance
Two-proportion z-test for conversion rate, z-test on means for revenue. Tells you whether a difference is real or noise.
Sample size
Plan how many visitors and how long you need before launching, so you do not stop a test too early.
Simulation
Monte Carlo runs of your planned test to see how often it would actually reach significance.
A/B testing questions, answered
How do I know if my A/B test result is statistically significant?+
Enter the visitors and conversions for your control and variant above. The calculator runs a two-proportion z-test and returns the confidence level and p-value. A result is usually considered significant at 95% confidence (p-value below 0.05), meaning there is less than a 5% chance the difference is down to random noise.
How many visitors do I need for an A/B test?+
Use the Sample size tab. Enter your baseline conversion rate, the smallest uplift you want to detect (minimum detectable effect), your confidence level and power. The calculator returns the required visitors per variant and, from your weekly traffic, how long the test needs to run.
How long should I run an A/B test?+
Run it until each variant reaches the required sample size, and ideally for at least one to two full business cycles (typically two weeks) to even out day-of-week effects. Stopping early, as soon as a result looks significant, is the most common way to get a false positive.
What confidence level should I use for A/B testing?+
95% is the standard for ecommerce A/B testing - it balances catching real wins against false positives. Use 99% when a change is risky or expensive to ship, and 90% only for low-stakes, fast iterations where some extra false positives are acceptable.
What's the difference between a one-tailed and two-tailed test?+
A two-tailed test checks whether the variant is different from control in either direction (better or worse) and is the safer default. A one-tailed test only checks whether the variant is better, so it reaches significance faster but cannot tell you if the variant is actually hurting performance.
What is statistical power, and why does the simulator matter?+
Power is the chance your test detects a real effect when one exists (80% is the usual target). The Simulation tab runs thousands of Monte Carlo trials of your planned test so you can see how often it would actually reach significance - a quick way to spot an underpowered test before you waste traffic on it.
Free tool
Revenue Calculator
See how much extra revenue a conversion-rate lift is worth on your store. A good next step once a test wins.