How Long Should You Run an A/B Test?
Test duration is math, not a guess. It comes from your baseline rate, the effect you want to detect and your traffic - plus the "test bandwidth" that limits how many tests you can run a year.


"How long should we run this test?" usually gets answered with "two weeks." Two weeks is a decent floor, but the real answer comes from math - and for a lot of stores it's uncomfortably longer.
Duration is an output, not an input
You don't choose how long a test runs. Four things decide it for you:
- Baseline conversion rate - what the control converts at today.
- Minimum detectable effect (MDE) - the smallest uplift you care to catch. Detecting +2% takes vastly more traffic than +20%.
- Confidence and power - usually 95% confidence and 80% power (the chance of catching a real effect that's there).
- Traffic - how many visitors per week reach the page you're testing.
Feed those into a sample size calculator and you get the visitors-per-variant you need. Divide by weekly traffic and you have your duration. That's the whole method.
A worked example
Say your product page converts at 5%, you want to catch a 10% relative lift (to 5.5%), at 95% confidence and 80% power. You need roughly 8,000 visitors per variant - about 16,000 total. At 14,000 visitors a week, the test takes a little over a week. At 4,000 a week, it's a month.
Same test, same goal - the only thing that changed was traffic, and duration tripled. That's why "two weeks" is a guess.
Run to the sample size, then read once
The discipline that protects the result: decide the sample size first, run until you reach it (and past the business-cycle floor), and read the outcome a single time. Checking daily and stopping on the first green number is how you manufacture fake winners.
The business-cycle floor
Even when the math says four days, run for at least one to two full weeks. Shopper behaviour swings by day of week and by payday; a Monday-to-Thursday test can be quietly biased by who shops then. Let it cover full cycles so those effects average out. So the practical rule is: run for max(the sample-size duration, two weeks).
Test bandwidth: the number nobody calculates
Here's the uncomfortable part. If each test needs two weeks and you run one at a time on a page, you can run about 26 tests a year on it - and that's the optimistic case. Lower traffic or a smaller MDE and you might get 6 to 10 real tests a year.
That's your test bandwidth, and it changes how you should think about CRO. With ten test slots a year, you cannot waste one on a button colour - every slot has to go to a change likely to move revenue. Bandwidth is exactly why a research-led, well-prioritized process matters. Our sample-size tab estimates this directly: enter your traffic and it tells you how many tests per year your store supports.
What to do when you're low on traffic
Most stores are traffic-constrained. Your options, in order:
- Test bigger swings. A bold redesign or a price/offer change produces a larger effect that reaches significance faster than a subtle tweak.
- Raise your MDE. Stop trying to detect +2%. Decide the smallest lift that's actually worth shipping and size for that.
- Test higher-traffic pages. Home, top collection and PDP templates accumulate visitors fastest.
- Use variance-reduction (CUPED) or sequential methods if your tool offers them - they squeeze more signal from the same traffic.
- Concentrate, don't spread. One well-aimed test beats three underpowered ones running at once.
A planning checklist
- Pin the baseline rate and the MDE worth shipping before anything else.
- Get visitors-per-variant from a calculator; divide by weekly traffic for duration.
- Apply the two-week floor.
- Sanity-check your annual test bandwidth - it sets how selective you must be.
- Commit to the sample size and don't stop early.
The takeaway
Duration is sample size divided by traffic, floored at two weeks. Your annual bandwidth is smaller than you hope, which makes prioritization the highest-leverage skill in the discipline. Plan the test before you launch it - the calculator does the arithmetic in seconds.
Want help spending your test bandwidth on the right experiments? See how our CRO program works.
Related research
Statistical Power, Sample Size, and Minimum Detectable Effect in A/B Testing
A practitioner's review of statistical power, the 80% convention, minimum detectable effect, and the sample-size formula that sets how much traffic an A/B test needs.
Read the research ResearchAlways-Valid Inference in Conversion Testing
A practitioner review of sequential testing methods and the statistical cost of peeking at A/B test results.
Read the researchWant a team to run this for you?
See how we help