Take one real number before anything else. A product page converts at 3 percent. You want to know whether a change lifts that to 3.3 percent, a 10 percent relative improvement that would matter to the business. Run the standard two-proportion sample size formula at 95 percent significance and 80 percent power and the answer is roughly 52,000 sessions per variant, around 104,000 total. At 5,000 weekly sessions to that page, split evenly across two variants, that is close to five months of traffic for one test. Most teams call it after two weeks and roughly 10,000 sessions.
That gap between the sample size the math demands and the sample size most ecommerce tests actually get is where the majority of AB testing failure lives. Not in bad ideas, not in ugly variants, but in stopping the clock before the number was ever real.
The four inputs that actually set your sample size
The ab test sample size ecommerce stores need is not one number you look up. It is the output of four inputs, and the baseline rate and the target lift alone can move the requirement by an order of magnitude:
- Baseline conversion rate - the rate your control already converts at. Lower baseline rates need more traffic to detect the same relative lift, because the underlying variance is higher relative to the signal.
- Minimum detectable effect (MDE) - the smallest lift worth being able to see. This is a business call, not a statistical one, and it is the input teams get wrong most often by setting it too optimistic.
- Statistical significance level - conventionally 95 percent (alpha of 0.05), the tolerance for a false positive.
- Statistical power - conventionally 80 percent, the tolerance for missing a real effect (a false negative).
Feed real numbers into Evan Miller's AB testing sample size calculator rather than trusting a rule of thumb, because the relationship is not linear. Halve the MDE you are trying to detect and the required sample size does not double, it roughly quadruples, since required sample size scales with the inverse square of the effect size. That single fact explains why testing a homepage headline (high baseline traffic, often a large MDE) is a two-week exercise, and testing a shipping-page microcopy change on a mid-traffic PDP (lower baseline conversion, smaller realistic MDE) can be a multi-month one, even though both look like "one AB test" on a roadmap.
Why the calculator's default inputs mislead ecommerce teams
Most off-the-shelf sample size calculators, including the widely used ones built around the two-proportion formula, assume a clean binary outcome: did the visitor convert, yes or no. Revenue per visitor, average order value, and engagement scores do not fit that shape directly. A continuous metric like AOV can be tested with different statistical methods, but most teams do not have those tools wired up, and instead either force the metric into a threshold, something like "percentage of sessions that placed an order over 100 dollars," or skip the translation and run a standard proportion-based test directly on revenue data. That second option routinely produces noisy, unstable reads that never settle, because a handful of large orders can swing the average without changing the underlying behavior at all.
The second and more expensive mistake is running the sample size math on a number the tracking cannot actually deliver. If a duplicate purchase event, a missing add_to_cart trigger, or an unreconciled GA4-versus-Shopify gap is inflating or deflating the conversion count your test tool reads, the entire calculation is built on a false baseline. We cover the mechanics of that discovery process in our breakdown of what a CRO audit actually checks before anyone touches a test, and it is the same discipline behind a dedicated conversion rate optimization audit: confirm the number feeding the test is real before you decide how many visitors you need to trust it.
The traffic floor: when a formal test cannot work at all
Some pages should never get a formal AB test, and the honest move is to say so rather than run one that will never reach significance. Ton Wesseling's ROAR framework, presented at Emerce Conversion and published on Online Dialogue, lays out the practical floor in plain terms across four stages tied to monthly conversion volume:
- Risk - below roughly 1,000 conversions a month, Wesseling's framing is that formal AB testing is "still totally nonsensical." The sample size math simply will not close in a usable timeframe.
- Optimization - from roughly 1,000 conversions a month, testing becomes statistically viable, with something like 20 well-run tests a year realistic for a focused program.
- Automation - past roughly 10,000 conversions a month, testing stops being one person's side project and becomes core to how the organization ships changes, with automated and algorithmic testing becoming feasible.
- Rethink - the framework's fourth stage, positioned beyond Automation as an ongoing innovation phase rather than a fixed traffic threshold.
If a page sits in the Risk zone, the better use of time is qualitative: session recordings, heuristic review, and the kind of funnel diagnosis we walk through in finding where an ecommerce funnel actually leaks. Ship the fix directly, monitor the trend, and save formal testing for pages with the traffic to support it.
Peeking: the fastest way to manufacture a fake winner
The single most common way ecommerce teams break their own test math is peeking, checking results daily and stopping the moment the dashboard turns green. Evan Miller's widely cited analysis of this exact behavior puts a number on the damage: a tester who watches continuously and stops as soon as the result crosses the conventional p less than 0.05 threshold will call a false winner 26.1 percent of the time, not the 5 percent most people assume they are risking. That is more than five times the error rate the significance threshold implies.
A few hygiene rules keep this from happening in practice:
- Decide the sample size and stop date before the test launches, not after watching the trend.
- Run whole weeks only, one to four weeks depending on the calculated requirement, so weekday and weekend behavior both appear in the read.
- Test one challenger against one control at a time. Multiple simultaneous variants without a proper multivariate design multiply the ways a false positive can sneak in.
- If you must check early, use pre-registered sequential thresholds that get stricter over time rather than the same static 0.05 cutoff at every glance.
Winning tests overstate themselves
Even a test that runs to its full sample size and reaches genuine significance tends to overstate the win. This is the practical consequence of what statisticians call a Type M, or magnitude, error: among results that clear a significance threshold, the measured effect is on average larger than the true effect, because a real signal happened to line up with favorable noise on the way to significance. Andrew Gelman and John Carlin's paper on the subject, Beyond Power Calculations: Assessing Type S and Type M Errors, defines the exaggeration ratio directly: the expected ratio of a statistically significant estimated effect to the true effect size, which is greater than one whenever a study is underpowered relative to the effect it is trying to detect.
The practical takeaway for an ecommerce testing program is to discount headline win percentages, especially on tests that only barely cleared significance or ran with a smaller sample than the calculator called for. A program that sums up every "significant" test result at face value to justify its budget is very likely counting inflated numbers. Treat a fresh win as a hypothesis to reconfirm at scale, not a locked-in permanent lift.
A working checklist before you launch the next test
- Write down the baseline conversion rate for the exact page and segment being tested, not a site-wide average.
- Set the minimum detectable effect as a business decision first, then check what sample size that MDE actually requires.
- Run the number through a real calculator, such as Evan Miller's tool, and write the required sample size and expected duration on the test brief before launch.
- Confirm the conversion event feeding the test is accurate, not inflated or undercounted by a tracking gap.
- Check the page's monthly conversion volume against the roughly 1,000-conversion floor before committing to a formal test at all.
- Lock the stop date and sample size before launch, and resist checking daily for a green result.
- Treat a fresh, barely-significant win as provisional, and discount the headline lift when forecasting its revenue impact.
None of this replaces judgment. It sets the floor under it. A test that respects its own sample size requirement gives you a number you can actually act on; one that does not is just a coin flip with a p-value attached. If the honest math says a page cannot support a real test, that is not a reason to run one anyway. It is a signal to fix the page directly and put the testing budget where the traffic can back it up, work that a dedicated CRO service built around measurement rigor is built to scope correctly from the first audit.

