A/B Test Calculator
Size the test before it starts and evaluate significance once it ends. Results update as you fill in the fields.
Results
Over 90 days is too long for one test: the answer arrives late and the risk of contamination grows. Raise the MDE or send more traffic to the test.
Check the numbers: visitors and conversions must be greater than zero, and conversions cannot exceed visitors.
Results
How to use the calculator
- Pick a mode: plan the test, to find out how much sample you need, or evaluate the result, to check significance.
- Fill in the fields: when planning, the current conversion rate and the effect you want to detect; when evaluating, visitors and conversions for each variant.
- Read the results, updated as you type: sample size and duration in one mode; uplift, p-value and confidence intervals in the other.
- Decide with discipline: run the test until the planned sample size and only call a winner if it is significant at your chosen level.
Why sample size and significance matter
An underpowered A/B test lies to you: with too few visitors, ordinary fluctuation looks like a win, and today's champion is back to a tie next week. Calculating the sample size before you start defines when the test ends and keeps you from deciding on noise.
Significance does the other half of the job: it estimates how likely the observed difference is to be pure chance. Once the p-value drops below the threshold of your chosen confidence level, you can ship the variant with confidence; above it, the honest answer is to keep collecting data.
A quick A/B testing glossary
Five terms used by the calculator:
- MDE (minimum detectable effect)
The smallest uplift the test can detect. The smaller the MDE, the larger the sample you need: detecting +5% takes far more traffic than detecting +20%. - Confidence level
Protection against false positives. At 95%, if there is no real difference between the variants, the test wrongly declares a winner only 5% of the time. - Statistical power
Protection against false negatives. At 80% power, a real effect the size of the MDE is detected in 8 out of 10 tests. - p-value
The probability of seeing a difference as large as the observed one (or larger) if the variants were identical. Below 0.05, results are conventionally called significant at 95%. - Confidence interval
The range where each variant's true conversion rate most likely sits. Heavily overlapping intervals call for a cautious read.
Frequently asked questions
Why is 95% confidence the default?
It is the convention that balances rigor and practicality: it accepts a 5% false-positive chance. High-stakes tests (pricing, checkout) deserve 99%; low-risk tests can run at 90% and finish sooner.
Can I stop the test as soon as it turns significant?
No; stop at the planned sample size. Checking daily and stopping at the first significant reading (known as peeking) inflates the false-positive rate, because chance crosses your chosen threshold several times during a test. Fix the sample size up front and only decide once you reach it.
How many variants does the calculator support?
Two: A (control) and B (variation). With three or more, pairwise comparisons multiply the false-positive risk; you would need to correct the significance level (Bonferroni, for instance) or use the testing platform's own tool.
Why does the result differ from other calculators?
Each tool adopts its own assumptions: a one-tailed or two-tailed test, continuity correction, Wald or Wilson intervals. This calculator uses the two-tailed two-proportion test with Wald intervals, the same standard as the leading market tools. Small differences between calculators come from those choices.
Need support running your tests?
Marktech is a Google and Meta partner and manages campaigns and experiments for companies, chains and franchises.
Talk to a consultant