Free tool
A/B Test Significance Calculator
A version doing better than another proves nothing until you know whether the gap holds up statistically. Four numbers, and you have the answer.
Statistical significance in an A/B test is the probability that the observed gap between two versions is not down to chance. It is computed with a two-proportion test, from the visitors and conversions of each version. The commonly used threshold is 95 %.
The gap between B and A, with its margin of error
- Reliable gap at 95%
- Within the margin of error
What the calculation does
Rate A = Conversions A ÷ Visitors A Rate B = Conversions B ÷ Visitors B Relative uplift = (Rate B ÷ Rate A) − 1 Confidence = two-proportion test (two-tailed)
5,000 visitors and 250 conversions on one side (5 %), 5,000 and 300 on the other (6 %): relative uplift is +20 % and confidence is 97.2 %. Above 95 %, the gap counts as significant.
The same 20 % uplift on a tenth of the traffic would not be conclusive. That is the point of the calculation: it separates the size of the gap from the strength of the evidence, two things the eye consistently merges when reading a results table.
⚠️ The test is two-tailed: it answers "do the two versions differ", not "is B better". It is symmetric, so a variant losing by 20 % comes out with the same confidence as one winning by 20 %.
Three ways to get an A/B test wrong
- Stopping the test the moment it crosses 95 %. Checking every day, you eventually cross the threshold by chance. Test duration is fixed BEFORE launch, and you do not conclude early.
- Testing on less than a full cycle. Monday buying behaviour is not Saturday buying behaviour: a test that does not cover whole weeks compares two days, not two versions.
- Confusing significant with meaningful. With enough traffic, a 0.1 point gap becomes significant while changing nothing about revenue. Confidence says the gap is real, not that it is worth shipping.
Frequently asked questions
What does 95 % significance mean?
It means there is less than a 5 % chance of seeing a gap this large if the two versions were in fact identical. It is not the probability that B is better: it is the probability of such a gap in the absence of a real difference.
How many visitors do you need to conclude?
It depends on the base rate and the gap you want to detect. The lower the rate and the smaller the gap you are after, the more traffic it takes. A 20 % gap on a 5 % rate needs a few thousand visitors per version; a 5 % gap needs tens of thousands.
Can you test more than two versions at once?
Yes, but the calculation above compares two. With several variants, the chance that at least one crosses the threshold by luck rises with each added comparison, and the threshold has to be adjusted accordingly.
Creative tests never really settling anything?
That is almost always a volume-per-variant problem before it is a creative problem. The Meta creative testing calculator says how many you can seriously test on your budget.
Get Creative Tests That Decide

