Answer · Conversion Engineering
What’s Statistical Significance in A/B Testing?
The short answer
Statistical significance is the probability that an observed difference between variants isn’t due to random chance. The standard threshold in CRO is 95% confidence — meaning there’s less than 5% probability the “lift” is noise. Without hitting this threshold, an A/B test result is not a real win.
№ 01The longer answer
When you run an A/B test, the conversion rates of the two variants will differ even if there’s no real underlying difference — just from random sampling variation. Statistical significance measures how likely the observed difference is to be random vs real.
The standard threshold is 95% confidence (p-value < 0.05). This means: if there were truly no difference between variants, we’d expect to see a result this extreme only 5% of the time by random chance. Higher confidence (99%) requires more sample; lower confidence (90%) accepts more false positives.
Statistical significance is necessary but not sufficient. You also need statistical power (typically 80% — the probability the test detects a real effect if one exists). Underpowered tests have low ability to detect real effects, leading to false null results.
The dishonest game many agencies play: report a 12% lift without confidence intervals, or with a confidence level that doesn’t reach 95%. A 12% lift with a 90% CI of [-3%, +27%] is a coin flip in disguise. Always demand the confidence interval, not just the point estimate.
№ 02What’s a p-value?
Probability of seeing the observed result (or more extreme) if there were truly no difference between variants. p < 0.05 = significant at 95% confidence. p < 0.01 = significant at 99%. Higher p = less significant.
№ 03Why 95% and not 99%?
95% is industry standard because it balances false-positive rate (5%) against required sample size. 99% confidence is stricter but requires roughly 70% more sample. Most CRO teams accept 95% as the working threshold.
№ 04Can a test be significant but not meaningful?
Yes — with huge sample sizes, you can detect tiny lifts (0.1%) as “statistically significant.” The result is real but not worth the implementation cost. Always combine statistical significance with practical significance (is the lift big enough to matter?).
Go deeper
Related questions
Three Ways to Start · No Sales Pitch
Want this answered for your business?
$500 audit. 5-day delivery. Refundable on engagement.