A/B Test Significance Calculator: P-Value & Confidence

Check whether the difference between your control and variant conversion rates is statistically significant.

Frequently Asked Questions

How do you calculate A/B test statistical significance?

This uses a two-proportion z-test: it compares the conversion rates of your control and variant, accounting for sample size, to calculate a p-value and confidence level for the observed difference.

What does "95% confidence" mean?

It means there's only a 5% probability the observed difference happened by random chance alone. A result below the 95% threshold (p < 0.05) is the most commonly used bar for declaring a winner.

Why is my test not significant even though the variant did better?

Small sample sizes produce a lot of random noise. A real difference needs enough visitors and conversions before you can be statistically confident it isn't just chance.

How A/B test significance is calculated

This calculator runs a two-proportion z-test: it compares your control and variant conversion rates, accounting for sample size, to produce a z-score and p-value. A p-value below 0.05 (95% confidence) is the most commonly used threshold for declaring a statistically significant winner.

How to use this A/B test calculator

  1. Enter your control’s visitors and conversions.
  2. Enter your variant’s visitors and conversions.
  3. Click Calculate Significance to see your lift, confidence level, and p-value.

Why sample size matters more than „the variant looks better“

A variant showing a higher conversion rate in raw numbers doesn’t automatically mean it’s actually better — small differences in small samples are often just random noise. Statistical significance testing accounts for sample size directly, which is why the same percentage-point difference can be „significant“ with a large enough sample and „not significant“ with a small one.

Reading your results correctly

  • P-value: The probability the observed difference could have happened by chance alone if there were truly no real difference. Lower is stronger evidence of a real effect.
  • Confidence level: Simply 1 minus the p-value, expressed as a percentage — 95% confidence (p < 0.05) is the most common threshold for declaring a winner.
  • Relative lift: How much better (or worse) the variant performed as a percentage of the control’s rate, useful alongside significance for judging practical impact.

Tips for running valid A/B tests

Decide your sample size and test duration in advance rather than stopping the test as soon as it looks significant, since checking repeatedly and stopping early inflates your chance of a false positive. Run tests for at least one full business cycle to average out day-of-week effects on traffic and conversion behavior.

Expert insight: why „significant“ doesn’t always mean „true“

Ronny Kohavi, who ran experimentation platforms at Microsoft, Airbnb, and Amazon and is one of the most cited voices in A/B testing, points out that a statistically significant result can still be wrong more often than most people assume, once you account for how many experiments actually succeed industry-wide. His most emphasized pitfall: an underpowered test (too small a sample) can fail to detect a real effect just as easily as it can manufacture a fake one, which is why sample size matters as much as the p-value itself.

Related calculators

Disclaimer: This calculator is provided for general informational and educational purposes only and does not constitute financial, investment, tax, or legal advice. Results are estimates based on the values you enter and should not be relied upon as the sole basis for any financial or other decision. Past performance and projected figures are not a guarantee of future results. Always consult a qualified professional before making financial decisions. See our Legal Notice and Privacy Policy for more information.