A/B test significance calculator
Enter visitors and conversions for each arm. You get the lift with its confidence interval, the p-value, a sample ratio check, and a readout written in plain English.
Readout
Two-proportion z-test—relative lift
- Control rate
- —
- Variant rate
- —
- Difference
- —
- p-value
- —
Enter your results to see the readout.
Reading the result
- Relative lift is the variant's rate divided by control's, minus one. The bar under it is the confidence interval: the range of true lifts that fit your data.
- The interval matters more than the p-value. A significant +4% with an interval from +0.2% to +8% is a much weaker case than one from +3% to +5%.
- Not significant doesn't mean no effect. It means this test couldn't tell. If the interval still includes a lift you'd care about, the test was underpowered.
- Check the split first. If visitors didn't divide the way you intended, something in assignment or tracking is broken and the lift can't be trusted. This page runs that check automatically.
The method
The p-value comes from a two-sided two-proportion z-test with a pooled standard error. The interval for the difference in rates uses the unpooled standard error. The interval for relative lift uses the delta method on the log of the rate ratio, so it isn't symmetric around the point estimate.
The sample ratio check is a chi-square goodness-of-fit test against your intended split. It flags a mismatch when p < 0.001, a conservative threshold that keeps false alarms rare on large tests.
These are fixed-horizon statistics. They're valid when you decide the sample size up front and read the result once, at the end. Use the sample size calculator to set that horizon.
Review results like this together, every week.
Loupe puts each test's lift, interval and split check into your weekly roundtable, next to the team's theories and the call you made. Everything stays searchable.
Results that keep coming back flat?
An experimentation program review covers your backlog, test design and readouts, and leaves you with a ranked list of tests worth running. Audits start at $2,000.