Plans an A/B test (sample size and duration) and checks whether the difference in conversion rate between two or more variants is real or just random variation, with the size of the difference and its confidence interval. For anyone running an A/B test who needs a straight answer before shipping a variant.
How to use it
Choose Plan a test before you launch to learn how many visitors you need, or Analyze a test with the visitors and conversions of each variant as your testing or analytics tool reports them. The result updates as you type.
- Plan first: enter the baseline conversion rate and the smallest change worth detecting, and note the sample size per variant.
- Run the test until every variant reaches that size, using the same date range and the same conversion for all variants.
- Enter the numbers in Analyze, with the first variant as the control. Read the verdict, then the intervals and the findings: they say how big the difference probably is and whether the data can be trusted.
How Reavlo tests this
The calculator uses a frequentist method, in your browser, and declares no winner without statistical significance.
Planning. With baseline p1, target p2 (the baseline plus your minimum detectable effect, relative or absolute), significance α and power 1 − β, the visitors needed per variant are n = (z1−α/2 × √(2 p̄ (1 − p̄)) + z1−β × √(p1(1 − p1) + p2(1 − p2)))² ÷ (p2 − p1)², where p̄ = (p1 + p2) ÷ 2, rounded up. With more than two variants, α is divided by the number of comparisons with the control (Bonferroni), which keeps the plan valid for the Holm procedure used when analyzing. Duration is the total sample divided by your daily visitors, rounded up to whole days.
Analyzing. Each variant is compared with the control (the first variant):
- Conversion rate p = conversions ÷ visitors.
- Pooled two-proportion z-test: z = (pvariant − pcontrol) ÷ √(p̂ (1 − p̂) (1 ÷ ncontrol + 1 ÷ nvariant)), with the two-sided p-value from the standard normal distribution.
- Holm correction across all comparisons: sort the p-values, multiply the smallest by the number of comparisons, the next by one fewer, and so on, keeping the adjusted values non-decreasing. A variant is significant when its adjusted p-value is below 1 − confidence.
- The interval for the difference uses the unpooled standard error and the critical z of your confidence level. It is not corrected for multiple comparisons.
- Lift is (pvariant − pcontrol) ÷ pcontrol.
- The winner is the variant with the highest conversion rate among those significantly better than the control. If none is, there is no winner.
Sample ratio mismatch (SRM). A chi-square test compares the visitors per variant with the planned split (equal when you give none). A p-value below 0.001 is flagged as critical: the groups are not comparable and the other results should not be trusted.
Findings.
srm.detected: the SRM test above.sample.too-small: a variant has fewer visitors than the planned sample size, or fewer than 100 when none is given.normal-approximation.weak: a variant has fewer than 5 conversions or 5 non-conversions, where the normal approximation is unreliable.peeking.warning: you looked at the results more than once before the planned end. A fixed-sample test assumes one read at the end.
Distribution functions (normal, chi-square) are computed from the regularized incomplete gamma function and checked against SciPy and statsmodels to within 1e-6.
Questions
What does an adjusted p-value of 0.02 mean?
If there were no real difference between the variants, a difference at least this large would show up in about 2% of tests of this size, after allowing for the number of variants you compared. It doesn't mean there's a 98% chance the variant is better.
Why is my test not significant even though a variant looks better?
With few visitors, random variation alone produces differences of that size often enough. The interval shows the range of differences that are still consistent with your data. Plan the sample size first and keep the test running until you reach it.
Why not just stop the test when it turns significant?
Every time you look, noise has another chance to look like a win. Checking daily and stopping at the first significant result makes false positives far more common than the confidence level suggests. Decide the sample size up front and read the result once.