Choose Plan a test to estimate sample size and duration, or Analyze results to compare data from an existing test. You can switch between these views without losing your inputs during this visit.
Optional - Explore an Uplift Target
π If you need help choosing a target, expand the optional historical-data tool the planning calculators. Use its scenarios as a starting point alongside the smallest improvement that would matter to your team.
- β
This helps you avoid setting targets that are too aggressive or statistically improbable.
- β
Compare your expected uplift against historical trends to ensure feasibility before running a test.
Step 1 - Plan Your Sample Size and Duration
π In Plan a test, enter your baseline in the indicated units and your relative improvement as a percentage. A 10% baseline with a 5% relative improvement means a 10.5% target. Then enter total eligible daily participants and the percentage allocated to the test; allocated traffic is split equally between A and B.
The tool accounts for:
- πConfidence Level β The probability that your test results are not due to random chance (commonly set at 95%).
- πStatistical Power β The likelihood that your test correctly detects a real uplift (usually set at 80%).
- πExpected Uplift & Baseline Metric β These factors determine the minimum sample size needed to achieve statistical significance.
Why this matters:
- β
Helps you plan for enough participants to detect the effect you're testing for.
- β
Helps avoid false positives and negatives due to small sample sizes.
Step 2 - Run Your A/B Test
π’ Launch your A/B test and monitor the results:
Control Group (A) β Current Experience.
Variant Group (B) β New Experience.
πRun the test until a sufficient sample size is reached to ensure statistical validity.
β οΈImportant: Avoid stopping a test early just because results look promisingβit can lead to misleading conclusions.
Step 3 - Analyze Your Results
π Use the Statistical Significance Calculator to determine if your test results are meaningful.
The tool applies a frequentist approach and considers your:
- πP-value β Measures whether the observed difference is statistically significant or just due to random chance.
- πConfidence Intervals β Provides a range where the true uplift likely falls, helping you assess the test's reliability.
How to interpret results:
- β
P-value < 0.05? Your uplift is statistically significant, meaning you can confidently roll out the change.
- β
P-value > 0.05? The test results are inconclusiveβconsider gathering more data or refining your test.
Best Practices for Reliable A/B Tests
- β
Run tests for a sufficient duration β Avoid premature conclusions by ensuring enough data collection.
- β
Test one variable at a time β This isolates the true cause of any observed uplift.
- β
Monitor external factors β Seasonal events, major updates, or promotions can skew test results.
Use these steps to build a repeatable experimentation process, share what your team learns, and make informed decisions about the metrics that matter to your organization. π