On this page

For AI agents: a documentation index is available at /llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.

Power Analysis

Learn how to use Statsig's Power Analysis Calculator to determine experiment parameters needed for statistically significant results.

A power analysis estimates how long an experiment must run, and how much traffic it needs, to detect a given change in a metric. The Statsig Power Analysis Calculator takes a metric's historical mean, variance, and traffic volume and returns the trade-off between minimum detectable effect (MDE), duration or exposures, and allocation. Run it before you start an experiment to set the experiment's target duration. After launch, read the confidence intervals on the experiment's Results tab instead; refer to How to read experiment results.

Variables the calculator estimates

The most common variable to optimize is duration (how long your experiment runs), but the calculator supports other variables as well. Using the known mean and variance of a metric and the observed traffic volume, the Power Analysis Calculator estimates the relationship between three variables:

  • Minimum detectable effect (MDE): The smallest change in the metric that the experiment can reliably detect. For example, with an MDE of 1% and power of 80%, the experiment has an 80% chance of producing a statistically significant result if the true effect is 1%. If the true effect is smaller than 1%, the experiment is less likely to produce a statistically significant result, though it still can.
  • Number of days or exposures: How long the experiment is active and how many users enroll in it. Longer experiments typically have more observations, which leads to tighter confidence intervals and a smaller MDE. Statsig uses historical data to estimate the number of new users eligible for the experiment each day.
  • Allocation: The percentage of traffic that participates in the experiment. Larger allocation leads to a smaller MDE, so allocating as many users as possible often gives faster or more sensitive results. When there's a risk of negative impact or a need for mutually exclusive experiments, the calculator tells you the smallest allocation that can achieve your target MDE.

Run a power analysis

Open the Power Analysis Calculator from the tools menu, or from the link below the Experiment Duration field on the experiment setup page.

Power Analysis entry points

  1. Select the population that Statsig uses to determine the metric mean and variance and to estimate the number of exposures over time.
    • Everyone: Analysis uses the entire user base.
    • Targeting gate: Analysis covers only users who pass the selected feature gate. The gate must have been active for at least 7 days. Choose this option when you plan to use a targeting gate for the experiment.
    • Past experiment: Analysis uses data collected from a previous experiment. Choose this option when the new experiment affects a similar user base or part of the product as the previous one.
    • Qualifying event: Analysis covers only users who logged the specified event.
  2. Select a metric of interest (or multiple metrics for a targeting gate analysis).
  3. Click Run Power Analysis.

Power analysis form inputs

Your past power analyses appear on the Past Analyses tab.

Past power analyses list

To attach an existing power analysis to an experiment, use the dropdown menu on that analysis.

Power analysis attachment interface

Population types

The population you select determines the inputs of the analysis: mean, variance, and number of users. For reliable estimates, the metric values of the selected population should match those of the users you plan to target in the experiment.

Example: checkout flow with the Everyone population

Suppose you want to test a change in the checkout flow and estimate the MDE for total_purchases. If only about 10% of your daily users reach the checkout page and you use the Everyone population, the analysis is likely to:

  • Overestimate the number of users the experiment gets.
  • Underestimate the mean value of the total_purchases metric. The 90% of users who don't reach the checkout page have a value of zero, but in practice they aren't in the experiment and don't contribute to the metric.
  • Incorrectly estimate the variance of the total_purchases metric. Including the 90% of users with zero purchases changes the distribution of metric values.

When an experiment includes only a subset of users, the MDE and duration from an analysis of the whole user base may not be reliable. To address this bias, use data from a past experiment that targeted the same part of the product. For example, if a prior experiment also targeted the checkout page, that data provides better estimates of traffic volumes and metrics for that part of the product.

Inputs by population type

The following table shows how Statsig obtains the power analysis inputs from each population type.

Analysis types

On the results view, you can change inputs such as # of Groups, Control Group %, and the analysis type to update the power analysis results.

Power analysis results showing adjustable parameters

Fixed allocation analysis

If you already know the available allocation, fixed allocation analysis shows how the length of the experiment affects the MDE. The following example shows how the MDE for a page load metric shrinks over time in an experiment with 100% allocation. After 1 week, the expected user count per group is 5,200 with an MDE of 21.6%. By week 4, the user count per group increases to approximately 48,000 and the MDE drops to 7%.

Fixed allocation analysis results chart

Fixed MDE analysis

If you already know the effect size you want to measure, fixed MDE analysis shows the allocation and duration needed to achieve that MDE. Enter your target MDE as a percentage of the current metric value. For example, if a website gets 1,000 page loads per day, an MDE of 10% means the experiment can detect a change of 100 or more page loads per day.

The results show the minimum number of weeks needed to reach this MDE for different allocation percentages. In the following example, the experiment must run for at least 2 weeks with 65% allocation or 4 weeks with 50% allocation. You can't achieve the target MDE in 1 week, because doing so would require more than 100% allocation.

Fixed MDE analysis results table

Advanced options

Advanced power analysis options interface

Advanced settings for customizing the analysis:

  • Number of Experiment Groups: The total number of groups in the experiment, including control.
  • Control Group %: The percentage of users in the control group (for example, 50% if half of all users are in control).
  • Fixed Allocation or Fixed MDE Analysis: The type of analysis to run. Refer to Analysis types for more details.
  • One-sided or Two-sided test: The type of z-test to use for the analysis.
  • Significance level (α): The significance level for the test. Typically 0.05.
  • Power (1-β): The statistical power for the test. Typically 0.8.
  • Bonferroni Correction Per Variant: Whether to include α correction for multiple tests in the power analysis.

Calculation details

Statsig computes the relative percentage MDE for a given metric X using the following equation:

MDE calculation formula

  • X-bar is the mean metric value across all users.
  • var(X) is the population variance of the metric.
  • Ntest and Ncontrol are the estimated number of users in the test and control groups. These come from historical active user data along with experiment allocation and group size.
  • Z1-β is the standard Z-score for the selected power. Typically 1-β = 0.8 and Z1-β = 0.84.
  • Z1-α/2 is the standard Z-score for the selected significance level in a 2-sided test. Typically α = 0.05 and Z1-α/2 = 1.96.

This calculation relies on statistics computed across the entire user base of the project. It doesn't account for experiments that target only a subset of users, which may have different summary statistics for their key metrics. For example, the metric mean and variance can differ in an experiment that targets only Android users or one that exposes users at the lower part of an acquisition funnel.

Was this helpful?