On this page

For AI agents: a documentation index is available at /llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.

Frequentist Sequential Testing

Learn how sequential testing addresses the peeking problem in A/B tests and enables early decision making with statistical rigor.

Sequential testing adjusts p-values and confidence intervals so you can read experiment results before the target sample size without inflating the false positive rate. Statsig implements it as a mixture sequential probability ratio test (mSPRT): each metric's confidence interval widens by an amount that shrinks as data accumulates. Use it on frequentist experiments when you need to catch regressions early or ship early on a statistically significant result.

Statsig offers three analysis methods, which you choose with the Analytics Type setting under Advanced Settings on the experiment Setup tab. Use frequentist analysis with sequential testing for standard p-values and confidence intervals that stay valid when you read results early. Use SPRT when you want unlimited peeking and the option to accept the null hypothesis as well as reject it. Use Bayesian mode for chance-to-beat and expected-loss readouts, with optional priors from past experiments.

Why peeking inflates false positives

Traditional A/B tests (t-tests and z-tests) are fixed-horizon tests. When you design the experiment, you set the number of units to observe, and you read the metrics once, after you reach that target sample size. Sequential testing lets you read results and make valid decisions before that point.

Continuous monitoring of an experiment, or peeking, inflates the false positive rate above the significance level you chose. Every time you consider ending an experiment early, you decide whether to reject the null hypothesis, and that decision can be wrong. Metric values and p-values fluctuate because of noise from random unit assignment and unpredictable user behavior, so results can move into and out of statistical significance even when there is no real effect. Noise levels vary by test and decrease over time as more users join and you observe them longer.

When you decide based on an early snapshot, you're more likely to select a statistically significant result that wouldn't appear if you analyzed the data once at the planned end of the experiment. In frequentist procedures, early decisions can only increase the false positive rate, even when you intend to make a less biased decision.

How sequential testing works for an A/B test

Statsig automatically adjusts p-values and confidence intervals for each preliminary analysis window to compensate for the increased false positive rate that peeking introduces. The Results tab shows the adjustment:

Sequential testing results visualization

Statsig expands the confidence interval for each metric and marks the expansion with wings on the ends of the bar. The wings show that sequential testing is on and how much the interval expanded.

Statsig results table highlighting sequential testing adjusted confidence interval

In this example, the sequential testing adjustment determines whether Statsig declares the result significant.

Sequential testing supports early decisions when observations are strong enough to outweigh random fluctuations, while limiting the risk of false positives. Statsig discourages peeking in general, but regular monitoring with sequential testing is valuable in two cases:

  • Unexpected regressions: When an experiment has a bug or an unintended consequence that severely affects key metrics, sequential testing identifies the regression early and distinguishes significant effects from random fluctuations.
  • Opportunity cost: Delaying a decision can be costly, for example when you need to launch a feature ahead of a major event or fix a bug. Sequential testing can support an early decision if key metrics show improvement. Use caution: an early significant result on some metrics doesn't guarantee enough power to detect regressions in other metrics. Limit this approach to cases where only a small number of metrics matter to the decision.

You can use sequential testing anywhere you run an experimental analysis, including the experiment Results page and any custom queries.

Enable sequential testing

On the Setup tab of your experiment, set Analytics Type to Frequentist, then select Apply Sequential Testing under Analysis Settings. You can toggle this setting at any time during the experiment; you don't need to turn it on before the experiment starts.

Sequential testing configuration interface

Interpret sequential testing results

Click Edit at the top of the metrics section in Pulse to toggle sequential testing on or off.

Pulse metrics sequential testing toggle

When you turn on sequential testing, Statsig adjusts the results it calculates before the target completion date of the experiment.

Sequential testing confidence interval visualization

The dashed line is the expanded confidence interval that results from the adjustment. The solid bar is the standard confidence interval without adjustment. If the adjusted confidence interval overlaps zero, the metric delta isn't significant yet, and the experiment should continue as planned.

Early decisions often produce underpowered lift estimates with high uncertainty. If you need the correct ship decision, act on significant sequential testing results. If you need an accurate measurement of the effect size, wait for the full power that your pre-experiment power calculation estimates. Statsig doesn't calculate statistical power on post-hoc results; refer to the section "Post-hoc Power Calculations are Noisy and Misleading" in Kohavi, Deng, and Vermeer, A/B Testing Intuition Busters.

How Statsig implements sequential testing

Two-sided tests

Confidence intervals

Statsig uses mSPRT based on the approach that Zhao et al. propose. The two-sided sequential testing confidence interval with significance level α\alpha is:

CI∗(ΔX‾)=ΔX‾±Zα/2∗⋅VCI^*(\Delta \overline{X}) = \Delta \overline{X} \pm Z^*_{\alpha/2} \cdot \sqrt{V}

where:

  • Zα/2∗Z^*_{\alpha/2} is the z-critical value, modified for sequential testing:
Zα/2∗=(V+τ)τ(−2ln⁡(α/2)−ln⁡(VV+τ))Z^*_{\alpha/2} = \sqrt{\frac{(V+\tau)}{\tau}\left(-2\ln(\alpha/2)-\ln(\frac{V}{V+\tau})\right)}
  • VV is the standard variance of the delta of means. You can obtain it from the sample variance of the test and control group means:
V=var(ΔX‾)=var(X‾t)+var(X‾c)=var(Xt)Nt+var(Xc)NcV = var(\Delta \overline X) = var(\overline X_t) + var(\overline X_c) = \frac{var(X_t)}{N_t} + \frac{var(X_c)}{N_c}
  • τ\tau is the mixing parameter:
τ=(Zα/2)2⋅var(Xt)+var(Xc)Nt+Nc\tau =(Z_{\alpha/2})^2\cdot\frac{var(X_t)+var(X_c)}{N_t+N_c}
  • Zα/2Z_{\alpha/2} is the z-critical value used in the non-sequential test, for the significance level you want (1.96 for the standard α=0.05\alpha = 0.05).

Statsig has validated that this parameter satisfies the expected false positive rate and provides enough power to detect large effects early. For details, refer to Statsig's sequential testing analysis.

p-values

To produce sequential testing p-values that are consistent with the expanded confidence intervals, Statsig modifies the standard p-value methods.

The goal is to evaluate the mSPRT test so that the Type I error remains approximately equal to α\alpha, and so that the sequential testing p-value is consistent with the expanded confidence interval: a confidence interval that includes 0.0% has a p-value ≥ α\alpha, and one that excludes 0.0% has a p-value < α\alpha.

The observed z-statistic (z-score) ZZ remains unchanged. Instead of evaluating ZZ on a standard normal distribution N(0,1)N(0, 1), Statsig evaluates it against a normal distribution N(0,σ2)N(0, \sigma^2) with mean zero and standard deviation σ\sigma. For a two-sided test, to limit the probability that an observed ZZ exceeds Zα/2∗Z^*_{\alpha/2} under the null hypothesis to α\alpha, solve for σ\sigma:

σ=Zα/2∗2⋅erf−1(1−α)\sigma=\frac{Z_{\alpha/2}^*}{\sqrt{2} \cdot erf^{-1}(1-\alpha)}

where erf−1erf^{-1} is the inverse error function.

The two-sided sequential testing p-value is then:

p-value∗=2⋅12π∫−∞−∣Z∣e−t22σ2σdt\text{p-value}^* = 2 \cdot \frac{1}{\sqrt{2\pi}} \int \limits _{-\infty}^{-|Z|} \frac{e^{- \frac{t^2}{{2\sigma^2}}}}{\sigma}dt

where ZZ is the observed z-statistic.

One-sided tests

Statsig modifies each step for one-sided sequential testing.

CI∗(ΔX‾)={[ΔX‾−Zα∗⋅V,+∞)if right-sided test(−∞,ΔX‾+Zα∗⋅V:]if left-sided testCI^*(\Delta \overline{X}) = \begin{cases} \left[\Delta \overline{X} - Z^*_{\alpha} \cdot \sqrt{V}, \quad +\infty \right) & \text{if right-sided test} \\ \\ \left(- \infty, \quad \Delta \overline{X} + Z^*_{\alpha} \cdot \sqrt{V} :\right] & \text{if left-sided test} \\ \end{cases}
p-value∗={1−12π∫−∞Ze−t22σ2σdtif right-sided test12π∫−∞Ze−t22σ2σdtif left-sided test\text{p-value}^* = \begin{cases} 1 - \frac{1}{\sqrt{2\pi}} \int \limits _{-\infty}^{Z} \frac{e^{- \frac{t^2}{{2\sigma^2}}}}{\sigma}dt \quad \text{if right-sided test} \\ \\ \frac{1}{\sqrt{2\pi}} \int \limits _{-\infty}^{Z} \frac{e^{- \frac{t^2}{{2\sigma^2}}}}{\sigma}dt \quad \text{if left-sided test} \\ \end{cases}

where:

  • Zα∗Z^*_{\alpha} is the one-sided z-critical value, modified for sequential testing:
Zα∗=(V+τ)τ(−2ln⁡(α)−ln⁡(VV+τ))Z^*_{\alpha} = \sqrt{\frac{(V+\tau)}{\tau}\left(-2\ln(\alpha)-\ln(\frac{V}{V+\tau})\right)}
  • VV is the variance of the delta of means, var(ΔX‾)var(\Delta \overline X).

  • τ\tau is the mixing parameter:

τ=(Zα)2⋅var(Xt)+var(Xc)Nt+Nc\tau =(Z_{\alpha})^2\cdot\frac{var(X_t)+var(X_c)}{N_t+N_c}
  • ZαZ_{\alpha} is the one-sided z-critical value used in the non-sequential test, for the significance level you want (1.645 for the standard α=0.05\alpha = 0.05).

  • Statsig solves for σ\sigma with:

σ={Zα∗2⋅erf−1(1−2α)if right-sided test−Zα∗2⋅erf−1(2α−1)if left-sided test\sigma = \begin{cases} \frac{Z_{\alpha}^*}{\sqrt{2} \cdot erf^{-1}(1 - 2 \alpha)} & \text{if right-sided test} \\ \\ \frac{- Z_{\alpha}^*}{\sqrt{2} \cdot erf^{-1}(2 \alpha - 1)} & \text{if left-sided test} \end{cases}
  • ZZ is the signed observed z-statistic (z-score).

Was this helpful?