For AI agents: a documentation index is available at /llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.
Frequently Asked Questions on Using Pulse
Frequently asked questions about interpreting Statsig experiment results, including p-values, confidence intervals, lift, exposures, and SRM warnings.
This page answers common questions about reading Pulse results: significance that flips, missing metrics, exposure-count mismatches, and dimension breakouts. For a walkthrough of the Results tab and the Scorecard, go to How to read experiment results.
Why did my statistically significant result turn negative
Trust the current result: it incorporates more information about the users in your experiment. A result can flip for several reasons:
- Random noise, which decreases as the sample size grows.
- Within-week seasonality (for example, an effect that differs on Mondays), which normalizes with more data.
- The population that saw the experiment early differs from slower adopters. This is common: a daily user likely sees your experiment before someone who uses your product once a month. Use the time series view to check.
- A novelty effect made the experiment look meaningful early, but the effect faded. For example, users might click a changed button out of curiosity, then revert to prior behavior. Use the days since exposure view to check.
Pick a readout date when you launch your experiment, based on a power analysis, and disregard the statistical interpretation of results until then. Reading results multiple times before the readout date increases the false positive rate.
How should I start interpreting results
Start with your Scorecard metrics and a hypothesis about what the experiment should drive; your primary metrics should answer that hypothesis. The delta is the observed difference between the test and control groups, and the error bars show the confidence interval: the range of values the true difference plausibly falls in.
Treat results as statistical evidence, not facts:
- A result that isn't statistically significant means you don't have sufficient evidence to reject the null hypothesis: given your experiment design, the observed result is reasonably likely to have happened by chance. Treat it as a lack of evidence for your hypothesis, and keep in mind that underpowered tests can produce neutral results even when a true effect exists.
- A statistically significant result means the probability of observing this result, or one more extreme, if the two groups were identical is below the threshold you set. Treat it as evidence for your hypothesis, but be wary when you've made multiple comparisons (many metrics, rerunning an experiment, or grouping by dimensions), because each comparison raises the chance of a false positive. A significant result from a test that was extremely unlikely to succeed has a high chance of being a false positive; consider reproducing the result, running a back-test, or lowering your significance level.
After the Scorecard, use the all-metrics tab and custom Explore queries to gather more information. Significant movements in those views aren't necessarily statistically sound, because examining more metrics increases the chance of a false positive. Use them to look for unexpected large regressions and to generate follow-up hypotheses. For more guidance, refer to Best practices and avoiding false positives.
Why are results missing for some metrics
Missing metric results usually happen when your company uses the SDK or event imports and also imports precomputed metrics from your data warehouse. Because these import processes can run at different times, data availability can differ. Adjust your analysis date range to get a full view of your data.
Why does my external source show more exposure events than Statsig
Statsig doesn't count exposures on the last day (the day you made a decision). Filter out that day when you analyze your external data. The hours that define a "day" for your project depend on your project's timezone.
Why doesn't Pulse show breakouts for the categorical metadata I log
Pulse shows results for sub-groups of a metric (for example, iOS vs. Android) only when you configure that metadata as a dimension. Event dimensions, which you log in the value or metadata fields of a custom event, are the most common dimension type. Define them in your custom event setup.
Why do I see "No dimensions available for this time range"
Precomputed user dimensions load asynchronously in separate Explore queries after the main Scorecard results load, so you may see this message shortly after the first reload of the day. Wait a few minutes and refresh the page. For details, refer to Custom Explore queries.
Was this helpful?