For AI agents: a documentation index is available at /llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.
Reconciling Results Between Experimentation Platforms
Learn how to reconcile differences in experiment results between different analysis platforms.
Experiment results can differ between analysis platforms because the same data supports many analysis methodologies. Most differences trace back to one of three causes, and you can resolve them by working through the causes in order:
- The two platforms read or join the metric source data to exposure data differently, which invalidates every downstream step.
- Statistical features that one platform applies and the other doesn't, most often outlier trimming or pre-experiment bias reduction, are working as intended.
- A metric definition, or an advanced configuration on a metric or experiment, behaves differently than you expected.
Use this page when a Statsig result doesn't match a number from your in-house platform or another platform for the same experiment. If results within Statsig look wrong on their own, start with the Pulse FAQ instead.
Joining data
Differences in experiment results most often stem from how a platform joins exposure data with metric data.
ID formats
In some cases, systems log IDs in different formats to different places. For example, the binary ID 4TLCtqzctSqusYcQljJLJE maps to the UUID a0fb4ef0-9d9e-11eb-9462-7bfc2b9a6ff2, so a company might have the binary ID in its production environment while its data users work with the equivalent UUIDs.
Exposures logged with the binary ID can't join with metric data that uses the UUID, and results are empty. Check samples from both the metric source and the assignment source or diagnostic logstream to confirm that the identifiers are in the same format.
You can use ID Resolution to bridge ID type gaps, but Statsig didn't design it for this scenario. ID Resolution connects identifiers across logged-out and logged-in sessions, or other scenarios where users switch identifiers during the experiment.
Timestamps
Analyze metric data only after Statsig exposes a user to the experiment. Pre-experiment data should have no average treatment effect, so including it dilutes results.
Statsig Cloud uses a date-based join between exposures and metric data: experiments include metric data from the whole of the first exposure date for each experimental unit. While this can include some pre-experiment metric data, the average treatment effect of this dilution should be null. Statsig Warehouse Native uses a timestamp-based join, with an option for a date-based join for daily data. The following query shows the Cloud date-based join. The Warehouse Native join is identical except that the join condition compares metrics.timestamp >= exposures.first_timestamp instead of metrics.date_id >= exposures.first_date_id.
WITH
metrics as (...),
exposures as (...),
joined_data as (
SELECT
exposures.unit_id,
exposures.experiment_id,
exposures.group_id,
metrics.timestamp,
metrics.value
FROM exposures
JOIN metrics
ON (
exposures.unit_id = metrics.unit_id
AND metrics.date_id >= exposures.first_date_id
)
)
SELECT
group_id,
SUM(value) as value
FROM joined_data
GROUP BY group_id;
Statsig's exposure timestamps are always in UTC. If your metric data is in another timezone, adjust it so the join doesn't filter on mismatched dates.
Statsig supports timestamp-based joins for some Enterprise Cloud customers. Contact Statsig to learn more.
Exposure duplication
De-duplicate exposure data before joining so that each user has a single record. Many platforms also manage crossover users (users present in more than one experiment group) by removing them from the analysis or alerting when crossovers occur at high frequency.
SELECT
unit_id,
experiment_id,
MIN(timestamp) as first_timestamp,
COUNT(distinct group_id) as groups
FROM <exposures_table>
GROUP BY
unit_id,
experiment_id,
group_id
HAVING COUNT(distinct group_id) = 1;
Data availability
When you compare a platform analysis to an experiment analysis that ran in the past, the underlying data may have aged out of retention or no longer exist. Compare the table's retention policy to the analysis dates in your original experiment analysis to confirm that the data still exists. Also confirm that you configured your experiment in Statsig to analyze the same time range as your original analysis.
Validation
To validate the initial metric data and join, run the date-based join query from the Timestamps section on both platforms, adjusted for each platform's SQL dialect. Confirm that a target metric has the same totals per group across both platforms. Statsig provides the intermediate and result datasets it uses, and the queries behind its analysis, so you can trace where a gap arises. Warehouse Native projects have an advantage here because the SQL dialect and source data are generally the same in Statsig's queries and in your in-house code, which makes comparisons simpler.
Pick one metric of interest, validate that data, and resolve any differences before you check statistical and metric methodologies.
Statistical features
Choices in statistical methodology can change experiment results. The following features are common root causes of gaps. Read the queries Statsig runs closely to understand the particulars of its methodology.
Winsorization
Outlier trimming, or Winsorization, changes experiment outcomes by capping extreme values. Disable it in Statsig metrics when you compare across systems, unless you also apply it manually on the other platform.
CUPED
CUPED changes variances and observed deltas, especially when pre- and post-exposure data are highly correlated or when groups differ systematically in their pre-experiment data. You can configure CUPED at the metric level, and you can disable it for a Pulse result set after running the analysis.
Ratio metrics
For ratio metrics that use the delta method, Statsig includes only units with a non-zero denominator. Statsig calculates ratios and means as
and uses the delta method to correct for the cluster-based nature of these metrics.
Metric definitions
Users often misunderstand how Statsig calculates a given metric. Refer to the metrics guide for details.
Was this helpful?