---
title: Benjamini–Hochberg Procedure
description: "How Statsig applies the Benjamini-Hochberg procedure to control the false discovery rate when analyzing many metrics in an experiment scorecard."
product: general
lang: en
last_updated: 2025-09-18
token_estimate: 717
---
# Benjamini–Hochberg Procedure

> For AI agents: a documentation index is available at [/llms.txt](/llms.txt). Append `.md` to any page URL for markdown, or send `Accept: text/markdown`.

The Benjamini-Hochberg (BH) procedure adjusts the significance level when a scorecard tests many metrics, so that only a controlled share of the results you call significant are false positives. That share is the false discovery rate. [Bonferroni correction](https://docs.statsig.com/experiments/statistical-methods/methodologies/bonferroni-correction) instead controls the family-wise error rate, the chance of at least one false positive. BH therefore rejects more null hypotheses than Bonferroni for the same p-values. Use BH when a scorecard has many metrics and a small, controlled share of false positives is acceptable. Use Bonferroni when any single false positive is costly.

You can enable the BH procedure for individual experiments, or configure global _Experiment Settings_ to use it by default.

![Benjamini-Hochberg procedure configuration interface](https://docs.statsig.com/images/snippets/stats-methods/benjamini-hochberg-procedure/c865494e-0ae4-489c-a416-45848b4d10bc.png)

## How BH adjusts the significance level

The BH procedure replaces your pre-set significance level ($\alpha$) with a new one. Statsig calculates it as follows:

1. Sort the metric p-values in ascending order.
2. Pair each p-value with a threshold. The threshold is the target false discovery rate ($q$) divided by the number of comparisons ($m$), multiplied by the rank ($k$) of that p-value in the ordered list.
3. Take the largest threshold that is higher than its paired p-value. That threshold becomes the new significance level ($\alpha$).

Statsig can apply the BH procedure across one of the following sets of p-values:

- **Test groups** (multiple treatment hypotheses): For each metric, Statsig aggregates the p-values from each variant and runs the BH procedure on that list.
- **Metrics in the scorecard**: For each variant, Statsig aggregates the p-values from each metric and runs the BH procedure on that list.
- **Both test groups and metrics**: Statsig aggregates all p-values and runs the BH procedure once.

Statsig doesn't apply the BH procedure to the p-values of event-dimension or user-property breakdowns of an experiment metric. Statsig compares only the top-line metric results to the new significance level.

## How the adjusted significance level appears in experiment results

In the experiment scorecard, Statsig derives confidence intervals for applicable metrics from (1 - adjusted $\alpha$). Hovering over a confidence interval displays the adjusted $\alpha$ alongside other metric details.

In the experiment explore section, Statsig calculates a new adjusted $\alpha$ based on your selections, and the confidence intervals use (1 - adjusted $\alpha$).

