Answer

Bonferroni correction: when should you use it, and what are the alternatives?

Inspired by a question on Cross Validated ·

multiple comparisonshypothesis testingp-values

The short answer

Bonferroni keeps the chance of any false positive across a set of tests at or below alpha, at a large cost in power. Use a familywise correction (preferably Holm, which is never less powerful) for a planned family of tests where one false claim is costly; for exploratory screening, control the false discovery rate with Benjamini-Hochberg. In one simulation of 20 tests, Bonferroni cut the familywise error rate from 52.7% to 3.3%, but power fell from .754 to .337.

The short answer

Every test run at α = .05 has a 5% chance of a false positive when its null is true. Run many tests and those chances add up: with 15 true nulls tested independently, the chance of at least one false positive is 1 − 0.95¹⁵, about 54%. A Bonferroni correction fixes this by testing each of m hypotheses at α/m (or, equivalently, multiplying each p-value by m). That guarantees the familywise error rate (FWER), the probability of one or more false positives in the whole set, stays at or below α.

The criticism is real but it is about when to use it, not whether it works. Bonferroni answers the question "is anything in this set of tests non-null?" It controls that error very conservatively, so it misses many real effects, and a given result's verdict changes with how many other tests happen to be in the analysis. Three practical rules:

See it in R and Python

The code first adjusts one fixed set of ten p-values with Bonferroni, Holm and Benjamini-Hochberg (BH). It then simulates 2,000 studies, each running 20 two-sample t tests with 40 people per group, where 15 nulls are true and 5 tests have a real effect of d = 0.6. For each method it records how often a study had any false positive (FWER), the average share of significant results that were false (FDR), and the share of real effects detected (power).

R

set.seed(120362)

# 1. One family of 10 p-values: how the adjustments compare (no randomness)
p <- c(0.001, 0.004, 0.008, 0.012, 0.020, 0.030, 0.041, 0.120, 0.350, 0.700)
adj <- cbind(raw = p,
             bonferroni = p.adjust(p, "bonferroni"),
             holm = p.adjust(p, "holm"),
             BH = p.adjust(p, "BH"))
round(adj, 3)
colSums(adj < 0.05)   # number of tests significant at .05

# 2. Simulation: 20 two-sample t tests per study (40 per group);
#    15 true nulls and 5 real effects of d = 0.6
m <- 20; real <- 1:5; reps <- 2000
methods <- c("none", "bonferroni", "holm", "BH")
res <- replicate(reps, {
  p <- sapply(1:m, function(j) {
    d <- if (j %in% real) 0.6 else 0
    t.test(rnorm(40, d), rnorm(40), var.equal = TRUE)$p.value
  })
  sapply(methods, function(mt) {
    rej <- p.adjust(p, mt) < 0.05
    n_false <- sum(rej[-real])
    c(fwer  = n_false > 0,                                   # any false positive
      fdr   = if (any(rej)) n_false / sum(rej) else 0,      # share of hits that are false
      power = mean(rej[real]))                               # share of real effects found
  })
})
round(apply(res, c(1, 2), mean), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(120362)

# Same definitions as R's p.adjust()
def p_adjust(p, method):
    p = np.asarray(p, float); m = len(p)
    if method == "none":
        return p
    if method == "bonferroni":
        return np.minimum(1, p * m)
    o = np.argsort(p); out = np.empty(m)
    if method == "holm":
        out[o] = np.minimum(1, np.maximum.accumulate((m - np.arange(m)) * p[o]))
    else:  # "BH"
        q = np.minimum.accumulate((m / np.arange(m, 0, -1)) * p[o][::-1])[::-1]
        out[o] = np.minimum(1, q)
    return out

methods = ["none", "bonferroni", "holm", "BH"]

# 1. One family of 10 p-values: how the adjustments compare (no randomness)
p = [0.001, 0.004, 0.008, 0.012, 0.020, 0.030, 0.041, 0.120, 0.350, 0.700]
for mt in methods:
    adj = p_adjust(p, mt)
    print(mt, np.round(adj, 3), "significant:", int(np.sum(adj < 0.05)))

# 2. Simulation: 20 two-sample t tests per study (40 per group);
#    15 true nulls and 5 real effects of d = 0.6
m, n_real, reps = 20, 5, 2000
d = np.r_[np.full(n_real, 0.6), np.zeros(m - n_real)]
x = rng.normal(size=(reps, m, 40)) + d[None, :, None]
y = rng.normal(size=(reps, m, 40))
pv = stats.ttest_ind(x, y, axis=2).pvalue
for mt in methods:
    rej = np.array([p_adjust(row, mt) < 0.05 for row in pv])
    n_false = rej[:, n_real:].sum(axis=1); n_rej = rej.sum(axis=1)
    fdp = np.where(n_rej > 0, n_false / np.maximum(n_rej, 1), 0)
    print(mt, "fwer", round(float(np.mean(n_false > 0)), 3),
          "fdr", round(float(fdp.mean()), 3),
          "power", round(float(rej[:, :n_real].mean()), 3))
        raw bonferroni  holm    BH
 [1,] 0.001       0.01 0.010 0.010
 [2,] 0.004       0.04 0.036 0.020
 [3,] 0.008       0.08 0.064 0.027
 [4,] 0.012       0.12 0.084 0.030
 [5,] 0.020       0.20 0.120 0.040
 [6,] 0.030       0.30 0.150 0.050
 [7,] 0.041       0.41 0.164 0.059
 [8,] 0.120       1.00 0.360 0.150
 [9,] 0.350       1.00 0.700 0.389
[10,] 0.700       1.00 0.700 0.700
       raw bonferroni       holm         BH 
         7          2          2          5 
       none bonferroni  holm    BH
fwer  0.527      0.033 0.036 0.118
fdr   0.148      0.013 0.014 0.036
power 0.754      0.337 0.346 0.469

The simulated rates come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly (for example an FWER of 4.3% instead of 3.3% for Bonferroni, and power of .487 instead of .469 for BH), but they show the same pattern. The adjusted p-values in part 1 involve no randomness, and Python reproduces them exactly.

What the numbers show

Answering the main objections

"It tests a null nobody cares about." Partly true. Bonferroni protects against a false positive anywhere in the set, which is exactly the right target when a paper's conclusion would collapse if any one of its claims were wrong, and the wrong target when each test is its own separate question. Decide what the family is before you look at the data: the comparisons that jointly support one claim belong together; unrelated questions in the same dataset usually do not.

"A result's meaning depends on how many other tests you ran." That is the point of a familywise correction, but it is also why the family must be defined by the research question rather than by everything in the spreadsheet. Pre-registering a small number of primary tests keeps the correction mild; the rest can be labelled exploratory.

"It inflates Type II errors." Yes, as the simulation shows. If you will correct, plan the sample size for the corrected α (see how alpha and beta are related, which works through the power cost of a Bonferroni correction in detail). That page covers the sample-size side; this one covers when to correct and which method to pick.

"Just describe what you tested." Transparency is necessary in any case: report every test, not only the significant ones. But a reader cannot mentally correct 20 unadjusted p-values, so pair full reporting with an adjustment that matches your goal. Treating correlated outcomes as independent makes Bonferroni and Holm even more conservative; BH remains valid under positive dependence, which covers many common cases. Our test chooser and the DASS blog have more on picking and reporting tests.

How to report adjusted p-values in APA style (7th edition)

Name the correction and the size of the family in the method or results section, then give adjusted p-values (or say that the threshold was adjusted). Using the ten p-values from the example above:

"To control the familywise error rate across 10 comparisons, p-values were adjusted with the Holm method. Two comparisons remained significant (adjusted p = .010 and adjusted p = .036)."

"To control the false discovery rate at .05 across 10 outcomes, we applied the Benjamini-Hochberg procedure; five outcomes remained significant (adjusted ps = .010 to .040)." In a table, add a column of adjusted p-values next to the unadjusted ones and state the method in a table note.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.