Answer
Bonferroni correction: when should you use it, and what are the alternatives?
The short answer
Bonferroni keeps the chance of any false positive across a set of tests at or below alpha, at a large cost in power. Use a familywise correction (preferably Holm, which is never less powerful) for a planned family of tests where one false claim is costly; for exploratory screening, control the false discovery rate with Benjamini-Hochberg. In one simulation of 20 tests, Bonferroni cut the familywise error rate from 52.7% to 3.3%, but power fell from .754 to .337.
The short answer
Every test run at α = .05 has a 5% chance of a false positive when its null is true. Run many tests and those chances add up: with 15 true nulls tested independently, the chance of at least one false positive is 1 − 0.95¹⁵, about 54%. A Bonferroni correction fixes this by testing each of m hypotheses at α/m (or, equivalently, multiplying each p-value by m). That guarantees the familywise error rate (FWER), the probability of one or more false positives in the whole set, stays at or below α.
The criticism is real but it is about when to use it, not whether it works. Bonferroni answers the question "is anything in this set of tests non-null?" It controls that error very conservatively, so it misses many real effects, and a given result's verdict changes with how many other tests happen to be in the analysis. Three practical rules:
- Correct when the tests form one planned family and a single false claim would matter, for example several confirmatory comparisons that all support one conclusion, or all pairwise comparisons after an ANOVA.
- Use Holm instead of plain Bonferroni. It controls the same FWER under the same conditions and always rejects at least as many hypotheses.
- For exploratory screening of many outcomes, control the false discovery rate (Benjamini-Hochberg) instead, and report every test you ran either way.
See it in R and Python
The code first adjusts one fixed set of ten p-values with Bonferroni, Holm and Benjamini-Hochberg (BH). It then simulates 2,000 studies, each running 20 two-sample t tests with 40 people per group, where 15 nulls are true and 5 tests have a real effect of d = 0.6. For each method it records how often a study had any false positive (FWER), the average share of significant results that were false (FDR), and the share of real effects detected (power).
R
set.seed(120362)
# 1. One family of 10 p-values: how the adjustments compare (no randomness)
p <- c(0.001, 0.004, 0.008, 0.012, 0.020, 0.030, 0.041, 0.120, 0.350, 0.700)
adj <- cbind(raw = p,
bonferroni = p.adjust(p, "bonferroni"),
holm = p.adjust(p, "holm"),
BH = p.adjust(p, "BH"))
round(adj, 3)
colSums(adj < 0.05) # number of tests significant at .05
# 2. Simulation: 20 two-sample t tests per study (40 per group);
# 15 true nulls and 5 real effects of d = 0.6
m <- 20; real <- 1:5; reps <- 2000
methods <- c("none", "bonferroni", "holm", "BH")
res <- replicate(reps, {
p <- sapply(1:m, function(j) {
d <- if (j %in% real) 0.6 else 0
t.test(rnorm(40, d), rnorm(40), var.equal = TRUE)$p.value
})
sapply(methods, function(mt) {
rej <- p.adjust(p, mt) < 0.05
n_false <- sum(rej[-real])
c(fwer = n_false > 0, # any false positive
fdr = if (any(rej)) n_false / sum(rej) else 0, # share of hits that are false
power = mean(rej[real])) # share of real effects found
})
})
round(apply(res, c(1, 2), mean), 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(120362)
# Same definitions as R's p.adjust()
def p_adjust(p, method):
p = np.asarray(p, float); m = len(p)
if method == "none":
return p
if method == "bonferroni":
return np.minimum(1, p * m)
o = np.argsort(p); out = np.empty(m)
if method == "holm":
out[o] = np.minimum(1, np.maximum.accumulate((m - np.arange(m)) * p[o]))
else: # "BH"
q = np.minimum.accumulate((m / np.arange(m, 0, -1)) * p[o][::-1])[::-1]
out[o] = np.minimum(1, q)
return out
methods = ["none", "bonferroni", "holm", "BH"]
# 1. One family of 10 p-values: how the adjustments compare (no randomness)
p = [0.001, 0.004, 0.008, 0.012, 0.020, 0.030, 0.041, 0.120, 0.350, 0.700]
for mt in methods:
adj = p_adjust(p, mt)
print(mt, np.round(adj, 3), "significant:", int(np.sum(adj < 0.05)))
# 2. Simulation: 20 two-sample t tests per study (40 per group);
# 15 true nulls and 5 real effects of d = 0.6
m, n_real, reps = 20, 5, 2000
d = np.r_[np.full(n_real, 0.6), np.zeros(m - n_real)]
x = rng.normal(size=(reps, m, 40)) + d[None, :, None]
y = rng.normal(size=(reps, m, 40))
pv = stats.ttest_ind(x, y, axis=2).pvalue
for mt in methods:
rej = np.array([p_adjust(row, mt) < 0.05 for row in pv])
n_false = rej[:, n_real:].sum(axis=1); n_rej = rej.sum(axis=1)
fdp = np.where(n_rej > 0, n_false / np.maximum(n_rej, 1), 0)
print(mt, "fwer", round(float(np.mean(n_false > 0)), 3),
"fdr", round(float(fdp.mean()), 3),
"power", round(float(rej[:, :n_real].mean()), 3)) raw bonferroni holm BH
[1,] 0.001 0.01 0.010 0.010
[2,] 0.004 0.04 0.036 0.020
[3,] 0.008 0.08 0.064 0.027
[4,] 0.012 0.12 0.084 0.030
[5,] 0.020 0.20 0.120 0.040
[6,] 0.030 0.30 0.150 0.050
[7,] 0.041 0.41 0.164 0.059
[8,] 0.120 1.00 0.360 0.150
[9,] 0.350 1.00 0.700 0.389
[10,] 0.700 1.00 0.700 0.700
raw bonferroni holm BH
7 2 2 5
none bonferroni holm BH
fwer 0.527 0.033 0.036 0.118
fdr 0.148 0.013 0.014 0.036
power 0.754 0.337 0.346 0.469
The simulated rates come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly (for example an FWER of 4.3% instead of 3.3% for Bonferroni, and power of .487 instead of .469 for BH), but they show the same pattern. The adjusted p-values in part 1 involve no randomness, and Python reproduces them exactly.
What the numbers show
- Without a correction, false positives are routine. 52.7% of simulated studies reported at least one false positive, close to the 54% the formula predicts for 15 independent true nulls.
- Bonferroni and Holm bring that down to the target. Their FWER was 3.3% and 3.6%, below .05. Bonferroni is a bit conservative because the bound it uses is not tight.
- The price is power. The share of real effects detected fell from .754 with no correction to .337 with Bonferroni and .346 with Holm. Most real effects were missed.
- BH sits in between. It let at least one false positive through in 11.8% of studies, so it does not control FWER, but only 3.6% of its significant results were false on average, under the 5% it targets, and it found .469 of the real effects.
- On the fixed set of ten p-values, 7 were below .05 unadjusted, 2 survived Bonferroni or Holm, and 5 survived BH. Holm's adjusted values are never larger than Bonferroni's, which is why it is the safer default of the two.
Answering the main objections
"It tests a null nobody cares about." Partly true. Bonferroni protects against a false positive anywhere in the set, which is exactly the right target when a paper's conclusion would collapse if any one of its claims were wrong, and the wrong target when each test is its own separate question. Decide what the family is before you look at the data: the comparisons that jointly support one claim belong together; unrelated questions in the same dataset usually do not.
"A result's meaning depends on how many other tests you ran." That is the point of a familywise correction, but it is also why the family must be defined by the research question rather than by everything in the spreadsheet. Pre-registering a small number of primary tests keeps the correction mild; the rest can be labelled exploratory.
"It inflates Type II errors." Yes, as the simulation shows. If you will correct, plan the sample size for the corrected α (see how alpha and beta are related, which works through the power cost of a Bonferroni correction in detail). That page covers the sample-size side; this one covers when to correct and which method to pick.
"Just describe what you tested." Transparency is necessary in any case: report every test, not only the significant ones. But a reader cannot mentally correct 20 unadjusted p-values, so pair full reporting with an adjustment that matches your goal. Treating correlated outcomes as independent makes Bonferroni and Holm even more conservative; BH remains valid under positive dependence, which covers many common cases. Our test chooser and the DASS blog have more on picking and reporting tests.
How to report adjusted p-values in APA style (7th edition)
Name the correction and the size of the family in the method or results section, then give adjusted p-values (or say that the threshold was adjusted). Using the ten p-values from the example above:
"To control the familywise error rate across 10 comparisons, p-values were adjusted with the Holm method. Two comparisons remained significant (adjusted p = .010 and adjusted p = .036)."
"To control the false discovery rate at .05 across 10 outcomes, we applied the Benjamini-Hochberg procedure; five outcomes remained significant (adjusted ps = .010 to .040)." In a table, add a column of adjusted p-values next to the unadjusted ones and state the method in a table note.
Related tools and guides
- Check a reported p-value
- How are alpha and beta (Type I and Type II error rates) related?
- What do p values and t values mean?
- Does a non-significant result support the null?
- Which statistical test should I use?
- Multiple comparisons problem (Wikipedia)
More answered questions
- Is a non-significant result from a large study evidence for the null?
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
- What do p values and t values actually mean?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.