Answer
False positive risk: how likely is a significant finding to be wrong?
The short answer
Often more likely than alpha suggests. Alpha limits false alarms among true nulls, but the share of significant findings that are false also depends on power and on how many tested effects are real. If a quarter are real, about 31% of significant findings are false at 34% power and 16% at 80% power, and real effects that reach significance are overestimated. A stricter alpha cannot rescue a weak study; more power, pre-registration and replication can.
The short answer
A significance level of α = .05 promises one thing: when the null hypothesis is true, the test will wrongly reject it 5% of the time. It promises nothing about the question a reader actually cares about, which runs the other way: given that this result was significant, how likely is it that the effect is real?
That second probability depends on three things: α, the study's power, and the share of tested hypotheses that are true to begin with (often called the prior). Low power lowers the number of true effects that reach significance while leaving the number of false alarms unchanged, so a larger fraction of the significant results are false. So yes, a significant result from an underpowered study deserves less trust, for two reasons: it is more likely to be a false positive, and if the effect is real, its estimated size is probably too large.
Where the number comes from
Imagine a field that tests many hypotheses, where a proportion π of them are real effects. Out of all tests, a share α(1 − π) are true nulls that come out significant, and a share power × π are real effects that come out significant. The share of significant results that are false is therefore
false positive risk = α(1 − π) / [α(1 − π) + power × π]
This is sometimes called the false positive risk, or the science-wise false discovery rate, and it is the argument behind Ioannidis's (2005) paper "Why most published research findings are false". It is related to, but not the same as, the false discovery rate that the Benjamini-Hochberg procedure controls within one set of tests; for that, see when to use a Bonferroni correction. Two caveats: π is never known exactly, and the formula treats every p below .05 alike, although a p of .001 is much stronger evidence than a p of .049.
See it in R and Python
The code first evaluates the formula for several values of power and π (no randomness). It then simulates 20,000 two-group studies in which 25% test a real effect of d = 0.5, once with 20 people per group and once with 64, and runs a Student t test at α = .05 in each. Finally it asks what α a 20-per-group study would need for only 5% of its significant results to be false, recomputing power at each α.
R
set.seed(148603)
alpha <- 0.05
# Share of significant results that are false nulls
fpr <- function(power, prior, alpha = 0.05)
alpha * (1 - prior) / (alpha * (1 - prior) + power * prior)
# 1. Exact values (no randomness): rows = power, columns = prior share of real effects
tab <- outer(c(0.2, 0.5, 0.8, 0.95), c(0.1, 0.25, 0.5), fpr)
dimnames(tab) <- list(power = c(0.2, 0.5, 0.8, 0.95), prior = c(0.1, 0.25, 0.5))
round(tab, 3)
# 2. Simulate 20,000 two-group studies; 25% test a real effect of d = 0.5
sim <- function(n, reps = 20000, prior = 0.25, d = 0.5) {
real <- runif(reps) < prior
x <- matrix(rnorm(reps * n), reps) + d * real
y <- matrix(rnorm(reps * n), reps)
sp <- sqrt((apply(x, 1, var) + apply(y, 1, var)) / 2) # pooled SD
dhat <- (rowMeans(x) - rowMeans(y)) / sp # Cohen's d
p <- 2 * pt(-abs(dhat * sqrt(n / 2)), 2 * n - 2) # Student t test
sig <- p < alpha
pow <- power.t.test(n = n, delta = d)$power
c(n_per_group = n, power = pow, power_sim = mean(sig[real]),
false_share_theory = fpr(pow, prior), false_share_sim = mean(!real[sig]),
mean_d_real_hits = mean(dhat[sig & real]))
}
round(rbind(sim(20), sim(64)), 3)
# 3. Alpha that would make only 5% of hits false (n = 20, prior = .25);
# power is recomputed at each alpha, because a stricter alpha lowers it
f <- function(a) fpr(power.t.test(n = 20, delta = 0.5, sig.level = a)$power, 0.25, a) - 0.05
a_star <- uniroot(f, c(1e-5, 0.05), tol = 1e-10)$root
signif(c(alpha = a_star,
power = power.t.test(n = 20, delta = 0.5, sig.level = a_star)$power), 3)Python
import numpy as np
from scipy import stats, optimize
rng = np.random.default_rng(148603)
alpha = 0.05
# Share of significant results that are false nulls
def fpr(power, prior, alpha=0.05):
return alpha * (1 - prior) / (alpha * (1 - prior) + power * prior)
# Power of a two-sided two-sample t test, as in R's power.t.test()
def power_t(n, d, a=0.05):
df = 2 * n - 2
return stats.nct.sf(stats.t.ppf(1 - a / 2, df), df, d * np.sqrt(n / 2))
# 1. Exact values (no randomness): rows = power, columns = prior share of real effects
powers, priors = np.array([0.2, 0.5, 0.8, 0.95]), np.array([0.1, 0.25, 0.5])
print(np.round(fpr(powers[:, None], priors[None, :]), 3))
# 2. Simulate 20,000 two-group studies; 25% test a real effect of d = 0.5
def sim(n, reps=20000, prior=0.25, d=0.5):
real = rng.random(reps) < prior
x = rng.normal(size=(reps, n)) + d * real[:, None]
y = rng.normal(size=(reps, n))
sp = np.sqrt((x.var(ddof=1, axis=1) + y.var(ddof=1, axis=1)) / 2) # pooled SD
dhat = (x.mean(axis=1) - y.mean(axis=1)) / sp # Cohen's d
p = 2 * stats.t.cdf(-np.abs(dhat * np.sqrt(n / 2)), 2 * n - 2) # Student t test
sig = p < alpha
pow_ = power_t(n, d)
return {"n_per_group": n, "power": pow_, "power_sim": sig[real].mean(),
"false_share_theory": fpr(pow_, prior), "false_share_sim": (~real[sig]).mean(),
"mean_d_real_hits": dhat[sig & real].mean()}
for n in (20, 64):
print({k: round(float(v), 3) for k, v in sim(n).items()})
# 3. Alpha that would make only 5% of hits false (n = 20, prior = .25);
# power is recomputed at each alpha, because a stricter alpha lowers it
f = lambda a: fpr(power_t(20, 0.5, a), 0.25, a) - 0.05
a_star = optimize.brentq(f, 1e-5, 0.05, xtol=1e-12)
print("alpha", float(f"{a_star:.3g}"), "power", float(f"{power_t(20, 0.5, a_star):.3g}")) prior
power 0.1 0.25 0.5
0.2 0.692 0.429 0.200
0.5 0.474 0.231 0.091
0.8 0.360 0.158 0.059
0.95 0.321 0.136 0.050
n_per_group power power_sim false_share_theory false_share_sim
[1,] 20 0.338 0.334 0.308 0.304
[2,] 64 0.801 0.800 0.158 0.163
mean_d_real_hits
[1,] 0.868
[2,] 0.565
alpha power
0.000203 0.011600
The simulated figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated shares differ slightly (for example 29.9% false instead of 30.4% at 20 per group, and an average d of 0.865 instead of 0.868 among real effects that reached significance), but the pattern is the same. The formula table, the power values and the α in part 3 involve no randomness, and Python reproduces them exactly.
- Power and the prior both matter. With half of tested hypotheses true and 80% power, 5.9% of significant results are false. With 20% power that rises to 20.0%, and if only 10% of hypotheses are true it reaches 69.2%.
- The simulation matches the formula. At 20 per group (power .338), 30.4% of significant results were false nulls, against 30.8% predicted. At 64 per group (power .801), the figures were 16.3% and 15.8%. Note that good power does not bring this down to 5% when real effects are rare.
- Real effects that get through are inflated. The true effect was d = 0.5, but among real effects that reached significance the average estimate was 0.868 at 20 per group and 0.565 at 64 per group. A small study can only reach significance when it overestimates the effect.
- A stricter α is not a fix. For only 5% of significant results to be false at 20 per group, α would have to be about .0002, and power would then fall to 1.2%. The study would almost never find anything.
What should one researcher do about it?
- Fix power at the design stage. The most effective step is a sample size that gives good power for a realistic effect, ideally a smaller one than early studies reported, since those are likely to be inflated. The power tools listed under related tools below can help, and the test chooser helps you pick the test to plan for.
- Do not tune α after the fact. Choosing α from a guess about power and the prior, once the data are in, adds another flexible decision. If you want a stricter threshold, set it in advance; some authors have proposed .005 for claims of new discoveries (Benjamin et al., 2018).
- Report the estimate, not just the verdict. Give the effect size with a confidence interval and state the study's power for a plausible effect, so readers can weigh the result themselves. Publishing a well-described underpowered study is better than hiding it, because meta-analyses need it.
- Treat a single significant result as provisional. Pre-register the hypothesis and analysis, and replicate before building on a finding. See what to do after rejecting the null, which covers estimation and replication in more detail.
- Judge plausibility honestly. A surprising hypothesis has a low prior, so even a well-powered significant result for it is weaker evidence than the p value suggests.
For the mirror image, see is a non-significant result evidence for the null?, and for how α and power trade off, how alpha and beta are related. The DASS blog has more plain-language guides.
How to report a power analysis in APA style (7th edition)
Report the power analysis in the method section, naming the effect size it assumed, α, the number of tails and the target power. Using the designs from the example:
"An a priori power analysis for a two-tailed independent-samples t test (α = .05) indicated that 64 participants per group would give 80% power to detect an effect of d = 0.50."
If the study was smaller, say so in the discussion: "With 20 participants per group, the study had 34% power to detect an effect of d = 0.50, so significant estimates may overstate the true effect and should be replicated." Base the assumed effect size on prior research or the smallest effect of practical interest, never on the effect observed in the same study.
Related tools and guides
- Power and sample size calculators
- Check a reported p-value
- What should you do after rejecting the null hypothesis?
- Bonferroni correction: when should you use it?
- How are alpha and beta related?
- Which statistical test should I use?
- Why most published research findings are false (Wikipedia)
More answered questions
- Is a non-significant result from a large study evidence for the null?
- What should you do after rejecting the null hypothesis?
- How to interpret the F statistic and p value in a one-way ANOVA
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.