Answer

False positive risk: how likely is a significant finding to be wrong?

Inspired by a question on Cross Validated ·

power analysishypothesis testingp-values

The short answer

Often more likely than alpha suggests. Alpha limits false alarms among true nulls, but the share of significant findings that are false also depends on power and on how many tested effects are real. If a quarter are real, about 31% of significant findings are false at 34% power and 16% at 80% power, and real effects that reach significance are overestimated. A stricter alpha cannot rescue a weak study; more power, pre-registration and replication can.

The short answer

A significance level of α = .05 promises one thing: when the null hypothesis is true, the test will wrongly reject it 5% of the time. It promises nothing about the question a reader actually cares about, which runs the other way: given that this result was significant, how likely is it that the effect is real?

That second probability depends on three things: α, the study's power, and the share of tested hypotheses that are true to begin with (often called the prior). Low power lowers the number of true effects that reach significance while leaving the number of false alarms unchanged, so a larger fraction of the significant results are false. So yes, a significant result from an underpowered study deserves less trust, for two reasons: it is more likely to be a false positive, and if the effect is real, its estimated size is probably too large.

Where the number comes from

Imagine a field that tests many hypotheses, where a proportion π of them are real effects. Out of all tests, a share α(1 − π) are true nulls that come out significant, and a share power × π are real effects that come out significant. The share of significant results that are false is therefore

false positive risk = α(1 − π) / [α(1 − π) + power × π]

This is sometimes called the false positive risk, or the science-wise false discovery rate, and it is the argument behind Ioannidis's (2005) paper "Why most published research findings are false". It is related to, but not the same as, the false discovery rate that the Benjamini-Hochberg procedure controls within one set of tests; for that, see when to use a Bonferroni correction. Two caveats: π is never known exactly, and the formula treats every p below .05 alike, although a p of .001 is much stronger evidence than a p of .049.

See it in R and Python

The code first evaluates the formula for several values of power and π (no randomness). It then simulates 20,000 two-group studies in which 25% test a real effect of d = 0.5, once with 20 people per group and once with 64, and runs a Student t test at α = .05 in each. Finally it asks what α a 20-per-group study would need for only 5% of its significant results to be false, recomputing power at each α.

R

set.seed(148603)
alpha <- 0.05
# Share of significant results that are false nulls
fpr <- function(power, prior, alpha = 0.05)
  alpha * (1 - prior) / (alpha * (1 - prior) + power * prior)

# 1. Exact values (no randomness): rows = power, columns = prior share of real effects
tab <- outer(c(0.2, 0.5, 0.8, 0.95), c(0.1, 0.25, 0.5), fpr)
dimnames(tab) <- list(power = c(0.2, 0.5, 0.8, 0.95), prior = c(0.1, 0.25, 0.5))
round(tab, 3)

# 2. Simulate 20,000 two-group studies; 25% test a real effect of d = 0.5
sim <- function(n, reps = 20000, prior = 0.25, d = 0.5) {
  real <- runif(reps) < prior
  x <- matrix(rnorm(reps * n), reps) + d * real
  y <- matrix(rnorm(reps * n), reps)
  sp <- sqrt((apply(x, 1, var) + apply(y, 1, var)) / 2)   # pooled SD
  dhat <- (rowMeans(x) - rowMeans(y)) / sp                 # Cohen's d
  p <- 2 * pt(-abs(dhat * sqrt(n / 2)), 2 * n - 2)         # Student t test
  sig <- p < alpha
  pow <- power.t.test(n = n, delta = d)$power
  c(n_per_group = n, power = pow, power_sim = mean(sig[real]),
    false_share_theory = fpr(pow, prior), false_share_sim = mean(!real[sig]),
    mean_d_real_hits = mean(dhat[sig & real]))
}
round(rbind(sim(20), sim(64)), 3)

# 3. Alpha that would make only 5% of hits false (n = 20, prior = .25);
#    power is recomputed at each alpha, because a stricter alpha lowers it
f <- function(a) fpr(power.t.test(n = 20, delta = 0.5, sig.level = a)$power, 0.25, a) - 0.05
a_star <- uniroot(f, c(1e-5, 0.05), tol = 1e-10)$root
signif(c(alpha = a_star,
         power = power.t.test(n = 20, delta = 0.5, sig.level = a_star)$power), 3)

Python

import numpy as np
from scipy import stats, optimize

rng = np.random.default_rng(148603)
alpha = 0.05

# Share of significant results that are false nulls
def fpr(power, prior, alpha=0.05):
    return alpha * (1 - prior) / (alpha * (1 - prior) + power * prior)

# Power of a two-sided two-sample t test, as in R's power.t.test()
def power_t(n, d, a=0.05):
    df = 2 * n - 2
    return stats.nct.sf(stats.t.ppf(1 - a / 2, df), df, d * np.sqrt(n / 2))

# 1. Exact values (no randomness): rows = power, columns = prior share of real effects
powers, priors = np.array([0.2, 0.5, 0.8, 0.95]), np.array([0.1, 0.25, 0.5])
print(np.round(fpr(powers[:, None], priors[None, :]), 3))

# 2. Simulate 20,000 two-group studies; 25% test a real effect of d = 0.5
def sim(n, reps=20000, prior=0.25, d=0.5):
    real = rng.random(reps) < prior
    x = rng.normal(size=(reps, n)) + d * real[:, None]
    y = rng.normal(size=(reps, n))
    sp = np.sqrt((x.var(ddof=1, axis=1) + y.var(ddof=1, axis=1)) / 2)  # pooled SD
    dhat = (x.mean(axis=1) - y.mean(axis=1)) / sp                       # Cohen's d
    p = 2 * stats.t.cdf(-np.abs(dhat * np.sqrt(n / 2)), 2 * n - 2)      # Student t test
    sig = p < alpha
    pow_ = power_t(n, d)
    return {"n_per_group": n, "power": pow_, "power_sim": sig[real].mean(),
            "false_share_theory": fpr(pow_, prior), "false_share_sim": (~real[sig]).mean(),
            "mean_d_real_hits": dhat[sig & real].mean()}

for n in (20, 64):
    print({k: round(float(v), 3) for k, v in sim(n).items()})

# 3. Alpha that would make only 5% of hits false (n = 20, prior = .25);
#    power is recomputed at each alpha, because a stricter alpha lowers it
f = lambda a: fpr(power_t(20, 0.5, a), 0.25, a) - 0.05
a_star = optimize.brentq(f, 1e-5, 0.05, xtol=1e-12)
print("alpha", float(f"{a_star:.3g}"), "power", float(f"{power_t(20, 0.5, a_star):.3g}"))
      prior
power    0.1  0.25   0.5
  0.2  0.692 0.429 0.200
  0.5  0.474 0.231 0.091
  0.8  0.360 0.158 0.059
  0.95 0.321 0.136 0.050
     n_per_group power power_sim false_share_theory false_share_sim
[1,]          20 0.338     0.334              0.308           0.304
[2,]          64 0.801     0.800              0.158           0.163
     mean_d_real_hits
[1,]            0.868
[2,]            0.565
   alpha    power 
0.000203 0.011600

The simulated figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated shares differ slightly (for example 29.9% false instead of 30.4% at 20 per group, and an average d of 0.865 instead of 0.868 among real effects that reached significance), but the pattern is the same. The formula table, the power values and the α in part 3 involve no randomness, and Python reproduces them exactly.

What should one researcher do about it?

For the mirror image, see is a non-significant result evidence for the null?, and for how α and power trade off, how alpha and beta are related. The DASS blog has more plain-language guides.

How to report a power analysis in APA style (7th edition)

Report the power analysis in the method section, naming the effect size it assumed, α, the number of tails and the target power. Using the designs from the example:

"An a priori power analysis for a two-tailed independent-samples t test (α = .05) indicated that 64 participants per group would give 80% power to detect an effect of d = 0.50."

If the study was smaller, say so in the discussion: "With 20 participants per group, the study had 34% power to detect an effect of d = 0.50, so significant estimates may overstate the true effect and should be replicated." Base the assumed effect size on prior research or the smallest effect of practical interest, never on the effect observed in the same study.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.