Answer

Why does almost everything become statistically significant with a large sample?

Inspired by a question on Cross Validated ·

hypothesis testingp-valueseffect sizesample size

The short answer

A significance test is not biased by a large sample: if the null is exactly true, it rejects 5% of the time at any n. What changes is power. Real differences are rarely exactly zero, and with enough data even a trivial one is detected. In one simulated study with 20,000 per group and a true d of 0.05, p was about .000002 while the effect stayed negligible. Judge size with an effect size and its confidence interval, or test against a smallest effect that matters.

The short answer

The claim that tests become "biased toward rejection" in big samples mixes up two different things. When the null hypothesis is exactly true, a correctly applied test at α = .05 rejects it about 5% of the time whether you have 20 observations or 20 million. That is the Type I error rate, and sample size does not inflate it.

What sample size does change is power, the chance of detecting a difference that really exists. The standard error shrinks roughly in proportion to 1/√n, so a fixed difference produces a larger and larger test statistic as n grows. In practice two groups, or a coefficient and zero, are almost never exactly equal: tiny differences come from real but negligible effects, from small imbalances in who was sampled, or from slight measurement artifacts. With enough data the test correctly reports that the difference is not zero. The problem is reading "not zero" as "large" or "important".

See it in R and Python

The code does three things. It simulates 2,000 two-sample t tests with a true null at three sample sizes and records how often p < .05. It then computes the exact power to detect a tiny standardised difference of d = 0.05 at four sample sizes (no randomness involved). Finally it runs one large study, 20,000 people per group, with that tiny difference built in.

R

set.seed(108911)

# 1. When the null is exactly true, the false-positive rate stays at 5% for any n
for (n in c(20, 200, 2000)) {
  p <- replicate(2000, t.test(rnorm(n), rnorm(n), var.equal = TRUE)$p.value)
  cat("null true, n per group =", n, " rejection rate =", mean(p < 0.05), "\n")
}

# 2. A tiny real difference (d = 0.05): power grows with n (exact, no randomness)
for (n in c(50, 500, 5000, 20000)) {
  pw <- power.t.test(n = n, delta = 0.05, sd = 1, sig.level = 0.05,
                     strict = TRUE)$power   # count both rejection tails
  cat("d = 0.05, n per group =", n, " power =", round(pw, 3), "\n")
}

# 3. One large study with that tiny difference
n <- 20000
x <- rnorm(n, mean = 0.05); y <- rnorm(n, mean = 0)
tt <- t.test(x, y, var.equal = TRUE)
d <- (mean(x) - mean(y)) / sqrt((var(x) + var(y)) / 2)   # Cohen's d (equal n)
cat("t =", round(tt$statistic, 2), " df =", tt$parameter,
    " p =", signif(tt$p.value, 3), "\n")
cat("mean difference =", round(mean(x) - mean(y), 3),
    " 95% CI [", round(tt$conf.int[1], 3), ",", round(tt$conf.int[2], 3), "]\n")
cat("Cohen's d =", round(d, 3), "\n")

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(108911)

# 1. When the null is exactly true, the false-positive rate stays at 5% for any n
for n in [20, 200, 2000]:
    p = stats.ttest_ind(rng.normal(size=(2000, n)), rng.normal(size=(2000, n)), axis=1).pvalue
    print("null true, n per group =", n, " rejection rate =", np.mean(p < 0.05))

# 2. A tiny real difference (d = 0.05): power grows with n (exact, no randomness)
#    Same as R's power.t.test(..., strict = TRUE): both rejection tails counted
for n in [50, 500, 5000, 20000]:
    df = 2 * n - 2; ncp = 0.05 * np.sqrt(n / 2); crit = stats.t.ppf(0.975, df)
    pw = stats.nct.sf(crit, df, ncp) + stats.nct.cdf(-crit, df, ncp)
    print("d = 0.05, n per group =", n, " power =", round(pw, 3))

# 3. One large study with that tiny difference
n = 20000
x = rng.normal(0.05, 1, n); y = rng.normal(0, 1, n)
tt = stats.ttest_ind(x, y)
diff = x.mean() - y.mean()
sp = np.sqrt((x.var(ddof=1) + y.var(ddof=1)) / 2)        # pooled SD (equal n)
se = sp * np.sqrt(2 / n); crit = stats.t.ppf(0.975, 2 * n - 2)
print("t =", round(tt.statistic, 2), " df =", 2 * n - 2, " p =", f"{tt.pvalue:.3g}")
print("mean difference =", round(diff, 3),
      " 95% CI [", round(diff - crit * se, 3), ",", round(diff + crit * se, 3), "]")
print("Cohen's d =", round(diff / sp, 3))
null true, n per group = 20  rejection rate = 0.044 
null true, n per group = 200  rejection rate = 0.0515 
null true, n per group = 2000  rejection rate = 0.042 
d = 0.05, n per group = 50  power = 0.057 
d = 0.05, n per group = 500  power = 0.124 
d = 0.05, n per group = 5000  power = 0.705 
d = 0.05, n per group = 20000  power = 0.999 
t = 4.77  df = 39998  p = 1.89e-06 
mean difference = 0.048  95% CI [ 0.028 , 0.067 ]
Cohen's d = 0.048

The simulated figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated figures differ slightly (rejection rates of .0485 to .0545 under the null, and d = 0.062 with p far below .001 in the large study), but they show the same pattern. The power values involve no randomness, and Python reproduces them exactly.

What the numbers show

How to guard against it

A Bayesian analysis reaches a related conclusion from another direction: with a very large sample, a p-value just under .05 can go with a Bayes factor that favours the null (Lindley's paradox), because such a small test statistic is more consistent with a near-zero effect than with a meaningful one. Either way, the fix is to ask how big the effect is, not only whether it is zero. The reverse mistake, reading a non-significant result in a small study as proof of no effect, is covered in does a non-significant result support the null?. Our test chooser and the DASS blog cover choosing and reporting tests.

How to report a t test in APA style (7th edition)

Report the test, the exact p-value (or p < .001 when it is smaller), and an effect size with its interval, then interpret the size, not just the significance. Using the large study from the example above:

"The groups differed significantly, t(39998) = 4.77, p < .001, but the difference was negligible in size: the mean difference was 0.05, 95% CI [0.03, 0.07], d = 0.05."

If you set a smallest effect of interest in advance, add it to the method section (for example, "differences smaller than d = x.xx were considered too small to matter") and report the equivalence test against it in the same sentence as the t test.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.