Answer
Why does almost everything become statistically significant with a large sample?
The short answer
A significance test is not biased by a large sample: if the null is exactly true, it rejects 5% of the time at any n. What changes is power. Real differences are rarely exactly zero, and with enough data even a trivial one is detected. In one simulated study with 20,000 per group and a true d of 0.05, p was about .000002 while the effect stayed negligible. Judge size with an effect size and its confidence interval, or test against a smallest effect that matters.
The short answer
The claim that tests become "biased toward rejection" in big samples mixes up two different things. When the null hypothesis is exactly true, a correctly applied test at α = .05 rejects it about 5% of the time whether you have 20 observations or 20 million. That is the Type I error rate, and sample size does not inflate it.
What sample size does change is power, the chance of detecting a difference that really exists. The standard error shrinks roughly in proportion to 1/√n, so a fixed difference produces a larger and larger test statistic as n grows. In practice two groups, or a coefficient and zero, are almost never exactly equal: tiny differences come from real but negligible effects, from small imbalances in who was sampled, or from slight measurement artifacts. With enough data the test correctly reports that the difference is not zero. The problem is reading "not zero" as "large" or "important".
See it in R and Python
The code does three things. It simulates 2,000 two-sample t tests with a true null at three sample sizes and records how often p < .05. It then computes the exact power to detect a tiny standardised difference of d = 0.05 at four sample sizes (no randomness involved). Finally it runs one large study, 20,000 people per group, with that tiny difference built in.
R
set.seed(108911)
# 1. When the null is exactly true, the false-positive rate stays at 5% for any n
for (n in c(20, 200, 2000)) {
p <- replicate(2000, t.test(rnorm(n), rnorm(n), var.equal = TRUE)$p.value)
cat("null true, n per group =", n, " rejection rate =", mean(p < 0.05), "\n")
}
# 2. A tiny real difference (d = 0.05): power grows with n (exact, no randomness)
for (n in c(50, 500, 5000, 20000)) {
pw <- power.t.test(n = n, delta = 0.05, sd = 1, sig.level = 0.05,
strict = TRUE)$power # count both rejection tails
cat("d = 0.05, n per group =", n, " power =", round(pw, 3), "\n")
}
# 3. One large study with that tiny difference
n <- 20000
x <- rnorm(n, mean = 0.05); y <- rnorm(n, mean = 0)
tt <- t.test(x, y, var.equal = TRUE)
d <- (mean(x) - mean(y)) / sqrt((var(x) + var(y)) / 2) # Cohen's d (equal n)
cat("t =", round(tt$statistic, 2), " df =", tt$parameter,
" p =", signif(tt$p.value, 3), "\n")
cat("mean difference =", round(mean(x) - mean(y), 3),
" 95% CI [", round(tt$conf.int[1], 3), ",", round(tt$conf.int[2], 3), "]\n")
cat("Cohen's d =", round(d, 3), "\n")Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(108911)
# 1. When the null is exactly true, the false-positive rate stays at 5% for any n
for n in [20, 200, 2000]:
p = stats.ttest_ind(rng.normal(size=(2000, n)), rng.normal(size=(2000, n)), axis=1).pvalue
print("null true, n per group =", n, " rejection rate =", np.mean(p < 0.05))
# 2. A tiny real difference (d = 0.05): power grows with n (exact, no randomness)
# Same as R's power.t.test(..., strict = TRUE): both rejection tails counted
for n in [50, 500, 5000, 20000]:
df = 2 * n - 2; ncp = 0.05 * np.sqrt(n / 2); crit = stats.t.ppf(0.975, df)
pw = stats.nct.sf(crit, df, ncp) + stats.nct.cdf(-crit, df, ncp)
print("d = 0.05, n per group =", n, " power =", round(pw, 3))
# 3. One large study with that tiny difference
n = 20000
x = rng.normal(0.05, 1, n); y = rng.normal(0, 1, n)
tt = stats.ttest_ind(x, y)
diff = x.mean() - y.mean()
sp = np.sqrt((x.var(ddof=1) + y.var(ddof=1)) / 2) # pooled SD (equal n)
se = sp * np.sqrt(2 / n); crit = stats.t.ppf(0.975, 2 * n - 2)
print("t =", round(tt.statistic, 2), " df =", 2 * n - 2, " p =", f"{tt.pvalue:.3g}")
print("mean difference =", round(diff, 3),
" 95% CI [", round(diff - crit * se, 3), ",", round(diff + crit * se, 3), "]")
print("Cohen's d =", round(diff / sp, 3))null true, n per group = 20 rejection rate = 0.044
null true, n per group = 200 rejection rate = 0.0515
null true, n per group = 2000 rejection rate = 0.042
d = 0.05, n per group = 50 power = 0.057
d = 0.05, n per group = 500 power = 0.124
d = 0.05, n per group = 5000 power = 0.705
d = 0.05, n per group = 20000 power = 0.999
t = 4.77 df = 39998 p = 1.89e-06
mean difference = 0.048 95% CI [ 0.028 , 0.067 ]
Cohen's d = 0.048
The simulated figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated figures differ slightly (rejection rates of .0485 to .0545 under the null, and d = 0.062 with p far below .001 in the large study), but they show the same pattern. The power values involve no randomness, and Python reproduces them exactly.
What the numbers show
- No bias under a true null. The rejection rates were .044, .0515 and .042 at 20, 200 and 2,000 per group. With 2,000 simulated tests, rates within about one percentage point of .05 are ordinary simulation noise, and there is no upward drift as n grows.
- Power to find a trivial effect climbs toward 1. For d = 0.05, power was .057 with 50 per group, .124 with 500, .705 with 5,000 and .999 with 20,000. The effect never changed; only the precision did.
- Significant is not the same as large. The big study gave t(39998) = 4.77, p = .0000019, yet Cohen's d was 0.048, and the 95% CI for the mean difference, 0.028 to 0.067 on a scale whose standard deviation is 1, rules out anything but a very small effect. The tiny p-value says the difference is reliably above zero, not that it matters.
How to guard against it
- Always report an effect size with its confidence interval. The interval shows both the direction and how big the effect could plausibly be, which is what a reader needs. See how to interpret Cohen's d.
- Decide in advance what size of effect would matter (a smallest effect size of interest) from theory, cost or clinical relevance, not from the data.
- Test against that threshold, not against zero. An equivalence test (two one-sided tests, TOST) or a minimum-effect test asks whether the effect is smaller, or larger, than the threshold you care about. With large samples this answers the useful question directly.
- Remember that systematic error does not shrink with n. Sampling bias, confounding and measurement artifacts stay the same size while the standard error falls, so in very large datasets they can produce significant results on their own. Design and measurement quality matter more as n grows.
A Bayesian analysis reaches a related conclusion from another direction: with a very large sample, a p-value just under .05 can go with a Bayes factor that favours the null (Lindley's paradox), because such a small test statistic is more consistent with a near-zero effect than with a meaningful one. Either way, the fix is to ask how big the effect is, not only whether it is zero. The reverse mistake, reading a non-significant result in a small study as proof of no effect, is covered in does a non-significant result support the null?. Our test chooser and the DASS blog cover choosing and reporting tests.
How to report a t test in APA style (7th edition)
Report the test, the exact p-value (or p < .001 when it is smaller), and an effect size with its interval, then interpret the size, not just the significance. Using the large study from the example above:
"The groups differed significantly, t(39998) = 4.77, p < .001, but the difference was negligible in size: the mean difference was 0.05, 95% CI [0.03, 0.07], d = 0.05."
If you set a smallest effect of interest in advance, add it to the method section (for example, "differences smaller than d = x.xx were considered too small to matter") and report the equivalence test against it in the same sentence as the t test.
Related tools and guides
- Check a reported p-value
- Power and sample size calculators
- Does a non-significant result support the null?
- How to interpret Cohen's d
- How are alpha and beta (Type I and Type II error rates) related?
- What does 95% confidence mean?
- Lindley's paradox (Wikipedia)
More answered questions
- Bonferroni correction: when should you use it, and what are the alternatives?
- How are alpha and beta (Type I and Type II error rates) related?
- Is a non-significant result from a large study evidence for the null?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.