Answer
Is a non-significant result from a large study evidence for the null?
The short answer
It can be, but the p value alone cannot show it. A non-significant result supports "no meaningful effect" only if the study could have detected an effect you would care about. Define that smallest effect, then check that the confidence interval excludes it or run an equivalence test (TOST). No finite study can show the effect is exactly zero; it can show the effect is too small to matter.
The short answer
The textbook line "you can never accept the null" is half right. A significance test on its own is built to control one kind of mistake: rejecting a null hypothesis that is true. A large p value only says the data are not surprising if the effect is zero. They may be equally unsurprising if the effect is small but real, and the p value does not tell you which.
Your intuition is also right, though. A big study that finds nothing is more informative than a small one that finds nothing, because the big study would probably have caught an effect of any meaningful size. The way to turn that intuition into evidence is to say what "meaningful" means and test against it directly.
Why the p value cannot do the job
- The exact null is out of reach. An effect of 0 and an effect of 0.0001 produce practically the same data. No sample, however large, can separate them, so no study can confirm that an effect is exactly zero.
- The p value ignores the sample size's promise. A p of .39 from 20 people per group and a p of .32 from 1,000 per group look alike, yet the first study could barely detect anything while the second could detect small effects. The number that tells them apart is the width of the confidence interval (or the power), not the p value.
- Observed power does not fix it. Power computed from the effect you happened to observe is a direct function of the p value, so it adds no information. Power only helps when it is computed for an effect size you chose in advance.
In Bayesian terms the picture is similar: a null result from a well-powered study usually shifts belief toward "no meaningful effect", while the same result from a weak study hardly shifts it at all. Bayes factors quantify that shift, but they also require you to say which effect sizes the alternative hypothesis expects.
See it in R and Python
The code simulates two studies of a test score with a standard deviation of 10 points. The true difference between groups is 0.5 points, which we treat as negligible; the smallest difference that would matter is set at 2 points (0.2 SD). One study has 20 people per group and the other 1,000. Each gets the usual Welch t test and an equivalence test (two one-sided tests, or TOST) against bounds of ±2 points. The code then gives each design's power to detect a 2-point difference and, over 2,000 repeats, how often each design concludes equivalence.
R
set.seed(1)
bound <- 2 # smallest difference that would matter: 2 points (0.2 SD)
# Two studies of the same tiny true difference (0.5 points, SD = 10)
analyse <- function(n) {
a <- rnorm(n, 50.5, 10); b <- rnorm(n, 50, 10)
test <- t.test(a, b) # usual two-sided test
lower <- t.test(a, b, mu = -bound, alternative = "greater")$p.value
upper <- t.test(a, b, mu = bound, alternative = "less")$p.value
ci90 <- t.test(a, b, conf.level = 0.90)$conf.int # matches the TOST
c(n = n, diff = unname(mean(a) - mean(b)), t = unname(test$statistic),
df = unname(test$parameter), p = test$p.value,
ci95_lo = test$conf.int[1], ci95_hi = test$conf.int[2],
p_tost = max(lower, upper), ci90_lo = ci90[1], ci90_hi = ci90[2])
}
round(rbind(small = analyse(20), large = analyse(1000)), 3)
# Power of each design to detect a 2-point difference (d = 0.2), alpha = .05
round(sapply(c(n20 = 20, n1000 = 1000), function(n)
power.t.test(n = n, delta = bound, sd = 10)$power), 3)
# How often each design declares equivalence when the true difference is 0.5
reps <- 2000
eq_rate <- function(n) mean(replicate(reps, analyse(n)["p_tost"] < 0.05))
round(c(n20 = eq_rate(20), n1000 = eq_rate(1000)), 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(1)
bound = 2 # smallest difference that would matter: 2 points (0.2 SD)
def welch(a, b):
va, vb = a.var(ddof=1, axis=-1) / a.shape[-1], b.var(ddof=1, axis=-1) / b.shape[-1]
se = np.sqrt(va + vb)
df = (va + vb) ** 2 / (va ** 2 / (a.shape[-1] - 1) + vb ** 2 / (b.shape[-1] - 1))
return a.mean(axis=-1) - b.mean(axis=-1), se, df
# Two studies of the same tiny true difference (0.5 points, SD = 10)
def analyse(n, size=None):
shape = (n,) if size is None else (size, n)
a, b = rng.normal(50.5, 10, shape), rng.normal(50, 10, shape)
diff, se, df = welch(a, b)
t = diff / se
p = 2 * stats.t.sf(abs(t), df) # usual two-sided test
lower = stats.t.sf((diff + bound) / se, df) # H0: diff <= -bound
upper = stats.t.cdf((diff - bound) / se, df) # H0: diff >= +bound
p_tost = np.maximum(lower, upper)
if size: # simulation: TOST p values only
return p_tost
ci95 = diff + np.array([-1, 1]) * stats.t.ppf(0.975, df) * se
ci90 = diff + np.array([-1, 1]) * stats.t.ppf(0.95, df) * se # matches the TOST
return dict(n=n, diff=diff, t=t, df=df, p=p, ci95=ci95, p_tost=p_tost, ci90=ci90)
for name, n in (("small", 20), ("large", 1000)):
r = analyse(n)
print(name, {k: np.round(v, 3).tolist() for k, v in r.items()})
# Power of each design to detect a 2-point difference (d = 0.2), alpha = .05
def power(n, d, alpha=0.05):
df = 2 * n - 2
return float(stats.nct.sf(stats.t.ppf(1 - alpha / 2, df), df, d * np.sqrt(n / 2)))
print("power", {f"n{n}": round(power(n, 0.2), 3) for n in (20, 1000)})
# How often each design declares equivalence when the true difference is 0.5
reps = 2000
print("equivalence", {f"n{n}": round(float(np.mean(analyse(n, reps) < 0.05)), 3)
for n in (20, 1000)}) n diff t df p ci95_lo ci95_hi p_tost ci90_lo ci90_hi
small 20 2.470 0.875 37.917 0.387 -3.244 8.184 0.566 -2.289 7.229
large 1000 0.464 0.996 1997.905 0.319 -0.450 1.378 0.000 -0.303 1.231
n20 n1000
0.090 0.994
n20 n1000
0.000 0.954
The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated studies and rates differ slightly (for example 95.9% instead of 95.4% for the equivalence rate of the large design), but they show the same pattern. The power figures involve no randomness, and Python reproduces them exactly.
- Both studies are non-significant, with p = .387 for the small study and p = .319 for the large one. On p values alone they look the same.
- Their confidence intervals are not the same. The small study's 95% interval runs from −3.24 to 8.18 points, so it cannot rule out a difference of 8 points in either the helpful or harmful direction. The large study's interval, −0.45 to 1.38, sits well inside ±2.
- Only the large study supports "no meaningful difference". Its equivalence test gave p < .001 (its 90% interval, −0.30 to 1.23, lies inside ±2), while the small study's gave p = .566.
- That is what power buys. The small design had 9.0% power to detect a 2-point difference and declared equivalence in none of the 2,000 repeats; the large design had 99.4% power and declared equivalence 95.4% of the time.
How to argue for no meaningful effect
- Pick the smallest effect of interest before looking at the data. Base it on what would matter in practice (a clinically or educationally important change), on earlier studies, or on what your study was designed to detect.
- Report the confidence interval, not just "n.s." If the whole interval lies inside the range of trivial effects, the data point to no meaningful effect. If it reaches beyond that range, the result is inconclusive. See what a 95% confidence interval means.
- Run an equivalence test when the claim matters. In TOST you test against the upper and the lower bound separately; if both one-sided tests are significant at α = .05, the effect is inside the bounds. That is the same as the 90% confidence interval falling inside them. See equivalence test (Wikipedia).
- Plan the sample size for it. A study that should be able to show equivalence needs enough people for the interval to be narrower than the bounds, which usually means more people than a standard test needs (see does a t test need a minimum sample size?).
- Say "no evidence of a difference" when the result is inconclusive, and "no meaningful difference" only after an equivalence test or an interval that supports it.
For a refresher on what the p value measures, see what p values and t values mean. Not sure which test fits your design? Try the test chooser, or browse the plain-language guides on the DASS blog.
How to report a non-significant result in APA style (7th edition)
Report the test in full with its confidence interval, so readers can see how precise the estimate was. Using the large study from the example:
"The difference between groups was not statistically significant, t(1997.91) = 1.00, p = .319, with a mean difference of 0.46 points, 95% CI [−0.45, 1.38]. An equivalence test against bounds of ±2 points (0.2 SD) was significant, p < .001, indicating that any difference was smaller than the smallest difference of interest."
For an inconclusive result such as the small study, stop at the first sentence and add the limitation: "The confidence interval, [−3.24, 8.18], was compatible with both no difference and differences large enough to matter." In the method section, state the smallest effect of interest and how it was chosen, ideally before data collection.
Related tools and guides
- Check a reported p-value
- Power and sample size calculators
- What do p values and t values mean?
- What does a 95% confidence interval mean?
- Does a t test need a minimum sample size?
- Which statistical test should I use?
- Equivalence test (Wikipedia)
More answered questions
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
- What do p values and t values actually mean?
- Is there a 95% probability that your confidence interval covers the true mean?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.