Answer

Is a non-significant result from a large study evidence for the null?

Inspired by a question on Cross Validated ·

hypothesis testingp-valuespower analysisconfidence intervals

The short answer

It can be, but the p value alone cannot show it. A non-significant result supports "no meaningful effect" only if the study could have detected an effect you would care about. Define that smallest effect, then check that the confidence interval excludes it or run an equivalence test (TOST). No finite study can show the effect is exactly zero; it can show the effect is too small to matter.

The short answer

The textbook line "you can never accept the null" is half right. A significance test on its own is built to control one kind of mistake: rejecting a null hypothesis that is true. A large p value only says the data are not surprising if the effect is zero. They may be equally unsurprising if the effect is small but real, and the p value does not tell you which.

Your intuition is also right, though. A big study that finds nothing is more informative than a small one that finds nothing, because the big study would probably have caught an effect of any meaningful size. The way to turn that intuition into evidence is to say what "meaningful" means and test against it directly.

Why the p value cannot do the job

In Bayesian terms the picture is similar: a null result from a well-powered study usually shifts belief toward "no meaningful effect", while the same result from a weak study hardly shifts it at all. Bayes factors quantify that shift, but they also require you to say which effect sizes the alternative hypothesis expects.

See it in R and Python

The code simulates two studies of a test score with a standard deviation of 10 points. The true difference between groups is 0.5 points, which we treat as negligible; the smallest difference that would matter is set at 2 points (0.2 SD). One study has 20 people per group and the other 1,000. Each gets the usual Welch t test and an equivalence test (two one-sided tests, or TOST) against bounds of ±2 points. The code then gives each design's power to detect a 2-point difference and, over 2,000 repeats, how often each design concludes equivalence.

R

set.seed(1)
bound <- 2   # smallest difference that would matter: 2 points (0.2 SD)

# Two studies of the same tiny true difference (0.5 points, SD = 10)
analyse <- function(n) {
  a <- rnorm(n, 50.5, 10); b <- rnorm(n, 50, 10)
  test  <- t.test(a, b)                                   # usual two-sided test
  lower <- t.test(a, b, mu = -bound, alternative = "greater")$p.value
  upper <- t.test(a, b, mu =  bound, alternative = "less")$p.value
  ci90  <- t.test(a, b, conf.level = 0.90)$conf.int       # matches the TOST
  c(n = n, diff = unname(mean(a) - mean(b)), t = unname(test$statistic),
    df = unname(test$parameter), p = test$p.value,
    ci95_lo = test$conf.int[1], ci95_hi = test$conf.int[2],
    p_tost = max(lower, upper), ci90_lo = ci90[1], ci90_hi = ci90[2])
}
round(rbind(small = analyse(20), large = analyse(1000)), 3)

# Power of each design to detect a 2-point difference (d = 0.2), alpha = .05
round(sapply(c(n20 = 20, n1000 = 1000), function(n)
  power.t.test(n = n, delta = bound, sd = 10)$power), 3)

# How often each design declares equivalence when the true difference is 0.5
reps <- 2000
eq_rate <- function(n) mean(replicate(reps, analyse(n)["p_tost"] < 0.05))
round(c(n20 = eq_rate(20), n1000 = eq_rate(1000)), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(1)
bound = 2   # smallest difference that would matter: 2 points (0.2 SD)

def welch(a, b):
    va, vb = a.var(ddof=1, axis=-1) / a.shape[-1], b.var(ddof=1, axis=-1) / b.shape[-1]
    se = np.sqrt(va + vb)
    df = (va + vb) ** 2 / (va ** 2 / (a.shape[-1] - 1) + vb ** 2 / (b.shape[-1] - 1))
    return a.mean(axis=-1) - b.mean(axis=-1), se, df

# Two studies of the same tiny true difference (0.5 points, SD = 10)
def analyse(n, size=None):
    shape = (n,) if size is None else (size, n)
    a, b = rng.normal(50.5, 10, shape), rng.normal(50, 10, shape)
    diff, se, df = welch(a, b)
    t = diff / se
    p = 2 * stats.t.sf(abs(t), df)                          # usual two-sided test
    lower = stats.t.sf((diff + bound) / se, df)             # H0: diff <= -bound
    upper = stats.t.cdf((diff - bound) / se, df)            # H0: diff >= +bound
    p_tost = np.maximum(lower, upper)
    if size:                                                # simulation: TOST p values only
        return p_tost
    ci95 = diff + np.array([-1, 1]) * stats.t.ppf(0.975, df) * se
    ci90 = diff + np.array([-1, 1]) * stats.t.ppf(0.95, df) * se   # matches the TOST
    return dict(n=n, diff=diff, t=t, df=df, p=p, ci95=ci95, p_tost=p_tost, ci90=ci90)

for name, n in (("small", 20), ("large", 1000)):
    r = analyse(n)
    print(name, {k: np.round(v, 3).tolist() for k, v in r.items()})

# Power of each design to detect a 2-point difference (d = 0.2), alpha = .05
def power(n, d, alpha=0.05):
    df = 2 * n - 2
    return float(stats.nct.sf(stats.t.ppf(1 - alpha / 2, df), df, d * np.sqrt(n / 2)))

print("power", {f"n{n}": round(power(n, 0.2), 3) for n in (20, 1000)})

# How often each design declares equivalence when the true difference is 0.5
reps = 2000
print("equivalence", {f"n{n}": round(float(np.mean(analyse(n, reps) < 0.05)), 3)
                      for n in (20, 1000)})
         n  diff     t       df     p ci95_lo ci95_hi p_tost ci90_lo ci90_hi
small   20 2.470 0.875   37.917 0.387  -3.244   8.184  0.566  -2.289   7.229
large 1000 0.464 0.996 1997.905 0.319  -0.450   1.378  0.000  -0.303   1.231
  n20 n1000 
0.090 0.994 
  n20 n1000 
0.000 0.954

The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated studies and rates differ slightly (for example 95.9% instead of 95.4% for the equivalence rate of the large design), but they show the same pattern. The power figures involve no randomness, and Python reproduces them exactly.

How to argue for no meaningful effect

For a refresher on what the p value measures, see what p values and t values mean. Not sure which test fits your design? Try the test chooser, or browse the plain-language guides on the DASS blog.

How to report a non-significant result in APA style (7th edition)

Report the test in full with its confidence interval, so readers can see how precise the estimate was. Using the large study from the example:

"The difference between groups was not statistically significant, t(1997.91) = 1.00, p = .319, with a mean difference of 0.46 points, 95% CI [−0.45, 1.38]. An equivalence test against bounds of ±2 points (0.2 SD) was significant, p < .001, indicating that any difference was smaller than the smallest difference of interest."

For an inconclusive result such as the small study, stop at the first sentence and add the limitation: "The confidence interval, [−3.24, 8.18], was compatible with both no difference and differences large enough to matter." In the method section, state the smallest effect of interest and how it was chosen, ideally before data collection.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.