Answer

What should you do after rejecting the null hypothesis?

Inspired by a question on Cross Validated ·

hypothesis testingp-valueseffect sizeconfidence intervals

The short answer

Rejecting the null only tells you the data would be unusual if there were no effect. The next steps are to estimate how big the effect is (an effect size with a confidence interval), check that the result survives reasonable changes to the analysis and is not one of many tests you ran, and then replicate it. No single test proves the alternative; evidence builds up across estimates, checks and repeated studies.

The short answer

A significance test answers a narrow question: if the true effect were zero, how surprising would data like yours be? A small p value says "quite surprising", so you reject the null. It does not say how large the effect is, whether it matters in practice, or how likely it is to show up again. Those are the questions to work on next.

In practice the follow-up has three parts: estimate the effect, stress-test the result, and replicate it. You never reach proof in the mathematical sense. What you can reach is a result that is precise, robust and repeatable, which is what "conclusive" means in empirical research.

Step 1: estimate how big the effect is

Step 2: check that the result is robust

See it in R and Python

The code simulates a two-group study of a test score with a standard deviation of 10 points and a true difference of 3 points (d = 0.3), with 40 people per group. It runs a Welch t test, computes Cohen's d with an approximate 95% interval (the common large-sample formula), and gives the design's power. It then repeats the same study 5,000 times and compares the average d across all studies with the average among the significant ones.

R

set.seed(7)
# One study: test scores (SD = 10), true difference 3 points (d = 0.3)
n <- 40
treat <- rnorm(n, 53, 10); ctrl <- rnorm(n, 50, 10)
test <- t.test(treat, ctrl)                     # Welch two-sided test

# Effect size: Cohen's d with an approximate 95% CI
d_est <- function(a, b) {
  sp <- sqrt((var(a) + var(b)) / 2)             # pooled SD (equal n)
  (mean(a) - mean(b)) / sp
}
d  <- d_est(treat, ctrl)
se <- sqrt(2 / n + d^2 / (4 * n))
round(c(diff = mean(treat) - mean(ctrl), t = unname(test$statistic),
        df = unname(test$parameter), p = test$p.value,
        ci_lo = test$conf.int[1], ci_hi = test$conf.int[2],
        d = d, d_lo = d - 1.96 * se, d_hi = d + 1.96 * se), 3)

# Power of this design for the true effect
round(power.t.test(n = n, delta = 3, sd = 10)$power, 3)

# Winner's curse: repeat the study 5,000 times; keep the significant ones
reps <- 5000
sims <- t(replicate(reps, {
  a <- rnorm(n, 53, 10); b <- rnorm(n, 50, 10)
  c(p = t.test(a, b)$p.value, d = d_est(a, b))
}))
sig <- sims[, "p"] < 0.05
round(c(share_significant = mean(sig),
        mean_d_all = mean(sims[, "d"]),
        mean_d_significant = mean(sims[sig, "d"])), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(7)
# One study: test scores (SD = 10), true difference 3 points (d = 0.3)
n = 40
treat, ctrl = rng.normal(53, 10, n), rng.normal(50, 10, n)

def welch(a, b):
    va, vb = a.var(ddof=1, axis=-1) / n, b.var(ddof=1, axis=-1) / n
    se = np.sqrt(va + vb)
    df = (va + vb) ** 2 / (va ** 2 / (n - 1) + vb ** 2 / (n - 1))
    diff = a.mean(axis=-1) - b.mean(axis=-1)
    return diff, se, df, 2 * stats.t.sf(np.abs(diff / se), df)

# Effect size: Cohen's d with an approximate 95% CI
def d_est(a, b):
    sp = np.sqrt((a.var(ddof=1, axis=-1) + b.var(ddof=1, axis=-1)) / 2)  # pooled SD (equal n)
    return (a.mean(axis=-1) - b.mean(axis=-1)) / sp

diff, se_diff, df, p = welch(treat, ctrl)
ci = diff + np.array([-1, 1]) * stats.t.ppf(0.975, df) * se_diff
d = d_est(treat, ctrl)
se = np.sqrt(2 / n + d ** 2 / (4 * n))
print({k: round(float(v), 3) for k, v in dict(
    diff=diff, t=diff / se_diff, df=df, p=p, ci_lo=ci[0], ci_hi=ci[1],
    d=d, d_lo=d - 1.96 * se, d_hi=d + 1.96 * se).items()})

# Power of this design for the true effect
dfp = 2 * n - 2
print("power", round(float(stats.nct.sf(stats.t.ppf(0.975, dfp), dfp, 0.3 * np.sqrt(n / 2))), 3))

# Winner's curse: repeat the study 5,000 times; keep the significant ones
reps = 5000
a, b = rng.normal(53, 10, (reps, n)), rng.normal(50, 10, (reps, n))
p_sim = welch(a, b)[3]
d_sim = d_est(a, b)
sig = p_sim < 0.05
print({"share_significant": round(float(sig.mean()), 3),
       "mean_d_all": round(float(d_sim.mean()), 3),
       "mean_d_significant": round(float(d_sim[sig].mean()), 3)})
  diff      t     df      p  ci_lo  ci_hi      d   d_lo   d_hi 
 4.483  2.028 76.898  0.046  0.081  8.885  0.453  0.010  0.897 
[1] 0.263
 share_significant         mean_d_all mean_d_significant 
             0.250              0.297              0.579

The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its single simulated study comes out differently (in fact non-significant, which is what this design produces about three times in four), and its simulation gives slightly different shares (26.5% significant, average d of 0.588 among those). The pattern is the same, and the power figure, which involves no randomness, matches exactly.

Step 3: replicate and accumulate evidence

For the mirror-image question of what a non-significant result can tell you, see is a non-significant result evidence for the null?. Not sure which test fits your design? Try the test chooser, or browse the plain-language guides on the DASS blog.

How to report a significant result in APA style (7th edition)

Report the test, the estimate with its confidence interval, and an effect size, so readers can judge both significance and size. Using the single study from the example:

"Scores were higher in the treatment group than in the control group, t(76.90) = 2.03, p = .046, with a mean difference of 4.48 points, 95% CI [0.08, 8.89], d = 0.45, 95% CI [0.01, 0.90]."

Then say what the interval implies: "The confidence interval was wide and included differences too small to be of practical importance, so the size of the effect remains uncertain." In the method section, state whether the hypothesis and analysis were planned in advance, and report every test you ran, not only the significant ones.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.