Answer
What should you do after rejecting the null hypothesis?
The short answer
Rejecting the null only tells you the data would be unusual if there were no effect. The next steps are to estimate how big the effect is (an effect size with a confidence interval), check that the result survives reasonable changes to the analysis and is not one of many tests you ran, and then replicate it. No single test proves the alternative; evidence builds up across estimates, checks and repeated studies.
The short answer
A significance test answers a narrow question: if the true effect were zero, how surprising would data like yours be? A small p value says "quite surprising", so you reject the null. It does not say how large the effect is, whether it matters in practice, or how likely it is to show up again. Those are the questions to work on next.
In practice the follow-up has three parts: estimate the effect, stress-test the result, and replicate it. You never reach proof in the mathematical sense. What you can reach is a result that is precise, robust and repeatable, which is what "conclusive" means in empirical research.
Step 1: estimate how big the effect is
- Report the estimate and its confidence interval. For two means, that is the mean difference in the original units with a 95% interval, which shows the range of effects the data fit. An interval running from almost zero to very large tells you the direction and little else. See what a 95% confidence interval means.
- Add a standardized effect size such as Cohen's d when readers need to compare across scales or studies, again with an interval. See how to interpret Cohen's d.
- Ask whether the effect matters. Compare the interval with the smallest difference that would be important in your field. With a large sample, a trivial effect can be highly significant (see why large samples make tiny effects significant), so statistical and practical significance have to be judged separately.
Step 2: check that the result is robust
- Check the assumptions and the data. Look at residuals, outliers and data-entry errors. Does the conclusion hold with a robust alternative, such as a rank-based test or a bootstrap interval?
- Count your tests honestly. If this was one of many outcomes, subgroups or model versions you tried, a p value just under .05 is much weaker evidence than it looks. Correct for multiple comparisons, or label the analysis as exploratory (see when to use a Bonferroni correction).
- Rule out other explanations. In an observational study, confounding, selection effects or measurement problems can produce a real difference that has nothing to do with the cause you care about. Significance says the pattern is unlikely to be noise; it says nothing about why the pattern is there.
- Look at the power of the design. A significant result from a small, low-powered study is fragile. The simulation below shows why.
See it in R and Python
The code simulates a two-group study of a test score with a standard deviation of 10 points and a true difference of 3 points (d = 0.3), with 40 people per group. It runs a Welch t test, computes Cohen's d with an approximate 95% interval (the common large-sample formula), and gives the design's power. It then repeats the same study 5,000 times and compares the average d across all studies with the average among the significant ones.
R
set.seed(7)
# One study: test scores (SD = 10), true difference 3 points (d = 0.3)
n <- 40
treat <- rnorm(n, 53, 10); ctrl <- rnorm(n, 50, 10)
test <- t.test(treat, ctrl) # Welch two-sided test
# Effect size: Cohen's d with an approximate 95% CI
d_est <- function(a, b) {
sp <- sqrt((var(a) + var(b)) / 2) # pooled SD (equal n)
(mean(a) - mean(b)) / sp
}
d <- d_est(treat, ctrl)
se <- sqrt(2 / n + d^2 / (4 * n))
round(c(diff = mean(treat) - mean(ctrl), t = unname(test$statistic),
df = unname(test$parameter), p = test$p.value,
ci_lo = test$conf.int[1], ci_hi = test$conf.int[2],
d = d, d_lo = d - 1.96 * se, d_hi = d + 1.96 * se), 3)
# Power of this design for the true effect
round(power.t.test(n = n, delta = 3, sd = 10)$power, 3)
# Winner's curse: repeat the study 5,000 times; keep the significant ones
reps <- 5000
sims <- t(replicate(reps, {
a <- rnorm(n, 53, 10); b <- rnorm(n, 50, 10)
c(p = t.test(a, b)$p.value, d = d_est(a, b))
}))
sig <- sims[, "p"] < 0.05
round(c(share_significant = mean(sig),
mean_d_all = mean(sims[, "d"]),
mean_d_significant = mean(sims[sig, "d"])), 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(7)
# One study: test scores (SD = 10), true difference 3 points (d = 0.3)
n = 40
treat, ctrl = rng.normal(53, 10, n), rng.normal(50, 10, n)
def welch(a, b):
va, vb = a.var(ddof=1, axis=-1) / n, b.var(ddof=1, axis=-1) / n
se = np.sqrt(va + vb)
df = (va + vb) ** 2 / (va ** 2 / (n - 1) + vb ** 2 / (n - 1))
diff = a.mean(axis=-1) - b.mean(axis=-1)
return diff, se, df, 2 * stats.t.sf(np.abs(diff / se), df)
# Effect size: Cohen's d with an approximate 95% CI
def d_est(a, b):
sp = np.sqrt((a.var(ddof=1, axis=-1) + b.var(ddof=1, axis=-1)) / 2) # pooled SD (equal n)
return (a.mean(axis=-1) - b.mean(axis=-1)) / sp
diff, se_diff, df, p = welch(treat, ctrl)
ci = diff + np.array([-1, 1]) * stats.t.ppf(0.975, df) * se_diff
d = d_est(treat, ctrl)
se = np.sqrt(2 / n + d ** 2 / (4 * n))
print({k: round(float(v), 3) for k, v in dict(
diff=diff, t=diff / se_diff, df=df, p=p, ci_lo=ci[0], ci_hi=ci[1],
d=d, d_lo=d - 1.96 * se, d_hi=d + 1.96 * se).items()})
# Power of this design for the true effect
dfp = 2 * n - 2
print("power", round(float(stats.nct.sf(stats.t.ppf(0.975, dfp), dfp, 0.3 * np.sqrt(n / 2))), 3))
# Winner's curse: repeat the study 5,000 times; keep the significant ones
reps = 5000
a, b = rng.normal(53, 10, (reps, n)), rng.normal(50, 10, (reps, n))
p_sim = welch(a, b)[3]
d_sim = d_est(a, b)
sig = p_sim < 0.05
print({"share_significant": round(float(sig.mean()), 3),
"mean_d_all": round(float(d_sim.mean()), 3),
"mean_d_significant": round(float(d_sim[sig].mean()), 3)}) diff t df p ci_lo ci_hi d d_lo d_hi
4.483 2.028 76.898 0.046 0.081 8.885 0.453 0.010 0.897
[1] 0.263
share_significant mean_d_all mean_d_significant
0.250 0.297 0.579
The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its single simulated study comes out differently (in fact non-significant, which is what this design produces about three times in four), and its simulation gives slightly different shares (26.5% significant, average d of 0.588 among those). The pattern is the same, and the power figure, which involves no randomness, matches exactly.
- The single study rejects the null, t(76.90) = 2.03, p = .046. The estimated difference is 4.48 points, but the 95% interval runs from 0.08 to 8.89 points: anything from a negligible effect to a large one fits the data.
- The effect size is overestimated. The study's d is 0.45, against a true value of 0.3, and its interval (0.01 to 0.90) is very wide.
- The design was weak. Its power to detect the true effect was 26.3%, so an exact repeat of the study would reach significance only about one time in four.
- Significant results are inflated on average. Across 5,000 repeats, 25.0% were significant. The average d across all studies was 0.297, close to the truth, but among the significant ones it was 0.579, almost double. Filtering on significance selects the studies that overestimated the effect by chance. This is often called the winner's curse.
Step 3: replicate and accumulate evidence
- Run a replication planned in advance. Fix the hypothesis, the outcome and the analysis before collecting new data. Because a first significant estimate tends to be too large, size the replication for a smaller, more realistic effect, not the one you happened to observe.
- Combine the evidence. Pool the original and replication studies in a meta-analysis to get a single, more precise estimate. Precision, not a second p value below .05, is what makes a finding convincing.
- Consider a Bayesian summary if you want a direct statement of how probable the effect sizes are, given the data and a stated prior. It answers a different question from the p value and makes the role of prior knowledge explicit.
- Test what the theory predicts next. If the effect is real, what else should be true? Checking a dose-response pattern, a different population or a different measure strengthens the claim far more than repeating the original test.
For the mirror-image question of what a non-significant result can tell you, see is a non-significant result evidence for the null?. Not sure which test fits your design? Try the test chooser, or browse the plain-language guides on the DASS blog.
How to report a significant result in APA style (7th edition)
Report the test, the estimate with its confidence interval, and an effect size, so readers can judge both significance and size. Using the single study from the example:
"Scores were higher in the treatment group than in the control group, t(76.90) = 2.03, p = .046, with a mean difference of 4.48 points, 95% CI [0.08, 8.89], d = 0.45, 95% CI [0.01, 0.90]."
Then say what the interval implies: "The confidence interval was wide and included differences too small to be of practical importance, so the size of the effect remains uncertain." In the method section, state whether the hypothesis and analysis were planned in advance, and report every test you ran, not only the significant ones.
Related tools and guides
- Check a reported p-value
- Is a non-significant result evidence for the null?
- How to interpret Cohen's d
- Why do large samples make tiny effects significant?
- Which statistical test should I use?
- Replication crisis (Wikipedia)
More answered questions
- Why does almost everything become statistically significant with a large sample?
- Is a non-significant result from a large study evidence for the null?
- How to interpret the F statistic and p value in a one-way ANOVA
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.