Answer
Planned contrasts vs post hoc tests: why do they give different p-values?
The short answer
A planned contrast is a question you committed to before seeing the data, so it is tested on its own, at (or close to) the usual α. A post hoc test such as Tukey's HSD protects you while you search every pair for differences, so each comparison must clear a higher bar. Same difference, same standard error, different reference distribution: that is why the p-values disagree. With four groups of 15, a planned comparison needs |t| above 2.003, Tukey needs 2.648, and Scheffé needs 2.882.
The short answer
After a one-way ANOVA you can compare two group means in several ways, and they all start from the same ingredients: the difference between the two means and a standard error built from the pooled within-group variance (the ANOVA's mean square error). What changes is how big the test statistic has to be before you call the difference real. That threshold depends on how many comparisons the procedure is protecting, and whether you chose them before or after looking at the data.
- Planned contrast (a priori): you named this comparison before collecting or inspecting the data. It is one question, so its p-value comes straight from the t distribution with the ANOVA's error degrees of freedom. This gives the smallest p-value of the three.
- Tukey's HSD (post hoc): it treats your comparison as one of all k(k − 1)/2 possible pairs and keeps the chance of any false positive across all of them at α. Its p-value comes from the studentized range distribution, so it is larger.
- Bonferroni over the pairs: it multiplies each unadjusted p-value by the number of comparisons (capped at 1). With three groups and three pairs, the Bonferroni p is exactly three times the planned-contrast p. For all pairwise comparisons it is usually more conservative than Tukey, because it ignores that the comparisons share groups and are correlated.
So there is no flaw in getting three different answers. They are answers to three different questions: "is this one prespecified difference real?", "which of all the pairs differ?", and "which of these m tests survive a worst-case correction?"
See it in R and Python
The code does three things for a design with four groups of 15 people (56 error degrees of freedom). Part 1 computes the |t| each approach requires at α = .05. Part 2 simulates one data set (a control group with a true mean of 50 and three treatments with true means of 54, all with SD = 10) and tests the planned contrast "average treatment minus control" alongside Tukey's HSD. Part 3 simulates 5,000 studies where all four groups are equal, picks the biggest observed gap in each, and tests it either with an ordinary t test or against Tukey's threshold.
R
set.seed(342638)
k <- 4; n <- 15; df_err <- k * n - k # 4 groups, 15 per group, 56 error df
# 1. The |t| each approach needs before a difference counts (alpha = .05)
crit <- c(planned = qt(0.975, df_err),
tukey = qtukey(0.95, k, df_err) / sqrt(2),
scheffe = sqrt((k - 1) * qf(0.95, k - 1, df_err)))
round(crit, 3)
# 2. One data set: a control group and three treatments
grp <- factor(rep(c("control", "A", "B", "C"), each = n),
levels = c("control", "A", "B", "C"))
y <- rnorm(k * n, mean = rep(c(50, 54, 54, 54), each = n), sd = 10)
fit <- aov(y ~ grp)
mse <- sum(residuals(fit)^2) / df_err
means <- tapply(y, grp, mean)
round(means, 2)
w <- c(-1, 1/3, 1/3, 1/3) # planned: average treatment minus control
est <- sum(w * means); se <- sqrt(mse * sum(w^2) / n)
t_c <- est / se
round(c(estimate = est, SE = se, t = t_c,
lower = est - crit[["planned"]] * se,
upper = est + crit[["planned"]] * se), 3)
signif(2 * pt(-abs(t_c), df_err), 3) # p-value of the planned contrast
round(TukeyHSD(fit)$grp, 3) # all six pairwise differences, Tukey-adjusted
# 3. All four groups equal: test only the biggest observed gap
reps <- 5000
hits <- replicate(reps, {
x <- matrix(rnorm(n * k), n)
m <- colMeans(x); s2 <- mean(apply(x, 2, var))
t_max <- (max(m) - min(m)) / sqrt(2 * s2 / n)
c(unadjusted_t = t_max > crit[["planned"]], tukey = t_max > crit[["tukey"]])
})
rowMeans(hits) # false-positive rate of each rulePython
import numpy as np
from itertools import combinations
from scipy import stats
rng = np.random.default_rng(342638)
k, n = 4, 15
df_err = k * n - k # 4 groups, 15 per group, 56 error df
# 1. The |t| each approach needs before a difference counts (alpha = .05)
q95 = stats.studentized_range.ppf(0.95, k, df_err)
crit = {"planned": stats.t.ppf(0.975, df_err),
"tukey": q95 / np.sqrt(2),
"scheffe": np.sqrt((k - 1) * stats.f.ppf(0.95, k - 1, df_err))}
print({key: round(float(v), 3) for key, v in crit.items()})
# 2. One data set: a control group and three treatments
names = ["control", "A", "B", "C"]
y = rng.normal(np.repeat([50, 54, 54, 54], n)[:, None], 10, size=(k * n, 1)).ravel()
groups = y.reshape(k, n)
means = groups.mean(axis=1)
mse = groups.var(axis=1, ddof=1).mean() # pooled error variance (equal n)
print(dict(zip(names, np.round(means, 2))))
w = np.array([-1, 1/3, 1/3, 1/3]) # planned: average treatment minus control
est = w @ means; se = np.sqrt(mse * np.sum(w**2) / n)
t_c = est / se
print("estimate", round(est, 3), "SE", round(se, 3), "t", round(t_c, 3),
"CI", np.round([est - crit["planned"] * se, est + crit["planned"] * se], 3),
"p", float(f"{2 * stats.t.sf(abs(t_c), df_err):.3g}"))
for i, j in combinations(range(k), 2): # Tukey HSD, all six pairs
diff = means[j] - means[i]; half = q95 * np.sqrt(mse / n)
p = stats.studentized_range.sf(abs(diff) / np.sqrt(mse / n), k, df_err)
print(f"{names[j]}-{names[i]}", np.round([diff, diff - half, diff + half, p], 3))
# 3. All four groups equal: test only the biggest observed gap
reps = 5000
x = rng.normal(size=(reps, k, n))
m = x.mean(axis=2); s2 = x.var(axis=2, ddof=1).mean(axis=1)
t_max = (m.max(axis=1) - m.min(axis=1)) / np.sqrt(2 * s2 / n)
print("unadjusted_t", np.mean(t_max > crit["planned"]),
"tukey", np.mean(t_max > crit["tukey"]))planned tukey scheffe
2.003 2.648 2.882
control A B C
44.68 54.47 55.20 56.40
estimate SE t lower upper
10.676 2.989 3.572 4.689 16.663
[1] 0.000736
diff lwr upr p adj
A-control 9.790 0.098 19.483 0.047
B-control 10.520 0.828 20.213 0.028
C-control 11.718 2.025 21.410 0.012
B-A 0.730 -8.962 10.423 0.997
C-A 1.927 -7.765 11.620 0.952
C-B 1.197 -8.495 10.890 0.988
unadjusted_t tukey
0.2056 0.0534
All figures come from one seeded run of the R code. The critical values in part 1 involve no randomness, and Python reproduces them exactly. Python's random numbers differ from R's, so its single data set in part 2 is a different sample: there the control group happens to land above two of the treatments, and neither the contrast nor any Tukey pair is significant, a reminder that one small study can miss a real 4-point effect. Its simulated rates in part 3 differ slightly (19.8% and 4.9%) but show the same pattern.
What the numbers show
- The bar rises as the family grows. One planned comparison needs |t| > 2.003. Tukey, which covers all six pairs, needs 2.648. Scheffé, which covers every contrast you could ever invent after seeing the data, needs 2.882.
- A focused planned contrast is also more precise. Averaging the three treatments gives a standard error about 18% smaller than for a single pair (the weights give √(4/3) instead of √2). Here the contrast estimated a 10.68-point advantage for treatment, t(56) = 3.57, p < .001.
- One sample, not the truth. The control group's sample mean (44.68) fell below its true value of 50 by chance, which is why the observed gap is larger than the true 4 points.
- Tukey reaches the same conclusion, but only just. Each treatment differed from control (adjusted p = .047, .028 and .012), and the lower confidence limit for treatment A was only 0.10 points above zero. A slightly smaller gap would have left some pairs non-significant while the planned contrast stayed clear.
- Choosing the comparison after looking breaks unadjusted tests. With all groups truly equal, testing the biggest observed gap with an ordinary t test gave a false positive in 20.6% of studies, four times the nominal 5%. Tukey's threshold brought it back to 5.3%, within simulation noise of 5%.
Which one should you use?
Use planned contrasts when your hypotheses name specific comparisons in advance, for example "each treatment beats control" or "the two active conditions differ from the placebo condition on average". Write them down before the analysis (a preregistration is ideal). Keep the set small: many textbooks allow up to k − 1 contrasts, preferably orthogonal ones, at the usual α without adjustment, while stricter reviewers expect a Bonferroni or Holm correction across the planned set. Even then the correction is mild, because the set is small. You do not need a significant omnibus F first; the contrast is the test you care about.
Use a post hoc procedure when you had no specific prediction and want to know which groups differ. Tukey's HSD is the standard choice for all pairwise comparisons with roughly equal group sizes (the Tukey-Kramer version handles unequal sizes). Use Dunnett's test if the only comparisons of interest are each group against one control, and Scheffé if you will also test complex contrasts suggested by the data. For the general logic of familywise corrections, and why Holm beats plain Bonferroni, see our page on when to use a Bonferroni correction.
Do not relabel a post hoc test as planned. A comparison only counts as planned if you would have run it whatever the data showed. Choosing the most striking pair after inspecting the means and reporting it with an unadjusted p-value is exactly the situation in part 3, with a false-positive rate around 20% for four groups. If a comparison was planned, say so and report it as such; if it was not, adjust for the full set you were in effect searching through. For the omnibus test itself, see how to interpret the F statistic in ANOVA, and for picking an analysis in the first place, our test chooser and the DASS blog can help.
How to report a planned contrast in APA style (7th edition)
State in the method section that the contrast was specified in advance and give its weights. In the results, report the estimate, its confidence interval and the t test with the ANOVA's error degrees of freedom. Using the example above:
"A planned contrast compared the mean of the three treatment groups with the control group (weights −1, 1/3, 1/3, 1/3). Treatment scores were 10.68 points higher on average, 95% CI [4.69, 16.66], t(56) = 3.57, p < .001."
If the pairwise comparisons were post hoc, name the procedure and report the adjusted values: "Tukey HSD comparisons showed that each treatment group scored higher than the control group (differences of 9.79, 10.52 and 11.72 points for A, B and C; adjusted p = .047, .028 and .012), while the treatment groups did not differ from one another (all ps > .95)." With many pairs, put the differences, intervals and adjusted p-values in a table and name the adjustment in a table note.
Related tools and guides
- ANOVA power and sample size calculator
- APA 7 formatter for ANOVA
- Check a reported p-value
- Bonferroni correction: when should you use it?
- How to interpret the F statistic in ANOVA
- t test vs ANOVA with two groups
- Which statistical test should I use?
- Tukey's range test (Wikipedia)
More answered questions
- How to interpret the F statistic and p value in a one-way ANOVA
- Bonferroni correction: when should you use it, and what are the alternatives?
- Why does almost everything become statistically significant with a large sample?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.