Answer
t test vs ANOVA with two groups: are the assumptions really different?
The short answer
They are the same test, so they make the same assumptions. With two groups, the one-way ANOVA F equals the square of the pooled t statistic, and the p values are identical. Both assume independent observations, equal variances, and normal data within each group. "Normal within each group" and "normal residuals" are one assumption in two words, because a residual is just a score minus its own group mean.
The short answer
The apparent difference in assumptions comes from how the two tests are usually taught, not from the tests themselves. The t test is introduced as "compare two means", so its normality condition gets phrased as "the data are normal". ANOVA is introduced as a model with sums of squares, so its condition gets phrased as "the residuals are normal". Stated carefully, both say the same thing: within each group, the scores are normally distributed around that group's mean, with the same spread in both groups.
Once you see that, the equivalence is no surprise. The pooled (Student) two-sample t test, a one-way ANOVA with two groups, and a regression of the outcome on a 0/1 group indicator are three descriptions of one model. They share one set of assumptions and give one p value.
Why the two tests give the same answer
All three approaches fit the same model: each person's score equals their group's mean plus an error term, and the errors come from one normal distribution with mean 0 and a common variance. The pieces line up like this:
- The error variance. The t test's pooled variance and the ANOVA's mean square within groups (mean square error) are the same number: the sum of squared deviations from each group's own mean, divided by n₁ + n₂ − 2.
- The test statistic. With two groups the ANOVA's F equals t², exactly, for any data set.
- The reference distribution. If t follows a t distribution with ν degrees of freedom, then t² follows an F distribution with 1 and ν degrees of freedom. So the two-sided p value from the t test and the p value from the F test are identical.
- The regression view. If you code the groups 0 and 1, the intercept is the mean of group 0, the slope is the difference between the means, and the slope's t test is the same t test again.
One small asymmetry: squaring throws away the sign, so the F test can only ever be two-sided. If you have a genuine directional hypothesis, the t test lets you run it one-sided (see One-tailed vs two-tailed tests); the ANOVA cannot.
"Normal data" and "normal residuals" are the same assumption
A common reading of the t test's condition is that the whole data set should look normal. That is not what it requires. It requires each group's scores to be normal around that group's own mean. Subtract each group's mean from its scores and what is left are the residuals of the ANOVA or regression. Asking whether the residuals are normal is therefore the same question as asking whether each group is normal, with the groups lined up on a common center so you can look at them together.
The version that is genuinely wrong is checking the pooled raw scores, all groups mixed together. If the groups really differ, the pooled scores are a mixture of two distributions with different centers, and with a large enough difference they turn bimodal. The pooled data can then look strongly non-normal even though every assumption holds. In other words, how normal the pooled data look depends partly on the very effect you are testing.
Another point of confusion is the idea that the t test relies on the z (standard normal) distribution. It does not: it uses the t distribution, whose heavier tails account for estimating the standard deviation from the sample. The z test is the version for a known standard deviation, which almost never happens in practice. In both tests the normality condition matters most in small samples; with moderate or large groups the central limit theorem makes the difference in means close to normal anyway.
See it in R and Python
The code simulates two groups of 30 from normal populations with the same standard deviation, runs the pooled t test, the one-way ANOVA and the dummy regression, and compares the Welch versions too. It then checks normality two ways and repeats that check on 2,000 simulated data sets.
R
set.seed(1637)
# Two groups of 30, true means 50 and 62, same SD of 8
group <- factor(rep(c("control", "treatment"), each = 30))
score <- c(rnorm(30, 50, 8), rnorm(30, 62, 8))
round(tapply(score, group, mean), 2)
round(tapply(score, group, sd), 2)
# 1. Student t test, one-way ANOVA and regression on a dummy
tt <- t.test(score ~ group, var.equal = TRUE)
fit <- lm(score ~ group)
av <- anova(fit)
round(c(t = unname(tt$statistic), t_squared = unname(tt$statistic)^2,
F = av$`F value`[1], df_t = unname(tt$parameter), df_resid = av$Df[2]), 4)
signif(c(p_t = tt$p.value, p_F = av$`Pr(>F)`[1],
p_lm = summary(fit)$coefficients[2, 4]), 4)
round(coef(fit), 2) # slope = difference in means
round(tt$conf.int, 2) # CI for control minus treatment
# 2. Welch t test and Welch's ANOVA agree in the same way
wt <- t.test(score ~ group)
wa <- oneway.test(score ~ group)
round(c(welch_t_squared = unname(wt$statistic)^2, welch_F = unname(wa$statistic),
df_welch_t = unname(wt$parameter), df_welch_F = unname(wa$parameter[2])), 4)
# 3. "Normal data" means normal within each group, which is what residuals check
res <- residuals(fit) # each score minus its own group mean
all.equal(unname(res), score - ave(score, group))
round(c(pooled_raw = shapiro.test(score)$p.value, residuals = shapiro.test(res)$p.value), 3)
# Share of 2,000 simulated datasets where Shapiro-Wilk rejects at .05
reject_rate <- function(shift, reps = 2000) {
r <- replicate(reps, {
y <- c(rnorm(30, 50, 8), rnorm(30, 50 + shift, 8))
c(pooled_raw = shapiro.test(y)$p.value < 0.05,
residuals = shapiro.test(y - ave(y, group))$p.value < 0.05)
})
round(rowMeans(r), 3)
}
reject_rate(12) # the design above
reject_rate(35) # a large difference: pooled raw scores are bimodal
# Effect size: Cohen's d from the pooled SD, and eta squared from the ANOVA table
sp <- sqrt(sum(res^2) / (length(score) - 2))
round(c(d = unname(diff(tapply(score, group, mean))) / sp,
eta_sq = av$`Sum Sq`[1] / sum(av$`Sum Sq`)), 2)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(1637)
# Two groups of 30, true means 50 and 62, same SD of 8
ctl, trt = rng.normal(50, 8, 30), rng.normal(62, 8, 30)
score = np.concatenate([ctl, trt])
print("means", round(ctl.mean(), 2), round(trt.mean(), 2),
"sds", round(ctl.std(ddof=1), 2), round(trt.std(ddof=1), 2))
# 1. Student t test, one-way ANOVA and regression on a dummy
tt = stats.ttest_ind(ctl, trt, equal_var=True)
av = stats.f_oneway(ctl, trt)
X = np.column_stack([np.ones(60), np.repeat([0.0, 1.0], 30)])
coef = np.linalg.lstsq(X, score, rcond=None)[0]
res = score - X @ coef # each score minus its own group mean
se = np.sqrt(res @ res / 58 * np.linalg.inv(X.T @ X)[1, 1])
p_lm = 2 * stats.t.sf(abs(coef[1] / se), 58)
print("t", round(tt.statistic, 4), "t^2", round(tt.statistic**2, 4), "F", round(av.statistic, 4))
print("p", f"{tt.pvalue:.4g} {av.pvalue:.4g} {p_lm:.4g}", "coef", np.round(coef, 2))
# 2. Welch t test and Welch's ANOVA (by hand; for two groups the k-2 term drops out)
def welch_anova(*g):
n = np.array([len(x) for x in g]); m = np.array([x.mean() for x in g])
w = n / np.array([x.var(ddof=1) for x in g]); mw = (w * m).sum() / w.sum()
k = len(g); lam = 3 * (((1 - w / w.sum())**2) / (n - 1)).sum() / (k**2 - 1)
F = (w * (m - mw)**2).sum() / (k - 1) / (1 + 2 * lam * (k - 2) / 3)
return F, 1 / lam
wt = stats.ttest_ind(ctl, trt, equal_var=False)
print("welch t^2", round(wt.statistic**2, 4), "F and df", np.round(welch_anova(ctl, trt), 4))
# 3. "Normal data" means normal within each group, which is what residuals check
print("shapiro p", round(stats.shapiro(score).pvalue, 3), round(stats.shapiro(res).pvalue, 3))
# Share of 2,000 simulated datasets where Shapiro-Wilk rejects at .05
def reject_rate(shift, reps=2000):
raw = resid = 0
for _ in range(reps):
a, b = rng.normal(50, 8, 30), rng.normal(50 + shift, 8, 30)
raw += stats.shapiro(np.concatenate([a, b])).pvalue < 0.05
resid += stats.shapiro(np.concatenate([a - a.mean(), b - b.mean()])).pvalue < 0.05
return {"pooled_raw": round(raw / reps, 3), "residuals": round(resid / reps, 3)}
print(reject_rate(12)) # the design above
print(reject_rate(35)) # a large difference: pooled raw scores are bimodal
# Effect size: Cohen's d from the pooled SD, and eta squared
sp = np.sqrt(res @ res / 58)
eta_sq = 1 - (res @ res) / ((score - score.mean()) @ (score - score.mean()))
print("d", round((trt.mean() - ctl.mean()) / sp, 2), "eta_sq", round(eta_sq, 2))The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated sample and rates differ (its sample happens to show a smaller difference), but every identity holds exactly in both languages and the simulated rates show the same pattern. SciPy has no Welch ANOVA, so the Python tab computes it from the standard formula.
- The sample. Control: M = 48.73 (SD = 9.61); treatment: M = 61.34 (SD = 8.53), 30 per group.
- Same test three ways. The pooled t test gives t(58) = −5.38 (negative because R subtracts treatment from control), and t² = 28.9085. The ANOVA gives F(1, 58) = 28.9085. The t test, the ANOVA and the regression slope all give p = 1.414 × 10⁻⁶. The regression intercept is 48.73 (the control mean) and the slope is 12.61 (the difference in means).
- Welch versions match too. The squared Welch t and the F from Welch's ANOVA (
oneway.test()) are both 28.9085, with 57.19 degrees of freedom each. With equal group sizes the Welch and Student t statistics coincide; only the degrees of freedom differ. - Residuals are within-group deviations. R confirms the model residuals equal each score minus its group mean (
TRUE). - Checking the right thing. In the design above (a 12-point difference), Shapiro-Wilk rejected normality in 4.0% of pooled raw samples and 5.3% of residual samples, both near the nominal 5%. With a 35-point difference, it rejected the pooled raw scores in 98.8% of samples but the residuals in only 4.7%, even though every group was drawn from a normal distribution.
The last line is the practical lesson: judge normality from the residuals (or from each group separately), never from the pooled outcome. Better still, look at a plot rather than a test; see How to interpret a Q-Q plot and Is a normality test worth running on your data?.
So which should you use?
For two groups, choose whichever your readers expect; the result is the same. A t test is more natural when you want a confidence interval for the difference in means or a one-sided test. ANOVA or regression is more natural when the two-group comparison is one part of a bigger model, for example with covariates (ANCOVA) or additional factors.
If the variances may differ, the matching pair is Welch's t test and Welch's ANOVA, and they are equivalent in the same way. Welch is a sensible default for two groups; see Welch vs Student t test. For help choosing a test for your design, try the test chooser, and for more plain-language guides see the DASS blog.
How to report a t test and a one-way ANOVA in APA style (7th edition)
Report one test, not both: they are the same analysis, and reporting both suggests two pieces of evidence. With the example above, either of these works:
"Scores were higher in the treatment group (M = 61.34, SD = 8.53, n = 30) than in the control group (M = 48.73, SD = 9.61, n = 30), t(58) = 5.38, p < .001, d = 1.39. The mean difference was 12.61 points, 95% CI [7.92, 17.31]."
"A one-way ANOVA showed that scores differed between the treatment and control groups, F(1, 58) = 28.91, p < .001, η² = .33."
Report p values below .001 as p < .001 rather than giving the exact small value. Stating the direction in words lets you report t as a positive number. In the method section, name the test and say whether variances were pooled, for example "Groups were compared with an independent-samples t test assuming equal variances."
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- ANOVA power and sample size calculator
- Welch vs Student t test: should you just always use Welch?
- Is a normality test worth running on your data?
- How to interpret a Q-Q plot
- How to interpret regression output in R
- Which statistical test should I use?
More answered questions
- Does a t test need a minimum sample size?
- How do you interpret Cohen's d, and are 0.2, 0.5 and 0.8 a real standard?
- Welch vs Student t test: should you just always use Welch?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.