Answer

t test vs ANOVA with two groups: are the assumptions really different?

Inspired by a question on Cross Validated ·

t testanovanormality

The short answer

They are the same test, so they make the same assumptions. With two groups, the one-way ANOVA F equals the square of the pooled t statistic, and the p values are identical. Both assume independent observations, equal variances, and normal data within each group. "Normal within each group" and "normal residuals" are one assumption in two words, because a residual is just a score minus its own group mean.

The short answer

The apparent difference in assumptions comes from how the two tests are usually taught, not from the tests themselves. The t test is introduced as "compare two means", so its normality condition gets phrased as "the data are normal". ANOVA is introduced as a model with sums of squares, so its condition gets phrased as "the residuals are normal". Stated carefully, both say the same thing: within each group, the scores are normally distributed around that group's mean, with the same spread in both groups.

Once you see that, the equivalence is no surprise. The pooled (Student) two-sample t test, a one-way ANOVA with two groups, and a regression of the outcome on a 0/1 group indicator are three descriptions of one model. They share one set of assumptions and give one p value.

Why the two tests give the same answer

All three approaches fit the same model: each person's score equals their group's mean plus an error term, and the errors come from one normal distribution with mean 0 and a common variance. The pieces line up like this:

One small asymmetry: squaring throws away the sign, so the F test can only ever be two-sided. If you have a genuine directional hypothesis, the t test lets you run it one-sided (see One-tailed vs two-tailed tests); the ANOVA cannot.

"Normal data" and "normal residuals" are the same assumption

A common reading of the t test's condition is that the whole data set should look normal. That is not what it requires. It requires each group's scores to be normal around that group's own mean. Subtract each group's mean from its scores and what is left are the residuals of the ANOVA or regression. Asking whether the residuals are normal is therefore the same question as asking whether each group is normal, with the groups lined up on a common center so you can look at them together.

The version that is genuinely wrong is checking the pooled raw scores, all groups mixed together. If the groups really differ, the pooled scores are a mixture of two distributions with different centers, and with a large enough difference they turn bimodal. The pooled data can then look strongly non-normal even though every assumption holds. In other words, how normal the pooled data look depends partly on the very effect you are testing.

Another point of confusion is the idea that the t test relies on the z (standard normal) distribution. It does not: it uses the t distribution, whose heavier tails account for estimating the standard deviation from the sample. The z test is the version for a known standard deviation, which almost never happens in practice. In both tests the normality condition matters most in small samples; with moderate or large groups the central limit theorem makes the difference in means close to normal anyway.

See it in R and Python

The code simulates two groups of 30 from normal populations with the same standard deviation, runs the pooled t test, the one-way ANOVA and the dummy regression, and compares the Welch versions too. It then checks normality two ways and repeats that check on 2,000 simulated data sets.

R

set.seed(1637)

# Two groups of 30, true means 50 and 62, same SD of 8
group <- factor(rep(c("control", "treatment"), each = 30))
score <- c(rnorm(30, 50, 8), rnorm(30, 62, 8))
round(tapply(score, group, mean), 2)
round(tapply(score, group, sd), 2)

# 1. Student t test, one-way ANOVA and regression on a dummy
tt  <- t.test(score ~ group, var.equal = TRUE)
fit <- lm(score ~ group)
av  <- anova(fit)
round(c(t = unname(tt$statistic), t_squared = unname(tt$statistic)^2,
        F = av$`F value`[1], df_t = unname(tt$parameter), df_resid = av$Df[2]), 4)
signif(c(p_t = tt$p.value, p_F = av$`Pr(>F)`[1],
         p_lm = summary(fit)$coefficients[2, 4]), 4)
round(coef(fit), 2)          # slope = difference in means
round(tt$conf.int, 2)        # CI for control minus treatment

# 2. Welch t test and Welch's ANOVA agree in the same way
wt <- t.test(score ~ group)
wa <- oneway.test(score ~ group)
round(c(welch_t_squared = unname(wt$statistic)^2, welch_F = unname(wa$statistic),
        df_welch_t = unname(wt$parameter), df_welch_F = unname(wa$parameter[2])), 4)

# 3. "Normal data" means normal within each group, which is what residuals check
res <- residuals(fit)        # each score minus its own group mean
all.equal(unname(res), score - ave(score, group))
round(c(pooled_raw = shapiro.test(score)$p.value, residuals = shapiro.test(res)$p.value), 3)

# Share of 2,000 simulated datasets where Shapiro-Wilk rejects at .05
reject_rate <- function(shift, reps = 2000) {
  r <- replicate(reps, {
    y <- c(rnorm(30, 50, 8), rnorm(30, 50 + shift, 8))
    c(pooled_raw = shapiro.test(y)$p.value < 0.05,
      residuals  = shapiro.test(y - ave(y, group))$p.value < 0.05)
  })
  round(rowMeans(r), 3)
}
reject_rate(12)   # the design above
reject_rate(35)   # a large difference: pooled raw scores are bimodal

# Effect size: Cohen's d from the pooled SD, and eta squared from the ANOVA table
sp <- sqrt(sum(res^2) / (length(score) - 2))
round(c(d = unname(diff(tapply(score, group, mean))) / sp,
        eta_sq = av$`Sum Sq`[1] / sum(av$`Sum Sq`)), 2)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(1637)

# Two groups of 30, true means 50 and 62, same SD of 8
ctl, trt = rng.normal(50, 8, 30), rng.normal(62, 8, 30)
score = np.concatenate([ctl, trt])
print("means", round(ctl.mean(), 2), round(trt.mean(), 2),
      "sds", round(ctl.std(ddof=1), 2), round(trt.std(ddof=1), 2))

# 1. Student t test, one-way ANOVA and regression on a dummy
tt = stats.ttest_ind(ctl, trt, equal_var=True)
av = stats.f_oneway(ctl, trt)
X = np.column_stack([np.ones(60), np.repeat([0.0, 1.0], 30)])
coef = np.linalg.lstsq(X, score, rcond=None)[0]
res = score - X @ coef                       # each score minus its own group mean
se = np.sqrt(res @ res / 58 * np.linalg.inv(X.T @ X)[1, 1])
p_lm = 2 * stats.t.sf(abs(coef[1] / se), 58)
print("t", round(tt.statistic, 4), "t^2", round(tt.statistic**2, 4), "F", round(av.statistic, 4))
print("p", f"{tt.pvalue:.4g} {av.pvalue:.4g} {p_lm:.4g}", "coef", np.round(coef, 2))

# 2. Welch t test and Welch's ANOVA (by hand; for two groups the k-2 term drops out)
def welch_anova(*g):
    n = np.array([len(x) for x in g]); m = np.array([x.mean() for x in g])
    w = n / np.array([x.var(ddof=1) for x in g]); mw = (w * m).sum() / w.sum()
    k = len(g); lam = 3 * (((1 - w / w.sum())**2) / (n - 1)).sum() / (k**2 - 1)
    F = (w * (m - mw)**2).sum() / (k - 1) / (1 + 2 * lam * (k - 2) / 3)
    return F, 1 / lam

wt = stats.ttest_ind(ctl, trt, equal_var=False)
print("welch t^2", round(wt.statistic**2, 4), "F and df", np.round(welch_anova(ctl, trt), 4))

# 3. "Normal data" means normal within each group, which is what residuals check
print("shapiro p", round(stats.shapiro(score).pvalue, 3), round(stats.shapiro(res).pvalue, 3))

# Share of 2,000 simulated datasets where Shapiro-Wilk rejects at .05
def reject_rate(shift, reps=2000):
    raw = resid = 0
    for _ in range(reps):
        a, b = rng.normal(50, 8, 30), rng.normal(50 + shift, 8, 30)
        raw += stats.shapiro(np.concatenate([a, b])).pvalue < 0.05
        resid += stats.shapiro(np.concatenate([a - a.mean(), b - b.mean()])).pvalue < 0.05
    return {"pooled_raw": round(raw / reps, 3), "residuals": round(resid / reps, 3)}

print(reject_rate(12))   # the design above
print(reject_rate(35))   # a large difference: pooled raw scores are bimodal

# Effect size: Cohen's d from the pooled SD, and eta squared
sp = np.sqrt(res @ res / 58)
eta_sq = 1 - (res @ res) / ((score - score.mean()) @ (score - score.mean()))
print("d", round((trt.mean() - ctl.mean()) / sp, 2), "eta_sq", round(eta_sq, 2))

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated sample and rates differ (its sample happens to show a smaller difference), but every identity holds exactly in both languages and the simulated rates show the same pattern. SciPy has no Welch ANOVA, so the Python tab computes it from the standard formula.

The last line is the practical lesson: judge normality from the residuals (or from each group separately), never from the pooled outcome. Better still, look at a plot rather than a test; see How to interpret a Q-Q plot and Is a normality test worth running on your data?.

So which should you use?

For two groups, choose whichever your readers expect; the result is the same. A t test is more natural when you want a confidence interval for the difference in means or a one-sided test. ANOVA or regression is more natural when the two-group comparison is one part of a bigger model, for example with covariates (ANCOVA) or additional factors.

If the variances may differ, the matching pair is Welch's t test and Welch's ANOVA, and they are equivalent in the same way. Welch is a sensible default for two groups; see Welch vs Student t test. For help choosing a test for your design, try the test chooser, and for more plain-language guides see the DASS blog.

How to report a t test and a one-way ANOVA in APA style (7th edition)

Report one test, not both: they are the same analysis, and reporting both suggests two pieces of evidence. With the example above, either of these works:

"Scores were higher in the treatment group (M = 61.34, SD = 8.53, n = 30) than in the control group (M = 48.73, SD = 9.61, n = 30), t(58) = 5.38, p < .001, d = 1.39. The mean difference was 12.61 points, 95% CI [7.92, 17.31]."

"A one-way ANOVA showed that scores differed between the treatment and control groups, F(1, 58) = 28.91, p < .001, η² = .33."

Report p values below .001 as p < .001 rather than giving the exact small value. Stating the direction in words lets you report t as a positive number. In the method section, name the test and say whether variances were pooled, for example "Groups were compared with an independent-samples t test assuming equal variances."

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.