Answer

Welch vs Student t test: should you just always use Welch?

Inspired by a question on Cross Validated ·

t testhypothesis testing

The short answer

Yes, for comparing two independent means use Welch by default. With equal variances it matches Student's test almost exactly, so you lose almost nothing. With unequal variances and unequal group sizes, Student's test can badly over- or under-reject while Welch stays near 5%. Do not pre-test variances to choose. Mann-Whitney answers a different question, so choose it for its hypothesis, not as a fallback.

The short answer

When you compare the means of two independent groups (two versions in an A/B test, a treatment and a control group), there are two classic t tests. Student's version pools the two sample variances into one estimate, which is only correct if the two populations have the same spread. Welch's version keeps the two variances separate and adjusts the degrees of freedom to match.

Using Welch as your default is sound advice, and it is what R already does: t.test(x, y) runs Welch unless you add var.equal = TRUE. The cost of choosing Welch when the variances really are equal is tiny. The cost of choosing Student when they are not can be large, and it is largest in exactly the situation where you cannot tell: a small group with a bigger spread than a large group.

What the two tests actually do differently

Both tests divide the difference in sample means by an estimate of its standard error. They differ in that estimate:

The trouble starts with unequal group sizes. If the small group is the more variable one, pooling lets the large, tight group drag the variance estimate down. The standard error comes out too small, the t statistic too big, and Student's test rejects too often. If the large group is the more variable one, the error runs the other way and Student's test loses power. Welch avoids both problems because it never pools.

See it in R and Python

The code compares a small, variable group (n = 10) with a larger, tighter one (n = 18) using both tests, then simulates 10,000 pairs of samples from normal populations to count how often each test gives p < .05 under several designs.

R

set.seed(313471)

# 1. Two independent groups: a small, variable group A and a larger, tighter group B
a <- c(23.1, 27.4, 19.8, 31.2, 25.6, 28.9, 22.3, 34.0, 26.7, 29.5)
b <- c(20.4, 22.1, 18.9, 21.7, 23.5, 19.6, 20.8, 22.9, 21.2,
       19.3, 24.0, 20.1, 21.9, 22.6, 18.4, 23.1, 20.7, 21.4)
round(c(mean_a = mean(a), sd_a = sd(a), mean_b = mean(b), sd_b = sd(b)), 2)

student <- t.test(a, b, var.equal = TRUE)
welch   <- t.test(a, b)                      # R's default is Welch
show <- function(tt) round(c(t = unname(tt$statistic), df = unname(tt$parameter),
                             p = tt$p.value, lower = tt$conf.int[1], upper = tt$conf.int[2]), 4)
show(student)
show(welch)
signif(c(student_p = student$p.value, welch_p = welch$p.value), 3)

# Cohen's d with the pooled SD
sp <- sqrt(((length(a) - 1) * var(a) + (length(b) - 1) * var(b)) / (length(a) + length(b) - 2))
round((mean(a) - mean(b)) / sp, 2)

# 2. Simulation: share of p < .05 for each test (normal data, 10,000 samples)
rates <- function(n1, n2, sd1, sd2, shift = 0, reps = 10000) {
  res <- replicate(reps, {
    x <- rnorm(n1, shift, sd1); y <- rnorm(n2, 0, sd2)
    c(student = t.test(x, y, var.equal = TRUE)$p.value < 0.05,
      welch   = t.test(x, y)$p.value < 0.05)
  })
  round(rowMeans(res), 3)
}
rates(20, 20, 1, 1)          # equal n, equal SDs: both near 5%
rates(10, 30, 2, 1)          # small group has the larger SD: Student too liberal
rates(10, 30, 1, 2)          # large group has the larger SD: Student too conservative
rates(10, 30, 1, 1)          # unequal n, equal SDs: both near 5%
rates(15, 15, 1, 1, shift = 1)  # power when Student's assumption holds exactly

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(313471)

# 1. Two independent groups: a small, variable group A and a larger, tighter group B
a = np.array([23.1, 27.4, 19.8, 31.2, 25.6, 28.9, 22.3, 34.0, 26.7, 29.5])
b = np.array([20.4, 22.1, 18.9, 21.7, 23.5, 19.6, 20.8, 22.9, 21.2,
              19.3, 24.0, 20.1, 21.9, 22.6, 18.4, 23.1, 20.7, 21.4])
print("mean_a", round(a.mean(), 2), "sd_a", round(a.std(ddof=1), 2),
      "mean_b", round(b.mean(), 2), "sd_b", round(b.std(ddof=1), 2))

def show(equal_var):
    res = stats.ttest_ind(a, b, equal_var=equal_var)   # scipy's default is Student
    ci = res.confidence_interval(0.95)
    print("t", round(res.statistic, 4), "df", round(res.df, 4), "p", round(res.pvalue, 4),
          "CI", round(ci.low, 4), round(ci.high, 4), "p (3 s.f.)", f"{res.pvalue:.3g}")

show(True)    # Student
show(False)   # Welch

# Cohen's d with the pooled SD
n1, n2 = len(a), len(b)
sp = np.sqrt(((n1 - 1) * a.var(ddof=1) + (n2 - 1) * b.var(ddof=1)) / (n1 + n2 - 2))
print("d", round((a.mean() - b.mean()) / sp, 2))

# 2. Simulation: share of p < .05 for each test (normal data, 10,000 samples)
def rates(n1, n2, sd1, sd2, shift=0.0, reps=10000):
    x = rng.normal(shift, sd1, size=(reps, n1))
    y = rng.normal(0.0, sd2, size=(reps, n2))
    student = stats.ttest_ind(x, y, axis=1, equal_var=True).pvalue < 0.05
    welch = stats.ttest_ind(x, y, axis=1, equal_var=False).pvalue < 0.05
    return {"student": round(student.mean(), 3), "welch": round(welch.mean(), 3)}

print(rates(20, 20, 1, 1))            # equal n, equal SDs: both near 5%
print(rates(10, 30, 2, 1))            # small group has the larger SD: Student too liberal
print(rates(10, 30, 1, 2))            # large group has the larger SD: Student too conservative
print(rates(10, 30, 1, 1))            # unequal n, equal SDs: both near 5%
print(rates(15, 15, 1, 1, shift=1))   # power when Student's assumption holds exactly

The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 15.6% instead of 16.0% in the second design), but they show the same pattern. Note one trap when switching languages: R's t.test() defaults to Welch, while SciPy's ttest_ind() defaults to Student unless you pass equal_var=False.

Are there any drawbacks to always using Welch?

One practice to drop is testing for equal variances first (Levene's or an F test) and then choosing Student or Welch based on the result. The pre-test has little power in small samples, exactly where unequal variances do the most damage, and the two-step procedure as a whole has worse error rates than simply using Welch every time. This is the same problem as pre-testing for normality (see Is a normality test worth running on your data?).

Where does Mann-Whitney fit?

The Mann-Whitney U test (also called the Wilcoxon rank-sum test) is not simply a version of the t test for non-normal data. It asks whether a randomly chosen value from one group tends to be larger than one from the other. That is a different question from "are the means different?", and the two can disagree. For revenue, the mean is usually what matters, because total revenue is the mean times the number of users, so a rank test can miss exactly the effect you care about, for example when a change affects a few big spenders.

Mann-Whitney is also not immune to unequal spreads: when the two groups differ in variance or shape, it can reject even though neither group tends to be larger. So the choice is about the question, not a ladder of assumptions. If you want to compare means, use Welch (with a bootstrap check if the data are very skewed and the samples small). If you want to know whether one group tends to produce larger values, use a rank-based test and report an effect size that matches it, such as the probability of superiority.

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report a Welch t test in APA style (7th edition)

Name the test so readers know the variances were not pooled, give the degrees of freedom to two decimals, the t value, the exact p value, the confidence interval for the difference in means and an effect size. Using the example above:

"Scores were higher in group A (M = 26.85, SD = 4.32, n = 10) than in group B (M = 21.26, SD = 1.62, n = 18). Welch's t test, which does not assume equal variances, showed that the difference was statistically significant, t(10.43) = 3.95, p = .003, 95% CI [2.45, 8.73], d = 1.96."

Say which standardiser you used for d (here the pooled standard deviation) because, with unequal variances, other choices give different values. In the method section, one sentence is enough: "Group means were compared with Welch's t test for independent samples." You do not need to report a test of equal variances.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.