Answer
Welch vs Student t test: should you just always use Welch?
The short answer
Yes, for comparing two independent means use Welch by default. With equal variances it matches Student's test almost exactly, so you lose almost nothing. With unequal variances and unequal group sizes, Student's test can badly over- or under-reject while Welch stays near 5%. Do not pre-test variances to choose. Mann-Whitney answers a different question, so choose it for its hypothesis, not as a fallback.
The short answer
When you compare the means of two independent groups (two versions in an A/B test, a treatment and a control group), there are two classic t tests. Student's version pools the two sample variances into one estimate, which is only correct if the two populations have the same spread. Welch's version keeps the two variances separate and adjusts the degrees of freedom to match.
Using Welch as your default is sound advice, and it is what R already does: t.test(x, y) runs Welch unless you add var.equal = TRUE. The cost of choosing Welch when the variances really are equal is tiny. The cost of choosing Student when they are not can be large, and it is largest in exactly the situation where you cannot tell: a small group with a bigger spread than a large group.
What the two tests actually do differently
Both tests divide the difference in sample means by an estimate of its standard error. They differ in that estimate:
- Student: the standard error is sₚ × √(1/n₁ + 1/n₂), where sₚ is the pooled standard deviation, a weighted average in which the larger group counts more. Degrees of freedom are n₁ + n₂ − 2.
- Welch: the standard error is √(s₁²/n₁ + s₂²/n₂), so each group's variance is scaled by its own size. The degrees of freedom come from the Welch-Satterthwaite formula and are usually not a whole number; they lie between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2.
- **With equal group sizes the two t statistics are identical.** Only the degrees of freedom differ, and with similar variances Welch's df is close to n₁ + n₂ − 2, so the p values almost match.
The trouble starts with unequal group sizes. If the small group is the more variable one, pooling lets the large, tight group drag the variance estimate down. The standard error comes out too small, the t statistic too big, and Student's test rejects too often. If the large group is the more variable one, the error runs the other way and Student's test loses power. Welch avoids both problems because it never pools.
See it in R and Python
The code compares a small, variable group (n = 10) with a larger, tighter one (n = 18) using both tests, then simulates 10,000 pairs of samples from normal populations to count how often each test gives p < .05 under several designs.
R
set.seed(313471)
# 1. Two independent groups: a small, variable group A and a larger, tighter group B
a <- c(23.1, 27.4, 19.8, 31.2, 25.6, 28.9, 22.3, 34.0, 26.7, 29.5)
b <- c(20.4, 22.1, 18.9, 21.7, 23.5, 19.6, 20.8, 22.9, 21.2,
19.3, 24.0, 20.1, 21.9, 22.6, 18.4, 23.1, 20.7, 21.4)
round(c(mean_a = mean(a), sd_a = sd(a), mean_b = mean(b), sd_b = sd(b)), 2)
student <- t.test(a, b, var.equal = TRUE)
welch <- t.test(a, b) # R's default is Welch
show <- function(tt) round(c(t = unname(tt$statistic), df = unname(tt$parameter),
p = tt$p.value, lower = tt$conf.int[1], upper = tt$conf.int[2]), 4)
show(student)
show(welch)
signif(c(student_p = student$p.value, welch_p = welch$p.value), 3)
# Cohen's d with the pooled SD
sp <- sqrt(((length(a) - 1) * var(a) + (length(b) - 1) * var(b)) / (length(a) + length(b) - 2))
round((mean(a) - mean(b)) / sp, 2)
# 2. Simulation: share of p < .05 for each test (normal data, 10,000 samples)
rates <- function(n1, n2, sd1, sd2, shift = 0, reps = 10000) {
res <- replicate(reps, {
x <- rnorm(n1, shift, sd1); y <- rnorm(n2, 0, sd2)
c(student = t.test(x, y, var.equal = TRUE)$p.value < 0.05,
welch = t.test(x, y)$p.value < 0.05)
})
round(rowMeans(res), 3)
}
rates(20, 20, 1, 1) # equal n, equal SDs: both near 5%
rates(10, 30, 2, 1) # small group has the larger SD: Student too liberal
rates(10, 30, 1, 2) # large group has the larger SD: Student too conservative
rates(10, 30, 1, 1) # unequal n, equal SDs: both near 5%
rates(15, 15, 1, 1, shift = 1) # power when Student's assumption holds exactlyPython
import numpy as np
from scipy import stats
rng = np.random.default_rng(313471)
# 1. Two independent groups: a small, variable group A and a larger, tighter group B
a = np.array([23.1, 27.4, 19.8, 31.2, 25.6, 28.9, 22.3, 34.0, 26.7, 29.5])
b = np.array([20.4, 22.1, 18.9, 21.7, 23.5, 19.6, 20.8, 22.9, 21.2,
19.3, 24.0, 20.1, 21.9, 22.6, 18.4, 23.1, 20.7, 21.4])
print("mean_a", round(a.mean(), 2), "sd_a", round(a.std(ddof=1), 2),
"mean_b", round(b.mean(), 2), "sd_b", round(b.std(ddof=1), 2))
def show(equal_var):
res = stats.ttest_ind(a, b, equal_var=equal_var) # scipy's default is Student
ci = res.confidence_interval(0.95)
print("t", round(res.statistic, 4), "df", round(res.df, 4), "p", round(res.pvalue, 4),
"CI", round(ci.low, 4), round(ci.high, 4), "p (3 s.f.)", f"{res.pvalue:.3g}")
show(True) # Student
show(False) # Welch
# Cohen's d with the pooled SD
n1, n2 = len(a), len(b)
sp = np.sqrt(((n1 - 1) * a.var(ddof=1) + (n2 - 1) * b.var(ddof=1)) / (n1 + n2 - 2))
print("d", round((a.mean() - b.mean()) / sp, 2))
# 2. Simulation: share of p < .05 for each test (normal data, 10,000 samples)
def rates(n1, n2, sd1, sd2, shift=0.0, reps=10000):
x = rng.normal(shift, sd1, size=(reps, n1))
y = rng.normal(0.0, sd2, size=(reps, n2))
student = stats.ttest_ind(x, y, axis=1, equal_var=True).pvalue < 0.05
welch = stats.ttest_ind(x, y, axis=1, equal_var=False).pvalue < 0.05
return {"student": round(student.mean(), 3), "welch": round(welch.mean(), 3)}
print(rates(20, 20, 1, 1)) # equal n, equal SDs: both near 5%
print(rates(10, 30, 2, 1)) # small group has the larger SD: Student too liberal
print(rates(10, 30, 1, 2)) # large group has the larger SD: Student too conservative
print(rates(10, 30, 1, 1)) # unequal n, equal SDs: both near 5%
print(rates(15, 15, 1, 1, shift=1)) # power when Student's assumption holds exactlyThe figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 15.6% instead of 16.0% in the second design), but they show the same pattern. Note one trap when switching languages: R's t.test() defaults to Welch, while SciPy's ttest_ind() defaults to Student unless you pass equal_var=False.
- The example. Group A: M = 26.85 (SD = 4.32); group B: M = 21.26 (SD = 1.62). Student gives t(26) = 4.96, p < .001, 95% CI [3.28, 7.91]. Welch gives t(10.43) = 3.95, p = .003, 95% CI [2.45, 8.73]. Student's interval is too narrow, because pooling let the 18 tight values in group B shrink the estimated spread of group A.
- Equal sizes, equal SDs (20 vs 20). Student rejected in 4.9% of samples and Welch in 4.8%.
- **Small group more variable (10 with SD 2 vs 30 with SD 1).** Student rejected a true null in 16.0% of samples, over three times the nominal 5%. Welch stayed at 5.6%.
- **Large group more variable (10 with SD 1 vs 30 with SD 2).** Student rejected in only 1.1%, so it is overly cautious and loses power. Welch stayed at 5.0%.
- Unequal sizes, equal SDs (10 vs 30). Both rejected in 4.6% of samples.
- Power when variances are truly equal (15 vs 15, a shift of one SD). Student detected the effect in 75.3% of samples and Welch in 75.0%. That gap is the whole price of Welch in Student's best case.
Are there any drawbacks to always using Welch?
- A very small loss of power when variances are truly equal. It is negligible with moderate samples and noticeable only with a handful of cases per group, where every method is weak anyway.
- Fractional degrees of freedom. t(10.43) surprises some readers, but it is correct and should be reported as it is.
- It still compares means. Welch fixes unequal variances, not skewness or outliers. With large samples, as in most A/B tests, the central limit theorem usually protects the mean, but heavily skewed data in small groups still call for care, for example a bootstrap interval.
- More than two groups or covariates need a model. The same idea carries over: Welch's ANOVA for several groups, and heteroscedasticity-robust standard errors in regression.
One practice to drop is testing for equal variances first (Levene's or an F test) and then choosing Student or Welch based on the result. The pre-test has little power in small samples, exactly where unequal variances do the most damage, and the two-step procedure as a whole has worse error rates than simply using Welch every time. This is the same problem as pre-testing for normality (see Is a normality test worth running on your data?).
Where does Mann-Whitney fit?
The Mann-Whitney U test (also called the Wilcoxon rank-sum test) is not simply a version of the t test for non-normal data. It asks whether a randomly chosen value from one group tends to be larger than one from the other. That is a different question from "are the means different?", and the two can disagree. For revenue, the mean is usually what matters, because total revenue is the mean times the number of users, so a rank test can miss exactly the effect you care about, for example when a change affects a few big spenders.
Mann-Whitney is also not immune to unequal spreads: when the two groups differ in variance or shape, it can reject even though neither group tends to be larger. So the choice is about the question, not a ladder of assumptions. If you want to compare means, use Welch (with a bootstrap check if the data are very skewed and the samples small). If you want to know whether one group tends to produce larger values, use a rank-based test and report an effect size that matches it, such as the probability of superiority.
Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report a Welch t test in APA style (7th edition)
Name the test so readers know the variances were not pooled, give the degrees of freedom to two decimals, the t value, the exact p value, the confidence interval for the difference in means and an effect size. Using the example above:
"Scores were higher in group A (M = 26.85, SD = 4.32, n = 10) than in group B (M = 21.26, SD = 1.62, n = 18). Welch's t test, which does not assume equal variances, showed that the difference was statistically significant, t(10.43) = 3.95, p = .003, 95% CI [2.45, 8.73], d = 1.96."
Say which standardiser you used for d (here the pooled standard deviation) because, with unequal variances, other choices give different values. In the method section, one sentence is enough: "Group means were compared with Welch's t test for independent samples." You do not need to report a test of equal variances.
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Check a reported p-value
- Which statistical test should I use?
- What are degrees of freedom in statistics?
- Is a normality test worth running on your data?
- What do p values and t values actually mean?
- Welch's t-test (Wikipedia)
More answered questions
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
- What do p values and t values actually mean?
- What are degrees of freedom in statistics, really?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.