Answer
Why does the Welch t statistic only approximately follow a t distribution?
The short answer
A t distribution needs a standard error built from one scaled chi-square. The unequal-variance standard error adds two chi-squares with different weights, so its exact distribution depends on the unknown variance ratio (the Behrens-Fisher problem). Satterthwaite's fix matches the mean and variance of that sum to a chi-square with fractional df. In simulation the approximation keeps the error rate within 0.7 percentage points of 5%.
The short answer
A t distribution is, by definition, what you get when you divide a standard normal variable by the square root of an independent chi-square variable divided by its degrees of freedom. Student's two-sample test fits that recipe exactly when the two populations have equal variances. Welch's test, which lets the variances differ, does not fit it, so its statistic is only approximately t. The fractional degrees of freedom that Welch's test reports are the approximation: they choose the t distribution that best mimics the true one.
This page explains where the approximation comes from and checks how well it works by simulation. If your question is the practical one (should you use Welch or Student's test?), see Welch vs Student t test: should you just always use Welch?, which compares the two tests' error rates and power.
Why a weighted sum of chi-squares breaks the exact result
Take two independent samples of sizes n₁ and n₂ from normal populations with the same mean. The difference in sample means is normal with variance σ₁²/n₁ + σ₂²/n₂. If you knew the σ's, dividing by the square root of that variance would give an exact standard normal z. The whole problem is the estimated denominator.
- Each sample variance is a scaled chi-square. For normal data, (nᵢ − 1)sᵢ²/σᵢ² has a chi-square distribution with nᵢ − 1 degrees of freedom, and it is independent of the sample mean.
- Student's test assumes σ₁ = σ₂ and pools the two variances. The pooled variance is one common σ² times the sum of the two chi-squares, which is itself a chi-square with n₁ + n₂ − 2 degrees of freedom. The ratio is therefore exactly t with n₁ + n₂ − 2 df.
- Welch's test estimates the variance of the difference as s₁²/n₁ + s₂²/n₂. That is a weighted sum of two independent chi-squares, and the weights, σ₁²/(n₁(n₁ − 1)) and σ₂²/(n₂(n₂ − 1)), are usually different. A sum of chi-squares with unequal weights is not a scaled chi-square, so the statistic is not exactly t with any number of degrees of freedom.
- The true distribution depends on an unknown quantity. Its shape changes with the ratio σ₁²/σ₂², which you do not know. If one group's term dominates, the statistic behaves like t with that group's n − 1 df; if both terms matter, it sits somewhere in between. Finding a test for this setting that is exact whatever the variance ratio is the classic Behrens-Fisher problem.
There is one curious exception. If σ₁²/(n₁(n₁ − 1)) happens to equal σ₂²/(n₂(n₂ − 1)), the two weights are equal, the sum is a chi-square again, and the Welch statistic is exactly t with n₁ + n₂ − 2 df. The simulation below checks this case too.
Where the fractional degrees of freedom come from
Welch and Satterthwaite proposed the same fix: pretend the estimated variance is a scaled chi-square, and pick the degrees of freedom ν so that its mean and variance match those of the real sum. That gives
ν = (σ₁²/n₁ + σ₂²/n₂)² / [ (σ₁²/n₁)²/(n₁ − 1) + (σ₂²/n₂)²/(n₂ − 1) ].
In practice the σ's are unknown, so software plugs in the sample variances. That adds a second layer of approximation: the df you see in the output is itself an estimate that changes from sample to sample. It always lies between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2, and it is usually not a whole number. For more on what degrees of freedom represent, see What are degrees of freedom in statistics?
How good is the approximation? A simulation in R and Python
The code draws 100,000 pairs of normal samples with equal means, computes the Welch statistic and its estimated df for each pair, and compares the true distribution of the statistic with the t distributions that are supposed to approximate it.
R
set.seed(571580)
# Simulate the Welch statistic under a true null (equal means, normal data)
welch_sim <- function(n1, n2, sd1, sd2, reps = 100000) {
x <- matrix(rnorm(reps * n1, 0, sd1), reps)
y <- matrix(rnorm(reps * n2, 0, sd2), reps)
v1 <- (rowSums(x^2) - n1 * rowMeans(x)^2) / (n1 - 1) / n1 # s1^2 / n1
v2 <- (rowSums(y^2) - n2 * rowMeans(y)^2) / (n2 - 1) / n2 # s2^2 / n2
t <- (rowMeans(x) - rowMeans(y)) / sqrt(v1 + v2)
df <- (v1 + v2)^2 / (v1^2 / (n1 - 1) + v2^2 / (n2 - 1)) # Welch's estimated df
w1 <- sd1^2 / n1; w2 <- sd2^2 / n2
nu <- (w1 + w2)^2 / (w1^2 / (n1 - 1) + w2^2 / (n2 - 1)) # Satterthwaite df, true SDs
list(t = t, df = df, nu = nu)
}
# 1. Small, variable group (n = 5, SD = 3) vs larger, tighter group (n = 20, SD = 1)
s <- welch_sim(5, 20, 3, 1)
round(s$nu, 2) # df from the true variances
round(quantile(s$df, c(0, .25, .5, .75, 1)), 2) # df Welch actually estimates
round(c(true_cutoff = quantile(abs(s$t), .95, names = FALSE),
t_nu_cutoff = qt(.975, s$nu)), 3)
round(c(reject_true_nu = mean(abs(s$t) > qt(.975, s$nu)),
reject_welch = mean(abs(s$t) > qt(.975, s$df))), 4)
# 2. The one special case where the statistic is exactly t:
# sd1^2 / (n1 (n1 - 1)) equals sd2^2 / (n2 (n2 - 1)); here n1 = 5, n2 = 10
e <- welch_sim(5, 10, sqrt(20 / 90), 1)
round(e$nu, 2)
round(c(true_cutoff = quantile(abs(e$t), .95, names = FALSE),
t13_cutoff = qt(.975, 13)), 3)
# 3. Welch's type I error rate as the SD of the small group changes (n = 5 vs 20)
sapply(c(0.25, 0.5, 1, 2, 4), function(sd1) {
r <- welch_sim(5, 20, sd1, 1)
round(mean(abs(r$t) > qt(.975, r$df)), 4)
})Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(571580)
# Simulate the Welch statistic under a true null (equal means, normal data)
def welch_sim(n1, n2, sd1, sd2, reps=100000):
x = rng.normal(0, sd1, size=(reps, n1))
y = rng.normal(0, sd2, size=(reps, n2))
v1 = x.var(axis=1, ddof=1) / n1 # s1^2 / n1
v2 = y.var(axis=1, ddof=1) / n2 # s2^2 / n2
t = (x.mean(axis=1) - y.mean(axis=1)) / np.sqrt(v1 + v2)
df = (v1 + v2) ** 2 / (v1 ** 2 / (n1 - 1) + v2 ** 2 / (n2 - 1)) # Welch's estimated df
w1, w2 = sd1 ** 2 / n1, sd2 ** 2 / n2
nu = (w1 + w2) ** 2 / (w1 ** 2 / (n1 - 1) + w2 ** 2 / (n2 - 1)) # Satterthwaite df, true SDs
return t, df, nu
# 1. Small, variable group (n = 5, SD = 3) vs larger, tighter group (n = 20, SD = 1)
t, df, nu = welch_sim(5, 20, 3, 1)
print("nu", round(nu, 2))
print("df quantiles", np.round(np.quantile(df, [0, .25, .5, .75, 1]), 2))
print("true_cutoff", round(np.quantile(np.abs(t), .95), 3),
"t_nu_cutoff", round(stats.t.ppf(.975, nu), 3))
print("reject_true_nu", round(np.mean(np.abs(t) > stats.t.ppf(.975, nu)), 4),
"reject_welch", round(np.mean(np.abs(t) > stats.t.ppf(.975, df)), 4))
# 2. The one special case where the statistic is exactly t:
# sd1^2 / (n1 (n1 - 1)) equals sd2^2 / (n2 (n2 - 1)); here n1 = 5, n2 = 10
t, df, nu = welch_sim(5, 10, np.sqrt(20 / 90), 1)
print("nu", round(nu, 2))
print("true_cutoff", round(np.quantile(np.abs(t), .95), 3),
"t13_cutoff", round(stats.t.ppf(.975, 13), 3))
# 3. Welch's type I error rate as the SD of the small group changes (n = 5 vs 20)
rates = []
for sd1 in [0.25, 0.5, 1, 2, 4]:
t, df, nu = welch_sim(5, 20, sd1, 1)
rates.append(round(np.mean(np.abs(t) > stats.t.ppf(.975, df)), 4))
print(rates)[1] 4.22
0% 25% 50% 75% 100%
4.01 4.15 4.26 4.47 22.99
true_cutoff t_nu_cutoff
2.690 2.719
reject_true_nu reject_welch
0.0482 0.0546
[1] 13
true_cutoff t13_cutoff
2.155 2.160
[1] 0.0489 0.0506 0.0566 0.0552 0.0540
The output above is from one seeded run of the R code. Python's random numbers differ from R's, so its simulated figures differ slightly (for example a rejection rate of 5.32% instead of 5.46% in the first design), but they show the same pattern. Each simulated rate has a Monte Carlo margin of roughly ±0.15 percentage points.
- The df from the true variances is 4.22, close to the minimum of 4 (the small group's n − 1), because the small, variable group dominates the standard error.
- **The true distribution is not t(4.22).** The real 95% two-sided cutoff of the statistic was 2.690, while t(4.22) puts it at 2.719. Using the true ν would make the test slightly conservative: it rejected in 4.82% of samples.
- Estimating the df pushes the rate the other way. The estimated df ranged from 4.01 to 22.99 (median 4.26), and the usual Welch test rejected in 5.46% of samples.
- Why the estimate makes it liberal. The df estimate and the statistic are linked: when the small group's sample variance comes out low by chance, the standard error is too small (so |t| is larger) and the estimated df rises (so the cutoff is lower) at the same time.
- The exact special case checks out. With n₁ = 5, n₂ = 10 and σ₁² = (20/90)σ₂², the Satterthwaite df is exactly 13, and the simulated cutoff (2.155) matches t(13) (2.160) within simulation error.
- Across variance ratios the test stays close to 5%. With n₁ = 5 and n₂ = 20, and the small group's SD at 0.25, 0.5, 1, 2 and 4 times the other group's, Welch's test rejected in 4.89%, 5.06%, 5.66%, 5.52% and 5.40% of samples. Its errors are a fraction of a percentage point even with only five cases in one group, while Student's test can be off by a factor of three in similar designs.
What this means in practice
- "Approximate" does not mean "unreliable". For normal data the Welch test's actual error rate is close to the nominal one across the whole range of variance ratios, which is why it is a sensible default for two independent means.
- The approximation is weakest with very small groups. With a handful of cases in one group, expect the actual rate to drift by around half a percentage point, as in the simulation. With moderate samples the drift becomes negligible.
- Non-normal data are a separate issue. The argument above relies on each sample variance being a scaled chi-square, which is only true for normal data. Strong skewness in small samples affects both Welch's and Student's tests, so consider a bootstrap interval in that case.
- Report the df as computed. A non-integer df is not an error; it is the Satterthwaite approximation at work.
Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report Welch-Satterthwaite degrees of freedom in APA style (7th edition)
Name the test, so readers know the variances were not pooled, and give the fractional degrees of freedom as your software computed them, usually to two decimals. With placeholders for your own results:
"Group A (M = x.xx, SD = x.xx, n = xx) differed from group B (M = x.xx, SD = x.xx, n = xx), Welch's t(x.xx) = x.xx, p = .xxx, 95% CI [LL, UL], d = x.xx."
In the method section, one sentence is enough: "Group means were compared with Welch's t test, which does not assume equal variances and uses Welch-Satterthwaite degrees of freedom." Do not round the df to a whole number, and do not report a preliminary test of equal variances.
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Check a reported p-value
- Welch vs Student t test: should you just always use Welch?
- What are degrees of freedom in statistics?
- Which statistical test should I use?
- Welch-Satterthwaite equation (Wikipedia)
More answered questions
- Paired vs independent t test: which should you use?
- How are alpha and beta (Type I and Type II error rates) related?
- Welch vs Student t test: should you just always use Welch?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.