Answer

Why does the Welch t statistic only approximately follow a t distribution?

Inspired by a question on Cross Validated ·

t testhypothesis testing

The short answer

A t distribution needs a standard error built from one scaled chi-square. The unequal-variance standard error adds two chi-squares with different weights, so its exact distribution depends on the unknown variance ratio (the Behrens-Fisher problem). Satterthwaite's fix matches the mean and variance of that sum to a chi-square with fractional df. In simulation the approximation keeps the error rate within 0.7 percentage points of 5%.

The short answer

A t distribution is, by definition, what you get when you divide a standard normal variable by the square root of an independent chi-square variable divided by its degrees of freedom. Student's two-sample test fits that recipe exactly when the two populations have equal variances. Welch's test, which lets the variances differ, does not fit it, so its statistic is only approximately t. The fractional degrees of freedom that Welch's test reports are the approximation: they choose the t distribution that best mimics the true one.

This page explains where the approximation comes from and checks how well it works by simulation. If your question is the practical one (should you use Welch or Student's test?), see Welch vs Student t test: should you just always use Welch?, which compares the two tests' error rates and power.

Why a weighted sum of chi-squares breaks the exact result

Take two independent samples of sizes n₁ and n₂ from normal populations with the same mean. The difference in sample means is normal with variance σ₁²/n₁ + σ₂²/n₂. If you knew the σ's, dividing by the square root of that variance would give an exact standard normal z. The whole problem is the estimated denominator.

There is one curious exception. If σ₁²/(n₁(n₁ − 1)) happens to equal σ₂²/(n₂(n₂ − 1)), the two weights are equal, the sum is a chi-square again, and the Welch statistic is exactly t with n₁ + n₂ − 2 df. The simulation below checks this case too.

Where the fractional degrees of freedom come from

Welch and Satterthwaite proposed the same fix: pretend the estimated variance is a scaled chi-square, and pick the degrees of freedom ν so that its mean and variance match those of the real sum. That gives

ν = (σ₁²/n₁ + σ₂²/n₂)² / [ (σ₁²/n₁)²/(n₁ − 1) + (σ₂²/n₂)²/(n₂ − 1) ].

In practice the σ's are unknown, so software plugs in the sample variances. That adds a second layer of approximation: the df you see in the output is itself an estimate that changes from sample to sample. It always lies between the smaller of n₁ − 1 and n₂ − 1 and n₁ + n₂ − 2, and it is usually not a whole number. For more on what degrees of freedom represent, see What are degrees of freedom in statistics?

How good is the approximation? A simulation in R and Python

The code draws 100,000 pairs of normal samples with equal means, computes the Welch statistic and its estimated df for each pair, and compares the true distribution of the statistic with the t distributions that are supposed to approximate it.

R

set.seed(571580)

# Simulate the Welch statistic under a true null (equal means, normal data)
welch_sim <- function(n1, n2, sd1, sd2, reps = 100000) {
  x <- matrix(rnorm(reps * n1, 0, sd1), reps)
  y <- matrix(rnorm(reps * n2, 0, sd2), reps)
  v1 <- (rowSums(x^2) - n1 * rowMeans(x)^2) / (n1 - 1) / n1   # s1^2 / n1
  v2 <- (rowSums(y^2) - n2 * rowMeans(y)^2) / (n2 - 1) / n2   # s2^2 / n2
  t  <- (rowMeans(x) - rowMeans(y)) / sqrt(v1 + v2)
  df <- (v1 + v2)^2 / (v1^2 / (n1 - 1) + v2^2 / (n2 - 1))     # Welch's estimated df
  w1 <- sd1^2 / n1; w2 <- sd2^2 / n2
  nu <- (w1 + w2)^2 / (w1^2 / (n1 - 1) + w2^2 / (n2 - 1))     # Satterthwaite df, true SDs
  list(t = t, df = df, nu = nu)
}

# 1. Small, variable group (n = 5, SD = 3) vs larger, tighter group (n = 20, SD = 1)
s <- welch_sim(5, 20, 3, 1)
round(s$nu, 2)                                   # df from the true variances
round(quantile(s$df, c(0, .25, .5, .75, 1)), 2)  # df Welch actually estimates
round(c(true_cutoff = quantile(abs(s$t), .95, names = FALSE),
        t_nu_cutoff = qt(.975, s$nu)), 3)
round(c(reject_true_nu = mean(abs(s$t) > qt(.975, s$nu)),
        reject_welch   = mean(abs(s$t) > qt(.975, s$df))), 4)

# 2. The one special case where the statistic is exactly t:
#    sd1^2 / (n1 (n1 - 1)) equals sd2^2 / (n2 (n2 - 1)); here n1 = 5, n2 = 10
e <- welch_sim(5, 10, sqrt(20 / 90), 1)
round(e$nu, 2)
round(c(true_cutoff = quantile(abs(e$t), .95, names = FALSE),
        t13_cutoff = qt(.975, 13)), 3)

# 3. Welch's type I error rate as the SD of the small group changes (n = 5 vs 20)
sapply(c(0.25, 0.5, 1, 2, 4), function(sd1) {
  r <- welch_sim(5, 20, sd1, 1)
  round(mean(abs(r$t) > qt(.975, r$df)), 4)
})

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(571580)

# Simulate the Welch statistic under a true null (equal means, normal data)
def welch_sim(n1, n2, sd1, sd2, reps=100000):
    x = rng.normal(0, sd1, size=(reps, n1))
    y = rng.normal(0, sd2, size=(reps, n2))
    v1 = x.var(axis=1, ddof=1) / n1                               # s1^2 / n1
    v2 = y.var(axis=1, ddof=1) / n2                               # s2^2 / n2
    t = (x.mean(axis=1) - y.mean(axis=1)) / np.sqrt(v1 + v2)
    df = (v1 + v2) ** 2 / (v1 ** 2 / (n1 - 1) + v2 ** 2 / (n2 - 1))  # Welch's estimated df
    w1, w2 = sd1 ** 2 / n1, sd2 ** 2 / n2
    nu = (w1 + w2) ** 2 / (w1 ** 2 / (n1 - 1) + w2 ** 2 / (n2 - 1))  # Satterthwaite df, true SDs
    return t, df, nu

# 1. Small, variable group (n = 5, SD = 3) vs larger, tighter group (n = 20, SD = 1)
t, df, nu = welch_sim(5, 20, 3, 1)
print("nu", round(nu, 2))
print("df quantiles", np.round(np.quantile(df, [0, .25, .5, .75, 1]), 2))
print("true_cutoff", round(np.quantile(np.abs(t), .95), 3),
      "t_nu_cutoff", round(stats.t.ppf(.975, nu), 3))
print("reject_true_nu", round(np.mean(np.abs(t) > stats.t.ppf(.975, nu)), 4),
      "reject_welch", round(np.mean(np.abs(t) > stats.t.ppf(.975, df)), 4))

# 2. The one special case where the statistic is exactly t:
#    sd1^2 / (n1 (n1 - 1)) equals sd2^2 / (n2 (n2 - 1)); here n1 = 5, n2 = 10
t, df, nu = welch_sim(5, 10, np.sqrt(20 / 90), 1)
print("nu", round(nu, 2))
print("true_cutoff", round(np.quantile(np.abs(t), .95), 3),
      "t13_cutoff", round(stats.t.ppf(.975, 13), 3))

# 3. Welch's type I error rate as the SD of the small group changes (n = 5 vs 20)
rates = []
for sd1 in [0.25, 0.5, 1, 2, 4]:
    t, df, nu = welch_sim(5, 20, sd1, 1)
    rates.append(round(np.mean(np.abs(t) > stats.t.ppf(.975, df)), 4))
print(rates)
[1] 4.22
   0%   25%   50%   75%  100% 
 4.01  4.15  4.26  4.47 22.99 
true_cutoff t_nu_cutoff 
      2.690       2.719 
reject_true_nu   reject_welch 
        0.0482         0.0546 
[1] 13
true_cutoff  t13_cutoff 
      2.155       2.160 
[1] 0.0489 0.0506 0.0566 0.0552 0.0540

The output above is from one seeded run of the R code. Python's random numbers differ from R's, so its simulated figures differ slightly (for example a rejection rate of 5.32% instead of 5.46% in the first design), but they show the same pattern. Each simulated rate has a Monte Carlo margin of roughly ±0.15 percentage points.

What this means in practice

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report Welch-Satterthwaite degrees of freedom in APA style (7th edition)

Name the test, so readers know the variances were not pooled, and give the fractional degrees of freedom as your software computed them, usually to two decimals. With placeholders for your own results:

"Group A (M = x.xx, SD = x.xx, n = xx) differed from group B (M = x.xx, SD = x.xx, n = xx), Welch's t(x.xx) = x.xx, p = .xxx, 95% CI [LL, UL], d = x.xx."

In the method section, one sentence is enough: "Group means were compared with Welch's t test, which does not assume equal variances and uses Welch-Satterthwaite degrees of freedom." Do not round the df to a whole number, and do not report a preliminary test of equal variances.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.