Answer

t test or Mann-Whitney in small samples: how do you choose?

Inspired by a question on Cross Validated ·

t testnonparametricnormalitysample size

The short answer

Decide before seeing the data, from what you know about the outcome and the question you want answered. With small groups a normality check tells you little. If the outcome is plausibly near normal, use Welch's t test. If it is likely skewed or prone to outliers, Mann-Whitney usually has more power and asks which group tends to score higher. With 3 per group, a two-sided Mann-Whitney test cannot reach p < .05 at all.

The short answer

Make the choice in advance, from knowledge about the variable rather than from the sample in front of you. Reaction times, incomes, hospital stays and counts are usually right-skewed; scale totals built from many items are often close to symmetric. That prior knowledge is worth far more than a test or a plot of 8 or 10 observations.

Then match the test to the question. Welch's t test compares means. The Mann-Whitney (Wilcoxon rank-sum) test asks whether a randomly chosen score from one group tends to exceed a randomly chosen score from the other; it only becomes a test of medians if both groups have the same shape. A permutation test keeps whatever statistic you care about, such as the difference in means, and gets its p value by reshuffling group labels instead of from a normal-theory formula.

This page is about the two-group, small-sample choice. For the paired version and its symmetry assumption, see What are the assumptions of the Wilcoxon signed-rank test?.

Why "check normality, then pick a test" works poorly with small n

The same logic applies to variances: rather than testing equality of variances first, use Welch's version of the t test by default (see Welch vs Student t test).

See it in R and Python

The code first shows the smallest p value Mann-Whitney can give with 3 to 6 per group. It then simulates 4,000 studies with 10 per group for each of four settings and records how often each approach rejects at α = .05: Welch's t test, Mann-Whitney, and a two-stage rule that runs Shapiro-Wilk on both groups and switches to Mann-Whitney if either fails. Finally it analyses one small skewed study three ways, including an exact permutation test of the difference in means.

R

set.seed(121852)

# 1. Tiny groups: the smallest two-sided p value Mann-Whitney can ever give
n_per_group <- 3:6
round(2 / choose(2 * n_per_group, n_per_group), 4)
wilcox.test(1:3, 4:6)$p.value     # complete separation, 3 vs 3

# 2. Rejection rates at alpha = .05, 10 per group, 4,000 simulated studies each
reject <- function(gen, shift, n = 10, reps = 4000) {
  r <- replicate(reps, {
    x <- gen(n); y <- gen(n) + shift
    p_t  <- t.test(x, y)$p.value           # Welch t test
    p_mw <- wilcox.test(x, y)$p.value      # Mann-Whitney (exact)
    pre  <- shapiro.test(x)$p.value < .05 || shapiro.test(y)$p.value < .05
    c(welch_t = p_t < .05, mann_whitney = p_mw < .05,
      two_stage = ifelse(pre, p_mw, p_t) < .05)
  })
  round(rowMeans(r), 3)
}
skewed <- function(n) rlnorm(n, 0, 1)      # strongly right-skewed
reject(rnorm, 0)     # normal, no effect: Type I error
reject(rnorm, 1)     # normal, shift of 1 SD: power
reject(skewed, 0)    # skewed, no effect: Type I error
reject(skewed, 1)    # skewed, shift of 1: power

# 3. One small skewed study: three tests side by side
x <- rlnorm(8, 0, 1); y <- rlnorm(8, 0, 1) + 1.5
round(c(median_x = median(x), median_y = median(y)), 2)
wt <- t.test(y, x); mw <- wilcox.test(y, x)
all_y <- combn(16, 8)                     # every way to relabel 8 of 16 as "y"
z <- c(x, y)
perm_diff <- apply(all_y, 2, function(i) mean(z[i]) - mean(z[-i]))
obs <- mean(y) - mean(x)
round(c(welch_t = unname(wt$statistic), df = unname(wt$parameter),
        p_t = wt$p.value, U = unname(mw$statistic), p_mw = mw$p.value,
        p_perm = mean(abs(perm_diff) >= abs(obs) - 1e-12),
        n_perms = ncol(all_y), r_rb = 2 * unname(mw$statistic) / 64 - 1), 4)

Python

import itertools, math
import numpy as np
from scipy import stats

rng = np.random.default_rng(121852)

# 1. Tiny groups: the smallest two-sided p value Mann-Whitney can ever give
print([round(2 / math.comb(2 * n, n), 4) for n in range(3, 7)])
print(stats.mannwhitneyu([1, 2, 3], [4, 5, 6], method="exact").pvalue)

# Exact two-sided Mann-Whitney p value for every possible U (10 per group, no ties)
counts = np.zeros(101)
for idx in itertools.combinations(range(20), 10):      # ranks held by group y
    counts[sum(idx) - 45] += 1
cdf = np.cumsum(counts) / counts.sum()
P_MW = np.minimum(1, 2 * np.minimum(cdf, 1 - np.concatenate([[0], cdf[:-1]])))

# 2. Rejection rates at alpha = .05, 10 per group, 4,000 simulated studies each
def reject(gen, shift, n=10, reps=4000):
    x, y = gen((reps, n)), gen((reps, n)) + shift
    vx, vy = x.var(1, ddof=1) / n, y.var(1, ddof=1) / n
    t = (x.mean(1) - y.mean(1)) / np.sqrt(vx + vy)      # Welch t test
    df = (vx + vy) ** 2 / ((vx ** 2 + vy ** 2) / (n - 1))
    p_t = 2 * stats.t.sf(np.abs(t), df)
    p_mw = P_MW[(y[:, :, None] > x[:, None, :]).sum(axis=(1, 2))]   # U -> exact p
    pre = np.array([min(stats.shapiro(a).pvalue, stats.shapiro(b).pvalue) < .05
                    for a, b in zip(x, y)])
    rates = [p_t < .05, p_mw < .05, np.where(pre, p_mw, p_t) < .05]
    return [round(float(r.mean()), 3) for r in rates]   # welch, mann-whitney, two-stage

normal = lambda size: rng.normal(0, 1, size)
skewed = lambda size: rng.lognormal(0, 1, size)         # strongly right-skewed
print("normal, no effect ", reject(normal, 0))
print("normal, shift 1 SD", reject(normal, 1))
print("skewed, no effect ", reject(skewed, 0))
print("skewed, shift 1   ", reject(skewed, 1))

# 3. One small skewed study: three tests side by side
x, y = rng.lognormal(0, 1, 8), rng.lognormal(0, 1, 8) + 1.5
print("medians", round(np.median(x), 2), round(np.median(y), 2))
wt = stats.ttest_ind(y, x, equal_var=False)
mw = stats.mannwhitneyu(y, x, method="exact")
z = np.concatenate([x, y])
sum_y = z[np.array(list(itertools.combinations(range(16), 8)))].sum(1)  # all relabellings
diffs = sum_y / 8 - (z.sum() - sum_y) / 8
p_perm = np.mean(np.abs(diffs) >= abs(y.mean() - x.mean()) - 1e-12)
print("t", round(wt.statistic, 4), "p_t", round(wt.pvalue, 4), "U", mw.statistic,
      "p_mw", round(mw.pvalue, 4), "p_perm", round(p_perm, 4), "n_perms", len(diffs),
      "r_rb", round(2 * mw.statistic / 64 - 1, 4))

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly but show the same pattern. To keep the Python simulation fast it vectorises Welch's t and looks up exact Mann-Whitney p values from a precomputed table of all possible U values.

The two-stage rule did about as well as the better single test in these two extreme settings, because strongly skewed samples usually fail Shapiro-Wilk even at n = 10. Its value lies in exactly those easy cases; for milder departures the screen often passes and the rule simply runs the t test. Knowing in advance that the outcome is skewed gets you the same power without the extra step.

In the single skewed study (8 per group, medians 0.86 and 2.50), Welch's test gave t(10.82) = 1.52, p = .158, the exact permutation test of the mean difference over all 12,870 relabellings gave p = .154, and Mann-Whitney gave U = 50, p = .065. The permutation test agrees with the t test because both use the difference in means, which a few large values dominate. Python's sample from the same design happened to show a much smaller difference (Welch p = .698, permutation p = .711, Mann-Whitney p = .442), with the same ordering. That swing between two random samples of a real 1.5-unit shift is itself a reminder of how unstable results are with 8 per group.

A practical way to decide

  1. State the question. If you need a statement about means (for example, total cost), use a test of means: Welch's t test or a permutation test of the mean difference. If "which group tends to score higher" is the real question, Mann-Whitney fits it directly.
  2. Use prior knowledge of the outcome's shape. Previous studies, the way the variable is produced, or a bounded or count scale tell you more than a small sample can. If a log scale makes sense for the variable on substantive grounds, a t test on the logs compares geometric means, and you can plan that in advance.
  3. Prefer Welch to Student for the t test, without a variance pre-test.
  4. Check the floor. With 3 or fewer per group, a two-sided rank test cannot reach .05; plan a larger sample or a different design. A power analysis before data collection is the honest fix for most small-sample dilemmas.
  5. Pre-register or state the choice in your method section so readers can see it was not chosen after looking at the results.

For help matching a design to a test, try the test chooser. For what a small sample costs in power more generally, see Does a t test need a minimum sample size?, and for more plain-language guides see the DASS blog.

How to report a t test or a Mann-Whitney U test in APA style (7th edition)

Report the test you planned, with descriptive statistics that match it: means and standard deviations for a t test, medians for Mann-Whitney. With the single study above:

"Scores were higher in the treatment group (Mdn = 2.50, n = 8) than in the control group (Mdn = 0.86, n = 8), but a two-sided exact Mann-Whitney U test did not reach significance, U = 50, p = .065, rank-biserial r = .56."

"A Welch's t test found no significant difference in means, t(10.82) = 1.52, p = .158." Add the means, standard deviations and a 95% CI for the difference, written as [LL, UL] in the outcome's units.

In the method section, say which test was planned and why, for example "Because response times are typically right-skewed, groups were compared with a two-sided exact Mann-Whitney U test." Report only one primary test; running both and reporting the smaller p value inflates the error rate.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.