Answer
t test or Mann-Whitney in small samples: how do you choose?
The short answer
Decide before seeing the data, from what you know about the outcome and the question you want answered. With small groups a normality check tells you little. If the outcome is plausibly near normal, use Welch's t test. If it is likely skewed or prone to outliers, Mann-Whitney usually has more power and asks which group tends to score higher. With 3 per group, a two-sided Mann-Whitney test cannot reach p < .05 at all.
The short answer
Make the choice in advance, from knowledge about the variable rather than from the sample in front of you. Reaction times, incomes, hospital stays and counts are usually right-skewed; scale totals built from many items are often close to symmetric. That prior knowledge is worth far more than a test or a plot of 8 or 10 observations.
Then match the test to the question. Welch's t test compares means. The Mann-Whitney (Wilcoxon rank-sum) test asks whether a randomly chosen score from one group tends to exceed a randomly chosen score from the other; it only becomes a test of medians if both groups have the same shape. A permutation test keeps whatever statistic you care about, such as the difference in means, and gets its p value by reshuffling group labels instead of from a normal-theory formula.
This page is about the two-group, small-sample choice. For the paired version and its symmetry assumption, see What are the assumptions of the Wilcoxon signed-rank test?.
Why "check normality, then pick a test" works poorly with small n
- The check has little power. With 10 observations per group, a Shapiro-Wilk test or a Q-Q plot will miss many non-normal distributions, including some that matter. Passing the check is not evidence of normality; see Is a normality test worth running on your data?.
- The two steps interact. When the data decide which test you run, the final p value no longer has exactly its stated meaning, because the test is applied only to samples that passed (or failed) the screen. The overall error rate of the whole procedure is what counts, and it is not guaranteed to be 5%.
- Large-sample efficiency figures do not transfer. Under normality, Mann-Whitney's asymptotic relative efficiency against the t test is 3/π, about 0.955, and it never falls below about 0.864 for any continuous distribution. Those are limits as n grows. With tiny groups the rank test is also limited by how few distinct p values it can produce.
- Rank tests have a floor. With 3 per group, the smallest possible two-sided Mann-Whitney p value is .10, even when the groups do not overlap at all. With 4 per group it is about .029.
The same logic applies to variances: rather than testing equality of variances first, use Welch's version of the t test by default (see Welch vs Student t test).
See it in R and Python
The code first shows the smallest p value Mann-Whitney can give with 3 to 6 per group. It then simulates 4,000 studies with 10 per group for each of four settings and records how often each approach rejects at α = .05: Welch's t test, Mann-Whitney, and a two-stage rule that runs Shapiro-Wilk on both groups and switches to Mann-Whitney if either fails. Finally it analyses one small skewed study three ways, including an exact permutation test of the difference in means.
R
set.seed(121852)
# 1. Tiny groups: the smallest two-sided p value Mann-Whitney can ever give
n_per_group <- 3:6
round(2 / choose(2 * n_per_group, n_per_group), 4)
wilcox.test(1:3, 4:6)$p.value # complete separation, 3 vs 3
# 2. Rejection rates at alpha = .05, 10 per group, 4,000 simulated studies each
reject <- function(gen, shift, n = 10, reps = 4000) {
r <- replicate(reps, {
x <- gen(n); y <- gen(n) + shift
p_t <- t.test(x, y)$p.value # Welch t test
p_mw <- wilcox.test(x, y)$p.value # Mann-Whitney (exact)
pre <- shapiro.test(x)$p.value < .05 || shapiro.test(y)$p.value < .05
c(welch_t = p_t < .05, mann_whitney = p_mw < .05,
two_stage = ifelse(pre, p_mw, p_t) < .05)
})
round(rowMeans(r), 3)
}
skewed <- function(n) rlnorm(n, 0, 1) # strongly right-skewed
reject(rnorm, 0) # normal, no effect: Type I error
reject(rnorm, 1) # normal, shift of 1 SD: power
reject(skewed, 0) # skewed, no effect: Type I error
reject(skewed, 1) # skewed, shift of 1: power
# 3. One small skewed study: three tests side by side
x <- rlnorm(8, 0, 1); y <- rlnorm(8, 0, 1) + 1.5
round(c(median_x = median(x), median_y = median(y)), 2)
wt <- t.test(y, x); mw <- wilcox.test(y, x)
all_y <- combn(16, 8) # every way to relabel 8 of 16 as "y"
z <- c(x, y)
perm_diff <- apply(all_y, 2, function(i) mean(z[i]) - mean(z[-i]))
obs <- mean(y) - mean(x)
round(c(welch_t = unname(wt$statistic), df = unname(wt$parameter),
p_t = wt$p.value, U = unname(mw$statistic), p_mw = mw$p.value,
p_perm = mean(abs(perm_diff) >= abs(obs) - 1e-12),
n_perms = ncol(all_y), r_rb = 2 * unname(mw$statistic) / 64 - 1), 4)Python
import itertools, math
import numpy as np
from scipy import stats
rng = np.random.default_rng(121852)
# 1. Tiny groups: the smallest two-sided p value Mann-Whitney can ever give
print([round(2 / math.comb(2 * n, n), 4) for n in range(3, 7)])
print(stats.mannwhitneyu([1, 2, 3], [4, 5, 6], method="exact").pvalue)
# Exact two-sided Mann-Whitney p value for every possible U (10 per group, no ties)
counts = np.zeros(101)
for idx in itertools.combinations(range(20), 10): # ranks held by group y
counts[sum(idx) - 45] += 1
cdf = np.cumsum(counts) / counts.sum()
P_MW = np.minimum(1, 2 * np.minimum(cdf, 1 - np.concatenate([[0], cdf[:-1]])))
# 2. Rejection rates at alpha = .05, 10 per group, 4,000 simulated studies each
def reject(gen, shift, n=10, reps=4000):
x, y = gen((reps, n)), gen((reps, n)) + shift
vx, vy = x.var(1, ddof=1) / n, y.var(1, ddof=1) / n
t = (x.mean(1) - y.mean(1)) / np.sqrt(vx + vy) # Welch t test
df = (vx + vy) ** 2 / ((vx ** 2 + vy ** 2) / (n - 1))
p_t = 2 * stats.t.sf(np.abs(t), df)
p_mw = P_MW[(y[:, :, None] > x[:, None, :]).sum(axis=(1, 2))] # U -> exact p
pre = np.array([min(stats.shapiro(a).pvalue, stats.shapiro(b).pvalue) < .05
for a, b in zip(x, y)])
rates = [p_t < .05, p_mw < .05, np.where(pre, p_mw, p_t) < .05]
return [round(float(r.mean()), 3) for r in rates] # welch, mann-whitney, two-stage
normal = lambda size: rng.normal(0, 1, size)
skewed = lambda size: rng.lognormal(0, 1, size) # strongly right-skewed
print("normal, no effect ", reject(normal, 0))
print("normal, shift 1 SD", reject(normal, 1))
print("skewed, no effect ", reject(skewed, 0))
print("skewed, shift 1 ", reject(skewed, 1))
# 3. One small skewed study: three tests side by side
x, y = rng.lognormal(0, 1, 8), rng.lognormal(0, 1, 8) + 1.5
print("medians", round(np.median(x), 2), round(np.median(y), 2))
wt = stats.ttest_ind(y, x, equal_var=False)
mw = stats.mannwhitneyu(y, x, method="exact")
z = np.concatenate([x, y])
sum_y = z[np.array(list(itertools.combinations(range(16), 8)))].sum(1) # all relabellings
diffs = sum_y / 8 - (z.sum() - sum_y) / 8
p_perm = np.mean(np.abs(diffs) >= abs(y.mean() - x.mean()) - 1e-12)
print("t", round(wt.statistic, 4), "p_t", round(wt.pvalue, 4), "U", mw.statistic,
"p_mw", round(mw.pvalue, 4), "p_perm", round(p_perm, 4), "n_perms", len(diffs),
"r_rb", round(2 * mw.statistic / 64 - 1, 4))The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly but show the same pattern. To keep the Python simulation fast it vectorises Welch's t and looks up exact Mann-Whitney p values from a precomputed table of all possible U values.
- The floor. The smallest two-sided Mann-Whitney p values are .1000, .0286, .0079 and .0022 for 3, 4, 5 and 6 per group. Three scores against three completely separated scores give p = .10.
- Normal data, no effect. Rejection rates were .048 (Welch), .043 (Mann-Whitney) and .049 (two-stage). All three hold their error rate; the exact rank test is slightly conservative because its p values move in steps.
- Normal data, 1 SD shift. Power was .540 (Welch), .501 (Mann-Whitney) and .538 (two-stage). When the data really are normal, the rank test gives up only a little power.
- Skewed (lognormal) data, no effect. Rates were .031, .046 and .052. Welch's test became conservative; the two-stage rule landed slightly above 5%.
- Skewed data, shift of 1. Power was .292 (Welch), .510 (Mann-Whitney) and .509 (two-stage). Here the rank test nearly doubles the power of the t test.
The two-stage rule did about as well as the better single test in these two extreme settings, because strongly skewed samples usually fail Shapiro-Wilk even at n = 10. Its value lies in exactly those easy cases; for milder departures the screen often passes and the rule simply runs the t test. Knowing in advance that the outcome is skewed gets you the same power without the extra step.
In the single skewed study (8 per group, medians 0.86 and 2.50), Welch's test gave t(10.82) = 1.52, p = .158, the exact permutation test of the mean difference over all 12,870 relabellings gave p = .154, and Mann-Whitney gave U = 50, p = .065. The permutation test agrees with the t test because both use the difference in means, which a few large values dominate. Python's sample from the same design happened to show a much smaller difference (Welch p = .698, permutation p = .711, Mann-Whitney p = .442), with the same ordering. That swing between two random samples of a real 1.5-unit shift is itself a reminder of how unstable results are with 8 per group.
A practical way to decide
- State the question. If you need a statement about means (for example, total cost), use a test of means: Welch's t test or a permutation test of the mean difference. If "which group tends to score higher" is the real question, Mann-Whitney fits it directly.
- Use prior knowledge of the outcome's shape. Previous studies, the way the variable is produced, or a bounded or count scale tell you more than a small sample can. If a log scale makes sense for the variable on substantive grounds, a t test on the logs compares geometric means, and you can plan that in advance.
- Prefer Welch to Student for the t test, without a variance pre-test.
- Check the floor. With 3 or fewer per group, a two-sided rank test cannot reach .05; plan a larger sample or a different design. A power analysis before data collection is the honest fix for most small-sample dilemmas.
- Pre-register or state the choice in your method section so readers can see it was not chosen after looking at the results.
For help matching a design to a test, try the test chooser. For what a small sample costs in power more generally, see Does a t test need a minimum sample size?, and for more plain-language guides see the DASS blog.
How to report a t test or a Mann-Whitney U test in APA style (7th edition)
Report the test you planned, with descriptive statistics that match it: means and standard deviations for a t test, medians for Mann-Whitney. With the single study above:
"Scores were higher in the treatment group (Mdn = 2.50, n = 8) than in the control group (Mdn = 0.86, n = 8), but a two-sided exact Mann-Whitney U test did not reach significance, U = 50, p = .065, rank-biserial r = .56."
"A Welch's t test found no significant difference in means, t(10.82) = 1.52, p = .158." Add the means, standard deviations and a 95% CI for the difference, written as [LL, UL] in the outcome's units.
In the method section, say which test was planned and why, for example "Because response times are typically right-skewed, groups were compared with a two-sided exact Mann-Whitney U test." Report only one primary test; running both and reporting the smaller p value inflates the error rate.
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- APA 7 formatter for nonparametric tests
- What are the assumptions of the Wilcoxon signed-rank test?
- Welch vs Student t test: should you just always use Welch?
- Is a normality test worth running on your data?
- Does a t test need a minimum sample size?
- How do you compare a short ordinal rating scale between two groups?
More answered questions
- Does a t test need a minimum sample size?
- How are alpha and beta (Type I and Type II error rates) related?
- How do you compare a short ordinal rating scale between two groups?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.