Answer

Does a t test need a minimum sample size?

Inspired by a question on Cross Validated ·

t testsample sizepower analysisnormality

The short answer

No. With roughly normal data the t test keeps its stated error rate with as few as two or three observations, and 30 or 40 is a rule of thumb, not a requirement. What a small sample costs is power: with 15 per group you will usually miss a medium effect. Small samples also make skewed data and outliers harder to detect and more damaging. Justify your n with a power analysis, not a cut-off.

The short answer

There is no sample size below which a t test becomes invalid in the mathematical sense. The t distribution was worked out precisely for small samples: if the data come from a normal distribution, the test rejects a true null hypothesis 5% of the time at α = .05 whether you have 3 observations or 3,000. The popular cut-offs ("at least 30", "at least 40") are rules of thumb that mix together two separate worries, and neither of them is validity in that sense.

So the honest reply to "your sample is too small" is not a citation that some minimum is acceptable. It is to show whether the assumptions are plausible and how much power the design had.

See it in R and Python

The code simulates 10,000 samples at each size from 3 to 40 and counts how often the t test gives p < .05 when the null hypothesis is true. It uses normal data, skewed (exponential) data for a one-sample test, and two equal-sized groups drawn from the same skewed distribution. It then computes exact power for a two-sample test with 15 and 40 per group, and the size needed for 80% power.

R

set.seed(37993)
reps  <- 10000
sizes <- c(3, 5, 10, 15, 40)

# 1. Share of p < .05 when the null is true (one-sample t test)
reject_rate <- function(n, draw, mu) {
  mean(replicate(reps, t.test(draw(n), mu = mu)$p.value < 0.05))
}
normal_data <- sapply(sizes, reject_rate, draw = rnorm, mu = 0)  # normal, mean 0
skewed_data <- sapply(sizes, reject_rate, draw = rexp,  mu = 1)  # exponential, mean 1
round(rbind(n = sizes, normal_data, skewed_data), 3)

# 2. Two equal-sized groups drawn from the same skewed distribution (Welch t test)
two_groups <- sapply(sizes, function(n)
  mean(replicate(reps, t.test(rexp(n), rexp(n))$p.value < 0.05)))
round(rbind(n = sizes, two_groups), 3)

# 3. Power of a two-sample t test (exact, no simulation; two-sided, alpha = .05)
pw <- function(n, d) power.t.test(n = n, delta = d, sd = 1)$power
round(outer(c(n15 = 15, n40 = 40), c(d0.5 = 0.5, d0.8 = 0.8), Vectorize(pw)), 3)

# 4. Sample size per group needed for 80% power
ceiling(sapply(c(d0.5 = 0.5, d0.8 = 0.8), function(d)
  power.t.test(delta = d, sd = 1, power = 0.8)$n))

Python

import numpy as np
from scipy import stats, optimize

rng = np.random.default_rng(37993)
reps = 10000
sizes = [3, 5, 10, 15, 40]

# 1. Share of p < .05 when the null is true (one-sample t test)
def reject_rate(n, draw, mu):
    x = draw(size=(reps, n))
    return round(float(np.mean(stats.ttest_1samp(x, mu, axis=1).pvalue < 0.05)), 3)

print("normal_data", [reject_rate(n, rng.normal, 0) for n in sizes])       # normal, mean 0
print("skewed_data", [reject_rate(n, rng.exponential, 1) for n in sizes])  # exponential, mean 1

# 2. Two equal-sized groups drawn from the same skewed distribution (Welch t test)
def two_groups(n):
    x, y = rng.exponential(size=(reps, n)), rng.exponential(size=(reps, n))
    return round(float(np.mean(stats.ttest_ind(x, y, axis=1, equal_var=False).pvalue < 0.05)), 3)

print("two_groups", [two_groups(n) for n in sizes])

# 3. Power of a two-sample t test (exact; same one-tail formula as R's power.t.test)
def power(n, d, alpha=0.05):
    df = 2 * n - 2
    ncp = d * np.sqrt(n / 2)
    return float(stats.nct.sf(stats.t.ppf(1 - alpha / 2, df), df, ncp))

for n in (15, 40):
    print(f"n{n}", {f"d{d}": round(power(n, d), 3) for d in (0.5, 0.8)})

# 4. Sample size per group needed for 80% power
for d in (0.5, 0.8):
    n_needed = optimize.brentq(lambda n: power(n, d) - 0.8, 2, 1000)
    print(f"d{d}", int(np.ceil(n_needed)))
             [,1]  [,2]   [,3]   [,4]   [,5]
n           3.000 5.000 10.000 15.000 40.000
normal_data 0.046 0.051  0.049  0.050  0.046
skewed_data 0.119 0.118  0.099  0.092  0.064
            [,1]  [,2]   [,3]   [,4]   [,5]
n          3.000 5.000 10.000 15.000 40.000
two_groups 0.026 0.026  0.038  0.041  0.051
     d0.5  d0.8
n15 0.262 0.562
n40 0.598 0.942
d0.5 d0.8 
  64   26

The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 12.2% instead of 11.9% for skewed data with n = 3), but they show the same pattern. The power calculations involve no randomness, and Python reproduces them exactly.

Where do 30 and 40 come from?

The "30" rule comes from the central limit theorem: as samples grow, the sampling distribution of the mean looks more and more normal, so the normality assumption matters less. It is a reasonable reminder, but the number that is "large enough" depends on how skewed or heavy-tailed the data are. For roughly symmetric data, far fewer than 30 is fine; for very skewed data, 30 may not be enough. A requirement such as "at least 40" is usually really about power, and whether 40 is enough depends on the effect you expect, not on the test.

The same reasoning applies to the F test in an ANOVA, which behaves much like the t test (with two groups it is the same test). The F test for comparing two variances is the exception: it is very sensitive to non-normality at any sample size, so use Levene's test or a robust alternative for that question.

What to do if your sample is small

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report a t test in APA style (7th edition)

With a small sample, report the test in full and add a sentence that justifies the sample size. Replace the placeholders with your own values:

"Scores were higher in the intervention group (M = x.xx, SD = x.xx, n = 15) than in the comparison group (M = x.xx, SD = x.xx, n = 15), t(df) = x.xx, p = .xxx, 95% CI [LL, UL], d = x.xx."

In the method section, state how the sample size was set and what it could detect, using the power figures from the example above: "The sample size was limited by the number of eligible participants. With 15 participants per group, a two-sided independent-samples t test at α = .05 had 56% power to detect a large effect (d = 0.80) and 26% power to detect a medium effect (d = 0.50)." If the test was not significant, add that the study was not powered to rule out effects of that size.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.