Answer
Does a t test need a minimum sample size?
The short answer
No. With roughly normal data the t test keeps its stated error rate with as few as two or three observations, and 30 or 40 is a rule of thumb, not a requirement. What a small sample costs is power: with 15 per group you will usually miss a medium effect. Small samples also make skewed data and outliers harder to detect and more damaging. Justify your n with a power analysis, not a cut-off.
The short answer
There is no sample size below which a t test becomes invalid in the mathematical sense. The t distribution was worked out precisely for small samples: if the data come from a normal distribution, the test rejects a true null hypothesis 5% of the time at α = .05 whether you have 3 observations or 3,000. The popular cut-offs ("at least 30", "at least 40") are rules of thumb that mix together two separate worries, and neither of them is validity in that sense.
- Robustness. The guarantee above assumes normality. With skewed data or outliers, small samples leave the test's error rate less accurate, and they also give you too little data to see the problem.
- Power. A small sample can be perfectly valid and still have little chance of detecting an effect of realistic size. A non-significant result from a small study says very little.
So the honest reply to "your sample is too small" is not a citation that some minimum is acceptable. It is to show whether the assumptions are plausible and how much power the design had.
See it in R and Python
The code simulates 10,000 samples at each size from 3 to 40 and counts how often the t test gives p < .05 when the null hypothesis is true. It uses normal data, skewed (exponential) data for a one-sample test, and two equal-sized groups drawn from the same skewed distribution. It then computes exact power for a two-sample test with 15 and 40 per group, and the size needed for 80% power.
R
set.seed(37993)
reps <- 10000
sizes <- c(3, 5, 10, 15, 40)
# 1. Share of p < .05 when the null is true (one-sample t test)
reject_rate <- function(n, draw, mu) {
mean(replicate(reps, t.test(draw(n), mu = mu)$p.value < 0.05))
}
normal_data <- sapply(sizes, reject_rate, draw = rnorm, mu = 0) # normal, mean 0
skewed_data <- sapply(sizes, reject_rate, draw = rexp, mu = 1) # exponential, mean 1
round(rbind(n = sizes, normal_data, skewed_data), 3)
# 2. Two equal-sized groups drawn from the same skewed distribution (Welch t test)
two_groups <- sapply(sizes, function(n)
mean(replicate(reps, t.test(rexp(n), rexp(n))$p.value < 0.05)))
round(rbind(n = sizes, two_groups), 3)
# 3. Power of a two-sample t test (exact, no simulation; two-sided, alpha = .05)
pw <- function(n, d) power.t.test(n = n, delta = d, sd = 1)$power
round(outer(c(n15 = 15, n40 = 40), c(d0.5 = 0.5, d0.8 = 0.8), Vectorize(pw)), 3)
# 4. Sample size per group needed for 80% power
ceiling(sapply(c(d0.5 = 0.5, d0.8 = 0.8), function(d)
power.t.test(delta = d, sd = 1, power = 0.8)$n))Python
import numpy as np
from scipy import stats, optimize
rng = np.random.default_rng(37993)
reps = 10000
sizes = [3, 5, 10, 15, 40]
# 1. Share of p < .05 when the null is true (one-sample t test)
def reject_rate(n, draw, mu):
x = draw(size=(reps, n))
return round(float(np.mean(stats.ttest_1samp(x, mu, axis=1).pvalue < 0.05)), 3)
print("normal_data", [reject_rate(n, rng.normal, 0) for n in sizes]) # normal, mean 0
print("skewed_data", [reject_rate(n, rng.exponential, 1) for n in sizes]) # exponential, mean 1
# 2. Two equal-sized groups drawn from the same skewed distribution (Welch t test)
def two_groups(n):
x, y = rng.exponential(size=(reps, n)), rng.exponential(size=(reps, n))
return round(float(np.mean(stats.ttest_ind(x, y, axis=1, equal_var=False).pvalue < 0.05)), 3)
print("two_groups", [two_groups(n) for n in sizes])
# 3. Power of a two-sample t test (exact; same one-tail formula as R's power.t.test)
def power(n, d, alpha=0.05):
df = 2 * n - 2
ncp = d * np.sqrt(n / 2)
return float(stats.nct.sf(stats.t.ppf(1 - alpha / 2, df), df, ncp))
for n in (15, 40):
print(f"n{n}", {f"d{d}": round(power(n, d), 3) for d in (0.5, 0.8)})
# 4. Sample size per group needed for 80% power
for d in (0.5, 0.8):
n_needed = optimize.brentq(lambda n: power(n, d) - 0.8, 2, 1000)
print(f"d{d}", int(np.ceil(n_needed))) [,1] [,2] [,3] [,4] [,5]
n 3.000 5.000 10.000 15.000 40.000
normal_data 0.046 0.051 0.049 0.050 0.046
skewed_data 0.119 0.118 0.099 0.092 0.064
[,1] [,2] [,3] [,4] [,5]
n 3.000 5.000 10.000 15.000 40.000
two_groups 0.026 0.026 0.038 0.041 0.051
d0.5 d0.8
n15 0.262 0.562
n40 0.598 0.942
d0.5 d0.8
64 26
The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 12.2% instead of 11.9% for skewed data with n = 3), but they show the same pattern. The power calculations involve no randomness, and Python reproduces them exactly.
- Normal data: fine at every size. The false-positive rate stayed between 4.6% and 5.1% from n = 3 to n = 40. Small samples did not break the test.
- Skewed data, one sample: too many false positives when small. The rate was 11.9% at n = 3, 9.2% at n = 15 and 6.4% at n = 40. Larger samples help because the sample mean becomes closer to normal, but the improvement is gradual, not a switch at 30.
- Two groups with the same skewed shape: much less of a problem. With equal group sizes, the skewness largely cancels in the difference between means. The test was cautious rather than too liberal at small sizes (2.6% at n = 3 per group, 4.1% at 15) and on target at 40 (5.1%).
- Power is the real cost. With 15 per group, the chance of detecting a medium effect (d = 0.5) was 26.2%, and even a large effect (d = 0.8) only 56.2%. With 40 per group the figures rose to 59.8% and 94.2%. For 80% power you need 64 per group for a medium effect and 26 for a large one.
Where do 30 and 40 come from?
The "30" rule comes from the central limit theorem: as samples grow, the sampling distribution of the mean looks more and more normal, so the normality assumption matters less. It is a reasonable reminder, but the number that is "large enough" depends on how skewed or heavy-tailed the data are. For roughly symmetric data, far fewer than 30 is fine; for very skewed data, 30 may not be enough. A requirement such as "at least 40" is usually really about power, and whether 40 is enough depends on the effect you expect, not on the test.
The same reasoning applies to the F test in an ANOVA, which behaves much like the t test (with two groups it is the same test). The F test for comparing two variances is the exception: it is very sensitive to non-normality at any sample size, so use Levene's test or a robust alternative for that question.
What to do if your sample is small
- Look at the data. Plot each group (a dot plot or a Q-Q plot). A formal normality test has almost no power with 15 observations, so it cannot reassure you (see Is a normality test worth running on your data?).
- Use Welch's version for two groups so unequal variances do not distort the result (see Welch vs Student t test).
- Report a confidence interval and an effect size, not just the p value. A wide interval shows readers honestly how little a small sample can pin down.
- Add a check if the data are skewed or have outliers, for example a permutation test or a rank-based test, and say whether it agrees.
- Justify the sample size with power, ideally planned in advance. If the sample was fixed by circumstances (a small population, for instance), say so, report the smallest effect the design could detect with 80% power, and treat a non-significant result as inconclusive rather than as evidence of no effect.
Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report a t test in APA style (7th edition)
With a small sample, report the test in full and add a sentence that justifies the sample size. Replace the placeholders with your own values:
"Scores were higher in the intervention group (M = x.xx, SD = x.xx, n = 15) than in the comparison group (M = x.xx, SD = x.xx, n = 15), t(df) = x.xx, p = .xxx, 95% CI [LL, UL], d = x.xx."
In the method section, state how the sample size was set and what it could detect, using the power figures from the example above: "The sample size was limited by the number of eligible participants. With 15 participants per group, a two-sided independent-samples t test at α = .05 had 56% power to detect a large effect (d = 0.80) and 26% power to detect a medium effect (d = 0.50)." If the test was not significant, add that the study was not powered to rule out effects of that size.
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Power and sample size calculators
- Which statistical test should I use?
- Welch vs Student t test: should you just always use Welch?
- Is a normality test worth running on your data?
- How to interpret a QQ plot
- Student's t-test (Wikipedia)
More answered questions
- Welch vs Student t test: should you just always use Welch?
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
- Log-transforming data: when it helps, and what it changes
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.