Answer
Why does the central limit theorem work? An intuitive explanation
The short answer
When you add up many independent pieces, extreme results need many pieces to be extreme in the same direction, which is rare, while middling totals can be reached in a huge number of ways. Averaging also shrinks every feature of the original shape (such as skewness) faster than it shrinks the spread, so after enough pieces only the center and the spread are left, and the normal curve is the shape defined by those two alone.
The idea in one paragraph
The central limit theorem says that if you take many independent draws from almost any population with a finite variance and average them, the average follows a distribution that gets closer and closer to a normal curve as the number of draws grows. It is why a normal approximation works for the binomial (a count of successes is a sum of many 0/1 values), and why t tests and confidence intervals for means behave well even when the raw data are not normal.
The intuition rests on two facts: there are far more ways to land in the middle than at the edges, and averaging erases the quirks of the original shape faster than it erases the spread. The sections below take them one at a time.
Reason 1: there are many more ways to reach the middle
Roll one fair die and every face from 1 to 6 is equally likely: a flat distribution. Roll two and add them. A total of 2 needs both dice to show 1, which happens in only one way out of 36. A total of 7 can be made in six ways (1+6, 2+5, 3+4 and the reverse orders). The flat shape has already turned into a triangle.
With three dice the counts for totals 3 to 18 are 1, 3, 6, 10, 15, 21, 25, 27, 27, 25, 21, 15, 10, 6, 3, 1, which already looks like a bell. To get an extreme total, every die has to be extreme in the same direction at once. To get a middling total, the highs and lows just need to roughly balance, and there are enormously many combinations that do that. Adding more pieces makes this imbalance between the edges and the middle ever stronger.
Reason 2: averaging wipes out the shape but not the spread
That explains a hump, but not why the hump is specifically normal. The second idea does. Every distribution has a center (mean), a spread (variance), and then finer features: asymmetry (skewness), tail weight (kurtosis) and so on. When you average n independent values, these features shrink at different rates once you rescale to a common spread:
- The spread of the average is the original standard deviation divided by sqrt(n). This is the standard error, and after rescaling it stays fixed.
- The skewness of the average is the original skewness divided by sqrt(n), so it fades away.
- The excess kurtosis is divided by n, so it fades even faster, and higher-order features vanish faster still.
So as n grows, everything that made your population distinctive melts away, and only a center and a spread survive. The normal distribution is exactly the shape that is fully described by a mean and a variance, with zero skewness and zero excess kurtosis. It is what remains when all the other detail has been averaged out. It is also stable: the sum of independent normal variables is normal again, so once you arrive at it, more averaging keeps you there.
See it in R and Python
The script draws samples from an exponential population (mean 1, standard deviation 1, skewness 2, heavily right-skewed), computes 20,000 sample means for each sample size, and checks their spread, their skewness, and how much of each tail lies beyond 1.96 standard errors. A normal curve would put 0.025 in each tail. It also counts the dice totals exactly. The figures quoted come from one seeded run of the R code; Python's random numbers differ, so its figures differ slightly but show the same pattern (the dice counts are identical).
R
set.seed(3734)
reps <- 20000 # number of simulated samples for each sample size
skewness <- function(x) mean(((x - mean(x)) / sd(x))^3)
# Population: exponential with mean 1 and SD 1 (strongly right-skewed, skewness 2)
res <- t(sapply(c(1, 2, 5, 30), function(n) {
means <- colMeans(matrix(rexp(n * reps, rate = 1), nrow = n))
c(n = n,
sd_means = sd(means), theory_sd = 1 / sqrt(n),
skew_means = skewness(means), theory_skew = 2 / sqrt(n),
# a normal curve puts 0.025 in each tail beyond 1.96 SE
low_tail = mean(means < 1 - 1.96 / sqrt(n)),
high_tail = mean(means > 1 + 1.96 / sqrt(n)))
}))
round(res, 3)
# Exact: ways to roll each total with one, two and three fair dice
one <- rep(1, 6)
two <- convolve(one, rev(one), type = "open")
three <- convolve(two, rev(one), type = "open")
round(two) # totals 2 to 12: a triangle
round(three) # totals 3 to 18: already bell-shapedPython
import numpy as np
import pandas as pd
rng = np.random.default_rng(3734)
reps = 20000 # number of simulated samples for each sample size
def skewness(x):
z = (x - x.mean()) / x.std(ddof=1)
return np.mean(z**3)
# Population: exponential with mean 1 and SD 1 (strongly right-skewed, skewness 2)
rows = []
for n in [1, 2, 5, 30]:
means = rng.exponential(scale=1, size=(reps, n)).mean(axis=1)
rows.append({
"n": n,
"sd_means": means.std(ddof=1), "theory_sd": 1 / np.sqrt(n),
"skew_means": skewness(means), "theory_skew": 2 / np.sqrt(n),
# a normal curve puts 0.025 in each tail beyond 1.96 SE
"low_tail": np.mean(means < 1 - 1.96 / np.sqrt(n)),
"high_tail": np.mean(means > 1 + 1.96 / np.sqrt(n)),
})
print(pd.DataFrame(rows).round(3).to_string(index=False))
# Exact: ways to roll each total with one, two and three fair dice
one = np.ones(6, dtype=int)
two = np.convolve(one, one)
three = np.convolve(two, one)
print(two) # totals 2 to 12: a triangle
print(three) # totals 3 to 18: already bell-shapedOutput from one seeded run of the R code (same numbers on a second run):
n sd_means theory_sd skew_means theory_skew low_tail high_tail
[1,] 1 0.990 1.000 1.976 2.000 0.000 0.050
[2,] 2 0.704 0.707 1.381 1.414 0.000 0.048
[3,] 5 0.450 0.447 0.948 0.894 0.000 0.045
[4,] 30 0.183 0.183 0.387 0.365 0.014 0.035
[1] 1 2 3 4 5 6 5 4 3 2 1
[1] 1 3 6 10 15 21 25 27 27 25 21 15 10 6 3 1
The spread of the means tracks 1/sqrt(n) closely (0.183 at n = 30), and the skewness falls roughly as 2/sqrt(n), from 1.976 for single values to 0.387 for averages of 30. The tails tell the same story: for single values the whole 5% sits in the upper tail and none in the lower one, while at n = 30 the split is 0.014 below and 0.035 above. That is much closer to the normal 0.025 and 0.025, but still visibly lopsided, which is a useful reminder that convergence is gradual.
What the theorem does not promise
- It is about averages and sums, not your raw data. Collecting more observations does not make the data themselves more normal; it makes the distribution of the mean more normal.
- There is no magic sample size. The rule of thumb of 30 is only a convention. A symmetric, light-tailed population is well approximated with a handful of values, while a very skewed one, like the exponential above, still shows asymmetry at 30.
- The variance must be finite. For a population with extremely heavy tails, such as the Cauchy distribution, the average of many values is just as spread out as a single value, and no normal limit appears.
- The pieces must be roughly independent. Strong dependence, for example repeated measurements on the same person treated as independent, can break the approximation. Versions of the theorem allow some dependence and unequal distributions, but not anything whatsoever.
How to report a mean and confidence interval in APA style (7th edition)
The central limit theorem is usually the silent justification for a normal-based confidence interval around a mean. Report the mean, the standard deviation and the sample size, then the confidence level and the interval in square brackets, in the same units as the mean:
"Participants waited an average of x.xx minutes (SD = x.xx, n = xx), 95% CI [LL, UL]."
If the raw data were clearly skewed and you relied on the sample size for the normal approximation, say so in the method section, for example: "Because the sample was large, confidence intervals for the mean were based on the t distribution; a bootstrap interval gave similar results." Only write the second part if you actually ran that check.
Related tools and guides
- Which statistical test should I use?
- Where does the 1/sqrt(n) margin of error come from?
- Why does bootstrapping work?
- Central limit theorem (Wikipedia)
More answered questions
- Log-transforming data: when it helps, and what it changes
- What is the beta distribution, and what is it used for?
- Why does bootstrapping work? A plain-language explanation
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.