Answer

Why does the central limit theorem work? An intuitive explanation

Inspired by a question on Cross Validated ·

distributionssampling

The short answer

When you add up many independent pieces, extreme results need many pieces to be extreme in the same direction, which is rare, while middling totals can be reached in a huge number of ways. Averaging also shrinks every feature of the original shape (such as skewness) faster than it shrinks the spread, so after enough pieces only the center and the spread are left, and the normal curve is the shape defined by those two alone.

The idea in one paragraph

The central limit theorem says that if you take many independent draws from almost any population with a finite variance and average them, the average follows a distribution that gets closer and closer to a normal curve as the number of draws grows. It is why a normal approximation works for the binomial (a count of successes is a sum of many 0/1 values), and why t tests and confidence intervals for means behave well even when the raw data are not normal.

The intuition rests on two facts: there are far more ways to land in the middle than at the edges, and averaging erases the quirks of the original shape faster than it erases the spread. The sections below take them one at a time.

Reason 1: there are many more ways to reach the middle

Roll one fair die and every face from 1 to 6 is equally likely: a flat distribution. Roll two and add them. A total of 2 needs both dice to show 1, which happens in only one way out of 36. A total of 7 can be made in six ways (1+6, 2+5, 3+4 and the reverse orders). The flat shape has already turned into a triangle.

With three dice the counts for totals 3 to 18 are 1, 3, 6, 10, 15, 21, 25, 27, 27, 25, 21, 15, 10, 6, 3, 1, which already looks like a bell. To get an extreme total, every die has to be extreme in the same direction at once. To get a middling total, the highs and lows just need to roughly balance, and there are enormously many combinations that do that. Adding more pieces makes this imbalance between the edges and the middle ever stronger.

Reason 2: averaging wipes out the shape but not the spread

That explains a hump, but not why the hump is specifically normal. The second idea does. Every distribution has a center (mean), a spread (variance), and then finer features: asymmetry (skewness), tail weight (kurtosis) and so on. When you average n independent values, these features shrink at different rates once you rescale to a common spread:

So as n grows, everything that made your population distinctive melts away, and only a center and a spread survive. The normal distribution is exactly the shape that is fully described by a mean and a variance, with zero skewness and zero excess kurtosis. It is what remains when all the other detail has been averaged out. It is also stable: the sum of independent normal variables is normal again, so once you arrive at it, more averaging keeps you there.

See it in R and Python

The script draws samples from an exponential population (mean 1, standard deviation 1, skewness 2, heavily right-skewed), computes 20,000 sample means for each sample size, and checks their spread, their skewness, and how much of each tail lies beyond 1.96 standard errors. A normal curve would put 0.025 in each tail. It also counts the dice totals exactly. The figures quoted come from one seeded run of the R code; Python's random numbers differ, so its figures differ slightly but show the same pattern (the dice counts are identical).

R

set.seed(3734)
reps <- 20000   # number of simulated samples for each sample size

skewness <- function(x) mean(((x - mean(x)) / sd(x))^3)

# Population: exponential with mean 1 and SD 1 (strongly right-skewed, skewness 2)
res <- t(sapply(c(1, 2, 5, 30), function(n) {
  means <- colMeans(matrix(rexp(n * reps, rate = 1), nrow = n))
  c(n = n,
    sd_means = sd(means), theory_sd = 1 / sqrt(n),
    skew_means = skewness(means), theory_skew = 2 / sqrt(n),
    # a normal curve puts 0.025 in each tail beyond 1.96 SE
    low_tail = mean(means < 1 - 1.96 / sqrt(n)),
    high_tail = mean(means > 1 + 1.96 / sqrt(n)))
}))
round(res, 3)

# Exact: ways to roll each total with one, two and three fair dice
one <- rep(1, 6)
two <- convolve(one, rev(one), type = "open")
three <- convolve(two, rev(one), type = "open")
round(two)     # totals 2 to 12: a triangle
round(three)   # totals 3 to 18: already bell-shaped

Python

import numpy as np
import pandas as pd

rng = np.random.default_rng(3734)
reps = 20000   # number of simulated samples for each sample size

def skewness(x):
    z = (x - x.mean()) / x.std(ddof=1)
    return np.mean(z**3)

# Population: exponential with mean 1 and SD 1 (strongly right-skewed, skewness 2)
rows = []
for n in [1, 2, 5, 30]:
    means = rng.exponential(scale=1, size=(reps, n)).mean(axis=1)
    rows.append({
        "n": n,
        "sd_means": means.std(ddof=1), "theory_sd": 1 / np.sqrt(n),
        "skew_means": skewness(means), "theory_skew": 2 / np.sqrt(n),
        # a normal curve puts 0.025 in each tail beyond 1.96 SE
        "low_tail": np.mean(means < 1 - 1.96 / np.sqrt(n)),
        "high_tail": np.mean(means > 1 + 1.96 / np.sqrt(n)),
    })
print(pd.DataFrame(rows).round(3).to_string(index=False))

# Exact: ways to roll each total with one, two and three fair dice
one = np.ones(6, dtype=int)
two = np.convolve(one, one)
three = np.convolve(two, one)
print(two)     # totals 2 to 12: a triangle
print(three)   # totals 3 to 18: already bell-shaped
Output from one seeded run of the R code (same numbers on a second run):
      n sd_means theory_sd skew_means theory_skew low_tail high_tail
[1,]  1    0.990     1.000      1.976       2.000    0.000     0.050
[2,]  2    0.704     0.707      1.381       1.414    0.000     0.048
[3,]  5    0.450     0.447      0.948       0.894    0.000     0.045
[4,] 30    0.183     0.183      0.387       0.365    0.014     0.035
 [1] 1 2 3 4 5 6 5 4 3 2 1
 [1]  1  3  6 10 15 21 25 27 27 25 21 15 10  6  3  1

The spread of the means tracks 1/sqrt(n) closely (0.183 at n = 30), and the skewness falls roughly as 2/sqrt(n), from 1.976 for single values to 0.387 for averages of 30. The tails tell the same story: for single values the whole 5% sits in the upper tail and none in the lower one, while at n = 30 the split is 0.014 below and 0.035 above. That is much closer to the normal 0.025 and 0.025, but still visibly lopsided, which is a useful reminder that convergence is gradual.

What the theorem does not promise

How to report a mean and confidence interval in APA style (7th edition)

The central limit theorem is usually the silent justification for a normal-based confidence interval around a mean. Report the mean, the standard deviation and the sample size, then the confidence level and the interval in square brackets, in the same units as the mean:

"Participants waited an average of x.xx minutes (SD = x.xx, n = xx), 95% CI [LL, UL]."

If the raw data were clearly skewed and you relied on the sample size for the normal approximation, say so in the method section, for example: "Because the sample was large, confidence intervals for the mean were based on the t distribution; a bootstrap interval gave similar results." Only write the second part if you actually ran that check.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.