Answer
Standard error vs standard deviation: what is the difference?
The short answer
The standard deviation (σ, or s in a sample) measures how spread out individual observations are. The standard error of the mean, σ/sqrt(n), measures how much the average of n observations would move from sample to sample. A total of n independent values has a standard deviation of sqrt(n) × σ. So ask what quantity the question is about: a single value uses σ, an average uses σ/sqrt(n), and a sum uses sqrt(n) × σ.
Two different questions
The standard deviation answers: how far does a typical individual value sit from the mean? If adult heights have a standard deviation of 7 cm, a randomly chosen person is often several centimetres away from the average. Collecting more people does not change this number; it only lets you estimate it more precisely.
The standard error answers a different question: if I repeated the whole study, how much would my estimate change? For a sample mean it equals σ/sqrt(n). It is the standard deviation of the sampling distribution of the mean, not of the data. Because averaging lets high and low values cancel, it shrinks as the sample grows: four times as many observations halve it.
Every standard error is a standard deviation of something. The confusion goes away once you name that something: individual values, an average, or a total.
Where the sqrt(n) rules come from
Both formulas follow from one fact: for independent random variables, variances add. If each of n values has variance σ², their sum has variance nσ², so the sum's standard deviation is sqrt(n) × σ. The total of a box of items spreads out more than a single item does, because the small deviations of each item accumulate.
The mean is the sum divided by n. Dividing a variable by n divides its variance by n², so the mean's variance is nσ²/n² = σ²/n, and its standard deviation is σ/sqrt(n). That is the standard error. The mean of the sum is nμ and the mean of the average is μ.
- One value ("the chance a single item weighs less than ..."): use σ.
- **An average of n** ("the chance the average of n items is below ..."): use σ/sqrt(n) around μ.
- **A total of n** ("the chance a pack of n items weighs less than ..."): use sqrt(n) × σ around nμ.
The last two describe the same event on two scales. A pack total below some limit is exactly the same as the pack's average being below that limit divided by n, so both routes give the same probability as long as you rescale the mean and the spread together.
See it in R and Python
The script imagines cookies whose weights are normal with mean 50 g and standard deviation 4 g, packed 25 to a box. It simulates 20,000 boxes, compares the spread of single cookies, box totals and box means with the formulas, and computes the chance that a box weighs under 1,230 g in two ways. It then takes one growing sample to show what happens to the sample SD and SE as n increases. Figures quoted come from one seeded run of the R code; Python's random numbers differ, so its simulated figures differ slightly but show the same pattern (the exact probabilities are identical).
R
set.seed(48133)
mu <- 50; sigma <- 4; n <- 25 # cookies: mean 50 g, SD 4 g; 25 per box
reps <- 20000 # number of simulated boxes
# Each column is one box of 25 cookies
boxes <- matrix(rnorm(n * reps, mu, sigma), nrow = n)
totals <- colSums(boxes)
means <- colMeans(boxes)
# Spread of single cookies, box totals and box means vs theory
round(c(sd_single = sd(as.vector(boxes)), sd_total = sd(totals),
sd_mean = sd(means)), 3)
c(theory_total = sqrt(n) * sigma, theory_mean = sigma / sqrt(n))
# Same event, two scales: box under 1230 g <=> mean under 49.2 g
c(p_total = pnorm(1230, n * mu, sqrt(n) * sigma),
p_mean = pnorm(1230 / n, mu, sigma / sqrt(n)),
simulated = mean(totals < 1230))
# One real sample at a time: the SD settles, the SE keeps shrinking
x <- rnorm(400, mu, sigma)
res <- t(sapply(c(25, 100, 400), function(k) {
s <- sd(x[1:k])
c(n = k, SD = s, SE = s / sqrt(k))
}))
round(res, 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(48133)
mu, sigma, n = 50, 4, 25 # cookies: mean 50 g, SD 4 g; 25 per box
reps = 20000 # number of simulated boxes
# Each row is one box of 25 cookies
boxes = rng.normal(mu, sigma, size=(reps, n))
totals = boxes.sum(axis=1)
means = boxes.mean(axis=1)
# Spread of single cookies, box totals and box means vs theory
print({"sd_single": round(boxes.std(ddof=1), 3),
"sd_total": round(totals.std(ddof=1), 3),
"sd_mean": round(means.std(ddof=1), 3)})
print({"theory_total": np.sqrt(n) * sigma, "theory_mean": sigma / np.sqrt(n)})
# Same event, two scales: box under 1230 g <=> mean under 49.2 g
print({"p_total": stats.norm.cdf(1230, n * mu, np.sqrt(n) * sigma),
"p_mean": stats.norm.cdf(1230 / n, mu, sigma / np.sqrt(n)),
"simulated": np.mean(totals < 1230)})
# One real sample at a time: the SD settles, the SE keeps shrinking
x = rng.normal(mu, sigma, size=400)
for k in [25, 100, 400]:
s = x[:k].std(ddof=1)
print(f"n = {k:3d} SD = {s:.3f} SE = {s / np.sqrt(k):.3f}")Output from one seeded run of the R code (same numbers on a second run):
sd_single sd_total sd_mean
4.003 19.995 0.800
theory_total theory_mean
20.0 0.8
p_total p_mean simulated
0.1586553 0.1586553 0.1573500
n SD SE
[1,] 25 4.102 0.820
[2,] 100 4.180 0.418
[3,] 400 4.258 0.213
Single cookies vary with a standard deviation of about 4.003 g, box totals with about 19.995 g (theory: sqrt(25) × 4 = 20), and box means with 0.800 g (theory: 4/5 = 0.8). The chance of a box under 1,230 g is 0.1587 whether you work with the total (mean 1,250, SD 20) or with the average (mean 50, SE 0.8, limit 49.2), and the simulation agrees at 0.157. In the growing sample, the SD hovers around 4.1 to 4.3 at every size, because it describes the cookies, while the SE falls from 0.820 to 0.418 to 0.213 as n goes from 25 to 100 to 400, halving each time the sample quadruples.
Common mix-ups
- Error bars. A chart with SE bars looks much tighter than one with SD bars. Always say which you plotted: SD bars show how variable the data are, SE bars (or better, confidence intervals) show how precise the mean is.
- "More data reduces the standard deviation." It does not. A larger sample gives a better estimate of the SD, but the SD itself reflects the population. Only the standard error shrinks with n.
- Unknown σ. In practice you plug in the sample SD, s, and get an estimated standard error s/sqrt(n). For confidence intervals and tests you then use the t distribution rather than the normal.
- Independence matters. The sqrt(n) rules assume independent values. Repeated measurements on the same person, or items from the same batch that share a common error, make the true standard error larger than the formula says.
For the next step, from a standard error to an interval, see our explanation of what 95% confidence means, or try the calculators on the DASS tools page.
How to report a standard error in APA style (7th edition)
Describe your sample with the mean and the standard deviation, and convey precision with a confidence interval. If you report the standard error itself, label it clearly. Using the first sample of 25 cookies from the example (SD 4.10, SE 0.82), with the mean left as a placeholder:
"Cookies weighed x.xx g on average (SD = 4.10, n = 25, SE = 0.82), 95% CI [LL, UL]."
In tables, give M and SD in separate columns. In figures, state in the caption whether error bars show ±1 SD, ±1 SE, or a 95% CI, because readers cannot tell them apart by eye.
Related tools and guides
- APA 7 formatter for descriptive statistics
- Which statistical test should I use?
- Why does the central limit theorem work?
- What does 95% confidence actually mean?
- Standard error (Wikipedia)
More answered questions
- Why does bootstrapping work? A plain-language explanation
- Where does the 1/sqrt(n) margin of error come from?
- Why does standard deviation square the differences?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.