Answer
How do you interpret Cohen's d, and are 0.2, 0.5 and 0.8 a real standard?
The short answer
Cohen's d is the difference between two group means divided by their pooled standard deviation. Cohen's labels of 0.2 (small), 0.5 (medium) and 0.8 (large) are rough conventions he offered for when nothing better is available, not a standard. Cutoffs of .10, .30 and .50 belong to the correlation r, not to d. Judge d against typical effects in your own field, translate it into overlap between the groups, and always report its confidence interval.
The short answer
Cohen's d expresses a difference between two means in standard deviation units: d = (M₁ − M₂) / sₚ, where sₚ is the pooled standard deviation. A d of 0.5 means the average person in one group sits half a standard deviation above the average person in the other. Because it has no units, d lets you compare results measured on different scales.
The familiar labels come from Jacob Cohen's 1988 book on power analysis: 0.2 small, 0.5 medium, 0.8 large. Cohen presented them as a fallback for planning studies when a researcher has no better information, and he warned that what counts as large depends on the area of research. Psychology adopted them widely, but no field has an official standard, including sociology and the other social sciences.
One common mix-up: the cutoffs .10, .30 and .50 are Cohen's benchmarks for the correlation r, not for d. The two scales are linked but not equal. For two groups of the same size, r = d / √(d² + 4), so a d of 0.2 corresponds to an r of about .10, but a d of 0.5 corresponds to an r of only about .24.
What a d value looks like in practice
Labels are easier to judge once you translate d into what it says about two overlapping groups. The code below does this for four values of d, assuming normal distributions with equal spread, and then estimates d from one simulated study and from 5,000 repeated studies.
R
set.seed(23775)
# 1. What Cohen's benchmarks mean in terms of overlap (two normal groups, equal SD)
d <- c(0.2, 0.5, 0.8, 1.2)
round(data.frame(d,
U3 = pnorm(d), # share of group 1 above the mean of group 2
overlap = 2 * pnorm(-d / 2), # overlap of the two distributions
PS = pnorm(d / sqrt(2)), # chance a random member of group 1 scores higher
r = d / sqrt(d^2 + 4)), 3) # point-biserial r (equal group sizes)
# 2. One study: two groups of 50, true difference of half an SD
treat <- rnorm(50, mean = 105, sd = 10)
control <- rnorm(50, mean = 100, sd = 10)
n1 <- length(treat); n2 <- length(control)
sp <- sqrt(((n1 - 1) * var(treat) + (n2 - 1) * var(control)) / (n1 + n2 - 2))
d_hat <- (mean(treat) - mean(control)) / sp
g_hat <- d_hat * (1 - 3 / (4 * (n1 + n2) - 9)) # Hedges' small-sample correction
se_d <- sqrt((n1 + n2) / (n1 * n2) + d_hat^2 / (2 * (n1 + n2)))
round(c(M1 = mean(treat), SD1 = sd(treat), M2 = mean(control), SD2 = sd(control)), 2)
tt <- t.test(treat, control, var.equal = TRUE)
round(c(t = unname(tt$statistic), df = unname(tt$parameter), p = tt$p.value), 3)
round(c(d = d_hat, g = g_hat, lower = d_hat - 1.96 * se_d, upper = d_hat + 1.96 * se_d), 2)
# 3. How much d varies from study to study (true d = 0.5, 20 per group, 5,000 studies)
d_sim <- replicate(5000, {
x <- rnorm(20, 0.5); y <- rnorm(20, 0)
(mean(x) - mean(y)) / sqrt((var(x) + var(y)) / 2)
})
round(quantile(d_sim, c(0.025, 0.25, 0.5, 0.75, 0.975)), 2)
round(c(below_0.2 = mean(d_sim < 0.2), above_0.8 = mean(d_sim > 0.8)), 3)Python
import numpy as np
import pandas as pd
from scipy import stats
rng = np.random.default_rng(23775)
# 1. What Cohen's benchmarks mean in terms of overlap (two normal groups, equal SD)
d = np.array([0.2, 0.5, 0.8, 1.2])
print(pd.DataFrame({
"d": d,
"U3": stats.norm.cdf(d), # share of group 1 above the mean of group 2
"overlap": 2 * stats.norm.cdf(-d / 2), # overlap of the two distributions
"PS": stats.norm.cdf(d / np.sqrt(2)), # chance a random member of group 1 scores higher
"r": d / np.sqrt(d**2 + 4), # point-biserial r (equal group sizes)
}).round(3))
# 2. One study: two groups of 50, true difference of half an SD
treat = rng.normal(105, 10, 50)
control = rng.normal(100, 10, 50)
n1, n2 = len(treat), len(control)
sp = np.sqrt(((n1 - 1) * treat.var(ddof=1) + (n2 - 1) * control.var(ddof=1)) / (n1 + n2 - 2))
d_hat = (treat.mean() - control.mean()) / sp
g_hat = d_hat * (1 - 3 / (4 * (n1 + n2) - 9)) # Hedges' small-sample correction
se_d = np.sqrt((n1 + n2) / (n1 * n2) + d_hat**2 / (2 * (n1 + n2)))
print("M1", round(treat.mean(), 2), "SD1", round(treat.std(ddof=1), 2),
"M2", round(control.mean(), 2), "SD2", round(control.std(ddof=1), 2))
tt = stats.ttest_ind(treat, control, equal_var=True)
print("t", round(tt.statistic, 3), "df", tt.df, "p", round(tt.pvalue, 3))
print("d", round(d_hat, 2), "g", round(g_hat, 2),
"lower", round(d_hat - 1.96 * se_d, 2), "upper", round(d_hat + 1.96 * se_d, 2))
# 3. How much d varies from study to study (true d = 0.5, 20 per group, 5,000 studies)
x = rng.normal(0.5, 1, size=(5000, 20))
y = rng.normal(0.0, 1, size=(5000, 20))
d_sim = (x.mean(axis=1) - y.mean(axis=1)) / np.sqrt((x.var(axis=1, ddof=1) + y.var(axis=1, ddof=1)) / 2)
print(np.round(np.quantile(d_sim, [0.025, 0.25, 0.5, 0.75, 0.975]), 2))
print("below_0.2", round(np.mean(d_sim < 0.2), 3), "above_0.8", round(np.mean(d_sim > 0.8), 3))The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly. Python's random numbers differ from R's, so its simulated figures differ slightly (for example, 17.2% of simulated studies below 0.2 instead of 16.2%), but they show the same pattern.
- **A "small" d of 0.2.** The two distributions overlap by 92%, 57.9% of one group scores above the other group's mean, and a randomly chosen person from the higher group beats a randomly chosen person from the lower group 55.6% of the time, barely better than a coin flip.
- **A "medium" d of 0.5.** Overlap is 80.3%, 69.1% of one group is above the other's mean, and the probability of superiority is 63.8%.
- **A "large" d of 0.8.** Overlap is still 68.9%. 78.8% of one group is above the other's mean, and the probability of superiority is 71.4%. Even a large effect leaves many people in the lower group scoring above many in the higher one.
- One study. With 50 per group, the treatment group had M = 106.66 (SD = 8.58) and the control group M = 102.04 (SD = 9.40). That gives d = 0.51 with an approximate 95% confidence interval from 0.12 to 0.91, so the data are consistent with anything from a small effect to a large one. Hedges' g, which corrects a slight upward bias in d, also rounds to 0.51 at this sample size.
- Many studies. When the true d is exactly 0.5 and each study has 20 people per group, the middle 95% of estimates runs from −0.11 to 1.17. In 16.2% of studies the estimate fell below 0.2 and in 18.8% it exceeded 0.8. A single small study's label (small, medium or large) is often wrong.
Why fields use different benchmarks
What counts as a meaningful effect depends on what is being measured, how hard it is to change, and what it costs. A few examples show how far typical effects differ from Cohen's labels:
- Individual differences research. Gignac and Szodorai (2016) looked at hundreds of published correlations and found that the typical r was about .19. They suggested treating .10, .20 and .30 as relatively small, typical and relatively large, well below Cohen's .10, .30 and .50.
- Education interventions. Kraft (2020) argued that for rigorous studies of school programs measured on standardized achievement tests, effects of 0.20 standard deviations are already large and effects of 0.05 are not trivial, because such outcomes are hard to move.
- Sociology and much of social science. Results are more often reported as regression coefficients, odds ratios or differences in percentage points than as d, and there is no agreed set of cutoffs. The usual advice is to compare an effect with published estimates on similar outcomes and to state it in the outcome's own units as well.
So the question to ask is not "is this small or large by Cohen's rules?" but "how does this compare with other effects on this outcome, and would a difference this size matter to anyone?" A d of 0.2 on mortality or on reading scores across a whole school district can be very important, while a d of 0.8 on a lab task may be of little practical use.
Practical tips for reading and using d
- Check the standardiser. d for independent groups usually divides by the pooled SD. Some authors divide by the control group's SD (Glass's delta), and repeated-measures designs can divide by the SD of the difference scores, which can give a much larger number for the same data. Say which one you used.
- Report the interval, not just the point estimate. As the simulation shows, d from a small study is noisy. The interval tells readers how far to trust the label.
- **Use Hedges' g in small samples.** The correction matters with a few dozen participants or fewer; with 100 in total it barely changes the value.
- Plan sample size from a realistic effect. Powering a study for a "medium" 0.5 when effects in your field are typically 0.2 leaves it badly underpowered. See Does a t test need a minimum sample size?.
Not sure which test and effect size fit your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report Cohen's d in APA style (7th edition)
Give d alongside the test it goes with, with two decimals and a confidence interval. Keep the leading zero, because d can be larger than 1. Using the example above:
"The treatment group scored higher (M = 106.66, SD = 8.58, n = 50) than the control group (M = 102.04, SD = 9.40, n = 50), t(98) = 2.57, p = .012, d = 0.51, 95% CI [0.12, 0.91]."
If you describe the size in words, say what you compared it with, for example "a medium effect by Cohen's (1988) conventions, and larger than typical effects reported for similar interventions." In the method section, state how d was computed: "Effect sizes are Cohen's d using the pooled standard deviation, with approximate 95% confidence intervals."
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Which statistical test should I use?
- Welch vs Student t test: should you just always use Welch?
- Does a t test need a minimum sample size?
- What do p values and t values actually mean?
- Effect size (Wikipedia)
More answered questions
- Does a t test need a minimum sample size?
- Welch vs Student t test: should you just always use Welch?
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.