Answer

How do you interpret Cohen's d, and are 0.2, 0.5 and 0.8 a real standard?

Inspired by a question on Cross Validated ·

effect sizet test

The short answer

Cohen's d is the difference between two group means divided by their pooled standard deviation. Cohen's labels of 0.2 (small), 0.5 (medium) and 0.8 (large) are rough conventions he offered for when nothing better is available, not a standard. Cutoffs of .10, .30 and .50 belong to the correlation r, not to d. Judge d against typical effects in your own field, translate it into overlap between the groups, and always report its confidence interval.

The short answer

Cohen's d expresses a difference between two means in standard deviation units: d = (M₁ − M₂) / sₚ, where sₚ is the pooled standard deviation. A d of 0.5 means the average person in one group sits half a standard deviation above the average person in the other. Because it has no units, d lets you compare results measured on different scales.

The familiar labels come from Jacob Cohen's 1988 book on power analysis: 0.2 small, 0.5 medium, 0.8 large. Cohen presented them as a fallback for planning studies when a researcher has no better information, and he warned that what counts as large depends on the area of research. Psychology adopted them widely, but no field has an official standard, including sociology and the other social sciences.

One common mix-up: the cutoffs .10, .30 and .50 are Cohen's benchmarks for the correlation r, not for d. The two scales are linked but not equal. For two groups of the same size, r = d / √(d² + 4), so a d of 0.2 corresponds to an r of about .10, but a d of 0.5 corresponds to an r of only about .24.

What a d value looks like in practice

Labels are easier to judge once you translate d into what it says about two overlapping groups. The code below does this for four values of d, assuming normal distributions with equal spread, and then estimates d from one simulated study and from 5,000 repeated studies.

R

set.seed(23775)

# 1. What Cohen's benchmarks mean in terms of overlap (two normal groups, equal SD)
d <- c(0.2, 0.5, 0.8, 1.2)
round(data.frame(d,
  U3      = pnorm(d),              # share of group 1 above the mean of group 2
  overlap = 2 * pnorm(-d / 2),     # overlap of the two distributions
  PS      = pnorm(d / sqrt(2)),    # chance a random member of group 1 scores higher
  r       = d / sqrt(d^2 + 4)), 3) # point-biserial r (equal group sizes)

# 2. One study: two groups of 50, true difference of half an SD
treat   <- rnorm(50, mean = 105, sd = 10)
control <- rnorm(50, mean = 100, sd = 10)
n1 <- length(treat); n2 <- length(control)
sp <- sqrt(((n1 - 1) * var(treat) + (n2 - 1) * var(control)) / (n1 + n2 - 2))
d_hat <- (mean(treat) - mean(control)) / sp
g_hat <- d_hat * (1 - 3 / (4 * (n1 + n2) - 9))          # Hedges' small-sample correction
se_d  <- sqrt((n1 + n2) / (n1 * n2) + d_hat^2 / (2 * (n1 + n2)))
round(c(M1 = mean(treat), SD1 = sd(treat), M2 = mean(control), SD2 = sd(control)), 2)
tt <- t.test(treat, control, var.equal = TRUE)
round(c(t = unname(tt$statistic), df = unname(tt$parameter), p = tt$p.value), 3)
round(c(d = d_hat, g = g_hat, lower = d_hat - 1.96 * se_d, upper = d_hat + 1.96 * se_d), 2)

# 3. How much d varies from study to study (true d = 0.5, 20 per group, 5,000 studies)
d_sim <- replicate(5000, {
  x <- rnorm(20, 0.5); y <- rnorm(20, 0)
  (mean(x) - mean(y)) / sqrt((var(x) + var(y)) / 2)
})
round(quantile(d_sim, c(0.025, 0.25, 0.5, 0.75, 0.975)), 2)
round(c(below_0.2 = mean(d_sim < 0.2), above_0.8 = mean(d_sim > 0.8)), 3)

Python

import numpy as np
import pandas as pd
from scipy import stats

rng = np.random.default_rng(23775)

# 1. What Cohen's benchmarks mean in terms of overlap (two normal groups, equal SD)
d = np.array([0.2, 0.5, 0.8, 1.2])
print(pd.DataFrame({
    "d": d,
    "U3": stats.norm.cdf(d),                # share of group 1 above the mean of group 2
    "overlap": 2 * stats.norm.cdf(-d / 2),  # overlap of the two distributions
    "PS": stats.norm.cdf(d / np.sqrt(2)),   # chance a random member of group 1 scores higher
    "r": d / np.sqrt(d**2 + 4),             # point-biserial r (equal group sizes)
}).round(3))

# 2. One study: two groups of 50, true difference of half an SD
treat = rng.normal(105, 10, 50)
control = rng.normal(100, 10, 50)
n1, n2 = len(treat), len(control)
sp = np.sqrt(((n1 - 1) * treat.var(ddof=1) + (n2 - 1) * control.var(ddof=1)) / (n1 + n2 - 2))
d_hat = (treat.mean() - control.mean()) / sp
g_hat = d_hat * (1 - 3 / (4 * (n1 + n2) - 9))   # Hedges' small-sample correction
se_d = np.sqrt((n1 + n2) / (n1 * n2) + d_hat**2 / (2 * (n1 + n2)))
print("M1", round(treat.mean(), 2), "SD1", round(treat.std(ddof=1), 2),
      "M2", round(control.mean(), 2), "SD2", round(control.std(ddof=1), 2))
tt = stats.ttest_ind(treat, control, equal_var=True)
print("t", round(tt.statistic, 3), "df", tt.df, "p", round(tt.pvalue, 3))
print("d", round(d_hat, 2), "g", round(g_hat, 2),
      "lower", round(d_hat - 1.96 * se_d, 2), "upper", round(d_hat + 1.96 * se_d, 2))

# 3. How much d varies from study to study (true d = 0.5, 20 per group, 5,000 studies)
x = rng.normal(0.5, 1, size=(5000, 20))
y = rng.normal(0.0, 1, size=(5000, 20))
d_sim = (x.mean(axis=1) - y.mean(axis=1)) / np.sqrt((x.var(axis=1, ddof=1) + y.var(axis=1, ddof=1)) / 2)
print(np.round(np.quantile(d_sim, [0.025, 0.25, 0.5, 0.75, 0.975]), 2))
print("below_0.2", round(np.mean(d_sim < 0.2), 3), "above_0.8", round(np.mean(d_sim > 0.8), 3))

The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly. Python's random numbers differ from R's, so its simulated figures differ slightly (for example, 17.2% of simulated studies below 0.2 instead of 16.2%), but they show the same pattern.

Why fields use different benchmarks

What counts as a meaningful effect depends on what is being measured, how hard it is to change, and what it costs. A few examples show how far typical effects differ from Cohen's labels:

So the question to ask is not "is this small or large by Cohen's rules?" but "how does this compare with other effects on this outcome, and would a difference this size matter to anyone?" A d of 0.2 on mortality or on reading scores across a whole school district can be very important, while a d of 0.8 on a lab task may be of little practical use.

Practical tips for reading and using d

Not sure which test and effect size fit your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report Cohen's d in APA style (7th edition)

Give d alongside the test it goes with, with two decimals and a confidence interval. Keep the leading zero, because d can be larger than 1. Using the example above:

"The treatment group scored higher (M = 106.66, SD = 8.58, n = 50) than the control group (M = 102.04, SD = 9.40, n = 50), t(98) = 2.57, p = .012, d = 0.51, 95% CI [0.12, 0.91]."

If you describe the size in words, say what you compared it with, for example "a medium effect by Cohen's (1988) conventions, and larger than typical effects reported for similar interventions." In the method section, state how d was computed: "Effect sizes are Cohen's d using the pooled standard deviation, with approximate 95% confidence intervals."

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.