Answer

Should you check normality of the outcome variable or of the residuals?

Inspired by a question on Cross Validated ·

normalityregression

The short answer

Check the residuals. Linear models (regression, t tests, ANOVA) assume that the errors, the scatter of Y around its predicted value, are normal. They assume nothing about the overall shape of Y. A skewed predictor can make Y badly skewed while the residuals are perfectly normal, and skewed errors can hide inside a Y that looks normal. Testing Y can therefore send you the wrong way in either direction.

The short answer

Every linear model can be written as Y = (the part explained by the predictors) + error. The normality assumption is a statement about that error term: for any fixed set of predictor values, the outcome is scattered normally around its predicted value. In other words, it is the conditional distribution of Y given X that should be normal, not the marginal distribution of Y that you get by throwing all the outcome values into one histogram.

You never observe the errors directly, but the residuals (observed minus fitted values) are your estimates of them. So a check of the normality assumption is a check of the residuals. A histogram or normality test of the raw outcome answers a different question, and its answer can point the wrong way.

Why the outcome's shape tells you so little

The overall distribution of Y is a blend of two ingredients: the distribution of the predictors (passed through the regression coefficients) and the distribution of the errors. Either ingredient can dominate, which gives two ways to be misled:

The same logic covers group comparisons. In a t test or ANOVA the predictor is group membership, and a real difference between groups shifts their centers apart, so the pooled outcome can look lumpy or bimodal while each group is perfectly normal. That case is worked through on the page t test vs ANOVA with two groups; this page focuses on regression with a continuous predictor, where the problem is easier to miss.

See it in R and Python

The code builds two data sets of 200 cases from a known model, Y = 20 + 3X + error. In case A the predictor is exponential (strongly right-skewed) and the errors are normal, so the assumption holds. In case B the predictor is normal and the errors are exponential, shifted to have mean 0, so the assumption fails. For each case it reports the skewness and the Shapiro-Wilk p value of the outcome and of the residuals, then repeats both designs 1,000 times.

R

set.seed(60410)
n <- 200
skew <- function(v) mean((v - mean(v))^3) / mean((v - mean(v))^2)^1.5

# Case A: skewed predictor (e.g. weekly study hours), perfectly normal errors
x_a <- rexp(n, rate = 1 / 10)
y_a <- 20 + 3 * x_a + rnorm(n, 0, 5)
fit_a <- lm(y_a ~ x_a)

# Case B: normal predictor, right-skewed errors (exponential, shifted to mean 0)
x_b <- rnorm(n, 10, 5)
y_b <- 20 + 3 * x_b + (rexp(n, rate = 1 / 5) - 5)
fit_b <- lm(y_b ~ x_b)

check <- function(y, fit) c(
  skew_y = skew(y), skew_resid = skew(residuals(fit)),
  sw_p_y = shapiro.test(y)$p.value, sw_p_resid = shapiro.test(residuals(fit))$p.value)
signif(rbind(A = check(y_a, fit_a), B = check(y_b, fit_b)), 3)

# Case A regression results (the assumptions really do hold here)
s <- summary(fit_a)
round(s$coefficients, 3)
round(confint(fit_a)["x_a", ], 2)
round(c(R2 = s$r.squared, F = unname(s$fstatistic[1]), df2 = s$df[2]), 3)
round(unname(shapiro.test(residuals(fit_a))$statistic), 3)

# Repeat both cases 1,000 times: how often does Shapiro-Wilk reject at .05?
reject <- function(case, reps = 1000) {
  r <- replicate(reps, {
    if (case == "A") { x <- rexp(n, 1 / 10); e <- rnorm(n, 0, 5) }
    else             { x <- rnorm(n, 10, 5); e <- rexp(n, 1 / 5) - 5 }
    y <- 20 + 3 * x + e
    c(outcome = shapiro.test(y)$p.value < .05,
      residuals = shapiro.test(residuals(lm(y ~ x)))$p.value < .05)
  })
  rowMeans(r)
}
rbind(A = reject("A"), B = reject("B"))

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(60410)
n = 200

def resid(x, y):
    slope, intercept = np.polyfit(x, y, 1)
    return y - (intercept + slope * x)

# Case A: skewed predictor (e.g. weekly study hours), perfectly normal errors
x_a = rng.exponential(10, n)
y_a = 20 + 3 * x_a + rng.normal(0, 5, n)

# Case B: normal predictor, right-skewed errors (exponential, shifted to mean 0)
x_b = rng.normal(10, 5, n)
y_b = 20 + 3 * x_b + (rng.exponential(5, n) - 5)

for name, x, y in [("A", x_a, y_a), ("B", x_b, y_b)]:
    r = resid(x, y)
    print(name, "skew_y", round(stats.skew(y), 3), "skew_resid", round(stats.skew(r), 3),
          "sw_p_y", f"{stats.shapiro(y).pvalue:.3g}", "sw_p_resid", f"{stats.shapiro(r).pvalue:.3g}")

# Case A regression results
fit = stats.linregress(x_a, y_a)
tcrit = stats.t.ppf(0.975, n - 2)
print("slope", round(fit.slope, 3), "se", round(fit.stderr, 3), "t", round(fit.slope / fit.stderr, 3),
      "CI", np.round([fit.slope - tcrit * fit.stderr, fit.slope + tcrit * fit.stderr], 2),
      "R2", round(fit.rvalue**2, 3), "W", round(stats.shapiro(resid(x_a, y_a)).statistic, 3))

# Repeat both cases 1,000 times: how often does Shapiro-Wilk reject at .05?
def reject(case, reps=1000):
    out = np.zeros(2)
    for _ in range(reps):
        if case == "A":
            x, e = rng.exponential(10, n), rng.normal(0, 5, n)
        else:
            x, e = rng.normal(10, 5, n), rng.exponential(5, n) - 5
        y = 20 + 3 * x + e
        out += [stats.shapiro(y).pvalue < .05, stats.shapiro(resid(x, y)).pvalue < .05]
    return {"outcome": out[0] / reps, "residuals": out[1] / reps}

print("A", reject("A"))
print("B", reject("B"))

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its samples and rates differ slightly, but they show exactly the same pattern.

Even though the outcome in case A is heavily skewed, the regression recovers the true slope well: b = 3.01, SE = 0.04, 95% CI [2.94, 3.08], with R² = .97.

Groups: one residual plot or one per group?

In a t test or one-way ANOVA, the residuals are simply each score minus its own group mean, so a single Q-Q plot of all residuals is the natural first look. It works well when the groups share the same shape and spread, which is what the model assumes anyway.

When the groups may differ in shape, though, the pooled residuals can hide it. If one group is strongly right-skewed and another is close to normal, the combined residuals show a diluted, milder skew. So look at the residuals within each group too, especially when the groups are unequal in size or spread. If a group is genuinely skewed, a model that matches the data (a transformation that makes sense for the variable, a generalized linear model, or a rank-based or bootstrap comparison) is usually a better fix than hoping the pooled plot looks acceptable. Note also that groups with very different distributions often differ in variance as well, which is a bigger threat to a pooled t test than non-normality; see Welch vs Student t test.

How much does normality matter, and how to check it

Among the assumptions of a linear model, normality of the errors is usually the least important for estimates and tests. Independence of observations, a correctly specified mean (for example, a straight line when the relationship is straight) and roughly constant error variance matter more. The coefficients are unbiased whatever the error distribution, and with moderate or large samples the central limit theorem makes their tests and confidence intervals approximately right. Normality matters most in small samples, when there are extreme outliers, and for prediction intervals, which describe individual values rather than averages.

  1. Fit the model first, then save the residuals. In R, residuals(fit) or rstandard(fit); plot(fit) gives the standard diagnostic plots.
  2. Look at a Q-Q plot of the residuals. It shows the size and shape of any departure, which a p value cannot. See How to interpret a Q-Q plot.
  3. Treat a normality test as a supporting detail, not a gate. With large samples it flags trivial departures, and with small samples it misses large ones; see Is a normality test worth running on your data?.
  4. Do not transform the outcome just because its histogram is skewed. Transform when the residuals, or the meaning of the variable, call for it; see When to log transform data.

For more practical guides on checking a data set before analysis, see the DASS blog.

How to report normality of residuals in APA style (7th edition)

Assumption checks usually go in the method or results section in a sentence or two, just before the main result. With the case A example above:

"A Q-Q plot of the residuals showed no marked departure from normality, and a Shapiro-Wilk test of the residuals was not significant, W = .99, p = .163. The predictor was strongly related to the outcome, b = 3.01, SE = 0.04, 95% CI [2.94, 3.08], t(198) = 82.30, p < .001, R² = .97."

Say which variable you checked (the residuals, not the raw outcome), and how. If you relied on a plot alone, a sentence such as "Normality of the residuals was assessed with a Q-Q plot, which showed an acceptable fit" is enough. Report p values below .001 as p < .001.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.