Answer

How do you compare a short ordinal rating scale between two groups?

Inspired by a question on Cross Validated ·

nonparametrict testchi-square

The short answer

For one item with only a few ordered levels, the Mann-Whitney test is the natural default: it uses the order but not the spacing of the levels. A t test on the codes usually gives a very similar answer, but it assumes the levels are equally spaced. A chi-square test throws the order away and loses power. Proportional-odds (ordinal logistic) regression is the model-based choice when you need covariates.

The short answer

Suppose each person answers one item scored 1 to 4 (say, low, fairly low, fairly high, high), and you want to know whether two independent groups differ. Three tests are commonly suggested, and they answer slightly different questions:

If you need to adjust for other variables, or want an effect expressed as an odds ratio, fit a proportional-odds (ordinal logistic) regression instead, for example with MASS::polr() or the ordinal package in R.

Why the choice matters less than people fear, and where it does matter

Coding the levels as 1, 2, 3, 4 is a decision about spacing: it says the step from low to fairly low is as large as the step from fairly high to high. A rank-based test makes no such claim. In practice, when two groups differ by a general upward or downward shift, the t test and Mann-Whitney almost always reach the same conclusion, and the t test's false-positive rate stays close to 5% even though the data are far from normal. The simulation below shows both.

Where the choice matters is interpretation. Mann-Whitney is not a test of medians in general. With only four levels many people share the same score, and two groups can have the same median while one still tends to score higher (or different medians with no significant difference). The honest summary of a Mann-Whitney result is the probability of superiority: the chance that a person from group B scores higher than a person from group A, counting ties as half. The rank-biserial correlation is the same information rescaled to run from -1 to 1.

The chi-square test is the right tool only when the order is irrelevant to your question, for example if you expect one group to cluster at both extremes rather than shift up. For an ordered scale where "higher is better", it wastes information.

See it in R and Python

The code builds a small fixed data set (60 people per group), runs all three tests on it, computes the probability of superiority, and then simulates 2,000 studies with 40 people per group to compare how often each test rejects, first when the groups do not differ and then when group B is shifted upward.

R

set.seed(675711)

# 1. One 4-point item (1 = low ... 4 = high), 60 people per group
a <- rep(1:4, times = c(10, 20, 20, 10))   # group A
b <- rep(1:4, times = c(5, 14, 25, 16))    # group B
table(group = rep(c("A", "B"), each = 60), score = c(a, b))
round(c(mean_A = mean(a), mean_B = mean(b), median_A = median(a), median_B = median(b)), 2)

# Three ways to compare the groups
tt <- t.test(b, a)                          # Welch t test on the 1-4 codes
mw <- wilcox.test(b, a)                     # Mann-Whitney (normal approx., tie-corrected)
cs <- chisq.test(table(rep(c("A", "B"), each = 60), c(a, b)))  # ignores the order
round(c(t = unname(tt$statistic), df = unname(tt$parameter), p = tt$p.value), 4)
round(tt$conf.int, 2)
c(W = unname(mw$statistic), p = round(mw$p.value, 4))
round(c(X2 = unname(cs$statistic), df = unname(cs$parameter), p = cs$p.value), 4)

# Effect size: P(B > A) + 0.5 * P(tie), and the rank-biserial correlation
ps <- unname(mw$statistic) / (60 * 60)
round(c(prob_superiority = ps, rank_biserial = 2 * ps - 1), 3)

# 2. How often does each test reject at 5%? (40 per group, 2,000 samples)
pA <- c(10, 20, 20, 10) / 60
pB <- c(5, 14, 25, 16) / 60
reject <- function(p1, p2, reps = 2000, n = 40) {
  res <- replicate(reps, {
    x <- sample(1:4, n, TRUE, p1); y <- sample(1:4, n, TRUE, p2)
    g <- rep(1:2, each = n); s <- factor(c(x, y), levels = 1:4)
    tab <- table(g, s); tab <- tab[, colSums(tab) > 0, drop = FALSE]
    c(t_test = t.test(y, x)$p.value < .05,
      mann_whitney = wilcox.test(y, x, exact = FALSE)$p.value < .05,
      chi_square = suppressWarnings(chisq.test(tab)$p.value) < .05)
  })
  rowMeans(res)
}
round(reject(pA, pA), 3)   # no difference: false-positive rate
round(reject(pA, pB), 3)   # B shifted upward: power

Python

import numpy as np
import pandas as pd
from scipy import stats

rng = np.random.default_rng(675711)

# 1. One 4-point item (1 = low ... 4 = high), 60 people per group
a = np.repeat([1, 2, 3, 4], [10, 20, 20, 10])   # group A
b = np.repeat([1, 2, 3, 4], [5, 14, 25, 16])    # group B
tab = pd.crosstab(np.repeat(["A", "B"], 60), np.concatenate([a, b]))
print(tab)
print("means", round(a.mean(), 2), round(b.mean(), 2), "medians", np.median(a), np.median(b))

# Three ways to compare the groups
tt = stats.ttest_ind(b, a, equal_var=False)                  # Welch t test on the 1-4 codes
ci = tt.confidence_interval()
mw = stats.mannwhitneyu(b, a, method="asymptotic")           # tie-corrected, continuity-corrected
cs = stats.chi2_contingency(tab.values)                      # ignores the order
print("t", round(tt.statistic, 4), "df", round(tt.df, 4), "p", round(tt.pvalue, 4),
      "CI", round(ci.low, 2), round(ci.high, 2))
print("W", mw.statistic, "p", round(mw.pvalue, 4))
print("X2", round(cs.statistic, 4), "df", cs.dof, "p", round(cs.pvalue, 4))

# Effect size: P(B > A) + 0.5 * P(tie), and the rank-biserial correlation
ps = mw.statistic / (60 * 60)
print("prob_superiority", round(ps, 3), "rank_biserial", round(2 * ps - 1, 3))

# 2. How often does each test reject at 5%? (40 per group, 2,000 samples)
pA = np.array([10, 20, 20, 10]) / 60
pB = np.array([5, 14, 25, 16]) / 60
def reject(p1, p2, reps=2000, n=40):
    hits = np.zeros(3)
    for _ in range(reps):
        x = rng.choice([1, 2, 3, 4], n, p=p1); y = rng.choice([1, 2, 3, 4], n, p=p2)
        t = np.array([[np.sum(v == k) for k in range(1, 5)] for v in (x, y)])
        t = t[:, t.sum(axis=0) > 0]
        hits += [stats.ttest_ind(y, x, equal_var=False).pvalue < .05,
                 stats.mannwhitneyu(y, x, method="asymptotic").pvalue < .05,
                 stats.chi2_contingency(t)[1] < .05]
    return dict(zip(["t_test", "mann_whitney", "chi_square"], np.round(hits / reps, 3)))

print(reject(pA, pA))   # no difference: false-positive rate
print(reject(pA, pB))   # B shifted upward: power

The figures below come from one seeded run of the R code. The tests on the fixed data (part 1) involve no randomness, and the Python version reproduces them exactly. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 38.7% instead of 40.7% for the t test's power), but they show the same pattern.

What to do in practice

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report a Mann-Whitney U test in APA style (7th edition)

Give the medians (and ideally the distribution across levels), the test statistic, the exact p value and an effect size. Using the example above:

"Ratings were higher in group B (Mdn = 3, n = 60) than in group A (Mdn = 2.5, n = 60). A Mann-Whitney U test showed that this difference was statistically significant, U = 2185, p = .035, rank-biserial r = .21. A randomly chosen participant in group B rated higher than one in group A 61% of the time."

R labels the statistic W, which equals U for the first group named in wilcox.test(); other software may report the smaller U or a z value, so name the statistic you used. In the method section, state that the item was treated as ordinal, that a normal approximation with a tie correction was used, and, if you also ran a t test as a check, that it led to the same conclusion.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.