Answer
How do you compare a short ordinal rating scale between two groups?
The short answer
For one item with only a few ordered levels, the Mann-Whitney test is the natural default: it uses the order but not the spacing of the levels. A t test on the codes usually gives a very similar answer, but it assumes the levels are equally spaced. A chi-square test throws the order away and loses power. Proportional-odds (ordinal logistic) regression is the model-based choice when you need covariates.
The short answer
Suppose each person answers one item scored 1 to 4 (say, low, fairly low, fairly high, high), and you want to know whether two independent groups differ. Three tests are commonly suggested, and they answer slightly different questions:
- Mann-Whitney (Wilcoxon rank-sum) test. Uses only the order of the levels. It asks whether a randomly chosen person from one group tends to score higher than a randomly chosen person from the other. This is usually the best default for a single short ordinal item.
- **Welch t test on the codes 1-4.** Treats the codes as numbers with equal gaps and compares means. With moderate samples it is robust and usually agrees with Mann-Whitney, but a mean of 2.87 on a scale whose gaps may not be equal is harder to defend.
- Chi-square test of independence. Treats the four levels as unordered categories. It will notice any difference in the pattern of answers, but because it ignores the order it has less power to detect the typical case where one group simply scores higher.
If you need to adjust for other variables, or want an effect expressed as an odds ratio, fit a proportional-odds (ordinal logistic) regression instead, for example with MASS::polr() or the ordinal package in R.
Why the choice matters less than people fear, and where it does matter
Coding the levels as 1, 2, 3, 4 is a decision about spacing: it says the step from low to fairly low is as large as the step from fairly high to high. A rank-based test makes no such claim. In practice, when two groups differ by a general upward or downward shift, the t test and Mann-Whitney almost always reach the same conclusion, and the t test's false-positive rate stays close to 5% even though the data are far from normal. The simulation below shows both.
Where the choice matters is interpretation. Mann-Whitney is not a test of medians in general. With only four levels many people share the same score, and two groups can have the same median while one still tends to score higher (or different medians with no significant difference). The honest summary of a Mann-Whitney result is the probability of superiority: the chance that a person from group B scores higher than a person from group A, counting ties as half. The rank-biserial correlation is the same information rescaled to run from -1 to 1.
The chi-square test is the right tool only when the order is irrelevant to your question, for example if you expect one group to cluster at both extremes rather than shift up. For an ordered scale where "higher is better", it wastes information.
See it in R and Python
The code builds a small fixed data set (60 people per group), runs all three tests on it, computes the probability of superiority, and then simulates 2,000 studies with 40 people per group to compare how often each test rejects, first when the groups do not differ and then when group B is shifted upward.
R
set.seed(675711)
# 1. One 4-point item (1 = low ... 4 = high), 60 people per group
a <- rep(1:4, times = c(10, 20, 20, 10)) # group A
b <- rep(1:4, times = c(5, 14, 25, 16)) # group B
table(group = rep(c("A", "B"), each = 60), score = c(a, b))
round(c(mean_A = mean(a), mean_B = mean(b), median_A = median(a), median_B = median(b)), 2)
# Three ways to compare the groups
tt <- t.test(b, a) # Welch t test on the 1-4 codes
mw <- wilcox.test(b, a) # Mann-Whitney (normal approx., tie-corrected)
cs <- chisq.test(table(rep(c("A", "B"), each = 60), c(a, b))) # ignores the order
round(c(t = unname(tt$statistic), df = unname(tt$parameter), p = tt$p.value), 4)
round(tt$conf.int, 2)
c(W = unname(mw$statistic), p = round(mw$p.value, 4))
round(c(X2 = unname(cs$statistic), df = unname(cs$parameter), p = cs$p.value), 4)
# Effect size: P(B > A) + 0.5 * P(tie), and the rank-biserial correlation
ps <- unname(mw$statistic) / (60 * 60)
round(c(prob_superiority = ps, rank_biserial = 2 * ps - 1), 3)
# 2. How often does each test reject at 5%? (40 per group, 2,000 samples)
pA <- c(10, 20, 20, 10) / 60
pB <- c(5, 14, 25, 16) / 60
reject <- function(p1, p2, reps = 2000, n = 40) {
res <- replicate(reps, {
x <- sample(1:4, n, TRUE, p1); y <- sample(1:4, n, TRUE, p2)
g <- rep(1:2, each = n); s <- factor(c(x, y), levels = 1:4)
tab <- table(g, s); tab <- tab[, colSums(tab) > 0, drop = FALSE]
c(t_test = t.test(y, x)$p.value < .05,
mann_whitney = wilcox.test(y, x, exact = FALSE)$p.value < .05,
chi_square = suppressWarnings(chisq.test(tab)$p.value) < .05)
})
rowMeans(res)
}
round(reject(pA, pA), 3) # no difference: false-positive rate
round(reject(pA, pB), 3) # B shifted upward: powerPython
import numpy as np
import pandas as pd
from scipy import stats
rng = np.random.default_rng(675711)
# 1. One 4-point item (1 = low ... 4 = high), 60 people per group
a = np.repeat([1, 2, 3, 4], [10, 20, 20, 10]) # group A
b = np.repeat([1, 2, 3, 4], [5, 14, 25, 16]) # group B
tab = pd.crosstab(np.repeat(["A", "B"], 60), np.concatenate([a, b]))
print(tab)
print("means", round(a.mean(), 2), round(b.mean(), 2), "medians", np.median(a), np.median(b))
# Three ways to compare the groups
tt = stats.ttest_ind(b, a, equal_var=False) # Welch t test on the 1-4 codes
ci = tt.confidence_interval()
mw = stats.mannwhitneyu(b, a, method="asymptotic") # tie-corrected, continuity-corrected
cs = stats.chi2_contingency(tab.values) # ignores the order
print("t", round(tt.statistic, 4), "df", round(tt.df, 4), "p", round(tt.pvalue, 4),
"CI", round(ci.low, 2), round(ci.high, 2))
print("W", mw.statistic, "p", round(mw.pvalue, 4))
print("X2", round(cs.statistic, 4), "df", cs.dof, "p", round(cs.pvalue, 4))
# Effect size: P(B > A) + 0.5 * P(tie), and the rank-biserial correlation
ps = mw.statistic / (60 * 60)
print("prob_superiority", round(ps, 3), "rank_biserial", round(2 * ps - 1, 3))
# 2. How often does each test reject at 5%? (40 per group, 2,000 samples)
pA = np.array([10, 20, 20, 10]) / 60
pB = np.array([5, 14, 25, 16]) / 60
def reject(p1, p2, reps=2000, n=40):
hits = np.zeros(3)
for _ in range(reps):
x = rng.choice([1, 2, 3, 4], n, p=p1); y = rng.choice([1, 2, 3, 4], n, p=p2)
t = np.array([[np.sum(v == k) for k in range(1, 5)] for v in (x, y)])
t = t[:, t.sum(axis=0) > 0]
hits += [stats.ttest_ind(y, x, equal_var=False).pvalue < .05,
stats.mannwhitneyu(y, x, method="asymptotic").pvalue < .05,
stats.chi2_contingency(t)[1] < .05]
return dict(zip(["t_test", "mann_whitney", "chi_square"], np.round(hits / reps, 3)))
print(reject(pA, pA)) # no difference: false-positive rate
print(reject(pA, pB)) # B shifted upward: powerThe figures below come from one seeded run of the R code. The tests on the fixed data (part 1) involve no randomness, and the Python version reproduces them exactly. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 38.7% instead of 40.7% for the t test's power), but they show the same pattern.
- The example. Group A's answers are 10, 20, 20 and 10 at levels 1-4; group B's are 5, 14, 25 and 16. The means are 2.50 and 2.87, and the medians 2.5 and 3.
- Three tests, same data. The Welch t test gives t(117.60) = 2.14, p = .034, with a mean difference of 0.37, 95% CI [0.03, 0.71]. Mann-Whitney gives W = 2185, p = .035: almost the same answer. The chi-square test gives χ²(3) = 4.67, p = .198, and misses the difference because it does not use the order.
- Effect size. The probability of superiority is .607: a randomly chosen person from group B scores higher than one from group A about 61% of the time (ties counted as half). The rank-biserial correlation is .21.
- No real difference. When both groups were drawn from the same distribution, the t test rejected in 4.6% of samples, Mann-Whitney in 4.1% and chi-square in 4.4%. All three keep their false-positive rate near 5%, so the t test is not "wrong" for these data.
- A real upward shift. When group B was drawn from its shifted distribution, the t test detected it in 40.7% of samples, Mann-Whitney in 40.0% and chi-square in only 27.4%. Ignoring the order costs about a third of the power.
What to do in practice
- One item, two groups, no covariates: use Mann-Whitney and report the probability of superiority or the rank-biserial correlation. Show the full distribution (counts or percentages at each level) in a table or stacked bar chart, because with four levels the distribution says more than any single number.
- If readers expect means: a Welch t test on the codes is a defensible sensitivity check, and it usually agrees. Say that you treated the levels as equally spaced.
- If you need covariates or several predictors: fit a proportional-odds model and report odds ratios with confidence intervals.
- Several items on the same construct: consider whether they should be summed or averaged into a scale score first. A total of many items behaves much more like a continuous variable than any single item does.
- Many separate items: adjust for multiple comparisons, for example with the Holm method, and say so.
Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report a Mann-Whitney U test in APA style (7th edition)
Give the medians (and ideally the distribution across levels), the test statistic, the exact p value and an effect size. Using the example above:
"Ratings were higher in group B (Mdn = 3, n = 60) than in group A (Mdn = 2.5, n = 60). A Mann-Whitney U test showed that this difference was statistically significant, U = 2185, p = .035, rank-biserial r = .21. A randomly chosen participant in group B rated higher than one in group A 61% of the time."
R labels the statistic W, which equals U for the first group named in wilcox.test(); other software may report the smaller U or a z value, so name the statistic you used. In the method section, state that the item was treated as ordinal, that a normal approximation with a tie correction was used, and, if you also ran a t test as a check, that it led to the same conclusion.
Related tools and guides
- APA 7 formatter for nonparametric tests
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Which statistical test should I use?
- What are the assumptions of the Wilcoxon signed-rank test?
- Welch vs Student t test: which should you use?
- Mann-Whitney U test (Wikipedia)
- Ordered logit model (Wikipedia)
More answered questions
- Is the chi-square test one-tailed or two-tailed?
- t test vs ANOVA with two groups: are the assumptions really different?
- How do you interpret Cohen's d, and are 0.2, 0.5 and 0.8 a real standard?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.