Answer

Mann-Whitney effect size: is z/√N right, and which CI should you report?

Inspired by a question on Cross Validated ·

nonparametriceffect sizeconfidence intervals

The short answer

You can compute z/√N, but do not wrap a Pearson-style (Fisher z) interval around it: that formula was derived for Pearson correlations, and in a simulation it gave intervals that were too narrow. Report the rank-biserial correlation or the probability of superiority, which have a clear meaning, with a bootstrap 95% CI, and add the Hodges-Lehmann shift with its exact CI when the outcome's units mean something.

The short answer

A popular recipe turns the large-sample z from a Mann-Whitney (Wilcoxon rank-sum) test into a correlation-like number by dividing by the square root of the total sample size, r = z/√N. It is easy to compute, but it has two weaknesses. It is not a Pearson correlation, so the usual Fisher z interval, tanh(atanh(r) ± 1.96/√(N − 3)), has no theoretical basis here. And it cannot reach 1: even when every score in one group beats every score in the other, z/√N tops out near 0.87 with equal groups, and lower with unequal ones.

Three effect sizes come straight from the same U statistic and are easier to interpret:

For an interval around the probability of superiority or the rank-biserial correlation, a bootstrap is a simple choice that worked well in the simulation below. Whether to run Mann-Whitney at all is a separate question, covered in t test or Mann-Whitney in small samples; this page assumes the rank test has been chosen and asks what to report alongside it.

Why z/√N and the Fisher interval do not fit together

See it in R and Python

The code uses a small fixed data set of response times (12 per group), computes every effect size above plus the Hodges-Lehmann shift, shows the ceiling on z/√N for three designs, and builds two 95% intervals for the rank-biserial correlation: a percentile bootstrap and the Fisher z formula. It then simulates 2,000 studies with 20 people per group, drawn from normal distributions half a standard deviation apart, and counts how often each interval contains the true rank-biserial correlation.

R

set.seed(124501)

# 1. Response times (seconds), 12 people per group, no ties
ctrl  <- c(1.12, 0.84, 1.95, 1.31, 0.97, 2.64, 1.08, 1.47, 0.76, 1.22, 3.10, 1.03)
treat <- c(1.58, 2.21, 1.36, 3.85, 1.74, 1.19, 2.47, 1.92, 4.30, 1.65, 2.05, 1.41)
n1 <- length(treat); n2 <- length(ctrl); N <- n1 + n2
c(median_treat = median(treat), median_ctrl = median(ctrl))
mw <- wilcox.test(treat, ctrl, conf.int = TRUE)   # exact test, Hodges-Lehmann CI
U  <- unname(mw$statistic)                        # pairs where treat > ctrl
z  <- (U - n1 * n2 / 2) / sqrt(n1 * n2 * (N + 1) / 12)  # large-sample z, no correction
r_rb <- 2 * U / (n1 * n2) - 1                     # rank-biserial correlation
round(c(U = U, p = mw$p.value, z = z, r_zN = z / sqrt(N),
        prob_sup = U / (n1 * n2), r_rb = r_rb), 4)
round(c(HL_shift = unname(mw$estimate), mw$conf.int), 2)

# 2. Largest possible z / sqrt(N): complete separation of the groups
r_max <- function(a, b) sqrt(3 * a * b / ((a + b) * (a + b + 1)))
round(c(n12_12 = r_max(12, 12), n6_18 = r_max(6, 18), n500_500 = r_max(500, 500)), 3)

# 3. Two 95% CIs for the rank-biserial correlation
rb <- function(X, Y) {   # rows of X = treatment samples, rows of Y = control samples
  X <- rbind(X); Y <- rbind(Y); s <- 0
  for (j in seq_len(ncol(Y))) s <- s + rowSums(X > Y[, j]) + 0.5 * rowSums(X == Y[, j])
  2 * s / (ncol(X) * ncol(Y)) - 1
}
resample <- function(v, B) matrix(sample(v, B * length(v), replace = TRUE), B)
boot <- rb(resample(treat, 4000), resample(ctrl, 4000))
round(c(boot = unname(quantile(boot, c(.025, .975))),
        fisher = tanh(atanh(r_rb) + c(-1, 1) * qnorm(.975) / sqrt(N - 3))), 3)

# 4. Coverage of each CI: normal data, 20 per group, shift of 0.5 SD
truth <- 2 * pnorm(0.5 / sqrt(2)) - 1             # true rank-biserial correlation
hits <- replicate(2000, {
  x <- rnorm(20, 0.5); y <- rnorm(20); r <- rb(x, y)
  bs <- quantile(rb(resample(x, 1000), resample(y, 1000)), c(.025, .975))
  fz <- tanh(atanh(r) + c(-1, 1) * qnorm(.975) / sqrt(37))
  c(boot = unname(bs[1] <= truth & truth <= bs[2]), fisher = fz[1] <= truth & truth <= fz[2])
})
round(c(truth = truth, rowMeans(hits)), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(124501)

# 1. Response times (seconds), 12 people per group, no ties
ctrl = np.array([1.12, 0.84, 1.95, 1.31, 0.97, 2.64, 1.08, 1.47, 0.76, 1.22, 3.10, 1.03])
treat = np.array([1.58, 2.21, 1.36, 3.85, 1.74, 1.19, 2.47, 1.92, 4.30, 1.65, 2.05, 1.41])
n1, n2 = len(treat), len(ctrl); N = n1 + n2
print("median_treat", np.median(treat), "median_ctrl", np.median(ctrl))
mw = stats.mannwhitneyu(treat, ctrl, method="exact")
U = mw.statistic                                       # pairs where treat > ctrl
z = (U - n1 * n2 / 2) / np.sqrt(n1 * n2 * (N + 1) / 12)  # large-sample z, no correction
r_rb = 2 * U / (n1 * n2) - 1                           # rank-biserial correlation
print("U", U, "p", round(mw.pvalue, 4), "z", round(z, 4), "r_zN", round(z / np.sqrt(N), 4),
      "prob_sup", round(U / (n1 * n2), 4), "r_rb", round(r_rb, 4))

# Hodges-Lehmann shift and its exact 95% CI (same rule as R's wilcox.test)
diffs = np.sort(np.subtract.outer(treat, ctrl).ravel())
f = np.zeros((n1 + 1, n2 + 1, n1 * n2 + 1)); f[0, :, 0] = f[:, 0, 0] = 1
for a in range(1, n1 + 1):                             # exact null distribution of U
    for b in range(1, n2 + 1):
        f[a, b, b:] += f[a - 1, b, :n1 * n2 + 1 - b]
        f[a, b] += f[a, b - 1]
cdf = np.cumsum(f[n1, n2]) / f[n1, n2].sum()
qu = max(int(np.argmax(cdf >= 0.025)), 1); ql = n1 * n2 - qu
print("HL_shift", round(np.median(diffs), 2), "CI", round(diffs[qu - 1], 2), round(diffs[ql], 2))

# 2. Largest possible z / sqrt(N): complete separation of the groups
r_max = lambda a, b: np.sqrt(3 * a * b / ((a + b) * (a + b + 1)))
print([round(r_max(*s), 3) for s in [(12, 12), (6, 18), (500, 500)]])

# 3. Two 95% CIs for the rank-biserial correlation
def rb(x, y):              # rows of x = treatment samples, rows of y = control samples
    x, y = np.atleast_2d(x), np.atleast_2d(y)
    d = x[:, :, None] - y[:, None, :]
    return 2 * ((d > 0).mean(axis=(1, 2)) + 0.5 * (d == 0).mean(axis=(1, 2))) - 1

boot = rb(rng.choice(treat, (4000, n1)), rng.choice(ctrl, (4000, n2)))
fz = np.tanh(np.arctanh(r_rb) + np.array([-1, 1]) * stats.norm.ppf(.975) / np.sqrt(N - 3))
print("boot", np.round(np.quantile(boot, [.025, .975]), 3), "fisher", np.round(fz, 3))

# 4. Coverage of each CI: normal data, 20 per group, shift of 0.5 SD
truth = 2 * stats.norm.cdf(0.5 / np.sqrt(2)) - 1      # true rank-biserial correlation
hits = np.zeros(2)
for _ in range(2000):
    x, y = rng.normal(0.5, 1, 20), rng.normal(0, 1, 20); r = rb(x, y)[0]
    bs = np.quantile(rb(rng.choice(x, (1000, 20)), rng.choice(y, (1000, 20))), [.025, .975])
    fz = np.tanh(np.arctanh(r) + np.array([-1, 1]) * stats.norm.ppf(.975) / np.sqrt(37))
    hits += [bs[0] <= truth <= bs[1], fz[0] <= truth <= fz[1]]
print("truth", round(truth, 3), "boot", hits[0] / 2000, "fisher", hits[1] / 2000)

The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly (including the exact Hodges-Lehmann interval, which it computes from the null distribution of U). Python's random numbers differ from R's, so its bootstrap and coverage figures differ slightly but show the same pattern.

Percentile bootstrap intervals have limits of their own: with very small groups, or when the groups barely overlap so that most resamples give a rank-biserial correlation of exactly 1, they can be too short. In those cases say so, and lean on the exact Hodges-Lehmann interval or report the effect size without a bootstrap interval.

Should the interval be one-sided or two-sided?

The interval should match the claim you make. A one-sided test at α = .05 corresponds to a one-sided 95% bound, which is the same as the lower (or upper) limit of a two-sided 90% interval. If you planned a directional test and want an interval that agrees with it, report that bound and label it clearly.

Most readers expect a two-sided 95% interval as a description of the plausible range of the effect, and reporting one is perfectly acceptable. Just be aware that it can include zero even when a one-sided test reached p < .05; that is not a contradiction, because the two procedures spend the 5% error rate differently. Decide which you will report before seeing the data. For more on the choice itself, see One-tailed vs two-tailed tests.

What to report in practice

  1. The test: U, which group it counts in favour of, and whether the p value is exact or from the normal approximation.
  2. One headline effect size: the rank-biserial correlation or the probability of superiority, with a 95% bootstrap interval. State the number of resamples and the bootstrap method.
  3. A shift in real units when it helps: the Hodges-Lehmann estimate with its interval. The interval is derived assuming the two distributions differ only by a shift; if their shapes clearly differ, describe it as the median of all pairwise differences.
  4. If a journal insists on z/√N: give it as a point value next to the rank-biserial correlation, note that its maximum is below 1, and do not attach a Fisher z interval to it.

For the parametric counterpart, see Effect size r for a t test; for a refresher on what the interval itself promises, see What does 95% confidence mean?. More plain-language guides are on the DASS blog.

How to report a Mann-Whitney U test effect size in APA style (7th edition)

Give the medians and group sizes, U and its p value, then the effect size with its interval. With the example above:

"Response times were longer in the treatment group (Mdn = 1.83 s, n = 12) than in the control group (Mdn = 1.17 s, n = 12), exact Mann-Whitney U = 112, p = .02. The rank-biserial correlation was .56, 95% CI [.125, .889] (percentile bootstrap, 4,000 resamples), and the Hodges-Lehmann estimate of the shift was 0.60 s, 95% CI [0.11, 1.16]."

Correlations and probabilities cannot exceed 1, so they take no leading zero; the shift is in seconds, so it keeps one. In the method section, name the effect size, explain how its interval was obtained, and state which group U counts in favour of, because programs differ: R labels the statistic W, which here equals U for the treatment group.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.