Answer
Mann-Whitney effect size: is z/√N right, and which CI should you report?
The short answer
You can compute z/√N, but do not wrap a Pearson-style (Fisher z) interval around it: that formula was derived for Pearson correlations, and in a simulation it gave intervals that were too narrow. Report the rank-biserial correlation or the probability of superiority, which have a clear meaning, with a bootstrap 95% CI, and add the Hodges-Lehmann shift with its exact CI when the outcome's units mean something.
The short answer
A popular recipe turns the large-sample z from a Mann-Whitney (Wilcoxon rank-sum) test into a correlation-like number by dividing by the square root of the total sample size, r = z/√N. It is easy to compute, but it has two weaknesses. It is not a Pearson correlation, so the usual Fisher z interval, tanh(atanh(r) ± 1.96/√(N − 3)), has no theoretical basis here. And it cannot reach 1: even when every score in one group beats every score in the other, z/√N tops out near 0.87 with equal groups, and lower with unequal ones.
Three effect sizes come straight from the same U statistic and are easier to interpret:
- Probability of superiority (also called the common-language effect size): U divided by n₁n₂, the number of treatment-control pairs. It is the chance that a randomly chosen person from one group scores higher than a randomly chosen person from the other, with ties counted as half.
- Rank-biserial correlation: 2 × (probability of superiority) − 1. It runs from −1 to 1 and reaches ±1 when the groups do not overlap at all.
- Hodges-Lehmann shift: the median of all treatment-minus-control differences, in the outcome's own units. R's
wilcox.test(..., conf.int = TRUE)gives it with an exact confidence interval.
For an interval around the probability of superiority or the rank-biserial correlation, a bootstrap is a simple choice that worked well in the simulation below. Whether to run Mann-Whitney at all is a separate question, covered in t test or Mann-Whitney in small samples; this page assumes the rank test has been chosen and asks what to report alongside it.
Why z/√N and the Fisher interval do not fit together
- What z/√N really measures. Without ties, and using the plain large-sample z, z/√N equals the rank-biserial correlation multiplied by √(3n₁n₂ / (N(N + 1))). That factor depends on how people are split between the groups: 0.849 for 12 + 12, 0.735 for 6 + 18 and 0.866 for 500 + 500 (the limit for equal groups is √3/2).
- Unbalanced designs look weaker. The same degree of separation earns a smaller "r" when the groups differ in size, so Cohen's .1/.3/.5 labels do not transfer cleanly.
- Different software, different z. Programs differ in whether they apply a tie correction or a continuity correction to z, so the same data can give slightly different values of z/√N. The U statistic itself does not change.
- Fisher's interval assumes something that is not true here. Its standard error, 1/√(N − 3) on the atanh scale, comes from the sampling distribution of Pearson's r for bivariate normal data.
- The spread of U depends on more than N. Even under the null hypothesis its variance is n₁n₂(N + 1)/12, which depends on the two group sizes separately, and away from the null it also depends on the shapes of the distributions. No formula in N alone can be right in general.
See it in R and Python
The code uses a small fixed data set of response times (12 per group), computes every effect size above plus the Hodges-Lehmann shift, shows the ceiling on z/√N for three designs, and builds two 95% intervals for the rank-biserial correlation: a percentile bootstrap and the Fisher z formula. It then simulates 2,000 studies with 20 people per group, drawn from normal distributions half a standard deviation apart, and counts how often each interval contains the true rank-biserial correlation.
R
set.seed(124501)
# 1. Response times (seconds), 12 people per group, no ties
ctrl <- c(1.12, 0.84, 1.95, 1.31, 0.97, 2.64, 1.08, 1.47, 0.76, 1.22, 3.10, 1.03)
treat <- c(1.58, 2.21, 1.36, 3.85, 1.74, 1.19, 2.47, 1.92, 4.30, 1.65, 2.05, 1.41)
n1 <- length(treat); n2 <- length(ctrl); N <- n1 + n2
c(median_treat = median(treat), median_ctrl = median(ctrl))
mw <- wilcox.test(treat, ctrl, conf.int = TRUE) # exact test, Hodges-Lehmann CI
U <- unname(mw$statistic) # pairs where treat > ctrl
z <- (U - n1 * n2 / 2) / sqrt(n1 * n2 * (N + 1) / 12) # large-sample z, no correction
r_rb <- 2 * U / (n1 * n2) - 1 # rank-biserial correlation
round(c(U = U, p = mw$p.value, z = z, r_zN = z / sqrt(N),
prob_sup = U / (n1 * n2), r_rb = r_rb), 4)
round(c(HL_shift = unname(mw$estimate), mw$conf.int), 2)
# 2. Largest possible z / sqrt(N): complete separation of the groups
r_max <- function(a, b) sqrt(3 * a * b / ((a + b) * (a + b + 1)))
round(c(n12_12 = r_max(12, 12), n6_18 = r_max(6, 18), n500_500 = r_max(500, 500)), 3)
# 3. Two 95% CIs for the rank-biserial correlation
rb <- function(X, Y) { # rows of X = treatment samples, rows of Y = control samples
X <- rbind(X); Y <- rbind(Y); s <- 0
for (j in seq_len(ncol(Y))) s <- s + rowSums(X > Y[, j]) + 0.5 * rowSums(X == Y[, j])
2 * s / (ncol(X) * ncol(Y)) - 1
}
resample <- function(v, B) matrix(sample(v, B * length(v), replace = TRUE), B)
boot <- rb(resample(treat, 4000), resample(ctrl, 4000))
round(c(boot = unname(quantile(boot, c(.025, .975))),
fisher = tanh(atanh(r_rb) + c(-1, 1) * qnorm(.975) / sqrt(N - 3))), 3)
# 4. Coverage of each CI: normal data, 20 per group, shift of 0.5 SD
truth <- 2 * pnorm(0.5 / sqrt(2)) - 1 # true rank-biserial correlation
hits <- replicate(2000, {
x <- rnorm(20, 0.5); y <- rnorm(20); r <- rb(x, y)
bs <- quantile(rb(resample(x, 1000), resample(y, 1000)), c(.025, .975))
fz <- tanh(atanh(r) + c(-1, 1) * qnorm(.975) / sqrt(37))
c(boot = unname(bs[1] <= truth & truth <= bs[2]), fisher = fz[1] <= truth & truth <= fz[2])
})
round(c(truth = truth, rowMeans(hits)), 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(124501)
# 1. Response times (seconds), 12 people per group, no ties
ctrl = np.array([1.12, 0.84, 1.95, 1.31, 0.97, 2.64, 1.08, 1.47, 0.76, 1.22, 3.10, 1.03])
treat = np.array([1.58, 2.21, 1.36, 3.85, 1.74, 1.19, 2.47, 1.92, 4.30, 1.65, 2.05, 1.41])
n1, n2 = len(treat), len(ctrl); N = n1 + n2
print("median_treat", np.median(treat), "median_ctrl", np.median(ctrl))
mw = stats.mannwhitneyu(treat, ctrl, method="exact")
U = mw.statistic # pairs where treat > ctrl
z = (U - n1 * n2 / 2) / np.sqrt(n1 * n2 * (N + 1) / 12) # large-sample z, no correction
r_rb = 2 * U / (n1 * n2) - 1 # rank-biserial correlation
print("U", U, "p", round(mw.pvalue, 4), "z", round(z, 4), "r_zN", round(z / np.sqrt(N), 4),
"prob_sup", round(U / (n1 * n2), 4), "r_rb", round(r_rb, 4))
# Hodges-Lehmann shift and its exact 95% CI (same rule as R's wilcox.test)
diffs = np.sort(np.subtract.outer(treat, ctrl).ravel())
f = np.zeros((n1 + 1, n2 + 1, n1 * n2 + 1)); f[0, :, 0] = f[:, 0, 0] = 1
for a in range(1, n1 + 1): # exact null distribution of U
for b in range(1, n2 + 1):
f[a, b, b:] += f[a - 1, b, :n1 * n2 + 1 - b]
f[a, b] += f[a, b - 1]
cdf = np.cumsum(f[n1, n2]) / f[n1, n2].sum()
qu = max(int(np.argmax(cdf >= 0.025)), 1); ql = n1 * n2 - qu
print("HL_shift", round(np.median(diffs), 2), "CI", round(diffs[qu - 1], 2), round(diffs[ql], 2))
# 2. Largest possible z / sqrt(N): complete separation of the groups
r_max = lambda a, b: np.sqrt(3 * a * b / ((a + b) * (a + b + 1)))
print([round(r_max(*s), 3) for s in [(12, 12), (6, 18), (500, 500)]])
# 3. Two 95% CIs for the rank-biserial correlation
def rb(x, y): # rows of x = treatment samples, rows of y = control samples
x, y = np.atleast_2d(x), np.atleast_2d(y)
d = x[:, :, None] - y[:, None, :]
return 2 * ((d > 0).mean(axis=(1, 2)) + 0.5 * (d == 0).mean(axis=(1, 2))) - 1
boot = rb(rng.choice(treat, (4000, n1)), rng.choice(ctrl, (4000, n2)))
fz = np.tanh(np.arctanh(r_rb) + np.array([-1, 1]) * stats.norm.ppf(.975) / np.sqrt(N - 3))
print("boot", np.round(np.quantile(boot, [.025, .975]), 3), "fisher", np.round(fz, 3))
# 4. Coverage of each CI: normal data, 20 per group, shift of 0.5 SD
truth = 2 * stats.norm.cdf(0.5 / np.sqrt(2)) - 1 # true rank-biserial correlation
hits = np.zeros(2)
for _ in range(2000):
x, y = rng.normal(0.5, 1, 20), rng.normal(0, 1, 20); r = rb(x, y)[0]
bs = np.quantile(rb(rng.choice(x, (1000, 20)), rng.choice(y, (1000, 20))), [.025, .975])
fz = np.tanh(np.arctanh(r) + np.array([-1, 1]) * stats.norm.ppf(.975) / np.sqrt(37))
hits += [bs[0] <= truth <= bs[1], fz[0] <= truth <= fz[1]]
print("truth", round(truth, 3), "boot", hits[0] / 2000, "fisher", hits[1] / 2000)The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly (including the exact Hodges-Lehmann interval, which it computes from the null distribution of U). Python's random numbers differ from R's, so its bootstrap and coverage figures differ slightly but show the same pattern.
- The test. Treatment median 1.83 s, control median 1.17 s. The exact Mann-Whitney test gives U = 112, meaning the treatment score was higher in 112 of the 144 treatment-control pairs, p = .02.
- Four effect sizes from one U. The large-sample z is 2.31, so z/√N = .47. The probability of superiority is .78 and the rank-biserial correlation is .56. As predicted, .47 is .56 times the 0.849 factor for a 12 + 12 design.
- In the outcome's units. The Hodges-Lehmann shift is 0.60 s, with an exact 95% CI of [0.11, 1.16] s.
- Two intervals for the rank-biserial correlation. The percentile bootstrap (4,000 resamples) gave [.125, .889]; the Fisher z formula gave the narrower [.196, .783].
- Which interval can you trust? The true rank-biserial correlation in the simulation was .276. The bootstrap interval contained it in 95.7% of the 2,000 studies, close to the promised 95%. The Fisher interval contained it in only 90.5%: it is too narrow. Python's run gave 93.9% and 90.0%, the same picture.
Percentile bootstrap intervals have limits of their own: with very small groups, or when the groups barely overlap so that most resamples give a rank-biserial correlation of exactly 1, they can be too short. In those cases say so, and lean on the exact Hodges-Lehmann interval or report the effect size without a bootstrap interval.
Should the interval be one-sided or two-sided?
The interval should match the claim you make. A one-sided test at α = .05 corresponds to a one-sided 95% bound, which is the same as the lower (or upper) limit of a two-sided 90% interval. If you planned a directional test and want an interval that agrees with it, report that bound and label it clearly.
Most readers expect a two-sided 95% interval as a description of the plausible range of the effect, and reporting one is perfectly acceptable. Just be aware that it can include zero even when a one-sided test reached p < .05; that is not a contradiction, because the two procedures spend the 5% error rate differently. Decide which you will report before seeing the data. For more on the choice itself, see One-tailed vs two-tailed tests.
What to report in practice
- The test: U, which group it counts in favour of, and whether the p value is exact or from the normal approximation.
- One headline effect size: the rank-biserial correlation or the probability of superiority, with a 95% bootstrap interval. State the number of resamples and the bootstrap method.
- A shift in real units when it helps: the Hodges-Lehmann estimate with its interval. The interval is derived assuming the two distributions differ only by a shift; if their shapes clearly differ, describe it as the median of all pairwise differences.
- If a journal insists on z/√N: give it as a point value next to the rank-biserial correlation, note that its maximum is below 1, and do not attach a Fisher z interval to it.
For the parametric counterpart, see Effect size r for a t test; for a refresher on what the interval itself promises, see What does 95% confidence mean?. More plain-language guides are on the DASS blog.
How to report a Mann-Whitney U test effect size in APA style (7th edition)
Give the medians and group sizes, U and its p value, then the effect size with its interval. With the example above:
"Response times were longer in the treatment group (Mdn = 1.83 s, n = 12) than in the control group (Mdn = 1.17 s, n = 12), exact Mann-Whitney U = 112, p = .02. The rank-biserial correlation was .56, 95% CI [.125, .889] (percentile bootstrap, 4,000 resamples), and the Hodges-Lehmann estimate of the shift was 0.60 s, 95% CI [0.11, 1.16]."
Correlations and probabilities cannot exceed 1, so they take no leading zero; the shift is in seconds, so it keeps one. In the method section, name the effect size, explain how its interval was obtained, and state which group U counts in favour of, because programs differ: R labels the statistic W, which here equals U for the treatment group.
Related tools and guides
- APA 7 formatter for nonparametric tests
- t test or Mann-Whitney in small samples: how do you choose?
- How do you compare a short ordinal rating scale between two groups?
- Effect size r for a t test
- Why does bootstrapping work?
- Mann-Whitney U test (Wikipedia)
More answered questions
- What should you do after rejecting the null hypothesis?
- Why does almost everything become statistically significant with a large sample?
- What does r mean when it is reported next to a t test, and how do you turn it into d?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.