Answer

What does r mean when it is reported next to a t test, and how do you turn it into d?

Inspired by a question on Cross Validated ·

effect sizet testcorrelation

The short answer

When a paper gives an r beside an independent-samples t test, it is an effect size: the correlation between group membership (coded 0/1) and the outcome, computed as r = √(t² / (t² + df)). r² is the share of outcome variance explained by group. To get Cohen's d, use d = t × √(1/n₁ + 1/n₂), or d ≈ 2r / √(1 − r²) when the groups are about the same size.

The short answer

An r printed after a t test is almost always an effect size, not a separate correlation between two measured variables. It answers the question "how strongly is the outcome related to which group a person is in?" on the familiar −1 to 1 scale of a correlation. Authors usually compute it straight from the test statistic, using a formula popularised by Robert Rosenthal for meta-analysis:

r = √(t² / (t² + df)), where df is the degrees of freedom of the t test (n₁ + n₂ − 2 for two independent groups).

This is not a loose analogy. For Student's (equal-variance) t test, the number this formula gives is exactly the ordinary Pearson correlation you would get by coding one group as 1, the other as 0, and correlating that code with the scores. Statisticians call it the point-biserial correlation. The formula only returns the size, so the sign is set by the direction of the difference.

Seeing it in one example

The code below simulates two groups of 40 with a true difference of half a standard deviation, runs a t test, and then computes r, r², and Cohen's d in several ways so you can see they agree. The last two parts involve no randomness: they show how r depends on the group split and what Cohen's benchmarks for r mean on the d scale.

R

set.seed(632153)

# Two independent groups of 40, true difference of half an SD
score <- c(rnorm(40, mean = 55, sd = 10), rnorm(40, mean = 50, sd = 10))
group <- rep(c(1, 0), each = 40)               # 1 = treatment, 0 = control
n1 <- 40; n2 <- 40

round(tapply(score, group, mean), 2); round(tapply(score, group, sd), 2)
tt <- t.test(score[group == 1], score[group == 0], var.equal = TRUE)
t  <- unname(tt$statistic); df <- unname(tt$parameter)
round(c(t = t, df = df, p = tt$p.value), 3)
round(tt$conf.int, 2)                           # 95% CI for the mean difference

# 1. r from t, and the same r as a plain correlation with the 0/1 group code
r_from_t <- sqrt(t^2 / (t^2 + df))
round(c(r_from_t = r_from_t, cor_with_group = cor(score, group)), 3)

# 2. r squared = share of variance explained by group (eta squared)
ss <- summary(aov(score ~ factor(group)))[[1]][["Sum Sq"]]
round(c(r_squared = r_from_t^2, eta_squared = ss[1] / sum(ss)), 3)

# 3. Cohen's d three ways
sp <- sqrt((39 * var(score[group == 1]) + 39 * var(score[group == 0])) / 78)
d_pooled  <- (mean(score[group == 1]) - mean(score[group == 0])) / sp
d_from_t  <- t * sqrt(1 / n1 + 1 / n2)
d_from_r  <- 2 * r_from_t / sqrt(1 - r_from_t^2)  # equal-n shortcut
round(c(d_pooled = d_pooled, d_from_t = d_from_t, d_from_r = d_from_r), 3)
round(c(mean_diff = mean(score[group == 1]) - mean(score[group == 0]), d = d_pooled), 2)

# 4. Same true d = 0.5, different group splits: r shrinks as groups get unequal
p <- c(0.5, 0.3, 0.2, 0.1)                      # share of people in group 1
round(data.frame(p, r = 0.5 * sqrt(p * (1 - p)) / sqrt(1 + 0.25 * p * (1 - p))), 3)

# 5. Cohen's r benchmarks translated to d (equal groups, large samples)
r_b <- c(0.10, 0.30, 0.50)
round(data.frame(r = r_b, d = 2 * r_b / sqrt(1 - r_b^2)), 2)

Python

import numpy as np
import pandas as pd
from scipy import stats

rng = np.random.default_rng(632153)

# Two independent groups of 40, true difference of half an SD
treat = rng.normal(55, 10, 40)
control = rng.normal(50, 10, 40)
score = np.concatenate([treat, control])
group = np.repeat([1, 0], 40)                 # 1 = treatment, 0 = control
n1, n2 = 40, 40

print("M", round(control.mean(), 2), round(treat.mean(), 2),
      "SD", round(control.std(ddof=1), 2), round(treat.std(ddof=1), 2))
tt = stats.ttest_ind(treat, control, equal_var=True)
t, df = tt.statistic, n1 + n2 - 2
ci = tt.confidence_interval(0.95)             # 95% CI for the mean difference
print("t", round(t, 3), "df", df, "p", round(tt.pvalue, 3),
      "CI", round(ci.low, 2), round(ci.high, 2))

# 1. r from t, and the same r as a plain correlation with the 0/1 group code
r_from_t = np.sqrt(t**2 / (t**2 + df))
print("r_from_t", round(r_from_t, 3), "cor_with_group", round(np.corrcoef(score, group)[0, 1], 3))

# 2. r squared = share of variance explained by group (eta squared)
ss_between = n1 * (treat.mean() - score.mean())**2 + n2 * (control.mean() - score.mean())**2
ss_total = ((score - score.mean())**2).sum()
print("r_squared", round(r_from_t**2, 3), "eta_squared", round(ss_between / ss_total, 3))

# 3. Cohen's d three ways
sp = np.sqrt((39 * treat.var(ddof=1) + 39 * control.var(ddof=1)) / 78)
d_pooled = (treat.mean() - control.mean()) / sp
d_from_t = t * np.sqrt(1 / n1 + 1 / n2)
d_from_r = 2 * r_from_t / np.sqrt(1 - r_from_t**2)  # equal-n shortcut
print("d_pooled", round(d_pooled, 3), "d_from_t", round(d_from_t, 3), "d_from_r", round(d_from_r, 3))
print("mean_diff", round(treat.mean() - control.mean(), 2), "d", round(d_pooled, 2))

# 4. Same true d = 0.5, different group splits: r shrinks as groups get unequal
p = np.array([0.5, 0.3, 0.2, 0.1])            # share of people in group 1
print(pd.DataFrame({"p": p, "r": 0.5 * np.sqrt(p * (1 - p)) / np.sqrt(1 + 0.25 * p * (1 - p))}).round(3))

# 5. Cohen's r benchmarks translated to d (equal groups, large samples)
r_b = np.array([0.10, 0.30, 0.50])
print(pd.DataFrame({"r": r_b, "d": 2 * r_b / np.sqrt(1 - r_b**2)}).round(2))

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated sample is different (its groups happen to be closer together, with r = .201 and d = 0.404), but every identity holds in the same way, and parts 4 and 5, which involve no randomness, match R exactly.

Converting between r and d

Both numbers describe the same difference, on different scales. d is a mean difference in standard deviation units; r is a correlation. The conversions for two independent groups are:

The benchmarks are not interchangeable. Cohen's labels of .10, .30 and .50 for r correspond to d values of 0.20, 0.63 and 1.15, so a "medium" r is a bigger effect than a "medium" d of 0.5. For more on what d values mean, see How do you interpret Cohen's d?.

Cautions before you compare r values

If you are choosing which test and effect size to report, the test chooser can help, and the DASS blog has more plain-language guides to reporting results.

How to report effect size r for a t test in APA style (7th edition)

Report r right after the test, with two decimals and no leading zero, because a correlation cannot exceed 1. Using the example above:

"Participants in the treatment group scored higher (M = 55.06, SD = 9.58, n = 40) than those in the control group (M = 49.81, SD = 12.33, n = 40), t(78) = 2.13, p = .037, r = .23, mean difference = 5.25, 95% CI [0.33, 10.17]."

If your readers expect d, give it instead or as well (d = 0.48, with a leading zero because d can exceed 1). In the method section, say how the effect size was computed, for example: "Effect sizes are point-biserial correlations computed as r = √(t² / (t² + df))."

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.