Answer

Confidence interval for eta squared: does the Fisher z shortcut work?

Inspired by a question on Cross Validated ·

anovaeffect sizeconfidence intervals

The short answer

Not reliably, and no choice of N rescues it. The Fisher z interval was built for a single Pearson correlation, so it ignores how many groups were compared and the upward bias of eta squared. In a simulation it held up with three or four groups but covered the true value only 64-80% of the time with eight small groups. The standard method inverts the noncentral F distribution, which needs only F, df1 and df2, and its coverage was 95% in every scenario.

The short answer

Eta squared (η²) is the share of the total variation in an outcome that is explained by group membership: the between-groups sum of squares divided by the total sum of squares. It is natural to want an interval around it, and when only a published F and its degrees of freedom are available, a tempting recipe is to take √η² as if it were a correlation, apply Fisher's z transformation (atanh), add ±1.96/√(N − 3), transform back with tanh and square the limits.

That recipe borrows a formula that does not belong here, so there is no correct N to put into it. The better route uses the same three numbers: F, df1 and df2. Each possible true effect size implies a noncentral F distribution for the test statistic, and the interval collects the effect sizes under which the observed F is not surprising. This is the standard method (described, for example, by Steiger, 2004, in Psychological Methods), it takes a few lines of R, and R packages such as MBESS implement it.

Why the Fisher z shortcut breaks down

So the question of which N to use (total observations, number of participants, or df2) has no good answer: changing N changes the width, but it cannot account for df1 or for the bias.

See it in R and Python

The code runs a one-way ANOVA on four simulated groups of 15, computes η², and builds three intervals: the exact 95% and 90% intervals from the noncentral F distribution, and the Fisher z version. It then checks the coverage of both 95% intervals over 2,000 simulated studies per scenario, drawing each study's F statistic from the noncentral F distribution implied by a known true η² (equivalent to simulating normal data with equal group sizes).

R

set.seed(96151)

# 1. One-way ANOVA: four groups of 15
g <- factor(rep(c("A", "B", "C", "D"), each = 15))
y <- rnorm(60, mean = rep(c(50, 52, 55, 59), each = 15), sd = 10)
tab <- anova(lm(y ~ g))
Fobs <- tab$`F value`[1]; df1 <- tab$Df[1]; df2 <- tab$Df[2]
eta2 <- tab$`Sum Sq`[1] / sum(tab$`Sum Sq`)
round(c(F = Fobs, df1 = df1, df2 = df2, p = tab$`Pr(>F)`[1], eta2 = eta2), 4)

# 2. Exact CI: find the noncentrality values that put the observed F
#    at the upper and lower tail cut-offs, then convert them to eta squared
ncp_for <- function(prob, F, df1, df2) {   # ncp at which P(F' <= F) = prob
  if (pf(F, df1, df2) < prob) return(0)
  uniroot(function(l) pf(F, df1, df2, ncp = l) - prob,
          c(0, 50 + 5 * F * df1), tol = 1e-10)$root
}
eta2_ci <- function(F, df1, df2, level = .95) {
  a <- (1 - level) / 2
  lam <- c(ncp_for(1 - a, F, df1, df2), ncp_for(a, F, df1, df2))
  lam / (lam + df1 + df2 + 1)               # df1 + df2 + 1 = N in a one-way design
}
round(rbind(exact95 = eta2_ci(Fobs, df1, df2, .95),
            exact90 = eta2_ci(Fobs, df1, df2, .90)), 3)

# 3. The Fisher z shortcut: treat sqrt(eta2) as a Pearson r
fisher_ci <- function(e2, N) {
  r <- tanh(atanh(sqrt(e2)) + c(-1, 1) * qnorm(.975) / sqrt(N - 3))
  pmax(r, 0)^2                               # a negative lower limit becomes 0
}
round(fisher_ci(eta2, 60), 3)

# 4. Coverage of both 95% CIs over 2,000 simulated studies
coverage <- function(k, n, true_eta2, n_sim = 2000) {
  N <- k * n; df1 <- k - 1; df2 <- N - k
  Fs <- rf(n_sim, df1, df2, ncp = N * true_eta2 / (1 - true_eta2))
  fz <- sapply(df1 * Fs / (df1 * Fs + df2), fisher_ci, N = N)
  ex <- sapply(Fs, eta2_ci, df1 = df1, df2 = df2)
  c(fisher = mean(fz[1, ] <= true_eta2 & true_eta2 <= fz[2, ]),
    exact  = mean(ex[1, ] <= true_eta2 & true_eta2 <= ex[2, ]))
}
round(rbind(k3_n20_eta.10 = coverage(3, 20, .10), k4_n15_eta.10 = coverage(4, 15, .10),
            k8_n5_eta.05 = coverage(8, 5, .05), k8_n5_eta.15 = coverage(8, 5, .15)), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(96151)

# 1. One-way ANOVA: four groups of 15
means = np.repeat([50, 52, 55, 59], 15)
y = rng.normal(means, 10)
groups = y.reshape(4, 15)
Fobs, p = stats.f_oneway(*groups)
df1, df2 = 3, 56
ss_between = (15 * (groups.mean(axis=1) - y.mean()) ** 2).sum()
eta2 = ss_between / ((y - y.mean()) ** 2).sum()
print("F", round(Fobs, 4), "df1", df1, "df2", df2, "p", round(p, 4), "eta2", round(eta2, 4))

# 2. Exact CI: find the noncentrality values that put the observed F
#    at the upper and lower tail cut-offs (vectorised bisection), then convert
def ncp_for(prob, F, df1, df2):          # ncp at which P(F' <= F) = prob
    F = np.atleast_1d(np.asarray(F, dtype=float))
    lo, hi = np.zeros_like(F), 50 + 5 * F * df1
    for _ in range(80):
        mid = (lo + hi) / 2
        up = stats.ncf.cdf(F, df1, df2, mid) > prob   # cdf falls as ncp rises
        lo, hi = np.where(up, mid, lo), np.where(up, hi, mid)
    return np.where(stats.f.cdf(F, df1, df2) < prob, 0.0, (lo + hi) / 2)

def eta2_ci(F, df1, df2, level=.95):
    a = (1 - level) / 2
    lam = np.array([ncp_for(1 - a, F, df1, df2), ncp_for(a, F, df1, df2)])
    return lam / (lam + df1 + df2 + 1)       # df1 + df2 + 1 = N in a one-way design

print("exact95", np.round(eta2_ci(Fobs, df1, df2, .95).ravel(), 3))
print("exact90", np.round(eta2_ci(Fobs, df1, df2, .90).ravel(), 3))
# The same functions applied to the F from the R run (no randomness involved)
print("R's F:", np.round(eta2_ci(3.1651, 3, 56, .95).ravel(), 3),
      np.round(eta2_ci(3.1651, 3, 56, .90).ravel(), 3))

# 3. The Fisher z shortcut: treat sqrt(eta2) as a Pearson r
def fisher_ci(e2, N):
    e2 = np.atleast_1d(e2)
    z = np.arctanh(np.sqrt(e2))
    se = stats.norm.ppf(.975) / np.sqrt(N - 3)
    r = np.tanh(np.array([z - se, z + se]))
    return np.maximum(r, 0) ** 2             # a negative lower limit becomes 0

print("fisher", np.round(fisher_ci(eta2, 60).ravel(), 3))

# 4. Coverage of both 95% CIs over 2,000 simulated studies
def coverage(k, n, true_eta2, n_sim=2000):
    N = k * n; df1, df2 = k - 1, N - k
    Fs = rng.noncentral_f(df1, df2, N * true_eta2 / (1 - true_eta2), n_sim)
    fz = fisher_ci(df1 * Fs / (df1 * Fs + df2), N)
    ex = eta2_ci(Fs, df1, df2)
    return (np.mean((fz[0] <= true_eta2) & (true_eta2 <= fz[1])),
            np.mean((ex[0] <= true_eta2) & (true_eta2 <= ex[1])))

for name, args in [("k3_n20_eta.10", (3, 20, .10)), ("k4_n15_eta.10", (4, 15, .10)),
                   ("k8_n5_eta.05", (8, 5, .05)), ("k8_n5_eta.15", (8, 5, .15))]:
    print(name, "fisher %.3f exact %.3f" % coverage(*args))

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its sample (F = 2.94 rather than 3.17) and its coverage figures differ slightly, but they show the same pattern. Given R's F value, the Python functions return exactly the same exact intervals, because that step involves no randomness.

90% or 95%, and what about repeated measures?

Because η² cannot be negative, the lower limit of the exact interval is often exactly zero. Steiger (2004) recommends reporting a 90% interval alongside an F test at α = .05: its lower limit is above zero exactly when p < .05, so the interval and the test always agree, as they do in the example. A 95% interval is also defensible; just say which level you used, and expect its lower limit to be zero whenever p is above .025.

For a repeated-measures ANOVA, the effect size usually reported is partial η², computed from the F test as df1 × F / (df1 × F + df2). The common practice is to plug the test's own df1 and df2 into the noncentral F method above, converting with df1 + df2 + 1 in place of N. For one factor with four levels and n participants, df1 = 3 and df2 = 3(n − 1). Treat the result as an approximation: it carries over the assumptions of the uncorrected F test, including sphericity. Also remember that partial η² from a within-subjects design leaves out the variance between people, so it is usually larger than η² from a between-subjects study of the same effect; generalized eta squared (Olejnik & Algina, 2003) was proposed to make the two comparable.

For what F itself tells you, see How to interpret the F statistic in ANOVA; for what the interval promises, see What does 95% confidence mean?. More plain-language guides are on the DASS blog.

How to report eta squared in APA style (7th edition)

Report the F test, then η² with its interval and the confidence level. With the example above:

"Scores differed across the four groups, F(3, 56) = 3.17, p = .031, η² = .145, 90% CI [.007, .254]."

If you report a 95% interval instead, write "η² = .145, 95% CI [.000, .282]". η² cannot exceed 1, so it takes no leading zero. In the method section, say how the interval was obtained, for example "Confidence intervals for η² were computed by inverting the noncentral F distribution (Steiger, 2004)." For a repeated-measures design, label the statistic as partial η² (η²p) so readers do not compare it with between-subjects values.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.