Answer

How to interpret the F statistic and p value in a one-way ANOVA

Inspired by a question on Cross Validated ·

anovahypothesis testingp-values

The short answer

F is a ratio of two variance estimates: how much the group means spread out (between groups) divided by how much scores vary inside the groups (within groups). If all population means are equal, both estimate the same noise and F lands near 1. The p value is the area of the F distribution to the right of your F, so F and p carry the same information once the degrees of freedom are fixed: p < .05 exactly when F exceeds the critical F. A significant F says at least one mean differs, not which one.

The short answer

A one-way ANOVA tests the null hypothesis that every group has the same population mean. Its test statistic, F, compares two estimates of the noise in your data:

F = MS between / MS within. When the null is true, the numerator and denominator estimate the same variance, so F hovers around 1. When the group means really differ, the numerator grows and F gets larger. There is no fixed cut-off such as "F above 4 is big": how large F must be to count as evidence depends on its two degrees of freedom, which is why software turns it into a p value.

How F, the critical F and the p value fit together

If the null hypothesis and the usual assumptions hold (independent observations, normal scores within each group, equal variances), F follows an F distribution with two degrees of freedom: df₁ = k − 1 for k groups, and df₂ = N − k for N observations in total. Everything else follows from that distribution:

With two groups, F is exactly the square of the pooled t statistic; see t test vs ANOVA with two groups. This page focuses on three or more groups, where ANOVA is no longer interchangeable with a single t test.

See it in R and Python

The code simulates three groups of 20 with true means of 50, 54 and 58 and a common standard deviation of 8. It prints the ANOVA table, rebuilds F from the sums of squares by hand, converts F to a p value, finds the critical F, and then simulates 5,000 data sets in which the null is true to show where F falls when there is nothing to find.

R

set.seed(12398)

# Three groups of 20, true means 50, 54 and 58, common SD of 8
group <- factor(rep(c("A", "B", "C"), each = 20))
score <- c(rnorm(20, 50, 8), rnorm(20, 54, 8), rnorm(20, 58, 8))
round(tapply(score, group, mean), 2)
round(tapply(score, group, sd), 2)

# 1. The ANOVA table
fit <- aov(score ~ group)
summary(fit)

# 2. Rebuild F by hand: between-group variance over within-group variance
grand  <- mean(score)
ss_b   <- sum(20 * (tapply(score, group, mean) - grand)^2)
ss_w   <- sum((score - ave(score, group))^2)
ms_b   <- ss_b / 2      # df between = k - 1 = 2
ms_w   <- ss_w / 57     # df within  = N - k = 57
F_stat <- ms_b / ms_w
round(c(SS_between = ss_b, SS_within = ss_w, MS_between = ms_b,
        MS_within = ms_w, F = F_stat), 2)

# 3. F and p carry the same information: p is the area to the right of F
p <- pf(F_stat, 2, 57, lower.tail = FALSE)
signif(p, 4)

# 4. Critical F at alpha = .05: reject H0 when F exceeds it
F_crit <- qf(0.95, 2, 57)
round(F_crit, 3)
F_stat > F_crit; p < 0.05

# 5. Under H0 (all means equal), F averages about df2 / (df2 - 2)
F_null <- replicate(5000, {
  y <- rnorm(60, 50, 8)
  summary(aov(y ~ group))[[1]]$`F value`[1]
})
round(c(mean_F = mean(F_null), theory = 57 / 55,
        share_above_crit = mean(F_null > F_crit)), 3)

# Effect size: eta squared
round(ss_b / (ss_b + ss_w), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(12398)

# Three groups of 20, true means 50, 54 and 58, common SD of 8
groups = [rng.normal(m, 8, 20) for m in (50, 54, 58)]
print("means", [round(g.mean(), 2) for g in groups])
print("sds", [round(g.std(ddof=1), 2) for g in groups])

# 1. The ANOVA F test
res = stats.f_oneway(*groups)
print("F", round(res.statistic, 2), "p", f"{res.pvalue:.4g}")

# 2. Rebuild F by hand: between-group variance over within-group variance
score = np.concatenate(groups)
grand = score.mean()
ss_b = sum(20 * (g.mean() - grand) ** 2 for g in groups)
ss_w = sum(((g - g.mean()) ** 2).sum() for g in groups)
ms_b, ms_w = ss_b / 2, ss_w / 57        # df = k - 1 and N - k
F_stat = ms_b / ms_w
print("SS_b", round(ss_b, 2), "SS_w", round(ss_w, 2),
      "MS_b", round(ms_b, 2), "MS_w", round(ms_w, 2), "F", round(F_stat, 2))

# 3. F and p carry the same information: p is the area to the right of F
p = stats.f.sf(F_stat, 2, 57)
print("p", f"{p:.4g}")

# 4. Critical F at alpha = .05: reject H0 when F exceeds it
F_crit = stats.f.ppf(0.95, 2, 57)
print("F_crit", round(F_crit, 3), F_stat > F_crit, p < 0.05)

# 5. Under H0 (all means equal), F averages about df2 / (df2 - 2)
F_null = np.array([stats.f_oneway(*rng.normal(50, 8, (3, 20))).statistic
                   for _ in range(5000)])
print("mean_F", round(F_null.mean(), 3), "theory", round(57 / 55, 3),
      "share_above_crit", round(float(np.mean(F_null > F_crit)), 3))

# Effect size: eta squared
print("eta_sq", round(ss_b / (ss_b + ss_w), 3))
    A     B     C 
48.60 52.43 61.14 
   A    B    C 
8.76 8.33 6.20 
            Df Sum Sq Mean Sq F value   Pr(>F)    
group        2   1652   825.9   13.42 1.67e-05 ***
Residuals   57   3507    61.5                     
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
SS_between  SS_within MS_between  MS_within          F 
   1651.79    3507.29     825.90      61.53      13.42 
[1] 1.673e-05
[1] 3.159
[1] TRUE
[1] TRUE
          mean_F           theory share_above_crit 
           1.033            1.036            0.047 
[1] 0.32

The output above is from one seeded run of the R code. Python's random numbers differ from R's, so its simulated groups and F differ (its sample happens to show smaller differences, with F = 6.32 and p = .003), but the pattern is the same: the hand-built F matches the software's, the critical F is identical (3.159, since it involves no randomness), and under the null F again averages close to 1.036 and exceeds the critical value in about 5% of data sets.

Reading the output

What a significant F does not tell you

The ANOVA F is an omnibus test. Rejecting the null means the data are hard to reconcile with all means being equal; it does not say which groups differ or by how much. In the example, group C clearly stands out, but the difference between A and B may or may not be distinguishable from noise. To find out, follow up with planned contrasts or a post hoc procedure such as Tukey's HSD (TukeyHSD(fit) in R), which compares every pair while controlling the familywise error rate. See Bonferroni correction: when to use it for why running many unadjusted pairwise t tests inflates false positives.

Two more cautions. A large F in a big sample can come from differences too small to matter, which is why the effect size belongs next to it (see why large samples make tiny effects significant). And a non-significant F does not show that the means are equal, only that this study could not tell them apart (see does a non-significant result support the null?). If group variances look clearly unequal, Welch's ANOVA (oneway.test() in R) is the safer choice. For help picking a test, try the test chooser; the DASS blog has more plain-language guides.

How to report a one-way ANOVA in APA style (7th edition)

Give F with both degrees of freedom in parentheses, then p and an effect size, and describe the group means (in the text or a table). With the example above:

"A one-way ANOVA showed that scores differed across the three groups, F(2, 57) = 13.42, p < .001, η² = .32. Mean scores were 48.60 (SD = 8.76) in group A, 52.43 (SD = 8.33) in group B, and 61.14 (SD = 6.20) in group C (n = 20 per group)."

Report p values below .001 as p < .001, round F to two decimals, and drop the leading zero on η² and p because neither can exceed 1. If you follow up with pairwise comparisons, name the procedure in the method section (for example "Pairwise differences were tested with Tukey's HSD") and report each comparison's mean difference with its adjusted p value or confidence interval.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.