Answer
How to interpret the F statistic and p value in a one-way ANOVA
The short answer
F is a ratio of two variance estimates: how much the group means spread out (between groups) divided by how much scores vary inside the groups (within groups). If all population means are equal, both estimate the same noise and F lands near 1. The p value is the area of the F distribution to the right of your F, so F and p carry the same information once the degrees of freedom are fixed: p < .05 exactly when F exceeds the critical F. A significant F says at least one mean differs, not which one.
The short answer
A one-way ANOVA tests the null hypothesis that every group has the same population mean. Its test statistic, F, compares two estimates of the noise in your data:
- Mean square between groups (MS between): how far the group means sit from the overall mean, scaled up by group size. It reflects noise plus any real differences between groups.
- Mean square within groups (MS within, also called mean square error or residual): the pooled variance of scores around their own group mean. It reflects noise only.
F = MS between / MS within. When the null is true, the numerator and denominator estimate the same variance, so F hovers around 1. When the group means really differ, the numerator grows and F gets larger. There is no fixed cut-off such as "F above 4 is big": how large F must be to count as evidence depends on its two degrees of freedom, which is why software turns it into a p value.
How F, the critical F and the p value fit together
If the null hypothesis and the usual assumptions hold (independent observations, normal scores within each group, equal variances), F follows an F distribution with two degrees of freedom: df₁ = k − 1 for k groups, and df₂ = N − k for N observations in total. Everything else follows from that distribution:
- The p value is the probability of an F at least as large as yours if the null were true: the area under the F(df₁, df₂) curve to the right of your value. In R that is
pf(F, df1, df2, lower.tail = FALSE). - The critical F is the value that cuts off the top α of that distribution. At α = .05 it is
qf(0.95, df1, df2)in R. You reject the null when your F is larger than it. - They are the same decision. For fixed degrees of freedom, a larger F always means a smaller p, so "F exceeds the critical value" and "p < .05" are two ways of saying the same thing. The p value is just more convenient: you do not need a table, and it tells you how far past the threshold you are.
- The test is one-sided in F but two-sided in the means. Only large F values count against the null, because any pattern of differences between means, in any direction, pushes F up. Very small F values are not evidence of anything except that the group means are unusually close together.
With two groups, F is exactly the square of the pooled t statistic; see t test vs ANOVA with two groups. This page focuses on three or more groups, where ANOVA is no longer interchangeable with a single t test.
See it in R and Python
The code simulates three groups of 20 with true means of 50, 54 and 58 and a common standard deviation of 8. It prints the ANOVA table, rebuilds F from the sums of squares by hand, converts F to a p value, finds the critical F, and then simulates 5,000 data sets in which the null is true to show where F falls when there is nothing to find.
R
set.seed(12398)
# Three groups of 20, true means 50, 54 and 58, common SD of 8
group <- factor(rep(c("A", "B", "C"), each = 20))
score <- c(rnorm(20, 50, 8), rnorm(20, 54, 8), rnorm(20, 58, 8))
round(tapply(score, group, mean), 2)
round(tapply(score, group, sd), 2)
# 1. The ANOVA table
fit <- aov(score ~ group)
summary(fit)
# 2. Rebuild F by hand: between-group variance over within-group variance
grand <- mean(score)
ss_b <- sum(20 * (tapply(score, group, mean) - grand)^2)
ss_w <- sum((score - ave(score, group))^2)
ms_b <- ss_b / 2 # df between = k - 1 = 2
ms_w <- ss_w / 57 # df within = N - k = 57
F_stat <- ms_b / ms_w
round(c(SS_between = ss_b, SS_within = ss_w, MS_between = ms_b,
MS_within = ms_w, F = F_stat), 2)
# 3. F and p carry the same information: p is the area to the right of F
p <- pf(F_stat, 2, 57, lower.tail = FALSE)
signif(p, 4)
# 4. Critical F at alpha = .05: reject H0 when F exceeds it
F_crit <- qf(0.95, 2, 57)
round(F_crit, 3)
F_stat > F_crit; p < 0.05
# 5. Under H0 (all means equal), F averages about df2 / (df2 - 2)
F_null <- replicate(5000, {
y <- rnorm(60, 50, 8)
summary(aov(y ~ group))[[1]]$`F value`[1]
})
round(c(mean_F = mean(F_null), theory = 57 / 55,
share_above_crit = mean(F_null > F_crit)), 3)
# Effect size: eta squared
round(ss_b / (ss_b + ss_w), 3)Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(12398)
# Three groups of 20, true means 50, 54 and 58, common SD of 8
groups = [rng.normal(m, 8, 20) for m in (50, 54, 58)]
print("means", [round(g.mean(), 2) for g in groups])
print("sds", [round(g.std(ddof=1), 2) for g in groups])
# 1. The ANOVA F test
res = stats.f_oneway(*groups)
print("F", round(res.statistic, 2), "p", f"{res.pvalue:.4g}")
# 2. Rebuild F by hand: between-group variance over within-group variance
score = np.concatenate(groups)
grand = score.mean()
ss_b = sum(20 * (g.mean() - grand) ** 2 for g in groups)
ss_w = sum(((g - g.mean()) ** 2).sum() for g in groups)
ms_b, ms_w = ss_b / 2, ss_w / 57 # df = k - 1 and N - k
F_stat = ms_b / ms_w
print("SS_b", round(ss_b, 2), "SS_w", round(ss_w, 2),
"MS_b", round(ms_b, 2), "MS_w", round(ms_w, 2), "F", round(F_stat, 2))
# 3. F and p carry the same information: p is the area to the right of F
p = stats.f.sf(F_stat, 2, 57)
print("p", f"{p:.4g}")
# 4. Critical F at alpha = .05: reject H0 when F exceeds it
F_crit = stats.f.ppf(0.95, 2, 57)
print("F_crit", round(F_crit, 3), F_stat > F_crit, p < 0.05)
# 5. Under H0 (all means equal), F averages about df2 / (df2 - 2)
F_null = np.array([stats.f_oneway(*rng.normal(50, 8, (3, 20))).statistic
for _ in range(5000)])
print("mean_F", round(F_null.mean(), 3), "theory", round(57 / 55, 3),
"share_above_crit", round(float(np.mean(F_null > F_crit)), 3))
# Effect size: eta squared
print("eta_sq", round(ss_b / (ss_b + ss_w), 3)) A B C
48.60 52.43 61.14
A B C
8.76 8.33 6.20
Df Sum Sq Mean Sq F value Pr(>F)
group 2 1652 825.9 13.42 1.67e-05 ***
Residuals 57 3507 61.5
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
SS_between SS_within MS_between MS_within F
1651.79 3507.29 825.90 61.53 13.42
[1] 1.673e-05
[1] 3.159
[1] TRUE
[1] TRUE
mean_F theory share_above_crit
1.033 1.036 0.047
[1] 0.32
The output above is from one seeded run of the R code. Python's random numbers differ from R's, so its simulated groups and F differ (its sample happens to show smaller differences, with F = 6.32 and p = .003), but the pattern is the same: the hand-built F matches the software's, the critical F is identical (3.159, since it involves no randomness), and under the null F again averages close to 1.036 and exceeds the critical value in about 5% of data sets.
Reading the output
- The table. The group means were 48.60, 52.43 and 61.14. The between-groups row has 2 degrees of freedom (3 groups − 1) and a mean square of 825.9; the residual row has 57 (60 − 3) and a mean square of 61.5. Their ratio is F = 13.42.
- By hand. The sums of squares are 1651.79 between and 3507.29 within; dividing each by its degrees of freedom and taking the ratio reproduces F = 13.42 exactly. Nothing in the ANOVA table is more than this arithmetic.
- F to p. The area of the F(2, 57) distribution to the right of 13.42 is 1.673 × 10⁻⁵, the same value
summary()printed in thePr(>F)column. - Critical F. At α = .05 the critical value for F(2, 57) is 3.159. Our F is well above it and our p is well below .05: the two checks agree, as they always will.
- What F looks like when there is nothing to find. Across 5,000 simulated data sets with equal population means, F averaged 1.033, close to the theoretical mean of df₂ / (df₂ − 2) = 57/55 ≈ 1.036, and it exceeded the critical value in 4.7% of them, close to the intended 5%.
- Effect size. Eta squared, the between-groups sum of squares as a share of the total, was .32: group membership accounts for about 32% of the variance in this sample. F and p say whether the differences are distinguishable from noise; η² says how large they are.
What a significant F does not tell you
The ANOVA F is an omnibus test. Rejecting the null means the data are hard to reconcile with all means being equal; it does not say which groups differ or by how much. In the example, group C clearly stands out, but the difference between A and B may or may not be distinguishable from noise. To find out, follow up with planned contrasts or a post hoc procedure such as Tukey's HSD (TukeyHSD(fit) in R), which compares every pair while controlling the familywise error rate. See Bonferroni correction: when to use it for why running many unadjusted pairwise t tests inflates false positives.
Two more cautions. A large F in a big sample can come from differences too small to matter, which is why the effect size belongs next to it (see why large samples make tiny effects significant). And a non-significant F does not show that the means are equal, only that this study could not tell them apart (see does a non-significant result support the null?). If group variances look clearly unequal, Welch's ANOVA (oneway.test() in R) is the safer choice. For help picking a test, try the test chooser; the DASS blog has more plain-language guides.
How to report a one-way ANOVA in APA style (7th edition)
Give F with both degrees of freedom in parentheses, then p and an effect size, and describe the group means (in the text or a table). With the example above:
"A one-way ANOVA showed that scores differed across the three groups, F(2, 57) = 13.42, p < .001, η² = .32. Mean scores were 48.60 (SD = 8.76) in group A, 52.43 (SD = 8.33) in group B, and 61.14 (SD = 6.20) in group C (n = 20 per group)."
Report p values below .001 as p < .001, round F to two decimals, and drop the leading zero on η² and p because neither can exceed 1. If you follow up with pairwise comparisons, name the procedure in the method section (for example "Pairwise differences were tested with Tukey's HSD") and report each comparison's mean difference with its adjusted p value or confidence interval.
Related tools and guides
- ANOVA power and sample size calculator
- APA 7 formatter for ANOVA
- Check a reported p-value
- t test vs ANOVA with two groups: are the assumptions really different?
- Bonferroni correction: when should you use it?
- What do p values and t values mean?
- What are degrees of freedom in statistics?
- Which statistical test should I use?
More answered questions
- Bonferroni correction: when should you use it, and what are the alternatives?
- Why does almost everything become statistically significant with a large sample?
- Is a non-significant result from a large study evidence for the null?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.