Answer

Is the chi-square test one-tailed or two-tailed?

Inspired by a question on Cross Validated ·

chi-squarehypothesis testing

The short answer

The usual chi-square test looks only at large values of its statistic, but because it squares every deviation, a large value can come from a difference in any direction. Its p value is already the non-directional answer, so never halve it. A directional test is possible only for a 2x2 table, as a planned one-sided z test of two proportions. The chi-square test for one variance is different: there, unusually small values count as evidence too.

The short answer

The confusion comes from two different meanings of "tail". One is about the distribution: which end of the chi-square curve holds the rejection region. The other is about the hypothesis: whether you are looking for a difference in one stated direction or in any direction.

For the usual chi-square tests (independence in a contingency table, and goodness of fit), the rejection region sits only in the upper tail of the distribution. But the hypothesis being tested is non-directional: any pattern that departs from what the null predicts pushes the statistic up. So the test is one-tailed in the first sense and two-sided (or, more precisely, omnibus) in the second. The practical rule follows directly: the p value your software reports is already the complete answer, so never halve it to get a "one-sided" p value.

Why squaring makes the test non-directional

The Pearson statistic adds up (observed - expected)² / expected over all cells. Because each gap is squared, a cell with too many cases and a cell with too few both add a positive amount. If group A passes more often than group B, the statistic grows; if B passes more often than A, it grows by exactly the same amount. The sign of the difference is thrown away before the statistic is formed.

A small chi-square value therefore means "close to what independence predicts", and a large one means "far from it, in some way". Evidence against the null only ever shows up as a large value, which is why the p value is the area to the right of the observed statistic. The same holds for the F test in ANOVA: it compares variation between group means with variation within groups, any spread of the means makes F large, and its upper-tail p value already covers every ordering of the groups.

What if you predicted which group would be higher?

In a 2x2 table, the chi-square statistic (without continuity correction) is exactly the square of the z statistic for comparing two proportions, and its p value equals the two-sided z test p value. The z statistic keeps the sign, so it can be tested in one direction. If you decided before seeing the data that only a higher rate in one particular group matters, a one-sided z test is legitimate, and its p value is half the chi-square p value when the difference points the predicted way (and above .5 when it does not).

That is different from taking a chi-square p value and dividing it by two after the fact. Halving throws away the check on direction, so it treats a difference either way as support, and it doubles the false-positive rate. For tables larger than 2x2 there is no single direction to test: the alternative covers many possible patterns at once, so a one-sided version does not exist. If you have an ordered prediction there (for example, rates rising across three ordered groups), use a test built for it, such as a trend test, rather than halving an omnibus p value. The general case for and against one-tailed tests is covered in One-tailed vs two-tailed tests.

Chi-square tests where small values count too

See it in R and Python

The code runs a chi-square test on a 2x2 table of pass/fail counts for two study-skills programs (80 students each), repeats it as a z test of two proportions, simulates 10,000 studies with no true difference to compare three decision rules, and finally runs a chi-square test for one variance on 15 fill weights, where the lower tail is what matters.

R

set.seed(22347)

# 1. A 2x2 table: pass/fail for two study-skills programs (80 students each)
tab <- matrix(c(48, 32,
                34, 46), nrow = 2, byrow = TRUE,
              dimnames = list(program = c("A", "B"), result = c("pass", "fail")))
ct <- chisq.test(tab, correct = FALSE)
round(c(chisq = unname(ct$statistic), p = ct$p.value), 4)

# The same comparison as a z test of two proportions
p1 <- 48 / 80; p2 <- 34 / 80; pool <- 82 / 160
z <- (p1 - p2) / sqrt(pool * (1 - pool) * (1 / 80 + 1 / 80))
round(c(z = z, z_squared = z^2,
        p_two_sided = 2 * pnorm(-abs(z)),
        p_one_sided_A_higher = pnorm(z, lower.tail = FALSE)), 4)
round(c(diff = p1 - p2), 3)

# 2. Null true (both pass rates 0.5): how often is each rule "significant"?
sims <- replicate(10000, {
  a <- rbinom(1, 80, 0.5); b <- rbinom(1, 80, 0.5)
  pa <- a / 80; pb <- b / 80; pp <- (a + b) / 160
  zz <- (pa - pb) / sqrt(pp * (1 - pp) * (2 / 80))
  p_chi <- pchisq(zz^2, df = 1, lower.tail = FALSE)
  c(chi = p_chi, halved = p_chi / 2, planned = pnorm(zz, lower.tail = FALSE))
})
round(rowMeans(sims < 0.05), 3)

# 3. A chi-square test that does use both tails: is a machine's variance 4?
x <- c(49.1, 50.8, 50.2, 48.7, 51.0, 49.6, 50.4, 49.9,
       51.3, 48.9, 50.6, 49.4, 50.1, 50.9, 49.2)   # 15 fill weights
stat <- (15 - 1) * var(x) / 4
round(c(p_upper_only = pchisq(stat, 14, lower.tail = FALSE)), 4)
round(c(var = var(x), stat = stat,
        p_lower = pchisq(stat, 14),
        p_two_sided = 2 * min(pchisq(stat, 14), pchisq(stat, 14, lower.tail = FALSE))), 4)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(22347)

# 1. A 2x2 table: pass/fail for two study-skills programs (80 students each)
tab = np.array([[48, 32],
                [34, 46]])
chi2, p, dof, expected = stats.chi2_contingency(tab, correction=False)
print("chisq", round(chi2, 4), "p", round(p, 4))

# The same comparison as a z test of two proportions
p1, p2, pool = 48 / 80, 34 / 80, 82 / 160
z = (p1 - p2) / np.sqrt(pool * (1 - pool) * (1 / 80 + 1 / 80))
print({"z": round(z, 4), "z_squared": round(z**2, 4),
       "p_two_sided": round(2 * stats.norm.sf(abs(z)), 4),
       "p_one_sided_A_higher": round(stats.norm.sf(z), 4)})
print("diff", round(p1 - p2, 3))

# 2. Null true (both pass rates 0.5): how often is each rule "significant"?
a = rng.binomial(80, 0.5, size=10000)
b = rng.binomial(80, 0.5, size=10000)
pa, pb, pp = a / 80, b / 80, (a + b) / 160
zz = (pa - pb) / np.sqrt(pp * (1 - pp) * (2 / 80))
p_chi = stats.chi2.sf(zz**2, df=1)
print({"chi": round(np.mean(p_chi < 0.05), 3),
       "halved": round(np.mean(p_chi / 2 < 0.05), 3),
       "planned": round(np.mean(stats.norm.sf(zz) < 0.05), 3)})

# 3. A chi-square test that does use both tails: is a machine's variance 4?
x = np.array([49.1, 50.8, 50.2, 48.7, 51.0, 49.6, 50.4, 49.9,
              51.3, 48.9, 50.6, 49.4, 50.1, 50.9, 49.2])   # 15 fill weights
stat = (15 - 1) * x.var(ddof=1) / 4
print("p_upper_only", round(stats.chi2.sf(stat, 14), 4))
print({"var": round(x.var(ddof=1), 4), "stat": round(stat, 4),
       "p_lower": round(stats.chi2.cdf(stat, 14), 4),
       "p_two_sided": round(2 * min(stats.chi2.cdf(stat, 14), stats.chi2.sf(stat, 14)), 4)})

The figures below come from one seeded run of the R code. Parts 1 and 3 involve no randomness, and the Python version reproduces them exactly. Python's random numbers differ from R's, so its simulated rates in part 2 differ slightly (5.1%, 10.0% and 5.2%), but they show the same pattern.

How to report a chi-square test in APA style (7th edition)

Give the degrees of freedom and the sample size in parentheses, the statistic to two decimals, and the p value as reported, without halving it. Describe the direction with the percentages, not with the test. Using the example above:

"Pass rates differed between the programs, χ²(1, N = 160) = 4.90, p = .027: 60.0% of students in Program A passed, compared with 42.5% in Program B."

If a directional hypothesis was set in advance for a 2x2 table, report the one-sided z test instead and say so: "As preregistered, a one-sided test of two proportions showed a higher pass rate in Program A, z = 2.21, p = .013 (one-tailed)." Put the full table of counts and percentages in a table when there are more than a few cells. Not sure which test fits your design? Try the test chooser, or see the DASS blog for more on planning and reporting analyses.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.