Answer

One-tailed vs two-tailed tests: why not just test the direction the data point to?

Inspired by a question on Cross Validated ·

hypothesis testingp-valuest test

The short answer

The 5% in a significance test is a property of the whole procedure, not of one result. If you look at the data first and then test in whichever direction the effect went, you reject whenever either tail is extreme, so your real false-positive rate is about 10%, not 5%. A one-tailed test is fine only when the direction is fixed before the data are seen and a reversed effect would lead to the same conclusion as no effect.

The short answer

Once you have a sample, you can of course see which way the effect went. The question a test answers is different: if there were no effect at all, how often would a procedure like mine declare one? That rate, alpha, is what the 5% promises, and it depends on every result that would have led you to claim an effect, not only on the result you happened to get.

A two-tailed test at 5% claims an effect when the result lands in the top 2.5% or the bottom 2.5% of what chance alone produces. A one-tailed test at 5% puts all 5% in one tail that was chosen in advance. If instead you pick the tail after looking, you are effectively claiming an effect when the result lands in the top 5% or the bottom 5%. That is a 10% procedure dressed up as a 5% one.

Why picking the tail afterwards is not the same as picking it first

It can feel as if the two routes give the same answer: in both, you end up running a one-tailed test in the direction of the effect. The difference is in the other half of the possible outcomes. With a direction fixed in advance, a large effect the "wrong" way is simply not significant. With a direction chosen afterwards, that same outcome becomes significant too, because you would have switched hypotheses. Counting both tails, chance hands you a significant result twice as often.

The same logic explains why you should not switch to a one-tailed test after a two-tailed one narrowly misses. Whenever the effect points in some direction, the one-tailed p value in that direction is exactly half the two-tailed one. Switching after the fact therefore halves every near-miss, which again means a real error rate closer to 10% than 5%. The threshold of 5% is a convention, but whatever threshold you use, it only means something if the rule for reaching it was fixed before the data arrived.

When a one-tailed test is legitimate

A one-tailed test is not cheating. It is the right tool when both of these hold:

The price is real: a one-tailed test can never report the opposite effect as significant, however large it is. If a reversed effect would be interesting, surprising or worrying (a teaching method that hurts learning, say), you want to be able to detect it, and a two-tailed test is the honest choice. This is why reviewers are often sceptical of one-tailed tests that were not justified in advance, and why two-tailed tests are the default in most fields.

See it in R and Python

The code does three things: it runs a one-sample t test on 20 paired differences with each of the three alternatives, simulates 10,000 studies with no true effect to count false positives under each strategy, and simulates the power of one- and two-tailed tests when there is a real effect of 0.5 standard deviations in each direction.

R

set.seed(347727)

# 1. One data set, three alternatives: 20 paired differences (after - before)
d <- c(1.2, -0.4, 0.9, 0.3, -1.1, 1.6, 0.5, -0.2, 0.8, 1.4,
       -0.6, 0.7, 0.1, 1.9, -0.9, 0.4, 1.1, -0.3, 0.6, -0.8)
round(c(mean = mean(d), sd = sd(d)), 2)
tt <- t.test(d, mu = 0)                       # two-tailed
round(unname(tt$statistic), 2)
round(c(two_tailed = tt$p.value,
        greater    = t.test(d, alternative = "greater")$p.value,
        less       = t.test(d, alternative = "less")$p.value), 4)
round(as.numeric(tt$conf.int), 2)

# 2. Type I error when the null is true (10,000 studies of n = 20)
p_greater <- replicate(10000, {
  x <- rnorm(20)                              # no effect at all
  t.test(x, alternative = "greater")$p.value
})
p_two  <- 2 * pmin(p_greater, 1 - p_greater)  # two-tailed p value
p_pick <- pmin(p_greater, 1 - p_greater)      # tail chosen after seeing the data
round(c(two_tailed      = mean(p_two < 0.05),
        one_tailed_plan = mean(p_greater < 0.05),
        tail_picked     = mean(p_pick < 0.05)), 3)

# 3. Power with a true effect of 0.5 SD (n = 20)
sim_p <- function(delta) replicate(10000,
  t.test(rnorm(20, mean = delta), alternative = "greater")$p.value)
p_up   <- sim_p(0.5)                          # effect in the predicted direction
p_down <- sim_p(-0.5)                         # effect in the other direction
round(c(two_tailed = mean(2 * pmin(p_up, 1 - p_up) < 0.05),
        one_tailed = mean(p_up < 0.05)), 3)
round(c(two_tailed = mean(2 * pmin(p_down, 1 - p_down) < 0.05),
        one_tailed = mean(p_down < 0.05)), 3)

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(347727)

# 1. One data set, three alternatives: 20 paired differences (after - before)
d = np.array([1.2, -0.4, 0.9, 0.3, -1.1, 1.6, 0.5, -0.2, 0.8, 1.4,
              -0.6, 0.7, 0.1, 1.9, -0.9, 0.4, 1.1, -0.3, 0.6, -0.8])
print("mean", round(d.mean(), 2), "sd", round(d.std(ddof=1), 2))
tt = stats.ttest_1samp(d, 0)                    # two-tailed
print("t", round(tt.statistic, 2))
print({alt: round(stats.ttest_1samp(d, 0, alternative=alt).pvalue, 4)
       for alt in ["two-sided", "greater", "less"]})
print("95% CI", np.round(tt.confidence_interval(0.95), 2))

# 2. Type I error when the null is true (10,000 studies of n = 20)
x = rng.normal(size=(10000, 20))                # no effect at all
p_greater = stats.ttest_1samp(x, 0, axis=1, alternative="greater").pvalue
p_two = 2 * np.minimum(p_greater, 1 - p_greater)   # two-tailed p value
p_pick = np.minimum(p_greater, 1 - p_greater)      # tail chosen after seeing the data
print({"two_tailed": round(np.mean(p_two < 0.05), 3),
       "one_tailed_plan": round(np.mean(p_greater < 0.05), 3),
       "tail_picked": round(np.mean(p_pick < 0.05), 3)})

# 3. Power with a true effect of 0.5 SD (n = 20)
def sim_p(delta):
    x = rng.normal(loc=delta, size=(10000, 20))
    return stats.ttest_1samp(x, 0, axis=1, alternative="greater").pvalue

for delta in [0.5, -0.5]:                       # predicted direction, then the other
    p = sim_p(delta)
    print(delta, {"two_tailed": round(np.mean(2 * np.minimum(p, 1 - p) < 0.05), 3),
                  "one_tailed": round(np.mean(p < 0.05), 3)})

The figures below come from one seeded run of the R code. Part 1 involves no randomness, and the Python version reproduces it exactly. Python's random numbers differ from R's, so its simulated rates differ slightly (for example 10.1% instead of 9.8% for the tail picked afterwards), but they show the same pattern.

What to do in practice

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report a one-tailed test in APA style (7th edition)

APA style expects two-tailed tests unless you say otherwise, so state clearly when a test is one-tailed and that the direction was set in advance. Using the example above, if a one-tailed test had been planned:

"As preregistered, we tested the directional hypothesis that scores would increase using a one-tailed paired-samples t test. The mean increase was 0.36 points (SD = 0.87), 95% CI [-0.05, 0.77], t(19) = 1.86, p = .040 (one-tailed)."

If no direction had been planned, the honest report is two-tailed: "The mean increase was 0.36 points (SD = 0.87), 95% CI [-0.05, 0.77], t(19) = 1.86, p = .079." In the method section, name the alpha level and whether tests were one- or two-tailed, and point to the preregistration if there is one.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.