Answer

How are alpha and beta (Type I and Type II error rates) related?

Inspired by a question on Cross Validated ·

hypothesis testingpower analysissample sizet test

The short answer

Lowering alpha always raises beta for a fixed design, but not in a straight line and not at a fixed ratio. How much beta grows depends on the effect size and the sample size. In one example, cutting alpha from .05 to .01 raised beta from .52 to .76. A Bonferroni correction therefore costs power on every test, and the only way to get it back is a larger sample or a larger effect.

The short answer

Alpha (α) is the Type I error rate: the probability of rejecting the null hypothesis when it is true. Beta (β) is the Type II error rate: the probability of failing to reject the null when a real effect of a given size exists. Power is 1 − β.

For a fixed study (same test, same sample size, same true effect), the two pull against each other. A smaller α moves the cut-off for significance further out, so fewer real effects clear it and β goes up. But the trade-off is curved, not linear, and the ratio α/β is not constant. Beta is not a property of the test alone: it depends on the true effect size and the sample size, while α is chosen by you.

See it in R and Python

The code uses a two-sided independent-samples t test. It computes β exactly for several values of α with 30 per group and a medium effect (d = 0.5), repeats this for 100 per group and for a large effect (d = 0.8), finds the sample size needed for 80% power at α = .05 and at α = .01, applies a large-sample shortcut, and finally checks the exact figures by simulation.

R

set.seed(59202)

# 1. Beta (Type II error rate) for a two-sample t test, d = 0.5, 30 per group
alphas <- c(0.10, 0.05, 0.01, 0.005, 0.001)
beta_at <- function(a, n = 30, d = 0.5) 1 - power.t.test(n = n, delta = d, sd = 1, sig.level = a)$power
beta <- sapply(alphas, beta_at)
round(rbind(alpha = alphas, beta = beta, ratio = alphas / beta), 3)

# 2. Same alphas with a larger sample and with a larger effect
round(rbind(alpha = alphas,
            n100_d0.5 = sapply(alphas, beta_at, n = 100),
            n30_d0.8  = sapply(alphas, beta_at, d = 0.8)), 3)

# 3. Bonferroni for 5 tests: n per group for 80% power at .05 vs .01
ceiling(sapply(c(a.05 = 0.05, a.01 = 0.01), function(a)
  power.t.test(delta = 0.5, sd = 1, power = 0.8, sig.level = a)$n))

# 4. Large-sample shortcut: a test with 80% power at alpha = .05, re-run at .01
shift <- qnorm(0.975) + qnorm(0.80)        # implied signal in z units
round(c(shift = shift, power_at_.01 = pnorm(shift - qnorm(0.995))), 3)

# 5. Simulation check: share of p < alpha with no effect and with d = 0.5
reps <- 10000
p_null <- replicate(reps, t.test(rnorm(30), rnorm(30), var.equal = TRUE)$p.value)
p_alt  <- replicate(reps, t.test(rnorm(30, 0.5), rnorm(30), var.equal = TRUE)$p.value)
round(rbind(type1 = c(a.05 = mean(p_null < 0.05), a.01 = mean(p_null < 0.01)),
            type2 = c(mean(p_alt >= 0.05), mean(p_alt >= 0.01))), 3)

Python

import numpy as np
from scipy import stats, optimize

rng = np.random.default_rng(59202)

# Power of a two-sided two-sample t test (same formula as R's power.t.test)
def power(n, d, alpha):
    df = 2 * n - 2
    ncp = d * np.sqrt(n / 2)
    return float(stats.nct.sf(stats.t.ppf(1 - alpha / 2, df), df, ncp))

# 1. Beta (Type II error rate) for d = 0.5, 30 per group
alphas = [0.10, 0.05, 0.01, 0.005, 0.001]
beta = [1 - power(30, 0.5, a) for a in alphas]
print("beta ", [round(b, 3) for b in beta])
print("ratio", [round(a / b, 3) for a, b in zip(alphas, beta)])

# 2. Same alphas with a larger sample and with a larger effect
print("n100_d0.5", [round(1 - power(100, 0.5, a), 3) for a in alphas])
print("n30_d0.8 ", [round(1 - power(30, 0.8, a), 3) for a in alphas])

# 3. Bonferroni for 5 tests: n per group for 80% power at .05 vs .01
for a in (0.05, 0.01):
    n_needed = optimize.brentq(lambda n: power(n, 0.5, a) - 0.8, 2, 1000)
    print(f"a{a}", int(np.ceil(n_needed)))

# 4. Large-sample shortcut: a test with 80% power at alpha = .05, re-run at .01
shift = stats.norm.ppf(0.975) + stats.norm.ppf(0.80)  # implied signal in z units
print("shift", round(float(shift), 3), "power_at_.01", round(float(stats.norm.cdf(shift - stats.norm.ppf(0.995))), 3))

# 5. Simulation check: share of p < alpha with no effect and with d = 0.5
reps = 10000
p_null = stats.ttest_ind(rng.normal(size=(reps, 30)), rng.normal(size=(reps, 30)), axis=1).pvalue
p_alt = stats.ttest_ind(rng.normal(0.5, size=(reps, 30)), rng.normal(size=(reps, 30)), axis=1).pvalue
print("type1", round(float(np.mean(p_null < 0.05)), 3), round(float(np.mean(p_null < 0.01)), 3))
print("type2", round(float(np.mean(p_alt >= 0.05)), 3), round(float(np.mean(p_alt >= 0.01)), 3))
       [,1]  [,2]  [,3]  [,4]  [,5]
alpha 0.100 0.050 0.010 0.005 0.001
beta  0.394 0.522 0.756 0.825 0.925
ratio 0.254 0.096 0.013 0.006 0.001
           [,1]  [,2]  [,3]  [,4]  [,5]
alpha     0.100 0.050 0.010 0.005 0.001
n100_d0.5 0.030 0.060 0.176 0.244 0.422
n30_d0.8  0.078 0.139 0.332 0.426 0.631
a.05 a.01 
  64   96 
       shift power_at_.01 
       2.802        0.589 
       a.05  a.01
type1 0.048 0.010
type2 0.520 0.755

The simulated rates come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated rates differ slightly (for example a Type I rate of 5.4% instead of 4.8% at α = .05), but they show the same pattern. Everything else involves no randomness, and Python reproduces it exactly.

What the numbers show

Why curved? For a large-sample test, power is roughly the normal probability that the test statistic lands beyond the critical value. Making α smaller pushes the critical value out along the tail of a bell curve, and the area you lose depends on where the true effect sits relative to that cut-off. That is a nonlinear function, so no single ratio or slope describes it.

What a Bonferroni correction does to power

A Bonferroni correction for five tests divides α by five, so each test is run at .01 instead of .05. That keeps the chance of at least one false positive across the five tests at or below .05. The price is paid in β on every test: the analysis does not just trade one error for the other at a fixed rate.

If all you know is that a test had 80% power at α = .05, you can still estimate its power at .01 with the large-sample shortcut: 80% power at .05 implies a signal of about 2.80 standard errors, and at α = .01 that gives power of about .589, so β rises from .20 to roughly .41. The exact figure for a particular t test depends on its degrees of freedom, but with a reasonable sample size this shortcut is close.

To keep 80% power for a medium effect after the correction, the example needs 96 per group instead of 64, which is 50% more participants. Whether that is worth it depends on how costly a false positive is compared with a missed effect. When the tests are confirmatory and each one matters, plan the sample for the corrected α from the start. A power analysis is where this decision belongs.

A note on the diagnostic-test analogy: 1 − α behaves like specificity and power like sensitivity, but only for one assumed true effect. A real collection of tests mixes true nulls and real effects of different sizes, so its observed true-positive and true-negative rates also depend on that mix, not on α alone. Also remember that failing to reject is not proof of no effect (see Does a non-significant result support the null?). For more plain-language guides, see the DASS blog, or use the test chooser.

How to report statistical power in APA style (7th edition)

Alpha, power and the effect size the study was planned for belong in the method section, usually where you justify the sample size. Using the numbers from the example above:

"To control the familywise error rate across five comparisons, we used a Bonferroni-corrected significance level of α = .01 for each test. An a priori power analysis for a two-sided independent-samples t test showed that 96 participants per group were needed to detect a medium effect (d = 0.50) with 80% power at this level, compared with 64 per group at α = .05."

If the sample was fixed in advance, report the power you actually had instead: "With 30 participants per group and α = .01, the study had 24% power to detect a medium effect (d = 0.50)." Replace the values with your own, and name the software used for the calculation.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.