Answer

Paired vs independent t test: which should you use?

Inspired by a question on Cross Validated ·

t testhypothesis testingstudy design

The short answer

Use the paired t test when each value in one condition is linked by design to one value in the other: the same person twice, matched pairs, or two versions run on the same day or site. Use Welch's independent t test when the groups share nothing. Any shared source of variation counts as pairing, not just repeated measures. Ignore real pairing and the shared variation can bury the difference.

The short answer

The choice is decided by how the data were collected, not by which test gives the smaller p value. Ask one question: is each observation in condition A naturally matched to one specific observation in condition B? If yes, the data are paired. If the two sets of values could be shuffled within each group without losing any information, they are independent.

Textbooks usually illustrate pairing with repeated measurements on the same person, which leaves many people thinking that is the only valid use. It is not. What matters is that the two members of a pair share something that pushes both values up or down together. A busy Saturday raises revenue in both arms of an A/B test, just as a strong student scores well both before and after training.

Why pairing helps so much

The paired t test is simply a one-sample t test on the differences, B − A, within each pair. Anything the two members share (the person's ability, the day's footfall, the batch's quality) cancels in that subtraction. Formally, the variance of a difference is Var(A) + Var(B) − 2 Cov(A, B), so the stronger the positive correlation between the pair members, the smaller the noise that is left.

The independent test cannot remove that shared variation. It compares the two group means against the full spread of the raw values, which includes the gap between quiet weekdays and busy weekends. A small but steady lift is buried under the day-to-day swings. Worse, the independent test assumes the two samples are unrelated, and that assumption is simply false when the values come in correlated pairs.

See it in R and Python

The code simulates a 14-day test in which both versions run every day. Each day has its own level, and weekends are about 15 units busier. Version B adds a true lift of 1.5. The same data are analysed with an independent (Welch) test and with a paired test, and then 10,000 simulated experiments show how often each test detects the lift and how often it raises a false alarm when there is no lift.

R

set.seed(477843)

# 1. Fourteen days; each day has its own level (weekends are busier),
#    and versions A and B both run on every day
n <- 14
weekend <- rep(c(0, 0, 0, 0, 0, 1, 1), 2)
day_level <- 50 + 15 * weekend + rnorm(n, 0, 4)   # shared by A and B on a day
a <- round(day_level + rnorm(n, 0, 1.5), 1)
b <- round(day_level + 1.5 + rnorm(n, 0, 1.5), 1)  # B is 1.5 higher on average
d <- b - a
round(c(mean_a = mean(a), sd_a = sd(a), mean_b = mean(b), sd_b = sd(b),
        mean_diff = mean(d), sd_diff = sd(d), r_ab = cor(a, b)), 2)

show <- function(tt) round(c(t = unname(tt$statistic), df = unname(tt$parameter),
                             p = tt$p.value, lower = tt$conf.int[1], upper = tt$conf.int[2]), 4)
show(t.test(b, a))                  # ignores the days (Welch, independent)
show(t.test(b, a, paired = TRUE))   # pairs by day
show(t.test(d))                     # same thing: one-sample t test on differences
signif(t.test(b, a, paired = TRUE)$p.value, 3)
round(mean(d) / sd(d), 2)           # Cohen's d_z

# 2. Simulation: how often each test detects the 1.5 lift (10,000 runs)
#    and how often it rejects when there is no lift
sim <- function(lift, reps = 10000) {
  res <- replicate(reps, {
    lvl <- 50 + 15 * weekend + rnorm(n, 0, 4)
    x <- lvl + rnorm(n, 0, 1.5); y <- lvl + lift + rnorm(n, 0, 1.5)
    c(independent = t.test(y, x)$p.value < 0.05,
      paired      = t.test(y, x, paired = TRUE)$p.value < 0.05)
  })
  round(rowMeans(res), 3)
}
sim(1.5)   # power
sim(0)     # false positive rate

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(477843)

# 1. Fourteen days; each day has its own level (weekends are busier),
#    and versions A and B both run on every day
n = 14
weekend = np.tile([0, 0, 0, 0, 0, 1, 1], 2)
day_level = 50 + 15 * weekend + rng.normal(0, 4, n)   # shared by A and B on a day
a = np.round(day_level + rng.normal(0, 1.5, n), 1)
b = np.round(day_level + 1.5 + rng.normal(0, 1.5, n), 1)  # B is 1.5 higher on average
d = b - a
print("mean_a", round(a.mean(), 2), "sd_a", round(a.std(ddof=1), 2),
      "mean_b", round(b.mean(), 2), "sd_b", round(b.std(ddof=1), 2),
      "mean_diff", round(d.mean(), 2), "sd_diff", round(d.std(ddof=1), 2),
      "r_ab", round(np.corrcoef(a, b)[0, 1], 2))

def show(res):
    ci = res.confidence_interval(0.95)
    print("t", round(res.statistic, 4), "df", round(res.df, 4), "p", round(res.pvalue, 4),
          "CI", round(ci.low, 4), round(ci.high, 4))

show(stats.ttest_ind(b, a, equal_var=False))   # ignores the days (Welch, independent)
show(stats.ttest_rel(b, a))                    # pairs by day
show(stats.ttest_1samp(d, 0))                  # same thing: one-sample t test on differences
print("d_z", round(d.mean() / d.std(ddof=1), 2))

# 2. Simulation: how often each test detects the 1.5 lift (10,000 runs)
#    and how often it rejects when there is no lift
def sim(lift, reps=10000):
    lvl = 50 + 15 * weekend + rng.normal(0, 4, size=(reps, n))
    x = lvl + rng.normal(0, 1.5, size=(reps, n))
    y = lvl + lift + rng.normal(0, 1.5, size=(reps, n))
    indep = stats.ttest_ind(y, x, axis=1, equal_var=False).pvalue < 0.05
    paired = stats.ttest_rel(y, x, axis=1).pvalue < 0.05
    return {"independent": round(indep.mean(), 3), "paired": round(paired.mean(), 3)}

print(sim(1.5))   # power
print(sim(0))     # false positive rate

The figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated data differ too. In its single 14-day sample the paired test gives p = .118, a reminder that one short experiment can miss a real lift. Its simulated rates (69.1% power, 5.1% false alarms for the paired test, 0% for the independent test) show the same pattern as R's.

When pairing is legitimate, and when it is not

You do not need to remove trends or cycles before testing. Taking the within-day difference already removes anything the two versions share on that day. If you have more than two conditions, or want to add covariates, the same idea becomes a repeated-measures ANOVA or a regression with day as a blocking factor; with exactly two conditions per day, that model gives the same answer as the paired t test. If the differences look skewed, a sign-flip permutation test or the Wilcoxon signed-rank test on the differences keeps the pairing (see Wilcoxon signed-rank test assumptions). For truly independent groups, use Welch's version of the independent test (see Welch vs Student t test).

Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.

How to report a paired t test in APA style (7th edition)

Say what formed the pairs, give the means and standard deviations of each condition, the mean difference with its confidence interval, the test result and an effect size. Using the example above:

"Daily revenue was higher under version B (M = 54.80, SD = 8.90) than under version A (M = 52.14, SD = 8.91). A paired-samples t test with days as pairs (n = 14) showed that the mean difference of 2.66, 95% CI [0.91, 4.40], was statistically significant, t(13) = 3.29, p = .006, d_z = 0.88."

Name the standardiser for the effect size: d_z divides the mean difference by the SD of the differences, and it is usually much larger than a d based on the raw SDs, so readers need to know which one you used (see How to interpret Cohen's d). In the method section, one sentence is enough: "Because both versions ran on each day, conditions were compared with a paired-samples t test on the daily differences."

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.