Answer
Paired vs independent t test: which should you use?
The short answer
Use the paired t test when each value in one condition is linked by design to one value in the other: the same person twice, matched pairs, or two versions run on the same day or site. Use Welch's independent t test when the groups share nothing. Any shared source of variation counts as pairing, not just repeated measures. Ignore real pairing and the shared variation can bury the difference.
The short answer
The choice is decided by how the data were collected, not by which test gives the smaller p value. Ask one question: is each observation in condition A naturally matched to one specific observation in condition B? If yes, the data are paired. If the two sets of values could be shuffled within each group without losing any information, they are independent.
- Paired: before and after scores for the same students; left and right eye of each patient; twins split between two diets; two versions of a web page run on the same days, giving one A value and one B value per day.
- Independent: one group of students taught with method A and a different, unrelated group taught with method B; customers randomly split into two arms and analysed as individuals with nothing linking a given A customer to a given B customer.
Textbooks usually illustrate pairing with repeated measurements on the same person, which leaves many people thinking that is the only valid use. It is not. What matters is that the two members of a pair share something that pushes both values up or down together. A busy Saturday raises revenue in both arms of an A/B test, just as a strong student scores well both before and after training.
Why pairing helps so much
The paired t test is simply a one-sample t test on the differences, B − A, within each pair. Anything the two members share (the person's ability, the day's footfall, the batch's quality) cancels in that subtraction. Formally, the variance of a difference is Var(A) + Var(B) − 2 Cov(A, B), so the stronger the positive correlation between the pair members, the smaller the noise that is left.
The independent test cannot remove that shared variation. It compares the two group means against the full spread of the raw values, which includes the gap between quiet weekdays and busy weekends. A small but steady lift is buried under the day-to-day swings. Worse, the independent test assumes the two samples are unrelated, and that assumption is simply false when the values come in correlated pairs.
See it in R and Python
The code simulates a 14-day test in which both versions run every day. Each day has its own level, and weekends are about 15 units busier. Version B adds a true lift of 1.5. The same data are analysed with an independent (Welch) test and with a paired test, and then 10,000 simulated experiments show how often each test detects the lift and how often it raises a false alarm when there is no lift.
R
set.seed(477843)
# 1. Fourteen days; each day has its own level (weekends are busier),
# and versions A and B both run on every day
n <- 14
weekend <- rep(c(0, 0, 0, 0, 0, 1, 1), 2)
day_level <- 50 + 15 * weekend + rnorm(n, 0, 4) # shared by A and B on a day
a <- round(day_level + rnorm(n, 0, 1.5), 1)
b <- round(day_level + 1.5 + rnorm(n, 0, 1.5), 1) # B is 1.5 higher on average
d <- b - a
round(c(mean_a = mean(a), sd_a = sd(a), mean_b = mean(b), sd_b = sd(b),
mean_diff = mean(d), sd_diff = sd(d), r_ab = cor(a, b)), 2)
show <- function(tt) round(c(t = unname(tt$statistic), df = unname(tt$parameter),
p = tt$p.value, lower = tt$conf.int[1], upper = tt$conf.int[2]), 4)
show(t.test(b, a)) # ignores the days (Welch, independent)
show(t.test(b, a, paired = TRUE)) # pairs by day
show(t.test(d)) # same thing: one-sample t test on differences
signif(t.test(b, a, paired = TRUE)$p.value, 3)
round(mean(d) / sd(d), 2) # Cohen's d_z
# 2. Simulation: how often each test detects the 1.5 lift (10,000 runs)
# and how often it rejects when there is no lift
sim <- function(lift, reps = 10000) {
res <- replicate(reps, {
lvl <- 50 + 15 * weekend + rnorm(n, 0, 4)
x <- lvl + rnorm(n, 0, 1.5); y <- lvl + lift + rnorm(n, 0, 1.5)
c(independent = t.test(y, x)$p.value < 0.05,
paired = t.test(y, x, paired = TRUE)$p.value < 0.05)
})
round(rowMeans(res), 3)
}
sim(1.5) # power
sim(0) # false positive ratePython
import numpy as np
from scipy import stats
rng = np.random.default_rng(477843)
# 1. Fourteen days; each day has its own level (weekends are busier),
# and versions A and B both run on every day
n = 14
weekend = np.tile([0, 0, 0, 0, 0, 1, 1], 2)
day_level = 50 + 15 * weekend + rng.normal(0, 4, n) # shared by A and B on a day
a = np.round(day_level + rng.normal(0, 1.5, n), 1)
b = np.round(day_level + 1.5 + rng.normal(0, 1.5, n), 1) # B is 1.5 higher on average
d = b - a
print("mean_a", round(a.mean(), 2), "sd_a", round(a.std(ddof=1), 2),
"mean_b", round(b.mean(), 2), "sd_b", round(b.std(ddof=1), 2),
"mean_diff", round(d.mean(), 2), "sd_diff", round(d.std(ddof=1), 2),
"r_ab", round(np.corrcoef(a, b)[0, 1], 2))
def show(res):
ci = res.confidence_interval(0.95)
print("t", round(res.statistic, 4), "df", round(res.df, 4), "p", round(res.pvalue, 4),
"CI", round(ci.low, 4), round(ci.high, 4))
show(stats.ttest_ind(b, a, equal_var=False)) # ignores the days (Welch, independent)
show(stats.ttest_rel(b, a)) # pairs by day
show(stats.ttest_1samp(d, 0)) # same thing: one-sample t test on differences
print("d_z", round(d.mean() / d.std(ddof=1), 2))
# 2. Simulation: how often each test detects the 1.5 lift (10,000 runs)
# and how often it rejects when there is no lift
def sim(lift, reps=10000):
lvl = 50 + 15 * weekend + rng.normal(0, 4, size=(reps, n))
x = lvl + rng.normal(0, 1.5, size=(reps, n))
y = lvl + lift + rng.normal(0, 1.5, size=(reps, n))
indep = stats.ttest_ind(y, x, axis=1, equal_var=False).pvalue < 0.05
paired = stats.ttest_rel(y, x, axis=1).pvalue < 0.05
return {"independent": round(indep.mean(), 3), "paired": round(paired.mean(), 3)}
print(sim(1.5)) # power
print(sim(0)) # false positive rateThe figures below come from one seeded run of the R code. Python's random numbers differ from R's, so its simulated data differ too. In its single 14-day sample the paired test gives p = .118, a reminder that one short experiment can miss a real lift. Its simulated rates (69.1% power, 5.1% false alarms for the paired test, 0% for the independent test) show the same pattern as R's.
- The raw values vary a lot. Version A: M = 52.14 (SD = 8.91); version B: M = 54.80 (SD = 8.90). Most of that spread is the weekday/weekend cycle, and the A and B values on the same day correlate at r = .94.
- The daily differences barely vary. Their mean is 2.66 and their SD is only 3.02, about a third of the raw SD. (The true lift is 1.5; this sample happened to come out higher, and the interval below still contains 1.5.)
- Independent (Welch) test: t(26.00) = 0.79, p = .437, 95% CI [−4.26, 9.57]. The interval is so wide that it cannot tell a gain from a loss.
- Paired test: t(13) = 3.29, p = .006, 95% CI [0.91, 4.40]. Running a one-sample t test on the 14 differences gives exactly the same result.
- Over 10,000 simulated experiments with a true lift, the paired test detected it 68.7% of the time; the independent test never did (0.0%).
- With no lift at all, the paired test raised a false alarm in 4.9% of runs, close to the nominal 5%. The independent test did so in 0.0%: because it wrongly treats correlated pairs as unrelated, it overstates the noise and becomes far too cautious, so its p values are not correct either.
When pairing is legitimate, and when it is not
- The pairs must be fixed by the design, before you see the outcomes. Pairing A and B by the day they ran, by the person measured, or by a matching variable chosen in advance is fine. Sorting both groups and pairing the smallest with the smallest is not; that manufactures correlation and gives misleading results.
- Both conditions must be exposed to the shared factor in the same way. In an A/B test, both versions must run at the same time every day, with users randomly assigned within each day. If B only ran in the second week, the day effect and the version effect are tangled together and no test can separate them.
- The pairs themselves should be roughly independent of each other. The paired test treats the differences as a simple random sample. If the difference drifts over time, for instance a novelty effect that fades, the average lift is a moving target and a time-series view of the daily differences is more informative than one t test.
- Pairing that turns out to be weak costs little. If the pair members are barely correlated, the paired test loses some degrees of freedom (n − 1 instead of 2n − 2) and a little power, but it remains valid. Ignoring strong real pairing costs far more, as the simulation shows.
You do not need to remove trends or cycles before testing. Taking the within-day difference already removes anything the two versions share on that day. If you have more than two conditions, or want to add covariates, the same idea becomes a repeated-measures ANOVA or a regression with day as a blocking factor; with exactly two conditions per day, that model gives the same answer as the paired t test. If the differences look skewed, a sign-flip permutation test or the Wilcoxon signed-rank test on the differences keeps the pairing (see Wilcoxon signed-rank test assumptions). For truly independent groups, use Welch's version of the independent test (see Welch vs Student t test).
Not sure which test fits your design? Try the test chooser. For more plain-language guides to planning and reporting analyses, see the DASS blog.
How to report a paired t test in APA style (7th edition)
Say what formed the pairs, give the means and standard deviations of each condition, the mean difference with its confidence interval, the test result and an effect size. Using the example above:
"Daily revenue was higher under version B (M = 54.80, SD = 8.90) than under version A (M = 52.14, SD = 8.91). A paired-samples t test with days as pairs (n = 14) showed that the mean difference of 2.66, 95% CI [0.91, 4.40], was statistically significant, t(13) = 3.29, p = .006, d_z = 0.88."
Name the standardiser for the effect size: d_z divides the mean difference by the SD of the differences, and it is usually much larger than a d based on the raw SDs, so readers need to know which one you used (see How to interpret Cohen's d). In the method section, one sentence is enough: "Because both versions ran on each day, conditions were compared with a paired-samples t test on the daily differences."
Related tools and guides
- t-test power and sample size calculator
- APA 7 formatter for t tests
- Check a reported p-value
- Which statistical test should I use?
- Welch vs Student t test: should you just always use Welch?
- Wilcoxon signed-rank test assumptions
- How to interpret Cohen's d
- Paired difference test (Wikipedia)
More answered questions
- How are alpha and beta (Type I and Type II error rates) related?
- Welch vs Student t test: should you just always use Welch?
- One-tailed vs two-tailed tests: why not just test the direction the data point to?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.