Answer
Regression to the mean vs the gambler's fallacy: what is the difference?
The short answer
They make different predictions. Regression to the mean says that after an extreme result the next one is expected to be closer to average, because part of the extreme was luck that will not repeat. The gambler's fallacy says the next result will swing the other way to balance things out. In a simulation, top scorers on a pure-guessing test averaged 58.12 and then 49.76 next time: back to the mean, not below it. After five heads in a row, a fair coin still came up heads about half the time.
The short answer
The two ideas sound alike, because both say that something unusual tends to be followed by something less unusual. But they make different predictions about the next result, and only one of them is correct.
- Regression to the mean says: if a score was extreme partly because of luck, the next score is expected to be closer to the average. The luck does not carry over, so the expectation goes back towards the mean. It never says the next score will overshoot to the other side.
- The gambler's fallacy says: after a run of one outcome, the opposite outcome becomes more likely, as if chance had to even things out. For independent events such as coin flips this is false. The coin has no memory, so the probability stays where it was.
So there is no conflict. After a good game, expecting a somewhat less good one is regression to the mean and is reasonable. Expecting a bad game, because you are now "due" for one, is the gambler's fallacy.
Why the two ideas do not contradict each other
Think of any observed score as stable part + luck. When we pick out people because their first score was extreme, we pick out mostly people whose luck happened to be good that day. On a second attempt their stable part is the same, but their luck is drawn afresh and averages zero. Their expected second score is therefore their stable part, which is less extreme than the score that got them selected.
That reasoning uses independence; it does not deny it. The gambler's fallacy, by contrast, assumes that past luck changes future luck. Regression to the mean only needs future luck to be unrelated to past luck.
How far scores move back depends on how much of the variation is stable. With standardised scores, the expected second score is the first score multiplied by the correlation r between the two occasions. If scores were pure luck (r = 0), an extreme group's expected second score is the overall mean: full regression. If they were pure skill (r = 1), there is no regression at all. Real tests sit in between. This is also why regression to the mean is a design problem: a group chosen for extreme baseline scores will move towards the mean at follow-up even if the treatment does nothing, which is why such studies need a control group selected in the same way.
See it in R and Python
The code simulates 5,000 students taking two tests in two settings: once where everyone guesses on every question of a true/false test with 100 items, and once where true ability varies and each test adds independent noise of the same size. It compares the top 10% on the first test with their own second scores. Finally it flips a fair coin a million times and looks at the flip that follows a run of five heads.
R
set.seed(204397)
n <- 5000
# 1. Pure luck: everyone guesses on all 100 true/false items, twice
g1 <- rbinom(n, 100, 0.5); g2 <- rbinom(n, 100, 0.5)
top <- g1 >= quantile(g1, 0.9)
round(c(top_test1 = mean(g1[top]), top_test2 = mean(g2[top]),
bottom_test2 = mean(g2[g1 <= quantile(g1, 0.1)]), r = cor(g1, g2)), 2)
# 2. Skill plus luck: true ability varies, each test adds independent noise
ability <- rnorm(n, 70, 8)
s1 <- ability + rnorm(n, 0, 8); s2 <- ability + rnorm(n, 0, 8)
top <- s1 >= quantile(s1, 0.9)
r <- cor(s1, s2)
predicted <- mean(s2) + r * sd(s2) / sd(s1) * (mean(s1[top]) - mean(s1))
round(c(all_mean = mean(s1), top_test1 = mean(s1[top]),
top_test2 = mean(s2[top]), predicted = predicted, r = r), 2)
ct <- cor.test(s1, s2)
round(c(r = unname(ct$estimate), ct$conf.int), 3)
# 3. A fair coin: chance of heads straight after five heads in a row
flips <- rbinom(1e6, 1, 0.5); N <- length(flips)
run5 <- flips[1:(N - 5)] + flips[2:(N - 4)] + flips[3:(N - 3)] +
flips[4:(N - 2)] + flips[5:(N - 1)]
nxt <- flips[6:N]
c(streaks = sum(run5 == 5), p_heads_after_5_heads = round(mean(nxt[run5 == 5]), 3),
p_heads_after_5_tails = round(mean(nxt[run5 == 0]), 3))Python
import numpy as np
from scipy import stats
rng = np.random.default_rng(204397)
n = 5000
# 1. Pure luck: everyone guesses on all 100 true/false items, twice
g1 = rng.binomial(100, 0.5, n); g2 = rng.binomial(100, 0.5, n)
top = g1 >= np.quantile(g1, 0.9)
print({"top_test1": round(g1[top].mean(), 2), "top_test2": round(g2[top].mean(), 2),
"bottom_test2": round(g2[g1 <= np.quantile(g1, 0.1)].mean(), 2),
"r": round(np.corrcoef(g1, g2)[0, 1], 2)})
# 2. Skill plus luck: true ability varies, each test adds independent noise
ability = rng.normal(70, 8, n)
s1 = ability + rng.normal(0, 8, n); s2 = ability + rng.normal(0, 8, n)
top = s1 >= np.quantile(s1, 0.9)
r = np.corrcoef(s1, s2)[0, 1]
predicted = s2.mean() + r * s2.std(ddof=1) / s1.std(ddof=1) * (s1[top].mean() - s1.mean())
print({"all_mean": round(s1.mean(), 2), "top_test1": round(s1[top].mean(), 2),
"top_test2": round(s2[top].mean(), 2), "predicted": round(predicted, 2), "r": round(r, 2)})
# 95% CI for r via Fisher's z, as in R's cor.test()
z, se = np.arctanh(r), 1 / np.sqrt(n - 3)
lo, hi = np.tanh(z + np.array([-1, 1]) * stats.norm.ppf(0.975) * se)
print({"r": round(r, 3), "lower": round(lo, 3), "upper": round(hi, 3)})
# 3. A fair coin: chance of heads straight after five heads in a row
flips = rng.binomial(1, 0.5, 10**6); N = len(flips)
run5 = sum(flips[k:N - 5 + k] for k in range(5))
nxt = flips[5:]
print({"streaks": int((run5 == 5).sum()),
"p_heads_after_5_heads": round(nxt[run5 == 5].mean(), 3),
"p_heads_after_5_tails": round(nxt[run5 == 0].mean(), 3)}) top_test1 top_test2 bottom_test2 r
58.12 49.76 49.94 -0.02
all_mean top_test1 top_test2 predicted r
69.90 89.85 79.08 79.54 0.50
r
0.498 0.477 0.519
streaks p_heads_after_5_heads p_heads_after_5_tails
31581.000 0.502 0.494
The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (for example, the top group's second-test mean is 79.71 instead of 79.08 in the skill-plus-luck setting), but the pattern is the same.
- Pure luck: full regression, no overshoot. The top 10% of guessers averaged 58.12 on the first test and 49.76 on the second, essentially the overall expected score of 50. The bottom 10% also ended up at 49.94. The correlation between the two tests was about zero (−0.02). Nobody who did well was "due" a bad score; they simply got an ordinary one.
- Skill plus luck: partial regression. Here half the variation was stable ability, and the test-retest correlation was 0.50. The top 10% averaged 89.85 on the first test (the overall mean was 69.90) and 79.08 on the second: roughly half-way back, close to the 79.54 that the regression line predicts. They stayed well above average, because their ability was real.
- The coin has no memory. Among about 31,600 runs of five heads, the next flip was heads 50.2% of the time; after five tails it was heads 49.4% of the time. Both are within about two standard errors (roughly 0.003) of one half. A gambler's-fallacy prediction, that tails becomes more likely after a run of heads, does not show up.
Where people go wrong
- Reading regression as a force. Nothing pulls extreme people back. The group's average moves because it was selected on a score that contained luck. Select people on their second score instead, and their first scores regress towards the mean too.
- Expecting regression for a sequence you did not select on. If you look at a coin's next flip without having chosen it because of an extreme past, there is nothing to regress. Regression to the mean is about the gap between a selected extreme and its expected value, not about future outcomes compensating for past ones.
- Crediting an intervention. Patients enrolled because their symptoms were at a peak, or schools flagged for a terrible year, tend to improve on their own. A before-after comparison without a similarly selected control group will mistake this for a treatment effect.
The test chooser can help you plan the right comparison, and the DASS blog has more on study design. For a related idea about how selection distorts estimates, see false positive risk in underpowered studies, where only effects that happened to come out large reach significance.
How to report a test-retest correlation in APA style (7th edition)
When you discuss regression to the mean in a write-up, the useful number is the correlation between the two occasions, because it tells readers how much an extreme group is expected to move back. Using the skill-plus-luck example:
"Scores on the two tests were positively correlated, r(4998) = .50, 95% CI [.48, .52], p < .001. Students in the top 10% on the first test (M = 89.85) scored lower on the second test (M = 79.08), in line with the 79.54 expected from regression to the mean."
Give the degrees of freedom (n − 2) in parentheses, drop the leading zero for r and p, and put the confidence interval in square brackets. If the extreme group received a treatment, report its change alongside the change in a control group selected the same way, so readers can separate the treatment effect from regression to the mean.
Related tools and guides
- APA 7 formatter for regression
- Correlation power and sample size calculator
- APA 7 formatter for correlations
- False positive risk: how likely is a significant finding to be wrong?
- How to interpret regression output in R
- Which statistical test should I use?
- Regression toward the mean (Wikipedia)
- Gambler's fallacy (Wikipedia)
More answered questions
- What goes wrong if you sort x and y separately before a regression?
- RMSE vs standard deviation: what does comparing them tell you?
- Should you check normality of the outcome variable or of the residuals?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.