Answer

Regression to the mean vs the gambler's fallacy: what is the difference?

Inspired by a question on Cross Validated ·

regressioncorrelationstudy design

The short answer

They make different predictions. Regression to the mean says that after an extreme result the next one is expected to be closer to average, because part of the extreme was luck that will not repeat. The gambler's fallacy says the next result will swing the other way to balance things out. In a simulation, top scorers on a pure-guessing test averaged 58.12 and then 49.76 next time: back to the mean, not below it. After five heads in a row, a fair coin still came up heads about half the time.

The short answer

The two ideas sound alike, because both say that something unusual tends to be followed by something less unusual. But they make different predictions about the next result, and only one of them is correct.

So there is no conflict. After a good game, expecting a somewhat less good one is regression to the mean and is reasonable. Expecting a bad game, because you are now "due" for one, is the gambler's fallacy.

Why the two ideas do not contradict each other

Think of any observed score as stable part + luck. When we pick out people because their first score was extreme, we pick out mostly people whose luck happened to be good that day. On a second attempt their stable part is the same, but their luck is drawn afresh and averages zero. Their expected second score is therefore their stable part, which is less extreme than the score that got them selected.

That reasoning uses independence; it does not deny it. The gambler's fallacy, by contrast, assumes that past luck changes future luck. Regression to the mean only needs future luck to be unrelated to past luck.

How far scores move back depends on how much of the variation is stable. With standardised scores, the expected second score is the first score multiplied by the correlation r between the two occasions. If scores were pure luck (r = 0), an extreme group's expected second score is the overall mean: full regression. If they were pure skill (r = 1), there is no regression at all. Real tests sit in between. This is also why regression to the mean is a design problem: a group chosen for extreme baseline scores will move towards the mean at follow-up even if the treatment does nothing, which is why such studies need a control group selected in the same way.

See it in R and Python

The code simulates 5,000 students taking two tests in two settings: once where everyone guesses on every question of a true/false test with 100 items, and once where true ability varies and each test adds independent noise of the same size. It compares the top 10% on the first test with their own second scores. Finally it flips a fair coin a million times and looks at the flip that follows a run of five heads.

R

set.seed(204397)
n <- 5000

# 1. Pure luck: everyone guesses on all 100 true/false items, twice
g1 <- rbinom(n, 100, 0.5); g2 <- rbinom(n, 100, 0.5)
top <- g1 >= quantile(g1, 0.9)
round(c(top_test1 = mean(g1[top]), top_test2 = mean(g2[top]),
        bottom_test2 = mean(g2[g1 <= quantile(g1, 0.1)]), r = cor(g1, g2)), 2)

# 2. Skill plus luck: true ability varies, each test adds independent noise
ability <- rnorm(n, 70, 8)
s1 <- ability + rnorm(n, 0, 8); s2 <- ability + rnorm(n, 0, 8)
top <- s1 >= quantile(s1, 0.9)
r <- cor(s1, s2)
predicted <- mean(s2) + r * sd(s2) / sd(s1) * (mean(s1[top]) - mean(s1))
round(c(all_mean = mean(s1), top_test1 = mean(s1[top]),
        top_test2 = mean(s2[top]), predicted = predicted, r = r), 2)
ct <- cor.test(s1, s2)
round(c(r = unname(ct$estimate), ct$conf.int), 3)

# 3. A fair coin: chance of heads straight after five heads in a row
flips <- rbinom(1e6, 1, 0.5); N <- length(flips)
run5 <- flips[1:(N - 5)] + flips[2:(N - 4)] + flips[3:(N - 3)] +
        flips[4:(N - 2)] + flips[5:(N - 1)]
nxt <- flips[6:N]
c(streaks = sum(run5 == 5), p_heads_after_5_heads = round(mean(nxt[run5 == 5]), 3),
  p_heads_after_5_tails = round(mean(nxt[run5 == 0]), 3))

Python

import numpy as np
from scipy import stats

rng = np.random.default_rng(204397)
n = 5000

# 1. Pure luck: everyone guesses on all 100 true/false items, twice
g1 = rng.binomial(100, 0.5, n); g2 = rng.binomial(100, 0.5, n)
top = g1 >= np.quantile(g1, 0.9)
print({"top_test1": round(g1[top].mean(), 2), "top_test2": round(g2[top].mean(), 2),
       "bottom_test2": round(g2[g1 <= np.quantile(g1, 0.1)].mean(), 2),
       "r": round(np.corrcoef(g1, g2)[0, 1], 2)})

# 2. Skill plus luck: true ability varies, each test adds independent noise
ability = rng.normal(70, 8, n)
s1 = ability + rng.normal(0, 8, n); s2 = ability + rng.normal(0, 8, n)
top = s1 >= np.quantile(s1, 0.9)
r = np.corrcoef(s1, s2)[0, 1]
predicted = s2.mean() + r * s2.std(ddof=1) / s1.std(ddof=1) * (s1[top].mean() - s1.mean())
print({"all_mean": round(s1.mean(), 2), "top_test1": round(s1[top].mean(), 2),
       "top_test2": round(s2[top].mean(), 2), "predicted": round(predicted, 2), "r": round(r, 2)})
# 95% CI for r via Fisher's z, as in R's cor.test()
z, se = np.arctanh(r), 1 / np.sqrt(n - 3)
lo, hi = np.tanh(z + np.array([-1, 1]) * stats.norm.ppf(0.975) * se)
print({"r": round(r, 3), "lower": round(lo, 3), "upper": round(hi, 3)})

# 3. A fair coin: chance of heads straight after five heads in a row
flips = rng.binomial(1, 0.5, 10**6); N = len(flips)
run5 = sum(flips[k:N - 5 + k] for k in range(5))
nxt = flips[5:]
print({"streaks": int((run5 == 5).sum()),
       "p_heads_after_5_heads": round(nxt[run5 == 5].mean(), 3),
       "p_heads_after_5_tails": round(nxt[run5 == 0].mean(), 3)})
   top_test1    top_test2 bottom_test2            r 
       58.12        49.76        49.94        -0.02 
 all_mean top_test1 top_test2 predicted         r 
    69.90     89.85     79.08     79.54      0.50 
    r             
0.498 0.477 0.519 
              streaks p_heads_after_5_heads p_heads_after_5_tails 
            31581.000                 0.502                 0.494

The figures come from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (for example, the top group's second-test mean is 79.71 instead of 79.08 in the skill-plus-luck setting), but the pattern is the same.

Where people go wrong

The test chooser can help you plan the right comparison, and the DASS blog has more on study design. For a related idea about how selection distorts estimates, see false positive risk in underpowered studies, where only effects that happened to come out large reach significance.

How to report a test-retest correlation in APA style (7th edition)

When you discuss regression to the mean in a write-up, the useful number is the correlation between the two occasions, because it tells readers how much an extreme group is expected to move back. Using the skill-plus-luck example:

"Scores on the two tests were positively correlated, r(4998) = .50, 95% CI [.48, .52], p < .001. Students in the top 10% on the first test (M = 89.85) scored lower on the second test (M = 79.08), in line with the 79.54 expected from regression to the mean."

Give the degrees of freedom (n − 2) in parentheses, drop the leading zero for r and p, and put the confidence interval in square brackets. If the extreme group received a treatment, report its change alongside the change in a control group selected the same way, so readers can separate the treatment effect from regression to the mean.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.