Answer

Training, validation and test sets: what is each one for?

Inspired by a question on Cross Validated ·

model selectionregression

The short answer

The training set fits each model. The validation set compares models and tunes settings, so it helps choose the model. The test set is kept aside and used once, at the end, to estimate how the chosen model will do on new data. Because you keep whichever model looks best on the validation data, its validation score flatters it; only an untouched test set is honest. This applies to any model, not just neural networks, and with small data cross-validation usually replaces the validation set.

The short answer

Each piece of data has one job, and the jobs differ in whether the data influences the final model:

If you only fit one model with no settings to tune, two pieces are enough: fit on one, evaluate on the other. A third piece becomes necessary as soon as you use held-out data to choose between models, because that data has then become part of the model-building process.

Why the validation score is too optimistic

Any score computed on a finite sample contains noise. When you compare ten candidate models on the same validation data and keep the one with the lowest error, you favour a model that is genuinely good and a little lucky on those particular cases. Picking the minimum of several noisy numbers pulls the result downward, the same winner's curse that makes the best-performing fund of last year look better than it will next year.

Training error has the same problem in a stronger form: the model was fitted to those exact points, so a flexible model can score well on them simply by bending towards the noise. The test set avoids both biases because nothing about the chosen model depended on it. That is also why it must be used once. If you check the test score, go back to adjust the model, and check again, the test set has quietly become a second validation set and its score is optimistic too.

Terminology varies, which causes much of the confusion. Some software and some fields (clinical prediction modelling, for example) say "validation" for what is described here as the test set, as in "external validation" on data from a different hospital. Ask what the data was used for, not what it was called.

See it in R and Python

The code simulates a curved relationship plus noise with a standard deviation of 0.5, so no model can achieve a root-mean-square error (RMSE) much below 0.5 on new data. It fits polynomials of degree 1 to 8 to 60 training cases, chooses the degree with the lowest RMSE on 30 validation cases, and then scores that model on 30 test cases. It shows one split, then repeats the whole procedure 2,000 times to see the average behaviour.

R

set.seed(19048)

# True curve plus noise (noise SD 0.5 is the best RMSE any model can reach)
sim <- function(n) {
  x <- runif(n, 0, 3)
  data.frame(x = x, y = sin(2 * x) + rnorm(n, sd = 0.5))
}
rmse <- function(fit, d) sqrt(mean((d$y - predict(fit, newdata = d))^2))

# One split: fit polynomials of degree 1-8 on the training set,
# pick the degree with the lowest validation RMSE, then score it once on the test set
split_once <- function(n_train = 60, n_valid = 30, n_test = 30) {
  train <- sim(n_train); valid <- sim(n_valid); test <- sim(n_test)
  fits <- lapply(1:8, function(k) lm(y ~ poly(x, k), data = train))
  tr <- sapply(fits, rmse, d = train)
  va <- sapply(fits, rmse, d = valid)
  best <- which.min(va)
  list(table = round(rbind(train = tr, validation = va), 3),
       result = c(degree = best, train = tr[best], validation = va[best],
                  test = rmse(fits[[best]], test)))
}

one <- split_once()
colnames(one$table) <- paste0("deg", 1:8)
one$table
round(one$result, 3)

# Repeat the whole procedure 2,000 times
res <- t(replicate(2000, split_once()$result))
round(colMeans(res[, c("train", "validation", "test")]), 3)
# Share of runs where the chosen model looked better on validation than on test
round(mean(res[, "validation"] < res[, "test"]), 3)
table(res[, "degree"])

Python

import numpy as np
from numpy.polynomial import Polynomial

rng = np.random.default_rng(19048)

# True curve plus noise (noise SD 0.5 is the best RMSE any model can reach)
def sim(n):
    x = rng.uniform(0, 3, n)
    return x, np.sin(2 * x) + rng.normal(0, 0.5, n)

def rmse(fit, x, y):
    return float(np.sqrt(np.mean((y - fit(x)) ** 2)))

# One split: fit polynomials of degree 1-8 on the training set,
# pick the degree with the lowest validation RMSE, then score it once on the test set
def split_once(n_train=60, n_valid=30, n_test=30):
    (xtr, ytr), (xva, yva), (xte, yte) = sim(n_train), sim(n_valid), sim(n_test)
    fits = [Polynomial.fit(xtr, ytr, k) for k in range(1, 9)]
    tr = np.array([rmse(f, xtr, ytr) for f in fits])
    va = np.array([rmse(f, xva, yva) for f in fits])
    best = int(np.argmin(va))
    return tr, va, {"degree": best + 1, "train": tr[best], "validation": va[best],
                    "test": rmse(fits[best], xte, yte)}

tr, va, result = split_once()
print("train     ", np.round(tr, 3))
print("validation", np.round(va, 3))
print({k: round(float(v), 3) for k, v in result.items()})

# Repeat the whole procedure 2,000 times
runs = [split_once()[2] for _ in range(2000)]
means = {k: round(float(np.mean([r[k] for r in runs])), 3) for k in ["train", "validation", "test"]}
print(means)
# Share of runs where the chosen model looked better on validation than on test
print(round(float(np.mean([r["validation"] < r["test"] for r in runs])), 3))
print(dict(sorted(zip(*np.unique([r["degree"] for r in runs], return_counts=True)))))
            deg1  deg2  deg3  deg4  deg5 deg6  deg7  deg8
train      0.613 0.584 0.475 0.466 0.466 0.46 0.459 0.458
validation 0.518 0.452 0.371 0.414 0.409 0.52 0.553 0.436
    degree      train validation       test 
     3.000      0.475      0.371      0.415 
     train validation       test 
     0.482      0.510      0.527 
[1] 0.559

  1   2   3   4   5   6   7   8 
 21  27 739 321 323 219 162 188

The output is from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (its single split also picks degree 3, and over 2,000 runs the averages are 0.482, 0.507 and 0.526, with the validation score beating the test score in 57.8% of runs), but they show the same pattern.

What the numbers show

The optimism here is modest because only eight similar models were compared. It grows when you compare many candidates, tune many settings, or use a small validation set, which is exactly when people are most tempted to report the validation score.

Practical advice

Information criteria such as AIC are another way to compare models without a separate validation set; see AIC vs BIC. For help choosing an analysis, try our test chooser or read the DASS blog.

How to report RMSE in APA style (7th edition)

Say how the data were split, what was chosen on which part, and report the test-set error as the headline figure. RMSE is an abbreviation, so it is not italicised, and it keeps its leading zero because it is in the units of the outcome and can exceed 1. Using the single split from the R example:

"Polynomial regression models of degree 1 to 8 were fitted to a training set (n = 60), and the degree was chosen by root-mean-square error (RMSE) on a separate validation set (n = 30). The selected cubic model had a validation RMSE of 0.371 and an RMSE of 0.415 on an independent test set (n = 30) that was not used for any modelling decision."

In the method section, state how cases were assigned to the sets (for example, at random or by date) and, if you used cross-validation, the number of folds and repeats. A table can list training, validation and test RMSE for each candidate model.

Related tools and guides

More answered questions

Working with your own data?

General answers only go so far. Send us your situation and we'll reply by email within two business days.

Ask your question

Written by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.