Answer
Training, validation and test sets: what is each one for?
The short answer
The training set fits each model. The validation set compares models and tunes settings, so it helps choose the model. The test set is kept aside and used once, at the end, to estimate how the chosen model will do on new data. Because you keep whichever model looks best on the validation data, its validation score flatters it; only an untouched test set is honest. This applies to any model, not just neural networks, and with small data cross-validation usually replaces the validation set.
The short answer
Each piece of data has one job, and the jobs differ in whether the data influences the final model:
- Training data estimates the model's parameters (regression coefficients, network weights, tree splits).
- Validation data helps you make choices the fitting procedure cannot make by itself: which predictors to keep, a polynomial degree, a penalty strength, how many training epochs to run before stopping. Each candidate is fitted on the training data and scored on the validation data, and the best-scoring one wins.
- Test data is locked away until every decision is made. You score the final model on it once and report that number as your estimate of performance on new cases.
If you only fit one model with no settings to tune, two pieces are enough: fit on one, evaluate on the other. A third piece becomes necessary as soon as you use held-out data to choose between models, because that data has then become part of the model-building process.
Why the validation score is too optimistic
Any score computed on a finite sample contains noise. When you compare ten candidate models on the same validation data and keep the one with the lowest error, you favour a model that is genuinely good and a little lucky on those particular cases. Picking the minimum of several noisy numbers pulls the result downward, the same winner's curse that makes the best-performing fund of last year look better than it will next year.
Training error has the same problem in a stronger form: the model was fitted to those exact points, so a flexible model can score well on them simply by bending towards the noise. The test set avoids both biases because nothing about the chosen model depended on it. That is also why it must be used once. If you check the test score, go back to adjust the model, and check again, the test set has quietly become a second validation set and its score is optimistic too.
Terminology varies, which causes much of the confusion. Some software and some fields (clinical prediction modelling, for example) say "validation" for what is described here as the test set, as in "external validation" on data from a different hospital. Ask what the data was used for, not what it was called.
See it in R and Python
The code simulates a curved relationship plus noise with a standard deviation of 0.5, so no model can achieve a root-mean-square error (RMSE) much below 0.5 on new data. It fits polynomials of degree 1 to 8 to 60 training cases, chooses the degree with the lowest RMSE on 30 validation cases, and then scores that model on 30 test cases. It shows one split, then repeats the whole procedure 2,000 times to see the average behaviour.
R
set.seed(19048)
# True curve plus noise (noise SD 0.5 is the best RMSE any model can reach)
sim <- function(n) {
x <- runif(n, 0, 3)
data.frame(x = x, y = sin(2 * x) + rnorm(n, sd = 0.5))
}
rmse <- function(fit, d) sqrt(mean((d$y - predict(fit, newdata = d))^2))
# One split: fit polynomials of degree 1-8 on the training set,
# pick the degree with the lowest validation RMSE, then score it once on the test set
split_once <- function(n_train = 60, n_valid = 30, n_test = 30) {
train <- sim(n_train); valid <- sim(n_valid); test <- sim(n_test)
fits <- lapply(1:8, function(k) lm(y ~ poly(x, k), data = train))
tr <- sapply(fits, rmse, d = train)
va <- sapply(fits, rmse, d = valid)
best <- which.min(va)
list(table = round(rbind(train = tr, validation = va), 3),
result = c(degree = best, train = tr[best], validation = va[best],
test = rmse(fits[[best]], test)))
}
one <- split_once()
colnames(one$table) <- paste0("deg", 1:8)
one$table
round(one$result, 3)
# Repeat the whole procedure 2,000 times
res <- t(replicate(2000, split_once()$result))
round(colMeans(res[, c("train", "validation", "test")]), 3)
# Share of runs where the chosen model looked better on validation than on test
round(mean(res[, "validation"] < res[, "test"]), 3)
table(res[, "degree"])Python
import numpy as np
from numpy.polynomial import Polynomial
rng = np.random.default_rng(19048)
# True curve plus noise (noise SD 0.5 is the best RMSE any model can reach)
def sim(n):
x = rng.uniform(0, 3, n)
return x, np.sin(2 * x) + rng.normal(0, 0.5, n)
def rmse(fit, x, y):
return float(np.sqrt(np.mean((y - fit(x)) ** 2)))
# One split: fit polynomials of degree 1-8 on the training set,
# pick the degree with the lowest validation RMSE, then score it once on the test set
def split_once(n_train=60, n_valid=30, n_test=30):
(xtr, ytr), (xva, yva), (xte, yte) = sim(n_train), sim(n_valid), sim(n_test)
fits = [Polynomial.fit(xtr, ytr, k) for k in range(1, 9)]
tr = np.array([rmse(f, xtr, ytr) for f in fits])
va = np.array([rmse(f, xva, yva) for f in fits])
best = int(np.argmin(va))
return tr, va, {"degree": best + 1, "train": tr[best], "validation": va[best],
"test": rmse(fits[best], xte, yte)}
tr, va, result = split_once()
print("train ", np.round(tr, 3))
print("validation", np.round(va, 3))
print({k: round(float(v), 3) for k, v in result.items()})
# Repeat the whole procedure 2,000 times
runs = [split_once()[2] for _ in range(2000)]
means = {k: round(float(np.mean([r[k] for r in runs])), 3) for k in ["train", "validation", "test"]}
print(means)
# Share of runs where the chosen model looked better on validation than on test
print(round(float(np.mean([r["validation"] < r["test"] for r in runs])), 3))
print(dict(sorted(zip(*np.unique([r["degree"] for r in runs], return_counts=True))))) deg1 deg2 deg3 deg4 deg5 deg6 deg7 deg8
train 0.613 0.584 0.475 0.466 0.466 0.46 0.459 0.458
validation 0.518 0.452 0.371 0.414 0.409 0.52 0.553 0.436
degree train validation test
3.000 0.475 0.371 0.415
train validation test
0.482 0.510 0.527
[1] 0.559
1 2 3 4 5 6 7 8
21 27 739 321 323 219 162 188
The output is from one seeded run of the R code. Python's random numbers differ from R's, so its figures differ slightly (its single split also picks degree 3, and over 2,000 runs the averages are 0.482, 0.507 and 0.526, with the validation score beating the test score in 57.8% of runs), but they show the same pattern.
What the numbers show
- Training error always rewards complexity. In the single split, training RMSE never rises as the degree goes up, falling from 0.613 for a straight line to 0.458 for degree 8, which is below the noise level of 0.5. No model can really predict new data that well; the extra terms are fitting noise.
- Validation error is what makes the choice. It is lowest for the cubic (0.371), so degree 3 is selected. Higher degrees do worse on the validation cases even though they fit the training cases better.
- The chosen model's validation score flatters it. Averaged over 2,000 repetitions, the selected model's validation RMSE is 0.510 but its test RMSE is 0.527, and it looked better on validation than on test in 55.9% of runs. The training RMSE of the same models averages 0.482, more optimistic still. Only the test figure is an honest estimate of performance on new data.
- The selection itself is noisy. Degree 3 won in 739 of the 2,000 runs, but degrees 6 to 8 won in 569 runs, purely because of which 30 cases happened to land in the validation set. A small validation set gives an unstable choice.
The optimism here is modest because only eight similar models were compared. It grows when you compare many candidates, tune many settings, or use a small validation set, which is exactly when people are most tempted to report the validation score.
Practical advice
- The validation set is optional, but validation is not. Any time you choose among models or tune a setting you need some held-out estimate to choose with. With plenty of data a single validation set is fine. With modest data, use k-fold cross-validation on the non-test data instead: every case takes a turn as validation data, which makes the choice far less noisy than one small split.
- It is not specific to neural networks. Networks use a validation set for early stopping (stop training when validation error starts rising), but the same logic applies to choosing a penalty in lasso regression, the depth of a tree, or the predictors in an ordinary regression.
- After choosing, refit on more data. Once the settings are fixed, it is common to refit the chosen model on training plus validation data before the final test.
- Keep all preprocessing inside the training data. Scaling constants, imputation values and selected features must be computed from training data only, then applied to the validation and test data. Otherwise information leaks from the held-out cases and every score is optimistic.
- If the data are too small to spare a test set, use nested cross-validation (an inner loop chooses the model, an outer loop scores the whole procedure) or the bootstrap; see why the bootstrap works.
- Split in the way new data will arrive. For time series, train on the past and test on the future; for clustered data such as patients within clinics, keep each cluster in a single set.
Information criteria such as AIC are another way to compare models without a separate validation set; see AIC vs BIC. For help choosing an analysis, try our test chooser or read the DASS blog.
How to report RMSE in APA style (7th edition)
Say how the data were split, what was chosen on which part, and report the test-set error as the headline figure. RMSE is an abbreviation, so it is not italicised, and it keeps its leading zero because it is in the units of the outcome and can exceed 1. Using the single split from the R example:
"Polynomial regression models of degree 1 to 8 were fitted to a training set (n = 60), and the degree was chosen by root-mean-square error (RMSE) on a separate validation set (n = 30). The selected cubic model had a validation RMSE of 0.371 and an RMSE of 0.415 on an independent test set (n = 30) that was not used for any modelling decision."
In the method section, state how cases were assigned to the sets (for example, at random or by date) and, if you used cross-validation, the number of folds and repeats. A table can list training, validation and test RMSE for each candidate model.
Related tools and guides
- APA 7 formatter for regression
- Why is accuracy a poor way to judge a classification model?
- AIC vs BIC: which should you use?
- Is R-squared useful or misleading?
- Training, validation, and test data sets (Wikipedia)
More answered questions
- AIC vs BIC: which model selection criterion should you use?
- Why is accuracy a poor way to judge a classification model?
- Fixed vs random effects in mixed models: what is the difference?
Working with your own data?
General answers only go so far. Send us your situation and we'll reply by email within two business days.
Ask your questionWritten by AskStats with AI assistance. This is general information, not advice for your specific data or study. When the results matter, check your approach with a qualified statistician.