Dictionary of Applied Machine Learning
Updated on 2026-09-30
Typeset PDF version — the authoritative form of this entry
Model selection chooses among several trained hypotheses by comparing the validation error each of them incurs on a validation set. Each candidate is the output of empirical risk minimization (ERM) run with a different hypothesis space or a different hyperparameter setting, such as the learning rate of a gradient-based method or the depth of a decision tree. The training error cannot serve as the criterion: it falls as the hypothesis space grows, while the risk eventually rises again. $k$-fold cross-validation ($k$-fold CV) supplies the validation error when the training set is too small to spare a validation set. A test set held out from every candidate and from the comparison itself is what remains for estimating the risk of the selected hypothesis.
Two forecasts of a day's maximum
daytime temperature $\truelabel$ from its morning minimum
$\feature$ have been fitted to the same six days of weather
recordings: a straight line and a degree-five polynomial
(Fig. 1). Each of them is a
hypothesis that empirical risk minimization (ERM) on the training set $\trainset$
of those six days delivers, the line from a linear model and the
polynomial from the larger hypothesis space of polynomial regression. The
polynomial passes through every day of $\trainset$, so its
training error, the average squared error loss over $\trainset$, is
zero; the line misses each of them and its training error is
$0.47$. The training error therefore ranks the polynomial first, and
following that ranking would pick the candidate that predicts worst
on a day it has not seen.
Model selection ranks instead by the validation error on a validation set
$\valset$ that ERM never saw, which reverses the order, and
it leaves a test set $\testset$ untouched for estimating the
risk of the winner. The twenty days held back from the fit are
split fourteen to six for these two roles, so the validation errors
below are averages over fourteen days, not over all twenty.
pythondemos/validation.py
The basic workflow has three stages (Fig. 2).
The test set $\testset$ is held out from all three stages and used only afterwards. The independent and identically distributed assumption (i.i.d. assumption) makes its role precise. Suppose the data points of $\trainset$, $\valset$, and $\testset$ are realizations of independent and identically distributed (i.i.d.) random variables (RVs) with a common probability distribution $\probdist^{(\datapoint)}$. The average loss on $\testset$ is then an unbiased estimator of the risk of $\learnthypothesis^{(c^{\star})}$, i.e., of its expected loss on a data point that no stage has seen. Selecting on $\testset$ destroys this property: a test set used repeatedly to keep the candidate with the smallest error underestimates the risk of the candidate finally chosen (Hastie et al., 2009, Sect. 7.2).
A/B testing is a further, live-deployment stage of model
selection: two short-listed hypotheses $\hypothesis^{(A)}$
and $\hypothesis^{(B)}$ are compared on data points drawn
from the operating machine learning system (ML system), which can differ from the
probability distribution behind $\trainset$, $\valset$, and $\testset$.
A/B testing is therefore a special case of model
selection restricted to two pre-trained candidates and live
data, not a synonym for model selection.
See also: validation, validation error, $k$-fold cross-validation, validation set, test set, hyperparameter, hypothesis space, overfitting, A/B testing.
@misc{dictml_modelsel,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {model selection},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-30},
url = {https://dictionaryofml.org/terms/modelsel.html}
}