Dictionary of Applied Machine Learning

model selection

Updated on 2026-09-30

Typeset PDF version — the authoritative form of this entry

Model selection chooses among several trained hypotheses by comparing the validation error each of them incurs on a validation set. Each candidate is the output of empirical risk minimization (ERM) run with a different hypothesis space or a different hyperparameter setting, such as the learning rate of a gradient-based method or the depth of a decision tree. The training error cannot serve as the criterion: it falls as the hypothesis space grows, while the risk eventually rises again. $k$-fold cross-validation ($k$-fold CV) supplies the validation error when the training set is too small to spare a validation set. A test set held out from every candidate and from the comparison itself is what remains for estimating the risk of the selected hypothesis.

Definition

Two forecasts of a day's maximum daytime temperature $\truelabel$ from its morning minimum $\feature$ have been fitted to the same six days of weather recordings: a straight line and a degree-five polynomial (Fig. 1). Each of them is a hypothesis that empirical risk minimization (ERM) on the training set $\trainset$ of those six days delivers, the line from a linear model and the polynomial from the larger hypothesis space of polynomial regression. The polynomial passes through every day of $\trainset$, so its training error, the average squared error loss over $\trainset$, is zero; the line misses each of them and its training error is $0.47$. The training error therefore ranks the polynomial first, and following that ranking would pick the candidate that predicts worst on a day it has not seen. Model selection ranks instead by the validation error on a validation set $\valset$ that ERM never saw, which reverses the order, and it leaves a test set $\testset$ untouched for estimating the risk of the winner. The twenty days held back from the fit are split fourteen to six for these two roles, so the validation errors below are averages over fourteen days, not over all twenty.

Figure 1 of the entry modelsel
Figure 1: Model selection between two candidates fitted to the same six days (filled circles). The degree-five polynomial $\learnthypothesis^{(2)}$ passes through every day of the training set, so its training error is zero against $0.47$ for the straight line $\learnthypothesis^{(1)}$. Averaged over the fourteen validation days (open triangles), the squared error loss is $2065$ for the polynomial and $7.06$ for the line, so the validation error selects the line. The six days of the test set (open squares) are held out from both ERM and the comparison, so the average squared error loss on them, $1.83$, estimates the risk of the selected hypothesis. Data generated by pythondemos/validation.py
Model selection is the step of a machine learning (ML) workflow that chooses among several trained hypotheses by comparing the validation error each of them incurs (Shalev-Shwartz and Ben-David, 2014, Sect. 11.2.2). Each candidate $\learnthypothesis^{(c)}$ is the output of ERM run with a different hypothesis space $\hypospace^{(c)}$, a different hyperparameter setting, or both. Typical hyperparameters include the learning rate of a gradient-based method, the strength $\regparam$ of a regularization penalty term, and the depth of a decision tree. The training error cannot serve as the criterion: it falls as the hypothesis space grows, while the risk eventually rises again (Hastie et al., 2009, Sect. 7.2). Model selection therefore follows model training, e.g., via ERM, and model validation, e.g., via $k$-fold cross-validation ($k$-fold CV): it re-runs training once per candidate and reads the validation error that validation returns.

The basic workflow has three stages (Fig. 2).

  1. For each candidate $c$, run ERM on the training set with hypothesis space $\hypospace^{(c)}$ and the corresponding hyperparameter setting. This yields a trained hypothesis $\learnthypothesis^{(c)} \in \hypospace^{(c)}$.
  2. Compute the validation error $\valerror\big(\learnthypothesis^{(c)}\big)$ of each $\learnthypothesis^{(c)}$ on a validation set held out from ERM. $k$-fold CV is the standard estimator of the validation error when the training set is too small to spare a validation set (Hastie et al., 2009, Sect. 7.10).
  3. Select the candidate with the smallest validation error. Its index is $c^{\star} = \arg\min_{c}\, \valerror\big(\learnthypothesis^{(c)}\big)$. The selected hypothesis is $\learnthypothesis^{(c^{\star})}$.
For example, in predicting customer churn, the candidates are obtained by running ERM once with hypothesis space $\hypospace^{(1)}$ = linear models and once with $\hypospace^{(2)}$ = decision trees; the candidate with smaller validation error is retained, and its loss is reported on the test set.

The test set $\testset$ is held out from all three stages and used only afterwards. The independent and identically distributed assumption (i.i.d. assumption) makes its role precise. Suppose the data points of $\trainset$, $\valset$, and $\testset$ are realizations of independent and identically distributed (i.i.d.) random variables (RVs) with a common probability distribution $\probdist^{(\datapoint)}$. The average loss on $\testset$ is then an unbiased estimator of the risk of $\learnthypothesis^{(c^{\star})}$, i.e., of its expected loss on a data point that no stage has seen. Selecting on $\testset$ destroys this property: a test set used repeatedly to keep the candidate with the smallest error underestimates the risk of the candidate finally chosen (Hastie et al., 2009, Sect. 7.2).

A/B testing is a further, live-deployment stage of model selection: two short-listed hypotheses $\hypothesis^{(A)}$ and $\hypothesis^{(B)}$ are compared on data points drawn from the operating machine learning system (ML system), which can differ from the probability distribution behind $\trainset$, $\valset$, and $\testset$. A/B testing is therefore a special case of model selection restricted to two pre-trained candidates and live data, not a synonym for model selection.

Figure 2 of the entry modelsel
Figure 2: The ML workflow with model selection. For each candidate hypothesis space and hyperparameter setting, $\hypospace^{(c)}$, ERM on the training set produces a candidate trained hypothesis $\learnthypothesis^{(c)}$. The candidate with the smallest validation error on the validation set is selected as $\learnthypothesis^{(c^{\star})}$. Its loss is then reported on the test set, which has remained unused until this point. A/B testing compares two short-listed hypotheses $\hypothesis^{(A)}$ and $\hypothesis^{(B)}$ on live data from the operating ML system

See also: validation, validation error, $k$-fold cross-validation, validation set, test set, hyperparameter, hypothesis space, overfitting, A/B testing.

References

  1. Shalev-Shwartz and Ben-David (2014). Understanding Machine Learning: From Theory to Algorithms. Cambridge Univ. Press. doi.org/10.1017/cbo9781107298019
  2. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7

Cite this entry

@misc{dictml_modelsel,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {model selection},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-30},
  url = {https://dictionaryofml.org/terms/modelsel.html}
}