Dictionary of Applied Machine Learning

validation error

Updated on 2026-09-29

Typeset PDF version — the authoritative form of this entry

Python demo — a script that recomputes what this entry states and prints one line per check

The validation error $\valerror$ of a hypothesis $\learnthypothesis$ is its average loss over a validation set $\valset$, computed with $\learnthypothesis$ fixed. Because no data point of $\valset$ entered the training, $\valerror$ estimates the risk of $\learnthypothesis$, while the training error understates it. For data points that are realizations of independent and identically distributed (i.i.d.) random variables (RVs), $\valerror$ is itself the realization of an RV whose expectation is the risk and whose fluctuation shrinks as $\valset$ grows. A training error far below the validation error signals overfitting, and comparing validation errors across candidate models is model selection.

Definition

B-defA straight line fitted to forty days of weather recordings predicts each day's maximum daytime temperature from its morning minimum. Applied to twenty held-back days, the line misses each day's recorded maximum by some amount. The squared error loss squares each miss, and the average of the twenty squared misses is the validation error of the line.

B-riskThe validation error $\valerror$ of a hypothesis $\learnthypothesis$ is its average loss over a validation set $\valset$, \[ \valerror = \frac{1}{|\valset|} \sum_{\datapoint \in \valset} \lossfunc{\datapoint}{\learnthypothesis} \text{,} \] computed with $\learnthypothesis$ fixed. Because no data point of $\valset$ entered the training, $\valerror$ estimates the risk of $\learnthypothesis$ (Hastie et al., 2009, Sect. 7.2); the training error, computed over the training set whose average loss empirical risk minimization (ERM) minimized, understates it. For the line of the opening example the training error is $2.45$, the validation error $2.87$ and the risk $2.61$.

B-spreadThe validation error is an average of $|\valset|$ individual losses. For data points that are realizations of independent and identically distributed (i.i.d.) random variables (RVs), $\valerror$ is itself the realization of an RV whose expectation is the risk of $\learnthypothesis$, since each summand has that expectation, and whose variance is the variance of a single loss divided by $|\valset|$, so the fluctuation around the risk shrinks as $\valset$ grows. Fig. 1 measures the fluctuation for one fixed line: recomputed over $300$ independent draws of $\valset$ for each size, the validation error scatters around the risk, with a spread that falls from $1.4$ at $|\valset| = 5$ to $0.3$ at $|\valset| = 80$.

Figure 1 of the entry valerr
Figure 1: The validation error of one fixed hypothesis, a line fitted to $80$ days of weather recordings, recomputed over $300$ independent draws of the validation set for each size $|\valset|$. Each bar spans one standard deviation either side of the mean. The mean sits on the risk (dashed) for every size, while the spread shrinks as $|\valset|$ grows. Data generated by pythondemos/valerr.py
Read beside the training error, the validation error diagnoses a machine learning (ML) method: a training error far below the validation error signals overfitting (Hastie et al., 2009, Sect. 7.2). Comparing the validation errors of several candidate models is model selection (Bishop, 2006, Sect. 1.3); a validation set that is reused many times for this purpose is itself overfitted, which is why a test set is kept apart for the final assessment (Hastie et al., 2009, Sect. 7.2).

See also: validation set, validation, training set, test set, training error, loss, risk, hypothesis, empirical risk minimization, overfitting, model selection.

References

  1. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7
  2. Bishop (2006). Pattern Recognition and Machine Learning. Springer Science+Business Media. doi.org/10.1007/978-0-387-45528-0

Cite this entry

@misc{dictml_valerr,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {validation error},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-29},
  url = {https://dictionaryofml.org/terms/valerr.html}
}