Dictionary of Applied Machine Learning

training error

Updated on 2026-09-29

Typeset PDF version — the authoritative form of this entry

Python demo — a script that recomputes what this entry states and prints one line per check

The training error $\trainerror$ of a hypothesis $\learnthypothesis$ is its average loss over the training set, the empirical risk of $\learnthypothesis$ on the data points used for training. For a hypothesis delivered by empirical risk minimization (ERM) it is the smallest empirical risk that any hypothesis of the hypothesis space attains. With the $0/1$ loss it is the fraction of misclassified training data points. Because the same data points selected $\learnthypothesis$, the training error is optimistic: it lies below the risk on average, by a gap that grows with the size of the hypothesis space, so a hypothesis with zero training error may be worthless on new data points. Read beside the validation error, a training error far below it signals overfitting and a large training error signals underfitting.

Definition

B-defA straight line is fitted to forty days of weather recordings to predict each day's maximum daytime temperature from its morning minimum. On each of the forty days the line misses the recorded maximum by some amount, between $0.0$ and $4.0$ degrees. The squared error loss squares each miss, and the average of the forty squared misses, $2.45$, is the training error of the line. No other line has a smaller average on these forty days: the line was chosen to minimize exactly this number. The numbers are computed by pythondemos/trainerr.py.

The training error $\trainerror$ of a hypothesis $\learnthypothesis$ is its average loss over the training set $\trainset = \{\datapoint^{(\sampleidx)}\}_{\sampleidx=1}^{\samplesize}$, \[ \trainerror = \frac{1}{\samplesize} \sum_{\sampleidx=1}^{\samplesize} \lossfunc{\datapoint^{(\sampleidx)}}{\learnthypothesis} \text{,} \] i.e., the empirical risk of $\learnthypothesis$ on the data points used for training. For a hypothesis delivered by empirical risk minimization (ERM), the training error is the smallest empirical risk that any hypothesis of the hypothesis space attains on $\trainset$, since ERM minimizes precisely this average. With the $0/1$ loss, the training error of a classifier is the fraction of training data points it misclassifies; with the logistic loss, it is the negative average log-likelihood of the training labels (Hastie et al., 2009, Sect. 7.2).

The training error is computed on the same data points that selected $\learnthypothesis$, and this is what limits its value as a measure of quality. The goal of a machine learning (ML) method is a small loss on data points it has not seen, formalized as a small risk for data points that are realizations of independent and identically distributed (i.i.d.) random variables (RVs). Because $\learnthypothesis$ was picked to make the average on $\trainset$ small, the training error is optimistic: on average over the draw of the training set it lies below the risk, and the more the hypothesis space offers to choose from, the larger the gap (Hastie et al., 2009, Sect. 7.4). A hypothesis with zero training error may therefore be worthless on new data points. The validation error, computed on data points that did not enter the training, has no such optimism and estimates the risk; the difference between the two is the generalization gap.

B-degreeFig. 1 shows both errors for polynomials of degree $0$ to $12$ fitted by ERM to the forty days. The hypothesis spaces are nested, a polynomial of degree $d$ being one of degree $d+1$ with a zero coefficient, so the training error never increases with the degree: it falls from $29.7$ for the constant to $2.45$ for the line and $1.84$ for degree $12$. The risk, computed on $200000$ fresh days, is smallest for the line, $2.61$, and reaches $24.8$ at degree $12$, where the polynomial follows the noise of the forty recordings.

Figure 1 of the entry trainerr
Figure 1: Training error and risk of the polynomial of each degree fitted by ERM to forty days of weather recordings. The training error never increases with the degree, since each hypothesis space contains the previous one, while the risk is smallest for the straight line and grows by an order of magnitude beyond degree $10$. Data generated by pythondemos/trainerr.py
Read beside the validation error, the training error diagnoses an ML method: a training error far below the validation error signals overfitting, and a training error that is itself large signals underfitting, a hypothesis space too small or an optimizer that stopped short of the minimum (Hastie et al., 2009, Sect. 7.2). The three routes of regularization all raise the training error on purpose, in exchange for a smaller risk.

See also: training set, empirical risk, empirical risk minimization, loss, validation error, risk, generalization gap, overfitting, underfitting, hypothesis.

References

  1. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7

Cite this entry

@misc{dictml_trainerr,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {training error},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-29},
  url = {https://dictionaryofml.org/terms/trainerr.html}
}