Dictionary of Applied Machine Learning

$k$-fold cross-validation ($k$-fold CV)

Updated on 2026-09-30

Typeset PDF version — the authoritative form of this entry

$k$-fold cross-validation ($k$-fold CV) estimates the risk of an empirical risk minimization (ERM)-based machine learning (ML) method by dividing a dataset into $k$ folds of equal size and using each fold once as the validation set while the other $k-1$ folds form the training set. The $k$ resulting validation errors are averaged into a single estimate. Every data point serves both for training and for validation, unlike a single split into training set and validation set. The choice of $k$ trades bias against variance: small $k$ leaves each training set a small fraction of the dataset and overestimates the risk, while large $k$ makes the $k$ training sets overlap in almost every data point and their validation errors highly correlated. Leave-one-out cross-validation (LOO-CV) is the extreme case with one data point per fold, and stratified $k$-fold cross-validation preserves the class proportions of the full dataset in every fold.

Definition

Weather recordings from fifteen days relate a day's maximum daytime temperature $\truelabel$ to its morning minimum $\feature$. Splitting them once into a training set $\trainset$ of twelve days and a validation set $\valset$ of three computes the validation error from three data points only. Which days are held back also changes the verdict: Fig. 1 shows two such splits, and the hypothesis that empirical risk minimization (ERM) on a linear model delivers has a different slope in each. $k$-fold CV avoids the choice. The fifteen days are divided into five folds of three, each fold serves once as $\valset$ while the other four form $\trainset$, and the validation error of each iteration enters an average. Those five validation errors are $0.77$, $1.03$, $0.90$, $1.05$, and $1.17$, and their average $0.99$ is the $k$-fold CV estimate of the expected squared error loss of ERM on the linear model.

Figure 1 of the entry kfoldcv
Figure 1: Two of the five iterations of $k$-fold CV on the same fifteen days, with $k=5$. In the iteration $\foldidx$, the three days of fold $\foldidx$ are the validation set (open triangles) and the remaining twelve are the training set (filled circles); the line is the hypothesis $\learnthypothesis^{(\foldidx)}$ that ERM on a linear model delivers from that training set. Its slope is $0.76$ for $\foldidx = 1$ and $0.60$ for $\foldidx = 4$, and the validation error is $0.77$ against $1.05$: both the learned hypothesis and its assessment move with the split
$k$-fold CV estimates the expected loss (or risk) of an ERM-based machine learning (ML) method (Hastie et al., 2009, Sect. 7.10; Stone, 1974). It divides a dataset $\dataset$ evenly into $k$ subsets, the folds $\dataset^{(1)},\,\ldots,\, \dataset^{(k)}$ (Fig. 2). For each fold $\foldidx = 1, \,\ldots, \,k$, the model is trained on the union of all folds except $\dataset^{(\foldidx)}$ and the resulting hypothesis is evaluated on the held-out fold $\dataset^{(\foldidx)}$. The $k$ per-fold validation errors are then averaged to obtain the CV estimate of the expected loss.
Figure 2 of the entry kfoldcv
Figure 2: An instance of $k$-fold CV with $k=5$: the available dataset $\dataset$ is evenly divided into five folds $\dataset^{(1)},\,\ldots,\,\dataset^{(5)}$. Each row corresponds to one iteration; the shaded cell marks the fold used as validation set, while the remaining $k-1$ folds form the training set
Compared to a single split into training set and validation set, $k$-fold CV uses each data point for both training and validation. In Fig. 2, for example, each data point has been used four times for training and once for validation after all five iterations. For a typical ERM-based method, the $k$-fold CV estimate is more accurate than one from a single validation set holding a fraction $1/k$ of the entire dataset (Blum et al., 1999).

The choice of $k$ trades off bias against variance of the $k$-fold CV estimate (Hastie et al., 2009, Sect. 7.10.1). Small $k$ (e.g., $k=2$) gives higher bias because each training set contains only a fraction $(k-1)/k$ of the full dataset; if the learning curve is steep at this size, $k$-fold CV overestimates the true expected loss. Choosing large $k$ tends to reduce bias but can increase variance because the $k$ training sets differ in only a few data points, making the per-fold estimates highly correlated (Bengio and Grandvalet, 2004).

Several variants of $k$-fold CV adapt it to specific dataset structures:

See also: validation, validation error, generalization gap, learning curve, leave-one-out cross-validation, stratified $k$-fold cross-validation, model selection.

References

  1. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7
  2. Stone (1974). Cross-Validatory Choice and Assessment of Statistical Predictions. Journal of the Royal Statistical Society. Series B (Methodological).
  3. Blum et al. (1999). Beating the Hold-Out: Bounds for K-fold and Progressive Cross-Validation. Proceedings of the 12th Annual Conference on Computational Learning Theory (COLT).
  4. Bengio and Grandvalet (2004). No Unbiased Estimator of the Variance of K-Fold Cross-Validation. Journal of Machine Learning Research.
  5. Kohavi (1995). A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection. Proceedings of the 14th International Joint Conference on Artificial Intelligence (IJCAI).
  6. Kim (2009). Estimating Classification Error Rate: Repeated Cross-Validation, Repeated Hold-Out and Bootstrap. Computational Statistics \& Data Analysis.

Cite this entry

@misc{dictml_kfoldcv,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {$k$-fold cross-validation ($k$-fold CV)},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-30},
  url = {https://dictionaryofml.org/terms/kfoldcv.html}
}