Dictionary of Applied Machine Learning
Updated on 2026-09-30
Typeset PDF version — the authoritative form of this entry
$k$-fold cross-validation ($k$-fold CV) estimates the risk of an empirical risk minimization (ERM)-based machine learning (ML) method by dividing a dataset into $k$ folds of equal size and using each fold once as the validation set while the other $k-1$ folds form the training set. The $k$ resulting validation errors are averaged into a single estimate. Every data point serves both for training and for validation, unlike a single split into training set and validation set. The choice of $k$ trades bias against variance: small $k$ leaves each training set a small fraction of the dataset and overestimates the risk, while large $k$ makes the $k$ training sets overlap in almost every data point and their validation errors highly correlated. Leave-one-out cross-validation (LOO-CV) is the extreme case with one data point per fold, and stratified $k$-fold cross-validation preserves the class proportions of the full dataset in every fold.
Weather recordings from fifteen days relate a day's maximum
daytime temperature $\truelabel$ to its morning minimum
$\feature$. Splitting them once into a training set $\trainset$
of twelve days and a validation set $\valset$ of three computes the
validation error from three data points only. Which days are held
back also changes the verdict: Fig. 1
shows two such splits, and the hypothesis that empirical risk minimization (ERM) on a
linear model delivers has a different slope in each. $k$-fold CV
avoids the choice. The fifteen days are divided into five folds of
three, each fold serves once as $\valset$ while the other four form
$\trainset$, and the validation error of each iteration enters an
average. Those five validation errors are $0.77$, $1.03$, $0.90$,
$1.05$, and $1.17$, and their average $0.99$ is the $k$-fold CV
estimate of the expected squared error loss of ERM on the
linear model.
The choice of $k$ trades off bias against variance of the $k$-fold CV estimate (Hastie et al., 2009, Sect. 7.10.1). Small $k$ (e.g., $k=2$) gives higher bias because each training set contains only a fraction $(k-1)/k$ of the full dataset; if the learning curve is steep at this size, $k$-fold CV overestimates the true expected loss. Choosing large $k$ tends to reduce bias but can increase variance because the $k$ training sets differ in only a few data points, making the per-fold estimates highly correlated (Bengio and Grandvalet, 2004).
Several variants of $k$-fold CV adapt it to specific dataset structures:
See also: validation, validation error, generalization gap, learning curve, leave-one-out cross-validation, stratified $k$-fold cross-validation, model selection.
@misc{dictml_kfoldcv,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {$k$-fold cross-validation ($k$-fold CV)},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-30},
url = {https://dictionaryofml.org/terms/kfoldcv.html}
}