Dictionary of Applied Machine Learning
Updated on 2026-10-07
See also data point feature label feature vector training set empirical risk minimization linear regression baseline test set
Data imputation replaces the missing feature or label values of data points by estimates computed from the values that were recorded. It is one response to missing data, alongside discarding the incomplete data points. The filling becomes a machine learning (ML) problem once it is read as a prediction task: the quantity to estimate is treated as a label and the values that survive around it as a feature vector, so the intact data points form a training set on which a hypothesis is learned. For a photograph with lost pixels, each pixel is a data point whose feature vector holds the colors of its neighbors and whose label is its own color. Where a lost value sits among other lost values, the hypothesis is applied repeatedly until the estimates stop changing, which is a fixed point of one sweep. An imputed value is an estimate and not a measurement, so a method trained on imputed data points reports a training error that flatters it.
B-sceneAn aerial photograph of the vineyards at Rossatz on the Danube comes back from the survey with a rectangular patch of its pixels lost, because the sensor dropped them. The photograph is still wanted, so those pixels need values. Data imputation is that filling in: it replaces the missing feature or label values of data points by estimates computed from the values that were recorded (Abayomi et al., 2008). It is one response to missing data, alongside discarding the incomplete data points or using a method that tolerates the gaps.
B-setThe filling becomes a machine learning (ML) problem once it is read as a
prediction task (Abayomi et al., 2008). For the
photograph, each pixel is a data point. Its feature vector
holds the red, green and blue values of the eight pixels
surrounding it, which is $24$ numbers, and its label is its
own triple of red, green and blue values
(Fig. 1). The intact pixels then
form a training set: each one supplies both a feature vector and
the label that belongs with it, so a hypothesis can be
learned from them by empirical risk minimization (ERM). Applying that hypothesis to a
lost pixel yields the estimate that fills it.
pythondemos/dataimputation.py reaches one after $413$
sweeps on a view of $64 \times 96$ pixels with $6\%$ of them lost.
B-fillWhat imputation can and cannot recover is visible in the numbers of
that experiment (Fig. 2). A
linear regression hypothesis learned from the intact pixels halves
the average squared error loss on the lost ones against the
baseline that fills them all with the average color. The gain
is not spread evenly: inside a region of nearly constant color the
error is a quarter of that baseline, while on the pixels where
the river meets the meadow it is as large. The filled river runs
through the gap, but its banks come out blurred.
pythondemos/dataimputation.py
Synonyms: imputation, missing-value imputation.
See also: missing data, data point, feature, label, feature vector, training set, empirical risk minimization, linear regression, fixed point, baseline, test set.
@misc{dictml_dataimputation,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {data imputation},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-08},
url = {https://dictionaryofml.org/terms/dataimputation.html}
}