Dictionary of Applied Machine Learning

data imputation

Updated on 2026-10-07

▶ Run the Python demo open in Colab Typeset PDF Cite this entry

See also data point feature label feature vector training set empirical risk minimization linear regression baseline test set

Data imputation replaces the missing feature or label values of data points by estimates computed from the values that were recorded. It is one response to missing data, alongside discarding the incomplete data points. The filling becomes a machine learning (ML) problem once it is read as a prediction task: the quantity to estimate is treated as a label and the values that survive around it as a feature vector, so the intact data points form a training set on which a hypothesis is learned. For a photograph with lost pixels, each pixel is a data point whose feature vector holds the colors of its neighbors and whose label is its own color. Where a lost value sits among other lost values, the hypothesis is applied repeatedly until the estimates stop changing, which is a fixed point of one sweep. An imputed value is an estimate and not a measurement, so a method trained on imputed data points reports a training error that flatters it.

Definition

B-sceneAn aerial photograph of the vineyards at Rossatz on the Danube comes back from the survey with a rectangular patch of its pixels lost, because the sensor dropped them. The photograph is still wanted, so those pixels need values. Data imputation is that filling in: it replaces the missing feature or label values of data points by estimates computed from the values that were recorded (Abayomi et al., 2008). It is one response to missing data, alongside discarding the incomplete data points or using a method that tolerates the gaps.

B-setThe filling becomes a machine learning (ML) problem once it is read as a prediction task (Abayomi et al., 2008). For the photograph, each pixel is a data point. Its feature vector holds the red, green and blue values of the eight pixels surrounding it, which is $24$ numbers, and its label is its own triple of red, green and blue values (Fig. 1). The intact pixels then form a training set: each one supplies both a feature vector and the label that belongs with it, so a hypothesis can be learned from them by empirical risk minimization (ERM). Applying that hypothesis to a lost pixel yields the estimate that fills it.

Figure 1 of the entry dataimputation
Figure 1: One pixel of the photograph as a data point. Its feature vector collects the red, green and blue values of the eight surrounding pixels, and its label is its own triple of those values. An intact pixel supplies both, and so belongs in the training set
A lost pixel usually has lost pixels among its neighbors, so its feature vector is not available and the hypothesis cannot be applied once and be done. The lost pixels instead start at the average color of the training set and are re-predicted from their current neighbors, sweep after sweep. Writing $\vu$ for the values currently held by the lost pixels, one sweep applies an operator $T$ to them, \[ \vu^{(\iteridx+1)} = T\big(\vu^{(\iteridx)}\big) \text{,} \] and the filling stops at a fixed point of $T$, where another sweep leaves the values unchanged. The experiment pythondemos/dataimputation.py reaches one after $413$ sweeps on a view of $64 \times 96$ pixels with $6\%$ of them lost.

B-fillWhat imputation can and cannot recover is visible in the numbers of that experiment (Fig. 2). A linear regression hypothesis learned from the intact pixels halves the average squared error loss on the lost ones against the baseline that fills them all with the average color. The gain is not spread evenly: inside a region of nearly constant color the error is a quarter of that baseline, while on the pixels where the river meets the meadow it is as large. The filled river runs through the gap, but its banks come out blurred.

Figure 2 of the entry dataimputation
Figure 2: Average squared error loss on the lost pixels. The first two bars compare filling them all with the average color against the linear regression hypothesis learned from the intact pixels. The last two split that learned result by where the lost pixel sits, inside a region of nearly constant color or on the boundary between two regions. Data generated by pythondemos/dataimputation.py
An imputed value is an estimate and not a measurement, and a method trained on imputed data points cannot tell the two apart. The training error it reports therefore flatters it, since part of the training set was produced by a hypothesis rather than observed. Recording which values were imputed, and keeping them out of the test set, is what preserves an honest estimate of the risk.

Synonyms: imputation, missing-value imputation.

See also: missing data, data point, feature, label, feature vector, training set, empirical risk minimization, linear regression, fixed point, baseline, test set.

References

  1. Abayomi et al. (2008). Diagnostics for Multivariate Imputations. J. Roy. Statist. Soc.: Ser. C (Appl. Statist.). doi.org/10.1111/j.1467-9876.2007.00613.x

Cite this entry

@misc{dictml_dataimputation,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {data imputation},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-08},
  url = {https://dictionaryofml.org/terms/dataimputation.html}
}