Dictionary of Applied Machine Learning
Updated on 2026-10-04
See also feature data point classification explainability explainable artificial intelligence interpretable machine learning
An explanation accompanies a prediction delivered by a machine learning (ML) method and states which properties of the data point led to that prediction. Explanations can be text, a score per feature, a simple hypothesis that approximates the learned one near the data point, or a heat map drawn over an image. Two requirements can be contradictory. On the one hand, an explanation must be faithful: it must reflect the computation by which the learned hypothesis arrives at the prediction. On the other hand, it must be understandable: its user must be able to anticipate the predictions it accompanies. An explanation that agreed with the learned hypothesis everywhere would be that hypothesis again, so a simpler one may have to give up faithfulness somewhere.
One approach to enhance the transparency of a machine learning (ML) method for its human user is to provide an explanation alongside each prediction delivered by the method. Explanations can take different forms. For instance, they may consist of human-readable text or quantitative indicators, such as feature importance scores for the individual features of a given data point (Molnar, 2025), obtained, e.g., via the weights of a local linear approximation (Ribeiro et al., 2016).
Fig. 1 illustrates two types of explanations for a learned hypothesis $\learnthypothesis(\feature)$ of a single feature $\feature$. The first is a local linear approximation $g(\feature)$ of the nonlinear $\learnthypothesis(\feature)$ around a specific feature value $\feature'$. The second form of explanation depicted in the figure is a sparse set of predictions $\learnthypothesis(\feature^{(1)}), \learnthypothesis(\feature^{(2)}), \learnthypothesis(\feature^{(3)})$ at selected feature values, possibly chosen by the user.
B-trainAnother form of explanation is a heat map: an intensity
map that scores each region of an image by how much it drove
the prediction, drawn over the image itself
(Selvaraju et al., 2017). For a convolutional neural network (CNN), those scores are computed from
the network's own activations, which is what a
class activation map (CAM) does.
Whatever form it takes, an explanation carries two requirements: it must be (i) understandable and (ii) faithful. Understandable means that the user can work with the explanation on their own. Faithful means that the explanation reflects the computation the learned hypothesis carries out, rather than one that is easier to present. An intensity map over the pixels of an image is faithful when manipulating a few of the pixels it highlights changes the prediction while changing dark ones does not (see explainable artificial intelligence (XAI)). These two requirements can be contradictory. An unfaithful explanation can even be understandable and let the user anticipate well, whenever it tracks the prediction without carrying anything about the computation. For some saliency methods, the map computed from a network with random weights is nearly the same as the map computed from the trained network. Such a map traces the edges of the input image and carries nothing about the prediction, yet it looks like a plausible explanation (Adebayo et al., 2018). An explanation that agreed with the learned hypothesis everywhere would be that hypothesis again, so a simpler explanation may have to give up faithfulness somewhere.
B-textFig. 2 shows both requirements at
work on a prediction made from weather radar. The features are
the hourly precipitation over a 120 km box around Krems an der Donau, and the prediction answers whether it will rain at Krems two hours later. A CNN fitted to these images is explained by a CAM, which scores each cell of the image by how much it contributed to the prediction (Selvaraju et al., 2017). For the hour drawn, the network is certain of rain, and the CAM puts its weight on the bands of precipitation south and south-east of the town rather than on the rain already overhead. Setting the precipitation to zero in the 360 cells with the highest CAM scores lowers the predicted score by 1.57, against 0.26 for the 360 cells with the lowest scores, and the same holds on 63 of the 72 held-out hours. The same content can be delivered as text. The sentence under the panels is generated from the CAM by computing the radius that holds the cells it scores highest and the direction of their center of mass, so it is faithful by construction and can be checked against the radar image on the left.
pythondemos/explanation.py
See also: machine learning, prediction, feature, data point, classification, explainability, explainable artificial intelligence, interpretable machine learning, local interpretable model-agnostic explanations, SHapley Additive exPlanations, counterfactual, class activation map.
@misc{dictml_explanation,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {explanation},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-05},
url = {https://dictionaryofml.org/terms/explanation.html}
}