Dictionary of Applied Machine Learning
Typeset PDF version — the authoritative form of this entry
The interpretability of a machine learning (ML) method is the extent to which a human user can comprehend its behavior. Interpretability is closely related to the notion of simulatability: the user should be able to anticipate the predictions of the method on an arbitrary test set. Interpretability differs from explainability, which concerns understanding specific predictions with the help of provided explanations.
Interpretability is often defined, informally, as the ability to explain, or to present in understandable terms, to a human (Doshi-Velez and Kim, 2017). The international terminology standard for artificial intelligence (AI) does not define interpretability (Standardization and Commission, 2022). Its closest notion is predictability, the property of an artificial intelligence system (AI system) that enables reliable assumptions by users about the predictions delivered by a machine learning (ML) method (Standardization and Commission, 2022, Sect. 3.5.8 and 5.15.7).
An ML method is interpretable for a human user if they can comprehend the computational process of the method. Interpretability requires simulatability which is the ability of a human to mentally simulate the behavior of the ML method (Chen et al., 2018; Colin et al., 2022; Doshi-Velez and Kim, 2017; Hase and Bansal, 2020; Lipton, 2018). Simulatability is a stronger requirement compared to predictability as it covers not only the input-output behavior but also the internal computational processes of an AI system (Standardization and Commission, 2022, Sect. 5.15.7).
Fig. 1 depicts a training set and a test set.
The training set is used by two different methods that learn
hypotheses $\learnthypothesis$ and $\learnthypothesis'$.
The ML method producing the hypothesis
$\learnthypothesis$ is interpretable to a human user familiar with the
concept of a linear map. Since $\learnthypothesis$ is a
linear map, the user can anticipate the predictions of
$\learnthypothesis$ on the test set. In contrast, the ML
method producing $\learnthypothesis'$ is less interpretable: the
behavior of $\learnthypothesis'$ deviates from the user's expectations.
Interpretability can also be obtained by decomposing a learned hypothesis into components that are themselves interpretable. This property is called decomposability (Lipton, 2018). The hypothesis $\learnthypothesis$ in Fig. 1 decomposes into a slope and an intercept: the slope states how the prediction changes with the feature. A trained deep net lacks such a decomposition, since the activation of an individual neuron typically does not correspond to a human-interpretable concept. Mechanistic interpretability aims to recover interpretable components from a trained deep net (Olah et al., 2020; Elhage et al., 2021) by decomposition techniques such as sparse coding of activations (Cunningham et al., 2024).
See also: explainability, trustworthy artificial intelligence, regularization, local interpretable model-agnostic explanations, interpretable machine learning, mechanistic interpretability, transparency.
@misc{dictml_interpretability,
author = {Jung, Alexander},
title = {interpretability},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-08-06},
url = {https://dictionaryofml.org/terms/interpretability.html}
}