Dictionary of Applied Machine Learning
Updated on 2026-10-07
See also artificial neural network gradient gradient descent
Backpropagation is an algorithm that computes the gradient of an objective function with respect to the model parameters of an artificial neural network (ANN). It runs in two passes. The forward pass sends a batch of data points through the ANN and evaluates the loss of the resulting predictions. The backward pass then applies the chain rule from the output back to the input, computing on the way the partial derivative with respect to every weight. Obtaining all of them this way costs a small multiple of one forward pass, where perturbing each weight separately would cost one pass per weight. Backpropagation is not itself an iteration: it yields one gradient, which a gradient step then consumes. Every training step of a deep net uses it.
An image classifier reads the
features of a photograph, passes them through layers
of neurons, and delivers at its output a
prediction of the label. Training it requires the
partial derivative of the loss with respect to each of
its millions of weights. Obtaining those one at a time, by
perturbing a single weight and evaluating the loss
again, would cost one pass through the network per weight.
Backpropagation delivers all of them in two passes
(Rumelhart et al., 1986; Goodfellow et al., 2016). It is an
algorithm that computes the gradient
$\nabla_{\weights} f(\weights)$ of an objective function $f(\weights)$
with respect to the model parameters $\weights$ of an artificial neural network (ANN),
which are the weights of its layers collected into one
vector. It applies the chain rule of calculus to evaluate the
partial derivatives of $f$ one layer at a time (see
Fig. 1).
Backpropagation is not itself an iteration: one forward and one backward pass yield one gradient. The iteration is gradient descent (GD), which consumes that gradient in a gradient step \[ \weights^{(\iteridx+1)} = \weights^{(\iteridx)} - \lrate \nabla f(\weights^{(\iteridx)}) \text{,} \] with step size $\lrate$, and calls backpropagation again at the new model parameters. The objective function $f(\weights)$ is the average loss on a training set, while the goal of training is a small loss on data points outside that training set. Backpropagation makes the search over $\weights$ affordable; it does not close that gap.
Every training step of a deep net uses backpropagation, once per batch. In image classification, for example, it computes the partial derivatives needed to update millions of weights spread over dozens of layers.
See also: artificial neural network, deep net, gradient, gradient descent, gradient step, layer, weight, model parameter, activation function, loss function.
@misc{dictml_backpropagation,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {backpropagation},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-07},
url = {https://dictionaryofml.org/terms/backpropagation.html}
}