Dictionary of Applied Machine Learning

backpropagation

Updated on 2026-10-07

Typeset PDF Cite this entry

See also artificial neural network gradient gradient descent

Backpropagation is an algorithm that computes the gradient of an objective function with respect to the model parameters of an artificial neural network (ANN). It runs in two passes. The forward pass sends a batch of data points through the ANN and evaluates the loss of the resulting predictions. The backward pass then applies the chain rule from the output back to the input, computing on the way the partial derivative with respect to every weight. Obtaining all of them this way costs a small multiple of one forward pass, where perturbing each weight separately would cost one pass per weight. Backpropagation is not itself an iteration: it yields one gradient, which a gradient step then consumes. Every training step of a deep net uses it.

Definition

An image classifier reads the features of a photograph, passes them through layers of neurons, and delivers at its output a prediction of the label. Training it requires the partial derivative of the loss with respect to each of its millions of weights. Obtaining those one at a time, by perturbing a single weight and evaluating the loss again, would cost one pass through the network per weight. Backpropagation delivers all of them in two passes (Rumelhart et al., 1986; Goodfellow et al., 2016). It is an algorithm that computes the gradient $\nabla_{\weights} f(\weights)$ of an objective function $f(\weights)$ with respect to the model parameters $\weights$ of an artificial neural network (ANN), which are the weights of its layers collected into one vector. It applies the chain rule of calculus to evaluate the partial derivatives of $f$ one layer at a time (see Fig. 1).

Figure 1 of the entry backpropagation
Figure 1: Backpropagation through an ANN with one hidden layer. The solid arrows are the forward pass: the features $\feature_{1},\feature_{2},\feature_{3}$ reach the neurons $h_{1},h_{2},h_{3}$ through the weights $\mW^{(1)}$, then the output $\predictedlabel$ through $\mW^{(2)}$, and the loss $\loss$ compares that output to the label. The dashed arrow is the backward pass, which runs the other way and yields the partial derivatives of $\loss$ with respect to the weights of both layers
In the forward pass, a batch of data points travels through the ANN: each layer applies its weights and its activation function, producing predictions at the output, and the loss function compares those predictions to the true labels. In the backward pass, the partial derivatives of the loss with respect to the weights of each layer are computed recursively from the output back to the input, which yields the whole gradient $\nabla_{\weights} f(\weights)$. Obtaining every partial derivative this way costs a small multiple of one forward pass, instead of one pass per weight.

Backpropagation is not itself an iteration: one forward and one backward pass yield one gradient. The iteration is gradient descent (GD), which consumes that gradient in a gradient step \[ \weights^{(\iteridx+1)} = \weights^{(\iteridx)} - \lrate \nabla f(\weights^{(\iteridx)}) \text{,} \] with step size $\lrate$, and calls backpropagation again at the new model parameters. The objective function $f(\weights)$ is the average loss on a training set, while the goal of training is a small loss on data points outside that training set. Backpropagation makes the search over $\weights$ affordable; it does not close that gap.

Every training step of a deep net uses backpropagation, once per batch. In image classification, for example, it computes the partial derivatives needed to update millions of weights spread over dozens of layers.

See also: artificial neural network, deep net, gradient, gradient descent, gradient step, layer, weight, model parameter, activation function, loss function.

References

  1. Rumelhart et al. (1986). Learning Representations by Back-Propagating Errors. Nature. doi.org/10.1038/323533a0
  2. Goodfellow et al. (2016). Deep Learning. MIT Press.

Cite this entry

@misc{dictml_backpropagation,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {backpropagation},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-07},
  url = {https://dictionaryofml.org/terms/backpropagation.html}
}