Dictionary of Applied Machine Learning

logistic loss

Updated on 2026-09-30

Typeset PDF version — the authoritative form of this entry

The logistic loss is a convex and smooth surrogate for the $0/1$ loss in binary classification, defined for a label $\truelabel \in \{-1, +1\}$ and a real-valued hypothesis $\hypothesis$ as a decreasing function of the margin $\truelabel \cdot \hypothesis(\featurevec)$. Like the hinge loss, it uses a data point only through that margin. Rescaled by its value at margin zero it is an upper bound on the $0/1$ loss, with equality there, so a small logistic loss forces a small $0/1$ loss. Unlike the hinge loss it is differentiable everywhere, which suits it to gradient-based optimization methods such as gradient descent (GD). Running empirical risk minimization (ERM) with a linear hypothesis and the logistic loss gives logistic regression.

Definition

A spam filter either flags an email or does not, so its mistakes are counted rather than measured: the $0/1$ loss charges $1$ for every wrong answer. That count is flat almost everywhere as a function of the model parameters, so it gives gradient descent (GD) no direction to move in. The logistic loss replaces it by a convex and smooth surrogate that lies above it.

Consider a binary classification problem with label $\truelabel \in \{-1, +1\}$ and a real-valued hypothesis $\hypothesis: \featurespace \rightarrow \reals$. On a data point $(\featurevec, \truelabel)$ the logistic loss is (Bishop, 2006) \begin{equation} \label{equ_log_loss_gls_dict} \lossfunc{(\featurevec, \truelabel)}{\hypothesis} \defeq \log\big(1 + \exp(-\truelabel \cdot \hypothesis(\featurevec))\big) \text{.} \end{equation} Like the hinge loss, the logistic loss depends on the margin $\truelabel \cdot \hypothesis(\featurevec)$ (see $0/1$ loss), and after division by $\log 2$ it is a convex upper bound on the $0/1$ loss, with equality at margin $0$ (see Fig. 1).

Figure 1 of the entry logloss
Figure 1: The logistic loss divided by $\log 2$ (solid) and the $0/1$ loss (dashed) as functions of the margin $\truelabel \cdot \hypothesis(\featurevec)$. The rescaled logistic loss lies on or above the $0/1$ loss for every margin, with equality at margin $0$; the rescaling is needed because $\log\big(1 + \exp(0)\big) = \log 2 < 1$
Unlike the hinge loss, the logistic loss is everywhere differentiable. It is therefore well suited for gradient-based optimization methods such as GD. Empirical risk minimization (ERM) with a linear hypothesis and the logistic loss yields logistic regression. For example, in medical diagnosis, logistic regression trained with the logistic loss predicts whether a patient has a disease ($\truelabel = +1$) or not ($\truelabel = -1$) based on clinical features.

Synonyms: log loss, binary cross-entropy loss.

See also: $\bf 0/1$ loss, hinge loss, classifier, logistic regression, gradient descent, convex, binary classification.

References

  1. Bishop (2006). Pattern Recognition and Machine Learning. Springer Science+Business Media. doi.org/10.1007/978-0-387-45528-0

Cite this entry

@misc{dictml_logloss,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {logistic loss},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-09-30},
  url = {https://dictionaryofml.org/terms/logloss.html}
}