Dictionary of Applied Machine Learning
Updated on 2026-10-07
See also logistic regression
The sigmoid function carries the
whole real line into the interval between $0$ and $1$, turning the
unbounded score that a classifier assigns to a
data point into a probability. It is the
function $\sigma: \reals \rightarrow (0,1)$ given by
\[
\sigma(z) \defeq \frac{1}{1 + e^{-z}}
\text{.}
\]
It rises from $\sigma(z) \rightarrow 0$ as $z \rightarrow -\infty$
to $\sigma(z) \rightarrow 1$ as $z \rightarrow +\infty$, and passes
through $\sigma(0) = 1/2$ (see Fig. 1). It is
also called the logistic function, and it is the softmax function for
two classes.
The sigmoid function turns a log-odds into a posterior probability. Let the binary label $\truelabel$ be either $1$ or $-1$. Write $a$ for the log-odds of the two classes at the feature vector $\featurevec$, \[ a \defeq \log \frac{\prob{\featurevec \mid \truelabel = 1} \, \prob{\truelabel = 1}} {\prob{\featurevec \mid \truelabel = -1} \, \prob{\truelabel = -1}} \text{.} \] The posterior probability of the label $1$ is then exactly $\prob{\truelabel = 1 \mid \featurevec} = \sigma(a)$, and the inverse recovers the log-odds, $a = \log\big(\sigma/(1 - \sigma)\big)$ (Bishop, 2006, Sect. 4.2). A linear model with model parameters $\weights$ scores the feature vector by $\weights^{\top} \featurevec$. Equating that score with $a$ is what turns $\sigma(\weights^{\top} \featurevec)$ into a posterior probability, so the reading is an assumption rather than a convention: that the log-odds is a linear function of the feature vector. The assumption holds exactly when the feature vectors of the two classes follow Gaussian random variables (Gaussian RVs) sharing one covariance matrix (Bishop, 2006, Sect. 4.2.1). Logistic regression instead fits $\weights$ to a training set without assuming any form for those feature vectors (Bishop, 2006, Sect. 4.3).
Synonyms: logistic function.
See also: activation function, rectified linear unit, softmax function, logistic regression, linear model, classifier, probability, smooth.
@misc{dictml_sigmoid,
author = {Jung, Alexander},
editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
title = {sigmoid function},
howpublished = {Dictionary of Applied Machine Learning (course edition)},
year = {2026},
doi = {10.5281/zenodo.21569296},
note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-07},
url = {https://dictionaryofml.org/terms/sigmoid.html}
}