Dictionary of Applied Machine Learning

sigmoid function

Updated on 2026-10-07

Typeset PDF Cite this entry

See also logistic regression

The sigmoid function carries the whole real line into the interval between $0$ and $1$, turning the unbounded score that a classifier assigns to a data point into a probability. It is the function $\sigma: \reals \rightarrow (0,1)$ given by \[ \sigma(z) \defeq \frac{1}{1 + e^{-z}} \text{.} \] It rises from $\sigma(z) \rightarrow 0$ as $z \rightarrow -\infty$ to $\sigma(z) \rightarrow 1$ as $z \rightarrow +\infty$, and passes through $\sigma(0) = 1/2$ (see Fig. 1). It is also called the logistic function, and it is the softmax function for two classes.

Figure 1 of the entry sigmoid
Figure 1: The sigmoid function $\sigma(z) = 1/(1 + e^{-z})$ carries the real line into the open interval $(0,1)$, with $\sigma(0) = 1/2$
The sigmoid function is smooth and strictly increasing, and it satisfies the symmetry $\sigma(-z) = 1 - \sigma(z)$ (Bishop, 2006, Sect. 4.2). Its derivative is available in closed form, and in terms of the function itself, \[ \sigma'(z) = \sigma(z)\big(1 - \sigma(z)\big) \text{.} \] Once $\sigma(z)$ has been evaluated, its derivative therefore costs one multiplication. That is what makes the sigmoid cheap to differentiate inside an artificial neural network (ANN). The product $\sigma(z)(1 - \sigma(z))$ is at most $1/4$, and it approaches $0$ for large $|z|$. A neuron whose score is large in magnitude therefore passes almost no gradient back. That is why the rectified linear unit (ReLU) replaced the sigmoid in the hidden layers of deep nets (Goodfellow et al., 2016, Sect. 6.3.2).

The sigmoid function turns a log-odds into a posterior probability. Let the binary label $\truelabel$ be either $1$ or $-1$. Write $a$ for the log-odds of the two classes at the feature vector $\featurevec$, \[ a \defeq \log \frac{\prob{\featurevec \mid \truelabel = 1} \, \prob{\truelabel = 1}} {\prob{\featurevec \mid \truelabel = -1} \, \prob{\truelabel = -1}} \text{.} \] The posterior probability of the label $1$ is then exactly $\prob{\truelabel = 1 \mid \featurevec} = \sigma(a)$, and the inverse recovers the log-odds, $a = \log\big(\sigma/(1 - \sigma)\big)$ (Bishop, 2006, Sect. 4.2). A linear model with model parameters $\weights$ scores the feature vector by $\weights^{\top} \featurevec$. Equating that score with $a$ is what turns $\sigma(\weights^{\top} \featurevec)$ into a posterior probability, so the reading is an assumption rather than a convention: that the log-odds is a linear function of the feature vector. The assumption holds exactly when the feature vectors of the two classes follow Gaussian random variables (Gaussian RVs) sharing one covariance matrix (Bishop, 2006, Sect. 4.2.1). Logistic regression instead fits $\weights$ to a training set without assuming any form for those feature vectors (Bishop, 2006, Sect. 4.3).

Synonyms: logistic function.

See also: activation function, rectified linear unit, softmax function, logistic regression, linear model, classifier, probability, smooth.

References

  1. Bishop (2006). Pattern Recognition and Machine Learning. Springer Science+Business Media. doi.org/10.1007/978-0-387-45528-0
  2. Goodfellow et al. (2016). Deep Learning. MIT Press.

Cite this entry

@misc{dictml_sigmoid,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {sigmoid function},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-07},
  url = {https://dictionaryofml.org/terms/sigmoid.html}
}