Dictionary of Applied Machine Learning

linear classifier

Updated on 2026-10-10

▶ Run the Python demo open in Colab Typeset PDF Cite this entry

See also binary classification logistic regression support vector machine empirical risk minimization

A linear classifier is a classifier whose decision regions are separated by hyperplanes. For binary classification it predicts the sign of a linear model hypothesis $\hypothesis(\featurevec) = \weights^{\top} \featurevec + \offset$, whose decision boundary is the hyperplane with normal vector $\weights$. The hypothesis value $\hypothesis(\featurevec)$ is the distance of $\featurevec$ from that hyperplane, multiplied by $\normgeneric{\weights}{2}$, and its sign names the side. That distance bounds the feature perturbations the prediction tolerates. The model parameters are learned by empirical risk minimization (ERM) with a convex surrogate for the $0/1$ loss: logistic regression uses the logistic loss and the support vector machine (SVM) the hinge loss. More than two label values are handled by one linear model hypothesis per value and the largest-score rule, which makes every decision region an intersection of halfspaces.

Definition

B-fitMapping the vineyards of a wine region starts from an aerial photograph. The photograph is cut into square patches of $25$ m side length, and each patch becomes a data point. Two numbers describe a patch: its contrast $\feature_{1}$, the standard deviation of the brightness of its pixels, and its greenness $\feature_{2}$, the relative difference between the average green and the average red channel value, in percent. The two numbers form the feature vector $\featurevec = (\feature_{1}, \feature_{2})^{\top} \in \reals^{2}$, and the label is $\truelabel = +1$ for a patch showing a vineyard and $\truelabel = -1$ otherwise. Vine rows alternate foliage with bare soil, so a vineyard patch is more contrasted and less green than the forest and meadow around it. A straight line separates most of the vineyard patches from the rest (Fig. 1).

Figure 1 of the entry linclass
Figure 1: Data points of the vineyard map: each patch of the aerial photograph is placed at its contrast ($\feature_{1}$) and its greenness ($\feature_{2}$). The solid line is the decision boundary $\hypothesis(\featurevec) = 0$ of a linear classifier learned from these data points, the arrow is its normal vector $\weights$, and the dotted segment drops one vineyard patch (filled black circle) perpendicularly onto the decision boundary. The length of that segment is $|\hypothesis(\featurevec)| / \normgeneric{\weights}{2}$. The feature vectors are drawn from a two-component Gaussian model whose means and covariance matrices are the patch statistics of an aerial photograph of the Wachau valley. Data generated by pythondemos/linclass.py
A linear classifier is a classifier whose decision regions are separated by hyperplanes in $\reals^{\featuredim}$ (Bishop, 2006, Sect. 4.1.1; Shalev-Shwartz and Ben-David, 2014, Sect. 9.1). For binary classification, with label space $\labelspace = \{-1, +1\}$, such a classifier is obtained from a linear model. A hypothesis \[ \hypothesis(\featurevec) = \weights^{\top} \featurevec + \offset \text{,} \] with model parameters $\weights \in \reals^{\featuredim} \setminus \{\mathbf{0}\}$ and $\offset \in \reals$, delivers the prediction $\predictedlabel = \operatorname{sign}(\hypothesis(\featurevec))$. Its decision boundary is the hyperplane \[ \{ \featurevec \in \reals^{\featuredim} : \weights^{\top} \featurevec + \offset = 0 \} \text{,} \] with normal vector $\weights$, and the two decision regions are the two halfspaces it bounds.

The number $\hypothesis(\featurevec)$ carries more than its sign. Consider a feature vector $\featurevec$ and any vector $\featurevec'$ of the decision boundary. Subtracting $\weights^{\top} \featurevec' + \offset = 0$ from $\hypothesis(\featurevec)$ gives \[ \hypothesis(\featurevec) = \weights^{\top} \big( \featurevec - \featurevec' \big) \text{.} \] The Cauchy-Schwarz inequality bounds the right-hand side in magnitude by $\normgeneric{\weights}{2} \normgeneric{\featurevec - \featurevec'}{2}$. Every vector of the decision boundary is therefore at distance at least $|\hypothesis(\featurevec)| / \normgeneric{\weights}{2}$ from $\featurevec$. That distance is attained by the orthogonal projection \[ \featurevec' = \featurevec - \frac{\hypothesis(\featurevec)} {\normgeneric{\weights}{2}^{2}} \, \weights \text{,} \] which satisfies $\weights^{\top} \featurevec' + \offset = 0$. The distance of $\featurevec$ from the decision boundary is thus $|\hypothesis(\featurevec)| / \normgeneric{\weights}{2}$, and the hypothesis value $\hypothesis(\featurevec)$ is that distance multiplied by $\normgeneric{\weights}{2}$, with the sign naming the side (Bishop, 2006, Sect. 4.1.1; Hastie et al., 2009, Sect. 4.5, Eq. (4.40)). Fig. 1 shows this distance for one vineyard patch.

B-scaleThe hypothesis value by itself does not measure that distance. Let $s$ be a positive number. Replacing the model parameters $(\weights, \offset)$ by $(s \weights, s \offset)$ leaves the decision boundary and every prediction unchanged, while every hypothesis value is multiplied by $s$. Dividing by $\normgeneric{\weights}{2}$ removes this freedom, which is why the distance and not $\hypothesis(\featurevec)$ is the quantity with a geometric meaning.

B-robustThe distance also bounds the feature perturbations that a prediction tolerates. Perturbing $\featurevec$ by a vector $\boldsymbol{\delta}$ changes the hypothesis value by $\weights^{\top} \boldsymbol{\delta}$, which the Cauchy-Schwarz inequality bounds in magnitude by $\normgeneric{\weights}{2} \normgeneric{\boldsymbol{\delta}}{2}$. The sign of $\hypothesis(\featurevec)$, and with it the prediction, therefore survives every perturbation with $\normgeneric{\boldsymbol{\delta}}{2} < |\hypothesis(\featurevec)| / \normgeneric{\weights}{2}$. A vineyard patch far from the decision boundary keeps its prediction under a measurement error that flips the prediction for a patch close to it. Multiplying $\hypothesis(\featurevec)$ by the label $\truelabel$ gives the margin, which is positive exactly on the correctly classified data points. For a linearly separable training set and a sufficiently small regularization parameter $\regparam$, the support vector machine (SVM) learns the linear classifier whose smallest margin is largest (Shalev-Shwartz and Ben-David, 2014, Sect. 15.1, Eq. (15.1)).

The model parameters $(\weights, \offset)$ are learned from a training set $\trainset = \{ (\featurevec^{(\sampleidx)}, \truelabel^{(\sampleidx)}) \}_{\sampleidx = 1}^{\samplesize}$ by empirical risk minimization (ERM). Counting the mistakes with the $0/1$ loss gives an objective function that is piecewise constant, and minimizing it is computationally hard outside the linearly separable case (Shalev-Shwartz and Ben-David, 2014, Sect. 9.1). Practical methods replace the $0/1$ loss by a convex loss of the margin: logistic regression uses the logistic loss and the SVM uses the hinge loss. Both learn a hypothesis from the same hypothesis space, the linear model, and differ only in the loss. Both objective functions are convex: gradient descent (GD) minimizes the smooth one of logistic regression, and subgradient descent with a subgradient the non-smooth one of the SVM. Adding a penalty term $\regparam \normgeneric{\weights}{2}^{2}$ to the average loss, one of the three routes that regularization distinguishes, keeps $\normgeneric{\weights}{2}$ small and the margins wide (Shalev-Shwartz and Ben-David, 2014, Sect. 15.2).

For more than two label values, with $\labelspace = \{1, \ldots, \nrcategories\}$, a linear classifier uses one linear model hypothesis $g_{\clusteridx}(\featurevec) = \big( \weights^{(\clusteridx)} \big)^{\top} \featurevec + \offset^{(\clusteridx)}$ per label value $\clusteridx \in \labelspace$ and predicts the value with the largest score, $\predictedlabel = \argmax_{\clusteridx \in \labelspace} g_{\clusteridx}(\featurevec)$. The decision region $\decreg{\clusteridx}$ is then an intersection of $\nrcategories - 1$ halfspaces, one per competing label value, and is therefore a convex set; the border between $\decreg{\clusteridx}$ and $\decreg{\clusteridx'}$ lies in the hyperplane on which the two scores agree (Bishop, 2006, Sect. 4.1.2, Eq. (4.10)). The resulting partition of the feature space is shown in Fig. 2.

Figure 2 of the entry linclass
Figure 2: A linear classifier for $\nrcategories = 3$ label values partitions the feature space $\reals^{2}$ into three decision regions $\decreg{1}, \decreg{2}, \decreg{3}$. Each border lies in the hyperplane on which the two competing scores agree, and each decision region is an intersection of halfspaces and therefore convex
A linear classifier is linear in the features, not in the raw measurements a data point comes with. Applying a feature transformation $\featuretrafo: \reals^{\featuredim} \to \reals^{\featuredim'}$ first and a linear classifier to $\featuretrafo(\featurevec)$ afterwards gives a decision boundary that is a hyperplane in $\reals^{\featuredim'}$ and a curved surface in the original feature space. A kernel method constructs such a feature transformation implicitly, and a deep net learns one from the training set, its last layer being a linear classifier on the features the earlier layers compute (Bishop and Bishop, 2024, Sect. 6.3.3).

See also: classifier, binary classification, decision boundary, decision region, hyperplane, halfspace, linear model, margin, logistic regression, support vector machine, empirical risk minimization, $\bf 0/1$ loss, feature transformation.

References

  1. Bishop (2006). Pattern Recognition and Machine Learning. Springer Science+Business Media. doi.org/10.1007/978-0-387-45528-0
  2. Shalev-Shwartz and Ben-David (2014). Understanding Machine Learning: From Theory to Algorithms. Cambridge Univ. Press. doi.org/10.1017/cbo9781107298019
  3. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7
  4. Bishop and Bishop (2024). Deep Learning: Foundations and Concepts. Springer Nature. doi.org/10.1007/978-3-031-45468-4

Cite this entry

@misc{dictml_linclass,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {linear classifier},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-10},
  url = {https://dictionaryofml.org/terms/linclass.html}
}