Dictionary of Applied Machine Learning

feature vector

Updated on 2026-10-07

▶ Run the Python demo open in Colab Typeset PDF Cite this entry

See also feature data point kernel method Hilbert space

A feature vector is a list of numbers describing a data point, with one entry for each of its features. A feature is an attribute of the data point that is measured or computed without human supervision, so the length of the vector is the number of such attributes. A feature transformation computes the feature vector from the data point and delivers a point of the feature space, which a hypothesis map then reads to predict the label. Non-numeric features enter as numbers: a binary attribute as $0$ or $1$, a categorical one through one-hot encoding. Most machine learning (ML) methods use feature vectors in a Euclidean space whose dimension is the number of features. A feature vector can more generally be an element of a Hilbert space, which is how a kernel method uses one without ever forming it, reaching it only through inner products.

Definition

B-houseA house offered for sale can be described to a machine learning (ML) method by three numbers: $85$ square meters of living area, $3$ rooms, and $1974$ as the year of construction. That list $\big(85, 3, 1974\big)$ is the feature vector of the house. In general, a feature vector gathers into one vector the features of a data point, those attributes of the data point that are measured or computed without human supervision (Duda et al., 2001). A list of $\nrfeatures$ real numbers, for a natural number $\nrfeatures$, is an element of the Euclidean space $\reals^{\nrfeatures}$ (Strang, 2016, Sect. 1.1).

B-encodeThe feature vector of a data point $\datapoint$ is written $\featurevec = \big(\feature_{1}, \,\ldots, \,\feature_{\nrfeatures}\big)^{\top}$, and its entry $\feature_{\featureidx}$ is one feature of that data point. It is obtained from the data point by a feature transformation $\featuretrafovec : \datapointspace \rightarrow \featurespace$, $\featurevec = \featuretrafovec(\datapoint)$, which views the data point itself as its raw features and delivers a point of the feature space $\featurespace$ (Fig. 1). A hypothesis map reads the feature vector and delivers a prediction of the label of the data point.

Figure 1 of the entry featurevec
Figure 1: A feature vector $\featurevec$ stacks the individual features $\feature_{1}, \,\ldots, \,\feature_{\nrfeatures}$ of a data point $\datapoint$, computed by the feature transformation $\featuretrafovec$. It is a point in the feature space $\featurespace$
Non-numeric features can also sit in the entries of a feature vector, once their values are encoded as real numbers (Hastie et al., 2009; Goodfellow et al., 2016). A binary feature, such as whether the house has a balcony, becomes $0$ or $1$. A categorical feature with $K$ possible values, such as the heating type (gas, oil, or electric), is one-hot encoded: $K$ entries of the feature vector carry the value $0$ except for a single entry equal to $1$, whose position names the value.

B-textA feature transformation also reaches data points that are not records of numbers. A text becomes a feature vector by counting how often each entry of a fixed word list appears in it, which gives one entry per listed word. A digital image is an array of pixels, and the red, green and blue intensities of those pixels become one entry each. An audio recording becomes a feature vector through the values its signal takes at successive instants, one entry per instant. The feature transformation differs in each case, and so does the dimension $\nrfeatures$ of the resulting feature space: the house above needs seven entries once its balcony and heating type are encoded, while an audio recording read at $256$ instants needs $256$ (Fig. 2).

Figure 2 of the entry featurevec
Figure 2: The feature vector of an audio recording read at $256$ instants. Entry $\feature_{\featureidx}$ is the value its signal takes at instant $\featureidx$, so this data point is a point of $\reals^{256}$. Data generated by pythondemos/featurevec.py
Most ML methods use feature vectors in a finite-dimensional Euclidean space $\reals^{\nrfeatures}$, whose dimension $\nrfeatures$ is the number of features. More generally, a feature vector can be an element of an abstract Hilbert space $\hilbertspace$. One way to build a feature vector of that kind starts from a kernel $\kernel: \featurespace \times \featurespace \rightarrow \reals$ on a given feature space $\featurespace$. The kernel assigns to each $\featurevec \in \featurespace$ the function $\kernelmap{\featurevec}{\cdot}: \featurespace \rightarrow \reals$, whose value at $\featurevec' \in \featurespace$ is $\kernelmap{\featurevec}{\featurevec'}$. That function is itself a vector in a reproducing kernel Hilbert space (RKHS) $\hilbertspace$, a Hilbert space whose elements are functions on $\featurespace$ (Aronszajn, 1950; Schölkopf and Smola, 2002). A kernel method uses these feature vectors without ever forming them explicitly: it reaches them only through their inner products, which the reproducing property reduces to kernel evaluations, $\innerprodgeneric{\kernelmap{\featurevec}{\cdot}}{\kernelmap{\featurevec'}{\cdot}}{\hilbertspace} = \kernelmap{\featurevec}{\featurevec'}$ (Schölkopf and Smola, 2002).

Synonyms: covariate vector, input vector, predictor vector.

See also: feature, feature space, feature transformation, data point, vector, Euclidean space, kernel, reproducing kernel Hilbert space, kernel method, Hilbert space.

References

  1. Duda et al. (2001). Pattern Classification. Wiley.
  2. Strang (2016). Introduction to Linear Algebra. Wellesley-Cambridge Press.
  3. Hastie et al. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Science+Business Media. doi.org/10.1007/978-0-387-84858-7
  4. Goodfellow et al. (2016). Deep Learning. MIT Press.
  5. Aronszajn (1950). Theory of Reproducing Kernels. Trans. Am. Math. Soc.. doi.org/10.1090/S0002-9947-1950-0051437-7
  6. Schölkopf and Smola (2002). Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT Press. doi.org/10.7551/mitpress/4175.001.0001

Cite this entry

@misc{dictml_featurevec,
  author = {Jung, Alexander},
  editor = {Olioumtsevits, Konstantina and Schnoor, Ekkehard},
  title = {feature vector},
  howpublished = {Dictionary of Applied Machine Learning (course edition)},
  year = {2026},
  doi = {10.5281/zenodo.21569296},
  note = {ISBN 978-952-64-3013-3, CC BY 4.0, retrieved 2026-10-07},
  url = {https://dictionaryofml.org/terms/featurevec.html}
}