{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "featurevec.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# feature vector \u2014 Python demo\n\nNumerical companion to the entry [feature vector](https://dictionaryofml.org/terms/featurevec.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nEvery block builds a feature vector as a one-dimensional numpy array, for a different kind of data point, and checks that the array is what the entry says a feature vector is: a point of a Euclidean space whose dimension is the number of features. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/featurevec.py`](https://dictionaryofml.org/terms/featurevec.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"featurevec.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nfeaturevec.py -- numerical companion to the glossary entry 'feature vector'.\n\nEvery block builds a feature vector as a one-dimensional numpy array, for a\ndifferent kind of data point, and checks that the array is what the entry\nsays a feature vector is: a point of a Euclidean space whose dimension is\nthe number of features. Self-contained (numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[B-house]  The entry's house offered for sale: living area, number of\n           rooms and year of construction give the feature vector\n           (85, 3, 1974), a point of R^3.\n[B-encode] Non-numeric features of the same house: a binary feature (does\n           it have a balcony) as 0 or 1, and a categorical feature with\n           K = 3 values (the heating type) one-hot encoded into K entries\n           of which exactly one equals 1. The house now sits in R^7.\n[B-text]   A text data point (a sentence of the Universal Declaration of\n           Human Rights): counting how often each word of a fixed word\n           list occurs gives one entry per word.\n[B-image]  An image data point: the red, green and blue intensities of its\n           pixels, one entry per intensity. The feature transformation\n           loses nothing here, so the image is recovered from its feature\n           vector by reshaping it.\n[B-audio]  An audio data point: the values its signal takes at successive\n           instants, one entry per instant.\n[B-dims]   The five feature vectors are points of Euclidean spaces of very\n           different dimension, although one feature transformation\n           produced each of them.\n\nOutputs\n-------\nfeaturevec_audio.csv : index j and entry x_j of the audio feature vector.\nfeaturevec_dims.csv  : kind of data point and dimension of its feature space.\nfeaturevec.png       : preview of all five feature vectors (checking only).\n\nData generated by pythondemos/featurevec.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\nrng = np.random.default_rng(20261007)\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\ndef is_feature_vector(x, d):\n    \"\"\"A feature vector is a one-dimensional array of d real numbers.\"\"\"\n    return (isinstance(x, np.ndarray) and x.ndim == 1 and x.shape == (d,)\n            and np.issubdtype(x.dtype, np.floating) and np.all(np.isfinite(x)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-house]** The entry's house offered for sale: living area, number of rooms and year of construction give the feature vector (85, 3, 1974), a point of R^3."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-house] a house offered for sale\")\n\nhouse = np.array([85.0, 3.0, 1974.0])\nprint(f\"    living area {house[0]:.0f} m^2, {house[1]:.0f} rooms, \"\n      f\"built {house[2]:.0f}  ->  x = {house}\")\ncheck(\"[B-house] the house is a point of R^3\", is_feature_vector(house, 3))\ncheck(\"[B-house] its entries are the three features in order\",\n      np.array_equal(house, np.array([85.0, 3.0, 1974.0])))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-encode]** Non-numeric features of the same house: a binary feature (does it have a balcony) as 0 or 1, and a categorical feature with K = 3 values (the heating type) one-hot encoded into K entries of which exactly one equals 1. The house now sits in R^7."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-encode] a binary and a categorical feature of the same house\")\n\nHEATING = [\"gas\", \"oil\", \"electric\"]\n\n\ndef encode_house(area, rooms, year, balcony, heating):\n    onehot = np.zeros(len(HEATING))\n    onehot[HEATING.index(heating)] = 1.0\n    return np.concatenate([[area, rooms, year, float(balcony)], onehot])\n\n\nhouse_enc = encode_house(85, 3, 1974, balcony=True, heating=\"oil\")\nonehot_part = house_enc[4:]\nprint(f\"    balcony -> {house_enc[3]:.0f}; heating 'oil' -> {onehot_part}\")\ncheck(\"[B-encode] the encoded house is a point of R^7\",\n      is_feature_vector(house_enc, 7))\ncheck(\"[B-encode] the binary feature is 0 or 1\", house_enc[3] in (0.0, 1.0))\ncheck(\"[B-encode] exactly one of the K = 3 one-hot entries equals 1\",\n      onehot_part.sum() == 1.0 and set(np.unique(onehot_part)) <= {0.0, 1.0})\ncheck(\"[B-encode] the position of that entry names the value\",\n      HEATING[int(np.argmax(onehot_part))] == \"oil\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-text]** A text data point (a sentence of the Universal Declaration of Human Rights): counting how often each word of a fixed word list occurs gives one entry per word."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-text] a text data point\")\n\nTEXT = (\"All human beings are born free and equal in dignity and rights. \"\n        \"They are endowed with reason and conscience and should act \"\n        \"towards one another in a spirit of brotherhood.\")\nWORD_LIST = [\"all\", \"and\", \"are\", \"born\", \"dignity\", \"equal\",\n             \"free\", \"human\", \"in\", \"of\", \"reason\", \"rights\"]\n\n\ndef text_features(sentence, words):\n    seen = sentence.lower().replace(\".\", \" \").replace(\",\", \" \").split()\n    return np.array([float(seen.count(w)) for w in words])\n\n\ntext_vec = text_features(TEXT, WORD_LIST)\nprint(f\"    {len(WORD_LIST)} words counted; \"\n      f\"'and' occurs {text_vec[WORD_LIST.index('and')]:.0f} times\")\ncheck(\"[B-text] the text is a point of R^12\",\n      is_feature_vector(text_vec, len(WORD_LIST)))\ncheck(\"[B-text] every entry is a non-negative whole count\",\n      np.all(text_vec >= 0) and np.all(text_vec == np.round(text_vec)))\ncheck(\"[B-text] the counts match a direct count of the words\",\n      text_vec.sum() == sum(TEXT.lower().replace(\".\", \" \").replace(\",\", \" \")\n                            .split().count(w) for w in WORD_LIST))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-image]** An image data point: the red, green and blue intensities of its pixels, one entry per intensity. The feature transformation loses nothing here, so the image is recovered from its feature vector by reshaping it."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-image] an image data point\")\n\nHEIGHT, WIDTH = 8, 8\nrows, cols = np.mgrid[0:HEIGHT, 0:WIDTH]\n# a bright disc on a darker ground, so the pixels show a shape rather than noise\ndisc = ((rows - 3.5) ** 2 + (cols - 3.5) ** 2) <= 2.5 ** 2\nimage = np.empty((HEIGHT, WIDTH, 3))\nimage[..., 0] = np.where(disc, 0.90, 0.15) + 0.04 * rng.normal(size=(HEIGHT, WIDTH))\nimage[..., 1] = np.where(disc, 0.75, 0.20) + 0.04 * rng.normal(size=(HEIGHT, WIDTH))\nimage[..., 2] = np.where(disc, 0.25, 0.45) + 0.04 * rng.normal(size=(HEIGHT, WIDTH))\nimage = np.clip(image, 0.0, 1.0)\nimage_vec = image.reshape(-1)\nprint(f\"    {HEIGHT} x {WIDTH} pixels, 3 intensities each  ->  \"\n      f\"{image_vec.size} entries\")\ncheck(\"[B-image] the image is a point of R^192\",\n      is_feature_vector(image_vec, HEIGHT * WIDTH * 3))\ncheck(\"[B-image] every intensity lies between 0 and 1\",\n      np.all(image_vec >= 0.0) and np.all(image_vec <= 1.0))\ncheck(\"[B-image] reshaping the feature vector recovers the image\",\n      np.array_equal(image_vec.reshape(HEIGHT, WIDTH, 3), image))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-audio]** An audio data point: the values its signal takes at successive instants, one entry per instant."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-audio] an audio data point\")\n\nNR_INSTANTS = 256\ninstants = np.arange(NR_INSTANTS)\naudio_vec = (0.6 * np.sin(2 * np.pi * 5 * instants / NR_INSTANTS)\n             + 0.25 * np.sin(2 * np.pi * 17 * instants / NR_INSTANTS)\n             + rng.normal(0.0, 0.03, NR_INSTANTS))\nprint(f\"    the signal is read at {NR_INSTANTS} instants  ->  \"\n      f\"{audio_vec.size} entries, largest magnitude \"\n      f\"{np.abs(audio_vec).max():.2f}\")\ncheck(\"[B-audio] the recording is a point of R^256\",\n      is_feature_vector(audio_vec, NR_INSTANTS))\ncheck(\"[B-audio] every entry stays within the range of the signal\",\n      np.all(np.abs(audio_vec) <= 1.0))\nrecomputed = (0.6 * np.sin(2 * np.pi * 5 * 40 / NR_INSTANTS)\n              + 0.25 * np.sin(2 * np.pi * 17 * 40 / NR_INSTANTS))\ncheck(\"[B-audio] entry j holds the value read at instant j\",\n      abs(audio_vec[40] - recomputed) < 0.15)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-dims]** The five feature vectors are points of Euclidean spaces of very different dimension, although one feature transformation produced each of them."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-dims] one construction, feature spaces of very different size\")\n\nKINDS = [(\"house\", house), (\"house, encoded\", house_enc),\n         (\"text\", text_vec), (\"image\", image_vec), (\"audio\", audio_vec)]\nfor name, vec in KINDS:\n    print(f\"    {name:16s} d = {vec.size}\")\ndims = [vec.size for _, vec in KINDS]\ncheck(\"[B-dims] each data point became a one-dimensional array\",\n      all(v.ndim == 1 for _, v in KINDS))\ncheck(\"[B-dims] the dimensions differ by about two orders of magnitude\",\n      min(dims) == 3 and max(dims) == 256)\n\nwith open(OUT_DIR / \"featurevec_audio.csv\", \"w\") as fh:\n    fh.write(\"j,x\\n\")\n    for j, value in enumerate(audio_vec):\n        fh.write(f\"{j},{value:.4f}\\n\")\n\nwith open(OUT_DIR / \"featurevec_dims.csv\", \"w\") as fh:\n    fh.write(\"kind,d\\n\")\n    for name, vec in KINDS:\n        fh.write(f\"{name.replace(',', '')},{vec.size}\\n\")\n\n\n# ---------------------------------------------------------------- preview\nfig, axs = plt.subplots(2, 2, figsize=(9.6, 6.2))\n\nax_house = axs[0, 0]\nlabels = [\"area\", \"rooms\", \"year\", \"balcony\", \"gas\", \"oil\", \"electric\"]\nax_house.bar(np.arange(house_enc.size), house_enc, color=\"0.35\")\nax_house.set_xticks(np.arange(house_enc.size))\nax_house.set_xticklabels(labels, rotation=45, ha=\"right\", fontsize=7)\nax_house.set_yscale(\"symlog\", linthresh=1.0)\nax_house.set_xlabel(\"feature\")\nax_house.set_ylabel(\"entry value (symmetric log scale)\")\nax_house.set_title(\"house: 3 numbers, a binary and a one-hot feature\",\n                   fontsize=9)\n\nax_text = axs[0, 1]\nax_text.bar(np.arange(text_vec.size), text_vec, color=\"0.35\")\nax_text.set_xticks(np.arange(text_vec.size))\nax_text.set_xticklabels(WORD_LIST, rotation=60, ha=\"right\", fontsize=7)\nax_text.set_xlabel(\"word of the fixed word list\")\nax_text.set_ylabel(\"number of occurrences\")\nax_text.set_title(\"text: one entry per word\", fontsize=9)\n\nax_img = axs[1, 0]\nax_img.imshow(image, interpolation=\"nearest\")\nax_img.set_xlabel(\"pixel column\")\nax_img.set_ylabel(\"pixel row\")\nax_img.set_title(f\"image: {HEIGHT}x{WIDTH} pixels give \"\n                 f\"{image_vec.size} entries\", fontsize=9)\n\nax_audio = axs[1, 1]\nax_audio.plot(instants, audio_vec, \"-\", color=\"black\", lw=0.9)\nax_audio.set_xlabel(\"instant j\")\nax_audio.set_ylabel(\"entry $x_j$\")\nax_audio.set_title(f\"audio: the signal read at {NR_INSTANTS} instants\",\n                   fontsize=9)\n\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"featurevec.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(\"wrote featurevec_audio.csv, featurevec_dims.csv, featurevec.png\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}