{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "feature.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# feature \u2014 Python demo\n\nNumerical companion to the entry [feature](https://dictionaryofml.org/terms/feature.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/feature.py`](https://dictionaryofml.org/terms/feature.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"feature.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nfeature.py \u2014 numerical companion to the glossary entry 'feature'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-featmap]   Features are assembled into a feature vector by a feature\n              map acting on the data point (its raw features): sampled\n              signal values of an audio-like waveform form the feature\n              vector x = Phi(z).\n[P-transform] New features from arithmetic transformations of existing\n              ones: augmenting a scalar feature x with x^2 turns an\n              unlearnable quadratic relation into one a linear hypothesis map\n              fits (the average training loss collapses).\n[P-dtft]      DFT-magnitude features are unchanged by time shifts of the\n              signal, while raw signal-value features change \u2014 the\n              shift-invariance claim, checked numerically.\n[P-activation] The activation of a neuron is a derived feature: a fixed\n              random one-layer network turns raw features into\n              activations, and a linear hypothesis map on these activations fits\n              a nonlinear relation better than on the raw features.\n[P-mechfeat]  The narrower mechanistic-interpretability sense: a feature\n              as the projection of the activation vector onto a\n              direction \u2014 the projection of activations onto a planted\n              direction recovers a planted concept (correlation with\n              the concept indicator is high).\n\nOutputs\n-------\nfeature.png : preview figure (checking only).\n\nData generated by pythondemos/feature.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-featmap]** Features are assembled into a feature vector by a feature map acting on the data point (its raw features): sampled signal values of an audio-like waveform form the feature vector x = Phi(z)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-featmap] feature map: data point -> feature vector\")\nt = np.arange(64)\nsignal = np.sin(2 * np.pi * 5 * t / 64) + 0.3 * np.sin(2 * np.pi * 11 * t / 64)\nphi = lambda z: z.copy()                          # signal values as features\nx = phi(signal)\ncheck(\"feature vector collects the d = 64 signal values\", x.size == 64)\ncheck(\"the map is deterministic: same data point, same features\",\n      np.array_equal(phi(signal), x))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-transform]** New features from arithmetic transformations of existing ones: augmenting a scalar feature x with x^2 turns an unlearnable quadratic relation into one a linear hypothesis map fits (the average training loss collapses)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-transform] transformed features make new features\")\nm = 120\nu = rng.uniform(-2, 2, m)\nyq = 1.5 * u**2 + 0.1 * rng.normal(size=m)        # quadratic relation\nfit_err = lambda X: np.mean((yq - np.c_[X, np.ones(m)] @ np.linalg.lstsq(\n    np.c_[X, np.ones(m)], yq, rcond=None)[0]) ** 2)\nerr_raw = fit_err(u[:, None])\nerr_aug = fit_err(np.c_[u, u**2])\nprint(f\"    average training loss raw: {err_raw:.3f}, with x^2 feature: \"\n      f\"{err_aug:.4f}\")\ncheck(\"adding the x^2 feature collapses the average training loss\",\n      err_aug < 0.05 * err_raw)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-dtft]** DFT-magnitude features are unchanged by time shifts of the signal, while raw signal-value features change \u2014 the shift-invariance claim, checked numerically."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-dtft] DFT magnitudes are shift-invariant features\")\nshifted = np.roll(signal, 17)                     # time shift\nmag = lambda z: np.abs(np.fft.rfft(z))\ncheck(\"raw signal-value features change under the shift\",\n      not np.allclose(signal, shifted))\ncheck(\"DFT-magnitude features are unchanged by the shift\",\n      np.allclose(mag(signal), mag(shifted), atol=1e-10))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-activation]** The activation of a neuron is a derived feature: a fixed random one-layer network turns raw features into activations, and a linear hypothesis map on these activations fits a nonlinear relation better than on the raw features."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-activation] neuron activations as derived features\")\nW1 = rng.normal(size=(40, 1))\nb1 = rng.normal(size=40)\nact = lambda X: np.maximum(W1 @ X.T + b1[:, None], 0).T   # ReLU layer\nerr_act = np.mean((yq - np.c_[act(u[:, None]), np.ones(m)] @\n                   np.linalg.lstsq(np.c_[act(u[:, None]), np.ones(m)],\n                                   yq, rcond=None)[0]) ** 2)\nprint(f\"    average training loss on activations: {err_act:.4f}\")\ncheck(\"linear hypothesis map on activations beats raw features\",\n      err_act < 0.1 * err_raw)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-mechfeat]** The narrower mechanistic-interpretability sense: a feature as the projection of the activation vector onto a direction \u2014 the projection of activations onto a planted direction recovers a planted concept (correlation with the concept indicator is high)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-mechfeat] feature = projection of activations onto a direction\")\nd_act = 30\nconcept = (rng.uniform(size=500) > 0.5).astype(float)   # planted concept\ndirection = rng.normal(size=d_act); direction /= np.linalg.norm(direction)\nacts = rng.normal(size=(500, d_act)) + 4.0 * concept[:, None] * direction\nproj = acts @ direction                            # scalar feature\ncorr = np.corrcoef(proj, concept)[0, 1]\nprint(f\"    corr(projection, concept) = {corr:.2f}\")\ncheck(\"projection onto the direction recovers the planted concept\",\n      corr > 0.8)\nother = rng.normal(size=d_act); other -= (other @ direction) * direction\nother /= np.linalg.norm(other)\ncheck(\"projection onto a perpendicular direction does not\",\n      abs(np.corrcoef(acts @ other, concept)[0, 1]) < 0.2)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 2, figsize=(8.4, 3.0))\nax[0].plot(t, signal, \"-\", lw=1, label=\"signal\")\nax[0].plot(t, shifted, \":\", lw=1, label=\"shifted\")\nax[0].set_xlabel(\"sample index\"); ax[0].set_ylabel(\"signal value\")\nax[0].legend(frameon=False); ax[0].set_title(\"[P-dtft] signals\")\nax[1].plot(mag(signal), \"o-\", ms=3, label=\"|DFT| original\")\nax[1].plot(mag(shifted), \"x\", ms=4, label=\"|DFT| shifted\")\nax[1].set_xlabel(\"frequency index\"); ax[1].set_ylabel(\"DFT magnitude\")\nax[1].legend(frameon=False); ax[1].set_title(\"equal magnitudes\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"feature.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}