{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "norm.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# norm \u2014 Python demo\n\nNumerical companion to the entry [norm](https://dictionaryofml.org/terms/norm.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/norm.py`](https://dictionaryofml.org/terms/norm.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"norm.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nnorm.py \u2014 numerical companion to the glossary entry 'norm'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-axioms] The three norm axioms \u2014 definiteness, homogeneity, triangle\n           inequality \u2014 hold for the l1-, l2-, and linf-norms on 10^4\n           randomly drawn pairs of vectors; the seminorm near miss\n           u -> |u_1| stays homogeneous and triangle but vanishes on a\n           non-zero vector.\n[P-metric] d(u, v) = ||u - v|| is a metric: symmetry, identity of\n           indiscernibles, and the triangle inequality checked on random\n           triples.\n[P-inner]  The inner product induces the l2-norm: ||u|| = sqrt(<u, u>)\n           for randomly drawn vectors.\n[P-lp]     The lp-norm family: the explicit sum formula matches\n           np.linalg.norm for p in {1, 2, inf}, the linf-norm is the\n           max absolute entry, and the unit spheres of l1/l2/linf are\n           nested (||x||_inf <= ||x||_2 <= ||x||_1) \u2014 the geometry of\n           the entry's unit-ball figure.\n[P-ml]     Norms define losses and regularizers: the squared error loss\n           is a squared l2-norm of the residual, and the ridge/Lasso\n           penalty terms evaluate norms of the model parameters.\n[P-fit]    The choice of norm determines the learned hypothesis: on a\n           training set of four data points (three on the line y = x,\n           one outlier), minimizing the l2-norm of the residual (the\n           squared norm has the same minimizer) gives slope 46/30,\n           pulled toward the outlier, while minimizing the l1-norm\n           gives slope 1 exactly \u2014 the entry's l2-vs-l1 fit figure.\n\nOutputs\n-------\nnorm.png           : preview figure (checking only).\nnorm_points.csv    : the four training data points (x, y).\nnorm_fits.csv      : the l2 and l1 fitted lines (x, l2, l1).\n\nData generated by pythondemos/norm.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nNORMS = {\"l1\": 1, \"l2\": 2, \"linf\": np.inf}\nU = rng.normal(size=(10**4, 5))\nV = rng.normal(size=(10**4, 5))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-axioms]** The three norm axioms \u2014 definiteness, homogeneity, triangle inequality \u2014 hold for the l1-, l2-, and linf-norms on 10^4 randomly drawn pairs of vectors; the seminorm near miss u -> |u_1| stays homogeneous and triangle but vanishes on a non-zero vector."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-axioms] definiteness, homogeneity, triangle inequality\")\nfor name, p in NORMS.items():\n    nu = np.linalg.norm(U, p, axis=1)\n    check(f\"{name}: norm(0) = 0 and norm(u) > 0 for u != 0\",\n          np.linalg.norm(np.zeros(5), p) == 0 and np.all(nu > 0))\n    a = rng.normal()\n    check(f\"{name}: homogeneity ||a u|| = |a| ||u||\",\n          np.allclose(np.linalg.norm(a * U, p, axis=1), abs(a) * nu))\n    check(f\"{name}: triangle ||u + v|| <= ||u|| + ||v||\",\n          np.all(np.linalg.norm(U + V, p, axis=1)\n                 <= nu + np.linalg.norm(V, p, axis=1) + 1e-12))\n# seminorm near miss: f(u) = |u_1| is homogeneous and satisfies the\n# triangle inequality but vanishes on the non-zero vector (0,...,0,1)\nf = lambda X: np.abs(np.atleast_2d(X)[:, 0])\na = rng.normal()\ncheck(\"near miss |u_1|: homogeneous and triangle hold\",\n      np.allclose(f(a * U), abs(a) * f(U))\n      and np.all(f(U + V) <= f(U) + f(V) + 1e-12))\ncheck(\"near miss |u_1|: not definite (zero on a non-zero vector)\",\n      np.isclose(f(np.eye(5)[4])[0], 0.0))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-metric]** d(u, v) = ||u - v|| is a metric: symmetry, identity of indiscernibles, and the triangle inequality checked on random triples."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-metric] d(u, v) = ||u - v|| is a metric\")\nW = rng.normal(size=(10**4, 5))\nd = lambda X, Y: np.linalg.norm(X - Y, 2, axis=1)\ncheck(\"symmetry d(u, v) = d(v, u)\", np.allclose(d(U, V), d(V, U)))\ncheck(\"d(u, u) = 0\", np.all(d(U, U) == 0))\ncheck(\"d(u, v) > 0 whenever u != v (identity of indiscernibles)\",\n      np.all(d(U, V)[np.any(U != V, axis=1)] > 0))\ncheck(\"triangle d(u, w) <= d(u, v) + d(v, w)\",\n      np.all(d(U, W) <= d(U, V) + d(V, W) + 1e-12))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-inner]** The inner product induces the l2-norm: ||u|| = sqrt(<u, u>) for randomly drawn vectors."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-inner] the inner product induces the l2-norm\")\ncheck(\"||u|| = sqrt(<u, u>)\",\n      np.allclose(np.linalg.norm(U, 2, axis=1),\n                  np.sqrt(np.sum(U * U, axis=1))))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-lp]** The lp-norm family: the explicit sum formula matches np.linalg.norm for p in {1, 2, inf}, the linf-norm is the max absolute entry, and the unit spheres of l1/l2/linf are nested (||x||_inf <= ||x||_2 <= ||x||_1) \u2014 the geometry of the entry's unit-ball figure."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-lp] the lp family and its unit-ball geometry\")\nx = rng.normal(size=(10**4, 5))\nlp_sum = lambda X, p: (np.sum(np.abs(X) ** p, axis=1)) ** (1 / p)\ncheck(\"sum formula matches np.linalg.norm for p = 1, 2\",\n      np.allclose(lp_sum(x, 1), np.linalg.norm(x, 1, axis=1))\n      and np.allclose(lp_sum(x, 2), np.linalg.norm(x, 2, axis=1)))\ncheck(\"linf-norm is the max absolute entry\",\n      np.allclose(np.linalg.norm(x, np.inf, axis=1),\n                  np.max(np.abs(x), axis=1)))\ncheck(\"||x||_inf <= ||x||_2 <= ||x||_1 (nested unit balls)\",\n      np.all(np.linalg.norm(x, np.inf, axis=1)\n             <= np.linalg.norm(x, 2, axis=1) + 1e-12)\n      and np.all(np.linalg.norm(x, 2, axis=1)\n                 <= np.linalg.norm(x, 1, axis=1) + 1e-12))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-ml]** Norms define losses and regularizers: the squared error loss is a squared l2-norm of the residual, and the ridge/Lasso penalty terms evaluate norms of the model parameters."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-ml] norms define losses and regularizers\")\nXf = rng.normal(size=(50, 3))\nyf = Xf @ np.array([1.0, 0.0, -0.5]) + 0.1 * rng.normal(size=50)\nw = np.linalg.lstsq(Xf, yf, rcond=None)[0]\nsq_loss = np.mean((yf - Xf @ w) ** 2)\ncheck(\"squared-error loss = (1/m) ||y - X w||_2^2\",\n      np.isclose(sq_loss, np.linalg.norm(yf - Xf @ w) ** 2 / 50))\ncheck(\"ridge regularizer alpha ||w||_2^2 and Lasso regularizer \"\n      \"alpha ||w||_1 are norm evaluations\",\n      np.isclose(np.linalg.norm(w, 2) ** 2, np.sum(w**2))\n      and np.isclose(np.linalg.norm(w, 1), np.sum(np.abs(w))))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-fit]** The choice of norm determines the learned hypothesis: on a training set of four data points (three on the line y = x, one outlier), minimizing the l2-norm of the residual (the squared norm has the same minimizer) gives slope 46/30, pulled toward the outlier, while minimizing the l1-norm gives slope 1 exactly \u2014 the entry's l2-vs-l1 fit figure."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-fit] the choice of norm determines the learned hypothesis\")\nx4 = np.array([1.0, 2.0, 3.0, 4.0])       # feature values\ny4 = np.array([1.0, 2.0, 3.0, 8.0])       # labels; the last one is an outlier\nw_l2 = float(x4 @ y4 / (x4 @ x4))         # minimizes ||y - w x||_2^2\nf1 = lambda w: np.sum(np.abs(y4 - np.multiply.outer(w, x4)), axis=-1)\ncand = y4 / x4                            # the l1 training error attains its\nw_l1 = float(cand[np.argmin(f1(cand))])   # minimum at one of these slopes\ngrid = np.linspace(0.5, 2.5, 20001)\ncheck(\"l1 fit: slope 1, fits the three regular points exactly and beats \"\n      \"every slope on a dense grid\",\n      np.isclose(w_l1, 1.0) and f1(w_l1) <= np.min(f1(grid)) + 1e-9)\ncheck(\"l2 fit: slope 46/30 > 1, pulled toward the outlier\",\n      np.isclose(w_l2, 46 / 30) and w_l2 > w_l1)\nf2 = lambda w: np.sum((y4 - w * x4) ** 2)\ncheck(\"each fit minimizes its own norm of the residual\",\n      f1(w_l1) < f1(w_l2) and f2(w_l2) < f2(w_l1))\nnp.savetxt(OUT_DIR / \"norm_points.csv\",\n           np.column_stack([x4, y4]),\n           header=\"x,y\", comments=\"\", delimiter=\",\", fmt=\"%.1f\")\nxs = np.array([0.0, 2.3, 4.6])\nnp.savetxt(OUT_DIR / \"norm_fits.csv\",\n           np.column_stack([xs, w_l2 * xs, w_l1 * xs]),\n           header=\"x,l2,l1\", comments=\"\", delimiter=\",\", fmt=\"%.4f\")\n\n# ------------------------------------------------------------ preview\nth = np.linspace(0, 2 * np.pi, 400)\ncirc = np.stack([np.cos(th), np.sin(th)])\nfig, (ax, axr) = plt.subplots(1, 2, figsize=(8.0, 4.0))\nax.plot(*(circ / np.linalg.norm(circ, 1, axis=0)), \"k--\", label=\"$\\\\ell_1$\")\nax.plot(*circ, \"k-\", label=\"$\\\\ell_2$\")\nax.plot(*(circ / np.linalg.norm(circ, np.inf, axis=0)), \"k:\",\n        label=\"$\\\\ell_\\\\infty$\")\nax.set_xlabel(\"$x_1$\"); ax.set_ylabel(\"$x_2$\")\nax.set_aspect(\"equal\"); ax.legend(frameon=False)\nax.set_title(\"[P-lp] unit spheres of the $\\\\ell_1$/$\\\\ell_2$/$\\\\ell_\\\\infty$ norms\")\naxr.plot(xs, w_l2 * xs, \"k-\",\n         label=f\"$\\\\ell_2$ fit: slope {w_l2:.2f}\")\naxr.plot(xs, w_l1 * xs, \"k--\",\n         label=f\"$\\\\ell_1$ fit: slope {w_l1:.0f}\")\naxr.plot(x4, y4, \"ko\", label=\"training set\")\naxr.annotate(\"outlier\", (x4[-1], y4[-1]), textcoords=\"offset points\",\n             xytext=(-8, 4), ha=\"right\")\naxr.set_xlabel(\"feature $x$\"); axr.set_ylabel(\"label $y$\")\naxr.legend(frameon=False)\naxr.set_title(\"[P-fit] the norm of the residual\\ndetermines the learned hypothesis\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"norm.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}