{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "regularization.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# regularization \u2014 Python demo\n\nNumerical companion to the entry [regularization](https://dictionaryofml.org/terms/regularization.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/regularization.py`](https://dictionaryofml.org/terms/regularization.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"regularization.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nregularization.py \u2014 numerical companion to the glossary entry\n'regularization'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-overfit] Plain ERM with a large model overfits: a degree-12\n            polynomial on m = 15 points has near-zero training error\n            and a large error on unseen data points.\n[P-routes]  The three routes to regularization, each improving the\n            error on unseen data points of the overfitting baseline:\n            1) model pruning \u2014 shrink the hypothesis space (here: fit\n               a smaller-degree polynomial, i.e., constrain higher\n               coefficients to zero);\n            2) loss penalization \u2014 add a penalty term (ridge);\n            3) data augmentation \u2014 enlarge the trainset with perturbed\n               copies of its data points.\n[P-equiv]   The routes can coincide: data augmentation with zero-mean\n            iid feature perturbations of variance sigma^2 yields\n            (asymptotically in the number of perturbed copies) the same\n            learned hypothesis as ridge regression with penalty\n            sigma^2 ||w||^2 \u2014 verified by comparing the two learned\n            parameter vectors.\n\nOutputs\n-------\nregularization.png : preview figure (checking only).\n\nData generated by pythondemos/regularization.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-overfit]** Plain ERM with a large model overfits: a degree-12 polynomial on m = 15 points has near-zero training error and a large error on unseen data points."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-overfit] plain ERM with a large model overfits\")\nm = 15\nxtr = np.sort(rng.uniform(-1, 1, m))\nytr = np.sin(2.5 * xtr) + 0.2 * rng.normal(size=m)\nxva = rng.uniform(-1, 1, 2000)\nyva = np.sin(2.5 * xva) + 0.2 * rng.normal(size=2000)\ndeg = 12\nV = np.vander(xtr, deg + 1)\nVva = np.vander(xva, deg + 1)\nw_erm = np.linalg.lstsq(V, ytr, rcond=None)[0]\ntr_erm = np.mean((ytr - V @ w_erm) ** 2)\nva_erm = np.mean((yva - Vva @ w_erm) ** 2)\nprint(f\"    ERM: error on the trainset {tr_erm:.4f}, \"\n      f\"on unseen data points {va_erm:.2f}\")\ncheck(\"training error near zero\", tr_erm < 0.01)\ncheck(\"error on unseen data points far larger (overfitting)\", va_erm > 10 * 0.04)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-routes]** The three routes to regularization, each improving the error on unseen data points of the overfitting baseline: 1) model pruning \u2014 shrink the hypothesis space (here: fit a smaller-degree polynomial, i.e., constrain higher coefficients to zero); 2) loss penalization \u2014 add a penalty term (ridge); 3) data augmentation \u2014 enlarge the trainset with perturbed copies of its data points."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-routes] three routes, all improving the error on unseen \"\n      \"data points\")\n# 1) model pruning: shrink the hypothesis space H to degree-3 polynomials\nc3 = np.polyfit(xtr, ytr, 3)\nva_prune = np.mean((yva - np.polyval(c3, xva)) ** 2)\n# 2) loss penalization: ridge on the degree-12 model\nlam = 1e-3\nw_ridge = np.linalg.solve(V.T @ V / m + lam * np.eye(deg + 1),\n                          V.T @ ytr / m)\nva_ridge = np.mean((yva - Vva @ w_ridge) ** 2)\n# 3) data augmentation: perturbed copies of the data points\nreps = 100\nsigma = 0.1\nx_aug = np.concatenate([xtr + sigma * rng.normal(size=m)\n                        for _ in range(reps)])\ny_aug = np.tile(ytr, reps)\nw_aug = np.linalg.lstsq(np.vander(x_aug, deg + 1), y_aug, rcond=None)[0]\nva_aug = np.mean((yva - Vva @ w_aug) ** 2)\nprint(f\"    error on unseen data points: ERM {va_erm:.2f} | \"\n      f\"prune {va_prune:.3f} | ridge {va_ridge:.3f} | augment {va_aug:.3f}\")\ncheck(\"1) model pruning improves the error on unseen data points\", va_prune < va_erm / 3)\ncheck(\"2) loss penalization improves the error on unseen data points\",\n      va_ridge < va_erm / 3)\ncheck(\"3) data augmentation improves the error on unseen data points\",\n      va_aug < va_erm / 3)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-equiv]** The routes can coincide: data augmentation with zero-mean iid feature perturbations of variance sigma^2 yields (asymptotically in the number of perturbed copies) the same learned hypothesis as ridge regression with penalty sigma^2 ||w||^2 \u2014 verified by comparing the two learned parameter vectors."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-equiv] augmentation with sigma^2-noise = ridge with sigma^2\")\n# linear regression, feature perturbations with covariance sigma^2 I\nmm, dd = 60, 3\nX = rng.normal(size=(mm, dd))\ny = X @ np.array([1.0, -0.5, 0.25]) + 0.1 * rng.normal(size=mm)\nsig = 0.4\nreps = 4000\nXa = np.vstack([X + sig * rng.normal(size=X.shape) for _ in range(reps)])\nya = np.tile(y, reps)\nw_aug2 = np.linalg.lstsq(Xa, ya, rcond=None)[0]\nw_ridge2 = np.linalg.solve(X.T @ X + sig**2 * mm * np.eye(dd),\n                           X.T @ y)      # (X^T X + m sig^2 I)^{-1} X^T y\ndiff = np.linalg.norm(w_aug2 - w_ridge2)\nprint(f\"    ||w_augment - w_ridge|| = {diff:.4f}\")\ncheck(\"the augmented-ERM solution matches the ridge solution\",\n      diff < 0.02)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(5.4, 3.2))\nxx = np.linspace(-1, 1, 300)\nax.plot(xtr, ytr, \"ko\", ms=4)\nax.plot(xx, np.polyval(w_erm, xx), \":\", label=\"plain ERM (deg 12)\")\nax.plot(xx, np.polyval(w_ridge, xx), \"-\", label=\"ridge\")\nax.plot(xx, np.polyval(c3, xx), \"--\", label=\"pruned (deg 3)\")\nax.set_ylim(-2, 2); ax.legend(frameon=False)\nax.set_xlabel(\"feature $x$\"); ax.set_ylabel(\"label $y$\")\nax.set_title(\"[P-routes] three routes to regularization\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"regularization.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}