{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "generalization.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# generalization \u2014 Python demo\n\nNumerical companion to the entry [generalization](https://dictionaryofml.org/terms/generalization.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/generalization.py`](https://dictionaryofml.org/terms/generalization.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"generalization.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ngeneralization.py \u2014 numerical companion to the glossary entry\n'generalization'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-goal]     Generalization = accurate predictions on data points not\n             used during training: a well-sized hypothesis space gives\n             a test error close to its empirical risk.\n[P-erm]      Low empirical risk does NOT guarantee generalization: a\n             degree-12 polynomial drives the empirical risk near zero\n             while its test error explodes; online (sequential) least\n             squares faces the same gap.\n[P-iid]      Under the iid assumption, the risk is the expected loss\n             and the generalization gap is risk minus empirical risk:\n             both estimated by Monte Carlo for the learned hypothesis.\n[P-event]    For a FIXED hypothesis h, the risk is deterministic while\n             the empirical risk is an RV over trainset draws: its\n             spread shrinks with m, so the probability of the event\n             |emprisk - risk| > eps decays as m grows (concentration).\n[P-stable]   The stability route: replacing one data point of the\n             trainset changes the loss of the learned hypothesis only\n             slightly on average, and the expected generalization gap\n             is bounded by (in fact equals) that average change --\n             verified by averaging over many trainset draws, with no\n             reference to the size of the hypothesis space.\n\nOutputs\n-------\ngeneralization.png : preview figure (checking only).\n\nData generated by pythondemos/generalization.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nf_true = lambda x: np.sin(2.0 * x)\ndef draw(m):\n    x = rng.uniform(-1, 1, m)\n    return x, f_true(x) + 0.2 * rng.normal(size=m)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-goal]** Generalization = accurate predictions on data points not used during training: a well-sized hypothesis space gives a test error close to its empirical risk."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-goal] accuracy beyond the trainset\")\nxtr, ytr = draw(40)\nxte, yte = draw(5000)\nc3 = np.polyfit(xtr, ytr, 3)\ntr3 = np.mean((ytr - np.polyval(c3, xtr)) ** 2)\nte3 = np.mean((yte - np.polyval(c3, xte)) ** 2)\nprint(f\"    degree 3: train {tr3:.3f}, test {te3:.3f}\")\ncheck(\"a well-sized model generalizes (test error close to the empirical risk)\",\n      te3 < 3 * tr3 and te3 < 0.08)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-erm]** Low empirical risk does NOT guarantee generalization: a degree-12 polynomial drives the empirical risk near zero while its test error explodes; online (sequential) least squares faces the same gap."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-erm] low empirical risk does not guarantee generalization\")\nc12 = np.polyfit(xtr, ytr, 12)\ntr12 = np.mean((ytr - np.polyval(c12, xtr)) ** 2)\nte12 = np.mean((yte - np.polyval(c12, xte)) ** 2)\nprint(f\"    degree 12: train {tr12:.4f}, test {te12:.2f}\")\ncheck(\"larger model: lower empirical risk\", tr12 < tr3)\ncheck(\"but much higher test error (no generalization guarantee)\",\n      te12 > 3 * te3)\n# online learning faces the same challenge\nw = np.zeros(13)\nV = np.vander(xtr, 13)\nfor t in range(40):\n    w += 0.05 * (ytr[t] - V[t] @ w) * V[t]\ntr_ol = np.mean((ytr - V @ w) ** 2)\nte_ol = np.mean((yte - np.vander(xte, 13) @ w) ** 2)\ncheck(\"online learning shows a generalization gap too\", te_ol > tr_ol)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-iid]** Under the iid assumption, the risk is the expected loss and the generalization gap is risk minus empirical risk: both estimated by Monte Carlo for the learned hypothesis."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-iid] risk, empirical risk, and the generalization gap\")\nrisk_hat = np.mean((yte - np.polyval(c3, xte)) ** 2)   # MC risk estimate\ngap = risk_hat - tr3\nprint(f\"    risk {risk_hat:.3f}, emprisk {tr3:.3f}, gap {gap:.3f}\")\ncheck(\"the generalization gap = risk - empirical risk is finite and \"\n      \"computable under the iid model\", np.isfinite(gap))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-event]** For a FIXED hypothesis h, the risk is deterministic while the empirical risk is an RV over trainset draws: its spread shrinks with m, so the probability of the event |emprisk - risk| > eps decays as m grows (concentration)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-event] concentration of the empirical risk for fixed h\")\nh_fix = c3                                          # a FIXED hypothesis\nrisk_fix = np.mean((f_true(np.linspace(-1, 1, 10**5))\n                    + 0.2 * rng.normal(size=10**5)\n                    - np.polyval(h_fix, np.linspace(-1, 1, 10**5))) ** 2)\neps = 0.02\ndef event_prob(m, reps=600):\n    hits = 0\n    for _ in range(reps):\n        x, y = draw(m)\n        emp = np.mean((y - np.polyval(h_fix, x)) ** 2)\n        hits += abs(emp - risk_fix) > eps\n    return hits / reps\nprobs = [event_prob(m) for m in (5, 20, 80)]\nprint(f\"    P(|emprisk - risk| > eps) at m = 5, 20, 80: \"\n      f\"{probs[0]:.2f}, {probs[1]:.2f}, {probs[2]:.2f}\")\ncheck(\"the probability of a large deviation decays with m\",\n      probs[0] > probs[1] > probs[2])\ncheck(\"at m = 80 the event is rare\", probs[2] < 0.05)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-stable]** The stability route: replacing one data point of the trainset changes the loss of the learned hypothesis only slightly on average, and the expected generalization gap is bounded by (in fact equals) that average change -- verified by averaging over many trainset draws, with no reference to the size of the hypothesis space."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-stable] stability bounds the expected generalization gap\")\nDEG_S, M_S, LAM, SIG = 3, 20, 0.1, 0.2\nxg = np.linspace(-1, 1, 2001)                     # dense grid for the risk\n\n\ndef ridge_fit(x, y):\n    V = np.vander(x, DEG_S + 1)\n    return np.linalg.solve(V.T @ V + LAM * np.eye(DEG_S + 1), V.T @ y)\n\n\ndef risk_of(c):                                   # E[(y - h(x))^2], exact MC-free\n    return float(np.mean((f_true(xg) - np.polyval(c, xg)) ** 2)) + SIG ** 2\n\n\ngaps, changes = [], []\nfor _ in range(400):\n    x, y = draw(M_S)\n    c = ridge_fit(x, y)\n    emp = float(np.mean((y - np.polyval(c, x)) ** 2))\n    gaps.append(risk_of(c) - emp)\n    i = int(rng.integers(M_S))                    # replace one data point\n    xi, yi = x.copy(), y.copy()\n    xi[i] = rng.uniform(-1, 1)\n    yi[i] = f_true(xi[i]) + SIG * rng.normal()\n    ci = ridge_fit(xi, yi)\n    changes.append(float((y[i] - np.polyval(ci, x[i])) ** 2\n                         - (y[i] - np.polyval(c, x[i])) ** 2))\nlhs, rhs = np.mean(gaps), np.mean(changes)\nprint(f\"    E[gap] {lhs:.4f}, average replace-one change of the loss {rhs:.4f}\")\ncheck(\"[P-stable] the expected gap matches the average replace-one change\",\n      abs(lhs - rhs) < 0.3 * abs(lhs) + 0.01)\ncheck(\"[P-stable] the average absolute change bounds the expected gap\",\n      lhs <= np.mean(np.abs(changes)) + 0.01)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(4.8, 3.2))\nax.semilogx([5, 20, 80], probs, \"o-\")\nax.set_xlabel(\"trainset size m\")\nax.set_ylabel(\"P(|emprisk $-$ risk| > eps)\")\nax.set_title(\"[P-event] concentration for a fixed hypothesis\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"generalization.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}