{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "modelsel.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# model selection \u2014 Python demo\n\nNumerical companion to the entry [model selection](https://dictionaryofml.org/terms/modelsel.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nBacks the entry's claims with numbers: the training error falls as the hypothesis space grows while the risk eventually rises again, so the training error cannot select; the validation error can; and a test set that is used to select underestimates the risk of the candidate finally chosen. The weather days are those of the entry's figure (generated exactly as in pythondemos/validation.py), so the entry's numbers 0.47, 2065, 7.06 and 1.83 are reproduced. Self-contained (numpy and matplotlib only), fixed seeds.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/modelsel.py`](https://dictionaryofml.org/terms/modelsel.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"modelsel.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nmodelsel.py \u2014 numerical companion to the glossary entry 'model\nselection'.\n\nPurpose\n-------\nBacks the entry's claims with numbers: the training error falls as the\nhypothesis space grows while the risk eventually rises again, so the\ntraining error cannot select; the validation error can; and a test set\nthat is used to select underestimates the risk of the candidate finally\nchosen.  The weather days are those of the entry's figure (generated\nexactly as in pythondemos/validation.py), so the entry's numbers 0.47,\n2065, 7.06 and 1.83 are reproduced.  Self-contained (numpy and\nmatplotlib only), fixed seeds.\n\nSetup\n-----\nSix training days and twenty held-back days with feature x = morning\nminimum temperature and label y = maximum daytime temperature; the\ntwenty are split fourteen to six into validation set and test set.\nCandidates c = 0, ..., 5 are the hypotheses that ERM with the squared\nerror loss delivers from polynomial regression of degree c, so the\nhypothesis spaces grow with c (degree 1 is the linear model).\n\nBlocks\n------\n[B-trainerr]  The training error falls monotonically with the degree and\n              reaches zero at degree five, where the polynomial passes\n              through all six days; the line's training error is 0.47.\n[B-valerr]    The validation error on the fourteen days falls from degree\n              0 to degree 1 and rises afterwards: 7.06 for the line,\n              2065 for the degree-five polynomial.  Selecting by\n              validation error picks the line; its average loss on the\n              six test days, 1.83, is reported as the risk estimate.\n[B-testreuse] Over 2000 draws of the twenty held-back days, the smallest\n              error among the six candidates on a six-day set used for\n              selecting is, on average, far below the error of that same\n              selected candidate on six fresh days: a test set reused for\n              selecting underestimates the risk.\n\nOutputs\n-------\nmodelsel_degrees.csv  : degree, training error, validation error (14\n                        days), test error (6 days) of each candidate.\nmodelsel_testreuse.csv: one row per draw: error of the selected candidate\n                        on the selecting set and on fresh days.\nmodelsel.png          : matplotlib preview of the two figures (checking\n                        only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\ndef days(n, seed, noise=1.5):\n    \"\"\"n days: morning minimum x and maximum daytime temperature y\n    (the generator of pythondemos/validation.py, same seeds).\"\"\"\n    gen = np.random.default_rng(seed)\n    xs = gen.uniform(-15.0, 5.0, n)\n    return xs, 4.0 + 0.8 * xs + gen.normal(0.0, noise, n)\n\n\ndef avg_sqerr(xs, ys, coef):\n    return float(np.mean((ys - np.polyval(coef, xs)) ** 2))\n\n\nxt, yt = days(6, seed=20260828)              # training set\nxv, yv = days(20, seed=31)                    # held back\nxval, yval, xte, yte = xv[:14], yv[:14], xv[14:], yv[14:]\nDEGREES = range(0, 6)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-trainerr]** The training error falls monotonically with the degree and reaches zero at degree five, where the polynomial passes through all six days; the line's training error is 0.47."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "coefs = {c: np.polyfit(xt, yt, c) for c in DEGREES}\ntr = {c: avg_sqerr(xt, yt, coefs[c]) for c in DEGREES}\ncheck(\"[B-trainerr]  training error by degree: \"\n      + \", \".join(f\"{c}: {tr[c]:.2f}\" for c in DEGREES),\n      all(tr[c + 1] <= tr[c] + 1e-12 for c in range(5)) and tr[5] < 1e-8\n      and abs(tr[1] - 0.47) < 0.005)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-valerr]** The validation error on the fourteen days falls from degree 0 to degree 1 and rises afterwards: 7.06 for the line, 2065 for the degree-five polynomial. Selecting by validation error picks the line; its average loss on the six test days, 1.83, is reported as the risk estimate."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "va = {c: avg_sqerr(xval, yval, coefs[c]) for c in DEGREES}\nte = {c: avg_sqerr(xte, yte, coefs[c]) for c in DEGREES}\nc_star = min(va, key=va.get)\ncheck(\"[B-valerr]    validation error by degree: \"\n      + \", \".join(f\"{c}: {va[c]:.2f}\" for c in DEGREES)\n      + f\"; selected degree {c_star}\",\n      c_star == 1 and va[0] > va[1] and all(va[c + 1] > va[c] for c in range(1, 5))\n      and abs(va[1] - 7.06) < 0.005 and abs(round(va[5]) - 2065) < 1)\ncheck(f\"[B-valerr]    average loss of the selected line on the six test \"\n      f\"days: {te[c_star]:.2f}\", abs(te[c_star] - 1.83) < 0.005)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-testreuse]** Over 2000 draws of the twenty held-back days, the smallest error among the six candidates on a six-day set used for selecting is, on average, far below the error of that same selected candidate on six fresh days: a test set reused for selecting underestimates the risk."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "rng = np.random.default_rng(0)\nR = 2000\nsel_err, fresh_err = [], []\nfor _ in range(R):\n    xs = rng.uniform(-15.0, 5.0, 12)\n    ys = 4.0 + 0.8 * xs + rng.normal(0.0, 1.5, 12)\n    x_sel, y_sel, x_new, y_new = xs[:6], ys[:6], xs[6:], ys[6:]\n    errs = {c: avg_sqerr(x_sel, y_sel, coefs[c]) for c in DEGREES}\n    c_sel = min(errs, key=errs.get)\n    sel_err.append(errs[c_sel])\n    fresh_err.append(avg_sqerr(x_new, y_new, coefs[c_sel]))\nsel_err, fresh_err = np.array(sel_err), np.array(fresh_err)\ncheck(f\"[B-testreuse] error of the selected candidate: {sel_err.mean():.2f} \"\n      f\"on the six days used for selecting, {fresh_err.mean():.2f} on six \"\n      f\"fresh days (average over {R} draws)\",\n      sel_err.mean() < 0.8 * fresh_err.mean())\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"modelsel_degrees.csv\", \"w\") as fh:\n    fh.write(\"degree,trainerr,valerr,testerr\\n\")\n    for c in DEGREES:\n        fh.write(f\"{c},{tr[c]:.4f},{va[c]:.4f},{te[c]:.4f}\\n\")\nwith open(OUT_DIR / \"modelsel_testreuse.csv\", \"w\") as fh:\n    fh.write(\"sel_err,fresh_err\\n\")\n    for a, b in zip(sel_err, fresh_err):\n        fh.write(f\"{a:.4f},{b:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.4, 3.9))\nds = list(DEGREES)\nax1.semilogy(ds, [tr[c] + 1e-3 for c in ds], \"ko-\", label=\"training error (6 days)\")\nax1.semilogy(ds, [va[c] for c in ds], \"ks--\", mfc=\"none\",\n             label=\"validation error (14 days)\")\nax1.semilogy(ds, [te[c] for c in ds], \"k^:\", mfc=\"none\", label=\"test error (6 days)\")\nax1.set_xlabel(\"degree of the polynomial (size of the hypothesis space)\")\nax1.set_ylabel(\"average squared error loss\")\nax1.set_title(\"training error falls, validation error rises again\")\nax1.legend(frameon=False, fontsize=8)\nbins = np.linspace(0.0, 8.0, 33)\nax2.hist(sel_err, bins=bins, histtype=\"step\", color=\"black\", ls=\"-\",\n         label=\"on the six days used for selecting\")\nax2.hist(fresh_err, bins=bins, histtype=\"step\", color=\"black\", ls=\"--\",\n         label=\"on six fresh days\")\nax2.set_xlabel(\"average squared error loss of the selected candidate\")\nax2.set_ylabel(\"number of draws\")\nax2.set_title(\"a test set reused for selecting reads too low\")\nax2.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"modelsel.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'modelsel_degrees.csv'}, {OUT_DIR / 'modelsel_testreuse.csv'}, \"\n      f\"{OUT_DIR / 'modelsel.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}