{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "trainerr.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# training error \u2014 Python demo\n\nNumerical companion to the entry [training error](https://dictionaryofml.org/terms/trainerr.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nThe entry's weather narrative, carried out: a straight line fitted by ERM to 40 days of synthetic weather recordings (morning minimum and maximum daytime temperature), its training error as the average of the 40 squared misses, the minimality of that average over all lines, and the training error against the risk for polynomials of growing degree fitted to the same 40 days. The days are generated exactly as in valerr.py (same seed), so the numbers of the two entries agree. Self-contained (numpy/matplotlib only), deterministic.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/trainerr.py`](https://dictionaryofml.org/terms/trainerr.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"trainerr.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ntrainerr.py -- numerical companion to the entry 'training error'.\n\nThe entry's weather narrative, carried out: a straight line fitted by ERM\nto 40 days of synthetic weather recordings (morning minimum and maximum\ndaytime temperature), its training error as the average of the 40 squared\nmisses, the minimality of that average over all lines, and the training\nerror against the risk for polynomials of growing degree fitted to the same\n40 days.  The days are generated exactly as in valerr.py (same seed), so the\nnumbers of the two entries agree.  Self-contained (numpy/matplotlib only),\ndeterministic.\n\nBlocks\n------\n[B-def]     Fit a line to 40 days and compute its training error: the\n            average squared error over the same 40 days.  Check that it is\n            the average of the per-day losses and that no other line has a\n            smaller one (ERM delivers the minimum).\n[B-degree]  Polynomials of degree 0 to 12 fitted by ERM to the same 40\n            days: the training error never increases with the degree,\n            while the risk (average loss on 200000 fresh days) falls and\n            then rises.  Writes the CSV behind the entry's figure.\n\nOutputs\n-------\npythondemos/trainerr_degree.csv : degree, training error and risk of the\n                                  polynomial fitted by ERM.\npythondemos/trainerr.png        : preview figure (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\ndef days(n, seed, noise=1.5):\n    \"\"\"n days: morning minimum x and maximum daytime temperature y.\"\"\"\n    gen = np.random.default_rng(seed)\n    x = gen.uniform(-15.0, 5.0, n)\n    y = 4.0 + 0.8 * x + gen.normal(0.0, noise, n)\n    return x, y\n\n\ndef avg_sqerr(x, y, coef):\n    return float(np.mean((y - np.polyval(coef, x)) ** 2))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-def]** Fit a line to 40 days and compute its training error: the average squared error over the same 40 days. Check that it is the average of the per-day losses and that no other line has a smaller one (ERM delivers the minimum)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-def] the training error of a line fitted by ERM\")\n\nx, y = days(60, seed=20260828)\ntr = np.arange(40)                         # the same 40 days as in valerr.py\nline = np.polyfit(x[tr], y[tr], 1)\nlosses = (y[tr] - np.polyval(line, x[tr])) ** 2\ntrainerr = avg_sqerr(x[tr], y[tr], line)\nprint(f\"    line: slope {line[0]:.3f}, offset {line[1]:.3f}; training error \"\n      f\"{trainerr:.3f} = average of {len(tr)} per-day squared errors \"\n      f\"(smallest {losses.min():.3f}, largest {losses.max():.3f})\")\ncheck(\"the training error is the average of the per-day losses\",\n      np.isclose(trainerr, float(np.mean(losses)), atol=1e-12))\ngen = np.random.default_rng(1)\nothers = line + gen.normal(0.0, [0.05, 0.5], size=(2000, 2))\nworse = [avg_sqerr(x[tr], y[tr], c) for c in others]\nprint(f\"    2000 other lines (slope and offset perturbed): smallest training \"\n      f\"error {min(worse):.3f}\")\ncheck(\"no other line has a smaller training error: ERM delivers the minimum\",\n      min(worse) >= trainerr)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-degree]** Polynomials of degree 0 to 12 fitted by ERM to the same 40 days: the training error never increases with the degree, while the risk (average loss on 200000 fresh days) falls and then rises. Writes the CSV behind the entry's figure."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-degree] training error and risk against the degree of the polynomial\")\n\nxf, yf = days(200_000, seed=99)\nDEGREES = list(range(0, 13))\nrows = []\nfor d in DEGREES:\n    coef = np.polyfit(x[tr], y[tr], d)\n    rows.append((d, avg_sqerr(x[tr], y[tr], coef), avg_sqerr(xf, yf, coef)))\n    print(f\"    degree {d:2d}: training error {rows[-1][1]:.3f}, risk {rows[-1][2]:.3f}\")\ntrain_curve = [r[1] for r in rows]\nrisk_curve = [r[2] for r in rows]\ncheck(\"the training error never increases with the degree\",\n      all(a >= b - 1e-9 for a, b in zip(train_curve, train_curve[1:])))\ncheck(\"the training error is below the risk for every degree from 1 on\",\n      all(r[1] < r[2] for r in rows[1:]))\nbest = int(np.argmin(risk_curve))\nprint(f\"    smallest risk at degree {best}; at degree 12 the training error is \"\n      f\"{train_curve[-1]:.3f} and the risk {risk_curve[-1]:.3f}\")\ncheck(\"the risk is smallest at degree 1 and larger at degree 12 than at degree 1\",\n      best == 1 and risk_curve[-1] > risk_curve[1])\n\nwith open(OUT_DIR / \"trainerr_degree.csv\", \"w\") as fh:\n    fh.write(\"degree,train,risk\\n\")\n    for d, t, r in rows:\n        fh.write(f\"{d},{t:.4f},{r:.4f}\\n\")\nprint(\"    wrote trainerr_degree.csv\")\n\n\n# --------------------------------------------------------------- preview\nfig, ax = plt.subplots(figsize=(6.4, 3.8))\nax.semilogy(DEGREES, train_curve, \"o-\", color=\"black\", ms=4, label=\"training error\")\nax.semilogy(DEGREES, risk_curve, \"s--\", color=\"0.4\", ms=4, label=\"risk (200000 fresh days)\")\nax.set_xlabel(\"degree of the polynomial\")\nax.set_ylabel(\"average squared error\")\nax.set_title(\"[B-degree] training error and risk of the polynomial fitted by ERM\",\n             fontsize=9)\nax.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"trainerr.png\", dpi=150)\nplt.close(fig)\nprint(\"wrote trainerr.png\")\n\npassed = sum(1 for _, ok in report if ok)\nprint(f\"\\n{passed}/{len(report)} checks pass\")"
  }
 ]
}