{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "valerr.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# validation error \u2014 Python demo\n\nNumerical companion to the entry [validation error](https://dictionaryofml.org/terms/valerr.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nThe entry's weather narrative, carried out: a straight line fitted to days of synthetic weather recordings (morning minimum and maximum daytime temperature), and its validation error computed, compared with the risk, and resampled over random splits. The entry's figure is drawn from the numbers computed here. Self-contained (numpy/matplotlib only), deterministic.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/valerr.py`](https://dictionaryofml.org/terms/valerr.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"valerr.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nvalerr.py -- numerical companion to the entry 'validation error'.\n\nThe entry's weather narrative, carried out: a straight line fitted to days\nof synthetic weather recordings (morning minimum and maximum daytime\ntemperature), and its validation error computed, compared with the risk, and\nresampled over random splits. The entry's figure is drawn from the numbers\ncomputed here. Self-contained (numpy/matplotlib only), deterministic.\n\nBlocks\n------\n[B-def]    Fit a line to a training set of 40 days and compute its\n           validation error: the average squared error over 20 held-back\n           days, with the line fixed.\n[B-risk]   The validation error estimates the risk (approximated by the\n           average loss on 200000 fresh days); the training error\n           understates it.\n[B-spread] The validation error is an average of the per-day losses, so it\n           is itself a random quantity: recomputed over 300 random splits\n           per validation-set size, its spread shrinks as the validation set\n           grows. Writes the CSV behind the entry's figure.\n\nOutputs\n-------\npythondemos/valerr_spread.csv : validation-set size, mean and spread of the\n                                validation error over 300 random splits.\npythondemos/valerr_risk.csv   : the risk of the line fitted to a fixed\n                                training set, for the figure's dashed level.\npythondemos/valerr.png        : preview figure (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\ndef days(n, seed, noise=1.5):\n    \"\"\"n days: morning minimum x and maximum daytime temperature y.\"\"\"\n    gen = np.random.default_rng(seed)\n    x = gen.uniform(-15.0, 5.0, n)\n    y = 4.0 + 0.8 * x + gen.normal(0.0, noise, n)\n    return x, y\n\n\ndef avg_sqerr(x, y, coef):\n    return float(np.mean((y - np.polyval(coef, x)) ** 2))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-def]** Fit a line to a training set of 40 days and compute its validation error: the average squared error over 20 held-back days, with the line fixed."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-def] the validation error of a fitted line\")\n\nx, y = days(60, seed=20260828)\ntr, va = np.arange(40), np.arange(40, 60)\nline = np.polyfit(x[tr], y[tr], 1)\nlosses = (y[va] - np.polyval(line, x[va])) ** 2\nvalerr = avg_sqerr(x[va], y[va], line)\ntrainerr = avg_sqerr(x[tr], y[tr], line)\nprint(f\"    validation error {valerr:.3f} = average of {len(va)} per-day \"\n      f\"squared errors; training error {trainerr:.3f}\")\ncheck(\"the validation error is the average of the per-day losses\",\n      np.isclose(valerr, float(np.mean(losses)), atol=1e-12))\ncheck(\"the line is fixed: computing it changed no coefficient\",\n      np.allclose(line, np.polyfit(x[tr], y[tr], 1)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-risk]** The validation error estimates the risk (approximated by the average loss on 200000 fresh days); the training error understates it."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-risk] the validation error estimates the risk; the training \"\n      \"error understates it\")\n\nxf, yf = days(200_000, seed=99)\nrisk = avg_sqerr(xf, yf, line)\nprint(f\"    validation error {valerr:.3f}, risk (200000 fresh days) \"\n      f\"{risk:.3f}, training error {trainerr:.3f}\")\ncheck(\"the validation error is within 25% of the risk\",\n      abs(valerr - risk) < 0.25 * risk)\ncheck(\"the training error is below the risk\", trainerr < risk)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-spread]** The validation error is an average of the per-day losses, so it is itself a random quantity: recomputed over 300 random splits per validation-set size, its spread shrinks as the validation set grows. Writes the CSV behind the entry's figure."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"\\n[B-spread] an average over more held-back days fluctuates less\")\n\nSIZES = [5, 10, 20, 40, 80]\nR = 300\nrows = []\nxl, yl = days(160, seed=3)          # a larger pool: 80 train + up to 80 val\ntrl = np.arange(80)\nlinel = np.polyfit(xl[trl], yl[trl], 1)\nriskl = avg_sqerr(xf, yf, linel)\nfor nv in SIZES:\n    errs = []\n    for r in range(R):\n        xr, yr = days(nv, seed=50_000 + 100 * nv + r)\n        errs.append(avg_sqerr(xr, yr, linel))\n    rows.append((nv, float(np.mean(errs)), float(np.std(errs))))\n    print(f\"    {nv:>2} held-back days: mean {rows[-1][1]:.3f}, spread \"\n          f\"{rows[-1][2]:.3f}\")\ncheck(\"the spread shrinks as the validation set grows\",\n      all(a[2] > b[2] for a, b in zip(rows, rows[1:])))\ncheck(\"each mean stays within 10% of the risk\",\n      all(abs(m - riskl) < 0.10 * riskl for _, m, _ in rows))\n\nwith open(OUT_DIR / \"valerr_spread.csv\", \"w\") as fh:\n    fh.write(\"size,mean,spread\\n\")\n    for nv, m, s in rows:\n        fh.write(f\"{nv},{m:.4f},{s:.4f}\\n\")\nwith open(OUT_DIR / \"valerr_risk.csv\", \"w\") as fh:\n    fh.write(\"size,risk\\n\")\n    fh.write(f\"{SIZES[0]},{riskl:.4f}\\n\")\n    fh.write(f\"{SIZES[-1]},{riskl:.4f}\\n\")\nprint(\"    wrote valerr_{spread,risk}.csv\")\n\n\n# --------------------------------------------------------------- preview\nfig, ax = plt.subplots(figsize=(6.4, 3.8))\nax.errorbar([r[0] for r in rows], [r[1] for r in rows],\n            yerr=[r[2] for r in rows], fmt=\"o\", ms=5, color=\"black\",\n            capsize=3, label=\"validation error (mean and spread)\")\nax.axhline(riskl, ls=\"--\", color=\"0.4\", label=\"risk of the fixed line\")\nax.set_xscale(\"log\")\nax.set_xticks(SIZES)\nax.set_xticklabels([str(s) for s in SIZES])\nax.set_xlabel(\"number of held-back days\")\nax.set_ylabel(\"validation error\")\nax.set_title(\"[B-spread] the validation error fluctuates less as the \"\n             \"validation set grows\", fontsize=9)\nax.legend(frameon=False, fontsize=8)\n\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"valerr.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(\"wrote valerr.png\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}