{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "kfoldcv.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# k-fold cross-validation (k-fold CV) \u2014 Python demo\n\nNumerical companion to the entry [k-fold cross-validation (k-fold CV)](https://dictionaryofml.org/terms/kfoldcv.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nCarries out the entry's fifteen-day weather example and backs its two claims about k-fold CV: a single split into training set and validation set yields a verdict that moves with the split, which the average over k folds avoids; and the choice of k trades bias against variance of the estimate. Self-contained (numpy and matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/kfoldcv.py`](https://dictionaryofml.org/terms/kfoldcv.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"kfoldcv.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nkfoldcv.py \u2014 numerical companion to the glossary entry 'k-fold\ncross-validation (k-fold CV)'.\n\nPurpose\n-------\nCarries out the entry's fifteen-day weather example and backs its two\nclaims about k-fold CV: a single split into training set and validation\nset yields a verdict that moves with the split, which the average over\nk folds avoids; and the choice of k trades bias against variance of the\nestimate.  Self-contained (numpy and matplotlib only), fixed seed.\n\nSetup\n-----\nFifteen days with feature x = morning minimum temperature (degrees C)\nand label y = maximum daytime temperature (degrees C), the fifteen\ncoordinates drawn in the entry's figure.  The method under study is\nERM on a linear model h(x) = w1 x + w2 with the squared error loss.\nFold b (b = 1, ..., 5) holds days b, b + 5 and b + 10 in the order of\nincreasing x, so each fold has three days.  For the choice of k, the\nsame kind of fifteen-day dataset is drawn 2000 times from y = 4 + 0.8 x\nplus noise, with x uniform on [-15, 5].\n\nBlocks\n------\n[B-split]  Two single splits (fold 1 and fold 4 held out): the slopes\n           of the learned hypotheses are 0.76 and 0.60, the validation\n           errors 0.77 and 1.05 \u2014 the verdict moves with the split.\n[B-cv]     Five-fold CV on the fifteen days: the five validation errors\n           0.77, 1.03, 0.90, 1.05, 1.17 and their average 0.99, as the\n           entry states them.\n[B-k]      The choice of k on 2000 drawn datasets, for k = 2, 3, 5 and\n           15 (leave-one-out): the estimate's bias with respect to the\n           expected loss of ERM on all fifteen days falls with k, since\n           each training set holds a fraction (k-1)/k of the days; the\n           spread of the estimate across datasets is reported alongside.\n\nOutputs\n-------\nkfoldcv_days.csv : the fifteen days with columns x, y, fold.\nkfoldcv_folds.csv: fold index, slope and intercept of the learned\n                   hypothesis, validation error of that fold.\nkfoldcv_k.csv    : k, bias of the k-fold CV estimate and its standard\n                   deviation across datasets.\nkfoldcv.png      : matplotlib preview of the two figures (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nx = np.array([0.4, 0.9, 1.3, 1.8, 2.2, 2.7, 3.2, 3.6, 4.1, 4.6, 5.0, 5.5,\n              6.0, 6.5, 7.0])\ny = np.array([7.99, 6.35, 8.24, 7.4, 9.58, 8.24, 9.9, 8.19, 10.15, 11.11,\n              9.5, 11.36, 10.52, 12.58, 11.04])\nm = len(x)\nK = 5\nfold = (np.arange(m) % K) + 1               # days b, b+5, b+10 form fold b\n\n\ndef erm_line(xs, ys):\n    \"\"\"ERM on the linear model with the squared error loss: slope, intercept.\"\"\"\n    A = np.c_[xs, np.ones_like(xs)]\n    return np.linalg.lstsq(A, ys, rcond=None)[0]\n\n\ndef avg_sqerr(xs, ys, w):\n    return float(np.mean((ys - (w[0] * xs + w[1])) ** 2))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-split]** Two single splits (fold 1 and fold 4 held out): the slopes of the learned hypotheses are 0.76 and 0.60, the validation errors 0.77 and 1.05 \u2014 the verdict moves with the split."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "per_fold = []\nfor b in range(1, K + 1):\n    tr, va = fold != b, fold == b\n    w = erm_line(x[tr], y[tr])\n    per_fold.append((b, w[0], w[1], avg_sqerr(x[va], y[va], w)))\ns1, s4 = per_fold[0][1], per_fold[3][1]\nv1, v4 = per_fold[0][3], per_fold[3][3]\ncheck(f\"[B-split] fold 1 held out: slope {s1:.2f}, validation error {v1:.2f};\"\n      f\" fold 4 held out: slope {s4:.2f}, validation error {v4:.2f}\",\n      abs(s1 - 0.76) < 0.005 and abs(s4 - 0.60) < 0.005\n      and abs(v1 - 0.77) < 0.005 and abs(v4 - 1.05) < 0.005)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-cv]** Five-fold CV on the fifteen days: the five validation errors 0.77, 1.03, 0.90, 1.05, 1.17 and their average 0.99, as the entry states them."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "valerrs = np.array([p[3] for p in per_fold])\ncv = float(valerrs.mean())\ncheck(f\"[B-cv]    five validation errors {np.round(valerrs, 2).tolist()}, \"\n      f\"average {cv:.2f}\",\n      np.allclose(np.round(valerrs, 2), [0.77, 1.03, 0.90, 1.05, 1.17])\n      and abs(cv - 0.99) < 0.005)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-k]** The choice of k on 2000 drawn datasets, for k = 2, 3, 5 and 15 (leave-one-out): the estimate's bias with respect to the expected loss of ERM on all fifteen days falls with k, since each training set holds a fraction (k-1)/k of the days; the spread of the estimate across datasets is reported alongside."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "rng = np.random.default_rng(0)\nR, N, NOISE = 2000, 15, 1.5\nks = [2, 3, 5, 15]\n\n\ndef draw(n):\n    xs = rng.uniform(-15.0, 5.0, n)\n    return xs, 4.0 + 0.8 * xs + rng.normal(0.0, NOISE, n)\n\n\nx_fresh, y_fresh = draw(20000)              # for the expected loss\nest = {k: [] for k in ks}\ntrue_risk = []\nfor _ in range(R):\n    xs, ys = draw(N)\n    true_risk.append(avg_sqerr(x_fresh, y_fresh, erm_line(xs, ys)))\n    for k in ks:\n        f = np.arange(N) % k\n        errs = []\n        for b in range(k):\n            w = erm_line(xs[f != b], ys[f != b])\n            errs.append(avg_sqerr(xs[f == b], ys[f == b], w))\n        est[k].append(float(np.mean(errs)))\nrisk = float(np.mean(true_risk))\nrows = []\nfor k in ks:\n    e = np.array(est[k])\n    rows.append((k, float(e.mean() - risk), float(e.std())))\nbias = {r[0]: r[1] for r in rows}\ncheck(f\"[B-k]     expected loss of ERM on 15 days {risk:.2f}; bias of the \"\n      f\"k-fold CV estimate \" + \", \".join(f\"k={r[0]}: {r[1]:+.2f}\" for r in rows),\n      bias[2] > bias[5] > bias[15] and bias[2] > 0.1 and abs(bias[15]) < 0.1)\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"kfoldcv_days.csv\", \"w\") as fh:\n    fh.write(\"x,y,fold\\n\")\n    for a, b, c in zip(x, y, fold):\n        fh.write(f\"{a:.2f},{b:.2f},{c}\\n\")\nwith open(OUT_DIR / \"kfoldcv_folds.csv\", \"w\") as fh:\n    fh.write(\"fold,slope,intercept,valerr\\n\")\n    for b, s, i, v in per_fold:\n        fh.write(f\"{b},{s:.4f},{i:.4f},{v:.4f}\\n\")\nwith open(OUT_DIR / \"kfoldcv_k.csv\", \"w\") as fh:\n    fh.write(\"k,bias,std\\n\")\n    for k, bi, sd in rows:\n        fh.write(f\"{k},{bi:.4f},{sd:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.4, 3.9))\nbs = [p[0] for p in per_fold]\nax1.bar(bs, valerrs, width=0.6, color=\"0.75\", edgecolor=\"black\",\n        label=\"validation error of fold $b$\")\nax1.axhline(cv, color=\"black\", ls=\"--\", label=f\"5-fold CV estimate {cv:.2f}\")\nax1.set_xlabel(\"fold $b$ held out as validation set\")\nax1.set_ylabel(\"average squared error loss\")\nax1.set_title(\"five-fold CV on the fifteen days\")\nax1.legend(frameon=False, fontsize=8)\nax2.plot(ks, [r[1] for r in rows], \"ko-\", label=\"bias of the estimate\")\nax2.plot(ks, [r[2] for r in rows], \"ks--\", mfc=\"none\",\n         label=\"standard deviation across datasets\")\nax2.set_xlabel(\"number of folds $k$\")\nax2.set_ylabel(\"value\")\nax2.set_title(\"choice of $k$ over 2000 drawn datasets\")\nax2.set_xticks(ks)\nax2.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"kfoldcv.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'kfoldcv_days.csv'}, {OUT_DIR / 'kfoldcv_folds.csv'}, \"\n      f\"{OUT_DIR / 'kfoldcv_k.csv'}, {OUT_DIR / 'kfoldcv.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}