{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "mlpipeline.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# machine learning pipeline (ML pipeline) \u2014 Python demo\n\nNumerical companion to the entry [machine learning pipeline (ML pipeline)](https://dictionaryofml.org/terms/mlpipeline.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nShows why a stage of an ML pipeline that has its own parameters must be fitted on the training set only. A hand-designed feature extraction stage (keeping the features that agree best with the label) fitted on all available data leaks information from the validation set into the learned hypothesis, so that the validation error underestimates the risk; the same stage fitted on the training set only gives a validation error that matches the risk. Self-contained (numpy and matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/mlpipeline.py`](https://dictionaryofml.org/terms/mlpipeline.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"mlpipeline.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nmlpipeline.py \u2014 numerical companion to the glossary entry 'ML pipeline'.\n\nPurpose\n-------\nShows why a stage of an ML pipeline that has its own parameters must be\nfitted on the training set only.  A hand-designed feature extraction\nstage (keeping the features that agree best with the label) fitted on\nall available data leaks information from the validation set into the\nlearned hypothesis, so that the validation error underestimates the\nrisk; the same stage fitted on the training set only gives a validation\nerror that matches the risk.  Self-contained (numpy and matplotlib\nonly), fixed seed.\n\nSetup\n-----\nRaw data: 500 features per data point, all drawn independently of the\nlabel, so that no hypothesis can do better than the risk 1.0 (the\nlabel's variance).  Training set of 100 data points, validation set of\n100, and a test set of 10000 data points on which the risk of a learned\nhypothesis is estimated.  Pipeline: feature extraction (keep the k\nfeatures whose agreement with the label is largest), then a linear\nmodel fitted on the training set, then prediction.  The feature\nextraction stage is fitted either on training and validation set\ntogether (leaked) or on the training set only (proper), for\nk = 1, 2, 5, 10, 20; every number is an average over 20 random draws\nof the data.\n\nBlocks\n------\n[B-leaked]  Feature extraction fitted on all 200 labelled data points:\n            the validation error of the pipeline drops below 0.9 for\n            k = 20, while the risk of the same hypothesis exceeds 1.4.\n[B-proper]  Feature extraction fitted on the training set only: the\n            validation error stays within 0.1 of the risk for every k.\n[B-compare] The amount by which the leaked validation error\n            underestimates the risk grows with k and exceeds 0.3 for\n            k >= 10; the proper pipeline's never exceeds 0.1.\n\nOutputs\n-------\nmlpipeline_errors.csv : per k, the validation error and the risk of the\n                        leaked pipeline and of the proper pipeline.\nmlpipeline.png        : matplotlib preview of that figure (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nrng = np.random.default_rng(0)\n\nD, M_TRAIN, M_VAL, M_TEST, REPS = 500, 100, 100, 10000, 20\nKS = [1, 2, 5, 10, 20]\ntr = slice(0, M_TRAIN); va = slice(M_TRAIN, M_TRAIN + M_VAL)\nte = slice(M_TRAIN + M_VAL, None)\n\n\ndef fit_extraction(Xf, yf, k):\n    \"\"\"Stage 1: keep the k features agreeing best with the label.\"\"\"\n    score = np.abs((Xf - Xf.mean(0)).T @ (yf - yf.mean())) / len(yf)\n    return np.argsort(score)[::-1][:k]\n\n\ndef fit_model(Xf, yf):\n    \"\"\"Stage 2: linear model with intercept, least squares.\"\"\"\n    A = np.c_[Xf, np.ones(len(yf))]\n    return np.linalg.lstsq(A, yf, rcond=None)[0]\n\n\ndef predict(w, Xf):\n    return np.c_[Xf, np.ones(len(Xf))] @ w\n\n\ndef run(X, y, k, extraction_rows):\n    cols = fit_extraction(X[extraction_rows], y[extraction_rows], k)\n    w = fit_model(X[tr][:, cols], y[tr])              # model: training set only\n    valerr = float(np.mean((y[va] - predict(w, X[va][:, cols])) ** 2))\n    risk = float(np.mean((y[te] - predict(w, X[te][:, cols])) ** 2))\n    return valerr, risk\n\n\nall_rows = np.arange(M_TRAIN + M_VAL)\ntrain_rows = np.arange(M_TRAIN)\nacc = np.zeros((len(KS), 4))                            # vl, rl, vp, rp\nfor _ in range(REPS):\n    X = rng.standard_normal((M_TRAIN + M_VAL + M_TEST, D))\n    y = rng.standard_normal(M_TRAIN + M_VAL + M_TEST)   # independent of X\n    for i, k in enumerate(KS):\n        acc[i, 0:2] += run(X, y, k, all_rows)\n        acc[i, 2:4] += run(X, y, k, train_rows)\nacc /= REPS\nleaked = [(acc[i, 0], acc[i, 1]) for i in range(len(KS))]\nproper = [(acc[i, 2], acc[i, 3]) for i in range(len(KS))]"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-leaked]** Feature extraction fitted on all 200 labelled data points: the validation error of the pipeline drops below 0.9 for k = 20, while the risk of the same hypothesis exceeds 1.4."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "v20, r20 = leaked[KS.index(20)]\ncheck(f\"[B-leaked]  extraction fitted on all 200 points, k=20: validation \"\n      f\"error {v20:.2f}, risk {r20:.2f}\", v20 < 0.9 and r20 > 1.4)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-proper]** Feature extraction fitted on the training set only: the validation error stays within 0.1 of the risk for every k."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "gap_proper = max(abs(v - r) for v, r in proper)\ncheck(f\"[B-proper]  extraction fitted on the training set only: validation \"\n      f\"error within {gap_proper:.2f} of the risk for every k\",\n      gap_proper < 0.1)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-compare]** The amount by which the leaked validation error underestimates the risk grows with k and exceeds 0.3 for k >= 10; the proper pipeline's never exceeds 0.1."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "under_leaked = [r - v for v, r in leaked]\nunder_proper = [r - v for v, r in proper]\ncheck(\"[B-compare] risk minus validation error per k: leaked \"\n      + \", \".join(f\"{u:.2f}\" for u in under_leaked) + \"; proper \"\n      + \", \".join(f\"{u:.2f}\" for u in under_proper),\n      all(a < b for a, b in zip(under_leaked, under_leaked[1:]))\n      and all(u > 0.3 for u, k in zip(under_leaked, KS) if k >= 10)\n      and all(abs(u) < 0.1 for u in under_proper))\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"mlpipeline_errors.csv\", \"w\") as fh:\n    fh.write(\"k,valerr_leaked,risk_leaked,valerr_proper,risk_proper\\n\")\n    for k, (vl, rl), (vp, rp) in zip(KS, leaked, proper):\n        fh.write(f\"{k},{vl:.4f},{rl:.4f},{vp:.4f},{rp:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, ax = plt.subplots(figsize=(5.6, 3.8))\nax.plot(KS, [v for v, _ in leaked], \"ks--\", ms=5, label=\"validation error, extraction fitted on all data\")\nax.plot(KS, [r for _, r in leaked], \"ks:\", ms=5, mfc=\"none\", label=\"risk of that hypothesis\")\nax.plot(KS, [v for v, _ in proper], \"ko-\", ms=5, label=\"validation error, extraction fitted on the training set\")\nax.plot(KS, [r for _, r in proper], \"ko:\", ms=5, mfc=\"none\", label=\"risk of that hypothesis\")\nax.set_xscale(\"log\"); ax.set_xticks(KS); ax.set_xticklabels([str(k) for k in KS])\nax.set_xlabel(\"number of features kept by the extraction stage, k\")\nax.set_ylabel(\"average squared error\")\nax.set_title(\"a stage fitted on all data makes the validation error optimistic\")\nax.legend(frameon=False, fontsize=7)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"mlpipeline.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'mlpipeline_errors.csv'}, {OUT_DIR / 'mlpipeline.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}