{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "model.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# model \u2014 Python demo\n\nNumerical companion to the entry [model](https://dictionaryofml.org/terms/model.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/model.py`](https://dictionaryofml.org/terms/model.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"model.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nmodel.py \u2014 numerical companion to the glossary entry 'model'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-hypospace] Sense 1 \u2014 the model as the hypothesis space: temperature\n              prediction commits to a set of candidate maps (three\n              linear maps, as in the entry's figure panel (a)); every\n              element is a map producing predictions.\n[P-trained]   Sense 2 \u2014 the trained model as the learned hypothesis:\n              the ML map A acts on a trainset and selects the single\n              element h-hat = A(D) of the hypothesis space\n              that best fits the trainset (panel (b)); software-library usage.\n[P-probmodel] Sense 3 \u2014 the probabilistic model as a family of\n              probability distributions: each member is a data\n              generator, and datasets generated from the two members of\n              the family are statistically distinguishable (panel (c)).\n\nOutputs\n-------\nmodel.png : preview figure (checking only).\n\nData generated by pythondemos/model.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-hypospace]** Sense 1 \u2014 the model as the hypothesis space: temperature prediction commits to a set of candidate maps (three linear maps, as in the entry's figure panel (a)); every element is a map producing predictions."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-hypospace] model as the hypothesis space (a set of maps)\")\nH = [lambda x: 0.8 * x + 5.0,\n     lambda x: 1.0 * x + 3.0,\n     lambda x: 1.2 * x + 1.0]                     # three candidate maps\nx_morning = 6.0    # note: at x = 10 all three maps happen to agree\npreds = [h(x_morning) for h in H]\ncheck(\"the model is a set of candidate maps (three linear maps)\",\n      len(H) == 3 and len(set(preds)) == 3)\ncheck(\"every element is a map: it delivers a prediction\",\n      all(np.isfinite(p) for p in preds))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-trained]** Sense 2 \u2014 the trained model as the learned hypothesis: the ML map A acts on a trainset and selects the single element h-hat = A(D) of the hypothesis space that best fits the trainset (panel (b)); software-library usage."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-trained] trained model = learned hypothesis A(D)\")\nm = 50\nx_tr = rng.uniform(5, 20, m)\ny_tr = 1.0 * x_tr + 3.0 + 0.3 * rng.normal(size=m)   # truth = H[1]\nemp_risk = [np.mean((y_tr - h(x_tr)) ** 2) for h in H]\nh_hat = H[int(np.argmin(emp_risk))]\ncheck(\"the ML map selects a single element of the hypothesis space\",\n      h_hat in H)\ncheck(\"A(D) picks the best-fitting map (the true map here)\",\n      int(np.argmin(emp_risk)) == 1)\ncheck(\"the trained model is itself a map (usable for prediction)\",\n      np.isfinite(h_hat(12.0)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-probmodel]** Sense 3 \u2014 the probabilistic model as a family of probability distributions: each member is a data generator, and datasets generated from the two members of the family are statistically distinguishable (panel (c))."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-probmodel] probabilistic model = family of distributions\")\nfamily = [{\"mean\": 0.0, \"std\": 1.0}, {\"mean\": 3.0, \"std\": 1.0}]\ndata0 = rng.normal(family[0][\"mean\"], family[0][\"std\"], 5000)\ndata1 = rng.normal(family[1][\"mean\"], family[1][\"std\"], 5000)\ncheck(\"each member of the family generates iid data points\",\n      data0.shape == (5000,) and data1.shape == (5000,))\ncheck(\"datasets from different members are distinguishable\",\n      abs(data0.mean() - data1.mean()) > 2.5)\ncheck(\"the empirical means identify the generating member\",\n      abs(data0.mean() - family[0][\"mean\"]) < 0.1\n      and abs(data1.mean() - family[1][\"mean\"]) < 0.1)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 3, figsize=(10, 2.8))\nxx = np.linspace(5, 20, 50)\nfor h, st in zip(H, (\"--\", \"-\", \":\")):\n    ax[0].plot(xx, h(xx), st)\nax[0].set_xlabel(\"feature $x$\"); ax[0].set_ylabel(\"prediction $h(x)$\")\nax[0].set_title(\"(a) hypothesis space\")\nax[1].plot(x_tr, y_tr, \"o\", ms=3, alpha=0.5)\nax[1].plot(xx, h_hat(xx), \"r-\")\nax[1].set_xlabel(\"feature $x$\"); ax[1].set_ylabel(\"label $y$\")\nax[1].set_title(\"(b) trained model $\\\\hat{h} = A(D)$\")\nax[2].hist(data0, bins=40, alpha=0.6, density=True)\nax[2].hist(data1, bins=40, alpha=0.6, density=True, hatch=\"//\")\nax[2].set_xlabel(\"observed value\"); ax[2].set_ylabel(\"probability density\")\nax[2].set_title(\"(c) probabilistic model\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"model.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}