{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "data.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# data \u2014 Python demo\n\nNumerical companion to the entry [data](https://dictionaryofml.org/terms/data.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/data.py`](https://dictionaryofml.org/terms/data.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"data.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ndata.py \u2014 numerical companion to the glossary entry 'data'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-record]    Data are representations of information recorded in a form\n              suitable for storage, communication, and processing: a\n              temperature reading with timestamp survives a round trip\n              through serialized bytes unchanged. The usefulness of a\n              learned hypothesis is limited by data quality and\n              quantity: test error decreases with the trainset size and\n              increases with measurement noise.\n[P-datapoint] The unit sense: a data point (x, y) bundles features and\n              a label; the entry's convention x easy to obtain, y the\n              quantity of interest.\n[P-dataset]   The collection sense: a dataset of m data points supports\n              model training and validation via a train/validation\n              split.\n\nOutputs\n-------\ndata.png : preview figure (checking only).\n\nData generated by pythondemos/data.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-record]** Data are representations of information recorded in a form suitable for storage, communication, and processing: a temperature reading with timestamp survives a round trip through serialized bytes unchanged. The usefulness of a learned hypothesis is limited by data quality and quantity: test error decreases with the trainset size and increases with measurement noise."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-record] recorded, storable, communicable representations\")\nreading = np.array([(1717.25, 23.4)],\n                   dtype=[(\"t\", \"f8\"), (\"temp\", \"f8\")])\nblob = reading.tobytes()                          # storage/communication\nrestored = np.frombuffer(blob, dtype=reading.dtype)\ncheck(\"a temperature reading survives the storage round trip\",\n      restored[\"temp\"][0] == 23.4 and restored[\"t\"][0] == 1717.25)\n# quality and quantity limit the learned hypothesis\ndef test_err(m, noise):\n    x = rng.uniform(0, 10, m)\n    y = 2.0 * x + 1.0 + noise * rng.normal(size=m)\n    c = np.polyfit(x, y, 1)\n    xt = rng.uniform(0, 10, 2000)\n    return np.mean((2.0 * xt + 1.0 - np.polyval(c, xt)) ** 2)\ne_small, e_large = np.mean([test_err(10, 1.0) for _ in range(40)]), \\\n                   np.mean([test_err(200, 1.0) for _ in range(40)])\ne_clean, e_noisy = np.mean([test_err(50, 0.2) for _ in range(40)]), \\\n                   np.mean([test_err(50, 2.0) for _ in range(40)])\ncheck(\"more data points -> lower test error (quantity)\",\n      e_large < e_small)\ncheck(\"noisier recordings -> higher test error (quality)\",\n      e_noisy > e_clean)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-datapoint]** The unit sense: a data point (x, y) bundles features and a label; the entry's convention x easy to obtain, y the quantity of interest."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-datapoint] the unit sense: features and label\")\nz = {\"x\": np.array([23.4, 55.0]), \"y\": 1.0}       # (features, label)\ncheck(\"a data point bundles features and a label\",\n      z[\"x\"].shape == (2,) and np.isscalar(z[\"y\"]))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-dataset]** The collection sense: a dataset of m data points supports model training and validation via a train/validation split."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-dataset] the collection sense: training and validation\")\nm = 100\nX = rng.uniform(0, 10, m)\nY = 2.0 * X + 1.0 + 0.5 * rng.normal(size=m)\ntr, va = slice(0, 80), slice(80, 100)\nc = np.polyfit(X[tr], Y[tr], 1)\nval_err = np.mean((Y[va] - np.polyval(c, X[va])) ** 2)\ncheck(\"the dataset supports training (fit on the training part)\",\n      np.all(np.isfinite(c)))\ncheck(\"and validation (error evaluated on held-out data points)\",\n      np.isfinite(val_err) and val_err < 1.0)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(4.8, 3.2))\nms = [10, 30, 100, 300]\nax.loglog(ms, [np.mean([test_err(mm, 1.0) for _ in range(30)])\n               for mm in ms], \"o-\")\nax.set_xlabel(\"trainset size m\"); ax.set_ylabel(\"test error\")\nax.set_title(\"[P-record] usefulness grows with data quantity\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"data.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}