{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "dataset.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# dataset \u2014 Python demo\n\nNumerical companion to the entry [dataset](https://dictionaryofml.org/terms/dataset.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/dataset.py`](https://dictionaryofml.org/terms/dataset.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"dataset.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ndataset.py \u2014 numerical companion to the glossary entry 'dataset'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-setsample] Two readings of 'dataset': strictly a set of distinct data\n              points (no order, no repetitions) vs the ML usage as a\n              sample (an indexed sequence that may repeat). A sequence\n              with a repeated data point has m = 4 entries but only 3\n              distinct elements, and reordering the sequence leaves the\n              underlying set unchanged.\n[P-table]     The relational-model reading: a table whose rows are data\n              points and whose columns are attributes (the cow table of\n              the entry). ML methods use the attribute columns as\n              features or the label; the row order is immaterial \u2014 the\n              least-squares fit computed from a row-shuffled table is\n              identical.\n\nOutputs\n-------\ndataset.png : preview figure (checking only).\n\nData generated by pythondemos/dataset.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-setsample]** Two readings of 'dataset': strictly a set of distinct data points (no order, no repetitions) vs the ML usage as a sample (an indexed sequence that may repeat). A sequence with a repeated data point has m = 4 entries but only 3 distinct elements, and reordering the sequence leaves the underlying set unchanged."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-setsample] set of distinct points vs indexed sample\")\nz1, z2, z3 = np.array([1.0, 2.0]), np.array([3.0, 1.0]), np.array([2.0, 2.0])\nsample = np.stack([z1, z2, z1, z3])              # position 1 and 3 repeat\nas_set = np.unique(sample, axis=0)\ncheck(\"the sample has m = 4 indexed entries\", sample.shape[0] == 4)\ncheck(\"the underlying set has only 3 distinct data points\",\n      as_set.shape[0] == 3)\nperm = rng.permutation(4)\ncheck(\"reordering the sample leaves the set unchanged\",\n      np.array_equal(np.unique(sample[perm], axis=0), as_set))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-table]** The relational-model reading: a table whose rows are data points and whose columns are attributes (the cow table of the entry). ML methods use the attribute columns as features or the label; the row order is immaterial \u2014 the least-squares fit computed from a row-shuffled table is identical."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-table] relational table: rows = data points, columns = attributes\")\n# the entry's cow table: Name, Weight, Age, Height, Stomach temperature\nnames = np.array([\"Zenzi\", \"Berta\", \"Resi\"])\ntable = np.array([[100.0, 4.0, 100.0, 25.0],\n                  [140.0, 3.0, 130.0, 23.0],\n                  [120.0, 4.0, 120.0, 31.0]])\nX = table[:, [0, 1, 2]]                          # features: weight, age, height\ny = table[:, 3]                                  # label: stomach temperature\ncheck(\"each row is one data point, each column one attribute\",\n      table.shape == (3, 4))\nw = np.linalg.lstsq(np.c_[X, np.ones(3)], y, rcond=None)[0]\nperm = rng.permutation(3)\nw_shuffled = np.linalg.lstsq(np.c_[X[perm], np.ones(3)], y[perm],\n                             rcond=None)[0]\ncheck(\"row order is immaterial: shuffled table gives the same fit\",\n      np.allclose(w, w_shuffled, atol=1e-8))\ncheck(\"attribute domains bound the columns (weights within [100, 140])\",\n      X[:, 0].min() >= 100 and X[:, 0].max() <= 140)\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(5.2, 3.0))\nax.axis(\"off\")\ncell_text = [[n] + [f\"{v:g}\" for v in row]\n             for n, row in zip(names, table)]\ntab = ax.table(cellText=cell_text,\n               colLabels=[\"Name\", \"Weight\", \"Age\", \"Height\", \"Temp.\"],\n               loc=\"center\")\ntab.scale(1, 1.4)\nax.set_title(\"[P-table] dataset as a relation\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"dataset.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}