{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "datapoint.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# data point \u2014 Python demo\n\nNumerical companion to the entry [data point](https://dictionaryofml.org/terms/datapoint.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/datapoint.py`](https://dictionaryofml.org/terms/datapoint.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"datapoint.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ndatapoint.py \u2014 numerical companion to the glossary entry 'data point'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-featlab]  A data point carries two categories of properties: features\n             (measurable/computable) and labels (higher-level facts). A\n             synthetic image data point yields pixel-intensity features\n             by computation, while its label (number of bright objects)\n             is fixed by construction \u2014 known to the \"expert\" that\n             generated the scene, not read off a sensor.\n[P-image]    The image example: color intensities of all pixels serve as\n             features x_1..x_d, and can be augmented with capture\n             metadata (timestamp, location) as further features.\n[P-choice]   Feature vs label is a design choice: the same attribute\n             (body weight) acts as a feature when predicting disease and\n             as the label when predicted from other attributes \u2014 both\n             predictions run on the same patient table.\n[P-labelnoise] Labels are error-prone proxies: one-sided label noise\n             (a fraction of positive training labels recorded as\n             negative, as when a diagnosis is missed) monotonically\n             depresses the fraction of positive test data points that\n             the learned hypothesis predicts correctly on clean test\n             data.\n[P-featnoise] Features are error-prone too: adding measurement noise to\n             the features at prediction time increases the error of a\n             fixed learned hypothesis monotonically in the sensor noise.\n\nOutputs\n-------\ndatapoint.png : preview figure (checking only).\n\nData generated by pythondemos/datapoint.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-featlab]** A data point carries two categories of properties: features (measurable/computable) and labels (higher-level facts). A synthetic image data point yields pixel-intensity features by computation, while its label (number of bright objects) is fixed by construction \u2014 known to the \"expert\" that generated the scene, not read off a sensor."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-featlab] features are computed; labels are known facts\")\ndef make_scene(n_objects):\n    img = 0.05 * rng.random((16, 16))\n    for _ in range(n_objects):\n        i, j = rng.integers(2, 14, size=2)\n        img[i - 1:i + 2, j - 1:j + 2] = 1.0\n    return img\n\nlabel_true = 3                                   # higher-level fact\nimg = make_scene(label_true)\nfeatures = np.array([img.mean(), img.std(), img.max()])\ncheck(\"features are computed from the data point itself\",\n      np.isclose(features[0], img.mean()))\ncheck(\"the label is not a pixel statistic (needs scene knowledge)\",\n      label_true not in np.round(features).astype(int)[:1])"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-image]** The image example: color intensities of all pixels serve as features x_1..x_d, and can be augmented with capture metadata (timestamp, location) as further features."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-image] pixel intensities + metadata as features\")\nx_pixels = img.flatten()\nx_meta = np.array([1717.0, 47.5])                # timestamp, latitude\nx = np.concatenate([x_pixels, x_meta])\ncheck(\"pixel features have length d = 256\", x_pixels.size == 16 * 16)\ncheck(\"metadata extends the features to d + 2\", x.size == 258)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-choice]** Feature vs label is a design choice: the same attribute (body weight) acts as a feature when predicting disease and as the label when predicted from other attributes \u2014 both predictions run on the same patient table."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-choice] feature vs label is a design choice\")\nm = 300\nweight = 60 + 20 * rng.random(m)\nage = 20 + 50 * rng.random(m)\ndisease = (0.03 * weight + 0.05 * age + rng.normal(0, 0.4, m) > 4.5)\n# design A: weight is a FEATURE for predicting the disease label\nXA = np.stack([weight, age], axis=1)\nwA = np.linalg.lstsq(np.c_[XA, np.ones(m)], disease.astype(float),\n                     rcond=None)[0]\naccA = np.mean((np.c_[XA, np.ones(m)] @ wA > 0.5) == disease)\n# design B: weight is the LABEL predicted from age and disease status\nXB = np.c_[age, disease.astype(float), np.ones(m)]\nwB = np.linalg.lstsq(XB, weight, rcond=None)[0]\ncheck(\"design A: weight used as a feature (prediction beats chance)\",\n      accA > 0.6)\ncheck(\"design B: weight used as the label (predicted better than by its average)\",\n      np.var(weight - XB @ wB) < np.var(weight))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-labelnoise]** Labels are error-prone proxies: one-sided label noise (a fraction of positive training labels recorded as negative, as when a diagnosis is missed) monotonically depresses the fraction of positive test data points that the learned hypothesis predicts correctly on clean test data."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-labelnoise] label noise degrades the learned hypothesis\")\ndef train_test_acc(flip, reps=25):\n    accs_r = []\n    for _ in range(reps):\n        Xtr, Xte = XA[:200], XA[200:]\n        ytr, yte = disease[:200].copy(), disease[200:]\n        pos = np.nonzero(ytr)[0]                  # missed diagnoses:\n        idx = rng.choice(pos, int(flip * pos.size), replace=False)\n        ytr[idx] = False                          # positives recorded negative\n        wn = np.linalg.lstsq(np.c_[Xtr, np.ones(200)],\n                             ytr.astype(float), rcond=None)[0]\n        pred = np.c_[Xte, np.ones(100)] @ wn > 0.5\n        accs_r.append(np.mean(pred[yte]))         # correct predictions on positive data points\n    return float(np.mean(accs_r))\n\naccs = [train_test_acc(f) for f in (0.0, 0.2, 0.4)]\nprint(f\"    correctly predicted positives at flip = 0, 0.2, 0.4: \"\n      f\"{accs[0]:.2f}, {accs[1]:.2f}, {accs[2]:.2f}\")\ncheck(\"correct positive predictions decrease with the label-noise level\",\n      accs[0] > accs[1] > accs[2])"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-featnoise]** Features are error-prone too: adding measurement noise to the features at prediction time increases the error of a fixed learned hypothesis monotonically in the sensor noise."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-featnoise] feature measurement noise degrades predictions\")\nw_clean = np.linalg.lstsq(np.c_[XA, np.ones(m)], disease.astype(float),\n                          rcond=None)[0]\nerrs = []\nfor s in (0.0, 5.0, 15.0):\n    Xn = XA + rng.normal(0, s, XA.shape)          # sensor uncertainty\n    errs.append(np.mean((np.c_[Xn, np.ones(m)] @ w_clean > 0.5)\n                        != disease))\nprint(f\"    error at sensor noise 0, 5, 15: \"\n      f\"{errs[0]:.2f}, {errs[1]:.2f}, {errs[2]:.2f}\")\ncheck(\"prediction error grows with feature noise\", errs[0] < errs[1] < errs[2])\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 2, figsize=(8.2, 3.0))\nax[0].imshow(img, cmap=\"gray\")\nax[0].set_title(f\"[P-featlab] data point (label = {label_true})\")\nax[1].plot([0, 0.2, 0.4], accs, \"o-\")\nax[1].set_xlabel(\"label-noise level\"); ax[1].set_ylabel(\"correct positive predictions\")\nax[1].set_title(\"[P-labelnoise]\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"datapoint.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}