{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "label.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# label \u2014 Python demo\n\nNumerical companion to the entry [label](https://dictionaryofml.org/terms/label.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/label.py`](https://dictionaryofml.org/terms/label.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"label.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nlabel.py \u2014 numerical companion to the glossary entry 'label'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-choice]   Feature vs label is a design choice: with the \"oncologist\n             available\" design the prognosis enters the feature vector\n             and improves prediction of the target; without it the\n             prognosis is the label to be predicted from the remaining\n             features.\n[P-labelfun] The labeling function h-bar acts on the data point itself:\n             a deterministic function assigns each (fully observed)\n             synthetic patient its label.\n[P-partial]  The feature vector captures only part of the data point:\n             two data points with identical features carry different\n             labels, so no hypothesis reading only the features predicts\n             perfectly \u2014 the empirical error of the best such hypothesis\n             stays bounded away from zero, while a hypothesis with\n             access to the full data point achieves zero error.\n[P-noise]    The observed label is a noisy proxy: flipping a growing\n             fraction of the training labels degrades the test error\n             of the learned hypothesis (label noise degrades training\n             and generalization).\n\nOutputs\n-------\nlabel.png : preview figure (checking only).\n\nData generated by pythondemos/label.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\n# synthetic patients: (weight, age, hidden_marker) fully describe the\n# data point; the label depends on all three.\nm = 600\nweight = 60 + 20 * rng.random(m)\nage = 20 + 50 * rng.random(m)\nhidden = rng.normal(size=m)                       # never recorded as feature\nz = np.stack([weight, age, hidden], axis=1)       # the data point itself\nlabel_fun = lambda Z: (0.03 * Z[:, 0] + 0.04 * Z[:, 1]\n                       + 1.5 * Z[:, 2] > 4.0)     # labeling function h-bar\ny = label_fun(z)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-choice]** Feature vs label is a design choice: with the \"oncologist available\" design the prognosis enters the feature vector and improves prediction of the target; without it the prognosis is the label to be predicted from the remaining features."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-choice] feature vs label as a design choice\")\nprognosis = (0.5 * hidden + 0.2 * rng.normal(size=m) > 0)  # physician's call\ntarget = y\ndef acc(X, yy):\n    Xb = np.c_[X, np.ones(m)]\n    w = np.linalg.lstsq(Xb, yy.astype(float), rcond=None)[0]\n    return np.mean((Xb @ w > 0.5) == yy)\nacc_without = acc(np.c_[weight, age], target)\nacc_with = acc(np.c_[weight, age, prognosis.astype(float)], target)\nprint(f\"    accuracy without / with the prognosis feature: \"\n      f\"{acc_without:.2f} / {acc_with:.2f}\")\ncheck(\"prognosis used as a FEATURE improves the prediction\",\n      acc_with > acc_without + 0.05)\ncheck(\"prognosis used as the LABEL is predictable from features \"\n      \"(better than chance)\",\n      acc(np.c_[weight, age], prognosis) >= 0.5)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-labelfun]** The labeling function h-bar acts on the data point itself: a deterministic function assigns each (fully observed) synthetic patient its label."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-labelfun] the labeling function acts on the data point\")\ncheck(\"h-bar is deterministic on data points\",\n      np.array_equal(label_fun(z), y))\ncheck(\"h-bar reads the whole data point (hidden part matters)\",\n      not np.array_equal(label_fun(z),\n                         label_fun(np.c_[z[:, :2], np.zeros(m)])))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-partial]** The feature vector captures only part of the data point: two data points with identical features carry different labels, so no hypothesis reading only the features predicts perfectly \u2014 the empirical error of the best such hypothesis stays bounded away from zero, while a hypothesis with access to the full data point achieves zero error."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-partial] identical features, different labels\")\nx_feat = np.round(np.stack([weight, age], axis=1))  # recorded features\n# find twin data points with equal features but different labels\n_, inv, counts = np.unique(x_feat, axis=0, return_inverse=True,\n                           return_counts=True)\ntwin_exists = any(len(set(y[inv == g])) > 1\n                  for g in np.nonzero(counts > 1)[0])\ncheck(\"two data points share features but differ in the label\",\n      twin_exists)\nerr_feat = 1 - acc(x_feat, y)\nXfull = np.c_[z, np.ones(m)]\nw_full = np.linalg.lstsq(Xfull, y.astype(float), rcond=None)[0]\nerr_full = np.mean((Xfull @ w_full > 0.5) != y)\nprint(f\"    error with features only: {err_feat:.2f}; \"\n      f\"with the full data point: {err_full:.2f}\")\ncheck(\"features-only hypothesis cannot be perfect\", err_feat > 0.05)\ncheck(\"full-data-point hypothesis is (nearly) perfect\", err_full < 0.02)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-noise]** The observed label is a noisy proxy: flipping a growing fraction of the training labels degrades the test error of the learned hypothesis (label noise degrades training and generalization)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-noise] label noise degrades training and generalization\")\nXtr, Xte = x_feat[:400], x_feat[400:]\nytr0, yte = y[:400], y[400:]\nerrs = []\nfor flip in (0.0, 0.2, 0.4):\n    errs_r = []\n    for _ in range(25):\n        ytr = ytr0.copy()\n        idx = rng.choice(400, int(flip * 400), replace=False)\n        ytr[idx] = ~ytr[idx]\n        Xb = np.c_[Xtr, np.ones(400)]\n        w = np.linalg.lstsq(Xb, ytr.astype(float), rcond=None)[0]\n        errs_r.append(np.mean((np.c_[Xte, np.ones(200)] @ w > 0.5)\n                              != yte))\n    errs.append(float(np.mean(errs_r)))\nprint(f\"    test error at flip = 0, 0.2, 0.4: \"\n      f\"{errs[0]:.2f}, {errs[1]:.2f}, {errs[2]:.2f}\")\ncheck(\"test error grows with the label-noise level\",\n      errs[0] < errs[1] < errs[2])\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(4.6, 3.2))\nax.plot([0, 0.2, 0.4], errs, \"o-\")\nax.set_xlabel(\"fraction of flipped labels\"); ax.set_ylabel(\"test error\")\nax.set_title(\"[P-noise] label noise degrades generalization\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"label.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}