{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "loss.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# loss \u2014 Python demo\n\nNumerical companion to the entry [loss](https://dictionaryofml.org/terms/loss.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/loss.py`](https://dictionaryofml.org/terms/loss.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"loss.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nloss.py \u2014 numerical companion to the glossary entry 'loss'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-def]     The loss of a hypothesis on a single data point: the entry's\n            temperature example (label 12 C, prediction 8 C, squared\n            error loss 16) reproduced exactly; the squared error loss\n            takes a continuum of values while the 0/1 loss is binary.\n[P-nonneg]  Losses built from norms are nonnegative, but nonnegativity\n            is not required: the logarithmic loss -log p(y) is negative\n            wherever a density exceeds one (narrow Gaussian). Losses\n            exist without labels: the clustering loss (distance to the\n            assigned centroid) and the autoencoder reconstruction error\n            (via a rank-1 PCA reconstruction) are computed from features\n            alone.\n[P-emprisk] The average loss over the trainset is the empirical risk\n            minimized by ERM: the least-squares fit achieves a lower\n            average squared error loss than every perturbed candidate.\n[P-access]  Supervised vs reinforcement learning access: with labels,\n            the loss is evaluable for EVERY hypothesis on the trainset;\n            in a bandit simulation only the loss of the action actually\n            taken is observed, so most action-loss pairs remain\n            unobserved.\n\nOutputs\n-------\nloss.png : preview figure (checking only).\n\nData generated by pythondemos/loss.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-def]** The loss of a hypothesis on a single data point: the entry's temperature example (label 12 C, prediction 8 C, squared error loss 16) reproduced exactly; the squared error loss takes a continuum of values while the 0/1 loss is binary."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-def] loss on a single data point (temperature example)\")\ny_true, y_pred = 12.0, 8.0\nsq_loss = (y_true - y_pred) ** 2\ncheck(\"squared error loss of the entry's example equals 16\",\n      sq_loss == 16.0)\npreds = np.linspace(0, 24, 200)\nsq_vals = (y_true - preds) ** 2\nzo_vals = (np.round(preds) != y_true).astype(float)\ncheck(\"squared error loss takes a continuum of values\",\n      len(np.unique(sq_vals)) > 100)\ncheck(\"0/1 loss is binary\", set(np.unique(zo_vals)) == {0.0, 1.0})"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-nonneg]** Losses built from norms are nonnegative, but nonnegativity is not required: the logarithmic loss -log p(y) is negative wherever a density exceeds one (narrow Gaussian). Losses exist without labels: the clustering loss (distance to the assigned centroid) and the autoencoder reconstruction error (via a rank-1 PCA reconstruction) are computed from features alone."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-nonneg] nonnegativity is typical but not required\")\ncheck(\"norm-based loss is bounded below by zero\", np.all(sq_vals >= 0))\nsigma = 0.1                                        # narrow density\nlog_loss = -np.log(np.exp(0) / (sigma * np.sqrt(2 * np.pi)))\ncheck(\"log loss is negative where the density exceeds one\",\n      log_loss < 0)\n# unsupervised losses: clustering and autoencoder reconstruction\nX = np.vstack([rng.normal(0, 0.5, (50, 2)), rng.normal(4, 0.5, (50, 2))])\ncentroids = np.array([[0.0, 0.0], [4.0, 4.0]])\nassign = np.argmin(((X[:, None, :] - centroids) ** 2).sum(-1), axis=1)\nclus_loss = np.mean(np.linalg.norm(X - centroids[assign], axis=1) ** 2)\nXc = X - X.mean(0)\nu = np.linalg.svd(Xc, full_matrices=False)[2][0]   # top principal direction\nrecon = np.outer(Xc @ u, u)\nae_loss = np.mean(np.linalg.norm(Xc - recon, axis=1) ** 2)\ncheck(\"clustering loss (distance to assigned centroid) needs no label\",\n      clus_loss > 0)\ncheck(\"autoencoder loss = reconstruction error (rank-1 PCA)\",\n      ae_loss < np.mean(np.linalg.norm(Xc, axis=1) ** 2))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-emprisk]** The average loss over the trainset is the empirical risk minimized by ERM: the least-squares fit achieves a lower average squared error loss than every perturbed candidate."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-emprisk] average loss over the trainset = empirical risk\")\nmt = 80\nxt = rng.uniform(0, 10, mt)\nyt = 2.0 * xt + 1.0 + 0.4 * rng.normal(size=mt)\nc = np.polyfit(xt, yt, 1)\nrisk_hat = np.mean((yt - np.polyval(c, xt)) ** 2)\ncheck(\"ERM's solution minimizes the average loss (vs perturbations)\",\n      all(np.mean((yt - np.polyval(c + d, xt)) ** 2) > risk_hat\n          for d in ([0.1, 0], [-0.1, 0], [0, 0.5], [0, -0.5])))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-access]** Supervised vs reinforcement learning access: with labels, the loss is evaluable for EVERY hypothesis on the trainset; in a bandit simulation only the loss of the action actually taken is observed, so most action-loss pairs remain unobserved."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-access] loss access: supervised vs reinforcement learning\")\n# supervised: every hypothesis's loss is computable from the trainset\nslopes = np.linspace(0, 4, 21)\nsup_risks = [np.mean((yt - s * xt - 1.0) ** 2) for s in slopes]\ncheck(\"supervised: the loss is evaluable for every hypothesis\",\n      len(sup_risks) == 21 and np.all(np.isfinite(sup_risks)))\n# bandit: only the taken action's loss is revealed\ntrue_losses = np.array([0.6, 0.4, 0.9])           # three actions\nobserved = np.full((300, 3), np.nan)\nfor t in range(300):\n    a = rng.integers(3)                            # agent picks one action\n    observed[t, a] = true_losses[a] + 0.1 * rng.normal()\nfrac_unobserved = np.isnan(observed).mean()\ncheck(\"reinforcement learning: only the chosen action's loss is observed \"\n      \"(2/3 of entries unobserved)\",\n      abs(frac_unobserved - 2 / 3) < 0.05)\nest = np.nanmean(observed, axis=0)\ncheck(\"action-loss estimates form only from observed losses\",\n      np.argmin(est) == np.argmin(true_losses))\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 2, figsize=(8.4, 3.0))\nax[0].plot(preds, sq_vals, label=\"squared error\")\nax[0].plot(preds, 20 * zo_vals, \":\", label=\"0/1 (scaled)\")\nax[0].axvline(y_true, c=\"k\", lw=0.5)\nax[0].set_xlabel(\"prediction\"); ax[0].set_ylabel(\"loss\")\nax[0].legend(frameon=False)\nax[0].set_title(\"[P-def] losses on one data point\")\nax[1].plot(slopes, sup_risks, \"-\")\nax[1].set_xlabel(\"hypothesis (slope)\"); ax[1].set_ylabel(\"empirical risk\")\nax[1].set_title(\"[P-emprisk]\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"loss.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}