{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "variance.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# variance \u2014 Python demo\n\nNumerical companion to the entry [variance](https://dictionaryofml.org/terms/variance.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/variance.py`](https://dictionaryofml.org/terms/variance.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"variance.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nvariance.py \u2014 numerical companion to the glossary entry 'variance'.\n\nOne block per paragraph of the entry (marked [P...]): each block verifies\nnumerically what the corresponding statement asserts. Self-contained\n(numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-def]    The variance of a real-valued RV y (such as the numeric label\n           of a data point) is E{(y - E{y})^2}: the empirical variance of\n           realizations of i.i.d. RVs with variance sigma^2 matches\n           sigma^2, and shifting\n           the RV leaves the variance unchanged (it measures spread\n           around the mean, not location). The entry's figure shows seven\n           measured daily maximum temperatures (see Data below): their\n           sample mean is exactly 20.5 degC, their sample variance\n           6.34/7, and the figure data is written to\n           variance_weather.csv. A dataset D of labels defines a\n           discrete RV y~ = y^(I) with I uniform on {1, ..., m}; the\n           variance of y~, computed from that uniform pmf, is the sample\n           variance of D (average squared deviation from the sample\n           mean).\n[P-min]    The variance is the minimum of the risk E{(y - b)^2} over\n           constants b, attained at b = E{y} (grid search).\n[P-erm]    The simplest ML problem: predict the label y without any\n           features. The hypothesis space consists of the constant maps\n           h(.) = b (a single bias term). ERM with the squared error\n           loss on the training set of measured temperatures amounts to\n           finding the constant that best predicts a label: the grid\n           minimizer of the training error is the sample mean 20.5 degC\n           (the dashed line of the entry's figure), the training error\n           of that learned hypothesis is the sample variance of the\n           labels, and at the population level every constant prediction\n           has risk >= variance \u2014 the baseline.\n[P-vector] For a random vector, E{||x - E{x}||^2} equals trace(C), the\n           sum of the per-entry variances (checked against the analytic\n           covariance matrix C = A A^T and against np.cov).\n\nOutputs\n-------\nvariance_weather.csv : figure data for the entry's pgfplots figure\n                       (day, tmax, mean, dev).\nvariance.png         : preview figure (checking only).\n\nData\n----\nDaily maximum temperatures (tmax), Helsinki Kaisaniemi weather station,\n24-30 August 2026, from the Finnish Meteorological Institute's open\ndata service (https://opendata.fmi.fi, stored query\nfmi::observations::weather::daily::simple), retrieved 2026-09-01.\n\nData generated by pythondemos/variance.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-def]** The variance of a real-valued RV y (such as the numeric label of a data point) is E{(y - E{y})^2}: the empirical variance of realizations of i.i.d. RVs with variance sigma^2 matches sigma^2, and shifting the RV leaves the variance unchanged (it measures spread around the mean, not location). The entry's figure shows seven measured daily maximum temperatures (see Data below): their sample mean is exactly 20.5 degC, their sample variance 6.34/7, and the figure data is written to variance_weather.csv. A dataset D of labels defines a discrete RV y~ = y^(I) with I uniform on {1, ..., m}; the variance of y~, computed from that uniform pmf, is the sample variance of D (average squared deviation from the sample mean)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-def] variance = E{(y - E y)^2}; sample variance = variance of \"\n      \"the dataset-induced discrete RV\")\nsigma = 1.7\nm = 10**6\ny = sigma * rng.standard_normal(m)\nvar_emp = np.mean((y - y.mean()) ** 2)\ncheck(\"empirical variance matches sigma^2 = 2.89\",\n      abs(var_emp - sigma**2) < 2e-2)\ncheck(\"shifting by +5 leaves the variance unchanged\",\n      abs(np.mean(((y + 5) - (y + 5).mean()) ** 2) - var_emp) < 1e-9)\nscales = [0.5, 1.0, 2.0]\nvars_scaled = [np.mean(((s * y) - (s * y).mean()) ** 2) for s in scales]\ncheck(\"scaling by s multiplies the variance by s^2\",\n      all(abs(v - s**2 * var_emp) < 1e-6 * max(1, s**2 * var_emp)\n          for s, v in zip(scales, vars_scaled)))\n# daily max temperatures y^(1), ..., y^(7), Helsinki Kaisaniemi,\n# 24-30 August 2026 (FMI open data; see the Data section above)\nd = np.array([20.1, 19.9, 22.2, 21.7, 20.3, 19.6, 19.7])\ndays = np.arange(1, d.size + 1)\ncheck(\"sample mean of the measured temperatures is 20.5 degC\",\n      np.isclose(d.mean(), 20.5))\ncheck(\"their sample variance is 6.34/7\",\n      np.isclose(np.mean((d - d.mean()) ** 2), 6.34 / 7))\nrows = np.column_stack([days, d, np.full(d.size, d.mean()), d - d.mean()])\nnp.savetxt(OUT_DIR / \"variance_weather.csv\", rows,\n           fmt=[\"%d\", \"%.1f\", \"%.4f\", \"%.4f\"], delimiter=\",\",\n           header=\"day,tmax,mean,dev\", comments=\"\")\np = np.full(d.size, 1 / d.size)               # uniform pmf of y~ = y^(I)\nmean_discrete = np.sum(p * d)\nvar_discrete = np.sum(p * (d - mean_discrete) ** 2)\nsample_var = np.mean((d - d.mean()) ** 2)\ncheck(\"pmf-based variance of y~ equals the sample variance\",\n      np.isclose(var_discrete, sample_var))\ncheck(\"matches np.var (ddof=0)\", np.isclose(sample_var, np.var(d)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-min]** The variance is the minimum of the risk E{(y - b)^2} over constants b, attained at b = E{y} (grid search)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-min] variance = min over constants b of E{(y - b)^2}\")\nys = y[:100_000]                               # draws from [P-def]\nb_grid = np.linspace(-3, 3, 601)\nrisk = np.array([np.mean((ys - b) ** 2) for b in b_grid])\nvar_ys = np.mean((ys - ys.mean()) ** 2)\ncheck(\"risk is minimized at the mean (grid resolution 0.01)\",\n      abs(b_grid[np.argmin(risk)] - ys.mean()) < 0.02)\ncheck(\"minimum of the risk equals the variance\",\n      abs(risk.min() - var_ys) < 1e-3)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-erm]** The simplest ML problem: predict the label y without any features. The hypothesis space consists of the constant maps h(.) = b (a single bias term). ERM with the squared error loss on the training set of measured temperatures amounts to finding the constant that best predicts a label: the grid minimizer of the training error is the sample mean 20.5 degC (the dashed line of the entry's figure), the training error of that learned hypothesis is the sample variance of the labels, and at the population level every constant prediction has risk >= variance \u2014 the baseline."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-erm] ERM over the constant maps h(.) = b: the learned bias \"\n      \"term is the sample mean; the variance is the baseline\")\ny_train = d                                    # temperatures as labels\nb_train_grid = np.linspace(18.0, 23.0, 501)\ntrain_err = np.array([np.mean((y_train - b) ** 2) for b in b_train_grid])\nb_hat = b_train_grid[np.argmin(train_err)]\ncheck(\"grid minimizer of the training error is the sample mean \"\n      \"(grid resolution 0.01)\", abs(b_hat - y_train.mean()) < 0.01)\ncheck(\"training error of the learned hypothesis = sample variance \"\n      \"of the labels\",\n      np.isclose(np.mean((y_train - y_train.mean()) ** 2), np.var(y_train)))\nyy = 1.5 * rng.standard_normal(200_000) + 3.0\nvar_yy = np.mean((yy - yy.mean()) ** 2)\nfor b in (2.0, 3.0, 4.5):\n    check(f\"constant prediction b = {b}: risk >= variance\",\n          np.mean((yy - b) ** 2) >= var_yy - 1e-9)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-vector]** For a random vector, E{||x - E{x}||^2} equals trace(C), the sum of the per-entry variances (checked against the analytic covariance matrix C = A A^T and against np.cov)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-vector] E{||x - E x||^2} = trace(C) = sum of entry variances\")\nA = np.array([[1.0, 0.0, 0.0], [0.5, 0.8, 0.0], [-0.2, 0.3, 1.1]])\nC = A @ A.T                                   # analytic covariance\nz = rng.standard_normal((10**6, 3))\nxv = z @ A.T                                  # zero-mean, covariance C\nsq_dev = np.mean(np.sum((xv - xv.mean(axis=0)) ** 2, axis=1))\ncheck(\"empirical E||x - Ex||^2 matches trace(C)\",\n      abs(sq_dev - np.trace(C)) < 2e-2)\nper_entry = np.array([np.mean((xv[:, j] - xv[:, j].mean()) ** 2)\n                      for j in range(3)])\ncheck(\"sum of per-entry variances equals trace(C)\",\n      abs(per_entry.sum() - np.trace(C)) < 2e-2)\ncheck(\"per-entry variances match diag(np.cov)\",\n      np.allclose(per_entry, np.diag(np.cov(xv.T, ddof=0)), atol=1e-9))\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 3, figsize=(12.5, 3.3))\nax[0].vlines(days, np.minimum(d, d.mean()), np.maximum(d, d.mean()),\n             colors=\"gray\", label=\"deviation\")\nax[0].plot(days, np.full(d.size, d.mean()), \"k--\", label=\"sample mean 20.5\")\nax[0].plot(days, d, \"o\", label=\"measurement\")\nax[0].set_xlabel(\"day (24-30 Aug 2026)\")\nax[0].set_ylabel(\"daily max temperature (degC)\")\nax[0].set_title(\"[P-def] Kaisaniemi temperatures around their mean\")\nax[0].legend(frameon=False)\nax[1].plot(b_train_grid, train_err)\nax[1].axvline(y_train.mean(), ls=\"--\", c=\"k\", label=\"sample mean\")\nax[1].axhline(np.var(y_train), ls=\":\", c=\"C1\", label=\"sample variance\")\nax[1].set_xlabel(\"constant prediction b\")\nax[1].set_ylabel(\"average of (y - b)$^2$ on the training set\")\nax[1].set_title(\"[P-erm] ERM over constants: minimum at the sample mean\")\nax[1].legend(frameon=False)\nax[2].bar(range(3), per_entry, label=\"per-entry variance\")\nax[2].axhline(np.trace(C), ls=\"--\", c=\"k\",\n              label=f\"trace(C) = {np.trace(C):.2f}\")\nax[2].plot([0, 1, 2], np.cumsum(per_entry), \"o-\", c=\"C1\",\n           label=\"cumulative sum\")\nax[2].set_xlabel(\"entry j\"); ax[2].set_ylabel(\"variance\")\nax[2].legend(frameon=False)\nax[2].set_title(\"[P-vector] variances sum to trace(C)\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"variance.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}