{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "explainability.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# explainability \u2014 Python demo\n\nNumerical companion to the entry [explainability](https://dictionaryofml.org/terms/explainability.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nQuantifies the (subjective) explainability of a trained hypothesis for a simulated user, via the two measures discussed in the entry: the deviation between the user-anticipated and the actual predictions on a test set, and the empirical conditional entropy of the predictions given the user's anticipations. Providing explanations (LIME-style local linear approximations) raises both measures. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/explainability.py`](https://dictionaryofml.org/terms/explainability.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"explainability.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nexplainability.py \u2014 numerical companion to the glossary entry\n'explainability'.\n\nPurpose\n-------\nQuantifies the (subjective) explainability of a trained hypothesis for a\nsimulated user, via the two measures discussed in the entry: the deviation\nbetween the user-anticipated and the actual predictions on a test set, and\nthe empirical conditional entropy of the predictions given the user's\nanticipations.  Providing explanations (LIME-style local linear\napproximations) raises both measures.  Self-contained (numpy/matplotlib\nonly), fixed seed.\n\nSetup\n-----\nTrained hypothesis: Gaussian-kernel ridge regression fit to m = 30 noisy\nsamples of a nonlinear function on [-3, 3] \u2014 opaque to a user who reasons\nin terms of linear maps.  The user is simulated as ridge-fitting a linear\nmap to a small set of labeled examples of the hypothesis' predictions\n(\"mental model\"), and anticipating predictions on a test set.\n\nBlocks\n------\n[B-lin]    For a LINEAR trained hypothesis, the simulated user anticipates\n           its test-set predictions almost exactly (mean squared deviation\n           < 1e-3): a linear hypothesis is highly explainable to this user.\n[B-opaque] For the kernel hypothesis, the user's anticipation deviates\n           strongly (mean squared deviation > 0.1): low explainability.\n[B-expl]   Explanations close the gap: given a LIME-style local linear\n           approximation around each test point (fit to perturbations near\n           the point), the user's anticipation error drops by a factor\n           of at least 10 compared with [B-opaque].\n[B-ent]    The empirical conditional entropy H(prediction | anticipation)\n           (both discretized into 8 bins) is smaller with explanations\n           than without: anticipations become more informative about the\n           predictions.\n[B-local]  One explanation shown in full: at the test point where the\n           user's anticipation without explanations deviates most, the\n           local linear approximation of the kernel hypothesis is written\n           out, and the anticipation read from it lies within 0.1 of the\n           prediction while the anticipation without explanations is off\n           by more than 0.3.\n\nOutputs\n-------\nexplainability_scatter.csv : test-set points with columns yhat (prediction\n                             of the kernel hypothesis), u_no (user\n                             anticipation without explanations), u_expl\n                             (with explanations), for the entry's pgfplots\n                             figure.\nexplainability_local.csv   : grid x with columns h (kernel hypothesis),\n                             u_glob (the user's linear mental model fitted\n                             without explanations) and g_loc (the local\n                             linear approximation at the point of [B-local]),\n                             for the left panel of the entry's figure.\nexplainability_point.csv   : one row with that point x0, the prediction\n                             yhat0 and the two anticipations u_no0, u_expl0.\nexplainability.png         : matplotlib preview of the two-panel figure\n                             (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nrng = np.random.default_rng(0)\n\n# ------------------------------------------------- trained hypotheses\nm = 30\nx_tr = np.sort(rng.uniform(-3.0, 3.0, m))\ny_tr = np.sin(1.5 * x_tr) + 0.5 * x_tr + 0.1 * rng.standard_normal(m)\n\nSIGMA, ALPHA = 0.6, 1e-2\n\n\ndef gauss_kernel(p, q):\n    return np.exp(-(p[:, None] - q[None, :]) ** 2 / (2.0 * SIGMA ** 2))\n\n\nbeta = np.linalg.solve(gauss_kernel(x_tr, x_tr) + ALPHA * m * np.eye(m),\n                       y_tr)\n\n\ndef h_kernel(p):\n    \"\"\"Opaque hypothesis: Gaussian-kernel ridge regression.\"\"\"\n    return gauss_kernel(np.atleast_1d(p), x_tr) @ beta\n\n\nw_lin = np.polyfit(x_tr, y_tr, 1)\n\n\ndef h_linear(p):\n    \"\"\"Interpretable hypothesis: linear map (with intercept).\"\"\"\n    return np.polyval(w_lin, p)\n\n\ndef user_linear_fit(h, x_examples):\n    \"\"\"The simulated user ridge-fits a linear map (slope, intercept) to\n    labeled examples of the hypothesis' predictions.\"\"\"\n    A = np.c_[x_examples, np.ones_like(x_examples)]\n    return np.linalg.solve(A.T @ A + 1e-6 * np.eye(2), A.T @ h(x_examples))\n\n\ndef user_anticipation(h, x_examples, x_test):\n    \"\"\"The simulated user fits a linear map to labeled examples of the\n    hypothesis' predictions and extrapolates to the test points.\"\"\"\n    w = user_linear_fit(h, x_examples)\n    return np.c_[x_test, np.ones_like(x_test)] @ w\n\n\nx_ex = np.linspace(-3.0, 3.0, 6)          # examples shown to the user\nx_te = np.linspace(-2.8, 2.8, 40)         # test set to anticipate"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-lin]** For a LINEAR trained hypothesis, the simulated user anticipates its test-set predictions almost exactly (mean squared deviation < 1e-3): a linear hypothesis is highly explainable to this user."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "dev_lin = float(np.mean(\n    (user_anticipation(h_linear, x_ex, x_te) - h_linear(x_te)) ** 2))\ncheck(f\"[B-lin]    linear hypothesis anticipated (msd {dev_lin:.1e})\",\n      dev_lin < 1e-3)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-opaque]** For the kernel hypothesis, the user's anticipation deviates strongly (mean squared deviation > 0.1): low explainability."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "u_no = user_anticipation(h_kernel, x_ex, x_te)\nyhat = h_kernel(x_te)\ndev_no = float(np.mean((u_no - yhat) ** 2))\ncheck(f\"[B-opaque] kernel hypothesis not anticipated (msd {dev_no:.2f})\",\n      dev_no > 0.1)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-expl]** Explanations close the gap: given a LIME-style local linear approximation around each test point (fit to perturbations near the point), the user's anticipation error drops by a factor of at least 10 compared with [B-opaque]."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "u_expl = np.empty_like(x_te)\nw_loc = np.empty((len(x_te), 2))             # local (slope, intercept)\nfor i, xt in enumerate(x_te):\n    x_loc = xt + 0.25 * rng.standard_normal(20)   # LIME-style perturbations\n    w_loc[i] = user_linear_fit(h_kernel, x_loc)\n    u_expl[i] = w_loc[i] @ np.array([xt, 1.0])\ndev_expl = float(np.mean((u_expl - yhat) ** 2))\ncheck(f\"[B-expl]   explanations shrink the deviation \"\n      f\"({dev_no:.2f} -> {dev_expl:.4f})\", dev_no > 10.0 * dev_expl)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-ent]** The empirical conditional entropy H(prediction | anticipation) (both discretized into 8 bins) is smaller with explanations than without: anticipations become more informative about the predictions."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def cond_entropy(target, given, bins=8):\n    \"\"\"Empirical conditional entropy H(target | given) in bits.\"\"\"\n    lo, hi = min(target.min(), given.min()), max(target.max(), given.max())\n    edges = np.linspace(lo, hi + 1e-9, bins + 1)\n    t = np.digitize(target, edges) - 1\n    g = np.digitize(given, edges) - 1\n    joint = np.zeros((bins, bins))\n    for ti, gi in zip(t, g):\n        joint[ti, gi] += 1\n    joint /= joint.sum()\n    pg = joint.sum(axis=0)\n    with np.errstate(divide=\"ignore\", invalid=\"ignore\"):\n        cond = joint / pg[None, :]\n        terms = np.where(joint > 0, joint * np.log2(cond), 0.0)\n    return float(-terms.sum())\n\n\nH_no = cond_entropy(yhat, u_no)\nH_expl = cond_entropy(yhat, u_expl)\ncheck(f\"[B-ent]    conditional entropy drops ({H_no:.2f} -> {H_expl:.2f} \"\n      f\"bits)\", H_expl < H_no)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-local]** One explanation shown in full: at the test point where the user's anticipation without explanations deviates most, the local linear approximation of the kernel hypothesis is written out, and the anticipation read from it lies within 0.1 of the prediction while the anticipation without explanations is off by more than 0.3."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "i0 = int(np.argmax(np.abs(u_no - yhat)))     # worst anticipation without\nx0, yhat0, u_no0, u_expl0 = x_te[i0], yhat[i0], u_no[i0], u_expl[i0]\nx_grid = np.linspace(-3.0, 3.0, 121)\nw_glob = user_linear_fit(h_kernel, x_ex)     # the user's mental model\nu_glob = np.c_[x_grid, np.ones_like(x_grid)] @ w_glob\ng_loc = np.c_[x_grid, np.ones_like(x_grid)] @ w_loc[i0]\ncheck(f\"[B-local]  at x0={x0:.2f}: explanation {abs(u_expl0 - yhat0):.3f} \"\n      f\"off, no explanation {abs(u_no0 - yhat0):.2f} off\",\n      abs(u_expl0 - yhat0) < 0.1 and abs(u_no0 - yhat0) > 0.3)\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"explainability_scatter.csv\", \"w\") as f:\n    f.write(\"yhat,u_no,u_expl\\n\")\n    for a, b, c in zip(yhat, u_no, u_expl):\n        f.write(f\"{a:.4f},{b:.4f},{c:.4f}\\n\")\nwith open(OUT_DIR / \"explainability_local.csv\", \"w\") as f:\n    f.write(\"x,h,u_glob,g_loc\\n\")\n    for a, b, c, d in zip(x_grid, h_kernel(x_grid), u_glob, g_loc):\n        f.write(f\"{a:.4f},{b:.4f},{c:.4f},{d:.4f}\\n\")\nwith open(OUT_DIR / \"explainability_point.csv\", \"w\") as f:\n    f.write(\"x0,yhat0,u_no0,u_expl0\\n\")\n    f.write(f\"{x0:.4f},{yhat0:.4f},{u_no0:.4f},{u_expl0:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (axl, ax) = plt.subplots(1, 2, figsize=(8.4, 4.2))\nwin = np.abs(x_grid - x0) <= 0.6\naxl.plot(x_grid, h_kernel(x_grid), \"k-\", lw=1.2,\n         label=\"prediction $\\\\hat{h}(x)$\")\naxl.plot(x_grid, u_glob, \"k--\", lw=0.8,\n         label=\"user's linear mental model\")\naxl.plot(x_grid[win], g_loc[win], \"-\", color=\"0.55\", lw=3,\n         label=\"explanation: local linear approximation\")\naxl.axvline(x0, color=\"k\", ls=\":\", lw=0.6)\naxl.plot([x0], [u_no0], \"ks\", mfc=\"none\", ms=6,\n         label=\"anticipation without explanation\")\naxl.plot([x0], [u_expl0], \"k.\", ms=9, label=\"anticipation with explanation\")\naxl.set_xlabel(\"feature $x$\")\naxl.set_ylabel(\"label $y$\")\naxl.set_title(f\"one explanation, at $x_0={x0:.2f}$\")\naxl.legend(frameon=False, fontsize=7, loc=\"lower right\")\nlim = [yhat.min() - 0.3, yhat.max() + 0.3]\nax.plot(lim, lim, \"k--\", lw=0.8)\nax.plot(yhat, u_no, \"ks\", mfc=\"none\", ms=4, label=\"without explanations\")\nax.plot(yhat, u_expl, \"k.\", ms=5, label=\"with explanations\")\nax.set_xlabel(\"prediction $\\\\hat{h}(x)$\")\nax.set_ylabel(\"user anticipation\")\nax.set_title(\"explanations align user anticipation\")\nax.legend(frameon=False, fontsize=8)\nax.set_aspect(\"equal\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"explainability.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'explainability_scatter.csv'}, \"\n      f\"{OUT_DIR / 'explainability_local.csv'}, \"\n      f\"{OUT_DIR / 'explainability_point.csv'}, \"\n      f\"{OUT_DIR / 'explainability.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}