{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "transparency.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# transparency \u2014 Python demo\n\nNumerical companion to the entry [transparency](https://dictionaryofml.org/terms/transparency.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nOne block per technical paragraph of the entry (marked [P...]): the entry is a regulation term, so its legal paragraphs (EU AI Act Arts. 13/26/50/86, documentation duties) are expository; the demo illustrates the paragraph on ML methods that inherently offer transparency and the three credit-scoring paragraphs of the entry's Fig. 1. Self-contained (numpy/matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/transparency.py`](https://dictionaryofml.org/terms/transparency.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"transparency.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\ntransparency.py \u2014 numerical companion to the glossary entry\n'transparency'.\n\nOne block per technical paragraph of the entry (marked [P...]): the\nentry is a regulation term, so its legal paragraphs (EU AI Act\nArts. 13/26/50/86, documentation duties) are expository; the demo\nillustrates the paragraph on ML methods that inherently offer\ntransparency and the three credit-scoring paragraphs of the entry's\nFig. 1. Self-contained (numpy/matplotlib only), fixed seed.\n\nBlocks\n------\n[P-methods] \"Some ML methods inherently offer transparency\": (a) a\n            classification method quantifies the confidence of a\n            classification via the distance |h(x)| of the feature\n            vector from the decision boundary \u2014 predictions far from\n            the boundary are empirically far more reliable than\n            near-boundary ones, so disclosing this distance (as the\n            entry's medical example requires) is informative; (b) a\n            depth-2 decision tree is printable as human-readable\n            if-then rules that exactly reproduce its predictions.\n[P-read]    Art. 13 in the entry's Fig. 1: the learned hypothesis\n            h(x) = 1 + 3.5(1 - exp(-0.35 x)) maps an applicant's\n            income x' = 2.3 to a predicted credit score below the\n            approval threshold 3.8, and the deployer can read off the\n            prediction and its distance from the threshold.\n[P-train]   Art. 11 in the entry's Fig. 1: the drawn hypothesis has a\n            far smaller average loss on the training set of completed\n            loans than a constant prediction, and the shaded income\n            range lies entirely outside the range covered by the\n            training set \u2014 the limitation the documentation must\n            state.\n[P-cf]      Art. 86 in the entry's Fig. 1: the counterfactual income\n            x'' at which h reaches the approval threshold is\n            ln(5)/0.35 = 4.5984..., matching the diamond drawn at 4.6;\n            predictions at incomes above x'' exceed the threshold, so\n            the stated change indeed flips the decision.\n\nOutputs\n-------\ntransparency.png : preview figure (checking only).\n\nData generated by pythondemos/transparency.py.\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-methods]** \"Some ML methods inherently offer transparency\": (a) a classification method quantifies the confidence of a classification via the distance |h(x)| of the feature vector from the decision boundary \u2014 predictions far from the boundary are empirically far more reliable than near-boundary ones, so disclosing this distance (as the entry's medical example requires) is informative; (b) a depth-2 decision tree is printable as human-readable if-then rules that exactly reproduce its predictions."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-methods] (a) classification: confidence via the distance \"\n      \"|h(x)| from the decision boundary\")\nm = 4000\nX = rng.normal(size=(m, 2))\nw_true = np.array([2.0, -1.5])\np = 1 / (1 + np.exp(-(X @ w_true)))\ny = (rng.uniform(size=m) < p).astype(float)\nw = np.zeros(2)\nfor _ in range(400):                               # train a linear hypothesis\n    s = 1 / (1 + np.exp(-(X @ w)))\n    w -= 0.5 * X.T @ (s - y) / m                   # decrease the average loss\nh = X @ w                                          # h(x) = w^T x\npred = (h > 0).astype(float)\nlo = np.abs(h) < np.quantile(np.abs(h), 0.3)       # low-confidence tercile\nhi = np.abs(h) > np.quantile(np.abs(h), 0.7)       # high-confidence tercile\nacc_lo, acc_hi = np.mean(pred[lo] == y[lo]), np.mean(pred[hi] == y[hi])\nprint(f\"    accuracy at low / high |h(x)|: {acc_lo:.2f} / {acc_hi:.2f}\")\ncheck(\"distance from the decision boundary quantifies reliability: \"\n      \"far-from-boundary predictions are far more accurate\",\n      acc_hi > acc_lo + 0.15)\ncheck(\"disclosing the distance separates confident from uncertain \"\n      \"predictions (medical-example requirement)\",\n      acc_hi > 0.9)\n\nprint(\"[P-methods] (b) decision tree: human-readable rules\")\nx1_split, x2_split = 0.0, 0.5\ndef tree_predict(X):\n    out = np.empty(len(X))\n    for i, (a, b) in enumerate(X):\n        if a <= x1_split:\n            out[i] = 0.0 if b <= x2_split else 1.0\n        else:\n            out[i] = 1.0 if b <= x2_split else 0.0\n    return out\nrules = [\n    f\"IF x1 <= {x1_split} AND x2 <= {x2_split} THEN predict 0\",\n    f\"IF x1 <= {x1_split} AND x2 >  {x2_split} THEN predict 1\",\n    f\"IF x1 >  {x1_split} AND x2 <= {x2_split} THEN predict 1\",\n    f\"IF x1 >  {x1_split} AND x2 >  {x2_split} THEN predict 0\",\n]\nfor r in rules:\n    print(\"      \" + r)\ndef rules_predict(X):\n    out = np.empty(len(X))\n    for i, (a, b) in enumerate(X):\n        if a <= x1_split and b <= x2_split: out[i] = 0.0\n        elif a <= x1_split: out[i] = 1.0\n        elif b <= x2_split: out[i] = 1.0\n        else: out[i] = 0.0\n    return out\nXt = rng.normal(size=(500, 2))\ncheck(\"the printed if-then rules exactly reproduce the tree's \"\n      \"predictions on every input\",\n      np.array_equal(tree_predict(Xt), rules_predict(Xt)))\ncheck(\"the rule list is small enough to read (4 rules, depth 2)\",\n      len(rules) == 4)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-read]** Art. 13 in the entry's Fig. 1: the learned hypothesis h(x) = 1 + 3.5(1 - exp(-0.35 x)) maps an applicant's income x' = 2.3 to a predicted credit score below the approval threshold 3.8, and the deployer can read off the prediction and its distance from the threshold."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-read] Art. 13: read off the prediction and its distance \"\n      \"from the approval threshold\")\ndef h_credit(x):\n    return 1 + 3.5 * (1 - np.exp(-0.35 * x))\ntau = 3.8                                          # approval threshold\nx_prime = 2.3                                      # applicant's income\nscore = h_credit(x_prime)\nprint(f\"    h({x_prime}) = {score:.2f}, threshold {tau}, \"\n      f\"distance {tau - score:.2f}\")\ncheck(\"the applicant's predicted score falls below the approval \"\n      \"threshold\", score < tau)\ncheck(\"the distance from the threshold is readable from h alone\",\n      np.isclose(tau - score, tau - h_credit(x_prime)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-train]** Art. 11 in the entry's Fig. 1: the drawn hypothesis has a far smaller average loss on the training set of completed loans than a constant prediction, and the shaded income range lies entirely outside the range covered by the training set \u2014 the limitation the documentation must state."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-train] Art. 11: the hypothesis fits the training set; the \"\n      \"shaded incomes are not covered by it\")\ntrain = np.array([(0.7, 1.5), (1.2, 2.5), (1.8, 2.4), (2.3, 3.2),\n                  (2.9, 3.0), (3.4, 3.7), (3.9, 3.4), (4.4, 4.0),\n                  (4.9, 3.7), (5.4, 4.2), (5.9, 3.9)])\nx_tr, y_tr = train[:, 0], train[:, 1]\nloss_h = np.mean((y_tr - h_credit(x_tr)) ** 2)\nloss_const = np.mean((y_tr - y_tr.mean()) ** 2)\nprint(f\"    average loss: hypothesis {loss_h:.3f} vs constant \"\n      f\"{loss_const:.3f}\")\ncheck(\"the drawn hypothesis has smaller average loss on the training \"\n      \"set than a constant prediction\", loss_h < 0.5 * loss_const)\nshaded_lo, shaded_hi = 6.8, 9.5                    # shaded region of Fig. 1\ncheck(\"the shaded income range lies outside the range covered by the \"\n      \"training set (documented limitation)\", shaded_lo > x_tr.max())"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[P-cf]** Art. 86 in the entry's Fig. 1: the counterfactual income x'' at which h reaches the approval threshold is ln(5)/0.35 = 4.5984..., matching the diamond drawn at 4.6; predictions at incomes above x'' exceed the threshold, so the stated change indeed flips the decision."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[P-cf] Art. 86: the counterfactual income at which h reaches \"\n      \"the threshold\")\nx_cf = np.log(5) / 0.35                            # h(x_cf) = tau exactly\nprint(f\"    x'' = ln(5)/0.35 = {x_cf:.4f}\")\ncheck(\"h reaches the approval threshold at the counterfactual income\",\n      np.isclose(h_credit(x_cf), tau))\ncheck(\"matches the diamond drawn at income 4.6 in Fig. 1\",\n      abs(x_cf - 4.6) < 0.01)\ncheck(\"the change flips the decision: every income above x'' is \"\n      \"predicted above the threshold\",\n      np.all(h_credit(np.linspace(x_cf + 1e-6, 9.5, 200)) > tau))\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(1, 2, figsize=(9.6, 3.2))\nbins = np.quantile(np.abs(h), np.linspace(0, 1, 9))\naccs = [np.mean(pred[(np.abs(h) >= a) & (np.abs(h) < b)]\n        == y[(np.abs(h) >= a) & (np.abs(h) < b)])\n        for a, b in zip(bins[:-1], bins[1:])]\nax[0].plot(0.5 * (bins[:-1] + bins[1:]), accs, \"o-\")\nax[0].set_xlabel(\"distance |h(x)| from the decision boundary\")\nax[0].set_ylabel(\"empirical accuracy\")\nax[0].set_title(\"[P-methods] distance from the boundary tracks reliability\")\nxs = np.linspace(0, 9.5, 200)\nax[1].plot(xs, h_credit(xs), \"k-\", label=\"learned hypothesis h\")\nax[1].axhline(tau, ls=\":\", c=\"k\", label=f\"approval threshold {tau}\")\nax[1].plot(x_tr, y_tr, \"o\", c=\"C0\", label=\"training set\")\nax[1].plot([x_prime], [h_credit(x_prime)], \"s\", c=\"C1\",\n           label=\"applicant x'\")\nax[1].plot([x_cf], [tau], \"D\", c=\"C2\", label=\"counterfactual x''\")\nax[1].axvspan(shaded_lo, shaded_hi, color=\"0.9\",\n              label=\"not covered by training set\")\nax[1].set_xlabel(\"income (feature x)\")\nax[1].set_ylabel(\"credit score (label y)\")\nax[1].set_title(\"[P-read/-train/-cf] the credit-scoring example of Fig. 1\")\nax[1].legend(frameon=False, fontsize=7)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"transparency.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}