{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "binclass.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# binary classification \u2014 Python demo\n\nNumerical companion to the entry [binary classification](https://dictionaryofml.org/terms/binclass.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nBacks the entry's claims numerically: the average zero-one loss of a classifier is 1 minus its accuracy; the decision boundary of a linear classifier is a straight line that cuts the feature space into the two decision regions; the zero-one loss is not convex in the model parameters, while the logistic loss (divided by log 2) and the hinge loss are convex and upper bound it; on imbalanced data the constant answer already has a high accuracy, so the confusion matrix with precision and recall says more than the accuracy alone. Self-contained (numpy and matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/binclass.py`](https://dictionaryofml.org/terms/binclass.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"binclass.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nbinclass.py \u2014 numerical companion to the glossary entry 'binary\nclassification'.\n\nPurpose\n-------\nBacks the entry's claims numerically: the average zero-one loss of a\nclassifier is 1 minus its accuracy; the decision boundary of a linear\nclassifier is a straight line that cuts the feature space into the two\ndecision regions; the zero-one loss is not convex in the model\nparameters, while the logistic loss (divided by log 2) and the hinge loss\nare convex and upper bound it; on imbalanced data the constant answer\nalready has a high accuracy, so the confusion matrix with precision and\nrecall says more than the accuracy alone.  Self-contained (numpy and\nmatplotlib only), fixed seed.\n\nSetup\n-----\nThe weather service of the entry: for m = 200 days, two features (the\nmorning temperature and the hours of sunshine forecast, both scaled to\n[0, 1]) and the label y = +1 if the maximum temperature of the day\nreaches 25 degrees Celsius, y = -1 otherwise.  The labels are imbalanced:\nabout 20 percent of the days are +1.  A linear classifier\nh(x) = w^T x + b is obtained by logistic regression, i.e. by minimizing\nthe average logistic loss (300 small steps against its slope).\n\nBlocks\n------\n[B-zeroone]   The average zero-one loss of the learned classifier equals\n              1 minus its accuracy on the m days.\n[B-boundary]  Every day on the +1 side of the line w^T x + b = 0 is\n              predicted +1 and every day on the other side -1: the\n              decision regions are the two half-planes.\n[B-convex]    Along a straight line through the model parameters, the\n              average zero-one loss is piecewise constant and not convex\n              (it violates the midpoint inequality), while the average\n              logistic loss / log 2 and the average hinge loss satisfy it\n              and lie on or above the zero-one loss at every point.\n[B-imbalance] The constant answer \"no\" reaches an accuracy of about 0.8\n              with recall 0; the learned classifier's confusion matrix\n              gives a recall at least 0.2 below its accuracy, a gap the\n              accuracy alone hides.\n\nOutputs\n-------\nbinclass_data.csv : x1, x2, y of the m days and the predicted label.\nbinclass_line.csv : t, average zero-one, logistic / log 2 and hinge loss\n                    along the line of model parameters w(t) = w_hat + t * d.\nbinclass_cm.csv   : the confusion matrix of the learned classifier and of\n                    the constant answer, with accuracy, precision, recall.\nbinclass.png      : matplotlib preview of both panels (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nrng = np.random.default_rng(0)\nLOG2 = np.log(2.0)\n\nm = 200\nX = rng.uniform(0.0, 1.0, (m, 2))\nscore = 1.6 * X[:, 0] + 1.0 * X[:, 1] - 1.9 + 0.12 * rng.standard_normal(m)\ny = np.where(score > 0, 1.0, -1.0)\nXb = np.c_[X, np.ones(m)]                       # (x1, x2, 1): w and b in one\n\n\ndef avg_losses(w):\n    margin = y * (Xb @ w)\n    return (float(np.mean(margin <= 0)),\n            float(np.mean(np.log1p(np.exp(-margin))) / LOG2),\n            float(np.mean(np.maximum(0.0, 1.0 - margin))))\n\n\n# logistic regression: minimize the average logistic loss\nw = np.zeros(3)\nfor _ in range(300):\n    margin = y * (Xb @ w)\n    slope = -(Xb * (y / (1.0 + np.exp(margin)))[:, None]).mean(axis=0)\n    w = w - 1.0 * slope\ny_hat = np.where(Xb @ w > 0, 1.0, -1.0)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-zeroone]** The average zero-one loss of the learned classifier equals 1 minus its accuracy on the m days."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "zo = float(np.mean(y_hat != y))\nacc = float(np.mean(y_hat == y))\ncheck(f\"[B-zeroone]   average zero-one loss {zo:.3f} = 1 - accuracy \"\n      f\"{1 - acc:.3f}\", abs(zo - (1 - acc)) < 1e-12)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-boundary]** Every day on the +1 side of the line w^T x + b = 0 is predicted +1 and every day on the other side -1: the decision regions are the two half-planes."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "side = np.sign(Xb @ w)\ncheck(f\"[B-boundary]  predicted label equals the side of the line \"\n      f\"w^T x + b = 0 for all {m} days; {int(np.sum(y_hat == 1))} days \"\n      f\"on the +1 side\", np.array_equal(side, y_hat))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-convex]** Along a straight line through the model parameters, the average zero-one loss is piecewise constant and not convex (it violates the midpoint inequality), while the average logistic loss / log 2 and the average hinge loss satisfy it and lie on or above the zero-one loss at every point."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "d = np.array([1.0, -1.5, 0.4]); d /= np.linalg.norm(d)\nts = np.linspace(-3.0, 3.0, 121)\nline = np.array([avg_losses(w + t * d) for t in ts])\nzo_l, lg_l, hg_l = line[:, 0], line[:, 1], line[:, 2]\n\n\ndef midpoint_violations(v):\n    # f((a+b)/2) <= (f(a)+f(b))/2 for all grid pairs a < b of equal spacing\n    n = len(v); bad = 0; tot = 0\n    for i in range(n):\n        for j in range(i + 2, n, 2):\n            k = (i + j) // 2; tot += 1\n            if v[k] > (v[i] + v[j]) / 2 + 1e-9:\n                bad += 1\n    return bad, tot\n\n\nzo_bad, tot = midpoint_violations(zo_l)\nlg_bad, _ = midpoint_violations(lg_l)\nhg_bad, _ = midpoint_violations(hg_l)\nlevels = len(np.unique(np.round(zo_l, 6)))\ncheck(f\"[B-convex]    along the line of model parameters the zero-one loss takes \"\n      f\"{levels} values and violates the midpoint inequality {zo_bad} of \"\n      f\"{tot} times; logistic/log 2: {lg_bad}, hinge: {hg_bad}; both \"\n      f\"surrogates >= zero-one loss everywhere\",\n      zo_bad > 0 and lg_bad == 0 and hg_bad == 0\n      and (lg_l >= zo_l - 1e-12).all() and (hg_l >= zo_l - 1e-12).all())"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-imbalance]** The constant answer \"no\" reaches an accuracy of about 0.8 with recall 0; the learned classifier's confusion matrix gives a recall at least 0.2 below its accuracy, a gap the accuracy alone hides."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def confusion(pred):\n    tp = int(np.sum((pred == 1) & (y == 1))); fn = int(np.sum((pred == -1) & (y == 1)))\n    fp = int(np.sum((pred == 1) & (y == -1))); tn = int(np.sum((pred == -1) & (y == -1)))\n    accuracy = (tp + tn) / m\n    precision = tp / (tp + fp) if tp + fp else 0.0\n    recall = tp / (tp + fn) if tp + fn else 0.0\n    return tp, fn, fp, tn, accuracy, precision, recall\n\n\ncm_learned = confusion(y_hat)\ncm_const = confusion(-np.ones(m))\nfrac_pos = float(np.mean(y == 1))\ncheck(f\"[B-imbalance] {frac_pos:.2f} of the days are +1; the constant \"\n      f\"answer -1 has accuracy {cm_const[4]:.2f} and recall \"\n      f\"{cm_const[6]:.2f}; the learned classifier has accuracy \"\n      f\"{cm_learned[4]:.2f}, precision {cm_learned[5]:.2f}, recall \"\n      f\"{cm_learned[6]:.2f}\",\n      0.15 <= frac_pos <= 0.25 and cm_const[4] >= 0.75 and cm_const[6] == 0\n      and 0.5 < cm_learned[6] < cm_learned[4] - 0.2)\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"binclass_data.csv\", \"w\") as fh:\n    fh.write(\"x1,x2,y,yhat\\n\")\n    for a, b, c, e in zip(X[:, 0], X[:, 1], y, y_hat):\n        fh.write(f\"{a:.4f},{b:.4f},{int(c)},{int(e)}\\n\")\nwith open(OUT_DIR / \"binclass_line.csv\", \"w\") as fh:\n    fh.write(\"t,zeroone,logistic_rescaled,hinge\\n\")\n    for t, a, b, c in zip(ts, zo_l, lg_l, hg_l):\n        fh.write(f\"{t:.3f},{a:.4f},{b:.4f},{c:.4f}\\n\")\nwith open(OUT_DIR / \"binclass_cm.csv\", \"w\") as fh:\n    fh.write(\"classifier,tp,fn,fp,tn,accuracy,precision,recall\\n\")\n    for name, cm in ((\"learned\", cm_learned), (\"constant_no\", cm_const)):\n        fh.write(f\"{name},{cm[0]},{cm[1]},{cm[2]},{cm[3]},{cm[4]:.4f},\"\n                 f\"{cm[5]:.4f},{cm[6]:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.6, 4.0))\npos, neg = y == 1, y == -1\nax1.plot(X[pos, 0], X[pos, 1], \"ko\", ms=4, label=\"$y = +1$ (reaches 25 degrees)\")\nax1.plot(X[neg, 0], X[neg, 1], \"ks\", mfc=\"none\", ms=4, label=\"$y = -1$\")\nxx = np.linspace(0.0, 1.0, 50)\nax1.plot(xx, -(w[0] * xx + w[2]) / w[1], \"k-\", lw=1.4, label=\"decision boundary\")\nwrong = y_hat != y\nax1.plot(X[wrong, 0], X[wrong, 1], \"kx\", ms=9, mew=1.5, label=\"misclassified\")\nax1.set_xlim(0, 1); ax1.set_ylim(0, 1)\nax1.set_xlabel(\"morning temperature $x_1$ (scaled)\")\nax1.set_ylabel(\"sunshine forecast $x_2$ (scaled)\")\nax1.set_title(\"linear classifier: boundary and the two decision regions\")\nax1.legend(frameon=False, fontsize=7, loc=\"lower left\")\nax2.plot(ts, zo_l, \"k-\", lw=1.4, label=\"average zero-one loss\")\nax2.plot(ts, lg_l, \"k--\", lw=1.2, label=\"average logistic loss / log 2\")\nax2.plot(ts, hg_l, \"k:\", lw=1.6, label=\"average hinge loss\")\nax2.set_xlabel(\"position $t$ along a line of model parameters\")\nax2.set_ylabel(\"average loss over the m days\")\nax2.set_title(\"zero-one loss is not convex, its surrogates are\")\nax2.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"binclass.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'binclass_data.csv'}, {OUT_DIR / 'binclass_line.csv'}, \"\n      f\"{OUT_DIR / 'binclass_cm.csv'}, {OUT_DIR / 'binclass.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}