{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "logloss.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# logistic loss \u2014 Python demo\n\nNumerical companion to the entry [logistic loss](https://dictionaryofml.org/terms/logloss.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nBacks the entry's claims about the logistic loss numerically: divided by log 2 it is an upper bound on the zero-one loss with equality at margin 0; it is convex and differentiable everywhere while the hinge loss has a kink at margin 1; the average zero-one loss is flat almost everywhere in the model parameters, so GD gets no direction from it, whereas GD on the average logistic loss decreases it at every step and thereby also drives the zero-one loss down. Self-contained (numpy and matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/logloss.py`](https://dictionaryofml.org/terms/logloss.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"logloss.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nlogloss.py \u2014 numerical companion to the glossary entry 'logistic loss'.\n\nPurpose\n-------\nBacks the entry's claims about the logistic loss numerically: divided by\nlog 2 it is an upper bound on the zero-one loss with equality at margin 0;\nit is convex and differentiable everywhere while the hinge loss has a kink\nat margin 1; the average zero-one loss is flat almost everywhere in the\nmodel parameters, so GD gets no direction from it, whereas GD on the\naverage logistic loss decreases it at every step and thereby also drives\nthe zero-one loss down.  Self-contained (numpy and matplotlib only), fixed\nseed.\n\nSetup\n-----\nA spam filter as binary classification: m = 200 emails, each with two\nfeatures (a count of suspicious words and a count of links, scaled to\n[0, 1]) and a label y in {-1, +1}.  The linear hypothesis h(x) = w^T x + b\nis trained by ERM with the logistic loss, using GD with 300 updates that\nsubtract 0.5 times the gradient of the average logistic loss.\n\nBlocks\n------\n[B-bound] On a grid of margins, log(1 + exp(-margin)) / log 2 is at least\n          the zero-one loss everywhere, equal to it at margin 0 and\n          strictly above it elsewhere.\n[B-kink]  The hinge loss has different one-sided slopes at margin 1 (-1 and\n          0); the logistic loss has the same one-sided slopes at every\n          margin (differentiable everywhere).\n[B-flat]  Starting from random model parameters, 200 small random changes\n          of the parameters leave the average zero-one loss unchanged in\n          at least 95 percent of the cases and the average logistic loss\n          unchanged in none: the zero-one loss is flat almost everywhere\n          and gives GD no direction.\n[B-gd]    GD on the average logistic loss decreases it at every one of the\n          300 updates; at the end the average zero-one loss is below 0.1\n          and lies below the rescaled logistic loss, as the bound promises.\n\nOutputs\n-------\nlogloss_margin.csv : margin, logistic loss divided by log 2, zero-one loss\n                     and hinge loss on a grid of margins.\nlogloss_gd.csv     : update index, average logistic loss divided by log 2\n                     and average zero-one loss along the GD run.\nlogloss.png        : matplotlib preview of both panels (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nrng = np.random.default_rng(0)\nLOG2 = np.log(2.0)\n\n\ndef logistic(margin):\n    return np.log1p(np.exp(-margin))\n\n\ndef zero_one(margin):\n    return (margin <= 0).astype(float)\n\n\ndef hinge(margin):\n    return np.maximum(0.0, 1.0 - margin)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-bound]** On a grid of margins, log(1 + exp(-margin)) / log 2 is at least the zero-one loss everywhere, equal to it at margin 0 and strictly above it elsewhere."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "mg = np.linspace(-4.0, 4.0, 161)\nrescaled = logistic(mg) / LOG2\ngap = rescaled - zero_one(mg)\nat_zero = abs(logistic(0.0) / LOG2 - 1.0)\ncheck(f\"[B-bound] logistic loss / log 2 >= zero-one loss on the grid \"\n      f\"(smallest gap {gap.min():.2e}), equality at margin 0 \"\n      f\"(|difference| {at_zero:.1e})\",\n      gap.min() >= -1e-12 and at_zero < 1e-12\n      and (gap[np.abs(mg) > 1e-9] > 0).all())"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-kink]** The hinge loss has different one-sided slopes at margin 1 (-1 and 0); the logistic loss has the same one-sided slopes at every margin (differentiable everywhere)."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "eps = 1e-6\nhinge_left = (hinge(1.0) - hinge(1.0 - eps)) / eps\nhinge_right = (hinge(1.0 + eps) - hinge(1.0)) / eps\nlog_slopes = [((logistic(t) - logistic(t - eps)) / eps,\n               (logistic(t + eps) - logistic(t)) / eps)\n              for t in (-2.0, 0.0, 1.0, 2.0)]\nmax_jump = max(abs(a - b) for a, b in log_slopes)\ncheck(f\"[B-kink]  hinge loss slopes at margin 1: {hinge_left:+.2f} (left) \"\n      f\"vs {hinge_right:+.2f} (right); logistic loss: largest one-sided \"\n      f\"slope difference {max_jump:.1e}\",\n      abs(hinge_left + 1.0) < 1e-4 and abs(hinge_right) < 1e-4\n      and max_jump < 1e-4)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-flat]** Starting from random model parameters, 200 small random changes of the parameters leave the average zero-one loss unchanged in at least 95 percent of the cases and the average logistic loss unchanged in none: the zero-one loss is flat almost everywhere and gives GD no direction."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "m = 200\nX = rng.uniform(0.0, 1.0, (m, 2))\ny = np.where(0.8 * X[:, 0] + 1.1 * X[:, 1] - 0.9\n             + 0.08 * rng.standard_normal(m) > 0, 1.0, -1.0)\nXb = np.c_[X, np.ones(m)]                      # (x1, x2, 1): w and b in one\n\n\ndef avg_zero_one(w):\n    return float(np.mean(zero_one(y * (Xb @ w))))\n\n\ndef avg_logistic(w):\n    return float(np.mean(logistic(y * (Xb @ w))))\n\n\nw_rand = rng.standard_normal(3)\nsame_zo = sum(avg_zero_one(w_rand + 1e-4 * rng.standard_normal(3))\n              == avg_zero_one(w_rand) for _ in range(200))\nsame_lg = sum(avg_logistic(w_rand + 1e-4 * rng.standard_normal(3))\n              == avg_logistic(w_rand) for _ in range(200))\ncheck(f\"[B-flat]  small random changes of the model parameters leave the \"\n      f\"average zero-one loss unchanged in {same_zo}/200 cases, the \"\n      f\"average logistic loss in {same_lg}/200\",\n      same_zo >= 190 and same_lg == 0)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-gd]** GD on the average logistic loss decreases it at every one of the 300 updates; at the end the average zero-one loss is below 0.1 and lies below the rescaled logistic loss, as the bound promises."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "w = np.zeros(3)\nn_updates = 300\ntrace_lg, trace_zo = [], []\nfor t in range(n_updates):\n    margin = y * (Xb @ w)\n    grad = -(Xb * (y / (1.0 + np.exp(margin)))[:, None]).mean(axis=0)\n    w = w - 0.5 * grad\n    trace_lg.append(avg_logistic(w) / LOG2)\n    trace_zo.append(avg_zero_one(w))\ndecreasing = all(b < a for a, b in zip(trace_lg[:-1], trace_lg[1:]))\ncheck(f\"[B-gd]    GD lowers the average logistic loss at every update \"\n      f\"({trace_lg[0]:.3f} -> {trace_lg[-1]:.3f}); final average zero-one \"\n      f\"loss {trace_zo[-1]:.3f} <= rescaled logistic loss {trace_lg[-1]:.3f}\",\n      decreasing and trace_zo[-1] < 0.1 and trace_zo[-1] <= trace_lg[-1])\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"logloss_margin.csv\", \"w\") as fh:\n    fh.write(\"margin,logistic_rescaled,zeroone,hinge\\n\")\n    for a, b, c, d in zip(mg, rescaled, zero_one(mg), hinge(mg)):\n        fh.write(f\"{a:.3f},{b:.4f},{c:.1f},{d:.4f}\\n\")\nwith open(OUT_DIR / \"logloss_gd.csv\", \"w\") as fh:\n    fh.write(\"update,logistic_rescaled,zeroone\\n\")\n    for t, (a, b) in enumerate(zip(trace_lg, trace_zo), start=1):\n        fh.write(f\"{t},{a:.4f},{b:.4f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.2, 3.8))\nax1.plot(mg, rescaled, \"k-\", lw=1.4, label=\"logistic loss / log 2\")\nax1.plot(mg, zero_one(mg), \"k--\", lw=1.2, label=\"zero-one loss\")\nax1.plot(mg, hinge(mg), \"k:\", lw=1.6, label=\"hinge loss\")\nax1.set_xlabel(\"margin $y \\\\cdot h(\\\\mathbf{x})$\")\nax1.set_ylabel(\"loss\")\nax1.set_ylim(-0.1, 4.2)\nax1.set_title(\"three losses as functions of the margin\")\nax1.legend(frameon=False, fontsize=8)\nax2.plot(range(1, n_updates + 1), trace_lg, \"k-\", lw=1.4,\n         label=\"average logistic loss / log 2\")\nax2.plot(range(1, n_updates + 1), trace_zo, \"k--\", lw=1.2,\n         label=\"average zero-one loss\")\nax2.set_xlabel(\"GD update\")\nax2.set_ylabel(\"average loss over the m emails\")\nax2.set_title(\"GD on the logistic loss also lowers the zero-one loss\")\nax2.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"logloss.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'logloss_margin.csv'}, {OUT_DIR / 'logloss_gd.csv'}, \"\n      f\"{OUT_DIR / 'logloss.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}