{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "policy.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# policy (reinforcement learning) \u2014 Python demo\n\nNumerical companion to the entry [policy (reinforcement learning)](https://dictionaryofml.org/terms/policy.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nShows what a policy is and what it is for: a function from the state to the action that reacts to the state it meets, as against a plan (a fixed sequence of actions); a stochastic policy that draws the action from a probability distribution given the state; and the value function of a policy, the expected cumulative reward obtained by following it from a state, which is what the policy is learned to maximize. Self-contained (numpy and matplotlib only), fixed seed.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/policy.py`](https://dictionaryofml.org/terms/policy.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"policy.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\npolicy.py \u2014 numerical companion to the glossary entry 'policy'.\n\nPurpose\n-------\nShows what a policy is and what it is for: a function from the state to\nthe action that reacts to the state it meets, as against a plan (a fixed\nsequence of actions); a stochastic policy that draws the action from a\nprobability distribution given the state; and the value function of a\npolicy, the expected cumulative reward obtained by following it from a\nstate, which is what the policy is learned to maximize.  Self-contained\n(numpy and matplotlib only), fixed seed.\n\nSetup\n-----\nA thermostat controls a room.  The state is the room temperature in\ndegrees C (rounded to one decimal), the action is heating on or off.\nEach time step the temperature moves by +0.6 degrees with heating on and\nby -0.4 degrees with heating off, plus a disturbance from outside that\nis drawn at random (standard normal times 0.3) and, for 10 of the 60\ntime steps of an episode, an open window that pulls the temperature\ndown by 0.8 degrees per step.  The reward per time step is\n-|temperature - 21| - 0.1 if heating is on, i.e. the thermostat is\nrewarded for staying near 21 degrees at little heating.\n\nBlocks\n------\n[B-plan]    A threshold policy (heat if the temperature is below 21) is\n            compared with a plan, the fixed on/off sequence that the\n            policy itself produced in an undisturbed episode, replayed\n            under the disturbances: the policy keeps the temperature\n            within 1 degree of 21 on average, the plan does not.\n[B-stoch]   A stochastic policy, P(heating on | temperature) =\n            1 / (1 + exp(k (temperature - 21))), draws the action from\n            that distribution; its expected cumulative reward over an\n            episode, estimated from 200 episodes, grows with k and\n            comes within 6 of that of the threshold policy at k = 8.\n[B-value]   The value function of the threshold policy, the expected\n            cumulative reward from each starting temperature, is\n            estimated from 200 episodes per state; it is largest near\n            21 degrees and falls off on both sides.  Among the\n            thresholds 18, 19, ..., 24 the one maximizing the value at\n            the starting state 19 degrees is 21.\n\nOutputs\n-------\npolicy_trace.csv : time step, temperature under the threshold policy and\n                   under the plan in one disturbed episode.\npolicy_value.csv : starting temperature, value of the threshold policy\n                   with threshold 21.\npolicy.png       : matplotlib preview of the two panels (checking only).\n\"\"\"\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nOUT_DIR = Path(__file__).parent\n\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\nrng = np.random.default_rng(0)\n\nTARGET, STEPS, N_EP = 21.0, 60, 200\nHEAT_ON, HEAT_OFF, WINDOW = 0.6, -0.4, -0.8\n\n\ndef disturbance(t):\n    d = 0.3 * rng.standard_normal()\n    return d + (WINDOW if 20 <= t < 30 else 0.0)\n\n\ndef reward(temp, on):\n    return -abs(temp - TARGET) - (0.1 if on else 0.0)\n\n\ndef episode(choose, start, disturbed=True):\n    \"\"\"Run one episode; choose(temp, t) -> action (True = heating on).\n    Returns the cumulative reward, the temperatures and the actions.\"\"\"\n    temp = start; ret = 0.0; temps = [temp]; acts = []\n    for t in range(STEPS):\n        on = choose(round(temp, 1), t)\n        temp = temp + (HEAT_ON if on else HEAT_OFF) + (disturbance(t) if disturbed else 0.0)\n        ret += reward(temp, on); temps.append(temp); acts.append(on)\n    return ret, np.array(temps), acts\n\n\ndef threshold_policy(th):\n    return lambda temp, t: temp < th"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-plan]** A threshold policy (heat if the temperature is below 21) is compared with a plan, the fixed on/off sequence that the policy itself produced in an undisturbed episode, replayed under the disturbances: the policy keeps the temperature within 1 degree of 21 on average, the plan does not."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "_, _, plan_actions = episode(threshold_policy(TARGET), 19.0, disturbed=False)\nplan = lambda temp, t: plan_actions[t]           # the fixed sequence\nrng = np.random.default_rng(1)\n_, temps_policy, _ = episode(threshold_policy(TARGET), 19.0)\nrng = np.random.default_rng(1)\n_, temps_plan, _ = episode(plan, 19.0)\ndev_policy = float(np.abs(temps_policy[20:] - TARGET).mean())\ndev_plan = float(np.abs(temps_plan[20:] - TARGET).mean())\ncheck(f\"[B-plan]  mean distance from 21 degrees after the window opens: \"\n      f\"policy {dev_policy:.2f}, plan {dev_plan:.2f}\",\n      dev_policy < 1.0 and dev_plan > 2.0 * dev_policy)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-stoch]** A stochastic policy, P(heating on | temperature) = 1 / (1 + exp(k (temperature - 21))), draws the action from that distribution; its expected cumulative reward over an episode, estimated from 200 episodes, grows with k and comes within 6 of that of the threshold policy at k = 8."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "rng = np.random.default_rng(2)\n\n\ndef stochastic_policy(k):\n    return lambda temp, t: rng.random() < 1.0 / (1.0 + np.exp(k * (temp - TARGET)))\n\n\ndef expected_return(choose, start):\n    return float(np.mean([episode(choose, start)[0] for _ in range(N_EP)]))\n\n\nks = [0.5, 1.0, 2.0, 4.0, 8.0]\nval_stoch = [expected_return(stochastic_policy(k), 19.0) for k in ks]\nval_thr = expected_return(threshold_policy(TARGET), 19.0)\ncheck(\"[B-stoch] expected cumulative reward from 19 degrees: \"\n      + \", \".join(f\"k={k}: {v:.1f}\" for k, v in zip(ks, val_stoch))\n      + f\"; threshold policy {val_thr:.1f}\",\n      all(a < b for a, b in zip(val_stoch, val_stoch[1:]))\n      and val_stoch[-1] > val_thr - 6.0)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-value]** The value function of the threshold policy, the expected cumulative reward from each starting temperature, is estimated from 200 episodes per state; it is largest near 21 degrees and falls off on both sides. Among the thresholds 18, 19, ..., 24 the one maximizing the value at the starting state 19 degrees is 21."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "starts = np.arange(16.0, 26.5, 0.5)\nvalue = np.array([expected_return(threshold_policy(TARGET), s) for s in starts])\nthresholds = np.arange(18.0, 25.0, 1.0)\nval_by_th = [expected_return(threshold_policy(th), 19.0) for th in thresholds]\nbest_th = float(thresholds[int(np.argmax(val_by_th))])\ncheck(f\"[B-value] value from 16..26 degrees: max {value.max():.1f} at \"\n      f\"{starts[int(np.argmax(value))]:.1f} degrees, min {value.min():.1f}; \"\n      f\"best threshold from 19 degrees: {best_th:.0f}\",\n      abs(starts[int(np.argmax(value))] - TARGET) <= 1.0 and best_th == TARGET)\n\n# ---------------------------------------------------------------- CSV\nwith open(OUT_DIR / \"policy_trace.csv\", \"w\") as fh:\n    fh.write(\"t,temp_policy,temp_plan\\n\")\n    for t in range(STEPS + 1):\n        fh.write(f\"{t},{temps_policy[t]:.3f},{temps_plan[t]:.3f}\\n\")\nwith open(OUT_DIR / \"policy_value.csv\", \"w\") as fh:\n    fh.write(\"start,value\\n\")\n    for s, v in zip(starts, value):\n        fh.write(f\"{s:.1f},{v:.3f}\\n\")\n\n# -------------------------------------------------------------- preview\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.4, 3.8))\nax1.plot(range(STEPS + 1), temps_policy, \"k-\", lw=1.4, label=\"threshold policy (reacts to the state)\")\nax1.plot(range(STEPS + 1), temps_plan, \"k--\", lw=1.2, label=\"plan (fixed action sequence)\")\nax1.axhline(TARGET, color=\"0.6\", lw=0.8, ls=\":\")\nax1.axvspan(20, 30, color=\"0.9\", lw=0, label=\"window open\")\nax1.set_xlabel(\"time step\")\nax1.set_ylabel(\"room temperature (degrees C)\")\nax1.set_title(\"policy versus plan under a disturbance\")\nax1.legend(frameon=False, fontsize=7)\nax2.plot(starts, value, \"k.-\", lw=1.2, ms=5, label=\"threshold policy, threshold 21\")\nax2.set_xlabel(\"starting temperature (state)\")\nax2.set_ylabel(\"expected cumulative reward\")\nax2.set_title(\"value function of the policy\")\nax2.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"policy.png\", dpi=110)\n\nn_ok = sum(ok for _, ok in report)\nprint(f\"\\n{n_ok}/{len(report)} checks pass\")\nprint(f\"wrote {OUT_DIR / 'policy_trace.csv'}, {OUT_DIR / 'policy_value.csv'}, \"\n      f\"{OUT_DIR / 'policy.png'}\")\nif n_ok != len(report):\n    raise SystemExit(1)"
  }
 ]
}