{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "onlinelearning.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# online learning \u2014 Python demo\n\nNumerical companion to the entry [online learning](https://dictionaryofml.org/terms/onlinelearning.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nNumerical companion to the glossary entry 'onlinelearning'. The hourly air temperature at Krems an der Donau in 2024 (GeoSphere Austria archive, station 3805) is the data stream. At hour t the data point is z^(t) = (x^(t), y^(t)): the feature vector x^(t) holds a constant one and the readings of the three previous hours, the label y^(t) is the reading of hour t. Before y^(t) is known, the current model parameters w^(t) issue the forecast <w^(t), x^(t)>; then y^(t) arrives, the squared error loss L(z^(t), w^(t)) = (y^(t) - <w^(t), x^(t)>)^2 is incurred, and the model parameters are updated. Two online methods run through the same rounds: online GD, one projected gradient step per round on the newest loss, and follow-the-leader, which after each round re-solves the least-squares problem over all data points seen so far (in closed form, by the recursive least-squares update).\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/onlinelearning.py`](https://dictionaryofml.org/terms/onlinelearning.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"onlinelearning.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"Online learning on an hourly temperature stream: the round protocol,\nonline gradient descent against follow-the-leader with full information,\nHedge against Exp3 with bandit feedback, and the regret per round\nagainst the best fixed choice in hindsight.\n\nPurpose\n-------\nNumerical companion to the glossary entry 'onlinelearning'.  The hourly\nair temperature at Krems an der Donau in 2024 (GeoSphere Austria archive,\nstation 3805) is the data stream.  At hour t the data point is\nz^(t) = (x^(t), y^(t)): the feature vector x^(t) holds a constant one\nand the readings of the three previous hours, the label y^(t) is the\nreading of hour t.  Before y^(t) is known, the current model parameters\nw^(t) issue the forecast <w^(t), x^(t)>; then y^(t) arrives, the squared\nerror loss L(z^(t), w^(t)) = (y^(t) - <w^(t), x^(t)>)^2 is incurred, and\nthe model parameters are updated.  Two online methods run through the\nsame rounds: online GD, one projected gradient step per round on the\nnewest loss, and follow-the-leader, which after each round re-solves\nthe least-squares problem over all data points seen so far (in closed\nform, by the recursive least-squares update).\n\nA second problem on the same stream has bandit feedback.  The actions\nare A = 5 fixed forecasters (persistence, the reading of the previous\nday, the mean of the last three readings, a linear extrapolation, and\nthe running mean of the year so far), the model parameters w^(t) are\nthe pmf over the forecasters from which one is drawn, and the loss of\nround t is the absolute error of the drawn forecaster's forecast,\nscaled to [0, 1].  With full information all five errors are observed\nafter each hour and Hedge (mirror descent on the simplex with the KL\ndivergence) updates the pmf; with bandit feedback only the error\nof the drawn forecaster is observed, and Exp3 updates the pmf from an\nimportance-weighted estimate of the loss vector.\n\nThe demo checks the entry's claims: every forecast uses only data\npoints of earlier hours; the regret of online GD after T rounds,\nrelative to the model parameters that are best in hindsight, stays\nbelow the bound D G sqrt(T) of the online GD entry, and the regret per\nround falls as T grows; follow-the-leader has a smaller regret on this\nstream; both forecasts beat the persistence forecast (the last reading)\non the accumulated loss; with full information, Hedge's regret against\nthe best single forecaster stays below sqrt(2 T log A); with bandit\nfeedback, Exp3's regret stays below 2 sqrt((e - 1) T A log A) and\nexceeds Hedge's, the price of observing one entry of the loss vector\ninstead of all of them; and the regret per round of both vanishes.\n\nDeterministic: fixed seed for Exp3's draws; the readings are fetched\nfrom the public archive.  Self-contained: numpy + matplotlib only.\n\nBlocks\n------\n[B-data]     8784 hourly readings of 2024 (station 3805), turned into\n             data points with three lagged readings as features.\n[B-rounds]   The round protocol for online GD: forecast, reveal, loss,\n             projected gradient step onto the ball of radius D/2.\n             Check that every forecast uses only earlier hours.\n[B-leader]   Follow-the-leader through the same rounds by the recursive\n             least-squares update; check it equals the least-squares\n             solution over all data points seen so far.\n[B-regret]   The regret of both methods after T rounds against the best\n             fixed model parameters in hindsight, for T up to 8784, and\n             the bound D G sqrt(T); check the bound, the vanishing\n             regret per round, and the comparison with persistence.\n[B-bandit]   Five fixed forecasters as actions; Hedge with full\n             information and Exp3 with bandit feedback through the same\n             rounds; regret against the best single forecaster in\n             hindsight and the two bounds; check both bounds, that\n             Exp3's regret exceeds Hedge's, and that both regrets per\n             round vanish.\n[B-figures]  Data for the entry's first two figures: three days of the\n             week with an offline forecaster trained once on the first\n             day and then fixed, against online GD started from zero\n             at hour 0 with a constant learning rate; and\n             the loss functions of five consecutive hours for a scalar\n             forecaster, with the online GD iterates on them.\n[B-plot]     Write the regret per round against T for all four methods\n             and a week of forecasts, plus the preview.\n\nOutputs\n-------\nonlinelearning_temperature.csv : hour, temperature -- the readings\nonlinelearning_regret.csv : T, ogd, leader, bound -- regret per round\n                       of online GD and follow-the-leader after T\n                       rounds, and D G / sqrt(T)\nonlinelearning_bandit.csv : T, hedge, exp3, hedgebound, exp3bound --\n                       regret per round of Hedge (full information) and\n                       Exp3 (bandit feedback) against the best single\n                       forecaster, and the two bounds divided by T\nonlinelearning_week.csv : hour, reading, ogd, leader -- one week of\n                       readings and forecasts\nonlinelearning_offline.csv : hour, reading, offline, online -- three\n                       days of readings, the forecasts of a least-squares\n                       forecaster trained once on the first day and then\n                       fixed (nan for the first day), and the forecasts of\n                       online GD started from zero at hour 0\nonlinelearning_lossfuncs.csv : w, f1, ..., f5 -- the squared error loss\n                       functions (y^(t) - w)^2 of five consecutive hours\n                       on a grid of the scalar forecast w\nonlinelearning_lossvalues.csv : t, y, w, f -- for those hours the\n                       reading, the online GD iterate w^(t) and the\n                       revealed value f^(t)(w^(t))\nonlinelearning.png : preview (checking only)\n\"\"\"\n\nimport json\nimport urllib.request\nfrom pathlib import Path\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nOUT_DIR = Path(__file__).parent\nARCHIVE = \"https://dataset.api.hub.geosphere.at/v1/station/historical/\"\nSTATION = 3805\n\nreport = []                         # collects (check name, pass/fail) pairs\n\n\ndef check(name, ok):                # records and prints one verification\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-data]** 8784 hourly readings of 2024 (station 3805), turned into data points with three lagged readings as features."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def fetch(resource, parameter, start, end):\n    \"\"\"One parameter of one station from the GeoSphere Austria archive.\"\"\"\n    url = (f\"{ARCHIVE}{resource}?parameters={parameter}\"\n           f\"&station_ids={STATION}&start={start}&end={end}\")\n    with urllib.request.urlopen(url, timeout=180) as resp:\n        payload = json.load(resp)\n    values = payload[\"features\"][0][\"properties\"][\"parameters\"]\n    key = list(values.keys())[0]\n    return (np.array(values[key][\"data\"], dtype=float),\n            [t[:13] for t in payload[\"timestamps\"]])\n\n\ntemp, stamps = fetch(\"klima-v2-1h\", \"tl\", \"2024-01-01T00:00\", \"2024-12-31T23:00\")\ntemp = np.where(np.isnan(temp), np.nanmean(temp), temp)      # a few gaps\nwith open(OUT_DIR / \"onlinelearning_temperature.csv\", \"w\") as f:\n    f.write(\"hour,temperature\\n\")\n    for h, v in zip(stamps, temp):\n        f.write(f\"{h},{v:.1f}\\n\")\nLAGS = 3\nX = np.column_stack([np.ones(len(temp) - LAGS)]\n                    + [temp[LAGS - k - 1:len(temp) - k - 1] for k in range(LAGS)])\ny = temp[LAGS:]                                   # label: the reading of hour t\nT_ALL = len(y)\nprint(f\"  {len(temp)} hourly readings of 2024 at station {STATION}; \"\n      f\"{T_ALL} data points with a constant and {LAGS} lagged readings as features\")\ncheck(\"[B-data] one data point per hour after the first three, with four features\",\n      X.shape == (T_ALL, LAGS + 1) and T_ALL > 8000)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-rounds]** The round protocol for online GD: forecast, reveal, loss, projected gradient step onto the ball of radius D/2. Check that every forecast uses only earlier hours."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "RADIUS = 30.0                                     # W = ball of radius D/2, D = 60\nD = 2.0 * RADIUS\n\n\ndef project(w):\n    n = np.linalg.norm(w)\n    return w if n <= RADIUS else w * (RADIUS / n)\n\n\ndef run_online_gd(X, y, lrate):\n    \"\"\"Forecast, reveal, loss, projected gradient step; returns the\n    forecasts, the losses, the largest gradient norm and the iterates.\"\"\"\n    d = X.shape[1]\n    w = np.zeros(d)\n    forecasts, losses, G = np.empty(len(y)), np.empty(len(y)), 0.0\n    for t in range(len(y)):\n        forecasts[t] = w @ X[t]                   # issued before y[t] is known\n        losses[t] = (y[t] - forecasts[t]) ** 2    # y[t] revealed, loss incurred\n        g = -2.0 * (y[t] - forecasts[t]) * X[t]   # gradient of the newest loss\n        G = max(G, float(np.linalg.norm(g)))\n        w = project(w - lrate(t + 1) * g)         # the update\n    return forecasts, losses, G\n\n\nforecasts_ogd, losses_ogd, G = run_online_gd(X, y, lambda t: 5e-3 / np.sqrt(t))\n# a second run whose stream is cut at hour 1000 must agree on the first 1000 forecasts\nf_cut, _, _ = run_online_gd(X[:1000], y[:1000], lambda t: 5e-3 / np.sqrt(t))\nprint(f\"  online GD: mean squared error of the forecasts {losses_ogd.mean():.2f}, \"\n      f\"largest gradient norm G = {G:.0f}\")\ncheck(\"[B-rounds] every forecast uses only data points of earlier hours: cutting the \"\n      \"stream at hour 1000 leaves the first 1000 forecasts unchanged\",\n      np.array_equal(forecasts_ogd[:1000], f_cut))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-leader]** Follow-the-leader through the same rounds by the recursive least-squares update; check it equals the least-squares solution over all data points seen so far."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def run_leader(X, y, delta=1e3):\n    \"\"\"After each round, the least-squares minimizer over all data points\n    seen so far, updated recursively (Sherman-Morrison); the forecast of\n    round t uses the minimizer over rounds 1..t-1.\"\"\"\n    d = X.shape[1]\n    w = np.zeros(d)\n    P = delta * np.eye(d)                          # inverse of (X^T X + I/delta), delta large\n    forecasts, losses = np.empty(len(y)), np.empty(len(y))\n    for t in range(len(y)):\n        forecasts[t] = w @ X[t]\n        losses[t] = (y[t] - forecasts[t]) ** 2\n        k = P @ X[t] / (1.0 + X[t] @ P @ X[t])\n        w = w + k * (y[t] - forecasts[t])\n        P = P - np.outer(k, X[t] @ P)\n    return forecasts, losses, w\n\n\nforecasts_ftl, losses_ftl, w_ftl = run_leader(X, y)\nw_ls = np.linalg.lstsq(X, y, rcond=None)[0]\nprint(f\"  follow-the-leader: mean squared error of the forecasts {losses_ftl.mean():.2f}; \"\n      f\"final parameters {np.round(w_ftl, 3).tolist()} vs the minimizer over all \"\n      f\"data points {np.round(w_ls, 3).tolist()}\")\ncheck(\"[B-leader] the recursive update ends at the least-squares solution over all \"\n      \"data points (up to the tiny initial value of P)\",\n      np.abs(w_ftl - w_ls).max() < 1e-2)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-regret]** The regret of both methods after T rounds against the best fixed model parameters in hindsight, for T up to 8784, and the bound D G sqrt(T); check the bound, the vanishing regret per round, and the comparison with persistence."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def hindsight_loss(T):\n    \"\"\"Accumulated loss of the best fixed w in W over the first T rounds.\"\"\"\n    w = np.linalg.lstsq(X[:T], y[:T], rcond=None)[0]\n    assert np.linalg.norm(w) <= RADIUS            # the minimizer lies inside W\n    return float(((y[:T] - X[:T] @ w) ** 2).sum())\n\n\nT_GRID = np.unique(np.concatenate([np.geomspace(24, T_ALL, 40).astype(int), [T_ALL]]))\ncum_ogd, cum_ftl = np.cumsum(losses_ogd), np.cumsum(losses_ftl)\nregret_ogd = np.array([cum_ogd[T - 1] - hindsight_loss(T) for T in T_GRID])\nregret_ftl = np.array([cum_ftl[T - 1] - hindsight_loss(T) for T in T_GRID])\nbound = D * G * np.sqrt(T_GRID)\nlosses_persist = (y - X[:, 1]) ** 2               # forecast = the last reading\nprint(f\"  after T = {T_ALL} rounds: regret of online GD {regret_ogd[-1]:.0f} \"\n      f\"(bound D G sqrt(T) = {bound[-1]:.0f}), of follow-the-leader {regret_ftl[-1]:.0f}; \"\n      f\"regret per round of online GD at T = 24, 168, 8784: \"\n      f\"{regret_ogd[0] / T_GRID[0]:.2f}, \"\n      f\"{regret_ogd[np.searchsorted(T_GRID, 168)] / T_GRID[np.searchsorted(T_GRID, 168)]:.2f}, \"\n      f\"{regret_ogd[-1] / T_ALL:.2f}; persistence forecast: mean squared error \"\n      f\"{losses_persist.mean():.2f}\")\ncheck(\"[B-regret] the regret of online GD stays below D G sqrt(T) for every T\",\n      np.all(regret_ogd <= bound))\ncheck(\"[B-regret] the regret per round of online GD after all rounds is below a \"\n      \"quarter of its value after the first day\",\n      regret_ogd[-1] / T_ALL < 0.25 * regret_ogd[0] / T_GRID[0])\ncheck(\"[B-regret] follow-the-leader has a smaller regret than online GD after all rounds\",\n      regret_ftl[-1] < regret_ogd[-1])\ncheck(\"[B-regret] both online methods beat the persistence forecast on the accumulated loss\",\n      cum_ogd[-1] < losses_persist.sum() and cum_ftl[-1] < losses_persist.sum())"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-bandit]** Five fixed forecasters as actions; Hedge with full information and Exp3 with bandit feedback through the same rounds; regret against the best single forecaster in hindsight and the two bounds; check both bounds, that Exp3's regret exceeds Hedge's, and that both regrets per round vanish."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "A_ACTIONS = 5\nSCALE_ERR = 10.0                                  # absolute error / 10, clipped to [0, 1]\nlast = X[:, 1]                                    # the reading of the previous hour\nday_ago = np.concatenate([last[:24], y[:-24]])    # the reading 24 hours earlier\nmean3 = X[:, 1:4].mean(axis=1)                    # mean of the last three readings\nextrap = 2.0 * X[:, 1] - X[:, 2]                  # linear extrapolation of the last two\nrunning = np.concatenate([[last[0]], np.cumsum(y)[:-1] / np.arange(1, T_ALL)])\nforecasters = np.column_stack([last, day_ago, mean3, extrap, running])\nL = np.minimum(np.abs(y[:, None] - forecasters) / SCALE_ERR, 1.0)   # loss vectors in [0,1]^A\nbest_action = int(L.sum(axis=0).argmin())\nprint(f\"  {A_ACTIONS} forecasters; accumulated losses {np.round(L.sum(axis=0), 1).tolist()}, \"\n      f\"best single forecaster in hindsight: action {best_action + 1}\")\n\n\ndef run_hedge(L, lrate):\n    \"\"\"Full information: the pmf is updated from the whole loss vector.\"\"\"\n    w = np.full(A_ACTIONS, 1.0 / A_ACTIONS)\n    expected = np.empty(len(L))\n    for t in range(len(L)):\n        expected[t] = w @ L[t]                    # f^(t)(w^(t)) = <l^(t), w^(t)>\n        w = w * np.exp(-lrate * L[t])\n        w /= w.sum()\n    return expected\n\n\ndef run_exp3(L, gamma, seed=0):\n    \"\"\"Bandit feedback: only the loss of the drawn action is observed.  The\n    pmf mixes the exponential weights of the accumulated loss estimates\n    with the uniform pmf (weight gamma), so every action keeps probability\n    at least gamma / A and the estimates stay bounded (Auer et al. 2002).\"\"\"\n    rng = np.random.default_rng(seed)\n    est_total = np.zeros(A_ACTIONS)               # accumulated estimated losses\n    expected = np.empty(len(L))\n    for t in range(len(L)):\n        weights = np.exp(-(gamma / A_ACTIONS) * (est_total - est_total.min()))\n        w = (1.0 - gamma) * weights / weights.sum() + gamma / A_ACTIONS\n        expected[t] = w @ L[t]                    # f^(t)(w^(t)) = <l^(t), w^(t)>\n        a = rng.choice(A_ACTIONS, p=w)            # the forecaster used this hour\n        est_total[a] += L[t, a] / w[a]            # unbiased estimate of l^(t)_a\n    return expected\n\n\neta_hedge = np.sqrt(2.0 * np.log(A_ACTIONS) / T_ALL)\ngamma_exp3 = min(1.0, np.sqrt(A_ACTIONS * np.log(A_ACTIONS) / ((np.e - 1.0) * T_ALL)))\nexp_hedge, exp_exp3 = run_hedge(L, eta_hedge), run_exp3(L, gamma_exp3)\ncum_hedge, cum_exp3, cum_best = np.cumsum(exp_hedge), np.cumsum(exp_exp3), np.cumsum(L, axis=0)\nregret_hedge = np.array([cum_hedge[T - 1] - cum_best[T - 1].min() for T in T_GRID])\nregret_exp3 = np.array([cum_exp3[T - 1] - cum_best[T - 1].min() for T in T_GRID])\nbound_hedge = np.sqrt(2.0 * T_GRID * np.log(A_ACTIONS))\nbound_exp3 = 2.0 * np.sqrt((np.e - 1.0) * T_GRID * A_ACTIONS * np.log(A_ACTIONS))\nprint(f\"  after T = {T_ALL} rounds: regret of Hedge {regret_hedge[-1]:.1f} \"\n      f\"(bound sqrt(2 T log A) = {bound_hedge[-1]:.1f}), of Exp3 {regret_exp3[-1]:.1f} \"\n      f\"(bound 2 sqrt((e-1) T A log A) = {bound_exp3[-1]:.1f}); regret per round of \"\n      f\"Exp3 at T = 168 and {T_ALL}: {regret_exp3[int(np.searchsorted(T_GRID, 168))] / 168:.3f}, \"\n      f\"{regret_exp3[-1] / T_ALL:.3f}\")\ncheck(\"[B-bandit] with full information, Hedge's regret stays below sqrt(2 T log A) for every T\",\n      np.all(regret_hedge <= bound_hedge))\ncheck(\"[B-bandit] with bandit feedback, Exp3's regret stays below 2 sqrt((e-1) T A log A) for every T\",\n      np.all(regret_exp3 <= bound_exp3))\ncheck(\"[B-bandit] bandit feedback costs regret: Exp3's regret after all rounds exceeds Hedge's\",\n      regret_exp3[-1] > regret_hedge[-1])\ni_week = int(np.searchsorted(T_GRID, 168))\ncheck(\"[B-bandit] the regret per round of Hedge and Exp3 after all rounds is below half \"\n      \"of its value after the first week\",\n      regret_hedge[-1] / T_ALL < 0.5 * regret_hedge[i_week] / T_GRID[i_week]\n      and regret_exp3[-1] / T_ALL < 0.5 * regret_exp3[i_week] / T_GRID[i_week])"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-figures]** Data for the entry's first two figures: three days of the week with an offline forecaster trained once on the first day and then fixed, against online GD started from zero at hour 0 with a constant learning rate; and the loss functions of five consecutive hours for a scalar forecaster, with the online GD iterates on them."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "WEEK = slice(24 * 200, 24 * 207)                  # a week in July\nDAYS3 = slice(WEEK.start, WEEK.start + 72)        # its first three days\nX3, y3 = X[DAYS3], y[DAYS3]\nw_offline = np.linalg.lstsq(X3[:24], y3[:24], rcond=None)[0]   # trained once on day 1\noffline = X3 @ w_offline                          # then fixed for days 2 and 3\n# online GD started from w = 0 at hour 0 of the window, constant learning rate:\n# it forecasts from the first hour, badly at first, and adapts within the day\nscratch, _, _ = run_online_gd(X3, y3, lambda t: 2e-4)\nwith open(OUT_DIR / \"onlinelearning_offline.csv\", \"w\") as f:\n    f.write(\"hour,reading,offline,online\\n\")\n    for h in range(72):\n        off = f\"{offline[h]:.2f}\" if h >= 24 else \"nan\"\n        f.write(f\"{h},{y3[h]:.2f},{off},{scratch[h]:.2f}\\n\")\nerr_offline = np.mean((y3[24:] - offline[24:]) ** 2)\nerr_online = np.mean((y3[24:] - scratch[24:]) ** 2)\nerr_first6 = np.mean((y3[:6] - scratch[:6]) ** 2)\nprint(f\"  three days of the week: online GD from scratch has mean squared error \"\n      f\"{err_first6:.0f} over the first six hours and {err_online:.2f} over days 2-3; \"\n      f\"the forecaster trained once on day 1 has {err_offline:.2f} over days 2-3\")\ncheck(\"[B-figures] online GD from scratch forecasts badly in the first six hours \"\n      \"and within a factor three of the once-trained forecaster on days 2-3\",\n      err_first6 > 10 * err_online and err_online < 3 * err_offline)\n\n# the scalar forecaster w of the onlineGD figure: loss (y^(t) - w)^2, update\n# w^(t+1) = w^(t) - 0.2 * 2 (w^(t) - y^(t)) = 0.6 w^(t) + 0.4 y^(t)\ny5 = y[WEEK][:5]\nw_scalar = [y5[0]]                                # start at the first reading\nfor t in range(4):\n    w_scalar.append(0.6 * w_scalar[-1] + 0.4 * y5[t])\nw_grid = np.linspace(y5.min() - 4.0, y5.max() + 4.0, 81)\nwith open(OUT_DIR / \"onlinelearning_lossfuncs.csv\", \"w\") as f:\n    f.write(\"w,\" + \",\".join(f\"f{t + 1}\" for t in range(5)) + \"\\n\")\n    for wv in w_grid:\n        f.write(f\"{wv:.3f},\" + \",\".join(f\"{(y5[t] - wv) ** 2:.3f}\" for t in range(5)) + \"\\n\")\nwith open(OUT_DIR / \"onlinelearning_lossvalues.csv\", \"w\") as f:\n    f.write(\"t,y,w,f\\n\")\n    for t in range(5):\n        f.write(f\"{t + 1},{y5[t]:.2f},{w_scalar[t]:.2f},{(y5[t] - w_scalar[t]) ** 2:.3f}\\n\")\ncheck(\"[B-figures] each loss function of an hour is zero exactly at that hour's reading\",\n      all(abs((y5[t] - y5[t]) ** 2) == 0.0 and (y5[t] - w_scalar[t]) ** 2 >= 0.0 for t in range(5)))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-plot]** Write the regret per round against T for all four methods and a week of forecasts, plus the preview."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "with open(OUT_DIR / \"onlinelearning_bandit.csv\", \"w\") as f:\n    f.write(\"T,hedge,exp3,hedgebound,exp3bound\\n\")\n    for T, a, b, c, d in zip(T_GRID, regret_hedge, regret_exp3, bound_hedge, bound_exp3):\n        f.write(f\"{T},{max(a, 1e-6) / T:.6f},{max(b, 1e-6) / T:.6f},{c / T:.6f},{d / T:.6f}\\n\")\nwith open(OUT_DIR / \"onlinelearning_regret.csv\", \"w\") as f:\n    f.write(\"T,ogd,leader,bound\\n\")\n    for T, a, b, c in zip(T_GRID, regret_ogd, regret_ftl, bound):\n        f.write(f\"{T},{a / T:.4f},{b / T:.4f},{c / T:.4f}\\n\")\nwith open(OUT_DIR / \"onlinelearning_week.csv\", \"w\") as f:\n    f.write(\"hour,reading,ogd,leader\\n\")\n    for h, (r, a, b) in enumerate(zip(y[WEEK], forecasts_ogd[WEEK], forecasts_ftl[WEEK])):\n        f.write(f\"{h},{r:.2f},{a:.2f},{b:.2f}\\n\")\n\nfig, axes = plt.subplots(1, 3, figsize=(15, 3.6))\nax = axes[0]\nhours = np.arange(WEEK.stop - WEEK.start)\nax.plot(hours, y[WEEK], \"-\", color=\"black\", label=\"reading $y^{(t)}$\")\nax.plot(hours, forecasts_ogd[WEEK], \"--\", color=\"0.5\", label=\"online GD forecast\")\nax.plot(hours, forecasts_ftl[WEEK], \":\", color=\"black\", label=\"follow-the-leader forecast\")\nax.set_xlabel(\"hour $t$ of the week\")\nax.set_ylabel(\"temperature (\u00b0C)\")\nax.set_title(\"One week of the stream and the forecasts issued\")\nax.legend(frameon=False, fontsize=8)\nax = axes[1]\nax.plot(T_GRID, regret_ogd / T_GRID, \"o-\", color=\"black\", markersize=3, label=\"online GD\")\nax.plot(T_GRID, regret_ftl / T_GRID, \"s--\", color=\"0.5\", markersize=3, label=\"follow-the-leader\")\nax.plot(T_GRID, bound / T_GRID, \":\", color=\"black\", label=\"bound $DG/\\\\sqrt{T}$\")\nax.set_xscale(\"log\")\nax.set_yscale(\"log\")\nax.set_xlabel(\"number of rounds $T$\")\nax.set_ylabel(\"regret per round\")\nax.set_title(\"Regret per round, full information\")\nax.legend(frameon=False, fontsize=8)\nax = axes[2]\nax.plot(T_GRID, np.maximum(regret_hedge, 1e-6) / T_GRID, \"o-\", color=\"black\", markersize=3,\n        label=\"Hedge (full information)\")\nax.plot(T_GRID, np.maximum(regret_exp3, 1e-6) / T_GRID, \"s--\", color=\"0.5\", markersize=3,\n        label=\"Exp3 (bandit feedback)\")\nax.plot(T_GRID, bound_hedge / T_GRID, \":\", color=\"black\", label=\"bound, full information\")\nax.plot(T_GRID, bound_exp3 / T_GRID, \"-.\", color=\"0.5\", label=\"bound, bandit feedback\")\nax.set_xscale(\"log\")\nax.set_yscale(\"log\")\nax.set_xlabel(\"number of rounds $T$\")\nax.set_ylabel(\"regret per round\")\nax.set_title(\"Regret per round, five forecasters as actions\")\nax.legend(frameon=False, fontsize=8)\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"onlinelearning.png\", dpi=150)\nplt.close(fig)\ncheck(\"[B-plot] the CSV files were written\",\n      all((OUT_DIR / f).exists() for f in\n          (\"onlinelearning_regret.csv\", \"onlinelearning_week.csv\", \"onlinelearning_bandit.csv\",\n           \"onlinelearning_offline.csv\", \"onlinelearning_lossfuncs.csv\",\n           \"onlinelearning_lossvalues.csv\")))\n\npassed = sum(1 for _, ok in report if ok)\nprint(f\"\\n{passed}/{len(report)} checks pass\")"
  }
 ]
}