Dictionary of Applied Machine Learning · policy (reinforcement learning)
Numerical companion to the entry policy (reinforcement learning): it recomputes what the entry states and prints one line per check
Shows what a policy is and what it is for: a function from the state to the action that reacts to the state it meets, as against a plan (a fixed sequence of actions); a stochastic policy that draws the action from a probability distribution given the state; and the value function of a policy, the expected cumulative reward obtained by following it from a state, which is what the policy is learned to maximize. Self-contained (numpy and matplotlib only), fixed seed.
Run it without installing anything:uv run https://dictionaryofml.org/terms/policy.py
uv downloads this script and the pinned NumPy and Matplotlib it needs, then runs it; the script fetches any input file it uses. To keep the output files, download policy.py into a folder and run uv run policy.py there. With NumPy and Matplotlib already installed, python3 policy.py, from any directory — it writes its output files into the current directory. Fixed seeds, so the printed numbers reproduce exactly. Download policy.py · Notebook · Open in Colab
One cell per block of the script: the code, and what that code printed when it last ran here
"""
policy.py — numerical companion to the glossary entry 'policy'.
Purpose
-------
Shows what a policy is and what it is for: a function from the state to
the action that reacts to the state it meets, as against a plan (a fixed
sequence of actions); a stochastic policy that draws the action from a
probability distribution given the state; and the value function of a
policy, the expected cumulative reward obtained by following it from a
state, which is what the policy is learned to maximize. Self-contained
(numpy and matplotlib only), fixed seed.
Setup
-----
A thermostat controls a room. The state is the room temperature in
degrees C (rounded to one decimal), the action is heating on or off.
Each time step the temperature moves by +0.6 degrees with heating on and
by -0.4 degrees with heating off, plus a disturbance from outside that
is drawn at random (standard normal times 0.3) and, for 10 of the 60
time steps of an episode, an open window that pulls the temperature
down by 0.8 degrees per step. The reward per time step is
-|temperature - 21| - 0.1 if heating is on, i.e. the thermostat is
rewarded for staying near 21 degrees at little heating.
Blocks
------
[B-plan] A threshold policy (heat if the temperature is below 21) is
compared with a plan, the fixed on/off sequence that the
policy itself produced in an undisturbed episode, replayed
under the disturbances: the policy keeps the temperature
within 1 degree of 21 on average, the plan does not.
[B-stoch] A stochastic policy, P(heating on | temperature) =
1 / (1 + exp(k (temperature - 21))), draws the action from
that distribution; its expected cumulative reward over an
episode, estimated from 200 episodes, grows with k and
comes within 6 of that of the threshold policy at k = 8.
[B-value] The value function of the threshold policy, the expected
cumulative reward from each starting temperature, is
estimated from 200 episodes per state; it is largest near
21 degrees and falls off on both sides. Among the
thresholds 18, 19, ..., 24 the one maximizing the value at
the starting state 19 degrees is 21.
Outputs
-------
policy_trace.csv : time step, temperature under the threshold policy and
under the plan in one disturbed episode.
policy_value.csv : starting temperature, value of the threshold policy
with threshold 21.
policy.png : matplotlib preview of the two panels (checking only).
"""
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from pathlib import Path
OUT_DIR = Path(__file__).parent
report = []
def check(name, ok):
report.append((name, bool(ok)))
print(f" [{'ok' if ok else 'FAIL'}] {name}")
rng = np.random.default_rng(0)
TARGET, STEPS, N_EP = 21.0, 60, 200
HEAT_ON, HEAT_OFF, WINDOW = 0.6, -0.4, -0.8
def disturbance(t):
d = 0.3 * rng.standard_normal()
return d + (WINDOW if 20 <= t < 30 else 0.0)
def reward(temp, on):
return -abs(temp - TARGET) - (0.1 if on else 0.0)
def episode(choose, start, disturbed=True):
"""Run one episode; choose(temp, t) -> action (True = heating on).
Returns the cumulative reward, the temperatures and the actions."""
temp = start; ret = 0.0; temps = [temp]; acts = []
for t in range(STEPS):
on = choose(round(temp, 1), t)
temp = temp + (HEAT_ON if on else HEAT_OFF) + (disturbance(t) if disturbed else 0.0)
ret += reward(temp, on); temps.append(temp); acts.append(on)
return ret, np.array(temps), acts
def threshold_policy(th):
return lambda temp, t: temp < th
A threshold policy (heat if the temperature is below 21) is compared with a plan, the fixed on/off sequence that the policy itself produced in an undisturbed episode, replayed under the disturbances: the policy keeps the temperature within 1 degree of 21 on average, the plan does not.
_, _, plan_actions = episode(threshold_policy(TARGET), 19.0, disturbed=False)
plan = lambda temp, t: plan_actions[t] # the fixed sequence
rng = np.random.default_rng(1)
_, temps_policy, _ = episode(threshold_policy(TARGET), 19.0)
rng = np.random.default_rng(1)
_, temps_plan, _ = episode(plan, 19.0)
dev_policy = float(np.abs(temps_policy[20:] - TARGET).mean())
dev_plan = float(np.abs(temps_plan[20:] - TARGET).mean())
check(f"[B-plan] mean distance from 21 degrees after the window opens: "
f"policy {dev_policy:.2f}, plan {dev_plan:.2f}",
dev_policy < 1.0 and dev_plan > 2.0 * dev_policy)
[ok] [B-plan] mean distance from 21 degrees after the window opens: policy 0.93, plan 7.26
A stochastic policy, P(heating on | temperature) = 1 / (1 + exp(k (temperature - 21))), draws the action from that distribution; its expected cumulative reward over an episode, estimated from 200 episodes, grows with k and comes within 6 of that of the threshold policy at k = 8.
rng = np.random.default_rng(2)
def stochastic_policy(k):
return lambda temp, t: rng.random() < 1.0 / (1.0 + np.exp(k * (temp - TARGET)))
def expected_return(choose, start):
return float(np.mean([episode(choose, start)[0] for _ in range(N_EP)]))
ks = [0.5, 1.0, 2.0, 4.0, 8.0]
val_stoch = [expected_return(stochastic_policy(k), 19.0) for k in ks]
val_thr = expected_return(threshold_policy(TARGET), 19.0)
check("[B-stoch] expected cumulative reward from 19 degrees: "
+ ", ".join(f"k={k}: {v:.1f}" for k, v in zip(ks, val_stoch))
+ f"; threshold policy {val_thr:.1f}",
all(a < b for a, b in zip(val_stoch, val_stoch[1:]))
and val_stoch[-1] > val_thr - 6.0)
[ok] [B-stoch] expected cumulative reward from 19 degrees: k=0.5: -96.8, k=1.0: -72.6, k=2.0: -59.4, k=4.0: -50.2, k=8.0: -46.4; threshold policy -42.8
The value function of the threshold policy, the expected cumulative reward from each starting temperature, is estimated from 200 episodes per state; it is largest near 21 degrees and falls off on both sides. Among the thresholds 18, 19, ..., 24 the one maximizing the value at the starting state 19 degrees is 21.
starts = np.arange(16.0, 26.5, 0.5)
value = np.array([expected_return(threshold_policy(TARGET), s) for s in starts])
thresholds = np.arange(18.0, 25.0, 1.0)
val_by_th = [expected_return(threshold_policy(th), 19.0) for th in thresholds]
best_th = float(thresholds[int(np.argmax(val_by_th))])
check(f"[B-value] value from 16..26 degrees: max {value.max():.1f} at "
f"{starts[int(np.argmax(value))]:.1f} degrees, min {value.min():.1f}; "
f"best threshold from 19 degrees: {best_th:.0f}",
abs(starts[int(np.argmax(value))] - TARGET) <= 1.0 and best_th == TARGET)
# ---------------------------------------------------------------- CSV
with open(OUT_DIR / "policy_trace.csv", "w") as fh:
fh.write("t,temp_policy,temp_plan\n")
for t in range(STEPS + 1):
fh.write(f"{t},{temps_policy[t]:.3f},{temps_plan[t]:.3f}\n")
with open(OUT_DIR / "policy_value.csv", "w") as fh:
fh.write("start,value\n")
for s, v in zip(starts, value):
fh.write(f"{s:.1f},{v:.3f}\n")
# -------------------------------------------------------------- preview
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(9.4, 3.8))
ax1.plot(range(STEPS + 1), temps_policy, "k-", lw=1.4, label="threshold policy (reacts to the state)")
ax1.plot(range(STEPS + 1), temps_plan, "k--", lw=1.2, label="plan (fixed action sequence)")
ax1.axhline(TARGET, color="0.6", lw=0.8, ls=":")
ax1.axvspan(20, 30, color="0.9", lw=0, label="window open")
ax1.set_xlabel("time step")
ax1.set_ylabel("room temperature (degrees C)")
ax1.set_title("policy versus plan under a disturbance")
ax1.legend(frameon=False, fontsize=7)
ax2.plot(starts, value, "k.-", lw=1.2, ms=5, label="threshold policy, threshold 21")
ax2.set_xlabel("starting temperature (state)")
ax2.set_ylabel("expected cumulative reward")
ax2.set_title("value function of the policy")
ax2.legend(frameon=False, fontsize=8)
fig.tight_layout()
fig.savefig(OUT_DIR / "policy.png", dpi=110)
n_ok = sum(ok for _, ok in report)
print(f"\n{n_ok}/{len(report)} checks pass")
print(f"wrote {OUT_DIR / 'policy_trace.csv'}, {OUT_DIR / 'policy_value.csv'}, "
f"{OUT_DIR / 'policy.png'}")
if n_ok != len(report):
raise SystemExit(1)
[ok] [B-value] value from 16..26 degrees: max -41.1 at 22.0 degrees, min -67.3; best threshold from 19 degrees: 21 3/3 checks pass wrote /Users/junga1/dictionaryappliedml/pythondemos/policy_trace.csv, /Users/junga1/dictionaryappliedml/pythondemos/policy_value.csv, /Users/junga1/dictionaryappliedml/pythondemos/policy.png

B-value writes when the script runs