Dictionary of Applied Machine Learning · learning rate

learning rate — Python demo

Numerical companion to the entry learning rate: it recomputes what the entry states and prints one line per check

Shows what the learning rate of a gradient step does to the HYPOTHESIS MAP and to the set of model parameters that a fixed number of gradient steps can reach. One gradient step of GD on the average squared error loss of a linear model moves the fitted line by an amount proportional to the learning rate; a learning rate above 2/L makes the loss grow; and T steps with learning rate eta stay inside a ball around the initial parameters whose radius grows with eta. Self-contained (numpy and matplotlib only), fixed seed.

Run it without installing anything:
uv run https://dictionaryofml.org/terms/learnrate.py
uv downloads this script and the pinned NumPy and Matplotlib it needs, then runs it; the script fetches any input file it uses. To keep the output files, download learnrate.py into a folder and run uv run learnrate.py there. With NumPy and Matplotlib already installed, python3 learnrate.py, from any directory — it writes its output files into the current directory. Fixed seeds, so the printed numbers reproduce exactly. Download learnrate.py · Notebook · Open in Colab

The script, block by block

One cell per block of the script: the code, and what that code printed when it last ran here

setup

"""
learnrate.py — numerical companion to the glossary entry 'learning rate'.

Purpose
-------
Shows what the learning rate of a gradient step does to the HYPOTHESIS
MAP and to the set of model parameters that a fixed number of gradient
steps can reach.  One gradient step of GD on the average squared error
loss of a linear model moves the fitted line by an amount proportional
to the learning rate; a learning rate above 2/L makes the loss grow; and
T steps with learning rate eta stay inside a ball around the initial
parameters whose radius grows with eta.  Self-contained (numpy and
matplotlib only), fixed seed.

Setup
-----
Training set: m = 4 consecutive days with feature x = morning minimum
temperature (degrees C) and label y = maximum daytime temperature
(degrees C); synthetic values drawn around y = 6.7 + 0.76 x.  Hypothesis
map h(x) = w1 x + w2, i.e. a linear model with the feature vector (x, 1).
Objective f(w) = average squared error loss on the training set, whose
Hessian is (2/m) X^T X with largest eigenvalue L.  Initial parameters
w = (0.3, 2.0), a poor line below the data.

Blocks
------
[B-data]  The training set, the initial hypothesis map and f at the
          initial parameters; L is reported.
[B-step]  One gradient step for three learning rates, 0.1/L, 1.5/L and
          2.5/L.  The change of the prediction at any x is eta times the
          product of the gradient with (x, 1), so the largest change of
          the line over the plotted range is exactly 15 times larger for
          1.5/L than for 0.1/L; the first two steps lower f, the third
          (above the threshold 2/L) raises it.
[B-reach] T = 10 gradient steps for the learning rates 0.1/L, 0.5/L and
          1.5/L.  Every iterate stays within the distance
          eta * sum_t ||grad f(w^(t))|| of the initial parameters, a
          radius that grows with eta; with 0.1/L the ERM solution lies
          outside that ball, with 1.5/L inside it.

Outputs
-------
learnrate_data.csv  : x, y of the 4 training days.
learnrate_lines.csv : x grid with h_before (initial line), h_small,
                      h_large, h_toolarge (after one step with 0.1/L,
                      1.5/L, 2.5/L).
learnrate_balls.csv : 91 points on each of the three circles of [B-reach]
                      (columns xs,ys / xm,ym / xl,yl) in the (w1, w2)
                      plane.
learnrate_reach.csv : one row per learning rate of [B-reach]: eta_over_L,
                      radius, w1_T, w2_T, plus the initial parameters and
                      the ERM solution as rows tagged in the first column.
learnrate.png       : matplotlib preview of the two figures (checking
                      only).
"""

import numpy as np
import matplotlib

matplotlib.use("Agg")
import matplotlib.pyplot as plt

from pathlib import Path

OUT_DIR = Path(__file__).parent

report = []


def check(name, ok):
    report.append((name, bool(ok)))
    print(f"  [{'ok' if ok else 'FAIL'}] {name}")


rng = np.random.default_rng(0)

B-data

The training set, the initial hypothesis map and f at the initial parameters; L is reported.

m = 4
x = np.array([2.1, 3.4, 4.0, 5.6])              # four consecutive days
y = 6.7 + 0.76 * x + 0.6 * rng.standard_normal(m)
X = np.c_[x, np.ones(m)]                        # feature vectors (x, 1)


def h(w, xs):
    return w[0] * xs + w[1]


def f(w):
    return float(np.mean((y - X @ w) ** 2))


def grad(w):
    return -(2.0 / m) * X.T @ (y - X @ w)


w0 = np.array([0.3, 2.0])
H = (2.0 / m) * X.T @ X
L = float(np.linalg.eigvalsh(H).max())
w_erm = np.linalg.lstsq(X, y, rcond=None)[0]
check(f"[B-data]  m={m} days, f(w0)={f(w0):.2f}, L={L:.1f}, "
      f"ERM solution w=({w_erm[0]:.2f}, {w_erm[1]:.2f})",
      f(w0) > 10.0 and L > 0)
  [ok] [B-data]  m=4 days, f(w0)=43.26, L=33.5, ERM solution w=(0.78, 6.75)

B-step

One gradient step for three learning rates, 0.1/L, 1.5/L and 2.5/L. The change of the prediction at any x is eta times the product of the gradient with (x, 1), so the largest change of the line over the plotted range is exactly 15 times larger for 1.5/L than for 0.1/L; the first two steps lower f, the third (above the threshold 2/L) raises it.

etas = {"small": 0.1 / L, "large": 1.5 / L, "toolarge": 2.5 / L}
g = grad(w0)
w_after = {k: w0 - e * g for k, e in etas.items()}
xg = np.linspace(0.0, 8.0, 81)
d = {k: np.abs(h(w_after[k], xg) - h(w0, xg)).max() for k in etas}
check(f"[B-step]  largest change of the line: {d['small']:.2f} (0.1/L) vs "
      f"{d['large']:.2f} (1.5/L), ratio {d['large'] / d['small']:.1f}",
      abs(d["large"] / d["small"] - 15.0) < 1e-6)
check(f"[B-step]  f: {f(w0):.2f} -> {f(w_after['small']):.2f} (0.1/L), "
      f"{f(w_after['large']):.2f} (1.5/L), {f(w_after['toolarge']):.2f} "
      f"(2.5/L, above 2/L: f grows)",
      f(w_after["small"]) < f(w0) and f(w_after["large"]) < f(w0)
      and f(w_after["toolarge"]) > f(w0))
  [ok] [B-step]  largest change of the line: 1.26 (0.1/L) vs 18.85 (1.5/L), ratio 15.0
  [ok] [B-step]  f: 43.26 -> 35.40 (0.1/L), 12.24 (1.5/L), 94.84 (2.5/L, above 2/L: f grows)

B-reach

T = 10 gradient steps for the learning rates 0.1/L, 0.5/L and 1.5/L. Every iterate stays within the distance eta * sum_t ||grad f(w^(t))|| of the initial parameters, a radius that grows with eta; with 0.1/L the ERM solution lies outside that ball, with 1.5/L inside it.

T = 10
reach = {}
for tag, c in (("s", 0.1), ("m", 0.5), ("l", 1.5)):
    eta = c / L
    w = w0.copy(); path = 0.0
    for _ in range(T):
        gt = grad(w); path += eta * np.linalg.norm(gt); w = w - eta * gt
    reach[tag] = (c, path, w)
dist_erm = float(np.linalg.norm(w_erm - w0))
check(f"[B-reach] T={T}: radius {reach['s'][1]:.2f} (0.1/L), "
      f"{reach['m'][1]:.2f} (0.5/L), {reach['l'][1]:.2f} (1.5/L); "
      f"||w_ERM - w0|| = {dist_erm:.2f}",
      all(np.linalg.norm(reach[t][2] - w0) <= reach[t][1] + 1e-9
          for t in reach)
      and reach["s"][1] < dist_erm < reach["l"][1])

# ---------------------------------------------------------------- CSV
with open(OUT_DIR / "learnrate_data.csv", "w") as fh:
    fh.write("x,y\n")
    for a, b in zip(x, y):
        fh.write(f"{a:.3f},{b:.3f}\n")
with open(OUT_DIR / "learnrate_lines.csv", "w") as fh:
    fh.write("x,h_before,h_small,h_large,h_toolarge\n")
    for i, a in enumerate(xg):
        fh.write(f"{a:.3f},{h(w0, a):.3f},{h(w_after['small'], a):.3f},"
                 f"{h(w_after['large'], a):.3f},{h(w_after['toolarge'], a):.3f}\n")
ang = np.linspace(0.0, 2.0 * np.pi, 91)
with open(OUT_DIR / "learnrate_balls.csv", "w") as fh:
    fh.write("xs,ys,xm,ym,xl,yl\n")
    for a in ang:
        row = []
        for t in ("s", "m", "l"):
            r = reach[t][1]
            row += [f"{w0[0] + r * np.cos(a):.4f}", f"{w0[1] + r * np.sin(a):.4f}"]
        fh.write(",".join(row) + "\n")
with open(OUT_DIR / "learnrate_reach.csv", "w") as fh:
    fh.write("tag,eta_over_L,radius,w1,w2\n")
    fh.write(f"init,0,0,{w0[0]:.4f},{w0[1]:.4f}\n")
    fh.write(f"erm,0,0,{w_erm[0]:.4f},{w_erm[1]:.4f}\n")
    for t in ("s", "m", "l"):
        c, r, w = reach[t]
        fh.write(f"{t},{c},{r:.4f},{w[0]:.4f},{w[1]:.4f}\n")

# -------------------------------------------------------------- preview
fig, (ax, ax2) = plt.subplots(1, 2, figsize=(9.6, 4.0))
ax.plot(x, y, "ko", ms=5, label="training set (4 days)")
ax.plot(xg, h(w0, xg), "k-", lw=1.4, label="before the step")
ax.plot(xg, h(w_after["small"], xg), "k--", lw=1.2, label="after, learning rate 0.1/L")
ax.plot(xg, h(w_after["large"], xg), "k:", lw=1.6, label="after, learning rate 1.5/L")
ax.plot(xg, h(w_after["toolarge"], xg), "-.", color="0.5", lw=1.4,
        label="after, learning rate 2.5/L (loss grows)")
ax.set_xlabel("morning minimum temperature $x$")
ax.set_ylabel("maximum daytime temperature $y$")
ax.set_title("one gradient step, three learning rates")
ax.legend(frameon=False, fontsize=7)
for t, ls in (("s", "-"), ("m", "--"), ("l", ":")):
    c, r, w = reach[t]
    ax2.plot(w0[0] + r * np.cos(ang), w0[1] + r * np.sin(ang), "k", ls=ls, lw=1.1,
             label=f"reachable in {T} steps, learning rate {c}/L")
ax2.plot([w0[0]], [w0[1]], "ks", ms=5, label="initial parameters")
ax2.plot([w_erm[0]], [w_erm[1]], "k*", ms=9, label="ERM solution")
ax2.set_xlabel("slope $w_1$")
ax2.set_ylabel("intercept $w_2$")
ax2.set_title("parameters reachable in 10 steps")
ax2.legend(frameon=False, fontsize=7)
fig.tight_layout()
fig.savefig(OUT_DIR / "learnrate.png", dpi=110)

n_ok = sum(ok for _, ok in report)
print(f"\n{n_ok}/{len(report)} checks pass")
print(f"wrote {OUT_DIR / 'learnrate_data.csv'}, {OUT_DIR / 'learnrate_lines.csv'}, "
      f"{OUT_DIR / 'learnrate_balls.csv'}, {OUT_DIR / 'learnrate_reach.csv'}, "
      f"{OUT_DIR / 'learnrate.png'}")
if n_ok != len(report):
    raise SystemExit(1)
  [ok] [B-reach] T=10: radius 1.02 (0.1/L), 1.61 (0.5/L), 4.82 (1.5/L); ||w_ERM - w0|| = 4.77

4/4 checks pass
wrote /Users/junga1/dictionaryappliedml/pythondemos/learnrate_data.csv, /Users/junga1/dictionaryappliedml/pythondemos/learnrate_lines.csv, /Users/junga1/dictionaryappliedml/pythondemos/learnrate_balls.csv, /Users/junga1/dictionaryappliedml/pythondemos/learnrate_reach.csv, /Users/junga1/dictionaryappliedml/pythondemos/learnrate.png
Preview figure produced by learnrate.py
The preview figure the block B-reach writes when the script runs