Dictionary of Applied Machine Learning · pretraining

pretraining — Python demo

Numerical companion to the entry pretraining: it recomputes what the entry states and prints one line per check

Numerical companion to the glossary entry 'pretraining'. A 2.4 x 1.2 km piece of the aerial photograph of the Wachau (assets/ wachau_ortho_center.jpg: the Danube bend at Duernstein, 0.79 m per pixel; orthophoto (c) basemap.at, CC BY 4.0) is cut into 4608 square patches of 32 px (25 m), and the parcel mask assets/ wachau_labels_center.png (INVEKOS Schlaege 2025-1, AgrarMarkt Austria, CC BY 3.0 AT) labels a patch as vineyard when more than half of its pixels lie on vineyard parcels. The western half of the piece is the training set, the eastern half the test set. At this resolution the rows of vines are visible stripes, and the mean color of a patch tells little: vineyards, meadows and gardens are all green.

Run it with python3 pretraining.py, from any directory — it writes its output files into the current directory. Requires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Download pretraining.py · Notebook · Open in Colab

The script, block by block

One cell per block of the script: the code, and what that code printed when it last ran here

setup

"""Pretraining the first layers of a convolutional network on a pretext
task without labels, then reusing them as frozen features for a linear
classifier of vineyard patches.

Purpose
-------
Numerical companion to the glossary entry 'pretraining'.  A 2.4 x 1.2 km
piece of the aerial photograph of the Wachau (assets/
wachau_ortho_center.jpg: the Danube bend at Duernstein, 0.79 m per
pixel; orthophoto (c) basemap.at, CC BY 4.0) is cut into 4608 square
patches of 32 px (25 m), and the parcel mask assets/
wachau_labels_center.png (INVEKOS Schlaege 2025-1, AgrarMarkt Austria,
CC BY 3.0 AT) labels a patch as vineyard when more than half of its
pixels lie on vineyard parcels.  The western half of the piece is the
training set, the eastern half the test set.  At this resolution the
rows of vines are visible stripes, and the mean color of a patch tells
little: vineyards, meadows and gardens are all green.

Pretraining needs no labels.  Two pretext tasks are run on the training
patches, each with the same network: two convolutional layers (3 x 3
filters, 8 channels each, ReLU) and a linear readout.  The first blanks
the central 8 x 8 pixel block of a patch and predicts the mean color of
the blanked pixels from their surroundings.  The second rotates a patch
by 0, 90, 180 or 270 degrees and predicts the rotation, which is
answered by the direction of the rows and of the shadows.  Only the
photograph is used, never the parcel mask.

Transfer.  The two convolutional layers are then frozen and applied to
the original patches; their output, average-pooled to a 4 x 4 grid,
gives 128 features per patch.  A linear classifier (logistic regression
with a small ridge penalty) is trained on these features of 32 to 2304
labeled western patches and evaluated on the eastern patches.  The
comparison: the same linear classifier on the mean color of the patch
(3 features), on the raw pixels (3072 features), and on the features of
the same two layers with their random initial weights, i.e., without
pretraining; and the majority baseline.  The last comparison is the
control: the color-pretrained layers beat the untrained ones when labels
are few and are matched by them when all labels are used, while the
rotation-pretrained layers beat the untrained ones at every number of
labels.

Deterministic: fixed seeds for the initial weights and the patches of each step.
Self-contained: numpy + matplotlib only.

Blocks
------
[B-patches]  Cut the photograph into 4608 patches of 32 px (25 m), label
             them by the parcel mask, and split them by easting into a
             western training half and an eastern test half.
[B-pretrain] Blank the center of every training patch and train the two
             convolutional layers plus a linear readout to predict the
             center color (Adam, 512 patches per step).  Check that the
             loss falls and that the trained network beats the patch's
             mean color as a prediction of the center on the test half.
[B-rotation] The second pretext task: predict the rotation of a patch
             (Adam, 512 patches per step).  Check that the rotation is
             recognized well above chance on the test half.
[B-transfer] Freeze the pretrained layers, pool their output to 128
             features, and fit a linear classifier of vineyard vs. no
             vineyard on 32 to 2304 labeled patches of the training
             half; compare its test accuracy with the same classifier on
             mean color, on raw pixels and on the features of the
             untrained layers, and with the majority baseline.
[B-plot]     Write the CSV files and images for the entry's figures and
             the preview.

Outputs
-------
pretraining_loss.csv : iteration, train, test -- the color pretext loss
                       (mean squared error of the predicted center
                       color, in units of the pixel range [0, 1]) on the
                       two halves
pretraining_accuracy.csv : labels, meancolor, rawpixels, untrained,
                       color, rotation, baseline -- the test accuracy of
                       the linear classifier on each feature set against
                       the number of labeled patches (mean over up to
                       five spread-out subsets), and the majority
                       baseline
pretraining_patches.png : eight test patches: the blanked input, the
                       network's prediction of the center painted in,
                       and the original
pretraining.png : preview (checking only)
"""

import math
from pathlib import Path

import numpy as np
import matplotlib

matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib.image import imread

OUT_DIR = Path(__file__).parent

report = []                         # collects (check name, pass/fail) pairs


def check(name, ok):                # records and prints one verification
    report.append((name, bool(ok)))
    print(f"  [{'ok' if ok else 'FAIL'}] {name}")

B-patches

Cut the photograph into 4608 patches of 32 px (25 m), label them by the parcel mask, and split them by easting into a western training half and an eastern test half.

SIDE = 32                           # 32 px = 25 m on the ground
photo = imread(OUT_DIR.parent / "assets" / "wachau_ortho_center.jpg") / 255.0
labels = np.asarray(imread(OUT_DIR.parent / "assets" / "wachau_labels_center.png"))
if labels.ndim == 3:
    labels = labels[:, :, 0]
if labels.max() <= 1.0 and labels.dtype != np.uint8:
    labels = np.rint(labels * 255).astype(np.uint8)
rows, cols = photo.shape[0] // SIDE, photo.shape[1] // SIDE
patches = photo[:rows * SIDE, :cols * SIDE].reshape(rows, SIDE, cols, SIDE, 3)
patches = patches.transpose(0, 2, 1, 3, 4).reshape(rows * cols, SIDE, SIDE, 3).astype(np.float32)
share = (labels[:rows * SIDE, :cols * SIDE] == 1).reshape(
    rows, SIDE, cols, SIDE).mean(axis=(1, 3)).ravel()
y = (share > 0.5).astype(np.float32)                  # 1 = vineyard
grid_c = np.arange(rows * cols) % cols
train = grid_c < cols // 2                            # western half
test = ~train
mean_color = patches[train].mean(axis=(0, 1, 2))      # training statistics only
X_img = patches - mean_color                          # centered pixels
print(f"  {len(y)} patches of {rows} x {cols}, {SIDE} x {SIDE} px each; vineyard share "
      f"{y.mean():.3f} (west {y[train].mean():.3f}, east {y[test].mean():.3f})")
check("[B-patches] 4608 patches, split into halves of equal size by easting",
      len(y) == 4608 and train.sum() == test.sum() == 2304)
check("[B-patches] between a fifth and two fifths of the patches are vineyard",
      0.2 < y.mean() < 0.4)

B-pretrain

Blank the center of every training patch and train the two convolutional layers plus a linear readout to predict the center color (Adam, 512 patches per step). Check that the loss falls and that the trained network beats the patch's mean color as a prediction of the center on the test half.

rng = np.random.default_rng(0)
CH = 8                              # channels of both convolutional layers
C0, C1 = 12, 20                     # the blanked center block: rows and columns 12..19
ITERS, BATCH, LR, B1, B2 = 300, 512, 3e-3, 0.9, 0.999


def conv_cols(x):
    """3 x 3 neighborhoods of every pixel (zero padding): (N, H, W, 9 C)."""
    n, h, w, c = x.shape
    xp = np.pad(x, ((0, 0), (1, 1), (1, 1), (0, 0)))
    return np.concatenate([xp[:, i:i + h, j:j + w, :] for i in range(3) for j in range(3)],
                          axis=3)


def conv_cols_back(dcols, shape):
    """Adjoint of conv_cols: scatter the neighborhood gradients back to pixels."""
    n, h, w, c = shape
    dxp = np.zeros((n, h + 2, w + 2, c), dtype=np.float32)
    for k, (i, j) in enumerate((i, j) for i in range(3) for j in range(3)):
        dxp[:, i:i + h, j:j + w, :] += dcols[..., k * c:(k + 1) * c]
    return dxp[:, 1:-1, 1:-1, :]


def init_params(n_in, n_out):
    """He initialization of the two convolutional layers and a linear readout."""
    w1 = (rng.standard_normal((27, CH)) * math.sqrt(2.0 / 27)).astype(np.float32)
    w2 = (rng.standard_normal((9 * CH, CH)) * math.sqrt(2.0 / (9 * CH))).astype(np.float32)
    wh = (rng.standard_normal((n_in * CH, n_out)) * math.sqrt(1.0 / (n_in * CH))).astype(np.float32)
    return {"w1": w1, "b1": np.zeros(CH, np.float32), "w2": w2,
            "b2": np.zeros(CH, np.float32), "wh": wh, "bh": np.zeros(n_out, np.float32)}


def layers(p, x):
    """The two convolutional layers with ReLU: their output (N, 32, 32, CH)."""
    c1 = conv_cols(x)
    a1 = np.maximum(c1 @ p["w1"] + p["b1"], 0.0)
    c2 = conv_cols(a1)
    a2 = np.maximum(c2 @ p["w2"] + p["b2"], 0.0)
    return a2, (c1, a1, c2)


def layers_back(p, da2, saved, grads):
    """Backpropagate a gradient at the layers' output into their weights."""
    c1, a1, c2 = saved
    grads["w2"] = c2.reshape(-1, 9 * CH).T @ da2.reshape(-1, CH)
    grads["b2"] = da2.sum((0, 1, 2))
    da1 = conv_cols_back(da2 @ p["w2"].T, a1.shape)
    da1 *= a1 > 0
    grads["w1"] = c1.reshape(-1, 27).T @ da1.reshape(-1, CH)
    grads["b1"] = da1.sum((0, 1, 2))
    return grads


def adam_step(p, grads, moments, step):
    for k in p:
        m, v = moments[k]
        m[...] = B1 * m + (1 - B1) * grads[k]
        v[...] = B2 * v + (1 - B2) * grads[k] ** 2
        p[k] -= LR * (m / (1 - B1 ** step)) / (np.sqrt(v / (1 - B2 ** step)) + 1e-8)


def color_forward(p, x):
    a2, saved = layers(p, x)
    block = a2[:, C0:C1, C0:C1, :].reshape(len(x), -1)
    return block @ p["wh"] + p["bh"], (a2, block, saved)


def color_backward(p, target, out, cache):
    """Gradients of the mean squared error of the center color."""
    a2, block, saved = cache
    n = len(out)
    dout = 2.0 * (out - target) / (n * 3)
    grads = {"wh": block.T @ dout, "bh": dout.sum(0)}
    da2 = np.zeros_like(a2)
    da2[:, C0:C1, C0:C1, :] = (dout @ p["wh"].T).reshape(n, C1 - C0, C1 - C0, CH)
    da2 *= a2 > 0
    return layers_back(p, da2, saved, grads)


def blank(x):
    """The pretext input: the center block set to zero (the mean color)."""
    xb = x.copy()
    xb[:, C0:C1, C0:C1, :] = 0.0
    return xb


def in_slices(fn, x, size=512):
    return np.concatenate([fn(x[i:i + size]) for i in range(0, len(x), size)])


X_blank = blank(X_img)
target = X_img[:, C0:C1, C0:C1, :].mean(axis=(1, 2))   # the color to predict
params_color = init_params((C1 - C0) ** 2, 3)
params_random = {k: v.copy() for k, v in params_color.items()}   # the untrained control
moments = {k: (np.zeros_like(v), np.zeros_like(v)) for k, v in params_color.items()}
train_idx = np.flatnonzero(train)
loss_log = []
for step in range(1, ITERS + 1):
    sel = np.sort(rng.choice(train_idx, BATCH, replace=False))
    out, cache = color_forward(params_color, X_blank[sel])
    adam_step(params_color, color_backward(params_color, target[sel], out, cache), moments, step)
    if step == 1 or step % 20 == 0:
        err = [float(((in_slices(lambda z: color_forward(params_color, z)[0], X_blank[part])
                        - target[part]) ** 2).mean()) for part in (train, test)]
        loss_log.append((step, err[0], err[1]))
        if step == 1 or step % 100 == 0:
            print(f"  iteration {step:3d}: color pretext loss train {err[0]:.5f}, test {err[1]:.5f}")
mse_net = loss_log[-1][2]
ring_mean = X_blank[test].sum(axis=(1, 2)) / (SIDE * SIDE - (C1 - C0) ** 2)
mse_ring = float(((ring_mean - target[test]) ** 2).mean())
print(f"  center color on the test half: mean squared error of the network {mse_net:.5f}, "
      f"of the patch's mean color outside the blank {mse_ring:.5f}")
check("[B-pretrain] the color pretext loss on the training half falls to below a "
      "tenth of its initial value", loss_log[-1][1] < 0.1 * loss_log[0][1])
check("[B-pretrain] on the test half the network predicts the center color better "
      "than the patch's mean color outside the blank", mse_net < mse_ring)

B-rotation

The second pretext task: predict the rotation of a patch (Adam, 512 patches per step). Check that the rotation is recognized well above chance on the test half.

def pool4(a2):
    """Average-pool the layers' output to a 4 x 4 grid: 128 numbers per patch."""
    n = len(a2)
    return a2.reshape(n, 4, SIDE // 4, 4, SIDE // 4, CH).mean(axis=(2, 4)).reshape(n, -1)


def rotation_forward(p, x):
    a2, saved = layers(p, x)
    pooled = pool4(a2)
    return pooled @ p["wh"] + p["bh"], (a2, pooled, saved)


def rotation_backward(p, target, logits, cache):
    """Gradients of the cross-entropy of the four rotations."""
    a2, pooled, saved = cache
    n = len(logits)
    q = np.exp(logits - logits.max(1, keepdims=True))
    q /= q.sum(1, keepdims=True)
    dout = (q - np.eye(4, dtype=np.float32)[target]) / n
    grads = {"wh": pooled.T @ dout, "bh": dout.sum(0)}
    dpool = (dout @ p["wh"].T).reshape(n, 4, 1, 4, 1, CH) / (SIDE // 4) ** 2
    da2 = np.broadcast_to(dpool, (n, 4, SIDE // 4, 4, SIDE // 4, CH)).reshape(a2.shape).copy()
    da2 *= a2 > 0
    return layers_back(p, da2, saved, grads)


def rotations(x):
    """Every patch in the four rotations, with the rotation as label."""
    return (np.concatenate([np.rot90(x, k, axes=(1, 2)) for k in range(4)]),
            np.repeat(np.arange(4), len(x)))


X_rot, rot_label = rotations(X_img[train])
params_rot = init_params(16, 4)
moments = {k: (np.zeros_like(v), np.zeros_like(v)) for k, v in params_rot.items()}
for step in range(1, ITERS + 1):
    sel = np.sort(rng.choice(len(X_rot), BATCH, replace=False))
    logits, cache = rotation_forward(params_rot, X_rot[sel])
    adam_step(params_rot, rotation_backward(params_rot, rot_label[sel], logits, cache),
              moments, step)
    if step == 1 or step % 100 == 0:
        print(f"  iteration {step:3d}: rotation recognized in "
              f"{(logits.argmax(1) == rot_label[sel]).mean():.3f} of the step's patches")
X_rot_test, rot_label_test = rotations(X_img[test])
rot_acc = float((in_slices(lambda z: rotation_forward(params_rot, z)[0], X_rot_test).argmax(1)
                 == rot_label_test).mean())
print(f"  rotation recognized in {rot_acc:.3f} of the rotated test patches (chance 0.25)")
check("[B-rotation] the rotation of a test patch is recognized in at least 35% of "
      "the cases, against 25% by chance", rot_acc >= 0.35)

B-transfer

Freeze the pretrained layers, pool their output to 128 features, and fit a linear classifier of vineyard vs. no vineyard on 32 to 2304 labeled patches of the training half; compare its test accuracy with the same classifier on mean color, on raw pixels and on the features of the untrained layers, and with the majority baseline.

def logistic_regression(x_tr, y_tr, ridge=1e-2, iters=30):
    """Logistic regression with a ridge penalty on standardized features."""
    mu, sd = x_tr.mean(0), x_tr.std(0) + 1e-6                # training statistics
    z = np.column_stack([np.ones(len(x_tr)), (x_tr - mu) / sd]).astype(np.float64)
    w = np.zeros(z.shape[1])
    reg = ridge * np.eye(z.shape[1])
    reg[0, 0] = 0.0
    for _ in range(iters):
        q = 1.0 / (1.0 + np.exp(-z @ w))
        grad = z.T @ (q - y_tr) / len(z) + reg @ w
        hess = (z * (q * (1 - q))[:, None]).T @ z / len(z) + reg
        w -= np.linalg.solve(hess, grad)
    return lambda x: (np.column_stack([np.ones(len(x)), (x - mu) / sd]) @ w > 0).astype(np.float32)


def frozen_features(p):
    return in_slices(lambda z: pool4(layers(p, z)[0]), X_img)


feature_sets = {
    "mean color": patches.mean(axis=(1, 2)),
    "raw pixels": X_img.reshape(len(y), -1),
    "untrained layers": frozen_features(params_random),
    "color-pretrained layers": frozen_features(params_color),
    "rotation-pretrained layers": frozen_features(params_rot),
}
majority = 1.0 if y[train].mean() >= 0.5 else 0.0    # the more frequent label
acc_base = float((majority == y[test]).mean())
SIZES = (32, 64, 128, 256, 512, 1024, 2304)           # labeled patches used
acc = {name: [] for name in feature_sets}             # test accuracy against the size
for m in SIZES:
    stride = len(train_idx) // m
    for name, f in feature_sets.items():
        runs = []
        for offset in range(min(5, stride)):          # up to five spread-out subsets
            idx = train_idx[offset::stride][:m]
            predict = logistic_regression(f[idx], y[idx])
            runs.append(float((predict(f[test]) == y[test]).mean()))
        acc[name].append(float(np.mean(runs)))
    print(f"  {m:4d} labeled patches: test accuracy "
          + ", ".join(f"{name} {acc[name][-1]:.3f}" for name in feature_sets))
print(f"  majority baseline (always 'no vineyard'): test accuracy {acc_base:.3f}")
check("[B-transfer] with 32 and 64 labeled patches the color-pretrained layers beat "
      "the untrained ones by at least 0.02; with all 2304 the two are within 0.015",
      all(p - u >= 0.02 for p, u in zip(acc["color-pretrained layers"][:2], acc["untrained layers"][:2]))
      and abs(acc["color-pretrained layers"][-1] - acc["untrained layers"][-1]) <= 0.015)
check("[B-transfer] the rotation-pretrained layers beat the untrained ones at every "
      "number of labeled patches, and are never below the color-pretrained ones",
      all(r > u and r >= c for r, u, c in zip(acc["rotation-pretrained layers"],
                                              acc["untrained layers"],
                                              acc["color-pretrained layers"])))
check("[B-transfer] both pretrained layer sets beat mean color and raw pixels at every "
      "number of labeled patches",
      all(min(c, r) > max(a, b) for c, r, a, b in
          zip(acc["color-pretrained layers"], acc["rotation-pretrained layers"],
              acc["mean color"], acc["raw pixels"])))
check("[B-transfer] the mean color of a patch stays within 0.05 of the majority "
      "baseline at every number of labeled patches: color alone does not tell a "
      "vineyard", all(abs(a - acc_base) <= 0.05 for a in acc["mean color"]))

B-plot

Write the CSV files and images for the entry's figures and the preview.

with open(OUT_DIR / "pretraining_loss.csv", "w") as f:
    f.write("iteration,train,test\n")
    for step, tr, te in loss_log:
        f.write(f"{step},{tr:.6f},{te:.6f}\n")
with open(OUT_DIR / "pretraining_accuracy.csv", "w") as f:
    f.write("labels,meancolor,rawpixels,untrained,color,rotation,baseline\n")
    for i, m in enumerate(SIZES):
        f.write(f"{m},{acc['mean color'][i]:.4f},{acc['raw pixels'][i]:.4f},"
                f"{acc['untrained layers'][i]:.4f},{acc['color-pretrained layers'][i]:.4f},"
                f"{acc['rotation-pretrained layers'][i]:.4f},{acc_base:.4f}\n")

show = np.flatnonzero(test)[::288][:8]                # eight test patches, spread out
pred_show, _ = color_forward(params_color, X_blank[show])
GAP = 3
strip = np.ones((3 * SIDE + 2 * GAP, 8 * SIDE + 7 * GAP, 3), np.float32)
for i, idx in enumerate(show):
    col = i * (SIDE + GAP)
    blanked = X_blank[idx] + mean_color
    blanked[C0:C1, C0:C1, :] = 1.0                    # the blank shown white
    filled = blanked.copy()
    filled[C0:C1, C0:C1, :] = pred_show[i] + mean_color
    for r, img in enumerate((blanked, filled, patches[idx])):
        strip[r * (SIDE + GAP):r * (SIDE + GAP) + SIDE, col:col + SIDE] = np.clip(img, 0, 1)
plt.imsave(OUT_DIR / "pretraining_patches.png", strip.repeat(3, axis=0).repeat(3, axis=1))

fig, axes = plt.subplots(1, 3, figsize=(14, 3.8))
ax = axes[0]
steps = [s for s, _, _ in loss_log]
ax.plot(steps, [tr for _, tr, _ in loss_log], "-", color="black", label="training half")
ax.plot(steps, [te for _, _, te in loss_log], "--", color="0.5", label="test half")
ax.axhline(mse_ring, color="black", linestyle=":", label="mean color outside the blank (test)")
ax.set_yscale("log")
ax.set_xlabel("Adam iteration")
ax.set_ylabel("mean squared error of the center color")
ax.set_title("Color pretext task: predict the blanked center")
ax.legend(frameon=False, fontsize=8)
ax = axes[1]
ax.imshow(strip)
ax.set_xticks([])
ax.set_yticks([SIDE // 2 + r * (SIDE + GAP) for r in range(3)])
ax.set_yticklabels(["blanked", "predicted", "original"], fontsize=8)
ax.set_xlabel("eight test patches of 25 m")
ax.set_ylabel("row")
ax.set_title("Center color predicted from the surroundings")
ax = axes[2]
styles = {"mean color": ("^:", "0.5"), "raw pixels": ("s--", "0.5"),
          "untrained layers": ("o-", "0.5"), "color-pretrained layers": ("D-", "black"),
          "rotation-pretrained layers": ("v-", "black")}
for name, (style, color) in styles.items():
    ax.semilogx(SIZES, acc[name], style, color=color, markersize=4, label=name)
ax.axhline(acc_base, color="black", linestyle=":", linewidth=0.8, label="majority baseline")
ax.set_xlabel("number of labeled patches")
ax.set_ylabel("test accuracy (eastern half)")
ax.set_title("Linear classifier: vineyard vs. no vineyard")
ax.legend(frameon=False, fontsize=7, loc="lower right")
fig.tight_layout()
fig.savefig(OUT_DIR / "pretraining.png", dpi=150)
plt.close(fig)
check("[B-plot] the CSV files and the patch strip were written",
      all((OUT_DIR / f).exists() for f in
          ("pretraining_loss.csv", "pretraining_accuracy.csv", "pretraining_patches.png")))

passed = sum(1 for _, ok in report if ok)
print(f"\n{passed}/{len(report)} checks pass")
Preview figure produced by pretraining.py
The preview figure the block B-plot writes when the script runs