{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "pretraining.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# pretraining \u2014 Python demo\n\nNumerical companion to the entry [pretraining](https://dictionaryofml.org/terms/pretraining.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nNumerical companion to the glossary entry 'pretraining'. A 2.4 x 1.2 km piece of the aerial photograph of the Wachau (assets/ wachau_ortho_center.jpg: the Danube bend at Duernstein, 0.79 m per pixel; orthophoto (c) basemap.at, CC BY 4.0) is cut into 4608 square patches of 32 px (25 m), and the parcel mask assets/ wachau_labels_center.png (INVEKOS Schlaege 2025-1, AgrarMarkt Austria, CC BY 3.0 AT) labels a patch as vineyard when more than half of its pixels lie on vineyard parcels. The western half of the piece is the training set, the eastern half the test set. At this resolution the rows of vines are visible stripes, and the mean color of a patch tells little: vineyards, meadows and gardens are all green.\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/pretraining.py`](https://dictionaryofml.org/terms/pretraining.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"pretraining.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"Pretraining the first layers of a convolutional network on a pretext\ntask without labels, then reusing them as frozen features for a linear\nclassifier of vineyard patches.\n\nPurpose\n-------\nNumerical companion to the glossary entry 'pretraining'.  A 2.4 x 1.2 km\npiece of the aerial photograph of the Wachau (assets/\nwachau_ortho_center.jpg: the Danube bend at Duernstein, 0.79 m per\npixel; orthophoto (c) basemap.at, CC BY 4.0) is cut into 4608 square\npatches of 32 px (25 m), and the parcel mask assets/\nwachau_labels_center.png (INVEKOS Schlaege 2025-1, AgrarMarkt Austria,\nCC BY 3.0 AT) labels a patch as vineyard when more than half of its\npixels lie on vineyard parcels.  The western half of the piece is the\ntraining set, the eastern half the test set.  At this resolution the\nrows of vines are visible stripes, and the mean color of a patch tells\nlittle: vineyards, meadows and gardens are all green.\n\nPretraining needs no labels.  Two pretext tasks are run on the training\npatches, each with the same network: two convolutional layers (3 x 3\nfilters, 8 channels each, ReLU) and a linear readout.  The first blanks\nthe central 8 x 8 pixel block of a patch and predicts the mean color of\nthe blanked pixels from their surroundings.  The second rotates a patch\nby 0, 90, 180 or 270 degrees and predicts the rotation, which is\nanswered by the direction of the rows and of the shadows.  Only the\nphotograph is used, never the parcel mask.\n\nTransfer.  The two convolutional layers are then frozen and applied to\nthe original patches; their output, average-pooled to a 4 x 4 grid,\ngives 128 features per patch.  A linear classifier (logistic regression\nwith a small ridge penalty) is trained on these features of 32 to 2304\nlabeled western patches and evaluated on the eastern patches.  The\ncomparison: the same linear classifier on the mean color of the patch\n(3 features), on the raw pixels (3072 features), and on the features of\nthe same two layers with their random initial weights, i.e., without\npretraining; and the majority baseline.  The last comparison is the\ncontrol: the color-pretrained layers beat the untrained ones when labels\nare few and are matched by them when all labels are used, while the\nrotation-pretrained layers beat the untrained ones at every number of\nlabels.\n\nDeterministic: fixed seeds for the initial weights and the patches of each step.\nSelf-contained: numpy + matplotlib only.\n\nBlocks\n------\n[B-patches]  Cut the photograph into 4608 patches of 32 px (25 m), label\n             them by the parcel mask, and split them by easting into a\n             western training half and an eastern test half.\n[B-pretrain] Blank the center of every training patch and train the two\n             convolutional layers plus a linear readout to predict the\n             center color (Adam, 512 patches per step).  Check that the\n             loss falls and that the trained network beats the patch's\n             mean color as a prediction of the center on the test half.\n[B-rotation] The second pretext task: predict the rotation of a patch\n             (Adam, 512 patches per step).  Check that the rotation is\n             recognized well above chance on the test half.\n[B-transfer] Freeze the pretrained layers, pool their output to 128\n             features, and fit a linear classifier of vineyard vs. no\n             vineyard on 32 to 2304 labeled patches of the training\n             half; compare its test accuracy with the same classifier on\n             mean color, on raw pixels and on the features of the\n             untrained layers, and with the majority baseline.\n[B-plot]     Write the CSV files and images for the entry's figures and\n             the preview.\n\nOutputs\n-------\npretraining_loss.csv : iteration, train, test -- the color pretext loss\n                       (mean squared error of the predicted center\n                       color, in units of the pixel range [0, 1]) on the\n                       two halves\npretraining_accuracy.csv : labels, meancolor, rawpixels, untrained,\n                       color, rotation, baseline -- the test accuracy of\n                       the linear classifier on each feature set against\n                       the number of labeled patches (mean over up to\n                       five spread-out subsets), and the majority\n                       baseline\npretraining_patches.png : eight test patches: the blanked input, the\n                       network's prediction of the center painted in,\n                       and the original\npretraining.png : preview (checking only)\n\"\"\"\n\nimport math\nfrom pathlib import Path\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\n\nOUT_DIR = Path(__file__).parent\n\nreport = []                         # collects (check name, pass/fail) pairs\n\n\ndef check(name, ok):                # records and prints one verification\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-patches]** Cut the photograph into 4608 patches of 32 px (25 m), label them by the parcel mask, and split them by easting into a western training half and an eastern test half."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "SIDE = 32                           # 32 px = 25 m on the ground\nphoto = imread(OUT_DIR.parent / \"assets\" / \"wachau_ortho_center.jpg\") / 255.0\nlabels = np.asarray(imread(OUT_DIR.parent / \"assets\" / \"wachau_labels_center.png\"))\nif labels.ndim == 3:\n    labels = labels[:, :, 0]\nif labels.max() <= 1.0 and labels.dtype != np.uint8:\n    labels = np.rint(labels * 255).astype(np.uint8)\nrows, cols = photo.shape[0] // SIDE, photo.shape[1] // SIDE\npatches = photo[:rows * SIDE, :cols * SIDE].reshape(rows, SIDE, cols, SIDE, 3)\npatches = patches.transpose(0, 2, 1, 3, 4).reshape(rows * cols, SIDE, SIDE, 3).astype(np.float32)\nshare = (labels[:rows * SIDE, :cols * SIDE] == 1).reshape(\n    rows, SIDE, cols, SIDE).mean(axis=(1, 3)).ravel()\ny = (share > 0.5).astype(np.float32)                  # 1 = vineyard\ngrid_c = np.arange(rows * cols) % cols\ntrain = grid_c < cols // 2                            # western half\ntest = ~train\nmean_color = patches[train].mean(axis=(0, 1, 2))      # training statistics only\nX_img = patches - mean_color                          # centered pixels\nprint(f\"  {len(y)} patches of {rows} x {cols}, {SIDE} x {SIDE} px each; vineyard share \"\n      f\"{y.mean():.3f} (west {y[train].mean():.3f}, east {y[test].mean():.3f})\")\ncheck(\"[B-patches] 4608 patches, split into halves of equal size by easting\",\n      len(y) == 4608 and train.sum() == test.sum() == 2304)\ncheck(\"[B-patches] between a fifth and two fifths of the patches are vineyard\",\n      0.2 < y.mean() < 0.4)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-pretrain]** Blank the center of every training patch and train the two convolutional layers plus a linear readout to predict the center color (Adam, 512 patches per step). Check that the loss falls and that the trained network beats the patch's mean color as a prediction of the center on the test half."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "rng = np.random.default_rng(0)\nCH = 8                              # channels of both convolutional layers\nC0, C1 = 12, 20                     # the blanked center block: rows and columns 12..19\nITERS, BATCH, LR, B1, B2 = 300, 512, 3e-3, 0.9, 0.999\n\n\ndef conv_cols(x):\n    \"\"\"3 x 3 neighborhoods of every pixel (zero padding): (N, H, W, 9 C).\"\"\"\n    n, h, w, c = x.shape\n    xp = np.pad(x, ((0, 0), (1, 1), (1, 1), (0, 0)))\n    return np.concatenate([xp[:, i:i + h, j:j + w, :] for i in range(3) for j in range(3)],\n                          axis=3)\n\n\ndef conv_cols_back(dcols, shape):\n    \"\"\"Adjoint of conv_cols: scatter the neighborhood gradients back to pixels.\"\"\"\n    n, h, w, c = shape\n    dxp = np.zeros((n, h + 2, w + 2, c), dtype=np.float32)\n    for k, (i, j) in enumerate((i, j) for i in range(3) for j in range(3)):\n        dxp[:, i:i + h, j:j + w, :] += dcols[..., k * c:(k + 1) * c]\n    return dxp[:, 1:-1, 1:-1, :]\n\n\ndef init_params(n_in, n_out):\n    \"\"\"He initialization of the two convolutional layers and a linear readout.\"\"\"\n    w1 = (rng.standard_normal((27, CH)) * math.sqrt(2.0 / 27)).astype(np.float32)\n    w2 = (rng.standard_normal((9 * CH, CH)) * math.sqrt(2.0 / (9 * CH))).astype(np.float32)\n    wh = (rng.standard_normal((n_in * CH, n_out)) * math.sqrt(1.0 / (n_in * CH))).astype(np.float32)\n    return {\"w1\": w1, \"b1\": np.zeros(CH, np.float32), \"w2\": w2,\n            \"b2\": np.zeros(CH, np.float32), \"wh\": wh, \"bh\": np.zeros(n_out, np.float32)}\n\n\ndef layers(p, x):\n    \"\"\"The two convolutional layers with ReLU: their output (N, 32, 32, CH).\"\"\"\n    c1 = conv_cols(x)\n    a1 = np.maximum(c1 @ p[\"w1\"] + p[\"b1\"], 0.0)\n    c2 = conv_cols(a1)\n    a2 = np.maximum(c2 @ p[\"w2\"] + p[\"b2\"], 0.0)\n    return a2, (c1, a1, c2)\n\n\ndef layers_back(p, da2, saved, grads):\n    \"\"\"Backpropagate a gradient at the layers' output into their weights.\"\"\"\n    c1, a1, c2 = saved\n    grads[\"w2\"] = c2.reshape(-1, 9 * CH).T @ da2.reshape(-1, CH)\n    grads[\"b2\"] = da2.sum((0, 1, 2))\n    da1 = conv_cols_back(da2 @ p[\"w2\"].T, a1.shape)\n    da1 *= a1 > 0\n    grads[\"w1\"] = c1.reshape(-1, 27).T @ da1.reshape(-1, CH)\n    grads[\"b1\"] = da1.sum((0, 1, 2))\n    return grads\n\n\ndef adam_step(p, grads, moments, step):\n    for k in p:\n        m, v = moments[k]\n        m[...] = B1 * m + (1 - B1) * grads[k]\n        v[...] = B2 * v + (1 - B2) * grads[k] ** 2\n        p[k] -= LR * (m / (1 - B1 ** step)) / (np.sqrt(v / (1 - B2 ** step)) + 1e-8)\n\n\ndef color_forward(p, x):\n    a2, saved = layers(p, x)\n    block = a2[:, C0:C1, C0:C1, :].reshape(len(x), -1)\n    return block @ p[\"wh\"] + p[\"bh\"], (a2, block, saved)\n\n\ndef color_backward(p, target, out, cache):\n    \"\"\"Gradients of the mean squared error of the center color.\"\"\"\n    a2, block, saved = cache\n    n = len(out)\n    dout = 2.0 * (out - target) / (n * 3)\n    grads = {\"wh\": block.T @ dout, \"bh\": dout.sum(0)}\n    da2 = np.zeros_like(a2)\n    da2[:, C0:C1, C0:C1, :] = (dout @ p[\"wh\"].T).reshape(n, C1 - C0, C1 - C0, CH)\n    da2 *= a2 > 0\n    return layers_back(p, da2, saved, grads)\n\n\ndef blank(x):\n    \"\"\"The pretext input: the center block set to zero (the mean color).\"\"\"\n    xb = x.copy()\n    xb[:, C0:C1, C0:C1, :] = 0.0\n    return xb\n\n\ndef in_slices(fn, x, size=512):\n    return np.concatenate([fn(x[i:i + size]) for i in range(0, len(x), size)])\n\n\nX_blank = blank(X_img)\ntarget = X_img[:, C0:C1, C0:C1, :].mean(axis=(1, 2))   # the color to predict\nparams_color = init_params((C1 - C0) ** 2, 3)\nparams_random = {k: v.copy() for k, v in params_color.items()}   # the untrained control\nmoments = {k: (np.zeros_like(v), np.zeros_like(v)) for k, v in params_color.items()}\ntrain_idx = np.flatnonzero(train)\nloss_log = []\nfor step in range(1, ITERS + 1):\n    sel = np.sort(rng.choice(train_idx, BATCH, replace=False))\n    out, cache = color_forward(params_color, X_blank[sel])\n    adam_step(params_color, color_backward(params_color, target[sel], out, cache), moments, step)\n    if step == 1 or step % 20 == 0:\n        err = [float(((in_slices(lambda z: color_forward(params_color, z)[0], X_blank[part])\n                        - target[part]) ** 2).mean()) for part in (train, test)]\n        loss_log.append((step, err[0], err[1]))\n        if step == 1 or step % 100 == 0:\n            print(f\"  iteration {step:3d}: color pretext loss train {err[0]:.5f}, test {err[1]:.5f}\")\nmse_net = loss_log[-1][2]\nring_mean = X_blank[test].sum(axis=(1, 2)) / (SIDE * SIDE - (C1 - C0) ** 2)\nmse_ring = float(((ring_mean - target[test]) ** 2).mean())\nprint(f\"  center color on the test half: mean squared error of the network {mse_net:.5f}, \"\n      f\"of the patch's mean color outside the blank {mse_ring:.5f}\")\ncheck(\"[B-pretrain] the color pretext loss on the training half falls to below a \"\n      \"tenth of its initial value\", loss_log[-1][1] < 0.1 * loss_log[0][1])\ncheck(\"[B-pretrain] on the test half the network predicts the center color better \"\n      \"than the patch's mean color outside the blank\", mse_net < mse_ring)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-rotation]** The second pretext task: predict the rotation of a patch (Adam, 512 patches per step). Check that the rotation is recognized well above chance on the test half."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def pool4(a2):\n    \"\"\"Average-pool the layers' output to a 4 x 4 grid: 128 numbers per patch.\"\"\"\n    n = len(a2)\n    return a2.reshape(n, 4, SIDE // 4, 4, SIDE // 4, CH).mean(axis=(2, 4)).reshape(n, -1)\n\n\ndef rotation_forward(p, x):\n    a2, saved = layers(p, x)\n    pooled = pool4(a2)\n    return pooled @ p[\"wh\"] + p[\"bh\"], (a2, pooled, saved)\n\n\ndef rotation_backward(p, target, logits, cache):\n    \"\"\"Gradients of the cross-entropy of the four rotations.\"\"\"\n    a2, pooled, saved = cache\n    n = len(logits)\n    q = np.exp(logits - logits.max(1, keepdims=True))\n    q /= q.sum(1, keepdims=True)\n    dout = (q - np.eye(4, dtype=np.float32)[target]) / n\n    grads = {\"wh\": pooled.T @ dout, \"bh\": dout.sum(0)}\n    dpool = (dout @ p[\"wh\"].T).reshape(n, 4, 1, 4, 1, CH) / (SIDE // 4) ** 2\n    da2 = np.broadcast_to(dpool, (n, 4, SIDE // 4, 4, SIDE // 4, CH)).reshape(a2.shape).copy()\n    da2 *= a2 > 0\n    return layers_back(p, da2, saved, grads)\n\n\ndef rotations(x):\n    \"\"\"Every patch in the four rotations, with the rotation as label.\"\"\"\n    return (np.concatenate([np.rot90(x, k, axes=(1, 2)) for k in range(4)]),\n            np.repeat(np.arange(4), len(x)))\n\n\nX_rot, rot_label = rotations(X_img[train])\nparams_rot = init_params(16, 4)\nmoments = {k: (np.zeros_like(v), np.zeros_like(v)) for k, v in params_rot.items()}\nfor step in range(1, ITERS + 1):\n    sel = np.sort(rng.choice(len(X_rot), BATCH, replace=False))\n    logits, cache = rotation_forward(params_rot, X_rot[sel])\n    adam_step(params_rot, rotation_backward(params_rot, rot_label[sel], logits, cache),\n              moments, step)\n    if step == 1 or step % 100 == 0:\n        print(f\"  iteration {step:3d}: rotation recognized in \"\n              f\"{(logits.argmax(1) == rot_label[sel]).mean():.3f} of the step's patches\")\nX_rot_test, rot_label_test = rotations(X_img[test])\nrot_acc = float((in_slices(lambda z: rotation_forward(params_rot, z)[0], X_rot_test).argmax(1)\n                 == rot_label_test).mean())\nprint(f\"  rotation recognized in {rot_acc:.3f} of the rotated test patches (chance 0.25)\")\ncheck(\"[B-rotation] the rotation of a test patch is recognized in at least 35% of \"\n      \"the cases, against 25% by chance\", rot_acc >= 0.35)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-transfer]** Freeze the pretrained layers, pool their output to 128 features, and fit a linear classifier of vineyard vs. no vineyard on 32 to 2304 labeled patches of the training half; compare its test accuracy with the same classifier on mean color, on raw pixels and on the features of the untrained layers, and with the majority baseline."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "def logistic_regression(x_tr, y_tr, ridge=1e-2, iters=30):\n    \"\"\"Logistic regression with a ridge penalty on standardized features.\"\"\"\n    mu, sd = x_tr.mean(0), x_tr.std(0) + 1e-6                # training statistics\n    z = np.column_stack([np.ones(len(x_tr)), (x_tr - mu) / sd]).astype(np.float64)\n    w = np.zeros(z.shape[1])\n    reg = ridge * np.eye(z.shape[1])\n    reg[0, 0] = 0.0\n    for _ in range(iters):\n        q = 1.0 / (1.0 + np.exp(-z @ w))\n        grad = z.T @ (q - y_tr) / len(z) + reg @ w\n        hess = (z * (q * (1 - q))[:, None]).T @ z / len(z) + reg\n        w -= np.linalg.solve(hess, grad)\n    return lambda x: (np.column_stack([np.ones(len(x)), (x - mu) / sd]) @ w > 0).astype(np.float32)\n\n\ndef frozen_features(p):\n    return in_slices(lambda z: pool4(layers(p, z)[0]), X_img)\n\n\nfeature_sets = {\n    \"mean color\": patches.mean(axis=(1, 2)),\n    \"raw pixels\": X_img.reshape(len(y), -1),\n    \"untrained layers\": frozen_features(params_random),\n    \"color-pretrained layers\": frozen_features(params_color),\n    \"rotation-pretrained layers\": frozen_features(params_rot),\n}\nmajority = 1.0 if y[train].mean() >= 0.5 else 0.0    # the more frequent label\nacc_base = float((majority == y[test]).mean())\nSIZES = (32, 64, 128, 256, 512, 1024, 2304)           # labeled patches used\nacc = {name: [] for name in feature_sets}             # test accuracy against the size\nfor m in SIZES:\n    stride = len(train_idx) // m\n    for name, f in feature_sets.items():\n        runs = []\n        for offset in range(min(5, stride)):          # up to five spread-out subsets\n            idx = train_idx[offset::stride][:m]\n            predict = logistic_regression(f[idx], y[idx])\n            runs.append(float((predict(f[test]) == y[test]).mean()))\n        acc[name].append(float(np.mean(runs)))\n    print(f\"  {m:4d} labeled patches: test accuracy \"\n          + \", \".join(f\"{name} {acc[name][-1]:.3f}\" for name in feature_sets))\nprint(f\"  majority baseline (always 'no vineyard'): test accuracy {acc_base:.3f}\")\ncheck(\"[B-transfer] with 32 and 64 labeled patches the color-pretrained layers beat \"\n      \"the untrained ones by at least 0.02; with all 2304 the two are within 0.015\",\n      all(p - u >= 0.02 for p, u in zip(acc[\"color-pretrained layers\"][:2], acc[\"untrained layers\"][:2]))\n      and abs(acc[\"color-pretrained layers\"][-1] - acc[\"untrained layers\"][-1]) <= 0.015)\ncheck(\"[B-transfer] the rotation-pretrained layers beat the untrained ones at every \"\n      \"number of labeled patches, and are never below the color-pretrained ones\",\n      all(r > u and r >= c for r, u, c in zip(acc[\"rotation-pretrained layers\"],\n                                              acc[\"untrained layers\"],\n                                              acc[\"color-pretrained layers\"])))\ncheck(\"[B-transfer] both pretrained layer sets beat mean color and raw pixels at every \"\n      \"number of labeled patches\",\n      all(min(c, r) > max(a, b) for c, r, a, b in\n          zip(acc[\"color-pretrained layers\"], acc[\"rotation-pretrained layers\"],\n              acc[\"mean color\"], acc[\"raw pixels\"])))\ncheck(\"[B-transfer] the mean color of a patch stays within 0.05 of the majority \"\n      \"baseline at every number of labeled patches: color alone does not tell a \"\n      \"vineyard\", all(abs(a - acc_base) <= 0.05 for a in acc[\"mean color\"]))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-plot]** Write the CSV files and images for the entry's figures and the preview."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "with open(OUT_DIR / \"pretraining_loss.csv\", \"w\") as f:\n    f.write(\"iteration,train,test\\n\")\n    for step, tr, te in loss_log:\n        f.write(f\"{step},{tr:.6f},{te:.6f}\\n\")\nwith open(OUT_DIR / \"pretraining_accuracy.csv\", \"w\") as f:\n    f.write(\"labels,meancolor,rawpixels,untrained,color,rotation,baseline\\n\")\n    for i, m in enumerate(SIZES):\n        f.write(f\"{m},{acc['mean color'][i]:.4f},{acc['raw pixels'][i]:.4f},\"\n                f\"{acc['untrained layers'][i]:.4f},{acc['color-pretrained layers'][i]:.4f},\"\n                f\"{acc['rotation-pretrained layers'][i]:.4f},{acc_base:.4f}\\n\")\n\nshow = np.flatnonzero(test)[::288][:8]                # eight test patches, spread out\npred_show, _ = color_forward(params_color, X_blank[show])\nGAP = 3\nstrip = np.ones((3 * SIDE + 2 * GAP, 8 * SIDE + 7 * GAP, 3), np.float32)\nfor i, idx in enumerate(show):\n    col = i * (SIDE + GAP)\n    blanked = X_blank[idx] + mean_color\n    blanked[C0:C1, C0:C1, :] = 1.0                    # the blank shown white\n    filled = blanked.copy()\n    filled[C0:C1, C0:C1, :] = pred_show[i] + mean_color\n    for r, img in enumerate((blanked, filled, patches[idx])):\n        strip[r * (SIDE + GAP):r * (SIDE + GAP) + SIDE, col:col + SIDE] = np.clip(img, 0, 1)\nplt.imsave(OUT_DIR / \"pretraining_patches.png\", strip.repeat(3, axis=0).repeat(3, axis=1))\n\nfig, axes = plt.subplots(1, 3, figsize=(14, 3.8))\nax = axes[0]\nsteps = [s for s, _, _ in loss_log]\nax.plot(steps, [tr for _, tr, _ in loss_log], \"-\", color=\"black\", label=\"training half\")\nax.plot(steps, [te for _, _, te in loss_log], \"--\", color=\"0.5\", label=\"test half\")\nax.axhline(mse_ring, color=\"black\", linestyle=\":\", label=\"mean color outside the blank (test)\")\nax.set_yscale(\"log\")\nax.set_xlabel(\"Adam iteration\")\nax.set_ylabel(\"mean squared error of the center color\")\nax.set_title(\"Color pretext task: predict the blanked center\")\nax.legend(frameon=False, fontsize=8)\nax = axes[1]\nax.imshow(strip)\nax.set_xticks([])\nax.set_yticks([SIDE // 2 + r * (SIDE + GAP) for r in range(3)])\nax.set_yticklabels([\"blanked\", \"predicted\", \"original\"], fontsize=8)\nax.set_xlabel(\"eight test patches of 25 m\")\nax.set_ylabel(\"row\")\nax.set_title(\"Center color predicted from the surroundings\")\nax = axes[2]\nstyles = {\"mean color\": (\"^:\", \"0.5\"), \"raw pixels\": (\"s--\", \"0.5\"),\n          \"untrained layers\": (\"o-\", \"0.5\"), \"color-pretrained layers\": (\"D-\", \"black\"),\n          \"rotation-pretrained layers\": (\"v-\", \"black\")}\nfor name, (style, color) in styles.items():\n    ax.semilogx(SIZES, acc[name], style, color=color, markersize=4, label=name)\nax.axhline(acc_base, color=\"black\", linestyle=\":\", linewidth=0.8, label=\"majority baseline\")\nax.set_xlabel(\"number of labeled patches\")\nax.set_ylabel(\"test accuracy (eastern half)\")\nax.set_title(\"Linear classifier: vineyard vs. no vineyard\")\nax.legend(frameon=False, fontsize=7, loc=\"lower right\")\nfig.tight_layout()\nfig.savefig(OUT_DIR / \"pretraining.png\", dpi=150)\nplt.close(fig)\ncheck(\"[B-plot] the CSV files and the patch strip were written\",\n      all((OUT_DIR / f).exists() for f in\n          (\"pretraining_loss.csv\", \"pretraining_accuracy.csv\", \"pretraining_patches.png\")))\n\npassed = sum(1 for _, ok in report if ok)\nprint(f\"\\n{passed}/{len(report)} checks pass\")"
  }
 ]
}