Dictionary of Applied Machine Learning · data point

data point — Python demo

Numerical companion to the entry data point: it recomputes what the entry states and prints one line per check

One block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.

Run it with python3 datapoint.py, from any directory — it writes its output files into the current directory. Requires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Download datapoint.py

The script, block by block

One cell per block of the script: the code, and what that code printed when it last ran here

setup

"""
datapoint.py — numerical companion to the glossary entry 'data point'.

One block per paragraph of the entry (marked [P...]): each block verifies
numerically what the corresponding statement asserts. Self-contained
(numpy/matplotlib only), fixed seed.

Blocks
------
[P-featlab]  A data point carries two categories of properties: features
             (measurable/computable) and labels (higher-level facts). A
             synthetic image data point yields pixel-intensity features
             by computation, while its label (number of bright objects)
             is fixed by construction — known to the "expert" that
             generated the scene, not read off a sensor.
[P-image]    The image example: color intensities of all pixels serve as
             features x_1..x_d, and can be augmented with capture
             metadata (timestamp, location) as further features.
[P-choice]   Feature vs label is a design choice: the same attribute
             (body weight) acts as a feature when predicting disease and
             as the label when predicted from other attributes — both
             predictions run on the same patient table.
[P-labelnoise] Labels are error-prone proxies: one-sided label noise
             (a fraction of positive training labels recorded as
             negative, as when a diagnosis is missed) monotonically
             depresses the fraction of correctly predicted positive test data points for the learned
             classifier on clean test data.
[P-featnoise] Features are error-prone too: adding measurement noise to
             the features at prediction time increases the error of a
             fixed learned hypothesis monotonically in the sensor noise.

Outputs
-------
datapoint.png : preview figure (checking only).

Data generated by pythondemos/datapoint.py.
"""

import numpy as np
import matplotlib

matplotlib.use("Agg")
import matplotlib.pyplot as plt

rng = np.random.default_rng(42)
report = []


def check(name, ok):
    report.append((name, bool(ok)))
    print(f"  [{'ok' if ok else 'FAIL'}] {name}")

P-featlab

A data point carries two categories of properties: features (measurable/computable) and labels (higher-level facts). A synthetic image data point yields pixel-intensity features by computation, while its label (number of bright objects) is fixed by construction — known to the "expert" that generated the scene, not read off a sensor.

print("[P-featlab] features are computed; labels are known facts")
def make_scene(n_objects):
    img = 0.05 * rng.random((16, 16))
    for _ in range(n_objects):
        i, j = rng.integers(2, 14, size=2)
        img[i - 1:i + 2, j - 1:j + 2] = 1.0
    return img

label_true = 3                                   # higher-level fact
img = make_scene(label_true)
features = np.array([img.mean(), img.std(), img.max()])
check("features are computed from the data point itself",
      np.isclose(features[0], img.mean()))
check("the label is not a pixel statistic (needs scene knowledge)",
      label_true not in np.round(features).astype(int)[:1])
[P-featlab] features are computed; labels are known facts
  [ok] features are computed from the data point itself
  [ok] the label is not a pixel statistic (needs scene knowledge)

P-image

The image example: color intensities of all pixels serve as features x_1..x_d, and can be augmented with capture metadata (timestamp, location) as further features.

print("[P-image] pixel intensities + metadata as features")
x_pixels = img.flatten()
x_meta = np.array([1717.0, 47.5])                # timestamp, latitude
x = np.concatenate([x_pixels, x_meta])
check("pixel features have length d = 256", x_pixels.size == 16 * 16)
check("metadata extends the features to d + 2", x.size == 258)
[P-image] pixel intensities + metadata as features
  [ok] pixel features have length d = 256
  [ok] metadata extends the features to d + 2

P-choice

Feature vs label is a design choice: the same attribute (body weight) acts as a feature when predicting disease and as the label when predicted from other attributes — both predictions run on the same patient table.

print("[P-choice] feature vs label is a design choice")
m = 300
weight = 60 + 20 * rng.random(m)
age = 20 + 50 * rng.random(m)
disease = (0.03 * weight + 0.05 * age + rng.normal(0, 0.4, m) > 4.5)
# design A: weight is a FEATURE for predicting the disease label
XA = np.stack([weight, age], axis=1)
wA = np.linalg.lstsq(np.c_[XA, np.ones(m)], disease.astype(float),
                     rcond=None)[0]
accA = np.mean((np.c_[XA, np.ones(m)] @ wA > 0.5) == disease)
# design B: weight is the LABEL predicted from age and disease status
XB = np.c_[age, disease.astype(float), np.ones(m)]
wB = np.linalg.lstsq(XB, weight, rcond=None)[0]
check("design A: weight used as a feature (prediction beats chance)",
      accA > 0.6)
check("design B: weight used as the label (predicted better than by its average)",
      np.var(weight - XB @ wB) < np.var(weight))
[P-choice] feature vs label is a design choice
  [ok] design A: weight used as a feature (prediction beats chance)
  [ok] design B: weight used as the label (predicted better than by its average)

P-labelnoise

Labels are error-prone proxies: one-sided label noise (a fraction of positive training labels recorded as negative, as when a diagnosis is missed) monotonically depresses the fraction of correctly predicted positive test data points for the learned classifier on clean test data.

print("[P-labelnoise] label noise degrades the learned hypothesis")
def train_test_acc(flip, reps=25):
    accs_r = []
    for _ in range(reps):
        Xtr, Xte = XA[:200], XA[200:]
        ytr, yte = disease[:200].copy(), disease[200:]
        pos = np.nonzero(ytr)[0]                  # missed diagnoses:
        idx = rng.choice(pos, int(flip * pos.size), replace=False)
        ytr[idx] = False                          # positives recorded negative
        wn = np.linalg.lstsq(np.c_[Xtr, np.ones(200)],
                             ytr.astype(float), rcond=None)[0]
        pred = np.c_[Xte, np.ones(100)] @ wn > 0.5
        accs_r.append(np.mean(pred[yte]))         # correct predictions on positive data points
    return float(np.mean(accs_r))

accs = [train_test_acc(f) for f in (0.0, 0.2, 0.4)]
print(f"    correctly predicted positives at flip = 0, 0.2, 0.4: "
      f"{accs[0]:.2f}, {accs[1]:.2f}, {accs[2]:.2f}")
check("correct positive predictions decrease with the label-noise level",
      accs[0] > accs[1] > accs[2])
[P-labelnoise] label noise degrades the learned hypothesis
    correctly predicted positives at flip = 0, 0.2, 0.4: 0.88, 0.75, 0.45
  [ok] correct positive predictions decrease with the label-noise level

P-featnoise

Features are error-prone too: adding measurement noise to the features at prediction time increases the error of a fixed learned hypothesis monotonically in the sensor noise.

print("[P-featnoise] feature measurement noise degrades predictions")
w_clean = np.linalg.lstsq(np.c_[XA, np.ones(m)], disease.astype(float),
                          rcond=None)[0]
errs = []
for s in (0.0, 5.0, 15.0):
    Xn = XA + rng.normal(0, s, XA.shape)          # sensor uncertainty
    errs.append(np.mean((np.c_[Xn, np.ones(m)] @ w_clean > 0.5)
                        != disease))
print(f"    error at sensor noise 0, 5, 15: "
      f"{errs[0]:.2f}, {errs[1]:.2f}, {errs[2]:.2f}")
check("prediction error grows with feature noise", errs[0] < errs[1] < errs[2])

# ------------------------------------------------------------ preview
fig, ax = plt.subplots(1, 2, figsize=(8.2, 3.0))
ax[0].imshow(img, cmap="gray")
ax[0].set_title(f"[P-featlab] data point (label = {label_true})")
ax[1].plot([0, 0.2, 0.4], accs, "o-")
ax[1].set_xlabel("label-noise level"); ax[1].set_ylabel("correct positive predictions")
ax[1].set_title("[P-labelnoise]")
fig.tight_layout()
fig.savefig("datapoint.png", dpi=110)
print(f"\n{sum(ok for _, ok in report)}/{len(report)} checks passed")
assert all(ok for _, ok in report)
[P-featnoise] feature measurement noise degrades predictions
    error at sensor noise 0, 5, 15: 0.13, 0.15, 0.25
  [ok] prediction error grows with feature noise

8/8 checks passed
Preview figure produced by datapoint.py
The preview figure the block P-featnoise writes when the script runs