Dictionary of Applied Machine Learning · data
Numerical companion to the entry data: it recomputes what the entry states and prints one line per check
One block per paragraph of the entry (marked [P...]): each block verifies numerically what the corresponding statement asserts. Self-contained (numpy/matplotlib only), fixed seed.
Run it without installing anything:uv run https://dictionaryofml.org/terms/data.py
uv downloads this script and the pinned NumPy and Matplotlib it needs, then runs it; the script fetches any input file it uses. To keep the output files, download data.py into a folder and run uv run data.py there. With NumPy and Matplotlib already installed, python3 data.py, from any directory — it writes its output files into the current directory. Fixed seeds, so the printed numbers reproduce exactly. Download data.py · Notebook · Open in Colab
One cell per block of the script: the code, and what that code printed when it last ran here
"""
data.py — numerical companion to the glossary entry 'data'.
One block per paragraph of the entry (marked [P...]): each block verifies
numerically what the corresponding statement asserts. Self-contained
(numpy/matplotlib only), fixed seed.
Blocks
------
[P-record] Data are representations of information recorded in a form
suitable for storage, communication, and processing: a
temperature reading with timestamp survives a round trip
through serialized bytes unchanged. The usefulness of a
learned hypothesis is limited by data quality and
quantity: test error decreases with the trainset size and
increases with measurement noise.
[P-datapoint] The unit sense: a data point (x, y) bundles features and
a label; the entry's convention x easy to obtain, y the
quantity of interest.
[P-dataset] The collection sense: a dataset of m data points supports
model training and validation via a train/validation
split.
Outputs
-------
data.png : preview figure (checking only).
Data generated by pythondemos/data.py.
"""
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from pathlib import Path
OUT_DIR = Path(__file__).parent
rng = np.random.default_rng(42)
report = []
def check(name, ok):
report.append((name, bool(ok)))
print(f" [{'ok' if ok else 'FAIL'}] {name}")
Data are representations of information recorded in a form suitable for storage, communication, and processing: a temperature reading with timestamp survives a round trip through serialized bytes unchanged. The usefulness of a learned hypothesis is limited by data quality and quantity: test error decreases with the trainset size and increases with measurement noise.
print("[P-record] recorded, storable, communicable representations")
reading = np.array([(1717.25, 23.4)],
dtype=[("t", "f8"), ("temp", "f8")])
blob = reading.tobytes() # storage/communication
restored = np.frombuffer(blob, dtype=reading.dtype)
check("a temperature reading survives the storage round trip",
restored["temp"][0] == 23.4 and restored["t"][0] == 1717.25)
# quality and quantity limit the learned hypothesis
def test_err(m, noise):
x = rng.uniform(0, 10, m)
y = 2.0 * x + 1.0 + noise * rng.normal(size=m)
c = np.polyfit(x, y, 1)
xt = rng.uniform(0, 10, 2000)
return np.mean((2.0 * xt + 1.0 - np.polyval(c, xt)) ** 2)
e_small, e_large = np.mean([test_err(10, 1.0) for _ in range(40)]), \
np.mean([test_err(200, 1.0) for _ in range(40)])
e_clean, e_noisy = np.mean([test_err(50, 0.2) for _ in range(40)]), \
np.mean([test_err(50, 2.0) for _ in range(40)])
check("more data points -> lower test error (quantity)",
e_large < e_small)
check("noisier recordings -> higher test error (quality)",
e_noisy > e_clean)
[P-record] recorded, storable, communicable representations [ok] a temperature reading survives the storage round trip [ok] more data points -> lower test error (quantity) [ok] noisier recordings -> higher test error (quality)
The unit sense: a data point (x, y) bundles features and a label; the entry's convention x easy to obtain, y the quantity of interest.
print("[P-datapoint] the unit sense: features and label")
z = {"x": np.array([23.4, 55.0]), "y": 1.0} # (features, label)
check("a data point bundles features and a label",
z["x"].shape == (2,) and np.isscalar(z["y"]))
[P-datapoint] the unit sense: features and label [ok] a data point bundles features and a label
The collection sense: a dataset of m data points supports model training and validation via a train/validation split.
print("[P-dataset] the collection sense: training and validation")
m = 100
X = rng.uniform(0, 10, m)
Y = 2.0 * X + 1.0 + 0.5 * rng.normal(size=m)
tr, va = slice(0, 80), slice(80, 100)
c = np.polyfit(X[tr], Y[tr], 1)
val_err = np.mean((Y[va] - np.polyval(c, X[va])) ** 2)
check("the dataset supports training (fit on the training part)",
np.all(np.isfinite(c)))
check("and validation (error evaluated on held-out data points)",
np.isfinite(val_err) and val_err < 1.0)
# ------------------------------------------------------------ preview
fig, ax = plt.subplots(figsize=(4.8, 3.2))
ms = [10, 30, 100, 300]
ax.loglog(ms, [np.mean([test_err(mm, 1.0) for _ in range(30)])
for mm in ms], "o-")
ax.set_xlabel("trainset size m"); ax.set_ylabel("test error")
ax.set_title("[P-record] usefulness grows with data quantity")
fig.tight_layout()
fig.savefig(OUT_DIR / "data.png", dpi=110)
print(f"\n{sum(ok for _, ok in report)}/{len(report)} checks passed")
assert all(ok for _, ok in report)
[P-dataset] the collection sense: training and validation [ok] the dataset supports training (fit on the training part) [ok] and validation (error evaluated on held-out data points) 6/6 checks passed

P-dataset writes when the script runs