Dictionary of Applied Machine Learning · large language model (LLM)
Numerical companion to the entry large language model (LLM): it recomputes what the entry states and prints one line per check
Blocks verify numerically what the entry's statements assert. Self-contained (numpy/matplotlib only), fixed seed. (The entry's scale claim — billions of parameters — is illustrated in miniature: the mechanisms, not the size.)
Run it without installing anything:uv run https://dictionaryofml.org/terms/llm.py
uv downloads this script and the pinned NumPy and Matplotlib it needs, then runs it; the script fetches any input file it uses. To keep the output files, download llm.py into a folder and run uv run llm.py there. With NumPy and Matplotlib already installed, python3 llm.py, from any directory — it writes its output files into the current directory. Fixed seeds, so the printed numbers reproduce exactly. Download llm.py · Notebook · Open in Colab
One cell per block of the script: the code, and what that code printed when it last ran here
"""
llm.py — numerical companion to the glossary entry
'large language model (LLM)'.
Blocks verify numerically what the entry's statements assert. Self-contained
(numpy/matplotlib only), fixed seed. (The entry's scale claim —
billions of parameters — is illustrated in miniature: the mechanisms,
not the size.)
Blocks
------
[B-selfsup] Self-supervised construction of data points from
raw text: masking words turns an unannotated corpus into
(context features, masked-word label) pairs -- one data
point per position, with zero human annotation effort.
[B-train] Training via ERM on these pairs: a small next-token model
(embedding + softmax over the vocabulary) trained by
GD on the training loss (the negative log probability
of the correct next token) drives the training loss
down and beats the uniform-guess baseline.
[B-nexttoken] A trained LLM maps an input token sequence to a
probability distribution over the next token: outputs are
nonnegative, sum to one, and the model assigns the
highest probability to continuations seen in the corpus;
sampling from the distribution generates text.
[B-local] Local execution in miniature: quantizing the model
parameters to 8-bit integers shrinks the memory to an
eighth while the next-token probabilities barely move
and the predicted next token agrees on 90% of contexts.
[B-agent] One step of an LLM agent in miniature: the same
architecture trained on task--code pairs emits, for the
task token 'sum', the token 'print(1+3)'; the wrapper
compiles the emitted text as Python source code, runs
it, captures the output '4', and appends it to the next
prompt. The model only ever outputs text -- the wrapper
is what turns the text into an action.
Outputs
-------
pythondemos/llm.png : preview figure (checking only).
Data generated by pythondemos/llm.py.
"""
from pathlib import Path
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
rng = np.random.default_rng(42)
report = []
def check(name, ok):
report.append((name, bool(ok)))
print(f" [{'ok' if ok else 'FAIL'}] {name}")
corpus = ("all human beings are born free and equal in dignity and "
"rights all human beings are endowed with reason and "
"conscience").split()
vocab = sorted(set(corpus))
V = len(vocab)
tok = {w: i for i, w in enumerate(vocab)}
ids = np.array([tok[w] for w in corpus])
Self-supervised construction of data points from raw text: masking words turns an unannotated corpus into (context features, masked-word label) pairs -- one data point per position, with zero human annotation effort.
print("[B-selfsup] masked words become labels, contexts become features")
pairs = [(ids[t], ids[t + 1]) for t in range(len(ids) - 1)]
check("one labeled pair per corpus position (no human annotation)",
len(pairs) == len(corpus) - 1)
check("labels are drawn from the text itself",
all(0 <= y < V for _, y in pairs))
[B-selfsup] masked words become labels, contexts become features [ok] one labeled pair per corpus position (no human annotation) [ok] labels are drawn from the text itself
Training via ERM on these pairs: a small next-token model (embedding + softmax over the vocabulary) trained by GD on the training loss (the negative log probability of the correct next token) drives the training loss down and beats the uniform-guess baseline.
print("[B-train] ERM on the constructed pairs")
d_emb = 8
E = 0.1 * rng.normal(size=(V, d_emb)) # embeddings
U = 0.1 * rng.normal(size=(d_emb, V)) # unembedding
def forward(x_ids):
logits = E[x_ids] @ U
p = np.exp(logits - logits.max(1, keepdims=True))
return p / p.sum(1, keepdims=True)
xs = np.array([x for x, _ in pairs])
ys = np.array([y for _, y in pairs])
def xent():
return -np.mean(np.log(forward(xs)[np.arange(len(ys)), ys] + 1e-12))
loss0 = xent()
for _ in range(800): # GD
P = forward(xs)
G = P.copy(); G[np.arange(len(ys)), ys] -= 1.0
gU = E[xs].T @ G / len(ys)
gE = np.zeros_like(E)
np.add.at(gE, xs, G @ U.T / len(ys))
U -= 2.0 * gU; E -= 2.0 * gE
loss1 = xent()
print(f" training loss: init {loss0:.2f} -> trained {loss1:.2f} "
f"(uniform baseline {np.log(V):.2f})")
check("training reduces the training loss", loss1 < loss0)
check("the trained model beats the uniform-guess baseline log|V|",
loss1 < np.log(V) - 0.5)
[B-train] ERM on the constructed pairs
training loss: init 2.71 -> trained 0.23 (uniform baseline 2.71)
[ok] training reduces the training loss
[ok] the trained model beats the uniform-guess baseline log|V|
A trained LLM maps an input token sequence to a probability distribution over the next token: outputs are nonnegative, sum to one, and the model assigns the highest probability to continuations seen in the corpus; sampling from the distribution generates text.
print("[B-nexttoken] input sequence -> distribution over the next token")
p_next = forward(np.array([tok["human"]]))[0]
check("the output is a probability distribution (nonneg, sums to 1)",
np.all(p_next >= 0) and np.isclose(p_next.sum(), 1.0))
check("'human' is followed by 'beings' in the corpus — and gets the "
"highest next-token probability",
vocab[int(np.argmax(p_next))] == "beings")
gen = [tok["all"]]
for _ in range(5): # sample a continuation
gen.append(int(rng.choice(V, p=forward(np.array([gen[-1]]))[0])))
check("sampling from the distributions generates a token sequence",
len(gen) == 6 and all(0 <= g < V for g in gen))
print(" generated:", " ".join(vocab[g] for g in gen))
[B-nexttoken] input sequence -> distribution over the next token
[ok] the output is a probability distribution (nonneg, sums to 1)
[ok] 'human' is followed by 'beings' in the corpus — and gets the highest next-token probability
[ok] sampling from the distributions generates a token sequence
generated: all human beings are endowed with
Local execution in miniature: quantizing the model parameters to 8-bit integers shrinks the memory to an eighth while the next-token probabilities barely move and the predicted next token agrees on 90% of contexts.
print("[B-local] the same model, quantized to 8 bits, runs locally")
sE = np.max(np.abs(E)) / 127.0
sU = np.max(np.abs(U)) / 127.0
Eq = (np.round(E / sE).astype(np.int8) * sE) # 8-bit stored, dequantized
Uq = (np.round(U / sU).astype(np.int8) * sU)
def forward_q(x_ids):
logits = Eq[x_ids] @ Uq
p = np.exp(logits - logits.max(1, keepdims=True))
return p / p.sum(1, keepdims=True)
agree = float(np.mean(np.argmax(forward_q(xs), 1) == np.argmax(forward(xs), 1)))
gap = float(np.max(np.abs(forward_q(xs) - forward(xs))))
mem = (E.nbytes + U.nbytes) / (E.size + U.size) # bytes per parameter
print(f" {E.size + U.size} model parameters at {mem:.0f} bytes each -> "
f"1 byte each after quantization (memory shrinks {mem:.0f}x)")
print(f" next-token probabilities move by at most {gap:.4f}; the "
f"predicted next token agrees on {100 * agree:.0f}% of contexts "
f"(the rest are corpus-ambiguous near-ties)")
check("8-bit quantization shrinks the memory to an eighth", mem == 8.0)
check("next-token probabilities move by less than 0.01", gap < 0.01)
check("the predicted next token agrees on at least 85% of contexts",
agree >= 0.85)
[B-local] the same model, quantized to 8 bits, runs locally
240 model parameters at 8 bytes each -> 1 byte each after quantization (memory shrinks 8x)
next-token probabilities move by at most 0.0076; the predicted next token agrees on 90% of contexts (the rest are corpus-ambiguous near-ties)
[ok] 8-bit quantization shrinks the memory to an eighth
[ok] next-token probabilities move by less than 0.01
[ok] the predicted next token agrees on at least 85% of contexts
One step of an LLM agent in miniature: the same architecture trained on task--code pairs emits, for the task token 'sum', the token 'print(1+3)'; the wrapper compiles the emitted text as Python source code, runs it, captures the output '4', and appends it to the next prompt. The model only ever outputs text -- the wrapper is what turns the text into an action.
print("[B-agent] the wrapper runs the model's text output as Python")
import contextlib
# tiny task->code corpus: after each task token, the code token follows
corpus_a = ("sum print(1+3) done double print(2*3) done "
"sum print(1+3) done double print(2*3) done").split()
vocab_a = sorted(set(corpus_a))
tok_a = {w: i for i, w in enumerate(vocab_a)}
ids_a = np.array([tok_a[w] for w in corpus_a])
Va = len(vocab_a)
Ea = 0.1 * rng.normal(size=(Va, d_emb))
Ua = 0.1 * rng.normal(size=(d_emb, Va))
xa = ids_a[:-1]
ya = ids_a[1:]
def forward_a(x_ids):
logits = Ea[x_ids] @ Ua
pr = np.exp(logits - logits.max(1, keepdims=True))
return pr / pr.sum(1, keepdims=True)
for _ in range(800): # GD
Pa = forward_a(xa)
Ga = Pa.copy(); Ga[np.arange(len(ya)), ya] -= 1.0
gUa = Ea[xa].T @ Ga / len(ya)
gEa = np.zeros_like(Ea)
np.add.at(gEa, xa, Ga @ Ua.T / len(ya))
Ua -= 2.0 * gUa; Ea -= 2.0 * gEa
import io as _io
transcript = ["sum"] # the prompt: a task
emitted = vocab_a[int(np.argmax(forward_a(np.array([tok_a["sum"]]))[0]))]
print(f" prompt 'sum' -> model emits the text '{emitted}'")
code_ok = True
try:
compiled = compile(emitted, "<llm-output>", "exec")
except SyntaxError:
code_ok = False
buf = _io.StringIO()
if code_ok:
with contextlib.redirect_stdout(buf):
exec(compiled) # the wrapper acts
result = buf.getvalue().strip()
transcript += [emitted, result] # result -> next prompt
print(f" wrapper runs it and captures '{result}'; "
f"next prompt: {transcript}")
emitted2 = vocab_a[int(np.argmax(forward_a(np.array([tok_a["double"]]))[0]))]
check("the emitted text compiles as Python source code", code_ok)
check("running the emitted code yields the task's answer",
result == "4")
check("a different task token yields different code",
emitted2 == "print(2*3)" and emitted2 != emitted)
check("the captured result is appended to the next prompt",
transcript == ["sum", "print(1+3)", "4"])
# ------------------------------------------------------------ preview
fig, ax = plt.subplots(figsize=(5.6, 3.0))
ax.bar(range(V), p_next)
ax.set_xticks(range(V))
ax.set_xticklabels(vocab, rotation=90, fontsize=6)
ax.set_xlabel("token")
ax.set_ylabel("P(next token | 'human')")
ax.set_title("[B-nexttoken] next-token distribution")
fig.tight_layout()
OUT_DIR = Path(__file__).parent
fig.savefig(OUT_DIR / "llm.png", dpi=110)
print(f"\n{sum(ok for _, ok in report)}/{len(report)} checks passed")
assert all(ok for _, ok in report)
[B-agent] the wrapper runs the model's text output as Python
prompt 'sum' -> model emits the text 'print(1+3)'
wrapper runs it and captures '4'; next prompt: ['sum', 'print(1+3)', '4']
[ok] the emitted text compiles as Python source code
[ok] running the emitted code yields the task's answer
[ok] a different task token yields different code
[ok] the captured result is appended to the next prompt
14/14 checks passed

B-agent writes when the script runs