{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "name": "python3",
   "display_name": "Python 3",
   "language": "python"
  },
  "language_info": {
   "name": "python"
  },
  "colab": {
   "name": "llm.ipynb"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "# large language model (LLM) \u2014 Python demo\n\nNumerical companion to the entry [large language model (LLM)](https://dictionaryofml.org/terms/llm.html) of the [Dictionary of Applied Machine Learning](https://dictionaryofml.org/): it recomputes what the entry states and prints one line per check.\n\nBlocks verify numerically what the entry's statements assert. Self-contained (numpy/matplotlib only), fixed seed. (The entry's scale claim \u2014 billions of parameters \u2014 is illustrated in miniature: the mechanisms, not the size.)\n\nRequires NumPy and Matplotlib only, and uses fixed seeds, so the printed numbers reproduce exactly. Generated from [`pythondemos/llm.py`](https://dictionaryofml.org/terms/llm.py); CC BY 4.0."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "# Notebook shim: the script resolves output paths relative to __file__,\n# which a notebook kernel does not define; everything lands in the\n# working directory instead.\nimport os\n__file__ = os.path.join(os.getcwd(), \"llm.py\")\nos.makedirs(\"pythondemos\", exist_ok=True)"
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "\"\"\"\nllm.py \u2014 numerical companion to the glossary entry\n'large language model (LLM)'.\n\nBlocks verify numerically what the entry's statements assert. Self-contained\n(numpy/matplotlib only), fixed seed. (The entry's scale claim \u2014\nbillions of parameters \u2014 is illustrated in miniature: the mechanisms,\nnot the size.)\n\nBlocks\n------\n[B-selfsup]   Self-supervised construction of data points from\n              raw text: masking words turns an unannotated corpus into\n              (context features, masked-word label) pairs -- one data\n              point per position, with zero human annotation effort.\n[B-train]     Training via ERM on these pairs: a small next-token model\n              (embedding + softmax over the vocabulary) trained by\n              GD on the training loss (the negative log probability\n              of the correct next token) drives the training loss\n              down and beats the uniform-guess baseline.\n[B-nexttoken] A trained LLM maps an input token sequence to a\n              probability distribution over the next token: outputs are\n              nonnegative, sum to one, and the model assigns the\n              highest probability to continuations seen in the corpus;\n              sampling from the distribution generates text.\n\n[B-local]     Local execution in miniature: quantizing the model\n              parameters to 8-bit integers shrinks the memory to an\n              eighth while the next-token probabilities barely move\n              and the predicted next token agrees on 90% of contexts.\n[B-agent]     One step of an LLM agent in miniature: the same\n              architecture trained on task--code pairs emits, for the\n              task token 'sum', the token 'print(1+3)'; the wrapper\n              compiles the emitted text as Python source code, runs\n              it, captures the output '4', and appends it to the next\n              prompt. The model only ever outputs text -- the wrapper\n              is what turns the text into an action.\n\nOutputs\n-------\npythondemos/llm.png : preview figure (checking only).\n\nData generated by pythondemos/llm.py.\n\"\"\"\n\nfrom pathlib import Path\n\nimport numpy as np\nimport matplotlib\n\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\n\nrng = np.random.default_rng(42)\nreport = []\n\n\ndef check(name, ok):\n    report.append((name, bool(ok)))\n    print(f\"  [{'ok' if ok else 'FAIL'}] {name}\")\n\n\ncorpus = (\"all human beings are born free and equal in dignity and \"\n          \"rights all human beings are endowed with reason and \"\n          \"conscience\").split()\nvocab = sorted(set(corpus))\nV = len(vocab)\ntok = {w: i for i, w in enumerate(vocab)}\nids = np.array([tok[w] for w in corpus])"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-selfsup]** Self-supervised construction of data points from raw text: masking words turns an unannotated corpus into (context features, masked-word label) pairs -- one data point per position, with zero human annotation effort."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-selfsup] masked words become labels, contexts become features\")\npairs = [(ids[t], ids[t + 1]) for t in range(len(ids) - 1)]\ncheck(\"one labeled pair per corpus position (no human annotation)\",\n      len(pairs) == len(corpus) - 1)\ncheck(\"labels are drawn from the text itself\",\n      all(0 <= y < V for _, y in pairs))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-train]** Training via ERM on these pairs: a small next-token model (embedding + softmax over the vocabulary) trained by GD on the training loss (the negative log probability of the correct next token) drives the training loss down and beats the uniform-guess baseline."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-train] ERM on the constructed pairs\")\nd_emb = 8\nE = 0.1 * rng.normal(size=(V, d_emb))              # embeddings\nU = 0.1 * rng.normal(size=(d_emb, V))              # unembedding\ndef forward(x_ids):\n    logits = E[x_ids] @ U\n    p = np.exp(logits - logits.max(1, keepdims=True))\n    return p / p.sum(1, keepdims=True)\nxs = np.array([x for x, _ in pairs])\nys = np.array([y for _, y in pairs])\ndef xent():\n    return -np.mean(np.log(forward(xs)[np.arange(len(ys)), ys] + 1e-12))\nloss0 = xent()\nfor _ in range(800):                               # GD\n    P = forward(xs)\n    G = P.copy(); G[np.arange(len(ys)), ys] -= 1.0\n    gU = E[xs].T @ G / len(ys)\n    gE = np.zeros_like(E)\n    np.add.at(gE, xs, G @ U.T / len(ys))\n    U -= 2.0 * gU; E -= 2.0 * gE\nloss1 = xent()\nprint(f\"    training loss: init {loss0:.2f} -> trained {loss1:.2f} \"\n      f\"(uniform baseline {np.log(V):.2f})\")\ncheck(\"training reduces the training loss\", loss1 < loss0)\ncheck(\"the trained model beats the uniform-guess baseline log|V|\",\n      loss1 < np.log(V) - 0.5)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-nexttoken]** A trained LLM maps an input token sequence to a probability distribution over the next token: outputs are nonnegative, sum to one, and the model assigns the highest probability to continuations seen in the corpus; sampling from the distribution generates text."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-nexttoken] input sequence -> distribution over the next token\")\np_next = forward(np.array([tok[\"human\"]]))[0]\ncheck(\"the output is a probability distribution (nonneg, sums to 1)\",\n      np.all(p_next >= 0) and np.isclose(p_next.sum(), 1.0))\ncheck(\"'human' is followed by 'beings' in the corpus \u2014 and gets the \"\n      \"highest next-token probability\",\n      vocab[int(np.argmax(p_next))] == \"beings\")\ngen = [tok[\"all\"]]\nfor _ in range(5):                                 # sample a continuation\n    gen.append(int(rng.choice(V, p=forward(np.array([gen[-1]]))[0])))\ncheck(\"sampling from the distributions generates a token sequence\",\n      len(gen) == 6 and all(0 <= g < V for g in gen))\nprint(\"    generated:\", \" \".join(vocab[g] for g in gen))"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-local]** Local execution in miniature: quantizing the model parameters to 8-bit integers shrinks the memory to an eighth while the next-token probabilities barely move and the predicted next token agrees on 90% of contexts."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-local] the same model, quantized to 8 bits, runs locally\")\nsE = np.max(np.abs(E)) / 127.0\nsU = np.max(np.abs(U)) / 127.0\nEq = (np.round(E / sE).astype(np.int8) * sE)       # 8-bit stored, dequantized\nUq = (np.round(U / sU).astype(np.int8) * sU)\ndef forward_q(x_ids):\n    logits = Eq[x_ids] @ Uq\n    p = np.exp(logits - logits.max(1, keepdims=True))\n    return p / p.sum(1, keepdims=True)\nagree = float(np.mean(np.argmax(forward_q(xs), 1) == np.argmax(forward(xs), 1)))\ngap = float(np.max(np.abs(forward_q(xs) - forward(xs))))\nmem = (E.nbytes + U.nbytes) / (E.size + U.size)    # bytes per parameter\nprint(f\"    {E.size + U.size} model parameters at {mem:.0f} bytes each -> \"\n      f\"1 byte each after quantization (memory shrinks {mem:.0f}x)\")\nprint(f\"    next-token probabilities move by at most {gap:.4f}; the \"\n      f\"predicted next token agrees on {100 * agree:.0f}% of contexts \"\n      f\"(the rest are corpus-ambiguous near-ties)\")\ncheck(\"8-bit quantization shrinks the memory to an eighth\", mem == 8.0)\ncheck(\"next-token probabilities move by less than 0.01\", gap < 0.01)\ncheck(\"the predicted next token agrees on at least 85% of contexts\",\n      agree >= 0.85)"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "**[B-agent]** One step of an LLM agent in miniature: the same architecture trained on task--code pairs emits, for the task token 'sum', the token 'print(1+3)'; the wrapper compiles the emitted text as Python source code, runs it, captures the output '4', and appends it to the next prompt. The model only ever outputs text -- the wrapper is what turns the text into an action."
  },
  {
   "cell_type": "code",
   "metadata": {},
   "execution_count": null,
   "outputs": [],
   "source": "print(\"[B-agent] the wrapper runs the model's text output as Python\")\nimport contextlib\n# tiny task->code corpus: after each task token, the code token follows\ncorpus_a = (\"sum print(1+3) done double print(2*3) done \"\n            \"sum print(1+3) done double print(2*3) done\").split()\nvocab_a = sorted(set(corpus_a))\ntok_a = {w: i for i, w in enumerate(vocab_a)}\nids_a = np.array([tok_a[w] for w in corpus_a])\nVa = len(vocab_a)\nEa = 0.1 * rng.normal(size=(Va, d_emb))\nUa = 0.1 * rng.normal(size=(d_emb, Va))\nxa = ids_a[:-1]\nya = ids_a[1:]\ndef forward_a(x_ids):\n    logits = Ea[x_ids] @ Ua\n    pr = np.exp(logits - logits.max(1, keepdims=True))\n    return pr / pr.sum(1, keepdims=True)\nfor _ in range(800):                               # GD\n    Pa = forward_a(xa)\n    Ga = Pa.copy(); Ga[np.arange(len(ya)), ya] -= 1.0\n    gUa = Ea[xa].T @ Ga / len(ya)\n    gEa = np.zeros_like(Ea)\n    np.add.at(gEa, xa, Ga @ Ua.T / len(ya))\n    Ua -= 2.0 * gUa; Ea -= 2.0 * gEa\nimport io as _io\ntranscript = [\"sum\"]                               # the prompt: a task\nemitted = vocab_a[int(np.argmax(forward_a(np.array([tok_a[\"sum\"]]))[0]))]\nprint(f\"    prompt 'sum' -> model emits the text '{emitted}'\")\ncode_ok = True\ntry:\n    compiled = compile(emitted, \"<llm-output>\", \"exec\")\nexcept SyntaxError:\n    code_ok = False\nbuf = _io.StringIO()\nif code_ok:\n    with contextlib.redirect_stdout(buf):\n        exec(compiled)                             # the wrapper acts\nresult = buf.getvalue().strip()\ntranscript += [emitted, result]                    # result -> next prompt\nprint(f\"    wrapper runs it and captures '{result}'; \"\n      f\"next prompt: {transcript}\")\nemitted2 = vocab_a[int(np.argmax(forward_a(np.array([tok_a[\"double\"]]))[0]))]\ncheck(\"the emitted text compiles as Python source code\", code_ok)\ncheck(\"running the emitted code yields the task's answer\",\n      result == \"4\")\ncheck(\"a different task token yields different code\",\n      emitted2 == \"print(2*3)\" and emitted2 != emitted)\ncheck(\"the captured result is appended to the next prompt\",\n      transcript == [\"sum\", \"print(1+3)\", \"4\"])\n\n# ------------------------------------------------------------ preview\nfig, ax = plt.subplots(figsize=(5.6, 3.0))\nax.bar(range(V), p_next)\nax.set_xticks(range(V))\nax.set_xticklabels(vocab, rotation=90, fontsize=6)\nax.set_xlabel(\"token\")\nax.set_ylabel(\"P(next token | 'human')\")\nax.set_title(\"[B-nexttoken] next-token distribution\")\nfig.tight_layout()\nOUT_DIR = Path(__file__).parent\nfig.savefig(OUT_DIR / \"llm.png\", dpi=110)\nprint(f\"\\n{sum(ok for _, ok in report)}/{len(report)} checks passed\")\nassert all(ok for _, ok in report)"
  }
 ]
}