How this book is built

The idea the dictionary is organized around, the practices applied to each entry, and the checks run over it.

The basic ML picture

The book's Introduction, rendered from its LaTeX source — the idea every entry elaborates.

The basic principle of machine learning (ML) is to fit a hypothesis $\hypothesis$ from a hypothesis space $\hypospace$ to a dataset $\dataset = \{ (\featurevec^{(\sampleidx)}, \truelabel^{(\sampleidx)}) \}_{\sampleidx=1}^{\samplesize}$, choosing the $\hypothesis \in \hypospace$ that minimizes the average loss on the data (Fig. 1). The core challenge is generalization: the learned hypothesis must perform well on new, unseen data points, not just on the training set.

Figure 1 of the entry intro
Figure 1: A dataset $\dataset = \{ (\feature^{(\sampleidx)}, \truelabel^{(\sampleidx)}) \}_{\sampleidx=1}^{5}$ (filled circles) and two hypotheses $\hypothesis, \hypothesis' \in \hypospace$. The vertical arrows show the loss incurred at data point $(\feature^{(3)}, \truelabel^{(3)})$.

The sections of this dictionary offer different levels of detail around this idea: from core concepts and algorithms, through federated learning, reinforcement learning, explainability, and ML systems, to the formal guarantees of statistical learning theory and the underlying mathematics. Entries on ML regulation and healthcare address domain-specific context.

The notation these entries share is collected in the List of symbols.

From one LaTeX entry to every output

Each term is one \newglossaryentry in a LaTeX chapter file of the book, and that entry is the single source of everything the term appears in. Its description is the definition: prose with every technical noun linked to its own entry by \gls, displayed equations written with the book's macros (one macro per quantity, so the feature vector, the training set and the learning rate render the same way in every entry), TikZ or pgfplots figures, citations into one shared bibliography, an optional block of synonyms, and a closing "See also" list. Its abstract is a self-contained summary of four to eight sentences whose first sentence is written to be read alone. The remaining fields fix how the term renders when another entry links it: the full form on first use, the abbreviation afterwards, and the plural.

\newglossaryentry{key}
{name={full name (ABBR)},
    description={First sentence defines the term. ... \gls{otherterm} ...
        \\
        See also: \gls{relatedterm1}, \gls{relatedterm2}.},
    abstract={Four to eight sentences, the first a standalone definition.},
    first={full name (ABBR)}, text={ABBR},
    type=CHAPTER_TYPE
}

From that source, scripts derive every output without a second copy of the text. The book itself is typeset by LuaLaTeX with the glossaries package. export_terms.py typesets each entry as a standalone publication — title, abstract, definition, cross references, the reference list resolved from the shared bibliography — and skips entries whose source has not changed, so a revision rebuilds one PDF. This site is produced by export_term_site.py, which converts the same LaTeX to HTML: prose and links directly, mathematics for MathJax, figures compiled by LuaLaTeX into SVG and cached by content, and the reference list into a numbered bibliography with DOIs. The same run writes the retrieval surfaces — llms.txt, a JSON export of all terms, and schema.org metadata on each page — and the notebooks and the LaTeX macro package of the public companion repository.

Where an entry states something that can be computed, a Python script named after the term computes it: fixed seeds, NumPy and Matplotlib only, its results written to CSV files that the entry's pgfplots figures read directly, so a figure in the PDF and its preview on the term page show the same numbers. The script is cut into blocks that the term page publishes as notebook cells beside the paragraphs they back. Every revision of an entry is one commit that names the term, so the history of a definition is one search of the log, and reviews arrive as annotated PDFs whose marks are extracted and answered one by one.

Evidence for the statements

A statement in an entry is backed by a citation, by a figure or derivation in the entry, or by a numerical experiment. Where it rests on an experiment, the script is published beside the term, so the computation can be repeated. A checker, check_claims, reads the entry sentence by sentence and lists statements for which it finds none of these. A second pass of the same checker asks of every sentence what it means precisely, and flags wrong quantifiers, glosses that conflict with the formal definition, and referents that are never named. A third pass reads the cited works themselves: it retrieves the passages of the source closest to the sentence and flags a claim the retrieved passages do not carry. Each list is a working list: it records what has not been backed yet, not a guarantee that nothing remains.

Citation pinpoints

A citation names a location in a work — a section, a numbered result, an example. A citation can name the right work and the wrong location in it, which a check of the bibliography does not detect. A checker, check_ref_accuracy, therefore locates the pinpoint in the text of the source, and the passage found there is compared with the statement it supports. Two corrections made this way: an identity displayed in Sect. 2.3.1 of Boyd and Vandenberghe had been cited to 2.5.1, where it is proved rather than stated; a definition of completeness had been cited to the Hilbert-space chapter of Bauschke and Combettes, which assumes it, rather than to Sect. 1.12, which states it. A complementary checker, check_canonical_refs, asks the question the others cannot: it names the most authoritative works for the concept — the paper that coined the term, the standard textbook, a survey — and reports which of them the entry does not cite.

Automated checks

More than forty checkers run over an entry, and one command runs the whole suite and renders the findings as one page per term, defects first. Every checker, every house rule and the screen that enforces it, are listed on the linters page, generated from the project's manifest at each build. Some checkers are mechanical scans of the LaTeX source: check_articles flags an article that disagrees with the rendered form of an abbreviation ("a LLM"), and re-reads the typeset PDF, where the first use of a term expands its abbreviation; check_plurals flags a plural inconsistent with its term; check_figures a figure without a caption or never referenced, a boxed legend, a gridded plot; check_macros and check_raw_notation notation written out where a macro exists; check_margins text overflowing the margin; check_bullshit filler vocabulary, vague intensifiers, and comparisons whose second term is never named; check_circular a sentence that argues for what the text already granted; check_types a property the book does not define for that object ("the hypothesis space is too rich"). Others read the prose with a language model: check_jargon applies a checkable-claim test to metaphors ("workhorse", "under the hood"); check_figures, in its second pass, judges grayscale renders of every figure, where color must never be the only channel separating two elements. Most checkers are advisory — they report candidates, and the decision stays with the author.

Composition

Three checkers judge quality rather than correctness, reviewing each entry against named doctrines of technical writing. check_bertsekas applies Bertsekas's rules for mathematical writing: organize in segments, keep notation and nomenclature consistent, give examples and counterexamples, visualize. check_winston applies Winston's heuristics for memorable teaching: one key idea that recurs, a fence against confusable neighbors, a near miss, and a handle — the symbol, slogan, or story by which a reader retrieves the entry. check_mcenerney applies McEnerney's doctrine of reader value: name the problem the entry resolves, and its cost, before the solution. Two lexical siblings, check_register and check_pruning, screen the register (padding constructions, stance adverbs, light-verb periphrasis) and prune sentences whose deletion loses nothing, with a stricter standard for abstracts, where every sentence must carry an essential point. These screens are the most deferential in the pipeline: the doctrines are adapted to the entry form (an entry has no introduction in which to preview itself, and cross-references stand in for hierarchical development), and house style overrides them where they conflict — the publisher's impersonal voice wins over Bertsekas's active-voice advice. What survives the adaptation is the useful core: a reviewer that asks, for every entry, what the reader gains, which neighboring concept they might confuse it with, and which sentence they will remember it by.

Terminology

One word per concept. Where the literature uses one word for two concepts, the entries separate them: a feature transformation maps the feature vector of a data point to another feature vector, while a feature map is the grid of activations produced by one channel of a convolutional layer. Both names are standard in their own literature, which is why the distinction is recorded. Two checkers hold the line: check_wording_consistency flags an entry that mixes variants of one concept (sample, instance and observation for data point), and check_word_repetition the same word repeated within a few lines — except for a term's own name, which is repeated deliberately instead of varied.

Licence

The course edition is licensed CC BY 4.0. It is distinct from the Dictionary of Applied Machine Learning to be published by Springer. Each term page carries a citation block, a typeset PDF, and the Python script for that term.

Back to the index