Being a candid mechanical account of the Large Language Model — what truly happens between your question and its answer — rendered in brass, in the manner of Mr. Charles Babbage.
Everything below actually runs: the gauges are moved by a real (if miniature) engine of probability computed live on this page. The plaques state, candidly, how the full-sized article differs — and terms under a dotted line yield their modern meaning to the pointer.
To operate: 1 — write a few opening words upon the card, or take a ready-punched card from the rack; 2 — pull Begin the Demonstration and follow the works, pressing Proceed at each station; 3 — thereafter Turn the Crank for each further word, or let the engine Run. Mind the boiler.
Your words are punched upon a card and fed to the reader. The engine does not hear you — it receives marks.
In truth — a language model receives nothing but text: no voice, no view of the world, no notion of who is asking beyond what the words themselves carry. Everything that follows is determined by these characters and by the settings of the machine’s parts.
A note on the manner of address. To a chat engine you write whole questions; this cabinet asks instead for a few opening words to continue. There is no contradiction — a chat engine is a continuation engine too: the clerks lay the conversation out as a transcript (Operator: … Engine: —) and the machine simply continues the document, which lands, every time, precisely upon the engine’s reply. The instructions that govern its manners — the system prompt — are more cards still, read before yours.
The card is read and matched to slugs — standard pieces from the engine’s type-case. Each slug bears a catalogue number, and from here onward the engine handles numbers only.
In truth — text is broken into tokens from a fixed catalogue of roughly 50,000–200,000 pieces. Common words are one token; rarer words are cut into fragments (un·fathom·able). The machine never sees letters or words as you do — only these numbers.
This cabinet’s catalogue holds — slugs; a slug marked ✶ was cut fresh and is a stranger to it. The rack holds only 24 slugs — its context window. Push in more and the oldest fall off the far end, forgotten. (A modern engine’s rack holds 200,000 or more.)
In truth — this drawer is the model’s entire vocabulary. Every reply it will ever compose is set from these sorts and no others; a word absent from the drawer cannot be said. A real engine’s case holds 50,000–200,000 sorts — larger, but every bit as closed. Sorts marked ✶ were cut fresh this session; sorts aglow are out at the composing line.
Each catalogue number is looked up in the Great Ledger, and the slug is exchanged for a column of dial-settings — its position upon the map of meaning.
In truth — every token becomes an embedding: a list of thousands of numbers (this cabinet shows eight dials; a large engine uses more than ten thousand). Nobody wrote the ledger by hand — the settings were learned, and they place words used in similar ways at similar settings. Meaning, to the machine, is position.
Here is the celebrated mechanism. Every slug is joined to every earlier slug by an adjustable linkage; the engine draws some taut and lets others hang slack. A word discovers what it means by choosing what to regard.
In truth — this is attention, the invention (2017) that makes these engines possible. For each position the model computes how much every earlier token should bear upon it — how it finds its noun, how bank leans toward the river or the counting-house. Dozens of heads work in parallel, each minding a different sort of relation: shown here, a nickel head that favours the recent, a copper head that favours old acquaintances.
Depicted schematically — in the full engine these weights are computed afresh, at every layer, from the dial-settings themselves.
The columns now descend through the Mill — floor after identical floor of linkage-and-gearwork, each refining the settings handed down by the last.
In truth — attention plus a small calculating network make one layer, and layers are stacked by the dozen (a large model, near a hundred). A token’s numbers are revised at every floor — coarse matters of spelling and grammar low down, subtler matters of sense and intent above, or so the anatomists report. The adjustable parts — the parameters — number in the hundreds of billions, and none was set by hand: in training, the engine read a vast library, guessed each next token, and every screw was turned a hair’s-breadth whenever it guessed wrong. Repeated a very great many times, that is the whole of its education. A finishing school follows — training on human preference — to make it helpful, honest, and polite.
This cabinet, for comparison, owns — counting-wheels, learned from a library of — sentences.
At the mill’s outfall the engine renders its judgment — not a word, but a pressure upon every gauge in the house: a score for every slug in the catalogue at once.
In truth — the model’s output at each step is a probability for every token in its vocabulary — all 50,000-odd, most of them minuscule. It does not decide the next word; it rates all possible ones. These gauges are live: they show this engine’s true arithmetic for the present context — and they obey the boiler lever at Station VII, should you care to move it.
One slug must be drawn. The lots are cast in proportion to the gauges — and the boiler decides how strictly. Run cold, the favourite is all but certain; run feverish, and outsiders take their chance.
In truth — the next token is chosen by weighted lottery — sampling — and temperature reshapes the odds exactly as shown: low temperature sharpens the distribution toward the favourite (predictable, sometimes repetitious); high temperature flattens it (inventive, and at the extreme, incoherent). This is why the same question, asked twice, may be answered differently. Try the lever — the gauges obey at once.
The chosen slug is stamped upon the tape — and carried straight back to the rack, where it joins your question. The whole engine then turns over again, from card to lottery, to choose the word after. One revolution, one word.
In truth — a language model writes autoregressively: one token per pass, each pass re-reading all that came before. There is no finished answer waiting inside; each word is minted the moment before you read it. And the engine stops as it speaks — by predicting a special stop token (printed here as ∎), whereupon it rests.
The full engine keeps notes — a cache — so as not to redo old arithmetic on every revolution; but in principle, every turn reads the whole rack afresh.