1Purpose first
Why a rough picture is worth having.
You do not need the maths to use an LLM. You do need a picture of what happens between your text going in and a reply coming out, because it explains why models behave the way they do: why they count badly, why long inputs cost more, why wording matters.
Every large language model you have used is built on the transformer. This page walks the pipeline in order, then zooms in on its key step, attention.
2The pipeline
From text to the next token.
The model then picks one token, adds it to the text, and runs the whole thing again. A reply is built one token at a time.
3Tokenizer
Text becomes pieces.
A model does not read letters or words. It reads tokens: pieces from a fixed vocabulary, often word parts. The tokenizer splits your text into the longest pieces it knows. This toy version shows the idea, and real tokenizers learn their vocabulary from data.
vocab = ["un", "believ", "able", "u", "n", "b"]
def tokenize(text):
out = []
while text:
for n in range(len(text), 0, -1): # longest match first
if text[:n] in vocab:
out.append(text[:n]); text = text[n:]; break
else:
out.append(text[0]); text = text[1:]
return out
print(tokenize("unbelievable"))
print(tokenize("unbelievably"))['un', 'believ', 'able'] ['un', 'believ', 'a', 'b', 'l', 'y']
A known word becomes a few tokens. A form the vocabulary lacks falls apart into smaller pieces. That is part of why models can stumble on rare words, spelling and counting letters.
4Embeddings
Pieces become numbers.
Each token is looked up in a table and replaced by a list of numbers, its embedding. Tokens with related meanings end up with similar numbers, which is what lets the model treat "invoice" and "bill" as near neighbours.
The same idea powers search by meaning in retrieval systems, where whole passages are turned into embeddings and compared.
5Positional encoding
Telling the model the order.
Attention looks at all tokens at once, so by itself it has no sense of order. "Dog bites man" and "man bites dog" would look the same. Positional encoding fixes that by adding a position signal to each token's numbers.
import math
def pos_enc(pos, d=4):
return [round(math.sin(pos / 10000 ** (i / d)) if i % 2 == 0
else math.cos(pos / 10000 ** ((i - 1) / d)), 3) for i in range(d)]
for p in range(3):
print(p, pos_enc(p))0 [0.0, 1.0, 0.0, 1.0] 1 [0.841, 0.54, 0.01, 1.0] 2 [0.909, -0.416, 0.02, 1.0]
This is the original sine and cosine scheme. Each position gets its own pattern. Many newer models use other position methods, but the purpose is the same: tell the model where each token sits.
6Same word, different attention
The core idea with an example.
This is the step that made transformers work. Take the same word in two sentences:
"Bank" is the same string both times. To know which meaning is intended, you look at the other words: "river" in the first, "deposit" and "cash" in the second. Attention is the machinery for doing that look, for every word at once. Each token decides how much every other token matters to it, then takes a weighted mix of what it finds.
7Queries, keys and values
The mechanism, with a runnable version.
Each token is turned into three lists of numbers:
| Name | Think of it as | Role |
|---|---|---|
| Query (Q) | What I am looking for | Asked by the current token |
| Key (K) | What I offer | Compared against every query |
| Value (V) | What I contribute | Mixed in when the key matches |
The steps: compare each query with every key to get a score, turn the scores into proportions with softmax, then use those proportions to blend the values. A library analogy fits: the query is your search, keys are book titles, values are book contents.
import numpy as np
def softmax(x):
e = np.exp(x - x.max(-1, keepdims=True))
return e / e.sum(-1, keepdims=True)
rng = np.random.default_rng(0)
X = rng.normal(size=(5, 4)) # 5 tokens, 4 numbers each
Wq, Wk, Wv = (rng.normal(size=(4, 4)) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv
scores = Q @ K.T / np.sqrt(4) # how well each token matches each other token
weights = softmax(scores) # each row becomes proportions that sum to 1
output = weights @ V # each token becomes a weighted mix of the values
print("weights shape:", weights.shape)
print("each row sums to:", weights.sum(axis=1).round(3))
print("output shape:", output.shape)weights shape: (5, 5) each row sums to: [1. 1. 1. 1. 1.] output shape: (5, 4)
Because the numbers are random, the weights mean nothing linguistically. Take the shape from it: five tokens in, a 5-by-5 table of weights out (every token against every token), each row adding to 1, and five blended outputs.
8Multi-head attention
Several lookups in parallel.
One set of weights can only capture one kind of relationship. Multi-head attention runs several attention steps side by side, each with its own Q, K and V, and joins the results. One head might track grammar, another might link a pronoun to its noun. Nobody assigns those jobs. They emerge in training.
Models stack many of these layers, so meaning is refined again and again.
9BERT, GPT and BART
One block, three arrangements.
| Variant | Attention | Good at | Examples |
|---|---|---|---|
| Encoder only | Each token sees tokens on both sides | Understanding: classification, search embeddings | BERT family |
| Decoder only | Each token sees only earlier tokens | Generating text, one token at a time | GPT and Claude style chat models |
| Encoder plus decoder | Encoder reads all; decoder writes and also attends to the encoder | Translation, summarising | T5, BART |
Different masks and arrangements, the same block. That is why one idea scaled into a whole family of models. The chat models you use day to day are decoder-only: trained to predict the next token, with chat behaviour built on top.
10What this explains
Behaviour you can now predict, and a check.
Why do models need positional encoding?
Attention sees all tokens at once with no order, so position has to be added to each token explicitly.
What do query, key and value do?
The query is what a token is looking for, keys are compared with it to give scores, and values are blended according to those scores.
Why multiple heads?
Each head can pick up a different relationship between tokens, and the results are combined.
What is the difference between BERT and GPT attention?
BERT looks both ways. GPT looks only backwards, so it can generate text.