← Home · Tutorial 06 · How LLMs work

How a transformer works

From text to the next token, and how attention lets words take meaning from each other.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where this comes from: the architecture is from the original paper, "Attention Is All You Need" (Vaswani et al., 2017). The explanation, analogies and code are my own, simplified on purpose. The numbers in the code are random, so they show the mechanics and not real language knowledge. Real models differ in many details.
Who this is for: anyone who uses LLMs and wants a working picture of what happens inside. You will learn the path from text to the next token, and how attention lets a word take its meaning from the words around it. No maths background needed.

1Purpose first

Why a rough picture is worth having.

You do not need the maths to use an LLM. You do need a picture of what happens between your text going in and a reply coming out, because it explains why models behave the way they do: why they count badly, why long inputs cost more, why wording matters.

Every large language model you have used is built on the transformer. This page walks the pipeline in order, then zooms in on its key step, attention.

2The pipeline

From text to the next token.

The pipeline
Textyour prompt
→
Tokenizertext to pieces
→
Embeddingspieces to numbers
→
+ Positionwhere each piece sits
→
Attention layerspieces look at each other
→
Next tokena probability for each

The model then picks one token, adds it to the text, and runs the whole thing again. A reply is built one token at a time.

3Tokenizer

Text becomes pieces.

A model does not read letters or words. It reads tokens: pieces from a fixed vocabulary, often word parts. The tokenizer splits your text into the longest pieces it knows. This toy version shows the idea, and real tokenizers learn their vocabulary from data.

Python
vocab = ["un", "believ", "able", "u", "n", "b"]

def tokenize(text):
    out = []
    while text:
        for n in range(len(text), 0, -1):       # longest match first
            if text[:n] in vocab:
                out.append(text[:n]); text = text[n:]; break
        else:
            out.append(text[0]); text = text[1:]
    return out

print(tokenize("unbelievable"))
print(tokenize("unbelievably"))
Output (run to check)
['un', 'believ', 'able']
['un', 'believ', 'a', 'b', 'l', 'y']

A known word becomes a few tokens. A form the vocabulary lacks falls apart into smaller pieces. That is part of why models can stumble on rare words, spelling and counting letters.

Why you care: you pay and are limited by tokens, not words. A rough rule for English is that a token is about four characters, and other languages can need more.

4Embeddings

Pieces become numbers.

Each token is looked up in a table and replaced by a list of numbers, its embedding. Tokens with related meanings end up with similar numbers, which is what lets the model treat "invoice" and "bill" as near neighbours.

The same idea powers search by meaning in retrieval systems, where whole passages are turned into embeddings and compared.

5Positional encoding

Telling the model the order.

Attention looks at all tokens at once, so by itself it has no sense of order. "Dog bites man" and "man bites dog" would look the same. Positional encoding fixes that by adding a position signal to each token's numbers.

Python
import math
def pos_enc(pos, d=4):
    return [round(math.sin(pos / 10000 ** (i / d)) if i % 2 == 0
                  else math.cos(pos / 10000 ** ((i - 1) / d)), 3) for i in range(d)]
for p in range(3):
    print(p, pos_enc(p))
Output (run to check)
0 [0.0, 1.0, 0.0, 1.0]
1 [0.841, 0.54, 0.01, 1.0]
2 [0.909, -0.416, 0.02, 1.0]

This is the original sine and cosine scheme. Each position gets its own pattern. Many newer models use other position methods, but the purpose is the same: tell the model where each token sits.

6Same word, different attention

The core idea with an example.

This is the step that made transformers work. Take the same word in two sentences:

"She sat by the river bank." and "He went to the bank to deposit cash."

"Bank" is the same string both times. To know which meaning is intended, you look at the other words: "river" in the first, "deposit" and "cash" in the second. Attention is the machinery for doing that look, for every word at once. Each token decides how much every other token matters to it, then takes a weighted mix of what it finds.

7Queries, keys and values

The mechanism, with a runnable version.

Each token is turned into three lists of numbers:

NameThink of it asRole
Query (Q)What I am looking forAsked by the current token
Key (K)What I offerCompared against every query
Value (V)What I contributeMixed in when the key matches

The steps: compare each query with every key to get a score, turn the scores into proportions with softmax, then use those proportions to blend the values. A library analogy fits: the query is your search, keys are book titles, values are book contents.

One attention step
Q · Kmatch scores
→
Softmaxscores to proportions
→
Blend Vweighted mix
Python (numpy)
import numpy as np
def softmax(x):
    e = np.exp(x - x.max(-1, keepdims=True))
    return e / e.sum(-1, keepdims=True)

rng = np.random.default_rng(0)
X = rng.normal(size=(5, 4))                  # 5 tokens, 4 numbers each
Wq, Wk, Wv = (rng.normal(size=(4, 4)) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv

scores  = Q @ K.T / np.sqrt(4)               # how well each token matches each other token
weights = softmax(scores)                    # each row becomes proportions that sum to 1
output  = weights @ V                        # each token becomes a weighted mix of the values

print("weights shape:", weights.shape)
print("each row sums to:", weights.sum(axis=1).round(3))
print("output shape:", output.shape)
Output (run to check)
weights shape: (5, 5)
each row sums to: [1. 1. 1. 1. 1.]
output shape: (5, 4)

Because the numbers are random, the weights mean nothing linguistically. Take the shape from it: five tokens in, a 5-by-5 table of weights out (every token against every token), each row adding to 1, and five blended outputs.

8Multi-head attention

Several lookups in parallel.

One set of weights can only capture one kind of relationship. Multi-head attention runs several attention steps side by side, each with its own Q, K and V, and joins the results. One head might track grammar, another might link a pronoun to its noun. Nobody assigns those jobs. They emerge in training.

Models stack many of these layers, so meaning is refined again and again.

A caution: talk of what each head "means" is an intuition from inspecting trained models, not a rule. Heads are often messy.

9BERT, GPT and BART

One block, three arrangements.

VariantAttentionGood atExamples
Encoder onlyEach token sees tokens on both sidesUnderstanding: classification, search embeddingsBERT family
Decoder onlyEach token sees only earlier tokensGenerating text, one token at a timeGPT and Claude style chat models
Encoder plus decoderEncoder reads all; decoder writes and also attends to the encoderTranslation, summarisingT5, BART

Different masks and arrangements, the same block. That is why one idea scaled into a whole family of models. The chat models you use day to day are decoder-only: trained to predict the next token, with chat behaviour built on top.

10What this explains

Behaviour you can now predict, and a check.

Next-token prediction is not looking things up. The model produces what is likely to come next given its training and your text. That is why it can sound right and be wrong.
Context cost grows. Because every token can attend to every other, longer inputs cost more to process.
Why do models need positional encoding?

Attention sees all tokens at once with no order, so position has to be added to each token explicitly.

What do query, key and value do?

The query is what a token is looking for, keys are compared with it to give scores, and values are blended according to those scores.

Why multiple heads?

Each head can pick up a different relationship between tokens, and the results are combined.

What is the difference between BERT and GPT attention?

BERT looks both ways. GPT looks only backwards, so it can generate text.

Try it: run the attention snippet above with a different random seed or more tokens. Check that every row of weights still adds to 1, and that the table grows to match the number of tokens.
One line to remember: "An LLM turns text into tokens, tokens into numbers, lets every token weigh every other, and predicts the next token, over and over."
Based on Attention Is All You Need (Vaswani et al., 2017). The tokenizer and attention code here are simplified teaching versions, not how a production model is implemented. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.