1Purpose first
Why the stages matter to you.
A chat model is not built in one step. Knowing the stages tells you what each one buys you, and which one to reach for when a model is not behaving. It also explains why "just train it on our data" is rarely the right first answer.
2The big picture
Four stages, one line each.
3Pre-training
Learning language from raw text.
The model is trained on a vast amount of text to predict the next token. Nobody labels this data. The text itself supplies the answers, which is why it scales. The result is a base model: fluent, knowledgeable, and not yet helpful. Ask it a question and it may continue with more questions, because that is what text often does.
4Instruction fine-tuning
Learning to follow requests.
Next, the model is trained on examples of instructions paired with good responses, so it learns to follow a request instead of just continuing the text. This is instruction fine-tuning, a form of supervised fine-tuning.
A risk worth knowing: training further on a narrow task can make a model worse at others. This is called catastrophic forgetting. Fine-tuning many tasks together, or using the lighter methods below, reduces it.
5RLHF
Learning what people prefer.
People rate or rank model answers. A reward model learns those preferences, and the language model is then trained to produce answers that score well. This is reinforcement learning from human feedback (RLHF). It is the stage that pushes a model toward being helpful, honest and harmless, and away from toxic or off-target replies.
Newer methods can reach a similar goal with less machinery, so RLHF is best seen as one way to teach preferences, not the only one.
6PEFT
Fine-tuning on a budget.
Updating every weight of a large model needs a lot of memory. Parameter-efficient fine-tuning (PEFT) freezes most of the model and trains a small extra part.
| Method | Idea |
|---|---|
| LoRA | Train small low-rank matrices added alongside the frozen weights |
| Prompt or soft-prompt tuning | Train a few extra learned vectors prepended to the input |
The result is a small add-on you can store and swap per task, and training fits on far less hardware.
7Which lever to pull
Matching the need to the method.
| Need | Try first | Why |
|---|---|---|
| Better instructions followed | Prompting and examples | Free and instant |
| Answers from your documents | Retrieval (RAG) | Facts stay in documents you can update |
| A consistent style or format | Few-shot, then fine-tuning | Fine-tuning fixes behaviour |
| A narrow task, at scale, cheaply | Fine-tuning (PEFT) | Smaller model, lower cost per call |
8Check yourself
Short questions and one line to remember.
What does pre-training produce?
A base model that predicts text well but does not reliably follow instructions.
What does RLHF add?
Alignment with human preferences, so answers are more helpful and less harmful.
Why use PEFT rather than full fine-tuning?
It trains a small part, so it needs far less memory and gives a small, swappable add-on.
Fine-tuning or retrieval for company policy documents?
Retrieval, because the facts change and should stay in documents you can update.