SI Glossary · Training & data
Fine-Tuning
Fine-tuning takes a general pretrained model and teaches it something more specific: following instructions, writing in a company’s voice, classifying insurance claims, or speaking a medical specialty’s jargon.
Common types
- Supervised fine-tuning (SFT): training on example inputs paired with ideal outputs, such as thousands of well-written assistant replies.
- Preference tuning: training on which of two answers people prefer, as in RLHF.
- Parameter-efficient fine-tuning (e.g., LoRA): adjusting a small add-on set of weights instead of the whole model, which is much cheaper.
Fine-tuning vs alternatives
| Need | Often best approach |
|---|---|
| Model needs up-to-date or private facts | Retrieval (RAG) |
| Consistent format, tone or narrow skill | Fine-tuning |
| Quick behaviour change | Prompt engineering |
Who can fine-tune
With open-weights models, anyone with suitable hardware can fine-tune. Some closed-model providers offer fine-tuning through their APIs. Thinking Machines Lab’s Tinker service (2025) was built specifically to make fine-tuning open models easier.
A safety caveat
Fine-tuning can also remove safety training from open models. That is one reason the release of very capable open weights is debated.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.