SI Glossary · Using SI
Retrieval-Augmented Generation (RAG)
On this page
Retrieval-augmented generation (RAG) gives a model an “open book”. Before answering, the system searches a collection of documents, such as a company wiki, product manuals, legal files or the web, retrieves the most relevant passages, and adds them to the model’s prompt. The model then answers using that material, often with citations.
How RAG works
- Index: documents are split into chunks and converted into embeddings, numerical representations of meaning, stored in a vector database.
- Retrieve: a user’s question is embedded too, and the closest-matching chunks are found (often combined with keyword search).
- Generate: the question and retrieved passages go to the LLM, which writes a grounded answer.
Why use it
- Fresh facts: works with information newer than the model’s training data.
- Private data: answers from your documents without retraining the model.
- Fewer hallucinations, and answers can cite their sources.
- Cheaper than fine-tuning for knowledge that changes.
RAG in the million-token era
With context windows of a million tokens, some tasks can skip retrieval and send everything at once. RAG remains the standard for large or frequently changing knowledge bases, and for keeping costs down.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., Facebook AI Research, 2020
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.