Skip to content
SI.info

SI Glossary · Using SI

Retrieval-Augmented Generation (RAG)

Published 1 min read
On this page
  1. How RAG works
  2. Why use it
  3. RAG in the million-token era

Retrieval-augmented generation (RAG) gives a model an “open book”. Before answering, the system searches a collection of documents, such as a company wiki, product manuals, legal files or the web, retrieves the most relevant passages, and adds them to the model’s prompt. The model then answers using that material, often with citations.

How RAG works

  1. Index: documents are split into chunks and converted into embeddings, numerical representations of meaning, stored in a vector database.
  2. Retrieve: a user’s question is embedded too, and the closest-matching chunks are found (often combined with keyword search).
  3. Generate: the question and retrieved passages go to the LLM, which writes a grounded answer.

Why use it

  • Fresh facts: works with information newer than the model’s training data.
  • Private data: answers from your documents without retraining the model.
  • Fewer hallucinations, and answers can cite their sources.
  • Cheaper than fine-tuning for knowledge that changes.

RAG in the million-token era

With context windows of a million tokens, some tasks can skip retrieval and send everything at once. RAG remains the standard for large or frequently changing knowledge bases, and for keeping costs down.

← Back to the SI Glossary

Sources

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., Facebook AI Research, 2020

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy