How RAG works
Retrieval-augmented generation lets a language model answer from your own documents instead of from memory. Scroll to follow a question from the documents to a cited answer.
- 1
Split the documents
Long documents are cut into short chunks, a few paragraphs each, so the system can retrieve just the part that answers a question.
- 2
Turn chunks into vectors
An embedding model turns each chunk into a list of numbers. Chunks with similar meaning land close together in this space.
- 3
Embed the question
The question goes through the same embedding model, so it lands in the same space as the chunks.
- 4
Find the nearest chunks
A vector search returns the chunks closest to the question. It's fast, but closeness is only a rough guess at relevance.
- 5
Rerank
A reranker reads the question and each candidate together and scores how well it actually answers. The order changes.
- 6
Answer with sources
The top chunks go into the prompt. The model answers from them and cites the chunk behind each claim.
Split the documents
Long documents are cut into short chunks, a few paragraphs each, so the system can retrieve just the part that answers a question.
Turn chunks into vectors
An embedding model turns each chunk into a list of numbers. Chunks with similar meaning land close together in this space.
Embed the question
The question goes through the same embedding model, so it lands in the same space as the chunks.
Find the nearest chunks
A vector search returns the chunks closest to the question. It's fast, but closeness is only a rough guess at relevance.
Rerank
A reranker reads the question and each candidate together and scores how well it actually answers. The order changes.
Answer with sources
The top chunks go into the prompt. The model answers from them and cites the chunk behind each claim.
In short
- RAG keeps answers grounded in documents you control, and you can update them without retraining the model.
- Retrieval decides answer quality: chunk size, the embedding model and reranking matter more than the prompt.
- Hybrid search, keyword matching (BM25) plus vectors, catches exact terms like product codes that embeddings miss.
- Keep a set of real questions with known answers, and check both what was retrieved and what was answered.