Skip to main content

AI coding · Advanced

Retrieval-Augmented Generation

Also called: RAG, Knowledge base Q&A, Vector search, Retrieval augmentation

First search a knowledge base for content related to the question, then put it in the prompt for the model to use. This cuts down on made-up answers and lets the model use up-to-date or private material.

In detail

RAG is like an open-book exam. The model doesn't answer purely from "memory." It first flips to the relevant pages, then answers using them. For example, to build a support bot that answers from company docs, you first split the docs into small chunks and store them. When a user asks something, you retrieve the most relevant chunks and hand them to the model.

How it relates to the Context Window: the context window is the limit on how much the model can see at once. RAG solves "there's too much material to fit, so only put in what's relevant." When AI coding tools index your codebase and find relevant files as needed, it's a similar idea.

Most of the time you don't need to build RAG yourself. You only need to learn about vector databases, chunking, and retrieval when you're building a product feature that "answers questions from a large amount of private material."

Developer info
Term ID
ai-rag
DOM selectors
No DOM cues. This concept isn't detected directly on a page.
Priority
1 · when several match at the same level, the higher priority wins
Version
v1 · updated Sep 29, 2026