RAG & Retrieval

RAG Explained: A Beginner's Guide to Retrieval-Augmented Generation

8 min readRuhi GargaGuideFoundations

A language model writes fluent answers about almost anything, but it can only draw on what it saw during training. Retrieval-augmented generation, or RAG, gives it the relevant text to read at the moment you ask the question.

The problem in one scene

Picture a company handbook and an assistant built on a language model. An employee asks, “How much time off do I get after having a baby?” The model has never seen this company’s handbook. At best it says it doesn’t know. At worst it answers from general patterns about parental leave, in the same confident tone it uses for facts it does know, and the answer is wrong for this company.

That gap, between what a model learned and what you need it to know, is the problem RAG addresses. The handbook in this guide is made up, and so is every number in it. It exists only to make the steps concrete.

Why RAG exists

A language model has three limits that matter here:

  • Its knowledge stops at a date. Whatever happened after training is unknown to it.
  • It has not seen your private material. Internal policies, product documents and tickets were never part of its training.
  • It can guess convincingly. When it lacks the facts it can still produce fluent, plausible text. This is often called hallucination.

There are three common ways to deal with that. You can paste your documents into the prompt, which works for a small amount of text but runs into the model’s input limit, costs more as the text grows, and buries the useful part among everything else. You can train the model further on your documents, which is generally a better way to change how a model behaves than a way to keep facts accurate and up to date. Or you can find the few passages that matter when the question arrives and give only those to the model. That third option is RAG.

Because the documents stay outside the model, updating a document changes the answers without retraining anything. And because the answer is built from specific passages, it can point to the ones it used.

The idea in one sentence

Find the relevant passages first, then ask the model to answer using them.

An open-book exam is a fair comparison. In a closed-book exam you rely on memory. In an open-book exam you look up the right page and then write your answer. RAG changes the model’s job from remembering to reading.

The architecture at a glance

A RAG system has two flows. Ingestion runs ahead of time and again whenever the documents change. It prepares the material for searching. The query flow runs every time someone asks a question. The two meet at the vector store, which holds the prepared passages.

RAG architecture: ingestion and query flowsIngestion: documents are chunked, embedded and stored in a vector store. Query: a question is embedded, similar chunks are retrieved from the store, combined with the question into a prompt, and passed to the language model, which returns a grounded answer.1 · INGESTION OFFLINE, REPEATED WHEN DOCUMENTS CHANGEDocumentsChunkEmbed chunksVector storechunk text + embedding + source detailsstore2 · QUERY ONLINE, EVERY QUESTIONQuestionEmbed questionRetrieveclosest chunksBuild promptquestion + contextLLMAnswergroundedsearch

Two flows meeting at the vector store. The names of the stages vary between implementations; the shape is common.

Flow 1: ingestion

Ingestion turns a pile of documents into something that can be searched by meaning.

  1. Load. Collect the documents, such as PDFs, wiki pages, web pages or tickets, and extract their text along with useful details like the title, source and date.
  2. Clean. Remove repeated headers and footers, navigation text and obvious duplicates, so they don’t crowd the results later.
  3. Chunk. Split each document into passages, commonly a few paragraphs or fewer.
  4. Embed. Turn each chunk into an embedding, explained in the next section.
  5. Store. Save the embedding, the original text and the source details together in a vector store, which may be a dedicated vector database or a search system with vector support.

Chunking exists for three reasons. The model cannot read a whole library for every question. Retrieval should return focused passages and not entire documents. And an embedding captures a passage best when the passage is about one thing. Chunk size is a trade-off: pieces that are too small lose their context, and pieces that are too large mix topics and waste space in the prompt. Many pipelines let neighbouring chunks overlap a little, so a sentence cut at a boundary is not lost. How to choose well is a production question, and What Makes a RAG System Production-Ready? covers it.

Embeddings and vector search without the maths

An embedding is a list of numbers that a model produces from a piece of text, arranged so that passages with similar meaning get similar numbers. It helps to think of coordinates on a map, where meaning is location. “Parental leave” and “time off after a birth” sit close together even though they share almost no words. “Expense claims” sits somewhere else.

Embedding space, simplified to two dimensionsPassages about leave, expenses and remote work form three groups. A question about time off after having a baby lands next to the leave passages.closer together = more similar in meaningparental leavetime off after a birthmaternity and paternityexpense claimsreimbursing travel costsworking remotelyhome-working approvalQuestion: how much time off after having a baby?The three nearest passages are retrieved.

A simplification. Real embeddings have hundreds or thousands of dimensions, not two, but the idea of closeness is the same.

Search works by embedding the question with the same model and then finding the stored passages nearest to it. Closeness is measured with a similarity score, and cosine similarity is a common choice. With many chunks, vector stores usually use an approximate nearest-neighbour index. It trades a small amount of exactness for a large gain in speed.

This is different from keyword search, which matches the words themselves. Vector search matches meaning, so it can find the leave policy for a question that never uses the word “leave”. It can also be weaker on exact terms such as part numbers or error codes, which is one reason some systems combine both approaches. One rule is easy to forget: use the same embedding model for the documents and the questions, or their numbers won’t be comparable.

Flow 2: the query

  1. Ask. The user types a question.
  2. Embed the question. The same embedding model turns it into numbers.
  3. Retrieve. The system finds the closest chunks, usually a handful, ranked by similarity.
  4. Build the prompt. It combines instructions, the retrieved chunks and the question into one prompt.
  5. Generate. The language model writes an answer from that prompt.
  6. Return. The answer is shown, usually with the sources of the chunks it used, so the reader can check it.

The word grounded describes an answer that is based on the supplied text and not on the model’s memory. The prompt tells the model to answer from the context and to say so when the context does not contain the answer. That instruction helps, but it is not a guarantee.

A small worked example

Here is the made-up handbook, already split into three chunks. The text is invented for this guide and is not a real policy.

ChunkSourceText
ALeave policyEmployees are entitled to 26 weeks of parental leave. Leave can begin up to 11 weeks before the expected birth date.
BExpensesExpense claims must be submitted within 30 days, with receipts attached.
CBenefits overviewBenefits include parental leave, health cover and a pension scheme.

The employee asks: “How much time off do I get after having a baby?” The question shares almost no words with chunk A. It still lands close to it in meaning, so retrieval ranks the chunks like this:

  1. Chunk A (Leave policy), closest in meaning.
  2. Chunk C (Benefits overview), related but less specific.
  3. Chunk B (Expenses), unrelated, so it is left out.

The system keeps the top results and builds the prompt:

InstructionsAnswer the question using only the context below. If the answer is not in the context, say that you don’t know.
Context: retrieved chunks[Leave policy] Employees are entitled to 26 weeks of parental leave. Leave can begin up to 11 weeks before the expected birth date.[Benefits overview] Benefits include parental leave, health cover and a pension scheme.
QuestionHow much time off do I get after having a baby?

The prompt the model receives. The user never sees this assembly.

Grounded answer

Employees are entitled to 26 weeks of parental leave, and it can begin up to 11 weeks before the expected birth date. Source: Leave policy.

The model did not remember this from training, because it was never trained on this handbook. It read it from the context it was given, and the source line lets the employee check it.

What RAG does not fix

RAG improves the odds that an answer is based on the right material. It does not make the system reliable by itself. The usual weak points are:

  • Retrieval can miss. If the right chunk is not retrieved, the model answers without it, and may still guess.
  • Chunk boundaries cause trouble. An answer can be split across two chunks, or a chunk can lose the context that gave it meaning.
  • The answers are only as good as the documents. Outdated or contradictory documents are retrieved just as faithfully as correct ones.
  • Grounded is not guaranteed. The model can still go beyond the context it was given.
  • There is a cost. Embedding, searching and longer prompts add latency and expense.
  • Permissions are not automatic. A vector store does not decide on its own who may see which document, and retrieval must respect that.

These are the points where a prototype and a dependable system start to differ. What Makes a RAG System Production-Ready? picks them up in depth.

When RAG fits, and when it doesn’t

SituationFit
Answers must come from a body of documents that changesGood fit for RAG
Users need to see where an answer came fromGood fit for RAG
The knowledge is small and stable enough to include in the promptRAG is probably unnecessary
You want to change the model’s tone, format or behaviourA different tool: prompting or fine-tuning
The question needs an exact lookup or calculation over structured dataQuery the database or API directly
The task needs several steps, tool choices or actionsBeyond RAG alone: see RAG vs Agentic AI

A starting point for the decision, not a rule.

A good habit is to begin with the simplest design that meets the need. If a question can be answered by one lookup, retrieval is enough. If the path to the answer has to be decided along the way, the design starts to look like an agent, and RAG becomes one of the tools it uses.

Key terms

ChunkA passage of a document, small enough to retrieve and read on its own.
EmbeddingA list of numbers that represents the meaning of a piece of text.
Vector storeA database or index that holds embeddings and finds the closest ones to a query.
Top-kThe k closest chunks returned by a search, where k is a small number.
GroundingBasing an answer on supplied source text and not on the model’s memory.

Where to go next

This guide covers the foundations. For what changes when the system has to be dependable, read What Makes a RAG System Production-Ready? For where RAG sits among agents, see RAG vs Agentic AI and From RAG to Agentic AI. The AI Knowledge Assistant for Enterprises case study shows a system built around these ideas.

Key takeaways

  1. RAG finds relevant passages first and asks the model to answer from them, so the model reads instead of remembering.
  2. It has two flows: ingestion prepares the documents, and the query flow retrieves and answers.
  3. Embeddings place text by meaning, so a question can match a passage that shares few of its words.
  4. RAG does not remove the need for good documents, good retrieval and access control.
  5. Use it for changing, private knowledge. Use something simpler, or something more capable, when the problem calls for it.

← Back to AI & Agentic AI