RAG in plain words

RAG lets AI answer from the company’s own documents and sources instead of generating text out of thin air.

A language model on its own knows nothing about your company. It is good at assembling plausible text, and if you ask it about your delivery times it will answer — confidently and wrongly.

RAG solves this in a straightforward way: before answering, the system finds a few relevant fragments in your documents and puts them into the request alongside the question. The model answers not from memory but from the text it has just been shown. Hence the name: retrieval-augmented generation.

What a source is

A source is a specific document, or part of one, that somebody is responsible for. The returns policy. A clause of the contract. A knowledge-base page. A service description. What matters is not the file format but that it has an owner and a clear shelf life.

A chat thread, a draft and “that file on Marina’s drive” are not sources, even when they contain the right answer. The first thing usually needed before launching RAG is to separate the documents the company is prepared to stand behind from everything else.

That separation has an unexpected consequence: the list of sources almost always turns out shorter than expected. Of a hundred files in a shared folder, a dozen or so survive to the state of “we can point at this”. The rest are old revisions, drafts and material nobody is prepared to stand behind. Starting with that dozen is the right move.

How to prepare documents

The search works not with a whole document but with fragments. So a fragment has to stand on its own: it should be clear what it is about without the neighbouring pages.

  • headings inside the document — they mark where sensible chunks begin
  • tables and scans — either turned into text or described separately
  • one topic per document, instead of collections titled “miscellaneous”
  • date and version next to the text, not only in the file name

This work looks dull, but it is what determines the quality. A badly chunked base returns fragments that are formally relevant and practically useless — and then the model starts filling the gaps itself.

Why showing the grounds matters

An answer with no reference to a source cannot be checked, which means it cannot be used with a client. So the interface always shows where the prompt came from: the document, the section, the passage.

The side effect turns out to be almost the main one: people start noticing that documents are out of date or contradict each other. RAG becomes a way of tidying the knowledge base.

Showing the grounds also changes how people relate to the system. A prompt with no reference reads as an opinion from nowhere, and gets either taken on faith or ignored wholesale. A prompt with a quotation from the policy is simply a fast way to find the right clause, and there is nothing to argue with: you can see where it came from.

Updating is worth thinking about from the start as well. If a policy is revised once a quarter, rebuilding the base by hand after each change is enough. If documents change weekly, you need a process: who updates it, what happens to the old version and how to confirm the system is already answering from the new one. Without that the base quietly drifts away from the originals within a couple of months.

Where RAG ends

RAG answers questions whose answers are written down somewhere. It does not calculate, does not reconcile data across systems and does not know what is not in the documents. If a rule lives only in the head of a department manager, no search will find it.

Questions that require assembling a picture from a dozen places at once are also hard: “how many do we have in progress right now” is a query to a system, not to a text. Those cases are solved by integration, not by document search.

Contradictions are a difficulty of their own. If two policies say different things, the search will honestly find both fragments, and the answer will depend on which ranked higher. Resolving that is a human job: mark one revision as superseded rather than hoping the system picks the right one.

How to judge quality

RAG quality splits in two: was the right fragment found, and did the model restate it correctly. It pays to measure these separately — they are fixed differently. The first is cured by chunking and metadata, the second by the wording of the request and the rules for answering.

You test on a set of real questions collected in advance, and repeat after every change to the base. Separately you watch how the system behaves when the answer is not in the documents: staying silent and handing over to a person is good, answering confidently is a reason to investigate.

The comparison worth making is not against perfection but against how it used to be. An operator without an assistant also makes mistakes, also fails to find the right clause and also sometimes answers from memory. The question is not whether RAG is flawless but whether things improved — and whether new ways to be wrong appeared that did not exist before.

Have a document base you need to answer from?

Describe where your policies live and who answers clients from them. We will work out what needs preparing and how to build checkable prompts for the operator.

Show us your process