Ask a general AI model about your returns policy, your price list or last year's contract, and it will either say it does not know or, worse, invent something plausible. Retrieval-augmented generation, usually shortened to RAG, is the standard way to fix that.

The problem RAG solves

A language model is trained on public text up to a certain date. It has never seen your documents, and it has no way to look anything up unless you give it one. Left alone, it answers from general knowledge and fills the gaps with guesses that read just as confidently as facts.

RAG changes the task. The model is no longer asked what it knows. It is handed the relevant passages from your own material and asked to answer from those.

How it works, in three steps

  • Index. Your documents are split into passages and stored in a searchable index, often a vector database, which can find passages by meaning and not only by matching words.
  • Retrieve. When someone asks a question, the system finds the handful of passages most relevant to it.
  • Generate. The question and those passages go to the model, with an instruction to answer only from what it has been given and to say where the answer came from.

The model does the reading and the writing. Your documents supply the facts.

Why not train the model on our data?

Training, or fine-tuning, a model on your documents sounds like the obvious route. It is usually the wrong one for this job.

  • It is slower and more expensive, and it has to be repeated when your content changes.
  • It does not reliably make a model recall specific facts. Fine-tuning shapes style and behaviour better than it stores knowledge.
  • A trained model cannot show you which document an answer came from.
  • Access rules are hard to apply once information is inside the model.

With RAG, updating the answer means updating the document. Sources can be shown next to every answer, and a user only gets passages they are allowed to see.

Where it works well

  • Customer support that answers from your help articles, policies and product information.
  • Internal search across procedures, manuals and past projects.
  • Document-heavy work such as contracts, tenders and compliance material.
  • Questions about a product catalogue or a technical specification.

The common thread is a body of written material that people already consult, and questions whose answers are in it.

What makes it go wrong

  • Poor content. Outdated, duplicated or contradictory documents produce outdated or contradictory answers.
  • Poor retrieval. If the right passage is not found, the best model in the world cannot answer correctly.
  • The wrong kind of question. Totals, comparisons across hundreds of records and calculations need a database query, not a reading exercise.
  • Missing permissions. Without access rules, a system can show one person what was meant for another.
  • No limits. A system that is never allowed to say it does not know will make something up.

How to tell if it is good enough

Collect a few dozen real questions from the people who will use the system, with the correct answers. Run them through, and check each answer and the passages it was built from. Repeat this every time something changes.

Then set the guardrails: what the system should decline, when it should hand over to a person, and what it must never reveal. After launch, keep watching cost, response time and answer quality, because all three move as usage grows.

What you need to start

  • The documents, in whatever form they are in today.
  • A list of real questions people ask.
  • Someone who owns the content and can correct it when an answer exposes a gap.

You do not need a large data project or a model of your own. Most useful systems of this kind start with material the business already has.

Our service

AI / ML Integration