← All articles

Making your own documents actually searchable

·2 min read·By Adrian

Also available in ES, RO

Every established company sits on an archive: proposals, specifications, reports, procedures, correspondence. The knowledge is there. Retrieval is the problem, because keyword search only finds documents containing the words you happened to guess.

What retrieval-augmented generation does

The mechanism is simpler than the name:

  1. Documents are split into passages
  2. Each passage is converted into a numeric representation of its meaning
  3. A question is converted the same way, and the closest passages are retrieved
  4. Those passages are given to a language model, which answers using them
  5. The answer cites the documents it came from

The important consequence: answers are grounded in your material, and every claim can be traced back to a source.

Why the citations matter most

An ungrounded model will produce a confident, plausible, wrong answer. One that must cite retrieved passages can be checked in seconds. For business use, that verifiability is the difference between a useful tool and a liability.

Where implementations go wrong

Splitting badly. Chopping documents at fixed character counts cuts tables in half and separates headings from content. Split on structure — sections, clauses, paragraphs.

Ignoring metadata. Date, author, client and document type must be filters. Without them, a superseded 2019 procedure ranks alongside the current one.

Skipping the ranking step. Retrieve twenty passages, re-rank them properly, pass the best five. This single step improves answer quality more than upgrading the model.

Feeding it everything. A curated set of current, authoritative documents beats a full archive dump. Rubbish retrieved is rubbish answered.

No feedback route. Users must be able to flag a bad answer, or you will never know which questions fail.

Realistic expectations

It answers questions whose answers exist in the documents. It will not reason across fifty files to produce novel analysis, and it will struggle where your documents genuinely disagree with each other — though surfacing that disagreement is itself valuable.

Start narrow

One document set, one team, one question type. Prove it on something like "what did we quote this client for, and when", then widen.

We build these systems under AI implementation.

← Blog