RAG vs Fine-Tuning

RAG or fine-tuning: which one your data actually needs

Almost every project that arrives asking about fine-tuning turns out to need retrieval instead. The two solve different problems, and picking the wrong one costs months. The short version: retrieval gives a model access to facts, fine-tuning teaches it behaviour.

Explore all services

What each one actually does

Retrieval-augmented generation leaves the model alone and changes what it can see. At question time the system finds the relevant passages in your own documents and hands them to the model with the question. The model answers from material it was given, and can cite where each claim came from.

Fine-tuning changes the model itself, training it on examples until it behaves differently — adopting a format, a tone, a classification scheme or a domain vocabulary. It does not teach facts reliably, and facts taught this way cannot be cited or updated without training again.

The question that decides it

Does your content change? If your policies, prices, contracts or product details are updated more than rarely, retrieval is the answer, because the alternative is retraining every time a document changes. This single question settles most projects.

Does the answer need a source? Anything a regulator, an auditor or a customer might challenge needs a citation, and retrieval gives you one by construction. A fine-tuned model produces fluent answers with nothing behind them.

When fine-tuning genuinely earns its place

A narrow, stable task repeated at volume where the model keeps getting the shape wrong — a specific output format it will not hold, a classification scheme with your own labels, a domain register that generic prompting cannot reach.

Even then it usually sits on top of retrieval rather than replacing it: the tuned model handles the behaviour, retrieval supplies the facts. Projects that skip straight to fine-tuning normally end up rebuilding retrieval later anyway.

What a RAG build really involves

The model is the easy part. The work is in the pipeline: getting documents out of whatever holds them, splitting them so a passage still makes sense on its own, keeping the index current as documents change, and respecting who is allowed to see what.

Access control is the one most often forgotten. If retrieval can reach a document the asker could not open themselves, you have built a data leak with a friendly interface. We scope permissions into the retrieval layer rather than bolting them on.

How we scope it

A short assessment: what questions people need answered, where the source material lives, how often it changes, who may see which parts, and what happens today when someone cannot find the answer.

The output is a written recommendation with options, costs and risks. If prompting alone solves it without either approach, that is in the document too, and it is the cheapest outcome on this page.

What we already build

11 real services,
ready today.

While we scope this together, here's what's already a proven, dedicated service.

Questions we get asked

Before you
brief anyone.

The answers are the same ones we give on a first call — including where the honest answer is "it depends, and here is what it depends on".

What is the difference between RAG and fine-tuning?

Retrieval gives a model access to facts at question time by finding relevant passages in your documents. Fine-tuning changes the model itself to behave differently. Retrieval is for knowledge, fine-tuning is for behaviour, and they are not substitutes.

Which should we use for our company documents?

Almost always retrieval. If your documents change and your answers need a source, fine-tuning is the wrong tool — you would be retraining every time a policy is updated, and the answers could not be cited.

Can we do both?

Yes, and that is often the right end state: a tuned model handles the output format or classification while retrieval supplies the facts. We would still build retrieval first and add tuning only if a specific behaviour problem remains.

Is RAG cheaper than fine-tuning?

Usually cheaper to start and cheaper to keep current, because updating means re-indexing a document rather than retraining. Its ongoing cost is the pipeline — keeping the index fresh and the permissions correct.

Why do RAG projects fail?

Rarely because of the model. Usually because documents were split so passages lost their meaning, the index went stale, or access control was added late and retrieval could reach material the asker should never see.

How long does a RAG build take?

A narrow pilot over one document set is weeks. A production system across several sources with real permissions and monitoring is longer, and the timeline is driven by the state of the documents rather than the AI.

A practical first step

Ready to talk it through?

Tell us what you're building and we'll be honest about whether we're the right fit.

READY WHEN YOU ARE

Start the conversation.

Tell us what you're building and we'll get back to you within one business day.