What each one actually does
Retrieval-augmented generation leaves the model alone and changes what it can see. At question time the system finds the relevant passages in your own documents and hands them to the model with the question. The model answers from material it was given, and can cite where each claim came from.
Fine-tuning changes the model itself, training it on examples until it behaves differently — adopting a format, a tone, a classification scheme or a domain vocabulary. It does not teach facts reliably, and facts taught this way cannot be cited or updated without training again.
The question that decides it
Does your content change? If your policies, prices, contracts or product details are updated more than rarely, retrieval is the answer, because the alternative is retraining every time a document changes. This single question settles most projects.
Does the answer need a source? Anything a regulator, an auditor or a customer might challenge needs a citation, and retrieval gives you one by construction. A fine-tuned model produces fluent answers with nothing behind them.
When fine-tuning genuinely earns its place
A narrow, stable task repeated at volume where the model keeps getting the shape wrong — a specific output format it will not hold, a classification scheme with your own labels, a domain register that generic prompting cannot reach.
Even then it usually sits on top of retrieval rather than replacing it: the tuned model handles the behaviour, retrieval supplies the facts. Projects that skip straight to fine-tuning normally end up rebuilding retrieval later anyway.
What a RAG build really involves
The model is the easy part. The work is in the pipeline: getting documents out of whatever holds them, splitting them so a passage still makes sense on its own, keeping the index current as documents change, and respecting who is allowed to see what.
Access control is the one most often forgotten. If retrieval can reach a document the asker could not open themselves, you have built a data leak with a friendly interface. We scope permissions into the retrieval layer rather than bolting them on.
How we scope it
A short assessment: what questions people need answered, where the source material lives, how often it changes, who may see which parts, and what happens today when someone cannot find the answer.
The output is a written recommendation with options, costs and risks. If prompting alone solves it without either approach, that is in the document too, and it is the cheapest outcome on this page.