Fine-Tuning vs Prompting vs RAG: How to Adapt a Model
Fundamentals

Fine-Tuning vs Prompting vs RAG: How to Adapt a Model

Three ways to make a general model do your specific job, with very different costs. Here is a decision rule based on what each one can and cannot actually change.

You have a general-purpose model and a specific job it does not quite do well. There are three ways to close the gap, and teams routinely pick the most expensive one first.

The useful distinction is not cost or effort. It is what each method can actually change: prompting changes instructions, retrieval changes available knowledge, fine-tuning changes learned behaviour. Diagnose which one you are missing and the choice mostly makes itself.

Prompting changes instructions

You supply the task description, constraints, output format and a few examples. Nothing about the model changes. Iteration is instant and free.

Prompting handles more than people expect. Tone, structure, output schema, step ordering, edge-case handling, refusal behaviour — all of it is reachable from the prompt, and few-shot examples are a surprisingly strong lever for format and style specifically.

Its limits are real though. Every instruction costs tokens on every request, so an elaborate system prompt is a permanent tax. Very long instruction sets degrade as the model attends unevenly across them. And prompting cannot install knowledge the model does not have, nor reliably teach a genuinely novel output format that looks nothing like anything in training.

Start here. Always. The cost of exhausting prompting is a few hours; the cost of skipping it is finding out after a fine-tuning run that a better prompt would have done it.

Retrieval changes available knowledge

RAG fetches relevant documents at request time and puts them in the prompt. The model does not learn anything; it reads.

This is the right tool whenever the gap is factual. Your internal documentation, your product catalogue, your customer records, anything that postdates the training cutoff. It also gives you two properties fine-tuning cannot: citations, because you know which document produced the answer, and access control, because you filter at retrieval time rather than hoping the model respects a rule.

Updating is cheap — re-index a changed document and it is live. Compare that with retraining.

The costs are infrastructure and latency. You need chunking, an embedding pipeline, an index and, if you want good results, reranking and evaluation. Retrieval quality also becomes a hard ceiling on answer quality: if the right chunk never surfaces, no model can use it.

Fine-tuning changes learned behaviour

Fine-tuning continues training on your examples, adjusting weights. It is the only one of the three that changes what the model has internalised.

Full fine-tuning updates every parameter and needs enough memory to hold optimiser state for the whole model. Almost nobody does this now. LoRA — Low-Rank Adaptation of Large Language Models, Hu et al., 2021 — freezes the pretrained weights and trains small low-rank matrices alongside them, cutting trainable parameters by orders of magnitude while retaining comparable accuracy. QLoRA (Dettmers et al., 2023) added 4-bit NormalFloat quantisation of the frozen base, double quantisation and paged optimisers, pushing 70B-parameter fine-tuning onto a single high-memory GPU.

What fine-tuning is genuinely good at:

  • Consistent style or voice that would otherwise need a long prompt on every call.
  • Rigid output formats, particularly domain-specific ones.
  • Narrow classification where a smaller fine-tuned model beats a larger prompted one on both accuracy and cost.
  • Prompt compression. Baking a 2,000-token system prompt into weights can pay for itself at volume.
  • Distillation. Training a small model on a large model outputs for one specific task.

What it is bad at: teaching facts. Fine-tuning on documents to make a model "know" them is the single most common mistake in this area. It teaches the model the shape and vocabulary of your documents, which often makes hallucination worse — the output now sounds authoritative in your house style while remaining wrong.

The decision rule

  1. Is the model missing information? Retrieval. Not fine-tuning. This includes anything private, changing, or after the cutoff.
  2. Is the model missing instructions? Prompting. Be explicit, give examples, constrain the output.
  3. Does the model understand the task but consistently do it in the wrong style or format, after serious prompting effort? Fine-tuning.
  4. Is a capable model too slow or too expensive at your volume? Fine-tune a smaller one on the specific task.

These combine. A well-built production system frequently uses all three: a fine-tuned model for consistent behaviour, retrieval for current facts, and a compact prompt to steer each request.

What fine-tuning actually costs

The GPU hours are usually the smallest line item, which surprises people.

Data is the real expense. You need hundreds to thousands of high-quality examples in the exact format you want out. Curating and cleaning them is human work, and the quality of that dataset dominates the result far more than hyperparameters do.

Evaluation is mandatory. Without a held-out set you cannot tell improvement from overfitting, and fine-tuned models degrade on tasks outside their training distribution in ways that are easy to miss.

Maintenance is permanent. Base models improve. Every upgrade means retraining and re-evaluating. Teams underestimate this and end up pinned to an older base model because moving is too much work.

Serving gets more complex. Self-hosted adapters need infrastructure, though multi-adapter serving has made hosting many LoRAs on one base model far cheaper than it once was.

A worked example

Suppose you are building support automation over an internal knowledge base.

The wrong approach is to fine-tune on all the support documents. You get a model that writes convincingly in your support voice and invents policies.

The right sequence: build retrieval over the documents first, because the answers must be grounded and current. Prompt carefully for the response structure and escalation rules. Measure. If after that the tone is still inconsistent and the prompt has grown unmanageably long, fine-tune on a few hundred approved response pairs to bake the style in — and keep retrieval doing the factual work.

Order matters because each step tells you whether the next is necessary. Most teams that start with fine-tuning discover afterwards that retrieval alone would have solved it, having spent weeks finding out.

Common questions

Can I fine-tune a model to know my company documents?

You can, and it usually backfires. Fine-tuning teaches style and structure far more reliably than facts, so the model learns to sound like your documentation while still fabricating details. Use retrieval for knowledge.

How many examples does fine-tuning need?

Hundreds to thousands, depending on how narrow the task is. Quality and consistency of the examples matter more than volume, and a small clean dataset routinely beats a large noisy one.

What is LoRA and why does everyone use it?

Low-rank adaptation freezes the base weights and trains small added matrices, cutting trainable parameters by orders of magnitude at comparable accuracy. QLoRA adds 4-bit quantisation of the frozen base, bringing large-model fine-tuning within reach of a single GPU.

Similar articles

Instruction Tuning: How a Text Predictor Becomes an Assistant
Fundamentals
Fundamentals·8 min read

Instruction Tuning: How a Text Predictor Becomes an Assistant

A pretrained model continues text; it does not answer questions. Instruction tuning is the small, cheap stage that turns one into the other.

Read
RAG vs Long Context: Retrieval Did Not Become Obsolete
Fundamentals
Fundamentals·8 min read

RAG vs Long Context: Retrieval Did Not Become Obsolete

Large context windows were supposed to kill retrieval. They did not. Here is how the two actually compare on cost, accuracy and latency — and when to use each.

Read
Attention Mechanisms Explained Without the Linear Algebra
Fundamentals
Fundamentals·9 min read

Attention Mechanisms Explained Without the Linear Algebra

What attention actually computes, why it made transformers work, and why its cost scaling explains almost every practical limit you hit with long context.

Read