RAG vs. Fine-Tuning for Domain Adaptation: When to Use Which
RAG and fine-tuning solve different problems. Learn the mechanical difference, see working code examples, and use a six-point framework to decide.
In this article, you will learn the mechanical difference between retrieval-augmented generation and fine-tuning, when each technique is the right tool, and how to decide which one — or both — your production system actually needs.
Topics covered include:
- What RAG and fine-tuning each do at a mechanical level, and what each one cannot do.
- Two complete, working code examples — one RAG pipeline for a knowledge-retrieval use case, and one LoRA fine-tuning setup for a structured-output use case.
- A concrete six-point decision framework for choosing between RAG, fine-tuning, or both.

RAG vs. fine-tuning is one of the most searched, most argued-about tradeoffs in applied LLM work right now, and most of the debate happens at the wrong level of abstraction. It gets framed as a single either/or decision, when the honest 2026 picture is that roughly 60% of production LLM deployments now use both together — not because teams couldn’t decide, but because retrieval-augmented generation and fine-tuning solve two genuinely different problems, and most real domain-adaptation projects have both problems at once.
This article breaks that false binary down properly: what each technique actually does at a mechanical level, two complete working examples — one for each approach — and a concrete decision framework for figuring out which one your specific project actually needs, and when the honest answer is both.
What RAG Actually Is (and Isn’t)
Retrieval-augmented generation doesn’t touch the model at all. The model’s weights never change; what changes is what the model sees in its context window at the moment it’s asked a question. A retrieval step searches a knowledge base, pulls back the most relevant documents, and hands them to the model alongside the user’s query, so the model is answering with a briefing document in front of it rather than from memory alone.
That mechanism is what makes RAG genuinely good at exactly one category of problem: information that’s large, that changes, or both. What RAG doesn’t fix is a model’s underlying behavior. If the model’s tone is inconsistent, if it won’t reliably follow a strict output format, if it uses your industry’s vocabulary incorrectly, feeding it more documents at inference time doesn’t touch any of that — because the problem was never a lack of information in the first place.
What Fine-Tuning Actually Is (and Isn’t)
Fine-tuning does the opposite: it changes the model itself, training the weights on real examples of the input/output behavior you want until that behavior becomes the model’s default, with no need to inject anything at inference time because the pattern is now baked in. LoRA and QLoRA are the standard approach for the large majority of projects, training a small adapter — often under 1% of the base model’s total parameters — rather than the full model, which brings a fine-tuning run down to a few hundred dollars and a few hours instead of a full retraining project.
Here’s the point worth landing hard, because it’s the single most common misunderstanding in this whole debate: fine-tuning doesn’t reliably add factual knowledge. A model fine-tuned on a pile of medical literature doesn’t “know” the facts in that literature the way a retrieval system genuinely does — it adjusts style, structure, and pattern recognition, but factual recall from training data is unreliable, especially for granular facts. Fine-tuning is a behavior tool, not a knowledge tool. Keep that distinction in mind, since it’s exactly what the two examples below are built to demonstrate directly rather than just assert.
Retrieval-Augmented Generation (RAG)
The scenario: an internal engineering team wants to ask natural-language questions against their incident runbooks and postmortems — documents that get added to and edited constantly as new incidents happen. This is a textbook RAG problem: the knowledge changes weekly, and every answer needs to be traceable back to a real source document for anyone debugging at 2 a.m.
Prerequisites:
- Python 3.10+
pip install scikit-learn anthropic- An Anthropic API key
First, the documents themselves — a small but real set of runbooks and postmortems:
# documents.py
DOCUMENTS = [
{
"id": "runbook-db-failover-001",
"title": "Database Failover Runbook",
"text": (
"When the primary Postgres instance becomes unresponsive, first check "
"replication lag on the standby via `SELECT now() - pg_last_xact_replay_timestamp()`. "
"If lag is under 30 seconds, promote the standby using `pg_ctl promote`. "
"Update the connection string in the config service immediately after promotion. "
"Do not attempt manual failover if replication lag exceeds 5 minutes, escalate "
"to the database team instead, since promoting a stale standby risks data loss."
),
},
{
"id": "postmortem-2026-03-outage",
"title": "Postmortem: March 2026 Checkout Outage",
"text": (
"Root cause was a connection pool exhaustion in the payments service after a "
"deploy reduced the pool size from 100 to 20 connections. Fix was reverting the "
"pool size and adding a minimum-pool-size alert. Action item: connection pool "
"changes now require a second reviewer from the platform team before merge."
),
},
{
"id": "runbook-oncall-escalation-003",
"title": "On-Call Escalation Policy",
"text": (
"Primary on-call has 15 minutes to acknowledge a page before it escalates to "
"secondary. Secondary has 10 minutes before escalating to the team lead. Any "
"incident affecting checkout or payments skips the normal escalation chain and "
"pages the team lead directly, regardless of acknowledgment status."
),
},
# additional documents omitted here for length
]
Now the chunking and retrieval index:
# retrieval.py
import re
from dataclasses import dataclass
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
@dataclass
class Chunk:
doc_id: str
title: str
text: str
The RAG approach keeps all knowledge outside the model, making it straightforward to update runbooks and postmortems without any retraining. The retrieval step surfaces the most relevant chunks at query time, and the model generates a grounded, traceable answer from the documents placed in its context window.
LoRA Fine-Tuning for Structured Output
The scenario: a team needs a model to reliably produce structured JSON output in a specific schema — a behavior problem, not a knowledge problem. No matter how many documents you retrieve, you can’t retrieval-augment a model into consistently following a strict output format it wasn’t trained to produce. This is where fine-tuning on input/output examples pays off.
Prerequisites:
- Python 3.10+
pip install transformers peft datasets torch- A GPU with at least 16 GB VRAM, or a cloud notebook instance
LoRA fine-tuning trains a small adapter on top of the frozen base model weights. The adapter learns the target behavior — in this case, consistent structured output — while the base model’s parameters remain unchanged. After training, the adapter can be merged back into the base model or applied at inference time, keeping the deployment footprint small.
The key distinction from the RAG example is that once fine-tuning is complete, the behavior is intrinsic to the model. There’s no retrieval step, no context injection, and no dependency on an external knowledge base at inference time. The model simply produces the target output format by default.
Six-Point Decision Framework
Use the following questions to decide between RAG, fine-tuning, or both:
-
Does the knowledge change frequently? If yes, RAG. Fine-tuning a new adapter every time your documents update is operationally expensive and slow. RAG lets you update the knowledge base without touching the model.
-
Do you need source attribution? If yes, RAG. Retrieval gives you document provenance by design. Fine-tuned models can’t tell you where a fact came from because the fact isn’t stored anywhere retrievable.
-
Is the problem about behavior, not knowledge? If yes, fine-tuning. Consistent output format, tone, domain-specific vocabulary, and structured response schemas are all behavior problems. More documents in context won’t fix them.
-
How much labeled input/output data do you have? Fine-tuning requires examples of the behavior you want. If you have fewer than a few hundred high-quality examples, fine-tuning results will be unreliable. RAG has no such requirement.
-
What are your latency and cost constraints? RAG adds a retrieval step and a longer context window at every inference call. Fine-tuning pays its cost upfront at training time and then runs at standard inference cost. For high-volume, low-latency applications, fine-tuning often has better economics once the behavior is stable.
-
Do you have both problems? Most production domain-adaptation projects do. A medical assistant needs both up-to-date clinical knowledge (RAG) and consistent output formatting for downstream systems (fine-tuning). Combining both is operationally more complex but often the honest answer.
Summary
RAG and fine-tuning are not competing solutions to the same problem — they target different failure modes. RAG solves knowledge access: it gives the model accurate, up-to-date, attributable information at inference time without changing the model itself. Fine-tuning solves behavior: it trains the model to reliably produce a specific style, structure, or output pattern without needing anything injected at inference time. Framing the choice as either/or is what leads teams to apply the wrong tool and then conclude the whole approach doesn’t work. Understand what each technique actually changes, match it to the failure mode you actually have, and the right answer — RAG, fine-tuning, or both — becomes straightforward.