Fine-tuning is not the answer to most problems — but for these three, it shipped 4x faster than RAG
A decision framework for when to LoRA vs. when to RAG
Fine-tuning is not the answer to most problems — but for these three, it shipped 4x faster than RAG
Every week I talk to teams that are drowning in RAG complexity. They’ve got chunking strategies, embedding models, vector databases, reranking pipelines, and a dozen hyperparameters that all whisper "tune me." The result? A system that works okay in demos but falls apart under real queries. Meanwhile, fine-tuning gets dismissed as "too expensive" or "too slow." But that’s cargo-cult thinking.
Fine-tuning, specifically LoRA (Low-Rank Adaptation), is not a silver bullet. For most problems — open-ended QA, summarization of arbitrary documents, or general knowledge lookup — RAG is the right tool. But there are exactly three patterns where a targeted LoRA fine-tuning ships faster, performs better, and costs less to maintain than any RAG pipeline I’ve seen.
Here’s the framework I use with my clients to decide, and the three patterns where fine-tuning wins every time.
The Decision Framework
Before you even think about training, ask three questions:
- Is the knowledge static or does it change weekly? If the answer set changes more often than you can retrain, use RAG.
- Does the model need to learn a new format or behavior, not just new facts? If you need to change how the model structures output, follow a strict template, or apply a specific reasoning chain, fine-tuning is likely the answer.
- Can you collect 50–500 high-quality examples? If yes, fine-tuning is on the table. If you have thousands of examples, even better.
If your problem hits all three — static knowledge, behavioral shift, and available examples — you’re in the sweet spot. Here are the three patterns that consistently ship 4x faster than building a RAG system.
Pattern 1: Structured Extraction from Heterogeneous Documents
You have a pile of invoices, contracts, or medical records — each from a different source, with different layouts, terminology, and quirks. RAG wants you to chunk them, embed them, and then ask "what is the total amount?" But the model has to understand the document structure, ignore headers and footers, and output a consistent JSON schema. That’s a formatting problem, not a retrieval problem.
I worked with a logistics company that processed 10,000+ PDF invoices per month from 50+ carriers. Each carrier used a different format. RAG with GPT-4 was hitting 70% accuracy on field extraction after weeks of prompt engineering and chunk tuning. A LoRA fine-tune on Llama 3 8B with 300 annotated invoices hit 96% accuracy in three days of training on a single A100. Inference cost dropped by 80% because they could use a smaller model.
Why it works: The model learns the behavior of extracting fields from messy layouts. It’s not looking up facts; it’s learning a transformation. RAG can’t teach that — it can only provide context.
How to do it:
- Collect 200–500 examples of input (PDF text or image) and output (JSON).
- Use a base model like Llama 3 8B or Mistral 7B.
- Apply LoRA with rank=16, alpha=32, target modules: q_proj, v_proj.
- Train for 3–5 epochs with a learning rate of 2e-4.
- Evaluate on a held-out set of 50 documents.
Pattern 2: Domain-Specific Classification with Subtle Boundaries
Classification tasks like "is this support ticket about billing or technical issue?" are easy. But what about "is this insurance claim likely fraudulent based on subtle patterns in the narrative?" RAG can’t help here because there’s no external document to retrieve — the answer is in the latent patterns of the text itself.
A fintech startup needed to classify transaction dispute narratives into 12 categories, each with nuanced legal definitions. Prompting GPT-4 with few-shot examples got them to 82% accuracy. A LoRA fine-tune on a 7B model with 1,000 labeled examples hit 94% in two days. The key was that the categories were defined by regulatory guidelines that didn’t change, and the model needed to internalize those boundaries.
Why it works: Fine-tuning teaches the model the decision boundary. RAG would require you to retrieve a "definition" for each category, but the model still has to apply it — and that’s where it fails.
How to do it:
- Use a balanced dataset of 500–2,000 examples per class.
- Format each example as
Input: <narrative> Output: <label>. - Train with LoRA on a model like BERT-large or DeBERTa for smaller tasks, or Llama 3 for longer text.
- Monitor validation accuracy and stop when it plateaus.
Pattern 3: Consistent Output Formatting for Downstream Systems
You need the LLM to output a specific JSON schema, XML, or even a custom DSL that another system consumes. RAG can provide examples, but the model will still hallucinate keys, miss nested fields, or format dates differently. Every inconsistency breaks the pipeline.
A healthcare analytics company needed to convert free-text clinical notes into FHIR R4 resources — a complex JSON standard with dozens of nested fields. Prompting with RAG context was a nightmare: the model would occasionally invent fields or skip required ones. A LoRA fine-tune on Mistral 7B with 500 annotated notes produced valid FHIR JSON 99% of the time. The RAG version never broke 85%.
Why it works: The model learns the exact output grammar. It’s not retrieving information; it’s generating structured data from unstructured input. Fine-tuning imprints the schema into the weights.
How to do it:
- Create a dataset where each input is the raw text and the output is the exact JSON you want.
- Use a base model that’s good at following instructions (Mistral, Llama 3, Qwen 2.5).
- LoRA with rank=32, alpha=64, target all linear layers.
- Train for 5–10 epochs with gradient checkpointing to save memory.
- Validate by parsing the output JSON and checking schema compliance.
Why Fine-Tuning Ships Faster
RAG seems fast because you don’t train a model. But the hidden cost is in the pipeline: chunking experiments, embedding model selection, vector DB tuning, reranker integration, and endless prompt engineering to make the model use the context correctly. Each of those is a rabbit hole.
Fine-tuning has a clear workflow: collect data, format it, train, evaluate, deploy. With LoRA, training takes hours, not days. A single A100 can fine-tune a 7B model on 1,000 examples in under 2 hours using Unsloth or Axolotl. Inference uses the same model — no extra infrastructure. You don’t need a vector database, an embedding service, or a reranker. Just a model and a simple API.
In all three patterns above, the teams went from problem definition to production in under a week. The RAG attempts had been dragging for months.
When Not to Fine-Tune
Let me be clear: fine-tuning is not for everyone. If your knowledge base changes weekly, or you need to answer questions about a large corpus of diverse documents, RAG is the right choice. Also, if you can’t collect at least 50 high-quality examples, fine-tuning will overfit or fail to generalize.
But if you’re in one of the three patterns — structured extraction, domain classification, or consistent formatting — stop building a RAG pipeline. Spend two days collecting examples, one day training, and ship. Your users will thank you.
The Bottom Line
Fine-tuning is not the answer to most problems. But these three patterns are the exceptions, and they’re more common than you think. The next time you’re about to add another chunking strategy to your RAG pipeline, ask yourself: "Am I trying to retrieve information, or am I trying to teach the model a behavior?" If it’s the latter, LoRA will ship faster, perform better, and cost less.
Try it on your next extraction or classification task. You’ll be surprised how far a small fine-tuned model can go.