Skip to content

Fine-Tuning LLMs

When and how to customize a pre-trained model on your own data.

Last reviewed · Download PDF

Prerequisites: LLM APIs · Evals

Related: Local LLMs · MLflow · RAG · Glossary


Overview

Challenge: A pre-trained LLM has broad general knowledge but no knowledge of an organization's tone, domain terminology, or required output formats.

Solution: fine-tuning continues training the model on examples specific to the use case. Given hundreds or thousands of (input, ideal output) pairs, the model adjusts its weights to produce outputs closer to those examples.

Pre-trained model:       Knows everything generally
Fine-tuned model:        Knows your specific thing very well

Examples of what fine-tuning fixes:
  "Always respond in SQL, never prose"
  "Use our internal table naming convention"
  "Output JSON that matches our exact schema"
  "Write in our company's brand voice"
  "Classify support tickets into our 40 internal categories"

Fine-tuning compared with RAG:

Use RAG when:
  - Your knowledge base changes frequently
  - You need source citations
  - You want to add new facts the model doesn't know

Use fine-tuning when:
  - You need a consistent output format the model ignores in prompts
  - You need a specific tone or style the model doesn't adopt
  - You're classifying into custom categories not in the base model
  - You need faster inference (smaller fine-tuned model > larger base model)
  - RAG works but the model still doesn't follow instructions reliably

Use both when:
  - Fine-tune for behavior/format, RAG for knowledge
flowchart TB
    Q{"Model needs to<br/>change behaviour or style?"} -->|"no: needs facts"| RAG["Use RAG"]
    Q -->|"yes"| P{"Prompting + examples<br/>enough?"}
    P -->|"yes"| PR["Improve the prompt"]
    P -->|"no"| FT["Fine-tune<br/>LoRA / API"]
    FT --> EV["Evaluate against baseline"]

On this page

Basic - Core Concepts - When Fine-Tuning Helps (and When It Doesn't) - Preparing Training Data

Intermediate - Fine-Tuning with OpenAI - Fine-Tuning with Hugging Face - LoRA / PEFT (Parameter-Efficient Fine-Tuning)

Advanced - Evaluating Fine-Tuned Models - Dataset Construction Patterns - Production Considerations - Common Pitfalls

Reference - Cheat Sheet - Interview Questions - Further Reading


Core Concepts

Concept Description
Full fine-tuning Update all model weights — most powerful, most expensive, needs multiple large GPUs for a 7B+ model
LoRA Update only a tiny fraction of weights via low-rank matrices — far cheaper, often close to full fine-tuning quality
PEFT Parameter-Efficient Fine-Tuning — umbrella term for LoRA and similar techniques
Training data (prompt, completion) pairs showing the model what good output looks like
Epochs How many times the model trains over your entire dataset
Overfitting Model memorizes training examples instead of learning the pattern — use a validation set
Base model The starting point — a pre-trained model you fine-tune from
Adapter A small trained add-on (LoRA) attached to the base model — easy to swap

When Fine-Tuning Helps (and When It Doesn't)

# Signs fine-tuning is the right call:
GOOD_CANDIDATES = [
    "Output format: model ignores JSON schema even with detailed prompts",
    "Style: model writes formally but you need casual/brand voice",
    "Classification: 40+ custom categories not in base model's vocabulary",
    "Extraction: model misses domain-specific entities (internal product names)",
    "Latency: need a smaller, faster model for high-volume inference",
    "Cost: a small fine-tuned model can cost less per token than a frontier API model at high volume",
]

# Signs fine-tuning won't help:
BAD_CANDIDATES = [
    "Knowledge: model doesn't know facts that change weekly → use RAG",
    "Hallucination: model makes things up → fine-tuning won't fix grounding",
    "Small dataset: < 50 examples → few-shot prompting is enough",
    "One-off task: task runs once per day → prompting is cheaper",
    "No eval: you can't measure if fine-tuning improved things",
]

Preparing Training Data

Training data quality matters far more than quantity. 50 excellent examples beat 5000 mediocre ones.

import json
from pathlib import Path

# OpenAI fine-tuning format: JSONL, one example per line
# Each example is a full conversation

def create_training_example(user_msg: str, assistant_msg: str,
                             system: str = None) -> dict:
    messages = []
    if system:
        messages.append({"role": "system", "content": system})
    messages.append({"role": "user",      "content": user_msg})
    messages.append({"role": "assistant", "content": assistant_msg})
    return {"messages": messages}

# Example: fine-tune for SQL generation
examples = [
    create_training_example(
        system="You are a SQL expert. Generate ANSI SQL only. No explanations.",
        user_msg="Count orders by status for the last 30 days",
        assistant_msg="""SELECT status, COUNT(*) AS order_count
FROM orders
WHERE created_at >= DATEADD(day, -30, CURRENT_TIMESTAMP())
GROUP BY status
ORDER BY order_count DESC;"""
    ),
    create_training_example(
        system="You are a SQL expert. Generate ANSI SQL only. No explanations.",
        user_msg="Show me the top 10 customers by revenue this year",
        assistant_msg="""SELECT c.customer_id, c.name, SUM(o.amount) AS total_revenue
FROM orders o
JOIN customers c ON o.customer_id = c.id
WHERE YEAR(o.created_at) = YEAR(CURRENT_DATE())
GROUP BY c.customer_id, c.name
ORDER BY total_revenue DESC
LIMIT 10;"""
    ),
    # ... add 50-200 more examples
]

# Split first, so no validation example also appears in the training file
# (shuffle before splitting if the examples are ordered)
split = int(len(examples) * 0.8)

with open("training_data.jsonl", "w") as f:
    for ex in examples[:split]:
        f.write(json.dumps(ex) + "\n")

# Validation file (10-20% of the data)
with open("validation_data.jsonl", "w") as f:
    for ex in examples[split:]:
        f.write(json.dumps(ex) + "\n")

Data quality checklist

def validate_training_data(filepath: str) -> dict:
    """Check training data for common problems."""
    issues = []
    examples = []
    total_tokens = 0

    with open(filepath) as f:
        for i, line in enumerate(f):
            try:
                ex = json.loads(line)
                examples.append(ex)
            except json.JSONDecodeError:
                issues.append(f"Line {i}: invalid JSON")
                continue

            msgs = ex.get("messages", [])

            # Must have user + assistant turn
            roles = [m["role"] for m in msgs]
            if "user" not in roles:
                issues.append(f"Line {i}: missing user message")
            if "assistant" not in roles:
                issues.append(f"Line {i}: missing assistant message")

            # Estimate token count
            text_length = sum(len(m["content"]) for m in msgs)
            total_tokens += text_length // 4  # rough estimate

    avg_tokens = total_tokens // max(len(examples), 1)

    return {
        "total_examples": len(examples),
        "issues":         issues,
        "avg_tokens":     avg_tokens,
        "ready":          len(issues) == 0 and len(examples) >= 10
    }

result = validate_training_data("training_data.jsonl")
print(result)

Fine-Tuning with OpenAI

from openai import OpenAI
import time

client = OpenAI()

# ── 1. Upload training data ────────────────────────────────────────────────────
training_file = client.files.create(
    file=open("training_data.jsonl", "rb"),
    purpose="fine-tune"
)
validation_file = client.files.create(
    file=open("validation_data.jsonl", "rb"),
    purpose="fine-tune"
)
print(f"Training file ID: {training_file.id}")

# ── 2. Create fine-tuning job ──────────────────────────────────────────────────
job = client.fine_tuning.jobs.create(
    training_file   = training_file.id,
    validation_file = validation_file.id,
    model           = "gpt-4.1-mini-2025-04-14",   # base model to fine-tune
    method = {                       # replaces the deprecated top-level `hyperparameters`
        "type": "supervised",
        "supervised": {
            "hyperparameters": {
                "n_epochs":        3,     # 3-5 is typical; more = higher overfitting risk
                "batch_size":      "auto",
                "learning_rate_multiplier": "auto",
            },
        },
    },
    suffix = "sql-generator"   # appears in the model name: ft:gpt-4.1-mini-...:my-org:sql-generator:<id>
)
print(f"Job ID: {job.id}, status: {job.status}")

# ── 3. Monitor progress ────────────────────────────────────────────────────────
while True:
    job = client.fine_tuning.jobs.retrieve(job.id)
    print(f"Status: {job.status}")
    if job.status in ("succeeded", "failed", "cancelled"):
        break
    # Check recent events
    for event in client.fine_tuning.jobs.list_events(job.id, limit=5).data:
        print(f"  [{event.created_at}] {event.message}")
    time.sleep(60)

print(f"Fine-tuned model: {job.fine_tuned_model}")
# ft:gpt-4.1-mini-2025-04-14:my-org:sql-generator:abc123

# ── 4. Use the fine-tuned model ────────────────────────────────────────────────
response = client.chat.completions.create(
    model=job.fine_tuned_model,
    messages=[
        {"role": "system",  "content": "You are a SQL expert. Generate ANSI SQL only."},
        {"role": "user",    "content": "Show total revenue by region for last quarter"}
    ]
)
print(response.choices[0].message.content)

Fine-Tuning with Hugging Face

For open-source models (Llama 3, Mistral, Gemma) on your own GPU or cloud VM.

pip install transformers datasets peft accelerate bitsandbytes trl
from datasets import Dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments
from peft import LoraConfig, get_peft_model
from trl import SFTConfig, SFTTrainer
import torch

MODEL_NAME = "meta-llama/Meta-Llama-3-8B-Instruct"   # gated: accept the licence on Hugging Face first; any causal LM works

# ── 1. Load model in 4-bit quantization (saves memory) ────────────────────────
from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    quantization_config=bnb_config,
    device_map="auto",
)

# ── 2. Apply LoRA ──────────────────────────────────────────────────────────────
lora_config = LoraConfig(
    r=16,               # rank — higher = more parameters, more capacity
    lora_alpha=32,      # scaling factor (usually 2*r)
    target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],  # which layers to adapt
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# Llama 3 8B, r=16, q/k/v/o_proj: trainable params: 13,631,488 (about 0.17% of the model).
# The 'all params' figure differs when the base model is loaded in 4-bit.

# ── 3. Prepare dataset ─────────────────────────────────────────────────────────
# Use the chat "messages" format: SFTTrainer then applies the model's own chat template, so
# training matches the prompt format used at inference. Hand-written special tokens
# (such as <|user|>) that the tokenizer does not define teach the model a format it never sees.
SYSTEM = "You are a SQL expert. Generate ANSI SQL only."

def to_messages(example):
    return {"messages": [
        {"role": "system",    "content": SYSTEM},
        {"role": "user",      "content": example["question"]},
        {"role": "assistant", "content": example["sql"]},
    ]}

raw_data = [
    {"question": "Count orders by status", "sql": "SELECT status, COUNT(*) FROM orders GROUP BY status;"},
    # ... more examples
]
dataset = Dataset.from_list(raw_data).map(to_messages, remove_columns=["question", "sql"])
splits  = dataset.train_test_split(test_size=0.1, seed=42)   # hold out data for eval

# ── 4. Train ───────────────────────────────────────────────────────────────────
# SFTConfig extends TrainingArguments with SFT-specific options.
# Tested with transformers 5.17, TRL 1.14 and PEFT 0.21; argument names change between major versions.
training_args = SFTConfig(
    output_dir          = "./fine-tuned-model",
    num_train_epochs    = 3,
    per_device_train_batch_size = 4,
    gradient_accumulation_steps = 4,
    warmup_steps        = 0.05,      # a float below 1 is a ratio of total steps (transformers 4.x: warmup_ratio)
    learning_rate       = 2e-4,
    bf16                = True,      # matches bnb_4bit_compute_dtype; use fp16=True on GPUs without bfloat16
    logging_steps       = 10,
    save_steps          = 100,
    eval_strategy       = "steps",   # was evaluation_strategy in older transformers
    eval_steps          = 100,
    max_length          = 2048,      # was max_seq_length in older TRL releases
)

trainer = SFTTrainer(
    model            = model,
    args             = training_args,
    train_dataset    = splits["train"],
    eval_dataset     = splits["test"],
    processing_class = tokenizer,
)

trainer.train()
trainer.save_model("./fine-tuned-model")

# ── 5. Merge LoRA weights into base model for deployment ──────────────────────
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained(MODEL_NAME, dtype=torch.float16)   # torch_dtype in transformers 4.x
merged = PeftModel.from_pretrained(base_model, "./fine-tuned-model")
merged = merged.merge_and_unload()
merged.save_pretrained("./merged-model")
tokenizer.save_pretrained("./merged-model")

LoRA / PEFT (Parameter-Efficient Fine-Tuning)

Why LoRA? Full fine-tuning keeps weights, gradients and optimizer state for every parameter: roughly 16 bytes per parameter with Adam in mixed precision, or about 110GB for a 7B model before activations. LoRA freezes the base weights and trains tiny added matrices, updating well under 1% of parameters, with quality that is comparable for many tasks at a fraction of the cost.

Full fine-tuning:
  Original weights W (7B params) → updated W' (7B params)
  GPU memory: ~16 bytes/param with Adam (~110GB for 7B)
  Training time: days

LoRA:
  Original weights W (frozen)
  Two small matrices A (r×d) and B (d×r) where r << d
  Update = W + A×B  (only A and B are trained)
  r=16 on the attention projections adds ~0.2% trainable parameters
  GPU memory: roughly 8-16GB for a 7-8B model with 4-bit quantization (QLoRA)
  Training time: hours
# LoRA hyperparameter guide
lora_config = LoraConfig(
    r=8,        # rank: 4-64; higher = more capacity, more memory
                # start with 8-16; increase if underfitting

    lora_alpha=16,   # scaling: usually 1-2x rank; controls magnitude of updates

    target_modules=["q_proj", "v_proj"],  # which attention layers to adapt
    # For most models: ["q_proj", "k_proj", "v_proj", "o_proj"]
    # For aggressive adaptation add: ["gate_proj", "up_proj", "down_proj"]

    lora_dropout=0.1,   # regularization: 0.05-0.1 typical
)

Evaluating Fine-Tuned Models

import anthropic
import json

def evaluate_model(model_fn, test_cases: list[dict]) -> dict:
    """
    model_fn: callable(prompt) → output string
    test_cases: [{"input": str, "expected": str, "check": callable}]
    """
    results = []
    for case in test_cases:
        output = model_fn(case["input"])
        passed = case["check"](output, case["expected"])
        results.append({"input": case["input"], "output": output, "passed": passed})

    pass_rate = sum(r["passed"] for r in results) / len(results)
    return {"pass_rate": pass_rate, "details": results}

# Test cases for SQL generation
sql_test_cases = [
    {
        "input": "Count orders by status",
        "expected": "SELECT status, COUNT(*)",
        "check": lambda output, expected: expected.lower() in output.lower()
    },
    {
        "input": "Top 5 customers by revenue",
        "expected": "LIMIT 5",
        "check": lambda output, expected: "LIMIT 5" in output.upper() and "ORDER BY" in output.upper()
    },
]

# Always evaluate:
# 1. On a held-out test set (not in training data)
# 2. Against the base model (did fine-tuning actually help?)
# 3. On adversarial inputs (does it handle edge cases?)
# 4. For regressions (did it lose general capability?)

Dataset Construction Patterns

# Pattern 1: Generate synthetic data with a stronger model
def generate_training_examples(task_description: str, n: int = 100) -> list[dict]:
    """Use Claude to generate (input, output) pairs for fine-tuning."""
    client = anthropic.Anthropic()
    response = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=4096,
        messages=[{"role": "user", "content": f"""
Generate {n} diverse training examples for this fine-tuning task:
{task_description}

Return as a JSON array: [{{"input": "...", "output": "..."}}]

Requirements:
- Vary complexity from simple to complex
- Include edge cases
- Outputs must be exactly correct
- No explanations in outputs — output only
"""}]
    )
    return json.loads(next(b.text for b in response.content if b.type == "text"))

# Pattern 2: Mine from existing system logs
def mine_from_logs(log_file: str) -> list[dict]:
    """Extract (user_query, good_response) pairs from production logs."""
    examples = []
    with open(log_file) as f:
        for line in f:
            entry = json.loads(line)
            if entry.get("feedback") == "thumbs_up":  # only use approved responses
                examples.append({
                    "input":  entry["user_message"],
                    "output": entry["assistant_message"]
                })
    return examples

# Pattern 3: Human-curated corrections
# Store: original_output, corrected_output, corrected_by, timestamp
# Fine-tune on: (input, corrected_output) pairs
# This is the highest quality data source

Production Considerations

# Serving a fine-tuned model
# Option 1: OpenAI fine-tuned model — just use the model ID
response = client.chat.completions.create(
    model="ft:gpt-4.1-mini-2025-04-14:my-org::abc123",   # returned as job.fine_tuned_model
    messages=[...]
)

# Option 2: Self-hosted with vLLM (open-source models)
# vllm serve ./merged-model --port 8000
# Then call like any OpenAI-compatible API:
from openai import OpenAI
local_client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
response = local_client.chat.completions.create(
    model="./merged-model",
    messages=[...]
)

# Cost comparison — work it out for your own volume rather than trusting rules of thumb:
#   API model:    daily_tokens / 1e6 × price_per_million        (input and output priced separately)
#                 e.g. 1M tokens/day at $2 per 1M  ≈  $2/day (look up current prices)
#   Fine-tuned:   training cost (one-off) + usually higher per-token inference price
#   Self-hosted:  GPU instance hours (paid even when idle) + engineering/ops time
# Self-hosting only wins at high, steady volume or when data can't leave your network.

Common Pitfalls

1. Fine-tuning to add knowledge
   Problem: model still hallucinates — fine-tuning teaches behavior, not facts
   Fix:     Use RAG for factual grounding; use fine-tuning for style/format

2. Too little data
   Problem: < 50 examples → model memorizes, doesn't generalize
   Fix:     Aim for 100-500 high-quality examples minimum; synthetic data helps

3. Low-quality training data
   Problem: inconsistent outputs → model learns inconsistency
   Fix:     Have a human review every training example; quality >> quantity

4. Skipping evaluation
   Problem: no way to know if fine-tuning helped or hurt
   Fix:     Build eval set before you start; measure base model vs fine-tuned

5. Over-training (too many epochs)
   Problem: model memorizes training set, fails on new inputs
   Fix:     Monitor validation loss; stop when it stops decreasing

6. No regression testing
   Problem: fine-tuning improves SQL but breaks summarization
   Fix:     Eval on diverse capabilities, not just the target task

7. Using fine-tuning as a shortcut for better prompting
   Problem: fine-tuning costs time and money
   Fix:     Exhaust prompt engineering first (few-shot, CoT, structured system prompt)

Cheat Sheet

Should you fine-tune? Try these first, in order

Step Fixes Cost
1. Better prompt + few-shot examples Format, tone, simple domain rules Minutes
2. Structured outputs / tool schemas Format reliability Minutes
3. RAG Missing or changing knowledge Days
4. A stronger model or higher effort Reasoning quality A config change
5. Fine-tuning Consistent style/format at scale, a narrow task on a small cheap model, lower latency Weeks: data, training, evals, hosting
Approach What changes Memory needed (7–8B model) When
Full fine-tune All weights Very high (multiple large GPUs) Large budgets; big domain shift
LoRA Small adapter matrices Moderate (one 24–48 GB GPU) Default for open models
QLoRA LoRA on a 4-bit quantized base Low (one 16–24 GB GPU) Limited hardware
Hosted API fine-tuning Provider-managed None Fastest path on supported models

LoRA knobs: r 8–64 (adapter capacity) · lora_alpha ≈ 2×r · lora_dropout 0.05–0.1 · target_modules attention projections (q_proj, v_proj, ...) · learning rate ~1e-4 to 2e-4 · 1–3 epochs

Training data checklist: hundreds to a few thousand high-quality examples · the same prompt format you'll use at inference · deduplicated · no test examples leaked into training · a held-out eval split · PII removed · edge cases and refusals included

Chat-format JSONL record

{"messages": [{"role": "system", "content": "You write ANSI SQL."}, {"role": "user", "content": "Count orders by status"}, {"role": "assistant", "content": "SELECT status, COUNT(*) FROM orders GROUP BY status;"}]}

Interview Questions

Q: What's the difference between fine-tuning and prompt engineering? A: Prompt engineering changes the input without modifying the model — fast, cheap, reversible. Fine-tuning changes the model's weights by training on examples — more powerful for consistent behavior and style, but requires data, compute, and eval infrastructure. Start with prompts; only fine-tune when prompts can't achieve the goal consistently.

Q: What is LoRA and why is it preferred over full fine-tuning? A: LoRA (Low-Rank Adaptation) freezes the pre-trained model weights and trains two small low-rank matrices that are added to the original weights. It updates well under 1% of parameters instead of 100%, cutting GPU memory from roughly 110GB to about 8-16GB for a 7-8B model when combined with 4-bit quantization (QLoRA), and training time from days to hours, with comparable quality on many tasks.

Q: When would you choose fine-tuning over RAG? A: RAG is better for knowledge (facts that change, need citations). Fine-tuning is better for behavior (consistent format, style, custom classifications, domain-specific extraction). Often the right answer is both: fine-tune the model for behavior, add RAG for knowledge grounding.


Further Reading


Previous: Claude Code · Next: AI Observability · Back to: Index