Home / Chapter 2 · Where a model’s knowledge comes from
    Last edited · 10 min read

    Use with AI

    Fine-tuning and LoRA

    Fine-tuning is further training of an existing model on your examples: it is good at changing how the model answers and poor at changing what it knows. LoRA does it cheaply by freezing the model and training a small correction, often under 1% of the weights.

    In plain wordsFine-tuning is on-the-job training for an experienced employee. Afterwards they write in the company’s format and tone, but they won’t memorise the price list; that’s what the binder on the desk is for, i.e. RAG. LoRA is corrections on a transparent sheet laid over the textbook: the original stays untouched, and the sheet can be lifted off or swapped for another.

    Pick a problem and see which rung to start from

    What it changes and what it doesn’t

    Types and where to train

    How LoRA works

    Change the rank and the model class. See how many weights LoRA trains and how much GPU memory it needs

    of the weights that full fine-tuning changes

    GPU memory for weights and optimiser state:

    Illustrative numbers. Architecture as in Llama 3.1 8B and 70B, LoRA on all seven linear matrices in every layer (Q, K, V, O and three in the MLP). Memory by rule of thumb: full fine-tuning with Adam in mixed precision is about 16 bytes per parameter (weights and gradients in BF16, an FP32 copy of the weights, two Adam moments), LoRA is 2 bytes per frozen weight in BF16, QLoRA about 0.5 bytes (4 bits), plus 16 bytes for every adapter parameter. Activations come on top: they grow with sequence length and batch size. Cards chosen with 20% headroom for activations.

    Data, evaluation, maintenance

    Check yourself

    When would you use fine-tuning instead of a prompt or RAG, and how does LoRA work?

    Order: prompt and context first, RAG for knowledge, fine-tuning last. Fine-tuning is good at behaviour: format, tone, a narrow task, or distilling a large model’s behaviour into a small one. It is poor at facts: they go in unreliably, without a source, every change means retraining, and they can increase confident errors. LoRA freezes the weights and learns a low-rank update ΔW = B·A of rank r, so r·(d + k) parameters instead of d·k. An adapter is tens of MB, so many adapters can be served on one base. QLoRA does the same on a 4-bit base.

    Po polsku

    Kolejność: prompt i kontekst, dla wiedzy RAG, fine-tuning na końcu. Fine-tuning dobrze uczy zachowania: formatu, tonu, wąskiego zadania, albo przenosi zachowanie dużego modelu do małego. Faktów uczy słabo: wchodzą niepewnie, bez źródła, każda zmiana to nowy trening, a do tego mogą zwiększyć liczbę pewnych siebie błędów. LoRA zamraża wagi i uczy poprawki ΔW = B·A o niskim ranku r, czyli r·(d + k) parametrów zamiast d·k. Adapter waży dziesiątki MB, więc na jednej bazie serwuje się wiele adapterów. QLoRA robi to samo na bazie w 4 bitach.

    Follow-up questions (5)
    You have 50k company documents that change every week. Fine-tuning or RAG?
    RAG. Facts from fine-tuning go in unreliably, have no source, and every change needs a new training run. Fine-tuning can come later to teach the model the answer format and how to cite passages, but the knowledge stays in the index.
    How do you choose the LoRA rank and the layers to apply it to?
    All linear matrices, including the MLP, because attention-only LoRA clearly lags behind. Rank is the adapter’s capacity: on small and medium datasets a low rank is enough, in RL even r = 1, while on large datasets LoRA starts losing to full fine-tuning. The learning rate is about 10 times higher than for full fine-tuning; compare a few ranks on the eval.
    You have 200 customers and each wants a model in their own style. How do you serve that?
    One base and one LoRA adapter per customer, without merging. vLLM picks the adapter for each request, keeps several active adapters in one batch next to a single copy of the base and loads more on the fly. The price is a little extra compute at every step instead of 200 full copies of the model.
    After fine-tuning the format is perfect, but the model does worse outside the task and is easier to talk into things it used to refuse. What happened?
    Catastrophic forgetting and weakened safeguards: training pushed the weights only towards your examples. Fewer epochs, a lower learning rate or LoRA instead of full fine-tuning help, as do general examples and refusals in the data, with an out-of-task set and safety tests in the eval.
    The provider retires the base model your fine-tune is built on. What do you do?
    First check whether the new base with just a prompt already passes the eval, because then fine-tuning is no longer needed. If not, retrain on the new base with the same pipeline. Data, configuration and the eval set are versioned, so it is a single run.

    Sources

    Report an error · Suggest a fix