Juan Pablo García
All writing
Explainer5 steps · scroll to play

How fine-tuning with LoRA works

Fine-tuning adapts a pretrained model to your task. LoRA does it by freezing the original weights and training two small matrices beside them.

  1. 1

    Start from a pretrained model

    A language model is mostly large weight matrices. This one, W, has 4096 × 4096 entries, about 16.8 million numbers, and it already knows language.

  2. 2

    Full fine-tuning changes everything

    Classic fine-tuning updates every weight. The optimizer also keeps extra state per weight, so memory and cost grow with the whole model.

  3. 3

    Freeze W, add two thin matrices

    LoRA leaves W untouched and adds B (4096 × 8) and A (8 × 4096). Their product has W's shape but only 65,536 trainable numbers.

  4. 4

    Train only A and B

    Gradients flow to A and B alone. The change B·A can only express a few directions, its rank, which is usually enough for a new task.

  5. 5

    Merge or swap

    Add B·A into W for zero extra cost when serving, or keep adapters separate and swap them per task on one base model.

Start from a pretrained model

A language model is mostly large weight matrices. This one, W, has 4096 × 4096 entries, about 16.8 million numbers, and it already knows language.

Full fine-tuning changes everything

Classic fine-tuning updates every weight. The optimizer also keeps extra state per weight, so memory and cost grow with the whole model.

Freeze W, add two thin matrices

LoRA leaves W untouched and adds B (4096 × 8) and A (8 × 4096). Their product has W's shape but only 65,536 trainable numbers.

Train only A and B

Gradients flow to A and B alone. The change B·A can only express a few directions, its rank, which is usually enough for a new task.

Merge or swap

Add B·A into W for zero extra cost when serving, or keep adapters separate and swap them per task on one base model.

In short

  • LoRA trains well under 1% of the parameters, so it fits on much smaller GPUs than full fine-tuning.
  • The rank (8 here) sets how much the adapter can change. Higher rank, more capacity and more memory.
  • QLoRA goes further by storing the frozen base in 4-bit, so even larger models fit on one GPU.
  • Try prompting and RAG first. Fine-tune when you need a consistent style, format or behavior that prompts can't hold.

Have an AI feature to build? Let's talk for 15 minutes.