How fine-tuning with LoRA works
Fine-tuning adapts a pretrained model to your task. LoRA does it by freezing the original weights and training two small matrices beside them.
- 1
Start from a pretrained model
A language model is mostly large weight matrices. This one, W, has 4096 × 4096 entries, about 16.8 million numbers, and it already knows language.
- 2
Full fine-tuning changes everything
Classic fine-tuning updates every weight. The optimizer also keeps extra state per weight, so memory and cost grow with the whole model.
- 3
Freeze W, add two thin matrices
LoRA leaves W untouched and adds B (4096 × 8) and A (8 × 4096). Their product has W's shape but only 65,536 trainable numbers.
- 4
Train only A and B
Gradients flow to A and B alone. The change B·A can only express a few directions, its rank, which is usually enough for a new task.
- 5
Merge or swap
Add B·A into W for zero extra cost when serving, or keep adapters separate and swap them per task on one base model.
Start from a pretrained model
A language model is mostly large weight matrices. This one, W, has 4096 × 4096 entries, about 16.8 million numbers, and it already knows language.
Full fine-tuning changes everything
Classic fine-tuning updates every weight. The optimizer also keeps extra state per weight, so memory and cost grow with the whole model.
Freeze W, add two thin matrices
LoRA leaves W untouched and adds B (4096 × 8) and A (8 × 4096). Their product has W's shape but only 65,536 trainable numbers.
Train only A and B
Gradients flow to A and B alone. The change B·A can only express a few directions, its rank, which is usually enough for a new task.
Merge or swap
Add B·A into W for zero extra cost when serving, or keep adapters separate and swap them per task on one base model.
In short
- LoRA trains well under 1% of the parameters, so it fits on much smaller GPUs than full fine-tuning.
- The rank (8 here) sets how much the adapter can change. Higher rank, more capacity and more memory.
- QLoRA goes further by storing the frozen base in 4-bit, so even larger models fit on one GPU.
- Try prompting and RAG first. Fine-tune when you need a consistent style, format or behavior that prompts can't hold.