LoRA, QLoRA and PEFT
Freeze the model, train a small patch beside it, and get most of the benefit of fine-tuning for a sliver of the cost.
Part of the How models are made track on lAItest.
You can adapt a model with hundreds of billions of frozen weights by training a patch small enough to email.
Full fine-tuning updates everything. PEFT updates a sliver.
Parameter-efficient fine-tuning is the family name: freeze the pretrained weights, add or adjust a small number of new ones, train only those. The saving is not only compute. You never hold a second full copy of the model in memory, and the thing you ship at the end is the sliver, not a whole new model.
LoRA: a low-rank patch alongside each weight matrix.
LoRA leaves the original weight matrix untouched and learns a pair of much smaller matrices beside it, whose product is added to the output. The bet is that the change one task needs can be expressed in far fewer numbers than the original matrix holds. Paper: arXiv 2106.09685. QLoRA, arXiv 2305.14314, adds quantization, squashing the frozen base down to 4-bit numbers so the whole arrangement fits on hardware you can actually rent.
A common misconception
Commonly believed: LoRA is the budget option. With enough money you would always do a full fine-tune.
Actually: Cost is only part of it. Because the base is frozen, adapters are swappable: one base model in memory, many adapters applied per request, one per customer or task. Full fine-tuning gives you a separate model per task and a serving bill to match. A frozen base is also much harder to damage, since the original capabilities are still literally there, unchanged.
Why can one server host many LoRA-tuned variants of a model at once?
Answer: The base weights are shared and frozen, so only the small adapters differ. The expensive thing, the frozen base, is loaded once and shared by every request. Each variant is a small extra set of numbers applied on top, so switching customers means switching a patch rather than reloading a model. Quantization is a separate idea that QLoRA adds; it is not what makes adapters swappable.
In one sentence
LoRA trains a small patch beside a frozen model, which makes tuning cheap and, more importantly, makes many tuned variants cheap to serve.