PEFT and LoRA¶
Learning contract override: Prerequisite: transfer learning, frozen-backbone evaluation, and PyTorch optimization. Time: 75–90 minutes with the PEFT lab. Evidence: trainable-parameter count, held-out comparison, and a defended adaptation strategy.
What This Is¶
Parameter-Efficient Fine-Tuning (PEFT) is the family of methods that adapt a pretrained model without updating all of its weights. LoRA (Low-Rank Adaptation) is one of its most widely used methods. The premise is simple: learn a small update to a large frozen base model.
This topic is a deeper treatment of the PEFT surface introduced in Transfer and Fine-Tuning. Where that topic covers the four-depth adaptation ladder end to end, this topic goes one level down into the PEFT design choices: which layers to adapt, what rank to pick, how to compose multiple adapters, and how to inspect one.
When You Use It¶
- you have a pretrained model worth gigabytes and an adaptation task worth megabytes of new parameters
- you need to serve many fine-tunes of the same base model without paying GPU memory for each
- you want to switch domain/persona/task by swapping an adapter file rather than loading a new model
- you have a small-to-medium dataset and want to compare a lower-cost adaptation against full fine-tuning
- compute is the binding constraint and a full fine-tune would not finish inside your budget
Do Not Use It When¶
- the task differs structurally from pretraining (e.g., adapting a text LLM to predict tabular regressions) — PEFT cannot rescue the wrong inductive bias
- the task requires broad changes throughout the model and a validated full fine-tune is worth its added cost
- the final deployment target is weight-merged and latency-critical and your library does not support merge (most do)
PEFT Taxonomy¶
| method | what it inserts | typical params | notes |
|---|---|---|---|
| LoRA | low-rank update ΔW = BA on chosen linear layers |
often a small fraction of base | widely supported starting point |
| QLoRA | LoRA on a 4-bit quantized base | same as LoRA | lets you fine-tune larger models on consumer GPUs |
| DoRA | decomposes weight into magnitude + direction, adapts the direction with LoRA | slightly > LoRA | alternative to evaluate, not a universal drop-in win |
| Adapters (Houlsby, Pfeiffer) | bottleneck MLP inserted between transformer sublayers | commonly a small fraction of base | changes the forward architecture |
| Prefix / Prompt tuning | learned vectors supplied to attention layers or the input | commonly a very small fraction of base | performance depends strongly on model and task |
| IA³ | learned per-channel gates on activations | 0.01% of base | extremely cheap, limited expressiveness |
| BitFit | train only bias terms | 0.01–0.1% | surprisingly strong baseline for small tasks |
For a new project: start with LoRA. If base-model memory is the constraint, start with QLoRA. The rest are worth knowing but rarely the default.
LoRA — The Math In One Paragraph¶
For a frozen pretrained weight W ∈ ℝ^{d_out × d_in}, LoRA learns two small matrices A ∈ ℝ^{r × d_in} and B ∈ ℝ^{d_out × r} with rank r << min(d_in, d_out). The adapted weight at inference is:
W' = W + (α / r) · B A
Key facts this encodes:
- parameters: only
r(d_in + d_out)instead ofd_in · d_out - in the original recipe,
Ais initialized randomly andBto zero, soΔW = 0and the adapted layer initially matches the base layer αis a scaling hyperparameter — a common default isα = 2rorα = r; changing it is approximately equivalent to changing the effective learning rate on the adapter- at serving time you can merge:
W ← W + (α/r) B A, and there is no inference cost; the merge is reversible
LoRA In PyTorch — A Self-Contained Implementation¶
The canonical form, readable in fewer than 40 lines:
import torch
from torch import nn
class LoRALinear(nn.Module):
"""Low-rank update around a frozen nn.Linear."""
def __init__(self, base_linear: nn.Linear, r: int = 8, alpha: int = 16, dropout: float = 0.0):
super().__init__()
self.base = base_linear
for p in self.base.parameters():
p.requires_grad = False
in_f = base_linear.in_features
out_f = base_linear.out_features
self.A = nn.Parameter(torch.randn(r, in_f) * (1.0 / r**0.5)) # small random init
self.B = nn.Parameter(torch.zeros(out_f, r)) # zero so ΔW = 0 at step 0
self.scale = alpha / r
self.dropout = nn.Dropout(dropout) if dropout > 0 else nn.Identity()
def forward(self, x):
return self.base(x) + self.dropout(x @ self.A.T) @ self.B.T * self.scale
For a simple zero-update initialization, initialize one factor to zero and the other nonzero. Random A, zero B is the recipe used in the original LoRA paper. Zero A, random B can also learn, but it gives different first-step gradients and is not optimization-equivalent. If both factors are zero, both gradients are zero at initialization and learning cannot start; if both are random, the adapter changes the base output at step 0.
Two inspections you should run the first time you use this:
- the output at step 0 must equal
self.base(x)exactly (becauseB = 0). If not, you have a bug. sum(p.numel() for p in self.parameters() if p.requires_grad)must be exactlyr(in_f + out_f). If it is larger, your base is accidentally trainable.
Which Layers To Adapt¶
Not every linear layer is equal. On transformer architectures, priority order from strongest per-parameter return:
q_proj,v_projin attention — a configuration evaluated in the original LoRA paper; small and worth testing first- add
k_proj,o_proj— modest gains, still cheap - add the MLP projections (
gate_proj,up_proj,down_projin modern architectures) — expands capacity further - add token embeddings or LM head — consider when the task or vocabulary requires output-space adaptation
The default heuristic: {q_proj, v_proj} first. Expand to {q, k, v, o} if you have headroom. Only add the MLPs if the evaluation is still moving when you do.
Adapting every linear layer increases capacity and parameter count. It can help on some tasks and hurt or waste compute on others, so compare target-module sets rather than treating either choice as universal.
Rank And Alpha — The Two Knobs¶
rank r |
scale α |
behavior |
|---|---|---|
| 4 | 8 | tiny, strong floor for simple task / style adaptation |
| 8 | 16 | common starting configuration to validate |
| 16 | 32 | when the task is harder or data is larger; diminishing returns after |
| 32+ | 64+ | higher-capacity adapter; compare cost and quality with fuller tuning |
Tuning advice:
- tune
rfirst. Double it until validation stops improving. - holding
α / rfixed isolates the effect of rank during a sweep; 1 or 2 are common starting ratios, not laws - if validation loss is unstable, sweep both adapter learning rate and
α / r; they interact
Multiple Adapters, Merged And Unmerged¶
Two deployment shapes worth knowing:
- Hot-swap: keep the base frozen, swap LoRA weights per request to serve multiple fine-tunes from one GPU. No merge. This is the common case.
- Merged: after training, fold the adapter into the base weights and ship a single model. Zero inference overhead. Good when you only serve one specialization per process.
You can compose adapters additively in special cases — W' = W + ΔW_1 + ΔW_2 — but most composition bugs come from treating this as algebraically safe when the two adapters were trained independently. Default to running one adapter at a time.
Training Recipe¶
A working LoRA recipe for instruction-tuning a 7B-class model on 10⁴–10⁵ examples:
- optimizer:
AdamW,lr=1e-4to3e-4on the adapter (higher than full fine-tune — only the adapter learns) - weight decay: start with 0 on LoRA parameters, then validate; low rank constrains capacity but is not equivalent to weight decay
- schedule: cosine decay with a 3% warmup, no restart
- batch size: effective batch of 32–128 via gradient accumulation
- epochs: select by held-out behavior; 1–3 is only an initial sweep for this example scale
- gradient checkpointing: on (you saved memory on weights — spend some of it on longer context)
- mixed precision: bf16 if the GPU supports it, else fp16 with loss scaling
If you are using QLoRA (base quantized to 4-bit), keep the adapter itself in higher precision (bf16) and use paged optimizers for the host-side state.
What To Inspect¶
- parameter count — confirm it matches
r(d_in + d_out) × number_of_adapted_layers. A surprise here means yourtarget_modulessetting is wrong. - step-0 output vs. base model — in evaluation mode and with identical inputs, it should match within numerical tolerance. If it does not, the update is nonzero or the forward path differs.
- gradient norms on
Avs.B— both should be non-trivial after a few steps. If either stays near zero, the scaling is wrong. - merge test — train, then compare outputs before and after merge. They should agree to numerical precision. A drift means the merge math is wrong for your layer type (common on quantized bases).
- adapter size on disk — should be a few MB, not GB. If it is GB, you saved the base by accident.
Failure Pattern¶
A common LoRA failure is silent convergence to the pretrained model. The loss looks flat; validation looks identical to baseline. Inspect:
- LR too low (the adapter needs 5–10× the LR of a full fine-tune)
- wrong
target_modules— e.g., a library naming mismatch means no layers are actually being LoRA'd - frozen flags misapplied — everything is trainable, so you are accidentally doing a full fine-tune with a tiny LR, which goes nowhere
The second failure is train loss drops, validation does not. This is the classic too-large-rank-on-too-small-dataset trap. Cut r in half and rerun.
Common Mistakes¶
- initializing both factors nonzero — breaks the zero-update-at-initialization property
- initializing both factors to zero — both gradients start at zero, so training cannot begin
- merging the adapter at inference without verifying outputs match the pre-merge model
- applying weight decay to LoRA parameters — slow, unnecessary
- using the same LR for the base and the adapter when both are trainable (they shouldn't both be trainable)
- choosing
target_modulesby name-matching strings in the wrong library (HuggingFace vs. Lit vs. nanoGPT all use different names) - re-tuning on top of a merged adapter — can work, but track the merge chain; otherwise reproducibility breaks
Decision: LoRA vs. Full Fine-Tune vs. Prompt-Only¶
| option | when it wins | trade |
|---|---|---|
| prompt-only / in-context | task spec fits in context, low volume, or rapidly changing | no training cost; limited steerability |
| LoRA (PEFT) | 10³–10⁵ examples; task style or domain shift; need to serve many variants | small artifacts; must learn the tooling |
| full fine-tune | 10⁵+ examples; deep task shift; single specialization | biggest lift; highest cost; heaviest artifact |
A useful rule: if you cannot tell whether LoRA or full fine-tune wins, LoRA is the right call — it is cheaper to iterate, easier to revert, and the evidence will accumulate.
Practice¶
- Take a 1B–2B parameter base LLM. Apply the
LoRALinearwrapper toq_projandv_proj. Train for one epoch on a small instruction set. Report the parameter count and the adapter file size. - Run a step-0 zero-update check in evaluation mode—confirm the fresh adapter matches the base model within numerical tolerance.
- Sweep
r ∈ {4, 8, 16, 32}withα = 2r. Plot validation loss againstr. Pick the value where the curve bends. - Train two LoRAs for two different personas. Swap them at inference. Confirm the outputs follow the adapter currently loaded.
- Merge one adapter into the base. Measure inference latency before and after merge on a fixed prompt — the merged model should be measurably faster than the hot-swap path because the adapter matmul is gone.
- Deliberately break the recipe: set
Bto random, or picktarget_modules=[]. Confirm that each break matches the failure signature described above.
Runnable Example¶
Run the offline adapter workflow from the repository root:
.venv/bin/python labs/optimization-regularization-and-peft/src/peft_workflow.py
Compare frozen-head, partial-unfreeze, full-fine-tune, and adapter strategies. Inspect trainable-parameter counts and held-out quality together before claiming parameter efficiency.
Longer Connection¶
PEFT lives at the intersection of three topics already in the academy:
- Transfer and Fine-Tuning gives the four-depth adaptation ladder; PEFT is the fourth rung
- Optimizers and Regularization — the LR / weight-decay / warmup recipe here is a specialized instance of the full-model recipe, with different defaults because only a small subset of parameters is trainable
- Attention and Transformers — knowing where the
q_proj/v_projlayers are in a transformer is what lets the "which layers to adapt" section make sense
For the decision surface — LoRA vs. full fine-tune vs. RAG vs. prompt — read Text Generation and Language Models and Retrieval-Augmented Generation back to back. PEFT is one option inside that larger decision; it should never be chosen in isolation.