USAAIO Lesson 101, from Phase 3, on LoRA - Low-Rank Adaptation - for parameter-efficient fine-tuning. It covers the low-rank decomposition delta_W=AB, initializing B to zero, the alpha-over-r scaling, which layers to apply LoRA to, and a full LoRALinear nn.Module implementation. All the parameter counts - 589,824 for a full 768×768 layer against 12,288 for rank-8 LoRA, a 48-fold reduction - and the forward-pass arithmetic were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 26 slides.
Subject: Machine Learning · 52 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 101 · Phase 3
Fine-tune a 175B-parameter model by training only 0.01% of the weights. LoRA decomposes the weight update into two tiny matrices and freezes everything else — same performance, 48× fewer trainable parameters on a 768×768 layer.
Objectives
delta_W = A @ B and derive the exact trainable parameter count for any (d_in, d_out, rank) tripleN(0, sigma), and what the model outputs at step 0LoRALinear nn.Module that wraps a frozen base weight with trainable A and B, applies the alpha/r scaling factor, and runs a correct forward passWarm-up
Discussion prompt
Before we open Lesson 101: LoRA — Low-Rank Adaptation: without looking back, what was the main idea of BERT Fine-Tuning Strategies, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
feature extraction vs full fine-tuning, layer-wise learning rates, two-stage domain adaptation, and catastrophic forgetting with elastic weight consolidation (EWC).
Section
Part 1 of 4
Concept
GPT-3 has 175 billion parameters. Full fine-tuning trains every one of them, storing optimizer states (momentum + variance) at ~3× the model size — roughly 2 TB of GPU memory for Adam. Most teams can't afford that.
| Model | Params | Adam states | Total GPU memory |
|---|---|---|---|
| GPT-2 (small) | 117 M | 234 M floats | ~1.3 GB |
| GPT-3 | 175 B | 350 B floats | ~2 TB |
| GPT-3 + LoRA r=8 | 175 B frozen | ~4.7 M trainable | ~19 MB for A/B |
LoRA (Hu et al. 2022) solves this by freezing the pretrained weights and injecting a trainable low-rank delta into selected layers. Only A and B are ever updated.
Comparison
Comparison matrix
From Full fine-tuning is parameter-prohibitive: refill the Params column from what you know. The rest of the table is as it appeared.
| Model | Params | Adam states | Total GPU memory |
|---|---|---|---|
| GPT-2 (small) | 117 M | 234 M floats | ~1.3 GB |
| GPT-3 | 175 B | 350 B floats | ~2 TB |
| GPT-3 + LoRA r=8 | 175 B frozen | ~4.7 M trainable | ~19 MB for A/B |
Intuition
The key insight (Aghajanyan et al. 2020): when you fine-tune a large pretrained model, the weight delta W_final - W_0 has very low intrinsic rank. Most of the information sits in a tiny subspace — the rest is near-zero noise.
If delta_W is low-rank, you can represent it exactly as a product of two thin matrices A (d × r) and B (r × d) where r << d. That's LoRA: instead of learning all d² entries of delta_W, you learn only 2·d·r entries.
\[ \text{rank}(\Delta W) \ll d \implies \Delta W \approx AB,\quad A \in \mathbb{R}^{d \times r},\ B \in \mathbb{R}^{r \times d} \]
Section
Part 2 of 4
Concept
\[ y = x W_0^\top + \frac{\alpha}{r}\, x A B \]
W_0 — pretrained weight matrix, frozen, never updatedA \in \mathbb{R}^{d_{in} \times r} — initialized N(0, 0.01), trainableB \in \mathbb{R}^{r \times d_{out}} — initialized zero, trainablealpha/r — scaling hyperparameter (default alpha = 2r, so scale = 2)At step 0: B = 0 so xAB = 0 and the model output is exactly xW_0 — the pretrained model is unchanged. Fine-tuning only deviates from pretrained behavior as B learns nonzero values.
Counterexample
Discussion prompt
At step 0: B = 0 so xAB = 0 and the model output is exactly xW_0 — the pretrained model is unchanged. Fine-tuning only deviates from pretrained behavior as B learns nonzero values.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
For a single weight matrix of shape (d_in, d_out) with LoRA rank r, the trainable parameter count is exactly d_in × r + r × d_out. Full fine-tuning uses d_in × d_out.
| Config | Full params | LoRA params (rank=8) | Reduction |
|---|---|---|---|
| 768×768 (GPT-2 Q) | 589,824 | 12,288 | 48.0x |
| 768×768 (rank=4) | 589,824 | 6,144 | 96.0x |
| 768×768 (rank=16) | 589,824 | 24,576 | 24.0x |
| 768×768 (rank=1) | 589,824 | 1,536 | 384.0x |
\[ \text{params}_{\text{LoRA}} = d_{\text{in}} \cdot r + r \cdot d_{\text{out}} = r(d_{\text{in}} + d_{\text{out}}) \]
Trade off
Comparison matrix
From Counting trainable parameters: every row here is a choice with a cost. Fill the Reduction column, then say which row you would actually pick and what you give up for it.
| Config | Full params | LoRA params (rank=8) | Reduction |
|---|---|---|---|
| 768×768 (GPT-2 Q) | 589,824 | 12,288 | 48.0x |
| 768×768 (rank=4) | 589,824 | 6,144 | 96.0x |
| 768×768 (rank=16) | 589,824 | 24,576 | 24.0x |
| 768×768 (rank=1) | 589,824 | 1,536 | 384.0x |
Ranking
Put in order
Put the moves of Worked example: param savings on a 768×768 weight into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. GPT-2 small attention Q projection: d_in = d_out = 768, rank r = 8, alpha = 16.
Worked example
Identify d_in, d_out, rank
Why: GPT-2 small attention Q projection: d_in = d_out = 768, rank r = 8, alpha = 16.
Compute full-tuning params: d_in × d_out = 768 × 768 = 589,824
Why: Every entry of W_0 would need a gradient and an optimizer state.
Compute LoRA params: A is 768×8 = 6,144; B is 8×768 = 6,144; total = 12,288
Why: r(d_in + d_out) = 8 × (768 + 768) = 8 × 1536 = 12,288.
Compute reduction: 589,824 / 12,288 = 48.0x; 97.9% fewer trainable params
Why: The scaling factor alpha/r = 16/8 = 2.0 does not add parameters — it is a fixed scalar multiplier.
| Quantity | Value |
|---|---|
| d_in = d_out | 768 |
| rank r | 8 |
| alpha | 16 |
| scale alpha/r | 2.0 |
| A params (768×8) | 6,144 |
| B params (8×768) | 6,144 |
| Total LoRA params | 12,288 |
| Full weight params | 589,824 |
| Reduction | 48.0x (97.9% fewer) |
Blank canvas
Draw it
Draw what Worked example: param savings on a 768×768 weight just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Initialize A and B both from N(0, 0.01), then add the LoRA adapter and start fine-tuning.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.
Initialize A ~ N(0, 0.01), B = zeros. Then add LoRA and start fine-tuning.
Why: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.
Trap
Initialize A and B both from N(0, 0.01), then add the LoRA adapter and start fine-tuning.
At step 0: y = xW0 + (alpha/r) * x @ A @ B
Why: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.
Result: the adapter corrupts the pretrained representation from step 0. Loss starts high and convergence is unstable — you've lost the quality of the pretrained model.
Initialize A ~ N(0, 0.01), B = zeros. Then add LoRA and start fine-tuning.
At step 0: y = xW0 + (alpha/r) * x @ A @ zeros = xW0
Why: B=0 guarantees delta_W = A@B = 0 exactly. The model starts from the pretrained output distribution and deviates only as the task gradient flows into B and A.
Result: fine-tuning starts from a stable pretrained baseline. Convergence is fast and the final model retains pretrained general capability while adapting to the task.
Break the constraint
Discussion prompt
The rule this trap just fixed:
B=0 guarantees delta_W = A@B = 0 exactly. The model starts from the pretrained output distribution and deviates only as the task gradient flows into B and A.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.
Section
Part 3 of 4
Concept
Hu et al. applied LoRA to the query (Q) and value (V) projection matrices in every attention head. These are where fine-tuning signal concentrates most efficiently in practice.
| Layer | Apply LoRA? | Reason |
|---|---|---|
| Attention Q, V projections | Yes (recommended) | Hu et al. default; most effective per-param |
| Attention K projection | Optional | Adds params, marginal gain |
| Attention output proj | Optional | Sometimes helpful on harder tasks |
| FFN layers | Optional | Larger matrices; helps on complex tasks |
| Embedding, LayerNorm | No | Very few params; full tuning is fine |
In the PEFT library (Lesson 102), target_modules=["q_proj", "v_proj"] sets this automatically for HuggingFace models.
Discrimination
Sort into buckets
Sort these by Apply LoRA?, from memory, without looking back at Which layers get LoRA adapters?. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Pattern
Predict first
The table runs: linear.weight (W0) | [16, 16] | False | 256 · A | [16, 4] | True | 64 · B | [4, 16] | True | 64 · Total | — | — | 384 · Trainable | — | True | 128 (33.3%)
In Implementing LoRALinear from scratch, given the rows so far: what is the next one — the row where Variable is Output y?
Correct: Output y | [3, 16] | — | —
| Variable | Shape | Requires grad | Num params |
|---|---|---|---|
| linear.weight (W0) | [16, 16] | False | 256 |
| A | [16, 4] | True | 64 |
| B | [4, 16] | True | 64 |
| Total | — | — | 384 |
| Trainable | — | True | 128 (33.3%) |
| Output y | [3, 16] | — | — |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. nn.Parameter registers A and B in the optimizer.
Worked example
Define LoRALinear with frozen base weight and trainable A, B
Why: nn.Parameter registers A and B in the optimizer. Setting requires_grad=False on the base linear's parameters prevents its update.
import torch
import torch.nn as nn
class LoRALinear(nn.Module):
def __init__(self, in_features, out_features, rank=8, alpha=16):
super().__init__()
self.linear = nn.Linear(in_features, out_features, bias=False)
self.A = nn.Parameter(torch.randn(in_features, rank) * 0.01)
self.B = nn.Parameter(torch.zeros(rank, out_features))
self.scale = alpha / rank
for p in self.linear.parameters():
p.requires_grad = False
def forward(self, x):
return self.linear(x) + (x @ self.A @ self.B) * self.scale
torch.manual_seed(42)
layer = LoRALinear(16, 16, rank=4, alpha=8)
total = sum(p.numel() for p in layer.parameters())
train = sum(p.numel() for p in layer.parameters() if p.requires_grad)
frozen = total - train
print(f'total={total}, trainable={train}, frozen={frozen}')
x = torch.randn(3, 16)
y = layer(x)
print(f'y.shape={y.shape}')Trace parameter counts: total=384, trainable=128 (A+B), frozen=256 (W0)
Why: Linear(16,16) has 16×16=256 params. A is 16×4=64, B is 4×16=64, so A+B=128. 128/384 = 33.3% trainable in this toy example; real models use much larger W0 so the ratio shrinks dramatically.
| Variable | Shape | Requires grad | Num params |
|---|---|---|---|
| linear.weight (W0) | [16, 16] | False | 256 |
| A | [16, 4] | True | 64 |
| B | [4, 16] | True | 64 |
| Total | — | — | 384 |
| Trainable | — | True | 128 (33.3%) |
| Output y | [3, 16] | — | — |
Error analysis
Annotate
Walk the callouts on Implementing LoRALinear from scratch. Each one is a place this is easy to get subtly wrong.
Pattern
Predict first
The table runs: x | [1.0, 0.0, -1.0, 0.5] · A@B max abs (step 0) | 0.000000 · base = xW0 | [-0.6614, 0.4710, -0.3133, 0.3568] · delta = x@A@B*scale | [0.0, 0.0, 0.0, 0.0]
In Forward pass trace: LoRALinear(4→4, rank=2, alpha=4), given the rows so far: what is the next one — the row where Term is y_init = base + delta?
Correct: y_init = base + delta | [-0.6614, 0.4710, -0.3133, 0.3568]
| Term | Value (4-vec, rounded) |
|---|---|
| x | [1.0, 0.0, -1.0, 0.5] |
| A@B max abs (step 0) | 0.000000 |
| base = xW0 | [-0.6614, 0.4710, -0.3133, 0.3568] |
| delta = x@A@B*scale | [0.0, 0.0, 0.0, 0.0] |
| y_init = base + delta | [-0.6614, 0.4710, -0.3133, 0.3568] |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. At step 0, delta_W = A@B = 4×2 @ 2×4 = 4×4 matrix of all zeros.
Worked example
Set up: d=4, r=2, alpha=4, scale=2.0; B=zeros at init
Why: At step 0, delta_W = A@B = 4×2 @ 2×4 = 4×4 matrix of all zeros. The frozen path xW0 dominates.
import torch
torch.manual_seed(0)
d, r, alpha = 4, 2, 4
W0 = torch.randn(d, d) # pretrained, frozen
A = torch.randn(d, r)*0.01 # trainable, small init
B = torch.zeros(r, d) # trainable, zero init
scale = alpha / r # = 2.0
x = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
base = x @ W0.T
delta = (x @ A @ B) * scale
y_init = base + delta
print(f'W0:\n{W0.numpy().round(3)}')
print(f'A@B max_abs: {(A@B).abs().max().item():.6f}')
print(f'base = {base.numpy().round(4)}')
print(f'delta = {delta.numpy().round(6)}')
print(f'y_init= {y_init.numpy().round(4)}')| Term | Value (4-vec, rounded) |
|---|---|
| x | [1.0, 0.0, -1.0, 0.5] |
| A@B max abs (step 0) | 0.000000 |
| base = xW0 | [-0.6614, 0.4710, -0.3133, 0.3568] |
| delta = x@A@B*scale | [0.0, 0.0, 0.0, 0.0] |
| y_init = base + delta | [-0.6614, 0.4710, -0.3133, 0.3568] |
Confirm: y_init == xW0 exactly because B=0 at step 0
Why: This is the zero-init guarantee: the adapter adds nothing at initialization, so the pretrained model's output distribution is perfectly preserved before any gradient step.
Comparison
Comparison matrix
From Forward pass trace: LoRALinear(4→4, rank=2, alpha=4): refill the Value (4-vec, rounded) column from what you know. The rest of the table is as it appeared.
| Term | Value (4-vec, rounded) |
|---|---|
| x | [1.0, 0.0, -1.0, 0.5] |
| A@B max abs (step 0) | 0.000000 |
| base = xW0 | [-0.6614, 0.4710, -0.3133, 0.3568] |
| delta = x@A@B*scale | [0.0, 0.0, 0.0, 0.0] |
| y_init = base + delta | [-0.6614, 0.4710, -0.3133, 0.3568] |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Pass all model parameters to the optimizer without filtering.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.
Filter to only requires_grad=True parameters before constructing the optimizer.
Why: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.
Trap
Pass all model parameters to the optimizer without filtering.
opt = torch.optim.Adam(model.parameters(), lr=1e-4)
Why: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.
Result: you've just re-implemented full fine-tuning. Memory savings from LoRA are completely negated.
Filter to only requires_grad=True parameters before constructing the optimizer.
opt = torch.optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), lr=1e-4)
Why: Only A and B (rank × 2 × d params per layer) enter the optimizer. For rank=8 on a 768×768 layer this is 12,288 params vs 589,824 — 48x fewer Adam states.
Result: the optimizer is tiny, backward pass skips frozen weights, and memory usage matches the LoRA param count — not the base model size.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
(d_in, d_out) with LoRA rank r, the trainable parameter count is exactly d_in × r + r × d_out. Full fine-tuning uses d_in × d_out.; In the PEFT library (Lesson 102), target_modules=["q_proj", "v_proj"] sets this automatically for HuggingFace models.N(0, 0.01), then add the LoRA adapter and start fine-tuning.; Pass all model parameters to the optimizer without filtering.Section
Part 4 of 4
Concept
The delta is scaled by alpha/r before being added to the base output. This decouples the learning-rate sensitivity from the rank choice: changing r without changing alpha adjusts both the delta magnitude and the effective learning rate.
\[ y = x W_0^\top + \frac{\alpha}{r}\, x A B \]
| alpha | r | scale = alpha/r | Effect |
|---|---|---|---|
| 16 | 8 | 2.0 | Hu et al. default; balanced update magnitude |
| 32 | 8 | 4.0 | Larger delta; effectively higher LR for adapter |
| 16 | 16 | 1.0 | Lower scale; smaller delta per gradient step |
| r | r | 1.0 | Common convention: set alpha = r to get scale 1 |
Constraint
Discussion prompt
Run LoRA implementation pattern with this step confiscated:
Initialize: A ~ N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weightrequires_grad = False on all .linear.parameters()N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0Adam(filter(lambda p: p.requires_grad, model.parameters()))W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overheadPattern
nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weightrequires_grad = False on all .linear.parameters()N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0Adam(filter(lambda p: p.requires_grad, model.parameters()))W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overheadEdge cases
Discussion prompt
LoRA implementation pattern works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weightrequires_grad = False on all .linear.parameters()N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0Adam(filter(lambda p: p.requires_grad, model.parameters()))W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overheadElimination
Eliminate the wrong options
A LoRA adapter is applied to a 512×512 weight matrix with rank r=8. How many trainable parameters does it add?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A is shape 512×8 = 4,096 params; B is shape 8×512 = 4,096 params; total = 8,192. The full matrix would be 512×512 = 262,144, a 32x reduction. Formula: r(d_in + d_out) = 8×(512+512) = 8,192.
Check
Work this out on paper before clicking.
Check your understanding
A LoRA adapter is applied to a 512×512 weight matrix with rank r=8. How many trainable parameters does it add?
Answer: A
Why: A is shape 512×8 = 4,096 params; B is shape 8×512 = 4,096 params; total = 8,192. The full matrix would be 512×512 = 262,144, a 32x reduction. Formula: r(d_in + d_out) = 8×(512+512) = 8,192.
Prediction
Predict first
A LoRA layer has alpha=32 and rank=8. At initialization (B=zeros), what does the layer output for input x?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: x @ W0 exactly, because delta_W = A@B = 0
Why: B=zeros so x@A@B = x@A@0 = 0 regardless of A. The alpha/r scale multiplies zero, still giving zero. The output is x@W0 exactly — the pretrained model is unchanged at step 0.
Check
Think about what the model outputs at step 0.
Check your understanding
A LoRA layer has alpha=32 and rank=8. At initialization (B=zeros), what does the layer output for input x?
Answer: A
Why: B=zeros so x@A@B = x@A@0 = 0 regardless of A. The alpha/r scale multiplies zero, still giving zero. The output is x@W0 exactly — the pretrained model is unchanged at step 0.
Elimination
Eliminate the wrong options
LoRA with rank=16 is applied to BOTH FFN matrices in one GPT-2 small block (768→3072 and 3072→768). How many total trainable parameters does this add?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Layer 1 (768→3072, rank=16): 768×16 + 16×3072 = 12,288 + 49,152 = 61,440. Layer 2 (3072→768, rank=16): 3072×16 + 16×768 = 49,152 + 12,288 = 61,440. Total = 61,440 + 61,440 = 122,880.
Check
GPT-2 small has two FFN layers per block: a 768→3072 and a 3072→768 linear. Work this out before clicking.
Check your understanding
LoRA with rank=16 is applied to BOTH FFN matrices in one GPT-2 small block (768→3072 and 3072→768). How many total trainable parameters does this add?
Answer: A
Why: Layer 1 (768→3072, rank=16): 768×16 + 16×3072 = 12,288 + 49,152 = 61,440. Layer 2 (3072→768, rank=16): 3072×16 + 16×768 = 49,152 + 12,288 = 61,440. Total = 61,440 + 61,440 = 122,880.
Concept
During training, every forward pass computes two separate matrix multiplications: xW0 and x@A@B. At inference, this is unnecessary — the delta is constant, so you can merge it into W0 once.
\[ W_{\text{eff}} = W_0 + \frac{\alpha}{r} A B \]
After merging, the model has exactly the same architecture as the original pretrained model — no adapter overhead, same latency, same memory. The LoRA structure only exists during fine-tuning.
Analogy
Discussion prompt
Explain Merging LoRA weights for zero-latency inference by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
During training, every forward pass computes two separate matrix multiplications: xW0 and x@A@B. At inference, this is unnecessary — the delta is constant, so you can merge it into W0 once.
Explain it to yourself
Discussion prompt
In Full LoRA fine-tuning pipeline (synthetic classification) this move is made:
Build a TinyModel with a LoRALinear(16→32, rank=4, alpha=8) first layer
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
The base linear is frozen; only A and B (16×4 + 4×32 = 192 trainable params out of 576 total for the first layer) are updated.
Worked example
Build a TinyModel with a LoRALinear(16→32, rank=4, alpha=8) first layer
Why: The base linear is frozen; only A and B (16×4 + 4×32 = 192 trainable params out of 576 total for the first layer) are updated.
import torch, torch.nn as nn
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
class LoRALinear(nn.Module):
def __init__(self, i, o, rank=4, alpha=8):
super().__init__()
self.linear = nn.Linear(i, o, bias=False)
self.A = nn.Parameter(torch.randn(i, rank)*0.01)
self.B = nn.Parameter(torch.zeros(rank, o))
self.scale = alpha / rank
for p in self.linear.parameters(): p.requires_grad = False
def forward(self, x):
return self.linear(x) + (x @ self.A @ self.B) * self.scale
class TinyModel(nn.Module):
def __init__(self):
super().__init__()
self.fc1 = LoRALinear(16, 32, rank=4, alpha=8)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(32, 2)
def forward(self, x): return self.fc2(self.relu(self.fc1(x)))
torch.manual_seed(7)
X_np, y_np = make_classification(500, 16, n_informative=8, random_state=42)
X = torch.tensor(X_np, dtype=torch.float32)
y = torch.tensor(y_np, dtype=torch.long)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
model = TinyModel()
opt = torch.optim.Adam(
filter(lambda p: p.requires_grad, model.parameters()), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
for ep in range(100):
model.train()
loss = loss_fn(model(X_tr), y_tr)
opt.zero_grad(); loss.backward(); opt.step()
model.eval()
with torch.no_grad():
acc = (model(X_te).argmax(1) == y_te).float().mean().item()
print(f'test_acc={acc:.4f}')| Epoch | Train loss | Note |
|---|---|---|
| 0 | 0.7011 | B=0, model behaves like pretrained baseline |
| 20 | 0.6572 | B learns nonzero values; delta activates |
| 40 | 0.6101 | Gradient flows through A and B only |
| 60 | 0.5605 | Loss converges toward 0.5 on balanced binary task |
| 80 | 0.5227 | Stable convergence |
| 100 (test) | acc=0.7300 | 73.0% test accuracy with only A+B updating |
Trade off
Comparison matrix
From Full LoRA fine-tuning pipeline (synthetic classification): every row here is a choice with a cost. Fill the Train loss column, then say which row you would actually pick and what you give up for it.
| Epoch | Train loss | Note |
|---|---|---|
| 0 | 0.7011 | B=0, model behaves like pretrained baseline |
| 20 | 0.6572 | B learns nonzero values; delta activates |
| 40 | 0.6101 | Gradient flows through A and B only |
| 60 | 0.5605 | Loss converges toward 0.5 on balanced binary task |
| 80 | 0.5227 | Stable convergence |
| 100 (test) | acc=0.7300 | 73.0% test accuracy with only A+B updating |
Concept
| Method | Trainable params | GPU memory | Performance | Best for |
|---|---|---|---|---|
| Feature extraction | Head only (~0.01%) | Smallest | Good if task ≈ pretrain domain | Tiny task datasets |
| LoRA (r=8) | ~0.1–2% of model | Small | Near full fine-tune on most tasks | Limited GPU, many tasks |
| Full fine-tuning | 100% | Largest (3× model for Adam) | Best on large task datasets | Large data, single task |
LoRA achieves near-full fine-tuning quality on most NLP benchmarks while using 48–384× fewer trainable parameters for typical transformer layers. It also enables multi-task deployment: store one frozen base model and dozens of tiny LoRA checkpoints, swapping them at serving time.
Comparison
Comparison matrix
From LoRA vs full fine-tuning vs feature extraction: refill the Trainable params column from what you know. The rest of the table is as it appeared.
| Method | Trainable params | GPU memory | Performance | Best for |
|---|---|---|---|---|
| Feature extraction | Head only (~0.01%) | Smallest | Good if task ≈ pretrain domain | Tiny task datasets |
| LoRA (r=8) | ~0.1–2% of model | Small | Near full fine-tune on most tasks | Limited GPU, many tasks |
| Full fine-tuning | 100% | Largest (3× model for Adam) | Best on large task datasets | Large data, single task |
Concept
In Lesson 88 (BERT), you saw that the pretrained transformer's Q/K/V attention projections carry rich contextual representations. LoRA injects task-specific deltas into exactly those projections — the minimum change needed to steer an already-capable model.
In Lesson 89 (GPT), you saw how decoder-only models autoregressively generate text; LoRA is the standard recipe for instruction-tuning GPT-style models (e.g. Alpaca, LLaMA-2-chat) without full 16-bit fine-tuning.
In Lesson 102, you will use the HuggingFace peft library to apply LoRA with two lines of config, wrapping the techniques built from scratch here.
Counterexample
Discussion prompt
In Lesson 102, you will use the HuggingFace peft library to apply LoRA with two lines of config, wrapping the techniques built from scratch here.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why LoRA? — The fine-tuning bottleneck · The LoRA Math — decomposition, scaling, init · Implementation — LoRALinear nn.Module · Full pipeline — LoRA fine-tuning a classifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
y = xW0 + (alpha/r) * xAB — W0 is frozen, only A and B are trainedrank r on a d×d layer → 2dr trainable params vs d² full; for 768×768 rank-8: 12,288 vs 589,824 (48x fewer)W_eff = W0 + (alpha/r)AB) for zero-latency inferenceWant this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.