Lesson 101: LoRA — Low-Rank Adaptation

USAAIO Lesson 101, from Phase 3, on LoRA - Low-Rank Adaptation - for parameter-efficient fine-tuning. It covers the low-rank decomposition delta_W=AB, initializing B to zero, the alpha-over-r scaling, which layers to apply LoRA to, and a full LoRALinear nn.Module implementation. All the parameter counts - 589,824 for a full 768×768 layer against 12,288 for rank-8 LoRA, a 48-fold reduction - and the forward-pass arithmetic were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 26 slides.

Subject: Machine Learning · 52 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. LoRA Low-Rank Adaptation

Title

USAAIO · Lesson 101 · Phase 3

Fine-tune a 175B-parameter model by training only 0.01% of the weights. LoRA decomposes the weight update into two tiny matrices and freezes everything else — same performance, 48× fewer trainable parameters on a 768×768 layer.

2. By the end of this lesson you can

Objectives

  1. Explain why full fine-tuning is compute-prohibitive for billion-parameter models and what property LoRA exploits
  2. State the LoRA parameterization delta_W = A @ B and derive the exact trainable parameter count for any (d_in, d_out, rank) triple
  3. Explain why B is initialized to zero and A to N(0, sigma), and what the model outputs at step 0
  4. Implement a LoRALinear nn.Module that wraps a frozen base weight with trainable A and B, applies the alpha/r scaling factor, and runs a correct forward pass
  5. State which layers LoRA is typically applied to (Q and V projections) and why FFN layers are optional

3. What survived from BERT Fine-Tuning Strategies?

Warm-up

Discussion prompt

Before we open Lesson 101: LoRA — Low-Rank Adaptation: without looking back, what was the main idea of BERT Fine-Tuning Strategies, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

feature extraction vs full fine-tuning, layer-wise learning rates, two-stage domain adaptation, and catastrophic forgetting with elastic weight consolidation (EWC).

4. Why LoRA? — The fine-tuning bottleneck

Section

Part 1 of 4

5. Full fine-tuning is parameter-prohibitive

Concept

GPT-3 has 175 billion parameters. Full fine-tuning trains every one of them, storing optimizer states (momentum + variance) at ~3× the model size — roughly 2 TB of GPU memory for Adam. Most teams can't afford that.

ModelParamsAdam statesTotal GPU memory
GPT-2 (small)117 M234 M floats~1.3 GB
GPT-3175 B350 B floats~2 TB
GPT-3 + LoRA r=8175 B frozen~4.7 M trainable~19 MB for A/B

LoRA (Hu et al. 2022) solves this by freezing the pretrained weights and injecting a trainable low-rank delta into selected layers. Only A and B are ever updated.

6. Fill in: Params for Full fine-tuning is parameter-prohibitive

Comparison

Comparison matrix

From Full fine-tuning is parameter-prohibitive: refill the Params column from what you know. The rest of the table is as it appeared.

ModelParamsAdam statesTotal GPU memory
GPT-2 (small)117 M234 M floats~1.3 GB
GPT-3175 B350 B floats~2 TB
GPT-3 + LoRA r=8175 B frozen~4.7 M trainable~19 MB for A/B

7. Weight updates live in a low-rank subspace

Intuition

The key insight (Aghajanyan et al. 2020): when you fine-tune a large pretrained model, the weight delta W_final - W_0 has very low intrinsic rank. Most of the information sits in a tiny subspace — the rest is near-zero noise.

If delta_W is low-rank, you can represent it exactly as a product of two thin matrices A (d × r) and B (r × d) where r << d. That's LoRA: instead of learning all d² entries of delta_W, you learn only 2·d·r entries.

\[ \text{rank}(\Delta W) \ll d \implies \Delta W \approx AB,\quad A \in \mathbb{R}^{d \times r},\ B \in \mathbb{R}^{r \times d} \]

8. The LoRA Math — decomposition, scaling, init

Section

Part 2 of 4

9. The LoRA forward pass

Concept

\[ y = x W_0^\top + \frac{\alpha}{r}\, x A B \]

At step 0: B = 0 so xAB = 0 and the model output is exactly xW_0 — the pretrained model is unchanged. Fine-tuning only deviates from pretrained behavior as B learns nonzero values.

10. Break it if you can: The LoRA forward pass

Counterexample

Discussion prompt

At step 0: B = 0 so xAB = 0 and the model output is exactly xW_0 — the pretrained model is unchanged. Fine-tuning only deviates from pretrained behavior as B learns nonzero values.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

11. Counting trainable parameters

Concept

For a single weight matrix of shape (d_in, d_out) with LoRA rank r, the trainable parameter count is exactly d_in × r + r × d_out. Full fine-tuning uses d_in × d_out.

ConfigFull paramsLoRA params (rank=8)Reduction
768×768 (GPT-2 Q)589,82412,28848.0x
768×768 (rank=4)589,8246,14496.0x
768×768 (rank=16)589,82424,57624.0x
768×768 (rank=1)589,8241,536384.0x

\[ \text{params}_{\text{LoRA}} = d_{\text{in}} \cdot r + r \cdot d_{\text{out}} = r(d_{\text{in}} + d_{\text{out}}) \]

12. What each one costs: Counting trainable parameters

Trade off

Comparison matrix

From Counting trainable parameters: every row here is a choice with a cost. Fill the Reduction column, then say which row you would actually pick and what you give up for it.

ConfigFull paramsLoRA params (rank=8)Reduction
768×768 (GPT-2 Q)589,82412,28848.0x
768×768 (rank=4)589,8246,14496.0x
768×768 (rank=16)589,82424,57624.0x
768×768 (rank=1)589,8241,536384.0x

13. What has to happen first: Worked example: param savings on a 768×768 weight

Ranking

Put in order

Put the moves of Worked example: param savings on a 768×768 weight into the order they have to happen.

  1. Identify d_in, d_out, rank
  2. Compute full-tuning params: d_in × d_out = 768 × 768 = 589,824
  3. Compute LoRA params: A is 768×8 = 6,144; B is 8×768 = 6,144; total = 12,288
  4. Compute reduction: 589,824 / 12,288 = 48.0x; 97.9% fewer trainable params

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. GPT-2 small attention Q projection: d_in = d_out = 768, rank r = 8, alpha = 16.

14. Worked example: param savings on a 768×768 weight

Worked example

Identify d_in, d_out, rank

Why: GPT-2 small attention Q projection: d_in = d_out = 768, rank r = 8, alpha = 16.

Compute full-tuning params: d_in × d_out = 768 × 768 = 589,824

Why: Every entry of W_0 would need a gradient and an optimizer state.

Compute LoRA params: A is 768×8 = 6,144; B is 8×768 = 6,144; total = 12,288

Why: r(d_in + d_out) = 8 × (768 + 768) = 8 × 1536 = 12,288.

Compute reduction: 589,824 / 12,288 = 48.0x; 97.9% fewer trainable params

Why: The scaling factor alpha/r = 16/8 = 2.0 does not add parameters — it is a fixed scalar multiplier.

QuantityValue
d_in = d_out768
rank r8
alpha16
scale alpha/r2.0
A params (768×8)6,144
B params (8×768)6,144
Total LoRA params12,288
Full weight params589,824
Reduction48.0x (97.9% fewer)

15. Draw the shape of it: Worked example: param savings on a 768×768…

Blank canvas

Draw it

Draw what Worked example: param savings on a 768×768 weight just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

16. Something is wrong here: forgetting B=0 init means the model changes at step 0

Anomaly

Predict first

A student writes this, and it looks reasonable:

Initialize A and B both from N(0, 0.01), then add the LoRA adapter and start fine-tuning.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.

Initialize A ~ N(0, 0.01), B = zeros. Then add LoRA and start fine-tuning.

Why: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.

17. Trap: forgetting B=0 init means the model changes at step 0

Trap

The trap

Initialize A and B both from N(0, 0.01), then add the LoRA adapter and start fine-tuning.

At step 0: y = xW0 + (alpha/r) * x @ A @ B

Why: With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.

Result: the adapter corrupts the pretrained representation from step 0. Loss starts high and convergence is unstable — you've lost the quality of the pretrained model.

The fix

Initialize A ~ N(0, 0.01), B = zeros. Then add LoRA and start fine-tuning.

At step 0: y = xW0 + (alpha/r) * x @ A @ zeros = xW0

Why: B=0 guarantees delta_W = A@B = 0 exactly. The model starts from the pretrained output distribution and deviates only as the task gradient flows into B and A.

Result: fine-tuning starts from a stable pretrained baseline. Convergence is fast and the final model retains pretrained general capability while adapting to the task.

18. Break it on purpose: forgetting B=0 init means the model changes…

Break the constraint

Discussion prompt

The rule this trap just fixed:

B=0 guarantees delta_W = A@B = 0 exactly. The model starts from the pretrained output distribution and deviates only as the task gradient flows into B and A.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

With B ~ N(0, 0.01), xAB is nonzero — the model immediately deviates from the pretrained checkpoint before seeing any task data.

19. Implementation — LoRALinear nn.Module

Section

Part 3 of 4

20. Which layers get LoRA adapters?

Concept

Hu et al. applied LoRA to the query (Q) and value (V) projection matrices in every attention head. These are where fine-tuning signal concentrates most efficiently in practice.

LayerApply LoRA?Reason
Attention Q, V projectionsYes (recommended)Hu et al. default; most effective per-param
Attention K projectionOptionalAdds params, marginal gain
Attention output projOptionalSometimes helpful on harder tasks
FFN layersOptionalLarger matrices; helps on complex tasks
Embedding, LayerNormNoVery few params; full tuning is fine

In the PEFT library (Lesson 102), target_modules=["q_proj", "v_proj"] sets this automatically for HuggingFace models.

21. Which is which, by Apply LoRA?

Discrimination

Sort into buckets

Sort these by Apply LoRA?, from memory, without looking back at Which layers get LoRA adapters?. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

Yes (recommended)
Attention Q, V projections
Optional
Attention K projection; Attention output proj; FFN layers
No
Embedding, LayerNorm
g1
Apply LoRA? is "Yes (recommended)" for Attention Q, V projections — that is what the table on "Which layers get LoRA adapters?" records, and it is the single property separating this group from the rest.
g2
Apply LoRA? is "Optional" for Attention K projection, Attention output proj, FFN layers — that is what the table on "Which layers get LoRA adapters?" records, and it is the single property separating this group from the rest.
g3
Apply LoRA? is "No" for Embedding, LayerNorm — that is what the table on "Which layers get LoRA adapters?" records, and it is the single property separating this group from the rest.

22. Predict the next row: Implementing LoRALinear from scratch

Pattern

Predict first

The table runs: linear.weight (W0) | [16, 16] | False | 256 · A | [16, 4] | True | 64 · B | [4, 16] | True | 64 · Total | — | — | 384 · Trainable | — | True | 128 (33.3%)

In Implementing LoRALinear from scratch, given the rows so far: what is the next one — the row where Variable is Output y?

Correct: Output y | [3, 16] | — | —

VariableShapeRequires gradNum params
linear.weight (W0)[16, 16]False256
A[16, 4]True64
B[4, 16]True64
Total——384
Trainable—True128 (33.3%)
Output y[3, 16]——

Why: The relationship between the columns, not the individual numbers, is what generates the next row. nn.Parameter registers A and B in the optimizer.

23. Implementing LoRALinear from scratch

Worked example

Define LoRALinear with frozen base weight and trainable A, B

Why: nn.Parameter registers A and B in the optimizer. Setting requires_grad=False on the base linear's parameters prevents its update.

import torch
import torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(self, in_features, out_features, rank=8, alpha=16):
        super().__init__()
        self.linear = nn.Linear(in_features, out_features, bias=False)
        self.A = nn.Parameter(torch.randn(in_features, rank) * 0.01)
        self.B = nn.Parameter(torch.zeros(rank, out_features))
        self.scale = alpha / rank
        for p in self.linear.parameters():
            p.requires_grad = False

    def forward(self, x):
        return self.linear(x) + (x @ self.A @ self.B) * self.scale

torch.manual_seed(42)
layer = LoRALinear(16, 16, rank=4, alpha=8)
total   = sum(p.numel() for p in layer.parameters())
train   = sum(p.numel() for p in layer.parameters() if p.requires_grad)
frozen  = total - train
print(f'total={total}, trainable={train}, frozen={frozen}')
x = torch.randn(3, 16)
y = layer(x)
print(f'y.shape={y.shape}')

Trace parameter counts: total=384, trainable=128 (A+B), frozen=256 (W0)

Why: Linear(16,16) has 16×16=256 params. A is 16×4=64, B is 4×16=64, so A+B=128. 128/384 = 33.3% trainable in this toy example; real models use much larger W0 so the ratio shrinks dramatically.

VariableShapeRequires gradNum params
linear.weight (W0)[16, 16]False256
A[16, 4]True64
B[4, 16]True64
Total——384
Trainable—True128 (33.3%)
Output y[3, 16]——

24. Inspect it line by line: Implementing LoRALinear from scratch

Error analysis

Annotate

Walk the callouts on Implementing LoRALinear from scratch. Each one is a place this is easy to get subtly wrong.

  • nn.Parameter registers A and B in the optimizer. Setting requires_grad=False on the base linear's parameters prevents its update.
  • Linear(16,16) has 16×16=256 params. A is 16×4=64, B is 4×16=64, so A+B=128. 128/384 = 33.3% trainable in this toy example; real models use much larger W0 so the ratio shrinks dramatically.

25. Predict the next row: Forward pass trace: LoRALinear(4→4, rank=2…

Pattern

Predict first

The table runs: x | [1.0, 0.0, -1.0, 0.5] · A@B max abs (step 0) | 0.000000 · base = xW0 | [-0.6614, 0.4710, -0.3133, 0.3568] · delta = x@A@B*scale | [0.0, 0.0, 0.0, 0.0]

In Forward pass trace: LoRALinear(4→4, rank=2, alpha=4), given the rows so far: what is the next one — the row where Term is y_init = base + delta?

Correct: y_init = base + delta | [-0.6614, 0.4710, -0.3133, 0.3568]

TermValue (4-vec, rounded)
x[1.0, 0.0, -1.0, 0.5]
A@B max abs (step 0)0.000000
base = xW0[-0.6614, 0.4710, -0.3133, 0.3568]
delta = x@A@B*scale[0.0, 0.0, 0.0, 0.0]
y_init = base + delta[-0.6614, 0.4710, -0.3133, 0.3568]

Why: The relationship between the columns, not the individual numbers, is what generates the next row. At step 0, delta_W = A@B = 4×2 @ 2×4 = 4×4 matrix of all zeros.

26. Forward pass trace: LoRALinear(4→4, rank=2, alpha=4)

Worked example

Set up: d=4, r=2, alpha=4, scale=2.0; B=zeros at init

Why: At step 0, delta_W = A@B = 4×2 @ 2×4 = 4×4 matrix of all zeros. The frozen path xW0 dominates.

import torch
torch.manual_seed(0)
d, r, alpha = 4, 2, 4
W0 = torch.randn(d, d)      # pretrained, frozen
A  = torch.randn(d, r)*0.01 # trainable, small init
B  = torch.zeros(r, d)      # trainable, zero init
scale = alpha / r            # = 2.0

x = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
base   = x @ W0.T
delta  = (x @ A @ B) * scale
y_init = base + delta
print(f'W0:\n{W0.numpy().round(3)}')
print(f'A@B max_abs: {(A@B).abs().max().item():.6f}')
print(f'base  = {base.numpy().round(4)}')
print(f'delta = {delta.numpy().round(6)}')
print(f'y_init= {y_init.numpy().round(4)}')
TermValue (4-vec, rounded)
x[1.0, 0.0, -1.0, 0.5]
A@B max abs (step 0)0.000000
base = xW0[-0.6614, 0.4710, -0.3133, 0.3568]
delta = x@A@B*scale[0.0, 0.0, 0.0, 0.0]
y_init = base + delta[-0.6614, 0.4710, -0.3133, 0.3568]

Confirm: y_init == xW0 exactly because B=0 at step 0

Why: This is the zero-init guarantee: the adapter adds nothing at initialization, so the pretrained model's output distribution is perfectly preserved before any gradient step.

27. Fill in: Value (4-vec, rounded) for Forward pass trace: LoRALinear(4→4, rank=2…

Comparison

Comparison matrix

From Forward pass trace: LoRALinear(4→4, rank=2, alpha=4): refill the Value (4-vec, rounded) column from what you know. The rest of the table is as it appeared.

TermValue (4-vec, rounded)
x[1.0, 0.0, -1.0, 0.5]
A@B max abs (step 0)0.000000
base = xW0[-0.6614, 0.4710, -0.3133, 0.3568]
delta = x@A@B*scale[0.0, 0.0, 0.0, 0.0]
y_init = base + delta[-0.6614, 0.4710, -0.3133, 0.3568]

28. Something is wrong here: including frozen base-weight params in the optimizer

Anomaly

Predict first

A student writes this, and it looks reasonable:

Pass all model parameters to the optimizer without filtering.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.

Filter to only requires_grad=True parameters before constructing the optimizer.

Why: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.

29. Trap: including frozen base-weight params in the optimizer

Trap

The trap

Pass all model parameters to the optimizer without filtering.

opt = torch.optim.Adam(model.parameters(), lr=1e-4)

Why: This includes the 175B frozen weights of the base model in Adam's state dict, allocating optimizer momentum and variance buffers for each — costing 2× model memory just for Adam states.

Result: you've just re-implemented full fine-tuning. Memory savings from LoRA are completely negated.

The fix

Filter to only requires_grad=True parameters before constructing the optimizer.

opt = torch.optim.Adam(filter(lambda p: p.requires_grad, model.parameters()), lr=1e-4)

Why: Only A and B (rank × 2 × d params per layer) enter the optimizer. For rank=8 on a 768×768 layer this is 12,288 params vs 589,824 — 48x fewer Adam states.

Result: the optimizer is tiny, backward pass skips frozen weights, and memory usage matches the LoRA param count — not the base model size.

30. Which of these survive contact with Lesson 101: LoRA — Low-Rank Adaptation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
LoRA (Hu et al. 2022) solves this by freezing the pretrained weights and injecting a trainable low-rank delta into selected layers. Only A and B are ever updated.; For a single weight matrix of shape (d_in, d_out) with LoRA rank r, the trainable parameter count is exactly d_in × r + r × d_out. Full fine-tuning uses d_in × d_out.; In the PEFT library (Lesson 102), target_modules=["q_proj", "v_proj"] sets this automatically for HuggingFace models.
Breaks
Initialize A and B both from N(0, 0.01), then add the LoRA adapter and start fine-tuning.; Pass all model parameters to the optimizer without filtering.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 101: LoRA — Low-Rank Adaptation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. Full pipeline — LoRA fine-tuning a classifier

Section

Part 4 of 4

32. The alpha/r scaling hyperparameter

Concept

The delta is scaled by alpha/r before being added to the base output. This decouples the learning-rate sensitivity from the rank choice: changing r without changing alpha adjusts both the delta magnitude and the effective learning rate.

\[ y = x W_0^\top + \frac{\alpha}{r}\, x A B \]

alpharscale = alpha/rEffect
1682.0Hu et al. default; balanced update magnitude
3284.0Larger delta; effectively higher LR for adapter
16161.0Lower scale; smaller delta per gradient step
rr1.0Common convention: set alpha = r to get scale 1

33. Without one step: LoRA implementation pattern

Constraint

Discussion prompt

Run LoRA implementation pattern with this step confiscated:

Initialize: A ~ N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Identify target layers: choose Q and V projections in each attention block (or all linear layers for domain shift tasks)
  2. Replace each nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weight
  3. Freeze base: set requires_grad = False on all .linear.parameters()
  4. Initialize: A ~ N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0
  5. Optimize only A, B: Adam(filter(lambda p: p.requires_grad, model.parameters()))
  6. Merge for inference: W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overhead

34. LoRA implementation pattern

Pattern

  1. Identify target layers: choose Q and V projections in each attention block (or all linear layers for domain shift tasks)
  2. Replace each nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weight
  3. Freeze base: set requires_grad = False on all .linear.parameters()
  4. Initialize: A ~ N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0
  5. Optimize only A, B: Adam(filter(lambda p: p.requires_grad, model.parameters()))
  6. Merge for inference: W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overhead

35. Where does it stop working: LoRA implementation pattern

Edge cases

Discussion prompt

LoRA implementation pattern works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Identify target layers: choose Q and V projections in each attention block (or all linear layers for domain shift tasks)
  2. Replace each nn.Linear(d_in, d_out) with LoRALinear(d_in, d_out, rank=r, alpha=alpha) — copy pretrained weights into .linear.weight
  3. Freeze base: set requires_grad = False on all .linear.parameters()
  4. Initialize: A ~ N(0, 0.01), B = zeros — guarantees delta_W = 0 at step 0
  5. Optimize only A, B: Adam(filter(lambda p: p.requires_grad, model.parameters()))
  6. Merge for inference: W_eff = W0 + (alpha/r) * A @ B — single matmul, zero latency overhead

36. Rule out three: Check 1: trainable parameter count

Elimination

Eliminate the wrong options

A LoRA adapter is applied to a 512×512 weight matrix with rank r=8. How many trainable parameters does it add?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 8,192
  • B. 4,096
  • C. 16,384
  • D. 262,144

Survives elimination: A

Why: A is shape 512×8 = 4,096 params; B is shape 8×512 = 4,096 params; total = 8,192. The full matrix would be 512×512 = 262,144, a 32x reduction. Formula: r(d_in + d_out) = 8×(512+512) = 8,192.

37. Check 1: trainable parameter count

Check

Work this out on paper before clicking.

Check your understanding

A LoRA adapter is applied to a 512×512 weight matrix with rank r=8. How many trainable parameters does it add?

  • A. 8,192 (correct)
  • B. 4,096
  • C. 16,384
  • D. 262,144

Answer: A

Why: A is shape 512×8 = 4,096 params; B is shape 8×512 = 4,096 params; total = 8,192. The full matrix would be 512×512 = 262,144, a 32x reduction. Formula: r(d_in + d_out) = 8×(512+512) = 8,192.

Why B tempts people
4,096 counts only one of A or B — the formula r×d_in, missing the B matrix (r×d_out).
Why C tempts people
16,384 = 2 × 8,192; this doubles the answer as if each matrix were 512×16 instead of 512×8.
Why D tempts people
262,144 = 512×512 is the full weight matrix — the number of params you would train with full fine-tuning, not with LoRA.

38. Answer it before you see the options: Check 2: scaling factor and init

Prediction

Predict first

A LoRA layer has alpha=32 and rank=8. At initialization (B=zeros), what does the layer output for input x?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: x @ W0 exactly, because delta_W = A@B = 0

Why: B=zeros so x@A@B = x@A@0 = 0 regardless of A. The alpha/r scale multiplies zero, still giving zero. The output is x@W0 exactly — the pretrained model is unchanged at step 0.

39. Check 2: scaling factor and init

Check

Think about what the model outputs at step 0.

Check your understanding

A LoRA layer has alpha=32 and rank=8. At initialization (B=zeros), what does the layer output for input x?

  • A. x @ W0 exactly, because delta_W = A@B = 0 (correct)
  • B. x @ W0 scaled by alpha/r = 4.0
  • C. x @ A @ B × 4.0, because A is random
  • D. Zero, because both A and B are initialized to 0

Answer: A

Why: B=zeros so x@A@B = x@A@0 = 0 regardless of A. The alpha/r scale multiplies zero, still giving zero. The output is x@W0 exactly — the pretrained model is unchanged at step 0.

Why B tempts people
The scale alpha/r=4 multiplies the delta term x@A@B, not W0. When the delta is zero (B=0), no scaling of W0 occurs.
Why C tempts people
A is nonzero (random N(0,0.01)), but x@A@B still collapses to 0 because B=0; the product of any matrix with the zero matrix is zero.
Why D tempts people
A is initialized to N(0,0.01), not zeros. Only B is zeroed. Setting A=0 would be wrong — then gradients through A would be zero at step 0, blocking learning entirely.

40. Rule out three: Check 3: GPT-2 FFN LoRA param count

Elimination

Eliminate the wrong options

LoRA with rank=16 is applied to BOTH FFN matrices in one GPT-2 small block (768→3072 and 3072→768). How many total trainable parameters does this add?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 122,880
  • B. 61,440
  • C. 98,304
  • D. 245,760

Survives elimination: A

Why: Layer 1 (768→3072, rank=16): 768×16 + 16×3072 = 12,288 + 49,152 = 61,440. Layer 2 (3072→768, rank=16): 3072×16 + 16×768 = 49,152 + 12,288 = 61,440. Total = 61,440 + 61,440 = 122,880.

41. Check 3: GPT-2 FFN LoRA param count

Check

GPT-2 small has two FFN layers per block: a 768→3072 and a 3072→768 linear. Work this out before clicking.

Check your understanding

LoRA with rank=16 is applied to BOTH FFN matrices in one GPT-2 small block (768→3072 and 3072→768). How many total trainable parameters does this add?

  • A. 122,880 (correct)
  • B. 61,440
  • C. 98,304
  • D. 245,760

Answer: A

Why: Layer 1 (768→3072, rank=16): 768×16 + 16×3072 = 12,288 + 49,152 = 61,440. Layer 2 (3072→768, rank=16): 3072×16 + 16×768 = 49,152 + 12,288 = 61,440. Total = 61,440 + 61,440 = 122,880.

Why B tempts people
61,440 is the count for only ONE of the two FFN matrices — the answer misses the second (symmetric) layer.
Why C tempts people
98,304 = 768×128, a rank-128 adapter on the 768-wide matrix only — not the correct formula for two layers at rank=16.
Why D tempts people
245,760 = 2 × 122,880 double-counts — as if both layers were applied twice each.

42. Merging LoRA weights for zero-latency inference

Concept

During training, every forward pass computes two separate matrix multiplications: xW0 and x@A@B. At inference, this is unnecessary — the delta is constant, so you can merge it into W0 once.

\[ W_{\text{eff}} = W_0 + \frac{\alpha}{r} A B \]

After merging, the model has exactly the same architecture as the original pretrained model — no adapter overhead, same latency, same memory. The LoRA structure only exists during fine-tuning.

43. By analogy: Merging LoRA weights for zero-latency inference

Analogy

Discussion prompt

Explain Merging LoRA weights for zero-latency inference by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

During training, every forward pass computes two separate matrix multiplications: xW0 and x@A@B. At inference, this is unnecessary — the delta is constant, so you can merge it into W0 once.

44. Why is this step legal: Build a TinyModel with a LoRALinear(16→32…

Explain it to yourself

Discussion prompt

In Full LoRA fine-tuning pipeline (synthetic classification) this move is made:

Build a TinyModel with a LoRALinear(16→32, rank=4, alpha=8) first layer

Why is that legal? Name the rule or definition it rests on before you read on.

Hint: If you can only say "because that is what you do", the rule is the thing to go and find.

Answer:

The base linear is frozen; only A and B (16×4 + 4×32 = 192 trainable params out of 576 total for the first layer) are updated.

45. Full LoRA fine-tuning pipeline (synthetic classification)

Worked example

Build a TinyModel with a LoRALinear(16→32, rank=4, alpha=8) first layer

Why: The base linear is frozen; only A and B (16×4 + 4×32 = 192 trainable params out of 576 total for the first layer) are updated.

import torch, torch.nn as nn
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

class LoRALinear(nn.Module):
    def __init__(self, i, o, rank=4, alpha=8):
        super().__init__()
        self.linear = nn.Linear(i, o, bias=False)
        self.A = nn.Parameter(torch.randn(i, rank)*0.01)
        self.B = nn.Parameter(torch.zeros(rank, o))
        self.scale = alpha / rank
        for p in self.linear.parameters(): p.requires_grad = False
    def forward(self, x):
        return self.linear(x) + (x @ self.A @ self.B) * self.scale

class TinyModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = LoRALinear(16, 32, rank=4, alpha=8)
        self.relu = nn.ReLU()
        self.fc2 = nn.Linear(32, 2)
    def forward(self, x): return self.fc2(self.relu(self.fc1(x)))

torch.manual_seed(7)
X_np, y_np = make_classification(500, 16, n_informative=8, random_state=42)
X = torch.tensor(X_np, dtype=torch.float32)
y = torch.tensor(y_np, dtype=torch.long)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
model = TinyModel()
opt = torch.optim.Adam(
    filter(lambda p: p.requires_grad, model.parameters()), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
for ep in range(100):
    model.train()
    loss = loss_fn(model(X_tr), y_tr)
    opt.zero_grad(); loss.backward(); opt.step()
model.eval()
with torch.no_grad():
    acc = (model(X_te).argmax(1) == y_te).float().mean().item()
print(f'test_acc={acc:.4f}')
EpochTrain lossNote
00.7011B=0, model behaves like pretrained baseline
200.6572B learns nonzero values; delta activates
400.6101Gradient flows through A and B only
600.5605Loss converges toward 0.5 on balanced binary task
800.5227Stable convergence
100 (test)acc=0.730073.0% test accuracy with only A+B updating

46. What each one costs: Full LoRA fine-tuning pipeline (synthetic…

Trade off

Comparison matrix

From Full LoRA fine-tuning pipeline (synthetic classification): every row here is a choice with a cost. Fill the Train loss column, then say which row you would actually pick and what you give up for it.

EpochTrain lossNote
00.7011B=0, model behaves like pretrained baseline
200.6572B learns nonzero values; delta activates
400.6101Gradient flows through A and B only
600.5605Loss converges toward 0.5 on balanced binary task
800.5227Stable convergence
100 (test)acc=0.730073.0% test accuracy with only A+B updating

47. LoRA vs full fine-tuning vs feature extraction

Concept

MethodTrainable paramsGPU memoryPerformanceBest for
Feature extractionHead only (~0.01%)SmallestGood if task ≈ pretrain domainTiny task datasets
LoRA (r=8)~0.1–2% of modelSmallNear full fine-tune on most tasksLimited GPU, many tasks
Full fine-tuning100%Largest (3× model for Adam)Best on large task datasetsLarge data, single task

LoRA achieves near-full fine-tuning quality on most NLP benchmarks while using 48–384× fewer trainable parameters for typical transformer layers. It also enables multi-task deployment: store one frozen base model and dozens of tiny LoRA checkpoints, swapping them at serving time.

48. Fill in: Trainable params for LoRA vs full fine-tuning vs feature…

Comparison

Comparison matrix

From LoRA vs full fine-tuning vs feature extraction: refill the Trainable params column from what you know. The rest of the table is as it appeared.

MethodTrainable paramsGPU memoryPerformanceBest for
Feature extractionHead only (~0.01%)SmallestGood if task ≈ pretrain domainTiny task datasets
LoRA (r=8)~0.1–2% of modelSmallNear full fine-tune on most tasksLimited GPU, many tasks
Full fine-tuning100%Largest (3× model for Adam)Best on large task datasetsLarge data, single task

49. Callback: LoRA in the Phase-3 context (Lessons 88–101)

Concept

In Lesson 88 (BERT), you saw that the pretrained transformer's Q/K/V attention projections carry rich contextual representations. LoRA injects task-specific deltas into exactly those projections — the minimum change needed to steer an already-capable model.

In Lesson 89 (GPT), you saw how decoder-only models autoregressively generate text; LoRA is the standard recipe for instruction-tuning GPT-style models (e.g. Alpaca, LLaMA-2-chat) without full 16-bit fine-tuning.

In Lesson 102, you will use the HuggingFace peft library to apply LoRA with two lines of config, wrapping the techniques built from scratch here.

50. Break it if you can: Callback: LoRA in the Phase-3 context (Lessons…

Counterexample

Discussion prompt

In Lesson 102, you will use the HuggingFace peft library to apply LoRA with two lines of config, wrapping the techniques built from scratch here.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

51. Connect it up: Lesson 101: LoRA — Low-Rank Adaptation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why LoRA? — The fine-tuning bottleneck · The LoRA Math — decomposition, scaling, init · Implementation — LoRALinear nn.Module · Full pipeline — LoRA fine-tuning a classifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

52. Lesson 101 recap: LoRA in five points

Recap

  1. Low-rank hypothesis: weight updates delta_W during fine-tuning live in a low-rank subspace, so delta_W ≈ AB with rank r << d
  2. Forward pass: y = xW0 + (alpha/r) * xAB — W0 is frozen, only A and B are trained
  3. Zero-init guarantee: B=zeros at init ensures delta_W=0 at step 0 — the pretrained model is perfectly preserved before gradient descent begins
  4. Parameter math: rank r on a d×d layer → 2dr trainable params vs d² full; for 768×768 rank-8: 12,288 vs 589,824 (48x fewer)
  5. Target layers: apply LoRA to Q and V projections first; FFN optional; never embed/LayerNorm; merge weights (W_eff = W0 + (alpha/r)AB) for zero-latency inference

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 101 — LoRA — Barron · USAAIO Round 2 Preparation, 2026
  2. Hu et al. 'LoRA: Low-Rank Adaptation of Large Language Models' (ICLR 2022) — arXiv:2106.09685
  3. LoRALinear nn.Module, forward pass math, parameter counts verified with torch 2.7.1+cpu, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108