USAAIO Lesson 12, from Week 4 on probability, fully worked. It derives MAP from Bayes' rule as MLE plus a log-prior, showing every logarithmic step, then turns a Gaussian prior into L2 or ridge and a Laplace prior into L1 or lasso, term by term. It gives the exact reason L1 zeros out coordinates, through the subgradient and the soft-threshold, computed one coordinate at a time, proves Beta-Binomial conjugacy by proportionality, and covers dropout as approximate Bayesian inference. It ends with a ridge-against-lasso sparsity project verified against scikit-learn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 115 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 12 · Week 4 (Probability)
Regularization is not a hack bolted onto the loss — it is a prior belief, written in log form. We derive MAP = MLE + log-prior from Bayes' rule with no skipped step, turn a Gaussian prior into ridge and a Laplace prior into lasso term by term, prove exactly why L1 drives weights to zero, and verify all of it against scikit-learn.
Objectives
arg max [log-likelihood + log-prior], every log step shownλ = σ²/τ²Beta(α,β) by inspectionsklearnWarm-up
Discussion prompt
Before we open Lesson 12: MAP, Priors & Regularization: without looking back, what was the main idea of Autoencoders — Standard, Denoising & Sparse, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the autoencoder as encoder -> bottleneck latent z -> decoder trained on reconstruction loss alone; undercomplete compression; the denoising autoencoder (corrupt the input, reconstruct the clean signal) and the sparse autoencoder (L1 penalty on activations, the workhorse of mechanistic interpretability); latent-space interpolation; and the key limitation that motivates the VAE — a standard AE's latent space has no prior, so you cannot sample from it. Build a convolutional autoencoder your-turn.
Section
Part 1 of 7 — from Bayes to a penalty
Concept
We reuse Lesson 7's five students — hours studied x and score y — and fit ŷ = w₀ + w₁·x. But now we ask a probability question: with so few points, what should keep the weights from swinging wildly?
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
The plain least-squares fit is w = [2.2, 0.6]. This lesson asks: what belief about w would nudge that fit — and where does that belief come from? The answer is a prior.
Comparison
Comparison matrix
From The running example: fit a line, but small data: refill the hours x column from what you know. The rest of the table is as it appeared.
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Concept
In Lesson 8 the maximum-likelihood estimate picked the θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.
\[ \theta_{\text{MLE}} = \arg\max_{\theta}\; p(\text{data}\mid\theta) \]
MLE uses only the data. With five points and two weights it is fine here — but with many features and few rows it overfits, chasing noise. We want a principled way to say "and I also believe the weights are small."
Counterexample
Discussion prompt
In Lesson 8 the maximum-likelihood estimate picked the θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Intuition
Before seeing a single point, you already suspect the weights are modest — a giant slope would be surprising. That suspicion is a prior distribution p(θ): it puts more probability on small weights than on huge ones.
MAP blends the two voices: the data (likelihood) says "here is what fits", the prior says "but stay reasonable". With little data the prior has a loud say; with mountains of data the likelihood drowns it out.
That single idea — belief as a distribution you can multiply in — is what turns a regularizer from a fudge factor into a probability statement.
Concept
Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:
\[ p(\theta \mid D) = \frac{p(D \mid \theta)\, p(\theta)}{p(D)} \]
posterior — The updated belief about θ AFTER seeing the data — proportional to (likelihood × prior). MAP picks the single θ at its peak; full Bayes keeps the whole distribution.
Analogy
Discussion prompt
Explain Bayes' rule, for parameters by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:
Ranking
Put in order
Put the moves of Derive MAP: drop the constant, take the log into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. MAP is the arg max of the posterior over θ — the most probable parameter given the data.
Worked example
Start from the posterior
Why: MAP is the arg max of the posterior over θ — the most probable parameter given the data.
\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; \frac{p(D\mid\theta)\,p(\theta)}{p(D)} \]
Drop p(D) — it does not depend on θ
Why: The denominator p(D) is a fixed number once the data is in; scaling by a positive constant never moves an arg max. So the posterior is PROPORTIONAL to likelihood × prior.
\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; p(D\mid\theta)\,p(\theta) \]
Take logs — log is monotonic
Why: log is strictly increasing, so arg max is unchanged; and log turns the product into a SUM, which is far easier to differentiate.
\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; \big[\log p(D\mid\theta) + \log p(\theta)\big] \]
Notation
Annotate
From Derive MAP: drop the constant, take the log — read this one piece at a time. What is each part doing?
On: \( \theta_{\text{MAP}} = \arg\max_{\theta}\; \frac{p(D\mid\theta)\,p(\theta)}{p(D)} \)
Concept
The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:
\[ \boxed{\;\theta_{\text{MAP}} = \arg\max_{\theta}\; \underbrace{\log p(D\mid\theta)}_{\text{MLE term}} + \underbrace{\log p(\theta)}_{\text{regularizer}}\;} \]
Flip the sign to turn maximize into minimize and −log p(θ) becomes a penalty added to the loss. Every regularizer you know is secretly some −log p(θ). The rest of Part 1 shows which prior gives which penalty.
Explain it
Discussion prompt
Explain The one-line summary: MAP = MLE + log-prior to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:
Intuition
Full Bayesian inference keeps the whole posterior p(θ|D), which needs the denominator p(D) = ∫ p(D|θ)p(θ) dθ — an integral that is usually intractable.
MAP dodges it. Since p(D) doesn't depend on θ, it cannot move the peak, so MAP finds the single most-probable θ without ever computing the integral — just maximize likelihood × prior.
The cost of the shortcut: you get a point, not a distribution — no uncertainty bars. That trade (a penalty you can optimize vs. a full posterior) is exactly the regularization story, and dropout in Part 5 buys some uncertainty back.
Step zero
Discussion prompt
A tiny MAP by hand — one weight — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Sums: Σxᵢ² = 1+4+9 = 14, Σxᵢyᵢ = 2.1+7.8+18.6 = 28.5
Answer:
Worked example
Take three points x = [1,2,3], y = [2.1, 3.9, 6.2] and the one-weight model ŷ = w·x with Gaussian noise σ² and a Gaussian prior w ~ N(0, τ²). The MAP weight has a clean closed form (derived fully in Part 2):
\[ w_{\text{MAP}} = \frac{\sum_i x_i y_i}{\sum_i x_i^2 + \sigma^2/\tau^2} \]
Sums: Σxᵢ² = 1+4+9 = 14, Σxᵢyᵢ = 2.1+7.8+18.6 = 28.5
Why: These are the same two sums OLS needs. Σxᵢyᵢ = 1(2.1)+2(3.9)+3(6.2) = 2.1+7.8+18.6 = 28.5.
MLE (no prior): w = 28.5 / 14 = 2.0357
Why: Setting σ²/τ² = 0 recovers plain least squares: 28.5/14 = 2.035714.
MAP with σ²=1, τ²=0.25 ⇒ λ = σ²/τ² = 4: w = 28.5 / (14+4) = 1.5833
Why: The prior adds 4 to the denominator, pulling the estimate from 2.0357 toward 0. Verified: 28.5/18 = 1.583333.
| estimator | denominator | w |
|---|---|---|
| MLE (λ=0) | 14 | 2.0357 |
| MAP (λ=4) | 14 + 4 = 18 | 1.5833 |
Trade off
Comparison matrix
From A tiny MAP by hand — one weight: every row here is a choice with a cost. Fill the denominator column, then say which row you would actually pick and what you give up for it.
| estimator | denominator | w |
|---|---|---|
| MLE (λ=0) | 14 | 2.0357 |
| MAP (λ=4) | 14 + 4 = 18 | 1.5833 |
Estimation
Predict first
Confirm the by-hand numbers with a runnable snippet. The MAP weight is the same two sums with σ²/τ² added to the denominator:
Commit before you compute: what does The tiny MAP in code come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: λ = σ²/τ² = 4 pulls the estimate from 2.0357 toward 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Adding λ to the denominator is exactly the 1-D case of (XᵀX + λI).
Worked example
Confirm the by-hand numbers with a runnable snippet. The MAP weight is the same two sums with σ²/τ² added to the denominator:
import numpy as np
xs = np.array([1., 2., 3.])
ys = np.array([2.1, 3.9, 6.2])
Sxx = np.sum(xs**2); Sxy = np.sum(xs*ys)
lam = 1.0/0.25 # sigma^2 / tau^2 = 1 / 0.25 = 4
print('Sxx, Sxy:', Sxx, Sxy)
print('MLE:', Sxy/Sxx)
print('MAP:', Sxy/(Sxx+lam))λ = σ²/τ² = 4 pulls the estimate from 2.0357 toward 0
Why: Adding λ to the denominator is exactly the 1-D case of (XᵀX + λI). Larger λ = larger denominator = more shrinkage toward the prior mean 0.
| value (verified) | |
|---|---|
| Sxx, Sxy | 14.0 28.5 |
| MLE | 2.035714 |
| MAP | 1.583333 |
Pattern
Step through it
Step through The tiny MAP in code one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 2 of 7 — term by term
Concept
Assume each target is the line plus independent Gaussian noise, yᵢ = xᵢᵀw + εᵢ with εᵢ ~ N(0, σ²). Then the negative log-likelihood is the familiar sum of squares over 2σ²:
\[ -\log p(y\mid X, w) = \frac{1}{2\sigma^2}\lVert Xw - y\rVert^2 + \text{const} \]
The constant collects n·½log(2πσ²) — no w in it, so it drops out of any minimization over w. Minimizing this alone is exactly OLS: MLE is least squares.
Step zero
Discussion prompt
The Gaussian log-prior is a negative quadratic — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Density: p(w) ∝ exp(−‖w‖² / (2τ²))
Answer:
Worked example
Put a zero-mean Gaussian prior on each weight, w ~ N(0, τ²I). Write its density and take −log:
Density: p(w) ∝ exp(−‖w‖² / (2τ²))
Why: A multivariate zero-mean Gaussian with covariance τ²I has this exponential kernel; the leading constant does not involve w.
\[ p(w) = \frac{1}{(2\pi\tau^2)^{d/2}}\exp\!\Big(-\frac{\lVert w\rVert^2}{2\tau^2}\Big) \]
−log p(w) = ‖w‖²/(2τ²) + const
Why: log of an exponential is its exponent; the −½log(2πτ²)ᵈ piece is constant in w. So the prior contributes a term PROPORTIONAL to ‖w‖².
\[ -\log p(w) = \frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]
Numeric check (τ²=0.5): the quadratic part is w²/(2τ²)
Why: At w=0,1,2 the −log density is 0.5724, 1.5724, 4.5724; subtract the const 0.5724 and you get 0, 1, 4 = w²/(2·0.5). The w-dependent part is exactly the quadratic.
| w | −log p(w) | quadratic part w²/(2τ²) | constant |
|---|---|---|---|
| 0 | 0.5724 | 0.0000 | 0.5724 |
| 1 | 1.5724 | 1.0000 | 0.5724 |
| 2 | 4.5724 | 4.0000 | 0.5724 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Numeric check (τ²=0.5): the quadratic part is w²/(2τ²)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Put a zero-mean Gaussian prior on each weight, w ~ N(0, τ²I). Write its density and take −log:
Ranking
Put in order
Put the moves of Add the two −logs → the ridge objective into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Flip the arg max of the sum of logs into an arg min of the negative sum.
Worked example
MAP minimizes −(log-likelihood + log-prior)
Why: Flip the arg max of the sum of logs into an arg min of the negative sum. Add the two negative-log pieces from the last two slides.
\[ -\log p(w\mid D) = \frac{1}{2\sigma^2}\lVert Xw-y\rVert^2 + \frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]
Multiply through by 2σ² (a positive constant)
Why: Scaling by 2σ² > 0 does not move the minimizer; it clears the fraction on the loss term and folds σ² into the penalty coefficient.
\[ \lVert Xw-y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \]
Name the ratio λ = σ²/τ² — this IS ridge
Why: The MAP objective is Lesson 7's ridge loss exactly. The regularization strength is the noise-to-prior variance ratio: noisier data or a tighter prior ⇒ larger λ ⇒ more shrinkage.
\[ \boxed{\;\min_w\; \lVert Xw-y\rVert^2 + \lambda\lVert w\rVert^2,\qquad \lambda = \frac{\sigma^2}{\tau^2}\;} \]
Translation
\( \lVert Xw-y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Missing information
Discussion prompt
Ridge's normal equations are (XᵀX + λI)w = Xᵀy (Lesson 7). Solve them with λ = 1 on the student data and compare to plain OLS. Runnable as written:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Adding np.eye(2) makes the Gaussian-prior term appear as +λI. The exact same solve returns a shrunken weight vector.
Worked example
Ridge's normal equations are (XᵀX + λI)w = Xᵀy (Lesson 7). Solve them with λ = 1 on the student data and compare to plain OLS. Runnable as written:
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
w_ridge = np.linalg.solve(X.T @ X + 1.0*np.eye(2), X.T @ y)
print('OLS :', w_ols)
print('ridge:', w_ridge)
print('norms:', np.linalg.norm(w_ols), np.linalg.norm(w_ridge))λI lifts every eigenvalue of XᵀX by λ, then solve
Why: Adding np.eye(2) makes the Gaussian-prior term appear as +λI. The exact same solve returns a shrunken weight vector.
| fit | w = [w₀, w₁] | ‖w‖ (verified) |
|---|---|---|
| OLS (λ=0) | [2.2, 0.6] | 2.2804 |
| ridge (λ=1) | [1.1712, 0.8649] | 1.4559 |
Intuition
The Gaussian prior is a smooth bump at 0: near the origin it is nearly flat, so it barely pulls once a weight is already small. Ridge therefore squeezes every coefficient toward 0 but leaves it a little nonzero.
On the students, ‖w‖ dropped from 2.28 to 1.46 — both entries got smaller, neither vanished. To get exact zeros we need a prior with a sharp spike at 0. That is the Laplace prior, next.
Intuition
In deep learning, the same L2 term is called weight decay: every gradient step, before applying the data gradient, you shrink w a little toward 0. That shrink IS the gradient of λ‖w‖².
So the optimizer setting you flip on out of habit (weight_decay=1e-4) is a Gaussian prior on every weight. Naming it 'decay' hides the probability story — but it is the identical −log p(w) term you just derived.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A tighter prior (small τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks.
Read it straight off the derivation: after clearing the loss fraction, the penalty coefficient is the noise variance over the prior variance.
Why: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks. τ²/σ² shrinks as τ²→0, giving less penalty exactly when you wanted more.
Trap
A tighter prior (small τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².
λ = τ²/σ² ✗
Why: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks. τ²/σ² shrinks as τ²→0, giving less penalty exactly when you wanted more.
Read it straight off the derivation: after clearing the loss fraction, the penalty coefficient is the noise variance over the prior variance.
λ = σ²/τ² ✓
Why: Small τ² (tight prior) ⇒ large λ ⇒ heavy shrinkage. Large σ² (noisy data, trust the fit less) ⇒ large λ too. Both directions match intuition; τ²→∞ (flat prior) ⇒ λ→0 ⇒ back to MLE.
Section
Part 3 of 7 — and why it is sparse
Picture it
Figure (svg): Two prior densities over w centered at zero: a smooth rounded Gaussian bump, and a Laplace density with a sharp cusp (a kink) at zero.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Swap the Gaussian bell for a Laplace (double-exponential) prior on each weight. It has a sharp peak at 0 and heavier-than-Gaussian tails:
Concept
Swap the Gaussian bell for a Laplace (double-exponential) prior on each weight. It has a sharp peak at 0 and heavier-than-Gaussian tails:
\[ p(w_j) = \frac{1}{2b}\exp\!\Big(-\frac{|w_j|}{b}\Big) \]
Figure (svg): Two prior densities over w centered at zero: a smooth rounded Gaussian bump, and a Laplace density with a sharp cusp (a kink) at zero.
Step zero
Discussion prompt
The Laplace log-prior is proportional to |w| — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: −log p(wⱼ) = |wⱼ|/b + log(2b)
Answer:
Worked example
−log p(wⱼ) = |wⱼ|/b + log(2b)
Why: Take −log of the density: the exponential's exponent |wⱼ|/b comes down linearly, and −log(1/(2b)) = log(2b) is the constant.
\[ -\log p(w) = \frac{1}{b}\sum_j |w_j| + \text{const} = \frac{1}{b}\lVert w\rVert_1 + \text{const} \]
Numeric check (b=1): the w-part is exactly |w|
Why: At w = −2,−1,0,1,2 the −log density is 2.6931, 1.6931, 0.6931, 1.6931, 2.6931. Subtract the const log(2)=0.6931 and you get 2,1,0,1,2 = |w|.
| w | −log p(w) | |w|/b | constant log(2b) |
|---|---|---|---|
| −2 | 2.6931 | 2.0000 | 0.6931 |
| −1 | 1.6931 | 1.0000 | 0.6931 |
| 0 | 0.6931 | 0.0000 | 0.6931 |
| 1 | 1.6931 | 1.0000 | 0.6931 |
| 2 | 2.6931 | 2.0000 | 0.6931 |
Add it to the Gaussian-noise NLL → lasso
Why: Same MAP recipe: loss + −log prior. The L1 term is the lasso penalty.
\[ \boxed{\;\min_w\; \lVert Xw-y\rVert^2 + \lambda\lVert w\rVert_1\;} \]
Invariant
Step through it
Step through The Laplace log-prior is proportional to |w| one row at a time. One of these columns never changes — find it, and say why it cannot.
Intuition
The Laplace prior does two things a Gaussian can't. Its spike at 0 says "most weights are probably exactly zero" — that is where sparsity comes from. Its heavier tails say "but the few nonzero ones can be genuinely large."
A Gaussian penalizes large weights harshly (quadratically), so it never lets a real signal grow. Laplace tolerates a handful of big coefficients while zeroing the rest — exactly the shape of a sparse truth, which is why lasso recovers {0,3,7} cleanly.
Sorting
Sort into buckets
These are the pieces of Lesson 12: MAP, Priors & Regularization, out of order. Put each one back under the part of the lesson it belongs to.
Intuition
Here is the heart of it. The gradient of the penalty tells you how hard it pushes a weight toward 0. For L2 that push shrinks with the weight; for L1 it is constant — the same shove at w = 0.001 as at w = 1.
So L2 eases off exactly when a weight gets small — it approaches 0 but never arrives. L1 keeps shoving with full force until the weight hits exactly 0 and stops. That constant pull is why lasso produces true zeros.
Pattern
Predict first
The table runs: 1.0 | 2.0000 | 1.0000 · 0.1 | 0.2000 | 1.0000 · 0.01 | 0.0200 | 1.0000
In Watch the two gradients near zero, given the rows so far: what is the next one — the row where w is 0.001?
Correct: 0.001 | 0.0020 | 1.0000
| w | L2 grad = 2λw | L1 grad = λ·sign(w) |
|---|---|---|
| 1.0 | 2.0000 | 1.0000 |
| 0.1 | 0.2000 | 1.0000 |
| 0.01 | 0.0200 | 1.0000 |
| 0.001 | 0.0020 | 1.0000 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The L2 restoring force vanishes near the origin (proportional shrink), so the coefficient asymptotes to a tiny nonzero value.
Worked example
Penalty gradients with λ = 1: L2 is d/dw (λw²) = 2λw; L1 is d/dw (λ|w|) = λ·sign(w), constant in magnitude. Print them as w → 0:
import numpy as np
for w in [1.0, 0.1, 0.01, 0.001]:
l2 = 2*1.0*w # d/dw (lambda w^2), lambda=1
l1 = 1.0*np.sign(w) # d/dw (lambda |w|), lambda=1
print(f'w={w:<6} L2 grad={l2:<8} L1 grad={l1}')L2 grad → 0 as w → 0; L1 grad stays at 1
Why: The L2 restoring force vanishes near the origin (proportional shrink), so the coefficient asymptotes to a tiny nonzero value. L1's force is constant, so it can carry the coefficient all the way to 0.
| w | L2 grad = 2λw | L1 grad = λ·sign(w) |
|---|---|---|
| 1.0 | 2.0000 | 1.0000 |
| 0.1 | 0.2000 | 1.0000 |
| 0.01 | 0.0200 | 1.0000 |
| 0.001 | 0.0020 | 1.0000 |
Concept
Each penalty is a constraint ball around the origin; the fit is the loss contour's first touch. L1's ball is a diamond with corners on the axes; L2's is a round circle.
Figure (svg): Left: a diamond-shaped L1 constraint whose corner on the vertical axis is where an elliptical loss contour first touches, making one weight zero. Right: a round L2 constraint whose tangent point with the loss contour is off the axes, so neither weight is zero.
Concept
|w| has a corner at 0 — it is not differentiable there. Instead of a single slope, at w=0 it has a whole set of valid slopes, the subgradient ∂|w| = [−1, 1].
\[ \partial\,|w| = \begin{cases} \{-1\} & w < 0 \\ [-1,\,1] & w = 0 \\ \{+1\} & w > 0 \end{cases} \]
A minimum can sit at w=0 whenever the data gradient there lands inside [−λ, λ] — the flat interval of the penalty absorbs it. A smooth L2 has no such interval (its slope at 0 is a single 0), so it can never pin a coordinate at exactly 0. The kink is the whole mechanism.
Fill the middle
Fill in the blanks
From One coordinate, exactly: shrink vs soft-threshold — one line has had its right-hand side removed. Put it back.
import numpy as np
def soft(rho, t):
return np.sign(rho)*max(abs(rho)-t, 0.0)
lam = 0.4
for rho in [1.0, 0.5, 0.15, 0.10, 0.05]:
ridge = rho/(1+lam)
lasso = soft(rho, lam/2) # threshold lam/2 = 0.2
print(f'rho=___ ridge=___ lasso=___')
Why: ridge is what everything below it consumes, so the wrong expression here fails later and somewhere else. Soft-thresholding subtracts λ/2 and clips at 0, so any small coordinate is killed outright.
Worked example
For a single standardized feature (so Σx²=1), the OLS coordinate is some ρ. Ridge and lasso each have a closed form for that one coordinate — and they differ in a decisive way:
\[ w_{\text{ridge}} = \frac{\rho}{1+\lambda}, \qquad w_{\text{lasso}} = \operatorname{sign}(\rho)\,\max\!\big(|\rho| - \tfrac{\lambda}{2},\, 0\big) \]
import numpy as np
def soft(rho, t):
return np.sign(rho)*max(abs(rho)-t, 0.0)
lam = 0.4
for rho in [1.0, 0.5, 0.15, 0.10, 0.05]:
ridge = rho/(1+lam)
lasso = soft(rho, lam/2) # threshold lam/2 = 0.2
print(f'rho={rho:<5} ridge={ridge:.4f} lasso={lasso:.4f}')Below the threshold |ρ| ≤ λ/2 = 0.2, lasso snaps to EXACTLY 0; ridge never does
Why: Soft-thresholding subtracts λ/2 and clips at 0, so any small coordinate is killed outright. Ridge only divides by (1+λ), which shrinks but preserves the sign and nonzero-ness.
| ρ (OLS coord) | ridge = ρ/(1+λ) | lasso = soft(ρ, 0.2) |
|---|---|---|
| 1.00 | 0.7143 | 0.8000 |
| 0.50 | 0.3571 | 0.3000 |
| 0.15 | 0.1071 | 0.0000 ← zero |
| 0.10 | 0.0714 | 0.0000 ← zero |
| 0.05 | 0.0357 | 0.0000 ← zero |
Discrimination
Sort into buckets
Sort these by lasso = soft(ρ, 0.2), from memory, without looking back at One coordinate, exactly: shrink vs soft-threshold. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero.
Ridge shrinks every coefficient smoothly; only L1 / lasso sets coordinates to exactly zero.
Why: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).
Trap
Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.
Expect many exact zeros from L2
Why: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).
Ridge shrinks every coefficient smoothly; only L1 / lasso sets coordinates to exactly zero.
Reach for lasso when you want sparsity / feature selection
Why: Verified in Part 5: on 10 features (only 0,3,7 real) ridge keeps 10 nonzero, lasso zeros 7 and recovers exactly the true support {0,3,7}.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Verified in Part 5: on 10 features (only 0,3,7 real) ridge keeps 10 nonzero, lasso zeros 7 and recovers exactly the true support {0,3,7}.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).
Section
Part 4 of 7 — updates by inspection
Concept
Multiplying a likelihood by a prior usually gives an ugly posterior you must integrate to normalize. A conjugate prior is chosen so the posterior is the same kind of distribution — the update becomes arithmetic on the parameters.
| likelihood | conjugate prior | posterior |
|---|---|---|
| Bernoulli / Binomial | Beta(α, β) | Beta(α+k, β+n−k) |
| Categorical / Multinomial | Dirichlet(α) | Dirichlet(α + counts) |
| Gaussian (known σ²) | Gaussian | Gaussian |
The star case for the exam is Beta–Binomial: observe k successes in n trials and the posterior is Beta(α+k, β+n−k). No integral. Let's prove it and use it.
Ranking
Put in order
Put the moves of Beta × Binomial = Beta, by matching exponents into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Binomial likelihood in the success prob p is pᵏ(1−p)ⁿ⁻ᵏ; Beta(α,β) prior kernel is p^(α−1)(1−p)^(β−1).
Worked example
Write the two kernels (drop constants)
Why: Binomial likelihood in the success prob p is pᵏ(1−p)ⁿ⁻ᵏ; Beta(α,β) prior kernel is p^(α−1)(1−p)^(β−1).
\[ p(D\mid p)\,p(p) \;\propto\; p^{k}(1-p)^{n-k}\cdot p^{\alpha-1}(1-p)^{\beta-1} \]
Add the exponents
Why: Same base p multiplies by adding exponents: k+(α−1) and (n−k)+(β−1).
\[ = p^{\,\alpha + k - 1}\,(1-p)^{\,\beta + (n-k) - 1} \]
Recognize the Beta(α+k, β+n−k) kernel
Why: That is exactly the kernel of a Beta with parameters α+k and β+n−k. The posterior is Beta again — conjugacy proved: add successes to α, failures to β.
\[ p(p\mid D) = \mathrm{Beta}(\alpha + k,\; \beta + n - k) \]
Notation
Annotate
From Beta × Binomial = Beta, by matching exponents — read this one piece at a time. What is each part doing?
On: \( p(p\mid D) = \mathrm{Beta}(\alpha + k,\; \beta + n - k) \)
Estimation
Predict first
Prior Beta(2,2), data k=7 successes in n=10. The claim: posterior is Beta(9,5), i.e. likelihood × prior is proportional to the Beta(9,5) kernel with the SAME ratio at every p. Check it:
Commit before you compute: what does Verify conjugacy numerically come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The ratio is a CONSTANT (1.0) across all p → same distribution
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. likelihood×prior and the Beta(9,5) kernel differ only by a normalizing constant, so their ratio is identical everywhere.
Worked example
Prior Beta(2,2), data k=7 successes in n=10. The claim: posterior is Beta(9,5), i.e. likelihood × prior is proportional to the Beta(9,5) kernel with the SAME ratio at every p. Check it:
import numpy as np
a0, b0, k, n = 2, 2, 7, 10
def unnorm_post(p): return p**k*(1-p)**(n-k) * p**(a0-1)*(1-p)**(b0-1)
def beta_kernel(p): return p**(a0+k-1)*(1-p)**(b0+(n-k)-1) # Beta(9,5)
for p in [0.3, 0.5, 0.7, 0.9]:
r = unnorm_post(p)/beta_kernel(p)
print(f'p={p}: ratio={r:.6f}')The ratio is a CONSTANT (1.0) across all p → same distribution
Why: likelihood×prior and the Beta(9,5) kernel differ only by a normalizing constant, so their ratio is identical everywhere. A constant ratio is precisely what 'same distribution up to normalization' means.
| p | likelihood×prior | Beta(9,5) kernel | ratio (verified) |
|---|---|---|---|
| 0.3 | 1.575e−05 | 1.575e−05 | 1.000000 |
| 0.5 | 2.441e−04 | 2.441e−04 | 1.000000 |
| 0.7 | 4.669e−04 | 4.669e−04 | 1.000000 |
| 0.9 | 4.305e−05 | 4.305e−05 | 1.000000 |
Scale up
Step through it
Step through Verify conjugacy numerically and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Concept
With Beta(9,5) the posterior mean is 9/14 = 0.6429 — sitting between the prior mean 0.5 and the data's MLE k/n = 0.7. The prior Beta(2,2) acts like 2 pseudo-successes and 2 pseudo-failures added to the real counts.
| quantity | formula | value (verified) |
|---|---|---|
| prior mean | α/(α+β) = 2/4 | 0.5000 |
| data MLE | k/n = 7/10 | 0.7000 |
| posterior mean | 9/14 | 0.6429 |
| posterior mode (MAP) | (9−1)/(14−2) | 0.6667 |
The MAP point estimate is the posterior mode (α−1)/(α+β−2) = 8/12 = 0.6667. More data (bigger n) drags the posterior toward the MLE; a stronger prior (bigger α+β) holds it near 0.5.
Fill the middle
Fill in the blanks
From Dirichlet–Categorical: the same trick, K outcomes — one line has had its right-hand side removed. Put it back.
import numpy as np
alpha0 = np.array([1., 1., 1.]) # flat Dirichlet prior
counts = np.array([12., 5., 3.]) # 20 rolls
post = alpha0 + counts # Dirichlet(13, 6, 4)
print('posterior alpha:', post)
print('posterior mean :', (post/post.sum()).round(4))
print('MLE :', (counts/counts.sum()).round(4))
Why: counts is what everything below it consumes, so the wrong expression here fails later and somewhere else. Just as Beta added successes/failures, Dirichlet adds each outcome's count to its α.
Worked example
The multi-outcome version generalizes Beta to Dirichlet. Roll a 3-sided die 20 times, seeing counts [12, 5, 3], under a flat prior Dirichlet(1,1,1). The update is just add the counts:
import numpy as np
alpha0 = np.array([1., 1., 1.]) # flat Dirichlet prior
counts = np.array([12., 5., 3.]) # 20 rolls
post = alpha0 + counts # Dirichlet(13, 6, 4)
print('posterior alpha:', post)
print('posterior mean :', (post/post.sum()).round(4))
print('MLE :', (counts/counts.sum()).round(4))Posterior = Dirichlet(α + counts) = Dirichlet(13, 6, 4)
Why: Just as Beta added successes/failures, Dirichlet adds each outcome's count to its α. The flat prior acts like one pseudo-count per face, nudging the estimate off the raw frequencies.
| face | count | MLE = count/n | posterior mean |
|---|---|---|---|
| 1 | 12 | 0.6000 | 0.5652 |
| 2 | 5 | 0.2500 | 0.2609 |
| 3 | 3 | 0.1500 | 0.1739 |
Pattern
Step through it
Step through Dirichlet–Categorical: the same trick, K outcomes one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 5 of 7 — same idea, deep nets
Concept
Dropout multiplies each activation by an independent Bernoulli(1−p) mask during training — a random subset of units is zeroed on every forward pass, then rescaled so the expected activation is unchanged.
\[ \tilde h = \frac{m \odot h}{1-p}, \qquad m_j \sim \mathrm{Bernoulli}(1-p) \]
Empirically it fights overfitting hard. The surprise is why it works — and the answer ties straight back to priors and posteriors.
Explain it
Discussion prompt
Explain Dropout: randomly delete neurons to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Empirically it fights overfitting hard. The surprise is why it works — and the answer ties straight back to priors and posteriors.
Intuition
Every dropout mask defines a different thinned network sharing the same weights. Over a training run you visit exponentially many of them, and prediction with the full net (weights scaled) approximates averaging their outputs — a giant implicit ensemble.
Gal & Ghahramani showed this is variational Bayes: dropout training approximately minimizes the KL divergence to a posterior over the weights. The random masks are samples from an approximate posterior — regularization and Bayesian inference, once again the same object seen from two sides.
Concept
A layer of n units with independent on/off masks has 2ⁿ possible thinned sub-networks. Dropout trains a sample of them, all sharing weights — an astronomically large implicit ensemble at the cost of one network.
| units n | distinct sub-networks 2ⁿ | roughly |
|---|---|---|
| 10 | 2¹⁰ | 1,024 |
| 100 | 2¹⁰⁰ | 1.27 × 10³⁰ |
| 1000 | 2¹⁰⁰⁰ | > atoms in the universe |
You can never enumerate them, but you don't have to: averaging with the weights scaled by (1−p) approximates the ensemble prediction — the same integrate-out-the-uncertainty move that Bayes makes, done cheaply.
Comparison
Comparison matrix
From How big is that ensemble?: refill the distinct sub-networks 2ⁿ column from what you know. The rest of the table is as it appeared.
| units n | distinct sub-networks 2ⁿ | roughly |
|---|---|---|
| 10 | 2¹⁰ | 1,024 |
| 100 | 2¹⁰⁰ | 1.27 × 10³⁰ |
| 1000 | 2¹⁰⁰⁰ | > atoms in the universe |
Concept
Because dropout ≈ sampling the posterior, you can keep it on at test time and run the network many times: the spread of the predictions estimates model uncertainty. This is Monte-Carlo (MC) dropout.
So the same log-prior story scales all the way up: a penalty on weights (Parts 2–3), a conjugate posterior you can write down (Part 4), and here an approximate posterior you sample — three faces of belief plus data.
Section
Part 6 of 7 — verified vs sklearn
Fill the middle
Fill in the blanks
From Sparse-truth data: only 3 of 10 features matter — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print('true nonzero idx:', np.flatnonzero(true))
print('X shape:', X.shape)
Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. The true weight vector is zero everywhere except three slots.
Worked example
Build 50 samples of 10 features where the target truly depends on only features 0, 3, 7 (weights 3, −2, 1.5), plus small noise. Fixed RNG seed so you see these exact numbers:
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print('true nonzero idx:', np.flatnonzero(true))
print('X shape:', X.shape)Only indices 0, 3, 7 carry signal; the other 7 are pure noise
Why: The true weight vector is zero everywhere except three slots. A good sparse learner should discover this support and zero the rest.
| object | value (verified) |
|---|---|
| true nonzero indices | [0, 3, 7] |
| true weights there | [3.0, −2.0, 1.5] |
| X shape | (50, 10) |
Missing information
Discussion prompt
Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Lasso recovers the true sparse support; ridge keeps all 10 coefficients small-but-nonzero (its largest 'noise' coefficient is still ≈0.065).
Worked example
Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:
from sklearn.linear_model import Ridge, Lasso
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0)))
print('lasso keeps:', np.flatnonzero(lasso.coef_))ridge: 0 exact zeros; lasso: 7 exact zeros, keeps exactly {0, 3, 7}
Why: Lasso recovers the true sparse support; ridge keeps all 10 coefficients small-but-nonzero (its largest 'noise' coefficient is still ≈0.065).
| feature | true | ridge coef | lasso coef |
|---|---|---|---|
| 0 | 3.00 | 2.9312 | 2.9291 |
| 3 | −2.00 | −1.9570 | −1.9212 |
| 7 | 1.50 | 1.4944 | 1.4284 |
| other 7 | 0 | up to ±0.065 (nonzero) | 0.0000 (exact) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
ridge: 0 exact zeros; lasso: 7 exact zeros, keeps exactly {0, 3, 7}
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:
Estimation
Predict first
λ (sklearn's alpha) is the prior strength — crank it up and more coordinates snap to zero. Sweep it and count the survivors:
Commit before you compute: what does The lasso path: λ is the sparsity dial come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Too small (0.01) keeps noise features; the sweet spot 0.1–1.0 recovers exactly {0,3,7}; too big (3.0) kills everything
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A weak prior barely thresholds, so noise coordinates survive; a strong prior over-shrinks and zeros even the true features.
Worked example
λ (sklearn's alpha) is the prior strength — crank it up and more coordinates snap to zero. Sweep it and count the survivors:
from sklearn.linear_model import Lasso
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
for a in [0.01, 0.1, 0.5, 1.0, 3.0]:
L = Lasso(alpha=a).fit(X, y)
print(f'alpha={a:<5} nonzero={int(np.sum(L.coef_ != 0))}',
'keeps', np.flatnonzero(L.coef_))Too small (0.01) keeps noise features; the sweet spot 0.1–1.0 recovers exactly {0,3,7}; too big (3.0) kills everything
Why: A weak prior barely thresholds, so noise coordinates survive; a strong prior over-shrinks and zeros even the true features. You tune λ on validation data to land in the middle band.
| alpha (λ) | # nonzero | features kept (verified) |
|---|---|---|
| 0.01 | 5 | {0, 2, 3, 7, 8} |
| 0.1 | 3 | {0, 3, 7} |
| 0.5 | 3 | {0, 3, 7} |
| 1.0 | 3 | {0, 3, 7} |
| 3.0 | 0 | { } — all zero |
Discrimination
Sort into buckets
Sort these by # nonzero, from memory, without looking back at The lasso path: λ is the sparsity dial. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
They are both just "regularization", so pick whichever — they do the same job, and I need to drop useless features, so ridge is fine.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients.
Match the penalty to the goal: L2 for stability/shrinkage, L1 for sparsity/feature selection.
Why: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients. If the goal is a sparse, interpretable model, L2 is the wrong tool: you get a dense weight vector, not a shortlist.
Trap
They are both just "regularization", so pick whichever — they do the same job, and I need to drop useless features, so ridge is fine.
Use ridge for feature selection
Why: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients. If the goal is a sparse, interpretable model, L2 is the wrong tool: you get a dense weight vector, not a shortlist.
Match the penalty to the goal: L2 for stability/shrinkage, L1 for sparsity/feature selection.
Lasso (L1) to select features; ridge (L2) to tame collinearity
Why: Different priors (Laplace vs Gaussian), different geometry (diamond vs circle), different solutions (7 zeros vs 0). Elastic Net blends both when you want some of each.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.; Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:; The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².; Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.Pattern
Predict first
The table runs: sklearn Ridge | 2.844950 | −0.030003 | 0.012108 · closed form | 2.844950 | −0.030003 | 0.012108
In Verify: the closed form is really MAP, given the rows so far: what is the next one — the row where source is allclose?
Correct: allclose | True | — | —
| source | coef[0] | coef[1] | coef[2] |
|---|---|---|---|
| sklearn Ridge | 2.844950 | −0.030003 | 0.012108 |
| closed form | 2.844950 | −0.030003 | 0.012108 |
| allclose | True | — | — |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. sklearn's alpha is our λ. Solving (XᵀX+λI)w = Xᵀy reproduces it exactly, closing the loop: the Gaussian log-prior and the L2 penalty are the same math.
Concept
One last sanity check that ridge = Gaussian-prior MAP: sklearn's Ridge(alpha=2, fit_intercept=False) must equal solving (XᵀX + 2I)w = Xᵀy by hand.
from sklearn.linear_model import Ridge
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
Rr = Ridge(alpha=2.0, fit_intercept=False).fit(X, y)
w_cf = np.linalg.solve(X.T @ X + 2.0*np.eye(10), X.T @ y)
print('sklearn [:3]:', Rr.coef_[:3].round(6))
print('closed [:3]:', w_cf[:3].round(6))
print('match:', np.allclose(Rr.coef_, w_cf))Identical to 6 decimals → the penalty IS the log-prior
Why: sklearn's alpha is our λ. Solving (XᵀX+λI)w = Xᵀy reproduces it exactly, closing the loop: the Gaussian log-prior and the L2 penalty are the same math.
| source | coef[0] | coef[1] | coef[2] |
|---|---|---|---|
| sklearn Ridge | 2.844950 | −0.030003 | 0.012108 |
| closed form | 2.844950 | −0.030003 | 0.012108 |
| allclose | True | — | — |
Analogy
Discussion prompt
Explain Verify: the closed form is really MAP by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One last sanity check that ridge = Gaussian-prior MAP: sklearn's Ridge(alpha=2, fit_intercept=False) must equal solving (XᵀX + 2I)w = Xᵀy by hand.
Concept
Honest caveat: here n=50 ≥ d=10 with tiny noise, so plain OLS barely overfits and has the lowest test error. Regularization is not free — it trades a little bias for variance reduction, and pays off most when data is scarce or features collinear.
| model | test MSE (1000 held-out, verified) | ‖coef − true‖ |
|---|---|---|
| OLS | 0.0118 | 0.0498 |
| ridge (α=1) | 0.0217 | 0.1077 |
| lasso (α=0.1) | 0.0270 | 0.1280 |
So why reach for lasso here? Not for accuracy — for interpretability. It is the only fit that hands you the exact 3-feature story {0,3,7} instead of 10 muddy coefficients. Sparsity is a modeling goal, distinct from raw predictive error.
Pattern
Step through it
Step through When does regularization actually help? one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
These are the steps of Regularization-as-prior recipe, scrambled. Put them back in order before the next slide shows you.
−log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)−log p(θ) — the belief, in log formλ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zerosλ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zerosλ → 0 = back to MLEWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
−log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)−log p(θ) — the belief, in log formλ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zerosλ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zerosλ → 0 = back to MLEEdge cases
Discussion prompt
Regularization-as-prior recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
−log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)−log p(θ) — the belief, in log formλ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zerosλ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zerosλ → 0 = back to MLEElimination
Eliminate the wrong options
L2 regularization (ridge) is the MAP estimate under which prior on the weights?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: −log of a zero-mean Gaussian is proportional to ‖w‖², so a Gaussian prior adds exactly the L2 penalty. Minimizing loss + λ‖w‖² is ridge, and its strength is λ = σ²/τ².
Check
Match the penalty to its prior distribution.
Check your understanding
L2 regularization (ridge) is the MAP estimate under which prior on the weights?
Answer: A
Why: −log of a zero-mean Gaussian is proportional to ‖w‖², so a Gaussian prior adds exactly the L2 penalty. Minimizing loss + λ‖w‖² is ridge, and its strength is λ = σ²/τ².
Prediction
Predict first
Why does L1 (lasso) set coefficients to EXACTLY zero while L2 (ridge) does not?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The L1 penalty's gradient magnitude stays constant (λ) as w → 0, so it pushes a small weight all the way to 0; L2's gradient 2λw fades to 0 and never finishes the job
Why: The subgradient of λ|w| has constant magnitude λ (its sign), so the pull toward 0 never weakens — the soft-threshold snaps small coordinates to exactly 0. L2's gradient 2λw shrinks with w and vanishes at the origin, so ridge only asymptotes toward 0.
Check
Think about the gradient of the penalty as w → 0.
Check your understanding
Why does L1 (lasso) set coefficients to EXACTLY zero while L2 (ridge) does not?
Answer: A
Why: The subgradient of λ|w| has constant magnitude λ (its sign), so the pull toward 0 never weakens — the soft-threshold snaps small coordinates to exactly 0. L2's gradient 2λw shrinks with w and vanishes at the origin, so ridge only asymptotes toward 0.
Prediction
Predict first
With a Beta(α, β) prior and k successes in n Bernoulli trials, the posterior is…
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Beta(α + k, β + n − k)
Why: Multiplying pᵏ(1−p)ⁿ⁻ᵏ by the Beta kernel p^(α−1)(1−p)^(β−1) adds exponents: α gets the successes k, β gets the failures n−k. The posterior is Beta(α+k, β+n−k) — that's conjugacy.
Check
Recall the conjugate update: successes go to α, failures to β.
Check your understanding
With a Beta(α, β) prior and k successes in n Bernoulli trials, the posterior is…
Answer: A
Why: Multiplying pᵏ(1−p)ⁿ⁻ᵏ by the Beta kernel p^(α−1)(1−p)^(β−1) adds exponents: α gets the successes k, β gets the failures n−k. The posterior is Beta(α+k, β+n−k) — that's conjugacy.
Elimination
Eliminate the wrong options
In the ridge MAP objective ‖Xw−y‖² + λ‖w‖² with λ = σ²/τ², what happens as the prior variance τ² → ∞?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: λ = σ²/τ² → 0 as τ² → ∞. A prior with infinite variance is flat (uninformative), so its log-term is constant and MAP becomes MLE — plain least squares, w = [2.2, 0.6] on our data.
Check
What happens to MAP as the prior gets flat?
Check your understanding
In the ridge MAP objective ‖Xw−y‖² + λ‖w‖² with λ = σ²/τ², what happens as the prior variance τ² → ∞?
Answer: A
Why: λ = σ²/τ² → 0 as τ² → ∞. A prior with infinite variance is flat (uninformative), so its log-term is constant and MAP becomes MLE — plain least squares, w = [2.2, 0.6] on our data.
Section
Part 7 of 7 — the project
Concept
Build data with a known sparse support, fit ridge and lasso, and confirm that only lasso recovers the zeros — watching a Laplace prior turn into feature selection. Every piece is derived; now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | Data where only features 0, 3, 7 matter | numpy default_rng + a sparse true w |
| 2 | Fit Ridge(alpha=1) and Lasso(alpha=0.1) | sklearn.linear_model |
| 3 | Count exact zeros and compare to true support | np.sum(coef == 0) |
Build rules: type every line, set the RNG seed default_rng(1) so results are reproducible, run after each block, and compare coefficients to the true support [0, 3, 7].
Counterexample
Discussion prompt
Build rules: type every line, set the RNG seed default_rng(1) so results are reproducible, run after each block, and compare coefficients to the true support [0, 3, 7].
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: make a true weight vector with nonzeros only at 0, 3, 7, then generate y = X·w_true + small noise. Predict how many of the 10 true weights are zero before you print.
Hint: true = np.zeros(10); true[[0,3,7]] = [3, -2, 1.5], then y = X @ true + 0.1*rng.normal(size=50). Seed with np.random.default_rng(1).
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print(np.flatnonzero(true))
print(X.shape)| check | value (verified) |
|---|---|
| true nonzero indices | [0, 3, 7] |
| number of true zeros | 7 of 10 |
| X shape | (50, 10) |
Worked example
Your turn: fit ridge and lasso on the same X, y. Predict which one's coefficient vector contains exact zeros before printing.
Hint: Ridge(alpha=1.0).fit(X, y) and Lasso(alpha=0.1).fit(X, y); read .coef_ and .round(3).
import numpy as np
from sklearn.linear_model import Ridge, Lasso
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge:', ridge.coef_.round(3))
print('lasso:', lasso.coef_.round(3))| model | features 0 / 3 / 7 | the other 7 |
|---|---|---|
| ridge | 2.931 / −1.957 / 1.494 | small but nonzero |
| lasso | 2.929 / −1.921 / 1.428 | exactly 0 |
Worked example
Your turn: count exact zeros in each model and compare to the true support. How many should lasso zero out?
Hint: ridge coefficients are never truly 0, so test np.abs(ridge.coef_) < 1e-6; lasso really hits 0, so lasso.coef_ == 0 works.
import numpy as np
from sklearn.linear_model import Ridge, Lasso
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0)), 'of 10')
print('lasso keeps:', np.flatnonzero(lasso.coef_))| model | exact zeros (verified) | support recovered? |
|---|---|---|
| ridge | 0 of 10 | no — keeps all 10 |
| lasso | 7 of 10 | yes — keeps {0, 3, 7} |
Trade off
Comparison matrix
From Milestone 3 — count the zeros: every row here is a choice with a cost. Fill the exact zeros (verified) column, then say which row you would actually pick and what you give up for it.
| model | exact zeros (verified) | support recovered? |
|---|---|---|
| ridge | 0 of 10 | no — keeps all 10 |
| lasso | 7 of 10 | yes — keeps {0, 3, 7} |
Concept
import numpy as np
from sklearn.linear_model import Ridge, Lasso
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50) # 1. sparse-truth data
ridge = Ridge(alpha=1.0).fit(X, y) # 2. Gaussian-prior MAP
lasso = Lasso(alpha=0.1).fit(X, y) # Laplace-prior MAP
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0))) # 3. count -> 7
print('lasso keeps:', np.flatnonzero(lasso.coef_))| printed line | value (verified) |
|---|---|
| ridge zeros: | 0 |
| lasso zeros: | 7 |
| lasso keeps: | [0 3 7] |
If lasso reports 7 zeros and keeps exactly features 0, 3, 7 while ridge keeps all 10 — you have watched a prior become a penalty, and a Laplace prior perform feature selection.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| ridge zeros: | 0 |
| lasso zeros: | 7 |
| lasso keeps: | [0 3 7] |
Concept
Slides closed, out loud: explain (1) how log p(θ) becomes an additive penalty and why λ = σ²/τ², (2) why the L1 diamond produces zeros while the L2 circle only shrinks, and (3) what conjugacy saves you from computing.
Stretch: derive the one-coordinate soft-threshold formula soft(ρ, λ/2) from ∂/∂w [ (w−ρ)² + λ|w| ] = 0, then confirm your lasso coefficients on features 0/3/7 match the sklearn values above. This thread continues into Elastic Net (Week 21), dropout-as-Bayes (Week 18), and Adam's decoupled weight decay (Week 35).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — MAP: a prior is a belief · Gaussian prior → L2 / ridge · Laplace prior → L1 / lasso · Conjugate priors · Dropout as approximate Bayes · Ridge vs Lasso — the real fit. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
p(D), take the log, split the product into a sumλ‖w‖² with λ = σ²/τ², and a Laplace prior into lasso λ‖w‖₁sklearn exactly| idea | the one thing to remember |
|---|---|
| MAP | MLE + log-prior — a penalty added to the loss |
| ridge | Gaussian prior → L2 → shrink, never zero; λ = σ²/τ² |
| lasso | Laplace prior → L1 → constant pull → exact zeros |
| conjugacy | posterior stays in the family; add counts (Beta, Dirichlet) |
| dropout | random masks = samples from an approximate posterior |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.