Lesson 12: MAP, Priors & Regularization

USAAIO Lesson 12, from Week 4 on probability, fully worked. It derives MAP from Bayes' rule as MLE plus a log-prior, showing every logarithmic step, then turns a Gaussian prior into L2 or ridge and a Laplace prior into L1 or lasso, term by term. It gives the exact reason L1 zeros out coordinates, through the subgradient and the soft-threshold, computed one coordinate at a time, proves Beta-Binomial conjugacy by proportionality, and covers dropout as approximate Bayesian inference. It ends with a ridge-against-lasso sparsity project verified against scikit-learn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 115 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. MAP, Priors & Regularization

Title

USAAIO · Lesson 12 · Week 4 (Probability)

Regularization is not a hack bolted onto the loss — it is a prior belief, written in log form. We derive MAP = MLE + log-prior from Bayes' rule with no skipped step, turn a Gaussian prior into ridge and a Laplace prior into lasso term by term, prove exactly why L1 drives weights to zero, and verify all of it against scikit-learn.

2. By the end of this lesson you can

Objectives

  1. State Bayes' rule for parameters and derive MAP as arg max [log-likelihood + log-prior], every log step shown
  2. Turn a zero-mean Gaussian prior into the L2 / ridge penalty and read off λ = σ²/τ²
  3. Turn a Laplace prior into the L1 / lasso penalty, and explain why L1 is sparse using the subgradient and the soft-threshold formula
  4. Recognize conjugate priors (Beta–Binomial, Dirichlet–Categorical) and update Beta(α,β) by inspection
  5. Read dropout as approximate Bayesian inference, and demonstrate ridge-vs-lasso sparsity from scratch — matching sklearn

3. What survived from Autoencoders — Standard, Denoising & Sparse?

Warm-up

Discussion prompt

Before we open Lesson 12: MAP, Priors & Regularization: without looking back, what was the main idea of Autoencoders — Standard, Denoising & Sparse, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the autoencoder as encoder -> bottleneck latent z -> decoder trained on reconstruction loss alone; undercomplete compression; the denoising autoencoder (corrupt the input, reconstruct the clean signal) and the sparse autoencoder (L1 penalty on activations, the workhorse of mechanistic interpretability); latent-space interpolation; and the key limitation that motivates the VAE — a standard AE's latent space has no prior, so you cannot sample from it. Build a convolutional autoencoder your-turn.

4. MAP: a prior is a belief

Section

Part 1 of 7 — from Bayes to a penalty

5. The running example: fit a line, but small data

Concept

We reuse Lesson 7's five students — hours studied x and score y — and fit ŷ = w₀ + w₁·x. But now we ask a probability question: with so few points, what should keep the weights from swinging wildly?

studenthours xscore y
112
224
335
444
555

The plain least-squares fit is w = [2.2, 0.6]. This lesson asks: what belief about w would nudge that fit — and where does that belief come from? The answer is a prior.

6. Fill in: hours x for The running example: fit a line, but small…

Comparison

Comparison matrix

From The running example: fit a line, but small data: refill the hours x column from what you know. The rest of the table is as it appeared.

studenthours xscore y
112
224
335
444
555

7. MLE recap: maximize the likelihood

Concept

In Lesson 8 the maximum-likelihood estimate picked the θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.

\[ \theta_{\text{MLE}} = \arg\max_{\theta}\; p(\text{data}\mid\theta) \]

MLE uses only the data. With five points and two weights it is fine here — but with many features and few rows it overfits, chasing noise. We want a principled way to say "and I also believe the weights are small."

8. Break it if you can: MLE recap: maximize the likelihood

Counterexample

Discussion prompt

In Lesson 8 the maximum-likelihood estimate picked the θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. A prior is what you believe before the data

Intuition

Before seeing a single point, you already suspect the weights are modest — a giant slope would be surprising. That suspicion is a prior distribution p(θ): it puts more probability on small weights than on huge ones.

MAP blends the two voices: the data (likelihood) says "here is what fits", the prior says "but stay reasonable". With little data the prior has a loud say; with mountains of data the likelihood drowns it out.

That single idea — belief as a distribution you can multiply in — is what turns a regularizer from a fudge factor into a probability statement.

10. Bayes' rule, for parameters

Concept

Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:

\[ p(\theta \mid D) = \frac{p(D \mid \theta)\, p(\theta)}{p(D)} \]

posterior — The updated belief about θ AFTER seeing the data — proportional to (likelihood × prior). MAP picks the single θ at its peak; full Bayes keeps the whole distribution.

11. By analogy: Bayes' rule, for parameters

Analogy

Discussion prompt

Explain Bayes' rule, for parameters by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:

12. What has to happen first: Derive MAP: drop the constant, take the log

Ranking

Put in order

Put the moves of Derive MAP: drop the constant, take the log into the order they have to happen.

  1. Start from the posterior
  2. Drop p(D) — it does not depend on θ
  3. Take logs — log is monotonic

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. MAP is the arg max of the posterior over θ — the most probable parameter given the data.

13. Derive MAP: drop the constant, take the log

Worked example

Start from the posterior

Why: MAP is the arg max of the posterior over θ — the most probable parameter given the data.

\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; \frac{p(D\mid\theta)\,p(\theta)}{p(D)} \]

Drop p(D) — it does not depend on θ

Why: The denominator p(D) is a fixed number once the data is in; scaling by a positive constant never moves an arg max. So the posterior is PROPORTIONAL to likelihood × prior.

\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; p(D\mid\theta)\,p(\theta) \]

Take logs — log is monotonic

Why: log is strictly increasing, so arg max is unchanged; and log turns the product into a SUM, which is far easier to differentiate.

\[ \theta_{\text{MAP}} = \arg\max_{\theta}\; \big[\log p(D\mid\theta) + \log p(\theta)\big] \]

14. Decode the notation: Derive MAP: drop the constant, take the log

Notation

Annotate

From Derive MAP: drop the constant, take the log — read this one piece at a time. What is each part doing?

On: \( \theta_{\text{MAP}} = \arg\max_{\theta}\; \frac{p(D\mid\theta)\,p(\theta)}{p(D)} \)

  • MAP is the arg max of the posterior over θ — the most probable parameter given the data.
  • The denominator p(D) is a fixed number once the data is in; scaling by a positive constant never moves an arg max. So the posterior is PROPORTIONAL to likelihood × prior.
  • log is strictly increasing, so arg max is unchanged; and log turns the product into a SUM, which is far easier to differentiate.

15. The one-line summary: MAP = MLE + log-prior

Concept

The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:

\[ \boxed{\;\theta_{\text{MAP}} = \arg\max_{\theta}\; \underbrace{\log p(D\mid\theta)}_{\text{MLE term}} + \underbrace{\log p(\theta)}_{\text{regularizer}}\;} \]

Flip the sign to turn maximize into minimize and −log p(θ) becomes a penalty added to the loss. Every regularizer you know is secretly some −log p(θ). The rest of Part 1 shows which prior gives which penalty.

16. Teach it back: The one-line summary: MAP = MLE + log-prior

Explain it

Discussion prompt

Explain The one-line summary: MAP = MLE + log-prior to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:

17. MAP is a shortcut past the hard integral

Intuition

Full Bayesian inference keeps the whole posterior p(θ|D), which needs the denominator p(D) = ∫ p(D|θ)p(θ) dθ — an integral that is usually intractable.

MAP dodges it. Since p(D) doesn't depend on θ, it cannot move the peak, so MAP finds the single most-probable θ without ever computing the integral — just maximize likelihood × prior.

The cost of the shortcut: you get a point, not a distribution — no uncertainty bars. That trade (a penalty you can optimize vs. a full posterior) is exactly the regularization story, and dropout in Part 5 buys some uncertainty back.

18. Plan first: A tiny MAP by hand — one weight

Step zero

Discussion prompt

A tiny MAP by hand — one weight — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Sums: Σxᵢ² = 1+4+9 = 14, Σxᵢyᵢ = 2.1+7.8+18.6 = 28.5

Answer:

  1. Sums: Σxᵢ² = 1+4+9 = 14, Σxᵢyᵢ = 2.1+7.8+18.6 = 28.5
  2. MLE (no prior): w = 28.5 / 14 = 2.0357
  3. MAP with σ²=1, τ²=0.25 ⇒ λ = σ²/τ² = 4: w = 28.5 / (14+4) = 1.5833

19. A tiny MAP by hand — one weight

Worked example

Take three points x = [1,2,3], y = [2.1, 3.9, 6.2] and the one-weight model ŷ = w·x with Gaussian noise σ² and a Gaussian prior w ~ N(0, τ²). The MAP weight has a clean closed form (derived fully in Part 2):

\[ w_{\text{MAP}} = \frac{\sum_i x_i y_i}{\sum_i x_i^2 + \sigma^2/\tau^2} \]

Sums: Σxᵢ² = 1+4+9 = 14, Σxᵢyᵢ = 2.1+7.8+18.6 = 28.5

Why: These are the same two sums OLS needs. Σxᵢyᵢ = 1(2.1)+2(3.9)+3(6.2) = 2.1+7.8+18.6 = 28.5.

MLE (no prior): w = 28.5 / 14 = 2.0357

Why: Setting σ²/τ² = 0 recovers plain least squares: 28.5/14 = 2.035714.

MAP with σ²=1, τ²=0.25 ⇒ λ = σ²/τ² = 4: w = 28.5 / (14+4) = 1.5833

Why: The prior adds 4 to the denominator, pulling the estimate from 2.0357 toward 0. Verified: 28.5/18 = 1.583333.

estimatordenominatorw
MLE (λ=0)142.0357
MAP (λ=4)14 + 4 = 181.5833

20. What each one costs: A tiny MAP by hand — one weight

Trade off

Comparison matrix

From A tiny MAP by hand — one weight: every row here is a choice with a cost. Fill the denominator column, then say which row you would actually pick and what you give up for it.

estimatordenominatorw
MLE (λ=0)142.0357
MAP (λ=4)14 + 4 = 181.5833

21. Guess the shape of the answer: The tiny MAP in code

Estimation

Predict first

Confirm the by-hand numbers with a runnable snippet. The MAP weight is the same two sums with σ²/τ² added to the denominator:

Commit before you compute: what does The tiny MAP in code come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: λ = σ²/τ² = 4 pulls the estimate from 2.0357 toward 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Adding λ to the denominator is exactly the 1-D case of (XᵀX + λI).

22. The tiny MAP in code

Worked example

Confirm the by-hand numbers with a runnable snippet. The MAP weight is the same two sums with σ²/τ² added to the denominator:

import numpy as np
xs = np.array([1., 2., 3.])
ys = np.array([2.1, 3.9, 6.2])
Sxx = np.sum(xs**2); Sxy = np.sum(xs*ys)
lam = 1.0/0.25            # sigma^2 / tau^2 = 1 / 0.25 = 4
print('Sxx, Sxy:', Sxx, Sxy)
print('MLE:', Sxy/Sxx)
print('MAP:', Sxy/(Sxx+lam))

λ = σ²/τ² = 4 pulls the estimate from 2.0357 toward 0

Why: Adding λ to the denominator is exactly the 1-D case of (XᵀX + λI). Larger λ = larger denominator = more shrinkage toward the prior mean 0.

printvalue (verified)
Sxx, Sxy14.0 28.5
MLE2.035714
MAP1.583333

23. Watch it run: The tiny MAP in code

Pattern

Step through it

Step through The tiny MAP in code one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: print is Sxx, Sxy
  2. Step 2: print is MLE
  3. Step 3: print is MAP

24. Gaussian prior → L2 / ridge

Section

Part 2 of 7 — term by term

25. The Gaussian noise likelihood (from Lesson 8)

Concept

Assume each target is the line plus independent Gaussian noise, yᵢ = xᵢᵀw + εᵢ with εᵢ ~ N(0, σ²). Then the negative log-likelihood is the familiar sum of squares over 2σ²:

\[ -\log p(y\mid X, w) = \frac{1}{2\sigma^2}\lVert Xw - y\rVert^2 + \text{const} \]

The constant collects n·½log(2πσ²) — no w in it, so it drops out of any minimization over w. Minimizing this alone is exactly OLS: MLE is least squares.

26. Plan first: The Gaussian log-prior is a negative quadratic

Step zero

Discussion prompt

The Gaussian log-prior is a negative quadratic — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Density: p(w) ∝ exp(−‖w‖² / (2τ²))

Answer:

  1. Density: p(w) ∝ exp(−‖w‖² / (2τ²))
  2. −log p(w) = ‖w‖²/(2τ²) + const
  3. Numeric check (τ²=0.5): the quadratic part is w²/(2τ²)

27. The Gaussian log-prior is a negative quadratic

Worked example

Put a zero-mean Gaussian prior on each weight, w ~ N(0, τ²I). Write its density and take −log:

Density: p(w) ∝ exp(−‖w‖² / (2τ²))

Why: A multivariate zero-mean Gaussian with covariance τ²I has this exponential kernel; the leading constant does not involve w.

\[ p(w) = \frac{1}{(2\pi\tau^2)^{d/2}}\exp\!\Big(-\frac{\lVert w\rVert^2}{2\tau^2}\Big) \]

−log p(w) = ‖w‖²/(2τ²) + const

Why: log of an exponential is its exponent; the −½log(2πτ²)ᵈ piece is constant in w. So the prior contributes a term PROPORTIONAL to ‖w‖².

\[ -\log p(w) = \frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]

Numeric check (τ²=0.5): the quadratic part is w²/(2τ²)

Why: At w=0,1,2 the −log density is 0.5724, 1.5724, 4.5724; subtract the const 0.5724 and you get 0, 1, 4 = w²/(2·0.5). The w-dependent part is exactly the quadratic.

w−log p(w)quadratic part w²/(2τ²)constant
00.57240.00000.5724
11.57241.00000.5724
24.57244.00000.5724

28. Work backwards from the answer: The Gaussian log-prior is a negative…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Numeric check (τ²=0.5): the quadratic part is w²/(2τ²)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Put a zero-mean Gaussian prior on each weight, w ~ N(0, τ²I). Write its density and take −log:

29. What has to happen first: Add the two −logs → the ridge objective

Ranking

Put in order

Put the moves of Add the two −logs → the ridge objective into the order they have to happen.

  1. MAP minimizes −(log-likelihood + log-prior)
  2. Multiply through by 2σ² (a positive constant)
  3. Name the ratio λ = σ²/τ² — this IS ridge

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Flip the arg max of the sum of logs into an arg min of the negative sum.

30. Add the two −logs → the ridge objective

Worked example

MAP minimizes −(log-likelihood + log-prior)

Why: Flip the arg max of the sum of logs into an arg min of the negative sum. Add the two negative-log pieces from the last two slides.

\[ -\log p(w\mid D) = \frac{1}{2\sigma^2}\lVert Xw-y\rVert^2 + \frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]

Multiply through by 2σ² (a positive constant)

Why: Scaling by 2σ² > 0 does not move the minimizer; it clears the fraction on the loss term and folds σ² into the penalty coefficient.

\[ \lVert Xw-y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \]

Name the ratio λ = σ²/τ² — this IS ridge

Why: The MAP objective is Lesson 7's ridge loss exactly. The regularization strength is the noise-to-prior variance ratio: noisier data or a tighter prior ⇒ larger λ ⇒ more shrinkage.

\[ \boxed{\;\min_w\; \lVert Xw-y\rVert^2 + \lambda\lVert w\rVert^2,\qquad \lambda = \frac{\sigma^2}{\tau^2}\;} \]

31. Say it in words: Add the two −logs → the ridge objective

Translation

\( \lVert Xw-y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

32. What has to be given first: Solve ridge on the 5 students — closed form

Missing information

Discussion prompt

Ridge's normal equations are (XᵀX + λI)w = Xᵀy (Lesson 7). Solve them with λ = 1 on the student data and compare to plain OLS. Runnable as written:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Adding np.eye(2) makes the Gaussian-prior term appear as +λI. The exact same solve returns a shrunken weight vector.

33. Solve ridge on the 5 students — closed form

Worked example

Ridge's normal equations are (XᵀX + λI)w = Xᵀy (Lesson 7). Solve them with λ = 1 on the student data and compare to plain OLS. Runnable as written:

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
w_ridge = np.linalg.solve(X.T @ X + 1.0*np.eye(2), X.T @ y)
print('OLS  :', w_ols)
print('ridge:', w_ridge)
print('norms:', np.linalg.norm(w_ols), np.linalg.norm(w_ridge))

λI lifts every eigenvalue of XᵀX by λ, then solve

Why: Adding np.eye(2) makes the Gaussian-prior term appear as +λI. The exact same solve returns a shrunken weight vector.

fitw = [w₀, w₁]‖w‖ (verified)
OLS (λ=0)[2.2, 0.6]2.2804
ridge (λ=1)[1.1712, 0.8649]1.4559

34. Ridge shrinks — it never zeros

Intuition

The Gaussian prior is a smooth bump at 0: near the origin it is nearly flat, so it barely pulls once a weight is already small. Ridge therefore squeezes every coefficient toward 0 but leaves it a little nonzero.

On the students, ‖w‖ dropped from 2.28 to 1.46 — both entries got smaller, neither vanished. To get exact zeros we need a prior with a sharp spike at 0. That is the Laplace prior, next.

35. This is exactly "weight decay"

Intuition

In deep learning, the same L2 term is called weight decay: every gradient step, before applying the data gradient, you shrink w a little toward 0. That shrink IS the gradient of λ‖w‖².

So the optimizer setting you flip on out of habit (weight_decay=1e-4) is a Gaussian prior on every weight. Naming it 'decay' hides the probability story — but it is the identical −log p(w) term you just derived.

36. Something is wrong here: is λ = τ²/σ² or σ²/τ²?

Anomaly

Predict first

A student writes this, and it looks reasonable:

A tighter prior (small τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks.

Read it straight off the derivation: after clearing the loss fraction, the penalty coefficient is the noise variance over the prior variance.

Why: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks. τ²/σ² shrinks as τ²→0, giving less penalty exactly when you wanted more.

37. Trap: is λ = τ²/σ² or σ²/τ²?

Trap

The trap

A tighter prior (small τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².

λ = τ²/σ² ✗

Why: This gets it backwards. A TIGHTER prior (small τ²) is a STRONGER belief that weights are small, so it must penalize MORE — λ should grow as τ² shrinks. τ²/σ² shrinks as τ²→0, giving less penalty exactly when you wanted more.

The fix

Read it straight off the derivation: after clearing the loss fraction, the penalty coefficient is the noise variance over the prior variance.

λ = σ²/τ² ✓

Why: Small τ² (tight prior) ⇒ large λ ⇒ heavy shrinkage. Large σ² (noisy data, trust the fit less) ⇒ large λ too. Both directions match intuition; τ²→∞ (flat prior) ⇒ λ→0 ⇒ back to MLE.

38. Laplace prior → L1 / lasso

Section

Part 3 of 7 — and why it is sparse

39. Picture it first: The Laplace prior spikes at zero

Picture it

Figure (svg): Two prior densities over w centered at zero: a smooth rounded Gaussian bump, and a Laplace density with a sharp cusp (a kink) at zero.

Gaussian (smooth top) vs Laplace (sharp cusp at 0). The kink is what forces exact zeros.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Swap the Gaussian bell for a Laplace (double-exponential) prior on each weight. It has a sharp peak at 0 and heavier-than-Gaussian tails:

40. The Laplace prior spikes at zero

Concept

Swap the Gaussian bell for a Laplace (double-exponential) prior on each weight. It has a sharp peak at 0 and heavier-than-Gaussian tails:

\[ p(w_j) = \frac{1}{2b}\exp\!\Big(-\frac{|w_j|}{b}\Big) \]

Figure (svg): Two prior densities over w centered at zero: a smooth rounded Gaussian bump, and a Laplace density with a sharp cusp (a kink) at zero.

Gaussian (smooth top) vs Laplace (sharp cusp at 0). The kink is what forces exact zeros.

41. Plan first: The Laplace log-prior is proportional to |w|

Step zero

Discussion prompt

The Laplace log-prior is proportional to |w| — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: −log p(wⱼ) = |wⱼ|/b + log(2b)

Answer:

  1. −log p(wⱼ) = |wⱼ|/b + log(2b)
  2. Numeric check (b=1): the w-part is exactly |w|
  3. Add it to the Gaussian-noise NLL → lasso

42. The Laplace log-prior is proportional to |w|

Worked example

−log p(wⱼ) = |wⱼ|/b + log(2b)

Why: Take −log of the density: the exponential's exponent |wⱼ|/b comes down linearly, and −log(1/(2b)) = log(2b) is the constant.

\[ -\log p(w) = \frac{1}{b}\sum_j |w_j| + \text{const} = \frac{1}{b}\lVert w\rVert_1 + \text{const} \]

Numeric check (b=1): the w-part is exactly |w|

Why: At w = −2,−1,0,1,2 the −log density is 2.6931, 1.6931, 0.6931, 1.6931, 2.6931. Subtract the const log(2)=0.6931 and you get 2,1,0,1,2 = |w|.

w−log p(w)|w|/bconstant log(2b)
−22.69312.00000.6931
−11.69311.00000.6931
00.69310.00000.6931
11.69311.00000.6931
22.69312.00000.6931

Add it to the Gaussian-noise NLL → lasso

Why: Same MAP recipe: loss + −log prior. The L1 term is the lasso penalty.

\[ \boxed{\;\min_w\; \lVert Xw-y\rVert^2 + \lambda\lVert w\rVert_1\;} \]

43. What stays fixed: The Laplace log-prior is proportional to |w|

Invariant

Step through it

Step through The Laplace log-prior is proportional to |w| one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: w is −2
  2. Step 2: w is −1
  3. Step 3: w is 0
  4. Step 4: w is 1
  5. Step 5: w is 2

44. Spike at zero, fat tails — the best of both

Intuition

The Laplace prior does two things a Gaussian can't. Its spike at 0 says "most weights are probably exactly zero" — that is where sparsity comes from. Its heavier tails say "but the few nonzero ones can be genuinely large."

A Gaussian penalizes large weights harshly (quadratically), so it never lets a real signal grow. Laplace tolerates a handful of big coefficients while zeroing the rest — exactly the shape of a sparse truth, which is why lasso recovers {0,3,7} cleanly.

45. Where does each piece belong: Lesson 12: MAP, Priors & Regularization

Sorting

Sort into buckets

These are the pieces of Lesson 12: MAP, Priors & Regularization, out of order. Put each one back under the part of the lesson it belongs to.

MAP: a prior is a belief
The running example: fit a line, but small data; MLE recap: maximize the likelihood; A prior is what you believe before the data
Gaussian prior → L2 / ridge
The Gaussian noise likelihood (from Lesson 8); The Gaussian log-prior is a negative quadratic; Add the two −logs → the ridge objective
Laplace prior → L1 / lasso
The Laplace prior spikes at zero; The Laplace log-prior is proportional to |w|; Spike at zero, fat tails — the best of both
s1
MAP: a prior is a belief is where Lesson 12: MAP, Priors & Regularization puts The running example: fit a line, but small data, MLE recap: maximize the likelihood, A prior is what you believe before the data. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Gaussian prior → L2 / ridge is where Lesson 12: MAP, Priors & Regularization puts The Gaussian noise likelihood (from Lesson 8), The Gaussian log-prior is a negative quadratic, Add the two −logs → the ridge objective. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Laplace prior → L1 / lasso is where Lesson 12: MAP, Priors & Regularization puts The Laplace prior spikes at zero, The Laplace log-prior is proportional to |w|, Spike at zero, fat tails — the best of both. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

46. The pull that never fades

Intuition

Here is the heart of it. The gradient of the penalty tells you how hard it pushes a weight toward 0. For L2 that push shrinks with the weight; for L1 it is constant — the same shove at w = 0.001 as at w = 1.

So L2 eases off exactly when a weight gets small — it approaches 0 but never arrives. L1 keeps shoving with full force until the weight hits exactly 0 and stops. That constant pull is why lasso produces true zeros.

47. Predict the next row: Watch the two gradients near zero

Pattern

Predict first

The table runs: 1.0 | 2.0000 | 1.0000 · 0.1 | 0.2000 | 1.0000 · 0.01 | 0.0200 | 1.0000

In Watch the two gradients near zero, given the rows so far: what is the next one — the row where w is 0.001?

Correct: 0.001 | 0.0020 | 1.0000

wL2 grad = 2λwL1 grad = λ·sign(w)
1.02.00001.0000
0.10.20001.0000
0.010.02001.0000
0.0010.00201.0000

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The L2 restoring force vanishes near the origin (proportional shrink), so the coefficient asymptotes to a tiny nonzero value.

48. Watch the two gradients near zero

Worked example

Penalty gradients with λ = 1: L2 is d/dw (λw²) = 2λw; L1 is d/dw (λ|w|) = λ·sign(w), constant in magnitude. Print them as w → 0:

import numpy as np
for w in [1.0, 0.1, 0.01, 0.001]:
    l2 = 2*1.0*w            # d/dw (lambda w^2), lambda=1
    l1 = 1.0*np.sign(w)     # d/dw (lambda |w|), lambda=1
    print(f'w={w:<6} L2 grad={l2:<8} L1 grad={l1}')

L2 grad → 0 as w → 0; L1 grad stays at 1

Why: The L2 restoring force vanishes near the origin (proportional shrink), so the coefficient asymptotes to a tiny nonzero value. L1's force is constant, so it can carry the coefficient all the way to 0.

wL2 grad = 2λwL1 grad = λ·sign(w)
1.02.00001.0000
0.10.20001.0000
0.010.02001.0000
0.0010.00201.0000

49. The geometry: diamond vs circle

Concept

Each penalty is a constraint ball around the origin; the fit is the loss contour's first touch. L1's ball is a diamond with corners on the axes; L2's is a round circle.

Figure (svg): Left: a diamond-shaped L1 constraint whose corner on the vertical axis is where an elliptical loss contour first touches, making one weight zero. Right: a round L2 constraint whose tangent point with the loss contour is off the axes, so neither weight is zero.

L1 contours meet the diamond AT a corner (a coordinate is 0). L2 meets the circle off-axis (only shrinkage).

50. The kink at zero and the subgradient

Concept

|w| has a corner at 0 — it is not differentiable there. Instead of a single slope, at w=0 it has a whole set of valid slopes, the subgradient ∂|w| = [−1, 1].

\[ \partial\,|w| = \begin{cases} \{-1\} & w < 0 \\ [-1,\,1] & w = 0 \\ \{+1\} & w > 0 \end{cases} \]

A minimum can sit at w=0 whenever the data gradient there lands inside [−λ, λ] — the flat interval of the penalty absorbs it. A smooth L2 has no such interval (its slope at 0 is a single 0), so it can never pin a coordinate at exactly 0. The kink is the whole mechanism.

51. Restore the missing line: One coordinate, exactly: shrink vs soft-threshold

Fill the middle

Fill in the blanks

From One coordinate, exactly: shrink vs soft-threshold — one line has had its right-hand side removed. Put it back.

import numpy as np
def soft(rho, t):
return np.sign(rho)*max(abs(rho)-t, 0.0)
lam = 0.4
for rho in [1.0, 0.5, 0.15, 0.10, 0.05]:
ridge = rho/(1+lam)
lasso = soft(rho, lam/2) # threshold lam/2 = 0.2
print(f'rho=___ ridge=___ lasso=___')

Why: ridge is what everything below it consumes, so the wrong expression here fails later and somewhere else. Soft-thresholding subtracts λ/2 and clips at 0, so any small coordinate is killed outright.

52. One coordinate, exactly: shrink vs soft-threshold

Worked example

For a single standardized feature (so Σx²=1), the OLS coordinate is some ρ. Ridge and lasso each have a closed form for that one coordinate — and they differ in a decisive way:

\[ w_{\text{ridge}} = \frac{\rho}{1+\lambda}, \qquad w_{\text{lasso}} = \operatorname{sign}(\rho)\,\max\!\big(|\rho| - \tfrac{\lambda}{2},\, 0\big) \]

import numpy as np
def soft(rho, t):
    return np.sign(rho)*max(abs(rho)-t, 0.0)
lam = 0.4
for rho in [1.0, 0.5, 0.15, 0.10, 0.05]:
    ridge = rho/(1+lam)
    lasso = soft(rho, lam/2)   # threshold lam/2 = 0.2
    print(f'rho={rho:<5} ridge={ridge:.4f}  lasso={lasso:.4f}')

Below the threshold |ρ| ≤ λ/2 = 0.2, lasso snaps to EXACTLY 0; ridge never does

Why: Soft-thresholding subtracts λ/2 and clips at 0, so any small coordinate is killed outright. Ridge only divides by (1+λ), which shrinks but preserves the sign and nonzero-ness.

ρ (OLS coord)ridge = ρ/(1+λ)lasso = soft(ρ, 0.2)
1.000.71430.8000
0.500.35710.3000
0.150.10710.0000 ← zero
0.100.07140.0000 ← zero
0.050.03570.0000 ← zero

53. Which is which, by lasso = soft(ρ, 0.2)

Discrimination

Sort into buckets

Sort these by lasso = soft(ρ, 0.2), from memory, without looking back at One coordinate, exactly: shrink vs soft-threshold. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.8000
1.00
0.3000
0.50
0.0000 ← zero
0.15; 0.10; 0.05
g1
lasso = soft(ρ, 0.2) is "0.8000" for 1.00 — that is what the table on "One coordinate, exactly: shrink vs…" records, and it is the single property separating this group from the rest.
g2
lasso = soft(ρ, 0.2) is "0.3000" for 0.50 — that is what the table on "One coordinate, exactly: shrink vs…" records, and it is the single property separating this group from the rest.
g3
lasso = soft(ρ, 0.2) is "0.0000 ← zero" for 0.15, 0.10, 0.05 — that is what the table on "One coordinate, exactly: shrink vs…" records, and it is the single property separating this group from the rest.

54. Something is wrong here: does ridge produce sparsity?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero.

Ridge shrinks every coefficient smoothly; only L1 / lasso sets coordinates to exactly zero.

Why: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).

55. Trap: does ridge produce sparsity?

Trap

The trap

Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.

Expect many exact zeros from L2

Why: L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).

The fix

Ridge shrinks every coefficient smoothly; only L1 / lasso sets coordinates to exactly zero.

Reach for lasso when you want sparsity / feature selection

Why: Verified in Part 5: on 10 features (only 0,3,7 real) ridge keeps 10 nonzero, lasso zeros 7 and recovers exactly the true support {0,3,7}.

56. Break it on purpose: does ridge produce sparsity?

Break the constraint

Discussion prompt

The rule this trap just fixed:

Verified in Part 5: on 10 features (only 0,3,7 real) ridge keeps 10 nonzero, lasso zeros 7 and recovers exactly the true support {0,3,7}.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

L2's gradient 2λw shrinks PROPORTIONALLY and vanishes as w→0, so it never reaches zero. On our 10-feature demo ridge leaves ALL 10 coefficients nonzero (the smallest is still about ±0.003).

57. Conjugate priors

Section

Part 4 of 7 — updates by inspection

58. When the posterior stays in the family

Concept

Multiplying a likelihood by a prior usually gives an ugly posterior you must integrate to normalize. A conjugate prior is chosen so the posterior is the same kind of distribution — the update becomes arithmetic on the parameters.

likelihoodconjugate priorposterior
Bernoulli / BinomialBeta(α, β)Beta(α+k, β+n−k)
Categorical / MultinomialDirichlet(α)Dirichlet(α + counts)
Gaussian (known σ²)GaussianGaussian

The star case for the exam is Beta–Binomial: observe k successes in n trials and the posterior is Beta(α+k, β+n−k). No integral. Let's prove it and use it.

59. What has to happen first: Beta × Binomial = Beta, by matching exponents

Ranking

Put in order

Put the moves of Beta × Binomial = Beta, by matching exponents into the order they have to happen.

  1. Write the two kernels (drop constants)
  2. Add the exponents
  3. Recognize the Beta(α+k, β+n−k) kernel

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Binomial likelihood in the success prob p is pᵏ(1−p)ⁿ⁻ᵏ; Beta(α,β) prior kernel is p^(α−1)(1−p)^(β−1).

60. Beta × Binomial = Beta, by matching exponents

Worked example

Write the two kernels (drop constants)

Why: Binomial likelihood in the success prob p is pᵏ(1−p)ⁿ⁻ᵏ; Beta(α,β) prior kernel is p^(α−1)(1−p)^(β−1).

\[ p(D\mid p)\,p(p) \;\propto\; p^{k}(1-p)^{n-k}\cdot p^{\alpha-1}(1-p)^{\beta-1} \]

Add the exponents

Why: Same base p multiplies by adding exponents: k+(α−1) and (n−k)+(β−1).

\[ = p^{\,\alpha + k - 1}\,(1-p)^{\,\beta + (n-k) - 1} \]

Recognize the Beta(α+k, β+n−k) kernel

Why: That is exactly the kernel of a Beta with parameters α+k and β+n−k. The posterior is Beta again — conjugacy proved: add successes to α, failures to β.

\[ p(p\mid D) = \mathrm{Beta}(\alpha + k,\; \beta + n - k) \]

61. Decode the notation: Beta × Binomial = Beta, by matching exponents

Notation

Annotate

From Beta × Binomial = Beta, by matching exponents — read this one piece at a time. What is each part doing?

On: \( p(p\mid D) = \mathrm{Beta}(\alpha + k,\; \beta + n - k) \)

  • Binomial likelihood in the success prob p is pᵏ(1−p)ⁿ⁻ᵏ; Beta(α,β) prior kernel is p^(α−1)(1−p)^(β−1).
  • Same base p multiplies by adding exponents: k+(α−1) and (n−k)+(β−1).
  • That is exactly the kernel of a Beta with parameters α+k and β+n−k. The posterior is Beta again — conjugacy proved: add successes to α, failures to β.

62. Guess the shape of the answer: Verify conjugacy numerically

Estimation

Predict first

Prior Beta(2,2), data k=7 successes in n=10. The claim: posterior is Beta(9,5), i.e. likelihood × prior is proportional to the Beta(9,5) kernel with the SAME ratio at every p. Check it:

Commit before you compute: what does Verify conjugacy numerically come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: The ratio is a CONSTANT (1.0) across all p → same distribution

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. likelihood×prior and the Beta(9,5) kernel differ only by a normalizing constant, so their ratio is identical everywhere.

63. Verify conjugacy numerically

Worked example

Prior Beta(2,2), data k=7 successes in n=10. The claim: posterior is Beta(9,5), i.e. likelihood × prior is proportional to the Beta(9,5) kernel with the SAME ratio at every p. Check it:

import numpy as np
a0, b0, k, n = 2, 2, 7, 10
def unnorm_post(p):  return p**k*(1-p)**(n-k) * p**(a0-1)*(1-p)**(b0-1)
def beta_kernel(p):  return p**(a0+k-1)*(1-p)**(b0+(n-k)-1)  # Beta(9,5)
for p in [0.3, 0.5, 0.7, 0.9]:
    r = unnorm_post(p)/beta_kernel(p)
    print(f'p={p}: ratio={r:.6f}')

The ratio is a CONSTANT (1.0) across all p → same distribution

Why: likelihood×prior and the Beta(9,5) kernel differ only by a normalizing constant, so their ratio is identical everywhere. A constant ratio is precisely what 'same distribution up to normalization' means.

plikelihood×priorBeta(9,5) kernelratio (verified)
0.31.575e−051.575e−051.000000
0.52.441e−042.441e−041.000000
0.74.669e−044.669e−041.000000
0.94.305e−054.305e−051.000000

64. What happens as it grows: Verify conjugacy numerically

Scale up

Step through it

Step through Verify conjugacy numerically and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: p is 0.3
  2. Step 2: p is 0.5
  3. Step 3: p is 0.7
  4. Step 4: p is 0.9

65. The posterior blends prior and data

Concept

With Beta(9,5) the posterior mean is 9/14 = 0.6429 — sitting between the prior mean 0.5 and the data's MLE k/n = 0.7. The prior Beta(2,2) acts like 2 pseudo-successes and 2 pseudo-failures added to the real counts.

quantityformulavalue (verified)
prior meanα/(α+β) = 2/40.5000
data MLEk/n = 7/100.7000
posterior mean9/140.6429
posterior mode (MAP)(9−1)/(14−2)0.6667

The MAP point estimate is the posterior mode (α−1)/(α+β−2) = 8/12 = 0.6667. More data (bigger n) drags the posterior toward the MLE; a stronger prior (bigger α+β) holds it near 0.5.

66. Restore the missing line: Dirichlet–Categorical: the same trick, K outcomes

Fill the middle

Fill in the blanks

From Dirichlet–Categorical: the same trick, K outcomes — one line has had its right-hand side removed. Put it back.

import numpy as np
alpha0 = np.array([1., 1., 1.]) # flat Dirichlet prior
counts = np.array([12., 5., 3.]) # 20 rolls
post = alpha0 + counts # Dirichlet(13, 6, 4)
print('posterior alpha:', post)
print('posterior mean :', (post/post.sum()).round(4))
print('MLE :', (counts/counts.sum()).round(4))

Why: counts is what everything below it consumes, so the wrong expression here fails later and somewhere else. Just as Beta added successes/failures, Dirichlet adds each outcome's count to its α.

67. Dirichlet–Categorical: the same trick, K outcomes

Worked example

The multi-outcome version generalizes Beta to Dirichlet. Roll a 3-sided die 20 times, seeing counts [12, 5, 3], under a flat prior Dirichlet(1,1,1). The update is just add the counts:

import numpy as np
alpha0 = np.array([1., 1., 1.])   # flat Dirichlet prior
counts = np.array([12., 5., 3.])  # 20 rolls
post = alpha0 + counts            # Dirichlet(13, 6, 4)
print('posterior alpha:', post)
print('posterior mean :', (post/post.sum()).round(4))
print('MLE            :', (counts/counts.sum()).round(4))

Posterior = Dirichlet(α + counts) = Dirichlet(13, 6, 4)

Why: Just as Beta added successes/failures, Dirichlet adds each outcome's count to its α. The flat prior acts like one pseudo-count per face, nudging the estimate off the raw frequencies.

facecountMLE = count/nposterior mean
1120.60000.5652
250.25000.2609
330.15000.1739

68. Watch it run: Dirichlet–Categorical: the same trick, K outcomes

Pattern

Step through it

Step through Dirichlet–Categorical: the same trick, K outcomes one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: face is 1
  2. Step 2: face is 2
  3. Step 3: face is 3

69. Dropout as approximate Bayes

Section

Part 5 of 7 — same idea, deep nets

70. Dropout: randomly delete neurons

Concept

Dropout multiplies each activation by an independent Bernoulli(1−p) mask during training — a random subset of units is zeroed on every forward pass, then rescaled so the expected activation is unchanged.

\[ \tilde h = \frac{m \odot h}{1-p}, \qquad m_j \sim \mathrm{Bernoulli}(1-p) \]

Empirically it fights overfitting hard. The surprise is why it works — and the answer ties straight back to priors and posteriors.

71. Teach it back: Dropout: randomly delete neurons

Explain it

Discussion prompt

Explain Dropout: randomly delete neurons to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Empirically it fights overfitting hard. The surprise is why it works — and the answer ties straight back to priors and posteriors.

72. Each mask is a sampled sub-network

Intuition

Every dropout mask defines a different thinned network sharing the same weights. Over a training run you visit exponentially many of them, and prediction with the full net (weights scaled) approximates averaging their outputs — a giant implicit ensemble.

Gal & Ghahramani showed this is variational Bayes: dropout training approximately minimizes the KL divergence to a posterior over the weights. The random masks are samples from an approximate posterior — regularization and Bayesian inference, once again the same object seen from two sides.

73. How big is that ensemble?

Concept

A layer of n units with independent on/off masks has 2ⁿ possible thinned sub-networks. Dropout trains a sample of them, all sharing weights — an astronomically large implicit ensemble at the cost of one network.

units ndistinct sub-networks 2ⁿroughly
102¹⁰1,024
1002¹⁰⁰1.27 × 10³⁰
10002¹⁰⁰⁰> atoms in the universe

You can never enumerate them, but you don't have to: averaging with the weights scaled by (1−p) approximates the ensemble prediction — the same integrate-out-the-uncertainty move that Bayes makes, done cheaply.

74. Fill in: distinct sub-networks 2ⁿ for How big is that ensemble?

Comparison

Comparison matrix

From How big is that ensemble?: refill the distinct sub-networks 2ⁿ column from what you know. The rest of the table is as it appeared.

units ndistinct sub-networks 2ⁿroughly
102¹⁰1,024
1002¹⁰⁰1.27 × 10³⁰
10002¹⁰⁰⁰> atoms in the universe

75. MC dropout gives uncertainty

Concept

Because dropout ≈ sampling the posterior, you can keep it on at test time and run the network many times: the spread of the predictions estimates model uncertainty. This is Monte-Carlo (MC) dropout.

So the same log-prior story scales all the way up: a penalty on weights (Parts 2–3), a conjugate posterior you can write down (Part 4), and here an approximate posterior you sample — three faces of belief plus data.

76. Ridge vs Lasso — the real fit

Section

Part 6 of 7 — verified vs sklearn

77. Restore the missing line: Sparse-truth data: only 3 of 10 features matter

Fill the middle

Fill in the blanks

From Sparse-truth data: only 3 of 10 features matter — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print('true nonzero idx:', np.flatnonzero(true))
print('X shape:', X.shape)

Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. The true weight vector is zero everywhere except three slots.

78. Sparse-truth data: only 3 of 10 features matter

Worked example

Build 50 samples of 10 features where the target truly depends on only features 0, 3, 7 (weights 3, −2, 1.5), plus small noise. Fixed RNG seed so you see these exact numbers:

import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print('true nonzero idx:', np.flatnonzero(true))
print('X shape:', X.shape)

Only indices 0, 3, 7 carry signal; the other 7 are pure noise

Why: The true weight vector is zero everywhere except three slots. A good sparse learner should discover this support and zero the rest.

objectvalue (verified)
true nonzero indices[0, 3, 7]
true weights there[3.0, −2.0, 1.5]
X shape(50, 10)

79. What has to be given first: Fit ridge and lasso, count the zeros

Missing information

Discussion prompt

Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Lasso recovers the true sparse support; ridge keeps all 10 coefficients small-but-nonzero (its largest 'noise' coefficient is still ≈0.065).

80. Fit ridge and lasso, count the zeros

Worked example

Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:

from sklearn.linear_model import Ridge, Lasso
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0)))
print('lasso keeps:', np.flatnonzero(lasso.coef_))

ridge: 0 exact zeros; lasso: 7 exact zeros, keeps exactly {0, 3, 7}

Why: Lasso recovers the true sparse support; ridge keeps all 10 coefficients small-but-nonzero (its largest 'noise' coefficient is still ≈0.065).

featuretrueridge coeflasso coef
03.002.93122.9291
3−2.00−1.9570−1.9212
71.501.49441.4284
other 70up to ±0.065 (nonzero)0.0000 (exact)

81. Work backwards from the answer: Fit ridge and lasso, count the zeros

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

ridge: 0 exact zeros; lasso: 7 exact zeros, keeps exactly {0, 3, 7}

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fit Ridge(alpha=1.0) and Lasso(alpha=0.1) on the same data and count how many coefficients each drives to exactly zero:

82. Guess the shape of the answer: The lasso path: λ is the sparsity dial

Estimation

Predict first

λ (sklearn's alpha) is the prior strength — crank it up and more coordinates snap to zero. Sweep it and count the survivors:

Commit before you compute: what does The lasso path: λ is the sparsity dial come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Too small (0.01) keeps noise features; the sweet spot 0.1–1.0 recovers exactly {0,3,7}; too big (3.0) kills everything

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A weak prior barely thresholds, so noise coordinates survive; a strong prior over-shrinks and zeros even the true features.

83. The lasso path: λ is the sparsity dial

Worked example

λ (sklearn's alpha) is the prior strength — crank it up and more coordinates snap to zero. Sweep it and count the survivors:

from sklearn.linear_model import Lasso
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
for a in [0.01, 0.1, 0.5, 1.0, 3.0]:
    L = Lasso(alpha=a).fit(X, y)
    print(f'alpha={a:<5} nonzero={int(np.sum(L.coef_ != 0))}',
          'keeps', np.flatnonzero(L.coef_))

Too small (0.01) keeps noise features; the sweet spot 0.1–1.0 recovers exactly {0,3,7}; too big (3.0) kills everything

Why: A weak prior barely thresholds, so noise coordinates survive; a strong prior over-shrinks and zeros even the true features. You tune λ on validation data to land in the middle band.

alpha (λ)# nonzerofeatures kept (verified)
0.015{0, 2, 3, 7, 8}
0.13{0, 3, 7}
0.53{0, 3, 7}
1.03{0, 3, 7}
3.00{ } — all zero

84. Which is which, by # nonzero

Discrimination

Sort into buckets

Sort these by # nonzero, from memory, without looking back at The lasso path: λ is the sparsity dial. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

5
0.01
3
0.1; 0.5; 1.0
0
3.0
g1
# nonzero is "5" for 0.01 — that is what the table on "The lasso path: λ is the sparsity dial" records, and it is the single property separating this group from the rest.
g2
# nonzero is "3" for 0.1, 0.5, 1.0 — that is what the table on "The lasso path: λ is the sparsity dial" records, and it is the single property separating this group from the rest.
g3
# nonzero is "0" for 3.0 — that is what the table on "The lasso path: λ is the sparsity dial" records, and it is the single property separating this group from the rest.

85. Something is wrong here: are L1 and L2 interchangeable?

Anomaly

Predict first

A student writes this, and it looks reasonable:

They are both just "regularization", so pick whichever — they do the same job, and I need to drop useless features, so ridge is fine.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients.

Match the penalty to the goal: L2 for stability/shrinkage, L1 for sparsity/feature selection.

Why: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients. If the goal is a sparse, interpretable model, L2 is the wrong tool: you get a dense weight vector, not a shortlist.

86. Trap: are L1 and L2 interchangeable?

Trap

The trap

They are both just "regularization", so pick whichever — they do the same job, and I need to drop useless features, so ridge is fine.

Use ridge for feature selection

Why: Ridge never eliminates a feature — on our data it returns all 10 nonzero coefficients. If the goal is a sparse, interpretable model, L2 is the wrong tool: you get a dense weight vector, not a shortlist.

The fix

Match the penalty to the goal: L2 for stability/shrinkage, L1 for sparsity/feature selection.

Lasso (L1) to select features; ridge (L2) to tame collinearity

Why: Different priors (Laplace vs Gaussian), different geometry (diamond vs circle), different solutions (7 zeros vs 0). Elastic Net blends both when you want some of each.

87. Which of these survive contact with Lesson 12: MAP, Priors & Regularization?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
In Lesson 8 the maximum-likelihood estimate picked the θ that makes the observed data most probable. For Gaussian noise on a linear model, that is exactly least squares.; Bayes' rule flips a conditional. Applied to a parameter θ given data D, the posterior is the likelihood times the prior, divided by a normalizing constant:; The boxed result is the spine of the whole lesson. MAP is MLE with one extra additive term — the log of the prior:
Breaks
A tighter prior (small τ²) should mean less penalty, so put the prior variance on top: λ = τ²/σ².; Ridge penalizes weight size, so it should zero out the useless features — just like lasso, only smoother.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 12: MAP, Priors & Regularization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

88. Predict the next row: Verify: the closed form is really MAP

Pattern

Predict first

The table runs: sklearn Ridge | 2.844950 | −0.030003 | 0.012108 · closed form | 2.844950 | −0.030003 | 0.012108

In Verify: the closed form is really MAP, given the rows so far: what is the next one — the row where source is allclose?

Correct: allclose | True | — | —

sourcecoef[0]coef[1]coef[2]
sklearn Ridge2.844950−0.0300030.012108
closed form2.844950−0.0300030.012108
allcloseTrue——

Why: The relationship between the columns, not the individual numbers, is what generates the next row. sklearn's alpha is our λ. Solving (XᵀX+λI)w = Xᵀy reproduces it exactly, closing the loop: the Gaussian log-prior and the L2 penalty are the same math.

89. Verify: the closed form is really MAP

Concept

One last sanity check that ridge = Gaussian-prior MAP: sklearn's Ridge(alpha=2, fit_intercept=False) must equal solving (XᵀX + 2I)w = Xᵀy by hand.

from sklearn.linear_model import Ridge
import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
Rr = Ridge(alpha=2.0, fit_intercept=False).fit(X, y)
w_cf = np.linalg.solve(X.T @ X + 2.0*np.eye(10), X.T @ y)
print('sklearn [:3]:', Rr.coef_[:3].round(6))
print('closed   [:3]:', w_cf[:3].round(6))
print('match:', np.allclose(Rr.coef_, w_cf))

Identical to 6 decimals → the penalty IS the log-prior

Why: sklearn's alpha is our λ. Solving (XᵀX+λI)w = Xᵀy reproduces it exactly, closing the loop: the Gaussian log-prior and the L2 penalty are the same math.

sourcecoef[0]coef[1]coef[2]
sklearn Ridge2.844950−0.0300030.012108
closed form2.844950−0.0300030.012108
allcloseTrue——

90. By analogy: Verify: the closed form is really MAP

Analogy

Discussion prompt

Explain Verify: the closed form is really MAP by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

One last sanity check that ridge = Gaussian-prior MAP: sklearn's Ridge(alpha=2, fit_intercept=False) must equal solving (XᵀX + 2I)w = Xᵀy by hand.

91. When does regularization actually help?

Concept

Honest caveat: here n=50 ≥ d=10 with tiny noise, so plain OLS barely overfits and has the lowest test error. Regularization is not free — it trades a little bias for variance reduction, and pays off most when data is scarce or features collinear.

modeltest MSE (1000 held-out, verified)‖coef − true‖
OLS0.01180.0498
ridge (α=1)0.02170.1077
lasso (α=0.1)0.02700.1280

So why reach for lasso here? Not for accuracy — for interpretability. It is the only fit that hands you the exact 3-feature story {0,3,7} instead of 10 muddy coefficients. Sparsity is a modeling goal, distinct from raw predictive error.

92. Watch it run: When does regularization actually help?

Pattern

Step through it

Step through When does regularization actually help? one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: model is OLS
  2. Step 2: model is ridge (α=1)
  3. Step 3: model is lasso (α=0.1)

93. Rebuild the recipe: Regularization-as-prior recipe

Ranking

Put in order

These are the steps of Regularization-as-prior recipe, scrambled. Put them back in order before the next slide shows you.

  1. Likelihood → loss: write −log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)
  2. Prior → penalty: add −log p(θ) — the belief, in log form
  3. Gaussian prior ⇒ λ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zeros
  4. Laplace prior ⇒ λ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zeros
  5. Conjugate prior (Beta, Dirichlet) ⇒ posterior in the same family, update by adding counts
  6. Tune λ = prior strength: large λ = confident prior = more shrinkage; λ → 0 = back to MLE

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

94. Regularization-as-prior recipe

Pattern

  1. Likelihood → loss: write −log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)
  2. Prior → penalty: add −log p(θ) — the belief, in log form
  3. Gaussian prior ⇒ λ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zeros
  4. Laplace prior ⇒ λ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zeros
  5. Conjugate prior (Beta, Dirichlet) ⇒ posterior in the same family, update by adding counts
  6. Tune λ = prior strength: large λ = confident prior = more shrinkage; λ → 0 = back to MLE

95. Where does it stop working: Regularization-as-prior recipe

Edge cases

Discussion prompt

Regularization-as-prior recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Likelihood → loss: write −log p(data|θ) (MSE for Gaussian noise, cross-entropy for classification)
  2. Prior → penalty: add −log p(θ) — the belief, in log form
  3. Gaussian prior ⇒ λ‖w‖² (ridge / weight decay), with λ = σ²/τ² — shrinks, never zeros
  4. Laplace prior ⇒ λ‖w‖₁ (lasso) — constant pull ⇒ soft-threshold ⇒ exact zeros
  5. Conjugate prior (Beta, Dirichlet) ⇒ posterior in the same family, update by adding counts
  6. Tune λ = prior strength: large λ = confident prior = more shrinkage; λ → 0 = back to MLE

96. Rule out three: Check yourself — which prior is ridge?

Elimination

Eliminate the wrong options

L2 regularization (ridge) is the MAP estimate under which prior on the weights?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. A zero-mean Gaussian prior
  • B. A Laplace prior
  • C. A uniform (flat) prior
  • D. A Beta prior

Survives elimination: A

Why: −log of a zero-mean Gaussian is proportional to ‖w‖², so a Gaussian prior adds exactly the L2 penalty. Minimizing loss + λ‖w‖² is ridge, and its strength is λ = σ²/τ².

97. Check yourself — which prior is ridge?

Check

Match the penalty to its prior distribution.

Check your understanding

L2 regularization (ridge) is the MAP estimate under which prior on the weights?

  • A. A zero-mean Gaussian prior (correct)
  • B. A Laplace prior
  • C. A uniform (flat) prior
  • D. A Beta prior

Answer: A

Why: −log of a zero-mean Gaussian is proportional to ‖w‖², so a Gaussian prior adds exactly the L2 penalty. Minimizing loss + λ‖w‖² is ridge, and its strength is λ = σ²/τ².

Why B tempts people
A Laplace prior gives the L1 penalty (lasso), whose −log is proportional to |w| — the sparse one, not ridge.
Why C tempts people
A uniform prior adds a constant (its log doesn't depend on w), so it contributes NO penalty — MAP collapses back to plain MLE / OLS.
Why D tempts people
A Beta prior lives on a probability in [0,1] (Bernoulli parameter), not on real-valued regression weights that can be negative or large.

98. Answer it before you see the options: Check yourself — why L1 is sparse

Prediction

Predict first

Why does L1 (lasso) set coefficients to EXACTLY zero while L2 (ridge) does not?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The L1 penalty's gradient magnitude stays constant (λ) as w → 0, so it pushes a small weight all the way to 0; L2's gradient 2λw fades to 0 and never finishes the job

Why: The subgradient of λ|w| has constant magnitude λ (its sign), so the pull toward 0 never weakens — the soft-threshold snaps small coordinates to exactly 0. L2's gradient 2λw shrinks with w and vanishes at the origin, so ridge only asymptotes toward 0.

99. Check yourself — why L1 is sparse

Check

Think about the gradient of the penalty as w → 0.

Check your understanding

Why does L1 (lasso) set coefficients to EXACTLY zero while L2 (ridge) does not?

  • A. The L1 penalty's gradient magnitude stays constant (λ) as w → 0, so it pushes a small weight all the way to 0; L2's gradient 2λw fades to 0 and never finishes the job (correct)
  • B. L1 uses a larger λ than L2 by default
  • C. L1 is non-convex, so it has multiple minima including zero
  • D. L2 rounds tiny coefficients up, keeping them nonzero

Answer: A

Why: The subgradient of λ|w| has constant magnitude λ (its sign), so the pull toward 0 never weakens — the soft-threshold snaps small coordinates to exactly 0. L2's gradient 2λw shrinks with w and vanishes at the origin, so ridge only asymptotes toward 0.

Why B tempts people
λ is a free hyperparameter for both; sparsity is about the SHAPE of the penalty near 0 (constant vs proportional gradient), not the size of λ.
Why C tempts people
The L1 penalty is convex (it's |w|), just non-differentiable at 0. Sparsity comes from that kink, not from non-convexity or multiple minima.
Why D tempts people
There is no rounding step. L2 keeps coefficients nonzero because its restoring force fades near 0, not because of any rounding.

100. Answer it before you see the options: Check yourself — Beta–Binomial update

Prediction

Predict first

With a Beta(α, β) prior and k successes in n Bernoulli trials, the posterior is…

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Beta(α + k, β + n − k)

Why: Multiplying pᵏ(1−p)ⁿ⁻ᵏ by the Beta kernel p^(α−1)(1−p)^(β−1) adds exponents: α gets the successes k, β gets the failures n−k. The posterior is Beta(α+k, β+n−k) — that's conjugacy.

101. Check yourself — Beta–Binomial update

Check

Recall the conjugate update: successes go to α, failures to β.

Check your understanding

With a Beta(α, β) prior and k successes in n Bernoulli trials, the posterior is…

  • A. Beta(α + k, β + n − k) (correct)
  • B. Beta(α + n, β + k)
  • C. Gaussian(k/n, 1)
  • D. Dirichlet(α + k)

Answer: A

Why: Multiplying pᵏ(1−p)ⁿ⁻ᵏ by the Beta kernel p^(α−1)(1−p)^(β−1) adds exponents: α gets the successes k, β gets the failures n−k. The posterior is Beta(α+k, β+n−k) — that's conjugacy.

Why B tempts people
Swaps the updates — α must receive the successes k and β the failures (n−k), not n and k. Beta(α+n, β+k) mislabels which count goes where.
Why C tempts people
The Beta–Bernoulli posterior stays Beta; conjugacy PRESERVES the family, it doesn't switch to a Gaussian.
Why D tempts people
Dirichlet is the conjugate prior for the CATEGORICAL/multinomial likelihood (many outcomes), not for a single Bernoulli probability.

102. Rule out three: Check yourself — MAP vs MLE

Elimination

Eliminate the wrong options

In the ridge MAP objective ‖Xw−y‖² + λ‖w‖² with λ = σ²/τ², what happens as the prior variance τ² → ∞?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. λ → 0, so MAP reduces to plain MLE / OLS
  • B. λ → ∞, so all weights are forced to 0
  • C. The prior becomes Laplace, giving lasso
  • D. The posterior becomes exactly the prior

Survives elimination: A

Why: λ = σ²/τ² → 0 as τ² → ∞. A prior with infinite variance is flat (uninformative), so its log-term is constant and MAP becomes MLE — plain least squares, w = [2.2, 0.6] on our data.

103. Check yourself — MAP vs MLE

Check

What happens to MAP as the prior gets flat?

Check your understanding

In the ridge MAP objective ‖Xw−y‖² + λ‖w‖² with λ = σ²/τ², what happens as the prior variance τ² → ∞?

  • A. λ → 0, so MAP reduces to plain MLE / OLS (correct)
  • B. λ → ∞, so all weights are forced to 0
  • C. The prior becomes Laplace, giving lasso
  • D. The posterior becomes exactly the prior

Answer: A

Why: λ = σ²/τ² → 0 as τ² → ∞. A prior with infinite variance is flat (uninformative), so its log-term is constant and MAP becomes MLE — plain least squares, w = [2.2, 0.6] on our data.

Why B tempts people
That is the OPPOSITE limit: λ → ∞ happens as τ² → 0 (a tight prior at 0), which crushes the weights. A wide prior barely penalizes.
Why C tempts people
Changing τ² only rescales the Gaussian; it never turns it into a Laplace. The penalty stays L2 for any Gaussian prior.
Why D tempts people
With τ² → ∞ the data dominates, so the posterior concentrates at the MLE — the prior's influence VANISHES rather than taking over.

104. Your turn: see the sparsity

Section

Part 7 of 7 — the project

105. Project: a prior becomes a penalty

Concept

Build data with a known sparse support, fit ridge and lasso, and confirm that only lasso recovers the zeros — watching a Laplace prior turn into feature selection. Every piece is derived; now assemble it.

#requirementtool
1Data where only features 0, 3, 7 matternumpy default_rng + a sparse true w
2Fit Ridge(alpha=1) and Lasso(alpha=0.1)sklearn.linear_model
3Count exact zeros and compare to true supportnp.sum(coef == 0)

Build rules: type every line, set the RNG seed default_rng(1) so results are reproducible, run after each block, and compare coefficients to the true support [0, 3, 7].

106. Break it if you can: Project: a prior becomes a penalty

Counterexample

Discussion prompt

Build rules: type every line, set the RNG seed default_rng(1) so results are reproducible, run after each block, and compare coefficients to the true support [0, 3, 7].

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

107. Milestone 1 — sparse data

Worked example

Your turn: make a true weight vector with nonzeros only at 0, 3, 7, then generate y = X·w_true + small noise. Predict how many of the 10 true weights are zero before you print.

Hint: true = np.zeros(10); true[[0,3,7]] = [3, -2, 1.5], then y = X @ true + 0.1*rng.normal(size=50). Seed with np.random.default_rng(1).

import numpy as np
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
print(np.flatnonzero(true))
print(X.shape)
checkvalue (verified)
true nonzero indices[0, 3, 7]
number of true zeros7 of 10
X shape(50, 10)

108. Milestone 2 — fit both models

Worked example

Your turn: fit ridge and lasso on the same X, y. Predict which one's coefficient vector contains exact zeros before printing.

Hint: Ridge(alpha=1.0).fit(X, y) and Lasso(alpha=0.1).fit(X, y); read .coef_ and .round(3).

import numpy as np
from sklearn.linear_model import Ridge, Lasso
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge:', ridge.coef_.round(3))
print('lasso:', lasso.coef_.round(3))
modelfeatures 0 / 3 / 7the other 7
ridge2.931 / −1.957 / 1.494small but nonzero
lasso2.929 / −1.921 / 1.428exactly 0

109. Milestone 3 — count the zeros

Worked example

Your turn: count exact zeros in each model and compare to the true support. How many should lasso zero out?

Hint: ridge coefficients are never truly 0, so test np.abs(ridge.coef_) < 1e-6; lasso really hits 0, so lasso.coef_ == 0 works.

import numpy as np
from sklearn.linear_model import Ridge, Lasso
rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)
ridge = Ridge(alpha=1.0).fit(X, y)
lasso = Lasso(alpha=0.1).fit(X, y)
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0)), 'of 10')
print('lasso keeps:', np.flatnonzero(lasso.coef_))
modelexact zeros (verified)support recovered?
ridge0 of 10no — keeps all 10
lasso7 of 10yes — keeps {0, 3, 7}

110. What each one costs: Milestone 3 — count the zeros

Trade off

Comparison matrix

From Milestone 3 — count the zeros: every row here is a choice with a cost. Fill the exact zeros (verified) column, then say which row you would actually pick and what you give up for it.

modelexact zeros (verified)support recovered?
ridge0 of 10no — keeps all 10
lasso7 of 10yes — keeps {0, 3, 7}

111. The full program

Concept

import numpy as np
from sklearn.linear_model import Ridge, Lasso

rng = np.random.default_rng(1)
X = rng.normal(size=(50, 10))
true = np.zeros(10); true[[0, 3, 7]] = [3., -2., 1.5]
y = X @ true + 0.1*rng.normal(size=50)   # 1. sparse-truth data

ridge = Ridge(alpha=1.0).fit(X, y)        # 2. Gaussian-prior MAP
lasso = Lasso(alpha=0.1).fit(X, y)        #    Laplace-prior MAP
print('ridge zeros:', int(np.sum(np.abs(ridge.coef_) < 1e-6)))
print('lasso zeros:', int(np.sum(lasso.coef_ == 0)))  # 3. count -> 7
print('lasso keeps:', np.flatnonzero(lasso.coef_))
printed linevalue (verified)
ridge zeros:0
lasso zeros:7
lasso keeps:[0 3 7]

If lasso reports 7 zeros and keeps exactly features 0, 3, 7 while ridge keeps all 10 — you have watched a prior become a penalty, and a Laplace prior perform feature selection.

112. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
ridge zeros:0
lasso zeros:7
lasso keeps:[0 3 7]

113. Show it off

Concept

Slides closed, out loud: explain (1) how log p(θ) becomes an additive penalty and why λ = σ²/τ², (2) why the L1 diamond produces zeros while the L2 circle only shrinks, and (3) what conjugacy saves you from computing.

Stretch: derive the one-coordinate soft-threshold formula soft(ρ, λ/2) from ∂/∂w [ (w−ρ)² + λ|w| ] = 0, then confirm your lasso coefficients on features 0/3/7 match the sklearn values above. This thread continues into Elastic Net (Week 21), dropout-as-Bayes (Week 18), and Adam's decoupled weight decay (Week 35).

114. Connect it up: Lesson 12: MAP, Priors & Regularization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — MAP: a prior is a belief · Gaussian prior → L2 / ridge · Laplace prior → L1 / lasso · Conjugate priors · Dropout as approximate Bayes · Ridge vs Lasso — the real fit. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

115. What you can do now

Recap

ideathe one thing to remember
MAPMLE + log-prior — a penalty added to the loss
ridgeGaussian prior → L2 → shrink, never zero; λ = σ²/τ²
lassoLaplace prior → L1 → constant pull → exact zeros
conjugacyposterior stays in the family; add counts (Beta, Dirichlet)
dropoutrandom masks = samples from an approximate posterior

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 12 (Week 4 — MAP & Regularization) — Barron · USAAIO Round 2 Preparation, 2026
  2. Christopher Bishop, Pattern Recognition and Machine Learning, Ch. 3 (Linear Models) & Ch. 2 (Conjugate priors) — Springer, 2006
  3. Kevin Murphy, Probabilistic Machine Learning: An Introduction, Ch. 11 (Linear regression: ridge, lasso, MAP) — MIT Press, 2022
  4. scikit-learn Ridge / Lasso user guide
  5. Gal & Ghahramani, Dropout as a Bayesian Approximation
  6. Every coefficient, zero-count, and posterior parameter produced by real execution — numpy 2.2.6 + scipy 1.16 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108