USAAIO Lesson 8, from Week 3 on probability, fully worked. It states maximum likelihood estimation as an optimization principle and scores it on a real coin, then justifies the log trick with a live underflow demo. It derives the Gaussian mean AND variance one calculus step at a time on a fixed 8-point dataset, including the second-derivative concavity check, and covers the Bernoulli, Poisson, and Exponential MLEs. It then delivers the big reveal that MSE and cross-entropy ARE negative log-likelihoods, checking both against torch, and closes with MAP as MLE plus a log-prior, from which ridge regression falls out of a Gaussian prior. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 62 slides.
Subject: Machine Learning · 119 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 8 · Week 3 (Probability)
The single principle behind every loss you'll ever minimize. We score parameters on a real coin, derive the Gaussian and Bernoulli MLEs one step at a time, then prove that MSE and cross-entropy ARE negative log-likelihoods — and check both against torch.
Objectives
log turns a product into a sum without moving the argmax1/n variance) and a Bernoulli (→ success fraction), skipping no calculus steptorchmaximum_likelihood_gaussian from scratchWarm-up
Discussion prompt
Before we open Lesson 8: Maximum Likelihood Estimation: without looking back, what was the main idea of Linear Systems & the Normal Equations, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
overdetermined systems written out equation by equation, the normal equations XᵀXw = Xᵀy derived twice (projection AND calculus) with no skipped steps, the 5-student 2×2 solved by hand, invertibility and rank, ridge derived from the penalized loss, R², and a from-scratch OLS verified against sklearn.
Section
Part 1 of 6 — what MLE asks
Intuition
You flip a coin 10 times and see 7 heads. Nobody told you the coin's bias p. But your gut already leans: this looks like a coin that favors heads.
A coin with p = 0.7 would produce 7-of-10 fairly often. A coin with p = 0.1 almost never would. So p = 0.7 explains what you saw far better than p = 0.1.
MLE is just that instinct made exact: among all candidate parameters, pick the one under which your data is most probable. The rest of Part 1 turns 'most probable' into calculus.
Counterexample
Discussion prompt
You flip a coin 10 times and see 7 heads. Nobody told you the coin's bias p. But your gut already leans: this looks like a coin that favors heads.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
A coin with p = 0.7 would produce 7-of-10 fairly often. A coin with p = 0.1 almost never would. So p = 0.7 explains what you saw far better than p = 0.1.
Concept
You observe i.i.d. data x₁,…,xₙ drawn from a family p(x; θ). Because the draws are independent, the probability of the whole dataset is the product of the per-point probabilities:
\[ \mathcal{L}(\theta) = p(x_1,\dots,x_n;\theta) = \prod_{i=1}^{n} p(x_i;\theta) \]
likelihood — The SAME formula as a probability, but read as a function of the parameter θ with the data held fixed. We are not scoring which data is likely — we are scoring which θ best explains the data we already have.
Analogy
Discussion prompt
Explain The likelihood function by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
You observe i.i.d. data x₁,…,xₙ drawn from a family p(x; θ). Because the draws are independent, the probability of the whole dataset is the product of the per-point probabilities:
Concept
A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.
So L(θ) need not integrate to 1 over θ, and 'the most likely θ' means 'the θ under which the data is most probable' — not 'the θ most probably true'. Turning θ into something with its own probability is what MAP does, in Part 5.
Explain it
Discussion prompt
Explain Likelihood is not probability of θ to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.
Concept
Maximum likelihood estimation picks the parameter that maximizes that function — the θ that makes exactly this dataset as probable as possible:
\[ \theta_{\text{MLE}} = \arg\max_{\theta}\; \prod_{i=1}^{n} p(x_i; \theta) \]
arg max returns the location of the maximum (the best θ), not its height. Everything today is a hunt for that location — usually by setting a derivative to zero.
Estimation
Predict first
Let's actually evaluate the likelihood for the coin. With k = 7 heads in n = 10, L(p) = C(10,7)·pᵏ(1−p)ⁿ⁻ᵏ. Compute it for three candidate p and read off the winner:
Commit before you compute: what does Score three coins on the 7-heads data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Peak likelihood is at p = 0.7 — the empirical head-fraction
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 0.7 assigns probability 0.266828 to the observed 7-of-10, beating 0.5 (0.117188) and 0.9 (0.057396).
Worked example
Let's actually evaluate the likelihood for the coin. With k = 7 heads in n = 10, L(p) = C(10,7)·pᵏ(1−p)ⁿ⁻ᵏ. Compute it for three candidate p and read off the winner:
from math import comb
k, n = 7, 10
for p in [0.1, 0.5, 0.7, 0.9]:
L = comb(n, k) * p**k * (1-p)**(n-k)
print(p, round(L, 6))Peak likelihood is at p = 0.7 — the empirical head-fraction
Why: 0.7 assigns probability 0.266828 to the observed 7-of-10, beating 0.5 (0.117188) and 0.9 (0.057396). The best explanation is exactly the observed frequency k/n.
| candidate p | L(p) = C(10,7)·p⁷(1−p)³ | verdict |
|---|---|---|
| 0.1 | 0.000009 | terrible |
| 0.5 | 0.117188 | so-so |
| 0.7 | 0.266828 | best (= k/n) |
| 0.9 | 0.057396 | over-shoots |
Comparison
Comparison matrix
From Score three coins on the 7-heads data: refill the L(p) = C(10,7)·p⁷(1−p)³ column from what you know. The rest of the table is as it appeared.
| candidate p | L(p) = C(10,7)·p⁷(1−p)³ | verdict |
|---|---|---|
| 0.1 | 0.000009 | terrible |
| 0.5 | 0.117188 | so-so |
| 0.7 | 0.266828 | best (= k/n) |
| 0.9 | 0.057396 | over-shoots |
Concept
That product ∏ p(xᵢ;θ) is exact, but it is a nightmare to work with. Multiplying n numbers below 1 underflows to 0.0 in floating point for large n, and the product rule makes its derivative a mess of n terms.
We need the same argmax without the product. The fix is one of the most reused tricks in machine learning — the next slide.
Missing information
Discussion prompt
The product-vs-sum problem is not hypothetical. Take 2000 i.i.d. Gaussian densities and multiply them versus summing their logs — the raw product collapses to 0.0:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
2000 numbers below 1 multiply down past the smallest float and flush to exactly 0.0 — the likelihood is destroyed. The log-sum stays a normal, finite, differentiable number.
Worked example
The product-vs-sum problem is not hypothetical. Take 2000 i.i.d. Gaussian densities and multiply them versus summing their logs — the raw product collapses to 0.0:
import numpy as np
np.random.seed(0)
z = np.random.randn(2000)
dens = np.exp(-z**2/2) / np.sqrt(2*np.pi)
print('product :', np.prod(dens))
print('sum logs:', np.sum(np.log(dens)))The product prints 0.0; the sum of logs prints a finite −2794.78
Why: 2000 numbers below 1 multiply down past the smallest float and flush to exactly 0.0 — the likelihood is destroyed. The log-sum stays a normal, finite, differentiable number.
| method | printed value | usable? |
|---|---|---|
| np.prod(densities) | 0.0 | no — underflowed |
| np.sum(np.log(densities)) | −2794.78 | yes — finite |
Trade off
Comparison matrix
From Watch the product underflow: every row here is a choice with a cost. Fill the printed value column, then say which row you would actually pick and what you give up for it.
| method | printed value | usable? |
|---|---|---|
| np.prod(densities) | 0.0 | no — underflowed |
| np.sum(np.log(densities)) | −2794.78 | yes — finite |
Concept
Take the logarithm. log is strictly increasing, so it never moves the location of a maximum — but by the log-of-a-product rule it turns the product into a sum:
\[ \ell(\theta) = \log \prod_{i=1}^{n} p(x_i;\theta) = \sum_{i=1}^{n} \log p(x_i;\theta) \]
Sums don't underflow, and their derivative is just the sum of the term derivatives. Maximizing ℓ(θ) gives the same θ as maximizing the raw likelihood — far more cheaply.
Intuition
Picture the likelihood as a hill. log stretches the vertical axis — it squashes tall values and drops everything below 1 into negatives — but it never reorders heights: if a > b then log a > log b.
So the tallest point stays the tallest point. Its height changes; its horizontal position — the θ we actually want — does not. That single fact is the entire license for the trick.
Pattern
Predict first
The table runs: 3.0 | −20.8967 | far left · 4.0 | −17.8967 | climbing · 5.0 | −16.8967 | PEAK = sample mean · 6.0 | −17.8967 | descending
In The log-likelihood peaks at the same place, given the rows so far: what is the next one — the row where μ is 7.0?
Correct: 7.0 | −20.8967 | far right
| μ | ℓ(μ) at σ²=4 | note |
|---|---|---|
| 3.0 | −20.8967 | far left |
| 4.0 | −17.8967 | climbing |
| 5.0 | −16.8967 | PEAK = sample mean |
| 6.0 | −17.8967 | descending |
| 7.0 | −20.8967 | far right |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The log-likelihood is a downward parabola in μ: −16.8967 at μ=5 beats −17.8967 at μ=4 and 6, and −20.8967 at μ=3 and 7.
Worked example
Concrete check on Gaussian data [2,4,4,4,5,5,7,9] with variance fixed at σ² = 4. Sweep the mean μ and print the log-likelihood — the peak should land at the sample mean 5:
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
def ll(mu, var=4.0):
n = len(data)
return -0.5*n*np.log(2*np.pi*var) - (1/(2*var))*np.sum((data-mu)**2)
for mu in [3., 4., 5., 6., 7.]:
print(mu, round(ll(mu), 4))ℓ(μ) rises to a single peak at μ = 5, then falls symmetrically
Why: The log-likelihood is a downward parabola in μ: −16.8967 at μ=5 beats −17.8967 at μ=4 and 6, and −20.8967 at μ=3 and 7. One smooth hump, maximum at the sample mean.
| μ | ℓ(μ) at σ²=4 | note |
|---|---|---|
| 3.0 | −20.8967 | far left |
| 4.0 | −17.8967 | climbing |
| 5.0 | −16.8967 | PEAK = sample mean |
| 6.0 | −17.8967 | descending |
| 7.0 | −20.8967 | far right |
Discrimination
Sort into buckets
Sort these by ℓ(μ) at σ²=4, from memory, without looking back at The log-likelihood peaks at the same place. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
log transforms the function and shrinks its value, so the maximizing θ must shift too.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).
log is strictly increasing, so it preserves the location of every maximum.
Why: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).
Trap
log transforms the function and shrinks its value, so the maximizing θ must shift too.
Treat argmax ℓ(θ) as a DIFFERENT θ from argmax ∏ p(xᵢ;θ)
Why: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).
Conclusion: 'I must maximize the raw product to be safe'
Why: So you fight underflow and the product rule for nothing — and on large n the product is literally 0.0, so there is no peak left to find.
log is strictly increasing, so it preserves the location of every maximum.
argmax_θ ℓ(θ) = argmax_θ ∏ p(xᵢ;θ)
Why: A monotone transform reorders nothing: the tallest point before is the tallest point after. Only the height changes.
So maximize the SUM of logs — same θ, no underflow, clean derivative
Why: The sweep on the previous slide confirmed it: the log-likelihood peaked at exactly μ = 5, the same place the raw likelihood does.
Section
Part 2 of 6 — mean & variance, every step
Concept
One dataset carries the whole Gaussian derivation — the same eight numbers we just swept over. Assume they are i.i.d. draws from N(μ, σ²) and we want the best μ and σ².
| i | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| xᵢ | 2 | 4 | 4 | 4 | 5 | 5 | 7 | 9 |
Keep n = 8 and Σxᵢ = 40 in view — they are the only summaries the answers depend on.
Worked example
One point's density is p(xᵢ) = (2πσ²)^(−1/2) · exp(−(xᵢ−μ)²/2σ²). Take the log of a single term first:
log of one point
Why: log turns the exp into its exponent and the (2πσ²)^(−1/2) into −½log(2πσ²).
\[ \log p(x_i;\mu,\sigma^2) = -\tfrac{1}{2}\log(2\pi\sigma^2) - \frac{(x_i-\mu)^2}{2\sigma^2} \]
Sum over all n points
Why: By the log trick, ℓ is the sum of the per-point logs. The constant term repeats n times.
\[ \ell(\mu,\sigma^2) = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(x_i-\mu)^2 \]
Notation
Annotate
From Write the Gaussian log-likelihood — read this one piece at a time. What is each part doing?
On: \( \log p(x_i;\mu,\sigma^2) = -\tfrac{1}{2}\log(2\pi\sigma^2) - \frac{(x_i-\mu)^2}{2\sigma^2} \)
Ranking
Put in order
Put the moves of Differentiate in μ — step by step into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The −(n/2)log(2πσ²) term has no μ in it, so its μ-derivative is 0.
Worked example
Only the sum-of-squares term contains μ
Why: The −(n/2)log(2πσ²) term has no μ in it, so its μ-derivative is 0. Focus on the second term.
Differentiate Σ(xᵢ−μ)² with the chain rule
Why: d/dμ (xᵢ−μ)² = 2(xᵢ−μ)·(−1) = −2(xᵢ−μ). The −1 is the inner derivative of (xᵢ−μ).
\[ \frac{\partial \ell}{\partial \mu} = -\frac{1}{2\sigma^2}\sum_i 2(x_i-\mu)(-1) = \frac{1}{\sigma^2}\sum_i (x_i-\mu) \]
The two minus signs and the 2 combine into +1/σ²
Why: (−1/2σ²)·(−2) = +1/σ². What survives is (1/σ²)·Σ(xᵢ−μ) — clean and linear in μ.
Blank canvas
Draw it
Draw what Differentiate in μ — step by step just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Step zero
Discussion prompt
Set ∂ℓ/∂μ = 0 → the sample mean — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: At the maximum the derivative is zero
Answer:
Worked example
At the maximum the derivative is zero
Why: σ² > 0, so the 1/σ² factor can never be zero — the bracket must vanish.
\[ \frac{1}{\sigma^2}\sum_i (x_i-\mu) = 0 \;\Longrightarrow\; \sum_i (x_i-\mu) = 0 \]
Split the sum: Σxᵢ − nμ = 0
Why: Σ(xᵢ−μ) = Σxᵢ − Σμ, and Σμ over n points is nμ. Solve for μ.
\[ \sum_i x_i - n\mu = 0 \;\Longrightarrow\; \mu_{\text{MLE}} = \frac{1}{n}\sum_i x_i \]
Plug in our numbers: μ = 40/8 = 5.0
Why: The MLE of a Gaussian mean is nothing but the SAMPLE MEAN. Σxᵢ = 40, n = 8, so μ_MLE = 5.0 — matching the sweep's peak exactly.
| quantity | value |
|---|---|
| n | 8 |
| Σxᵢ | 40 |
| μ_MLE = Σxᵢ / n | 5.0 |
Concept
Setting a derivative to zero only finds a flat point — we owe a check that it is a peak, not a valley. Differentiate ∂ℓ/∂μ = (1/σ²)Σ(xᵢ−μ) once more:
\[ \frac{\partial^2 \ell}{\partial \mu^2} = \frac{1}{\sigma^2}\sum_i (-1) = -\frac{n}{\sigma^2} < 0 \]
The second derivative is −n/σ², strictly negative for all μ. So ℓ(μ) is concave — a single downward hump — and the stationary point μ = x̄ is the global maximum, exactly as the sweep in Part 1 showed.
Fill the middle
Fill in the blanks
From Differentiate in σ² — step by step — finish the line. Write what belongs on the right of the equals sign before you look.
\frac-\frac{n}{2v} + \frac{1}{2v^2}\sum_i (x_i-\mu)^2___ = ___
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. −(1/2v)·S = −(S/2)·v⁻¹, and d/dv(v⁻¹) = −v⁻², so the derivative is +(S/2)v⁻² = S/(2v²).
Worked example
Now hold μ = 5 and differentiate ℓ with respect to the variable v = σ². Both terms contain v, so differentiate each:
Term 1: d/dv [−(n/2)log(2πv)] = −n/(2v)
Why: d/dv log(2πv) = 1/v, times −n/2.
Term 2: d/dv [−(1/2v)·S] = +S/(2v²), where S = Σ(xᵢ−μ)²
Why: −(1/2v)·S = −(S/2)·v⁻¹, and d/dv(v⁻¹) = −v⁻², so the derivative is +(S/2)v⁻² = S/(2v²).
\[ \frac{\partial \ell}{\partial v} = -\frac{n}{2v} + \frac{1}{2v^2}\sum_i (x_i-\mu)^2 \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Term 2: d/dv [−(1/2v)·S] = +S/(2v²), where S = Σ(xᵢ−μ)²
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Now hold μ = 5 and differentiate ℓ with respect to the variable v = σ². Both terms contain v, so differentiate each:
Step zero
Discussion prompt
Set ∂ℓ/∂σ² = 0 → the 1/n variance — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Set to zero and multiply through by 2v²
Answer:
Worked example
Set to zero and multiply through by 2v²
Why: Clearing the denominators turns the equation into −nv + Σ(xᵢ−μ)² = 0.
\[ -\frac{n}{2v} + \frac{S}{2v^2} = 0 \;\Longrightarrow\; -nv + S = 0 \]
Solve for v = σ²
Why: v = S/n = the average squared deviation from the MLE mean. Note the divisor is n, not n−1.
\[ \sigma^2_{\text{MLE}} = \frac{1}{n}\sum_i (x_i-\mu_{\text{MLE}})^2 \]
Plug in: deviations are [−3,−1,−1,−1,0,0,2,4]
Why: xᵢ − 5 for our data. Square them: [9,1,1,1,0,0,4,16], which sum to S = 32.
σ²_MLE = 32/8 = 4.0, so σ_MLE = 2.0
Why: The variance MLE is 4.0 and the standard-deviation MLE is its root, 2.0 — matching the σ² = 4 we fixed during the sweep.
Translation
\( \sigma^2_{\text{MLE}} = \frac{1}{n}\sum_i (x_i-\mu_{\text{MLE}})^2 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Worked example
Every term of S = Σ(xᵢ−5)², in one table — this is the full trace behind σ²_MLE = 4.0:
| xᵢ | xᵢ − 5 | (xᵢ − 5)² |
|---|---|---|
| 2 | −3 | 9 |
| 4 | −1 | 1 |
| 4 | −1 | 1 |
| 4 | −1 | 1 |
| 5 | 0 | 0 |
| 5 | 0 | 0 |
| 7 | 2 | 4 |
| 9 | 4 | 16 |
S = 9+1+1+1+0+0+4+16 = 32 → σ²_MLE = 32/8 = 4.0
Why: Sum the last column and divide by n = 8. The sample standard deviation the MLE reports is √4 = 2.0.
Pattern
Step through it
Step through Full trace of the variance sum one row at a time. What is driving the change, and what would the row after the last one be?
Concept
We solved for μ first, then plugged that μ into the variance equation. That is valid because ∂ℓ/∂μ = 0 gives the same μ = x̄ for every fixed σ² — the best mean never depends on the variance.
So the joint maximum factors: find μ_MLE, then find the σ² that maximizes ℓ at that mean. No circularity — the two conditions are compatible.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.
The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.
Why: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.
Trap
Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.
Report σ² = Σ(xᵢ−μ)² / (n−1) = 32/7 ≈ 4.571
Why: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.
Claim this is 'the MLE'
Why: It is not. Setting ∂ℓ/∂σ² = 0 gave a divisor of n. Calling the n−1 version the MLE is exactly the error the exam probes.
The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.
σ²_MLE = Σ(xᵢ−μ)² / n = 32/8 = 4.0
Why: This is what maximizing the Gaussian likelihood actually yields — verified in code as np.var(data) with its default ddof=0.
Know both: MLE = ÷n (biased), unbiased = ÷(n−1)
Why: np.var(data, ddof=1) = 4.571 gives the unbiased one. The MLE and the unbiased estimator are two different, both-correct answers to two different questions.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.
Estimation
Predict first
np.var defaults to ddof=0 — the 1/n MLE. Pass ddof=1 for the 1/(n−1) unbiased version. Both, side by side:
Commit before you compute: what does Confirm both variances in NumPy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: np.var(data) = 4.0 matches our by-hand MLE; ddof=1 = 4.571
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The default np.var IS the MLE. The Bessel-corrected 32/7 = 4.5714 only appears when you ask for ddof=1.
Worked example
np.var defaults to ddof=0 — the 1/n MLE. Pass ddof=1 for the 1/(n−1) unbiased version. Both, side by side:
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
var_mle = ((data - mu)**2).mean() # divides by n
print(mu, var_mle)
print(np.var(data)) # ddof=0 -> MLE
print(np.var(data, ddof=1)) # ddof=1 -> unbiasednp.var(data) = 4.0 matches our by-hand MLE; ddof=1 = 4.571
Why: The default np.var IS the MLE. The Bessel-corrected 32/7 = 4.5714 only appears when you ask for ddof=1.
| call | value | which estimator |
|---|---|---|
| ((data-mu)**2).mean() | 4.0 | MLE (÷n) |
| np.var(data) | 4.0 | MLE (ddof=0) |
| np.var(data, ddof=1) | 4.571 | unbiased (÷(n−1)) |
Section
Part 3 of 6 — and Poisson, Exponential
Intuition
Back to Part 1's coin. We saw by brute force that p = 0.7 maximized the likelihood of 7-heads-in-10. Now we get the same 0.7 from calculus — and see it holds for any k and n, not just this one sample.
The lesson: the sweep and the derivative are two views of the same peak. Once you trust the recipe, you skip the sweep and go straight to setting dℓ/dp = 0.
Ranking
Put in order
Put the moves of Derive the Bernoulli MLE into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Σ[xᵢ log p + (1−xᵢ)log(1−p)] collapses because Σxᵢ = k and Σ(1−xᵢ) = n−k.
Worked example
Data are 0/1 outcomes with k successes in n trials. One point's probability is p^{xᵢ}(1−p)^{1−xᵢ}. Its log summed over all points is:
Log-likelihood in p
Why: Σ[xᵢ log p + (1−xᵢ)log(1−p)] collapses because Σxᵢ = k and Σ(1−xᵢ) = n−k.
\[ \ell(p) = k\log p + (n-k)\log(1-p) \]
Differentiate: dℓ/dp = k/p − (n−k)/(1−p)
Why: d/dp log p = 1/p; d/dp log(1−p) = −1/(1−p), and the −1 flips the sign of the second term.
Set to zero and cross-multiply
Why: k/p = (n−k)/(1−p) ⇒ k(1−p) = (n−k)p ⇒ k − kp = np − kp ⇒ k = np.
\[ p_{\text{MLE}} = \frac{k}{n} \]
Fill the middle
Fill in the blanks
From Bernoulli MLE in code — one line has had its right-hand side removed. Put it back.
import numpy as np
b = np.array([1,0,1,1,0,1,1,0,1,1])
k = b.sum()
n = len(b)
p_mle = b.mean() # = k/n
print(k, n, p_mle)
Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The mean of a 0/1 array IS the success fraction.
Worked example
The success fraction, on a 10-outcome sample with 7 ones — confirming p_MLE = k/n:
import numpy as np
b = np.array([1,0,1,1,0,1,1,0,1,1])
k = b.sum()
n = len(b)
p_mle = b.mean() # = k/n
print(k, n, p_mle)k = 7, n = 10 → p_MLE = 0.7
Why: The mean of a 0/1 array IS the success fraction. It matches the coin experiment from Part 1 — the same k/n falls out whether you sweep candidates or take the derivative.
| k (successes) | n (trials) | p_MLE = k/n |
|---|---|---|
| 7 | 10 | 0.7 |
Sorting
Sort into buckets
These are the pieces of Lesson 8: Maximum Likelihood Estimation, out of order. Put each one back under the part of the lesson it belongs to.
Ranking
Put in order
These are the steps of The MLE recipe, scrambled. Put them back in order before the next slide shows you.
p(x; θ) for one data pointℓ(θ) = Σ log p(xᵢ; θ) — the log trick turns the product into a sumℓ with respect to each parameterΣxᵢ, k, Σxᵢ², …) to get a numberWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
p(x; θ) for one data pointℓ(θ) = Σ log p(xᵢ; θ) — the log trick turns the product into a sumℓ with respect to each parameterΣxᵢ, k, Σxᵢ², …) to get a numberEvery MLE today — Gaussian, Bernoulli, Poisson, Exponential — is this exact five-step loop. Learn the loop, not the formulas.
Fill the middle
Fill in the blanks
From Poisson and Exponential fall out too — one line has had its right-hand side removed. Put it back.
import numpy as np
pois = np.array([2,3,1,4,2,3])
ex = np.array([0.5,1.5,2.0,1.0])
print('Poisson lambda =', pois.mean()) # MLE = sample mean
print('Exp lambda =', 1/ex.mean()) # MLE = 1/sample mean
Why: ex is what everything below it consumes, so the wrong expression here fails later and somewhere else. d/dλ Σ(xᵢ log λ − λ) = Σxᵢ/λ − n = 0 ⇒ λ = Σxᵢ/n.
Worked example
Apply the same recipe to two more families and confirm the closed forms in code:
import numpy as np
pois = np.array([2,3,1,4,2,3])
ex = np.array([0.5,1.5,2.0,1.0])
print('Poisson lambda =', pois.mean()) # MLE = sample mean
print('Exp lambda =', 1/ex.mean()) # MLE = 1/sample meanPoisson: ℓ = Σ(xᵢ log λ − λ) ⇒ λ_MLE = mean = 2.5
Why: d/dλ Σ(xᵢ log λ − λ) = Σxᵢ/λ − n = 0 ⇒ λ = Σxᵢ/n. For [2,3,1,4,2,3], that is 15/6 = 2.5.
Exponential: ℓ = Σ(log λ − λxᵢ) ⇒ λ_MLE = 1/mean = 0.8
Why: d/dλ = n/λ − Σxᵢ = 0 ⇒ λ = n/Σxᵢ = 1/x̄. For mean 1.25, λ = 0.8.
| family | MLE | value |
|---|---|---|
| Poisson | λ = sample mean | 2.5 |
| Exponential | λ = 1 / sample mean | 0.8 |
Error analysis
Annotate
Walk the callouts on Poisson and Exponential fall out too. Each one is a place this is easy to get subtly wrong.
Elimination
Eliminate the wrong options
Maximizing ℓ(θ) = Σ log p(xᵢ;θ) gives the same θ as maximizing ∏ p(xᵢ;θ) because…
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A monotonically increasing transform changes the height of the peak but never its location, so the argmax is unchanged. As a bonus, log turns the product into a sum, which is why we use it.
Check
Why is it always safe to maximize the log-likelihood instead of the likelihood?
Check your understanding
Maximizing ℓ(θ) = Σ log p(xᵢ;θ) gives the same θ as maximizing ∏ p(xᵢ;θ) because…
Answer: A
Why: A monotonically increasing transform changes the height of the peak but never its location, so the argmax is unchanged. As a bonus, log turns the product into a sum, which is why we use it.
Prediction
Predict first
For Gaussian data, the maximum-likelihood estimate of the variance divides the sum of squared deviations by…
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: n (and is biased low)
Why: Setting ∂ℓ/∂σ² = 0 yields σ²_MLE = (1/n)Σ(xᵢ−μ)². It is biased low because the same data estimates both μ and σ. On our data that is 32/8 = 4.0.
Check
Careful with the denominator.
Check your understanding
For Gaussian data, the maximum-likelihood estimate of the variance divides the sum of squared deviations by…
Answer: A
Why: Setting ∂ℓ/∂σ² = 0 yields σ²_MLE = (1/n)Σ(xᵢ−μ)². It is biased low because the same data estimates both μ and σ. On our data that is 32/8 = 4.0.
Section
Part 4 of 6 — MSE & cross-entropy
Intuition
Here is the payoff. Every loss you have minimized — squared error for regression, cross-entropy for classification — is not an arbitrary choice someone handed you. Each one is the negative log-likelihood of a probabilistic model.
Pick a noise model, write its likelihood, negate the log — and out drops the loss. Change the noise model and the loss changes with it. Let's watch MSE fall out of a Gaussian.
Picture it
Figure (svg): A rising straight regression line with small vertical bell curves centered on the line at four x positions, and data points scattered near each bell.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture the regression line, and at each xᵢ a small Gaussian bell standing up off the line. The model says the true yᵢ was drawn from that bell — most likely near the line, occasionally far.
Intuition
Picture the regression line, and at each xᵢ a small Gaussian bell standing up off the line. The model says the true yᵢ was drawn from that bell — most likely near the line, occasionally far.
Figure (svg): A rising straight regression line with small vertical bell curves centered on the line at four x positions, and data points scattered near each bell.
The farther a point sits from the line, the lower its bell-height, the smaller its likelihood. Maximizing total likelihood pulls the line to where the points cluster — which is exactly least squares.
Concept
Assume each target is the model's prediction plus independent Gaussian noise:
\[ y_i = f(x_i) + \varepsilon_i, \qquad \varepsilon_i \sim \mathcal{N}(0,\sigma^2) \]
Then yᵢ is Gaussian centered at f(xᵢ) with variance σ², so its density is (2πσ²)^{−1/2} exp(−(yᵢ−f(xᵢ))²/2σ²). Now run the MLE recipe on that.
Step zero
Discussion prompt
From Gaussian likelihood to MSE — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Negative log-likelihood of the targets
Answer:
Worked example
Negative log-likelihood of the targets
Why: Same log as the Gaussian in Part 2, but centered at f(xᵢ) instead of μ, and negated.
\[ -\ell = \frac{n}{2}\log(2\pi\sigma^2) + \frac{1}{2\sigma^2}\sum_i \big(y_i - f(x_i)\big)^2 \]
The log term is constant in the model parameters
Why: (n/2)log(2πσ²) has no f in it, so it does not move the argmin over f's parameters — drop it.
Drop the positive scale 1/2σ²
Why: A positive constant multiplier never moves an argmin. What remains is exactly the sum of squared errors.
\[ \arg\min \;(-\ell) = \arg\min \sum_i \big(y_i - f(x_i)\big)^2 \]
This is the MSE objective — Lesson 7's OLS
Why: Least squares was never a convention. Its normal equations gave w = [2.2, 0.6] because MSE IS the MLE under Gaussian noise.
Blank canvas
Draw it
Draw what From Gaussian likelihood to MSE just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From The tie-back to Lesson 7, in code — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1.,2.,3.,4.,5.])
y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
sse = np.sum((y - X @ w)**2)
print(w, sse, sse/5)
Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Minimizing MSE gave the same w = [2.2, 0.6] as Lesson 7, and the noise-variance MLE is just the mean squared residual, 0.48.
Worked example
Re-fit the 5-student line from Lesson 7 and read off the noise-variance MLE σ² = SSE/n — the least-squares residuals ARE the Gaussian MLE:
import numpy as np
x = np.array([1.,2.,3.,4.,5.])
y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
sse = np.sum((y - X @ w)**2)
print(w, sse, sse/5)w = [2.2, 0.6], SSE = 2.4, σ²_MLE = 2.4/5 = 0.48
Why: Minimizing MSE gave the same w = [2.2, 0.6] as Lesson 7, and the noise-variance MLE is just the mean squared residual, 0.48.
| quantity | value |
|---|---|
| w (OLS / MLE) | [2.2, 0.6] |
| SSE = Σ(y − ŷ)² | 2.4 |
| σ²_MLE = SSE / n | 0.48 |
Concept
Now the classification side. Model P(yᵢ=1|xᵢ) = pᵢ (e.g. pᵢ = σ(wᵀxᵢ+b)). Each label is Bernoulli, so its log-likelihood is the Bernoulli log from Part 3:
\[ \ell = \sum_i \big[\, y_i \log p_i + (1-y_i)\log(1-p_i)\,\big] \]
The negative of that, averaged, is exactly the binary cross-entropy loss. So minimizing BCE literally IS maximizing the Bernoulli likelihood of the labels.
Pattern
Predict first
The table runs: (1, 0.9) | −log(0.9) | 0.1054 · (0, 0.2) | −log(0.8) | 0.2231 · (1, 0.7) | −log(0.7) | 0.3567
In BCE equals the negative log-likelihood, given the rows so far: what is the next one — the row where point (y, p) is mean of the three?
Correct: mean of the three | = F.binary_cross_entropy | 0.2284
| point (y, p) | −log-likelihood term | value |
|---|---|---|
| (1, 0.9) | −log(0.9) | 0.1054 |
| (0, 0.2) | −log(0.8) | 0.2231 |
| (1, 0.7) | −log(0.7) | 0.3567 |
| mean of the three | = F.binary_cross_entropy | 0.2284 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. y=1 point contributes −log(0.9)=0.1054; y=0 point contributes −log(1−0.2)=−log(0.8)=0.2231; y=1 point contributes −log(0.7)=0.3567.
Worked example
Confirm numerically that a hand-computed mean NLL matches torch's BCE for predictions [0.9, 0.2, 0.7], labels [1, 0, 1]:
import numpy as np, torch
import torch.nn.functional as F
y = np.array([1.,0.,1.]); p = np.array([0.9,0.2,0.7])
nll = -np.mean(y*np.log(p) + (1-y)*np.log(1-p))
tb = F.binary_cross_entropy(torch.tensor(p), torch.tensor(y)).item()
print(round(nll,4), round(tb,4))Per-point −log terms: 0.1054, 0.2231, 0.3567
Why: y=1 point contributes −log(0.9)=0.1054; y=0 point contributes −log(1−0.2)=−log(0.8)=0.2231; y=1 point contributes −log(0.7)=0.3567.
mean NLL = 0.6852/3 = 0.2284 = torch BCE
Why: The three terms sum to 0.6852; divided by n=3 gives 0.2284, exactly what F.binary_cross_entropy returns. BCE is the averaged negative Bernoulli log-likelihood — same formula.
| point (y, p) | −log-likelihood term | value |
|---|---|---|
| (1, 0.9) | −log(0.9) | 0.1054 |
| (0, 0.2) | −log(0.8) | 0.2231 |
| (1, 0.7) | −log(0.7) | 0.3567 |
| mean of the three | = F.binary_cross_entropy | 0.2284 |
Comparison
Comparison matrix
From BCE equals the negative log-likelihood: refill the −log-likelihood term column from what you know. The rest of the table is as it appeared.
| point (y, p) | −log-likelihood term | value |
|---|---|---|
| (1, 0.9) | −log(0.9) | 0.1054 |
| (0, 0.2) | −log(0.8) | 0.2231 |
| (1, 0.7) | −log(0.7) | 0.3567 |
| mean of the three | = F.binary_cross_entropy | 0.2284 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
MLE maximizes likelihood, and 'the loss is the likelihood', so training should maximize the cross-entropy.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This silently drops the negative sign.
The loss is the negative log-likelihood, so maximizing likelihood = minimizing loss.
Why: This silently drops the negative sign. Cross-entropy is the NEGATED log-likelihood, so climbing it drives the likelihood DOWN.
Trap
MLE maximizes likelihood, and 'the loss is the likelihood', so training should maximize the cross-entropy.
Run gradient ASCENT on the cross-entropy value
Why: This silently drops the negative sign. Cross-entropy is the NEGATED log-likelihood, so climbing it drives the likelihood DOWN.
Result: predictions get worse each step
Why: You are maximizing −ℓ, i.e. minimizing ℓ — the exact opposite of MLE. Loss curves shoot upward and the model diverges.
The loss is the negative log-likelihood, so maximizing likelihood = minimizing loss.
Minimize L = −ℓ with gradient DESCENT
Why: max ℓ ⇔ min(−ℓ). The single sign flip is what connects the probability story to the descent loop from Lesson 6.
Loss decreases, likelihood of the data increases
Why: Descending BCE raises P(labels | model) step by step — training and MLE are the same motion, read with opposite signs.
Commit first
Predict first
Minimizing mean squared error is equivalent to maximum likelihood under which assumption?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: The target noise is Gaussian with constant variance
Why: With y = f(x) + ε and ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² plus constants — exactly MSE.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Where does squared-error loss actually come from?
Check your understanding
Minimizing mean squared error is equivalent to maximum likelihood under which assumption?
Answer: A
Why: With y = f(x) + ε and ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² plus constants — exactly MSE.
Prediction
Predict first
Training a classifier by minimizing binary cross-entropy is equivalent to…
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: maximizing the Bernoulli log-likelihood of the labels
Why: BCE is the averaged NEGATIVE Bernoulli log-likelihood, so min BCE = max ℓ. Our numeric check gave the same 0.2284 for the hand NLL and torch's BCE.
Check
Connect the probability story to the training loop.
Check your understanding
Training a classifier by minimizing binary cross-entropy is equivalent to…
Answer: A
Why: BCE is the averaged NEGATIVE Bernoulli log-likelihood, so min BCE = max ℓ. Our numeric check gave the same 0.2284 for the hand NLL and torch's BCE.
Section
Part 5 of 6 — regularization = prior
Concept
MLE trusts the data alone. MAP (maximum a posteriori) adds a prior p(θ) — a belief about θ before seeing data — and maximizes the posterior p(θ|data) ∝ p(data|θ)·p(θ).
Take logs, and the product becomes a sum: the log-likelihood plus the log-prior. The prior turns into an additive term on the objective.
\[ \theta_{\text{MAP}} = \arg\max_\theta\; \big[\, \ell(\theta) + \log p(\theta) \,\big] \]
Ranking
Put in order
Put the moves of A Gaussian prior IS the ridge penalty into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. A belief that weights are small and centered at 0 — the same belief ridge encodes.
Worked example
Put a Gaussian prior on the weights: w ~ N(0, τ²I)
Why: A belief that weights are small and centered at 0 — the same belief ridge encodes.
\[ \log p(w) = -\frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]
MAP objective = log-likelihood + log-prior
Why: Add the Gaussian-noise log-likelihood (→ −½σ⁻²·SSE) to the log-prior (→ −½τ⁻²‖w‖²), then negate to get a loss to minimize.
\[ -\big[\ell + \log p(w)\big] \;\propto\; \lVert Xw - y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \]
Identify λ = σ²/τ² — this is ridge from Lesson 7
Why: The log-prior contributed exactly λ‖w‖². Ridge's penalty was never ad hoc — it is a Gaussian prior on the weights. A Laplace prior would give L1 instead.
Notation
Annotate
From A Gaussian prior IS the ridge penalty — read this one piece at a time. What is each part doing?
On: \( -\big[\ell + \log p(w)\big] \;\propto\; \lVert Xw - y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \)
Concept
A prior centered at 0 pulls the estimate toward 0 — the tighter the prior (smaller τ²), the harder the pull. As τ² → ∞ the prior goes flat and MAP relaxes back to plain MLE.
\[ \mu_{\text{MAP}} = \frac{n\,\bar x/\sigma^2}{\,n/\sigma^2 + 1/\tau^2\,} \]
This is the posterior mean for a Gaussian mean with a N(0, τ²) prior — a precision-weighted blend of the data mean x̄ and the prior mean 0.
Explain it
Discussion prompt
Explain The prior shrinks the estimate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A prior centered at 0 pulls the estimate toward 0 — the tighter the prior (smaller τ²), the harder the pull. As τ² → ∞ the prior goes flat and MAP relaxes back to plain MLE.
Missing information
Discussion prompt
Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
A flat prior recovers the sample mean 5. Tightening the prior toward 0 shrinks the estimate — 4.44, then 3.33. This is exactly what ridge does to regression weights.
Worked example
Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
n, xbar, s2 = len(data), data.mean(), 4.0
for tau2 in [np.inf, 4.0, 1.0]:
prec = n/s2 + (0.0 if np.isinf(tau2) else 1/tau2)
print(tau2, round((n*xbar/s2)/prec, 4))τ²=∞ → 5.0 (pure MLE); τ²=4 → 4.4444; τ²=1 → 3.3333
Why: A flat prior recovers the sample mean 5. Tightening the prior toward 0 shrinks the estimate — 4.44, then 3.33. This is exactly what ridge does to regression weights.
| prior variance τ² | prior strength | μ_MAP |
|---|---|---|
| ∞ | none (flat) → MLE | 5.0000 |
| 4.0 | moderate | 4.4444 |
| 1.0 | strong | 3.3333 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
τ²=∞ → 5.0 (pure MLE); τ²=4 → 4.4444; τ²=1 → 3.3333
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:
Anomaly
Predict first
A student writes this, and it looks reasonable:
MAP is Bayesian and MLE is frequentist, so MAP must need a totally separate machinery.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.
MAP is MLE plus one extra term: the log-prior. Same optimization, one added regularizer.
Why: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.
Trap
MAP is Bayesian and MLE is frequentist, so MAP must need a totally separate machinery.
Treat the prior as a bolt-on unrelated to the loss
Why: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.
Fail to see ridge/L2 as a prior
Why: So regularization looks like a hack you sprinkle on, rather than a modeling choice with a probabilistic meaning.
MAP is MLE plus one extra term: the log-prior. Same optimization, one added regularizer.
Objective = ℓ(θ) + log p(θ), minimize its negative
Why: The log-prior sits right beside the log-likelihood. Everything you know about maximizing ℓ still applies.
Gaussian prior ⇒ L2 (ridge); Laplace prior ⇒ L1
Why: Regularization strength λ = σ²/τ² comes straight from the prior's variance. Regularizing IS choosing a prior — nothing bolted on.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
p. But your gut already leans: this looks like a coin that favors heads.; A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.; Maximum likelihood estimation picks the parameter that maximizes that function — the θ that makes exactly this dataset as probable as possible:log transforms the function and shrinks its value, so the maximizing θ must shift too.; Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.Ranking
Put in order
These are the steps of Deriving a loss from a model, scrambled. Put them back in order before the next slide shows you.
p(x; θ) or p(y | x; θ)ℓ(θ) = Σ log pL = −ℓ (Gaussian noise → MSE, Bernoulli → BCE)−log p(θ) if you have a prior → the regularizer (Gaussian → L2, Laplace → L1)L by gradient descent — the Lesson 6 loopWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
p(x; θ) or p(y | x; θ)ℓ(θ) = Σ log pL = −ℓ (Gaussian noise → MSE, Bernoulli → BCE)−log p(θ) if you have a prior → the regularizer (Gaussian → L2, Laplace → L1)L by gradient descent — the Lesson 6 loopThis recipe reduces 'which loss?' and 'which regularizer?' to two modeling questions: what is the noise, and what do I believe about the parameters.
Edge cases
Discussion prompt
Deriving a loss from a model works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
This recipe reduces 'which loss?' and 'which regularizer?' to two modeling questions: what is the noise, and what do I believe about the parameters.
Elimination
Eliminate the wrong options
Adding an L2 (ridge) penalty λ‖w‖² to a squared-error loss is equivalent to MAP estimation with which prior on the weights?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: log of a zero-mean Gaussian prior is −‖w‖²/(2τ²) + const, which is exactly the L2 penalty with λ = σ²/τ². A tighter prior (smaller τ²) means larger λ and stronger shrinkage.
Check
Tie MAP back to Lesson 7's ridge.
Check your understanding
Adding an L2 (ridge) penalty λ‖w‖² to a squared-error loss is equivalent to MAP estimation with which prior on the weights?
Answer: A
Why: log of a zero-mean Gaussian prior is −‖w‖²/(2τ²) + const, which is exactly the L2 penalty with λ = σ²/τ². A tighter prior (smaller τ²) means larger λ and stronger shrinkage.
Section
Part 6 of 6 — the project
Concept
Write one function that returns the MLE μ and σ of a 1-D dataset, then prove it matches NumPy. You have derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | μ = sample mean | data.mean() |
| 2 | σ² = mean squared deviation (÷n) | ((data-μ)**2).mean() |
| 3 | Verify against np.var (ddof=0 = MLE) | np.var(data) |
Build rules: type every line yourself, run after each line, and confirm np.var defaults to ddof=0 (the 1/n MLE) before you trust it.
Analogy
Discussion prompt
Explain Project: maximum_likelihood_gaussian(data) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Write one function that returns the MLE μ and σ of a 1-D dataset, then prove it matches NumPy. You have derived every piece — now assemble it yourself.
Worked example
Your turn: compute μ for [2,4,4,4,5,5,7,9]. Say the value out loud from Σ/n before you print it.
Hint: the MLE mean is just data.mean() — no loop needed. Expect 40/8 = 5.0.
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
print(data.sum(), mu)| check | value |
|---|---|
| Σ data | 40.0 |
| μ = Σ/8 | 5.0 |
Worked example
Your turn: compute σ² as the mean squared deviation from μ. Will it be above or below the unbiased 4.571?
Hint: ((data-mu)**2).mean() divides by n (the MLE), so it comes out below the unbiased estimate. Expect 32/8 = 4.0.
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
var_mle = ((data - mu)**2).mean()
sigma = var_mle**0.5
print(((data-mu)**2).sum(), var_mle, sigma)| quantity | value |
|---|---|
| Σ(x−μ)² | 32.0 |
| σ²_MLE = 32/8 | 4.0 |
| σ_MLE = √4 | 2.0 |
Pattern
Step through it
Step through Milestone 2 — the variance one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: confirm your σ² matches np.var(data) (default ddof=0 = the 1/n MLE), and see how ddof=1 differs.
Hint: np.var(data) uses ddof=0 (÷n); np.var(data, ddof=1) uses ÷(n−1). Only the first should match your var_mle.
import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
var_mle = ((data - data.mean())**2).mean()
print(var_mle) # your MLE
print(np.var(data)) # ddof=0 -> MLE
print(np.var(data, ddof=1)) # ddof=1 -> unbiased| call | value | matches your MLE? |
|---|---|---|
| your var_mle | 4.0 | — |
| np.var(data) | 4.0 | ✓ (MLE) |
| np.var(data, ddof=1) | 4.571 | no (unbiased) |
Estimation
Predict first
Assemble the three milestones into one function and prove it agrees with NumPy in a single boolean:
Commit before you compute: what does The full function come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Returns (5.0, 2.0); np.allclose vs np.var prints True
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Your from-scratch MLE matches NumPy's default variance exactly.
Worked example
Assemble the three milestones into one function and prove it agrees with NumPy in a single boolean:
import numpy as np
def maximum_likelihood_gaussian(data):
data = np.asarray(data, dtype=float)
mu = data.mean() # MLE mean
var = ((data - mu)**2).mean() # MLE variance (1/n)
return mu, var**0.5
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu, sigma = maximum_likelihood_gaussian(data)
print(mu, sigma) # 5.0 2.0
print(np.allclose(sigma**2, np.var(data))) # TrueReturns (5.0, 2.0); np.allclose vs np.var prints True
Why: Your from-scratch MLE matches NumPy's default variance exactly. If you see (5.0, 2.0) and True, you have implemented maximum likelihood for the Gaussian.
| output line | value |
|---|---|
| mu, sigma | 5.0 2.0 |
| allclose vs np.var | True |
Trade off
Comparison matrix
From The full function: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| output line | value |
|---|---|
| mu, sigma | 5.0 2.0 |
| allclose vs np.var | True |
Concept
Slides closed, out loud: explain (1) why the log trick never moves the argmax, (2) why MSE and cross-entropy are both negative log-likelihoods, and (3) which prior turns MLE into ridge.
Stretch: prove the Poisson MLE is the sample mean (verify on [2,3,1,4,2,3] → 2.5) and the Exponential MLE is 1/mean (on [0.5,1.5,2,1] → 0.8). Both fall straight out of the five-step recipe.
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why the log trick never moves the argmax, (2) why MSE and cross-entropy are both negative log-likelihoods, and (3) which prior turns MLE into ridge.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Step back and see how much collapsed into a single idea. Each row below is the same recipe — model, log, negate — run on a different assumption:
| assumption | MLE gives | the loss it becomes |
|---|---|---|
| Gaussian noise on y | least squares | MSE |
| Bernoulli labels | success probabilities | cross-entropy (BCE) |
| Gaussian prior on w | shrunk weights | L2 / ridge penalty |
| Laplace prior on w | sparse weights | L1 / lasso penalty |
You never again memorize a loss. You state a probabilistic model and the loss is forced on you — that is the whole reason this lesson sits under every later one.
Comparison
Comparison matrix
From One principle, every loss: refill the the loss it becomes column from what you know. The rest of the table is as it appeared.
| assumption | MLE gives | the loss it becomes |
|---|---|---|
| Gaussian noise on y | least squares | MSE |
| Bernoulli labels | success probabilities | cross-entropy (BCE) |
| Gaussian prior on w | shrunk weights | L2 / ridge penalty |
| Laplace prior on w | sparse weights | L1 / lasso penalty |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The likelihood principle · MLE for the Gaussian · MLE for the Bernoulli · Losses are negative log-likelihoods · MAP: MLE plus a prior · Your turn: build the MLE. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
p=0.7 won the coin)1/n variance (μ=5, σ²=4); Bernoulli → success fraction (0.7); Poisson → mean; Exponential → 1/mean0.2284 against torchλ = σ²/τ²), L1 ↔ Laplace, and watch the estimate shrink from 5 toward 0maximum_likelihood_gaussian from scratch and match np.var| idea | the one thing to remember |
|---|---|
| MLE | argmax of the likelihood = best-explaining θ |
| log trick | monotone → same argmax; product → sum |
| loss | loss = −log-likelihood (minimize it) |
| MLE variance | divide by n, not n−1 (biased low) |
| MSE / BCE | Gaussian noise → MSE; Bernoulli → cross-entropy |
| MAP | MLE + log-prior; Gaussian prior → ridge |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.