Lesson 8: Maximum Likelihood Estimation

USAAIO Lesson 8, from Week 3 on probability, fully worked. It states maximum likelihood estimation as an optimization principle and scores it on a real coin, then justifies the log trick with a live underflow demo. It derives the Gaussian mean AND variance one calculus step at a time on a fixed 8-point dataset, including the second-derivative concavity check, and covers the Bernoulli, Poisson, and Exponential MLEs. It then delivers the big reveal that MSE and cross-entropy ARE negative log-likelihoods, checking both against torch, and closes with MAP as MLE plus a log-prior, from which ridge regression falls out of a Gaussian prior. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 62 slides.

Subject: Machine Learning · 119 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Maximum Likelihood Estimation

Title

USAAIO · Lesson 8 · Week 3 (Probability)

The single principle behind every loss you'll ever minimize. We score parameters on a real coin, derive the Gaussian and Bernoulli MLEs one step at a time, then prove that MSE and cross-entropy ARE negative log-likelihoods — and check both against torch.

2. By the end of this lesson you can

Objectives

  1. State MLE as "pick the θ that makes the observed data most probable" and evaluate the likelihood of competing parameters by hand
  2. Justify the log-likelihood trick — why log turns a product into a sum without moving the argmax
  3. Derive the MLE for a Gaussian (→ sample mean AND the 1/n variance) and a Bernoulli (→ success fraction), skipping no calculus step
  4. Show that MSE = MLE under Gaussian noise and BCE = MLE under a Bernoulli model, and verify BCE numerically against torch
  5. Read MAP = MLE + log-prior, deriving ridge from a Gaussian prior, and build maximum_likelihood_gaussian from scratch

3. What survived from Linear Systems & the Normal Equations?

Warm-up

Discussion prompt

Before we open Lesson 8: Maximum Likelihood Estimation: without looking back, what was the main idea of Linear Systems & the Normal Equations, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

overdetermined systems written out equation by equation, the normal equations XᵀXw = Xᵀy derived twice (projection AND calculus) with no skipped steps, the 5-student 2×2 solved by hand, invertibility and rank, ridge derived from the penalized loss, R², and a from-scratch OLS verified against sklearn.

4. The likelihood principle

Section

Part 1 of 6 — what MLE asks

5. A coin told you something

Intuition

You flip a coin 10 times and see 7 heads. Nobody told you the coin's bias p. But your gut already leans: this looks like a coin that favors heads.

A coin with p = 0.7 would produce 7-of-10 fairly often. A coin with p = 0.1 almost never would. So p = 0.7 explains what you saw far better than p = 0.1.

MLE is just that instinct made exact: among all candidate parameters, pick the one under which your data is most probable. The rest of Part 1 turns 'most probable' into calculus.

6. Break it if you can: A coin told you something

Counterexample

Discussion prompt

You flip a coin 10 times and see 7 heads. Nobody told you the coin's bias p. But your gut already leans: this looks like a coin that favors heads.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

A coin with p = 0.7 would produce 7-of-10 fairly often. A coin with p = 0.1 almost never would. So p = 0.7 explains what you saw far better than p = 0.1.

7. The likelihood function

Concept

You observe i.i.d. data x₁,…,xₙ drawn from a family p(x; θ). Because the draws are independent, the probability of the whole dataset is the product of the per-point probabilities:

\[ \mathcal{L}(\theta) = p(x_1,\dots,x_n;\theta) = \prod_{i=1}^{n} p(x_i;\theta) \]

likelihood — The SAME formula as a probability, but read as a function of the parameter θ with the data held fixed. We are not scoring which data is likely — we are scoring which θ best explains the data we already have.

8. By analogy: The likelihood function

Analogy

Discussion prompt

Explain The likelihood function by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

You observe i.i.d. data x₁,…,xₙ drawn from a family p(x; θ). Because the draws are independent, the probability of the whole dataset is the product of the per-point probabilities:

9. Likelihood is not probability of θ

Concept

A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.

So L(θ) need not integrate to 1 over θ, and 'the most likely θ' means 'the θ under which the data is most probable' — not 'the θ most probably true'. Turning θ into something with its own probability is what MAP does, in Part 5.

10. Teach it back: Likelihood is not probability of θ

Explain it

Discussion prompt

Explain Likelihood is not probability of θ to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.

11. MLE is an argmax

Concept

Maximum likelihood estimation picks the parameter that maximizes that function — the θ that makes exactly this dataset as probable as possible:

\[ \theta_{\text{MLE}} = \arg\max_{\theta}\; \prod_{i=1}^{n} p(x_i; \theta) \]

arg max returns the location of the maximum (the best θ), not its height. Everything today is a hunt for that location — usually by setting a derivative to zero.

12. Guess the shape of the answer: Score three coins on the 7-heads data

Estimation

Predict first

Let's actually evaluate the likelihood for the coin. With k = 7 heads in n = 10, L(p) = C(10,7)·pᵏ(1−p)ⁿ⁻ᵏ. Compute it for three candidate p and read off the winner:

Commit before you compute: what does Score three coins on the 7-heads data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Peak likelihood is at p = 0.7 — the empirical head-fraction

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 0.7 assigns probability 0.266828 to the observed 7-of-10, beating 0.5 (0.117188) and 0.9 (0.057396).

13. Score three coins on the 7-heads data

Worked example

Let's actually evaluate the likelihood for the coin. With k = 7 heads in n = 10, L(p) = C(10,7)·pᵏ(1−p)ⁿ⁻ᵏ. Compute it for three candidate p and read off the winner:

from math import comb
k, n = 7, 10
for p in [0.1, 0.5, 0.7, 0.9]:
    L = comb(n, k) * p**k * (1-p)**(n-k)
    print(p, round(L, 6))

Peak likelihood is at p = 0.7 — the empirical head-fraction

Why: 0.7 assigns probability 0.266828 to the observed 7-of-10, beating 0.5 (0.117188) and 0.9 (0.057396). The best explanation is exactly the observed frequency k/n.

candidate pL(p) = C(10,7)·p⁷(1−p)³verdict
0.10.000009terrible
0.50.117188so-so
0.70.266828best (= k/n)
0.90.057396over-shoots

14. Fill in: L(p) = C(10,7)·p⁷(1−p)³ for Score three coins on the 7-heads data

Comparison

Comparison matrix

From Score three coins on the 7-heads data: refill the L(p) = C(10,7)·p⁷(1−p)³ column from what you know. The rest of the table is as it appeared.

candidate pL(p) = C(10,7)·p⁷(1−p)³verdict
0.10.000009terrible
0.50.117188so-so
0.70.266828best (= k/n)
0.90.057396over-shoots

15. The product is the problem

Concept

That product ∏ p(xᵢ;θ) is exact, but it is a nightmare to work with. Multiplying n numbers below 1 underflows to 0.0 in floating point for large n, and the product rule makes its derivative a mess of n terms.

We need the same argmax without the product. The fix is one of the most reused tricks in machine learning — the next slide.

16. What has to be given first: Watch the product underflow

Missing information

Discussion prompt

The product-vs-sum problem is not hypothetical. Take 2000 i.i.d. Gaussian densities and multiply them versus summing their logs — the raw product collapses to 0.0:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

2000 numbers below 1 multiply down past the smallest float and flush to exactly 0.0 — the likelihood is destroyed. The log-sum stays a normal, finite, differentiable number.

17. Watch the product underflow

Worked example

The product-vs-sum problem is not hypothetical. Take 2000 i.i.d. Gaussian densities and multiply them versus summing their logs — the raw product collapses to 0.0:

import numpy as np
np.random.seed(0)
z = np.random.randn(2000)
dens = np.exp(-z**2/2) / np.sqrt(2*np.pi)
print('product :', np.prod(dens))
print('sum logs:', np.sum(np.log(dens)))

The product prints 0.0; the sum of logs prints a finite −2794.78

Why: 2000 numbers below 1 multiply down past the smallest float and flush to exactly 0.0 — the likelihood is destroyed. The log-sum stays a normal, finite, differentiable number.

methodprinted valueusable?
np.prod(densities)0.0no — underflowed
np.sum(np.log(densities))−2794.78yes — finite

18. What each one costs: Watch the product underflow

Trade off

Comparison matrix

From Watch the product underflow: every row here is a choice with a cost. Fill the printed value column, then say which row you would actually pick and what you give up for it.

methodprinted valueusable?
np.prod(densities)0.0no — underflowed
np.sum(np.log(densities))−2794.78yes — finite

19. The log-likelihood trick

Concept

Take the logarithm. log is strictly increasing, so it never moves the location of a maximum — but by the log-of-a-product rule it turns the product into a sum:

\[ \ell(\theta) = \log \prod_{i=1}^{n} p(x_i;\theta) = \sum_{i=1}^{n} \log p(x_i;\theta) \]

Sums don't underflow, and their derivative is just the sum of the term derivatives. Maximizing ℓ(θ) gives the same θ as maximizing the raw likelihood — far more cheaply.

20. Why a monotone map keeps the argmax

Intuition

Picture the likelihood as a hill. log stretches the vertical axis — it squashes tall values and drops everything below 1 into negatives — but it never reorders heights: if a > b then log a > log b.

So the tallest point stays the tallest point. Its height changes; its horizontal position — the θ we actually want — does not. That single fact is the entire license for the trick.

21. Predict the next row: The log-likelihood peaks at the same place

Pattern

Predict first

The table runs: 3.0 | −20.8967 | far left · 4.0 | −17.8967 | climbing · 5.0 | −16.8967 | PEAK = sample mean · 6.0 | −17.8967 | descending

In The log-likelihood peaks at the same place, given the rows so far: what is the next one — the row where μ is 7.0?

Correct: 7.0 | −20.8967 | far right

μℓ(μ) at σ²=4note
3.0−20.8967far left
4.0−17.8967climbing
5.0−16.8967PEAK = sample mean
6.0−17.8967descending
7.0−20.8967far right

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The log-likelihood is a downward parabola in μ: −16.8967 at μ=5 beats −17.8967 at μ=4 and 6, and −20.8967 at μ=3 and 7.

22. The log-likelihood peaks at the same place

Worked example

Concrete check on Gaussian data [2,4,4,4,5,5,7,9] with variance fixed at σ² = 4. Sweep the mean μ and print the log-likelihood — the peak should land at the sample mean 5:

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
def ll(mu, var=4.0):
    n = len(data)
    return -0.5*n*np.log(2*np.pi*var) - (1/(2*var))*np.sum((data-mu)**2)
for mu in [3., 4., 5., 6., 7.]:
    print(mu, round(ll(mu), 4))

ℓ(μ) rises to a single peak at μ = 5, then falls symmetrically

Why: The log-likelihood is a downward parabola in μ: −16.8967 at μ=5 beats −17.8967 at μ=4 and 6, and −20.8967 at μ=3 and 7. One smooth hump, maximum at the sample mean.

μℓ(μ) at σ²=4note
3.0−20.8967far left
4.0−17.8967climbing
5.0−16.8967PEAK = sample mean
6.0−17.8967descending
7.0−20.8967far right

23. Which is which, by ℓ(μ) at σ²=4

Discrimination

Sort into buckets

Sort these by ℓ(μ) at σ²=4, from memory, without looking back at The log-likelihood peaks at the same place. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

−20.8967
3.0; 7.0
−17.8967
4.0; 6.0
−16.8967
5.0
g1
ℓ(μ) at σ²=4 is "−20.8967" for 3.0, 7.0 — that is what the table on "The log-likelihood peaks at the same…" records, and it is the single property separating this group from the rest.
g2
ℓ(μ) at σ²=4 is "−17.8967" for 4.0, 6.0 — that is what the table on "The log-likelihood peaks at the same…" records, and it is the single property separating this group from the rest.
g3
ℓ(μ) at σ²=4 is "−16.8967" for 5.0 — that is what the table on "The log-likelihood peaks at the same…" records, and it is the single property separating this group from the rest.

24. Something is wrong here: does taking the log move the answer?

Anomaly

Predict first

A student writes this, and it looks reasonable:

log transforms the function and shrinks its value, so the maximizing θ must shift too.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).

log is strictly increasing, so it preserves the location of every maximum.

Why: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).

25. Trap: does taking the log move the answer?

Trap

The trap

log transforms the function and shrinks its value, so the maximizing θ must shift too.

Treat argmax ℓ(θ) as a DIFFERENT θ from argmax ∏ p(xᵢ;θ)

Why: Confuses changing the HEIGHT of the peak (log does shrink it, often to a negative number) with changing its LOCATION (log never touches that).

Conclusion: 'I must maximize the raw product to be safe'

Why: So you fight underflow and the product rule for nothing — and on large n the product is literally 0.0, so there is no peak left to find.

The fix

log is strictly increasing, so it preserves the location of every maximum.

argmax_θ ℓ(θ) = argmax_θ ∏ p(xᵢ;θ)

Why: A monotone transform reorders nothing: the tallest point before is the tallest point after. Only the height changes.

So maximize the SUM of logs — same θ, no underflow, clean derivative

Why: The sweep on the previous slide confirmed it: the log-likelihood peaked at exactly μ = 5, the same place the raw likelihood does.

26. MLE for the Gaussian

Section

Part 2 of 6 — mean & variance, every step

27. Our running dataset

Concept

One dataset carries the whole Gaussian derivation — the same eight numbers we just swept over. Assume they are i.i.d. draws from N(μ, σ²) and we want the best μ and σ².

i12345678
xᵢ24445579

Keep n = 8 and Σxᵢ = 40 in view — they are the only summaries the answers depend on.

28. Write the Gaussian log-likelihood

Worked example

One point's density is p(xᵢ) = (2πσ²)^(−1/2) · exp(−(xᵢ−μ)²/2σ²). Take the log of a single term first:

log of one point

Why: log turns the exp into its exponent and the (2πσ²)^(−1/2) into −½log(2πσ²).

\[ \log p(x_i;\mu,\sigma^2) = -\tfrac{1}{2}\log(2\pi\sigma^2) - \frac{(x_i-\mu)^2}{2\sigma^2} \]

Sum over all n points

Why: By the log trick, ℓ is the sum of the per-point logs. The constant term repeats n times.

\[ \ell(\mu,\sigma^2) = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(x_i-\mu)^2 \]

29. Decode the notation: Write the Gaussian log-likelihood

Notation

Annotate

From Write the Gaussian log-likelihood — read this one piece at a time. What is each part doing?

On: \( \log p(x_i;\mu,\sigma^2) = -\tfrac{1}{2}\log(2\pi\sigma^2) - \frac{(x_i-\mu)^2}{2\sigma^2} \)

  • log turns the exp into its exponent and the (2πσ²)^(−1/2) into −½log(2πσ²).
  • By the log trick, ℓ is the sum of the per-point logs. The constant term repeats n times.

30. What has to happen first: Differentiate in μ — step by step

Ranking

Put in order

Put the moves of Differentiate in μ — step by step into the order they have to happen.

  1. Only the sum-of-squares term contains μ
  2. Differentiate Σ(xᵢ−μ)² with the chain rule
  3. The two minus signs and the 2 combine into +1/σ²

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The −(n/2)log(2πσ²) term has no μ in it, so its μ-derivative is 0.

31. Differentiate in μ — step by step

Worked example

Only the sum-of-squares term contains μ

Why: The −(n/2)log(2πσ²) term has no μ in it, so its μ-derivative is 0. Focus on the second term.

Differentiate Σ(xᵢ−μ)² with the chain rule

Why: d/dμ (xᵢ−μ)² = 2(xᵢ−μ)·(−1) = −2(xᵢ−μ). The −1 is the inner derivative of (xᵢ−μ).

\[ \frac{\partial \ell}{\partial \mu} = -\frac{1}{2\sigma^2}\sum_i 2(x_i-\mu)(-1) = \frac{1}{\sigma^2}\sum_i (x_i-\mu) \]

The two minus signs and the 2 combine into +1/σ²

Why: (−1/2σ²)·(−2) = +1/σ². What survives is (1/σ²)·Σ(xᵢ−μ) — clean and linear in μ.

32. Draw the shape of it: Differentiate in μ — step by step

Blank canvas

Draw it

Draw what Differentiate in μ — step by step just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

33. Plan first: Set ∂ℓ/∂μ = 0 → the sample mean

Step zero

Discussion prompt

Set ∂ℓ/∂μ = 0 → the sample mean — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: At the maximum the derivative is zero

Answer:

  1. At the maximum the derivative is zero
  2. Split the sum: Σxᵢ − nμ = 0
  3. Plug in our numbers: μ = 40/8 = 5.0

34. Set ∂ℓ/∂μ = 0 → the sample mean

Worked example

At the maximum the derivative is zero

Why: σ² > 0, so the 1/σ² factor can never be zero — the bracket must vanish.

\[ \frac{1}{\sigma^2}\sum_i (x_i-\mu) = 0 \;\Longrightarrow\; \sum_i (x_i-\mu) = 0 \]

Split the sum: Σxᵢ − nμ = 0

Why: Σ(xᵢ−μ) = Σxᵢ − Σμ, and Σμ over n points is nμ. Solve for μ.

\[ \sum_i x_i - n\mu = 0 \;\Longrightarrow\; \mu_{\text{MLE}} = \frac{1}{n}\sum_i x_i \]

Plug in our numbers: μ = 40/8 = 5.0

Why: The MLE of a Gaussian mean is nothing but the SAMPLE MEAN. Σxᵢ = 40, n = 8, so μ_MLE = 5.0 — matching the sweep's peak exactly.

quantityvalue
n8
Σxᵢ40
μ_MLE = Σxᵢ / n5.0

35. Is that stationary point a maximum?

Concept

Setting a derivative to zero only finds a flat point — we owe a check that it is a peak, not a valley. Differentiate ∂ℓ/∂μ = (1/σ²)Σ(xᵢ−μ) once more:

\[ \frac{\partial^2 \ell}{\partial \mu^2} = \frac{1}{\sigma^2}\sum_i (-1) = -\frac{n}{\sigma^2} < 0 \]

The second derivative is −n/σ², strictly negative for all μ. So ℓ(μ) is concave — a single downward hump — and the stationary point μ = x̄ is the global maximum, exactly as the sweep in Part 1 showed.

36. Complete the line: Differentiate in σ² — step by step

Fill the middle

Fill in the blanks

From Differentiate in σ² — step by step — finish the line. Write what belongs on the right of the equals sign before you look.

\frac-\frac{n}{2v} + \frac{1}{2v^2}\sum_i (x_i-\mu)^2___ = ___

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. −(1/2v)·S = −(S/2)·v⁻¹, and d/dv(v⁻¹) = −v⁻², so the derivative is +(S/2)v⁻² = S/(2v²).

37. Differentiate in σ² — step by step

Worked example

Now hold μ = 5 and differentiate ℓ with respect to the variable v = σ². Both terms contain v, so differentiate each:

Term 1: d/dv [−(n/2)log(2πv)] = −n/(2v)

Why: d/dv log(2πv) = 1/v, times −n/2.

Term 2: d/dv [−(1/2v)·S] = +S/(2v²), where S = Σ(xᵢ−μ)²

Why: −(1/2v)·S = −(S/2)·v⁻¹, and d/dv(v⁻¹) = −v⁻², so the derivative is +(S/2)v⁻² = S/(2v²).

\[ \frac{\partial \ell}{\partial v} = -\frac{n}{2v} + \frac{1}{2v^2}\sum_i (x_i-\mu)^2 \]

38. Work backwards from the answer: Differentiate in σ² — step by step

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Term 2: d/dv [−(1/2v)·S] = +S/(2v²), where S = Σ(xᵢ−μ)²

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Now hold μ = 5 and differentiate ℓ with respect to the variable v = σ². Both terms contain v, so differentiate each:

39. Plan first: Set ∂ℓ/∂σ² = 0 → the 1/n variance

Step zero

Discussion prompt

Set ∂ℓ/∂σ² = 0 → the 1/n variance — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Set to zero and multiply through by 2v²

Answer:

  1. Set to zero and multiply through by 2v²
  2. Solve for v = σ²
  3. Plug in: deviations are [−3,−1,−1,−1,0,0,2,4]
  4. σ²_MLE = 32/8 = 4.0, so σ_MLE = 2.0

40. Set ∂ℓ/∂σ² = 0 → the 1/n variance

Worked example

Set to zero and multiply through by 2v²

Why: Clearing the denominators turns the equation into −nv + Σ(xᵢ−μ)² = 0.

\[ -\frac{n}{2v} + \frac{S}{2v^2} = 0 \;\Longrightarrow\; -nv + S = 0 \]

Solve for v = σ²

Why: v = S/n = the average squared deviation from the MLE mean. Note the divisor is n, not n−1.

\[ \sigma^2_{\text{MLE}} = \frac{1}{n}\sum_i (x_i-\mu_{\text{MLE}})^2 \]

Plug in: deviations are [−3,−1,−1,−1,0,0,2,4]

Why: xᵢ − 5 for our data. Square them: [9,1,1,1,0,0,4,16], which sum to S = 32.

σ²_MLE = 32/8 = 4.0, so σ_MLE = 2.0

Why: The variance MLE is 4.0 and the standard-deviation MLE is its root, 2.0 — matching the σ² = 4 we fixed during the sweep.

41. Say it in words: Set ∂ℓ/∂σ² = 0 → the 1/n variance

Translation

\( \sigma^2_{\text{MLE}} = \frac{1}{n}\sum_i (x_i-\mu_{\text{MLE}})^2 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

42. Full trace of the variance sum

Worked example

Every term of S = Σ(xᵢ−5)², in one table — this is the full trace behind σ²_MLE = 4.0:

xᵢxᵢ − 5(xᵢ − 5)²
2−39
4−11
4−11
4−11
500
500
724
9416

S = 9+1+1+1+0+0+4+16 = 32 → σ²_MLE = 32/8 = 4.0

Why: Sum the last column and divide by n = 8. The sample standard deviation the MLE reports is √4 = 2.0.

43. Watch it run: Full trace of the variance sum

Pattern

Step through it

Step through Full trace of the variance sum one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: xᵢ is 2
  2. Step 2: xᵢ is 4
  3. Step 3: xᵢ is 4
  4. Step 4: xᵢ is 4
  5. Step 5: xᵢ is 5
  6. Step 6: xᵢ is 5
  7. Step 7: xᵢ is 7
  8. Step 8: xᵢ is 9

44. Why μ then σ² is legitimate

Concept

We solved for μ first, then plugged that μ into the variance equation. That is valid because ∂ℓ/∂μ = 0 gives the same μ = x̄ for every fixed σ² — the best mean never depends on the variance.

So the joint maximum factors: find μ_MLE, then find the σ² that maximizes ℓ at that mean. No circularity — the two conditions are compatible.

45. Something is wrong here: 1/n versus 1/(n−1)

Anomaly

Predict first

A student writes this, and it looks reasonable:

Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.

The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.

Why: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.

46. Trap: 1/n versus 1/(n−1)

Trap

The trap

Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.

Report σ² = Σ(xᵢ−μ)² / (n−1) = 32/7 ≈ 4.571

Why: That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.

Claim this is 'the MLE'

Why: It is not. Setting ∂ℓ/∂σ² = 0 gave a divisor of n. Calling the n−1 version the MLE is exactly the error the exam probes.

The fix

The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.

σ²_MLE = Σ(xᵢ−μ)² / n = 32/8 = 4.0

Why: This is what maximizing the Gaussian likelihood actually yields — verified in code as np.var(data) with its default ddof=0.

Know both: MLE = ÷n (biased), unbiased = ÷(n−1)

Why: np.var(data, ddof=1) = 4.571 gives the unbiased one. The MLE and the unbiased estimator are two different, both-correct answers to two different questions.

47. Break it on purpose: 1/n versus 1/(n−1)

Break the constraint

Discussion prompt

The rule this trap just fixed:

The MLE divides by n. It is biased low, and that is the honest price of reusing the data to estimate μ.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

That is Bessel's UNBIASED sample variance — a different estimator, deliberately corrected for bias.

48. Guess the shape of the answer: Confirm both variances in NumPy

Estimation

Predict first

np.var defaults to ddof=0 — the 1/n MLE. Pass ddof=1 for the 1/(n−1) unbiased version. Both, side by side:

Commit before you compute: what does Confirm both variances in NumPy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: np.var(data) = 4.0 matches our by-hand MLE; ddof=1 = 4.571

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The default np.var IS the MLE. The Bessel-corrected 32/7 = 4.5714 only appears when you ask for ddof=1.

49. Confirm both variances in NumPy

Worked example

np.var defaults to ddof=0 — the 1/n MLE. Pass ddof=1 for the 1/(n−1) unbiased version. Both, side by side:

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
var_mle = ((data - mu)**2).mean()   # divides by n
print(mu, var_mle)
print(np.var(data))                 # ddof=0 -> MLE
print(np.var(data, ddof=1))         # ddof=1 -> unbiased

np.var(data) = 4.0 matches our by-hand MLE; ddof=1 = 4.571

Why: The default np.var IS the MLE. The Bessel-corrected 32/7 = 4.5714 only appears when you ask for ddof=1.

callvaluewhich estimator
((data-mu)**2).mean()4.0MLE (÷n)
np.var(data)4.0MLE (ddof=0)
np.var(data, ddof=1)4.571unbiased (÷(n−1))

50. MLE for the Bernoulli

Section

Part 3 of 6 — and Poisson, Exponential

51. The coin, finished

Intuition

Back to Part 1's coin. We saw by brute force that p = 0.7 maximized the likelihood of 7-heads-in-10. Now we get the same 0.7 from calculus — and see it holds for any k and n, not just this one sample.

The lesson: the sweep and the derivative are two views of the same peak. Once you trust the recipe, you skip the sweep and go straight to setting dℓ/dp = 0.

52. What has to happen first: Derive the Bernoulli MLE

Ranking

Put in order

Put the moves of Derive the Bernoulli MLE into the order they have to happen.

  1. Log-likelihood in p
  2. Differentiate: dℓ/dp = k/p − (n−k)/(1−p)
  3. Set to zero and cross-multiply

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Σ[xᵢ log p + (1−xᵢ)log(1−p)] collapses because Σxᵢ = k and Σ(1−xᵢ) = n−k.

53. Derive the Bernoulli MLE

Worked example

Data are 0/1 outcomes with k successes in n trials. One point's probability is p^{xᵢ}(1−p)^{1−xᵢ}. Its log summed over all points is:

Log-likelihood in p

Why: Σ[xᵢ log p + (1−xᵢ)log(1−p)] collapses because Σxᵢ = k and Σ(1−xᵢ) = n−k.

\[ \ell(p) = k\log p + (n-k)\log(1-p) \]

Differentiate: dℓ/dp = k/p − (n−k)/(1−p)

Why: d/dp log p = 1/p; d/dp log(1−p) = −1/(1−p), and the −1 flips the sign of the second term.

Set to zero and cross-multiply

Why: k/p = (n−k)/(1−p) ⇒ k(1−p) = (n−k)p ⇒ k − kp = np − kp ⇒ k = np.

\[ p_{\text{MLE}} = \frac{k}{n} \]

54. Restore the missing line: Bernoulli MLE in code

Fill the middle

Fill in the blanks

From Bernoulli MLE in code — one line has had its right-hand side removed. Put it back.

import numpy as np
b = np.array([1,0,1,1,0,1,1,0,1,1])
k = b.sum()
n = len(b)
p_mle = b.mean() # = k/n
print(k, n, p_mle)

Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The mean of a 0/1 array IS the success fraction.

55. Bernoulli MLE in code

Worked example

The success fraction, on a 10-outcome sample with 7 ones — confirming p_MLE = k/n:

import numpy as np
b = np.array([1,0,1,1,0,1,1,0,1,1])
k = b.sum()
n = len(b)
p_mle = b.mean()    # = k/n
print(k, n, p_mle)

k = 7, n = 10 → p_MLE = 0.7

Why: The mean of a 0/1 array IS the success fraction. It matches the coin experiment from Part 1 — the same k/n falls out whether you sweep candidates or take the derivative.

k (successes)n (trials)p_MLE = k/n
7100.7

56. Where does each piece belong: Lesson 8: Maximum Likelihood Estimation

Sorting

Sort into buckets

These are the pieces of Lesson 8: Maximum Likelihood Estimation, out of order. Put each one back under the part of the lesson it belongs to.

The likelihood principle
A coin told you something; The likelihood function; Likelihood is not probability of θ
MLE for the Gaussian
Our running dataset; Write the Gaussian log-likelihood; Differentiate in μ — step by step
MLE for the Bernoulli
The coin, finished; Derive the Bernoulli MLE; Bernoulli MLE in code
s1
The likelihood principle is where Lesson 8: Maximum Likelihood Estimation puts A coin told you something, The likelihood function, Likelihood is not probability of θ. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
MLE for the Gaussian is where Lesson 8: Maximum Likelihood Estimation puts Our running dataset, Write the Gaussian log-likelihood, Differentiate in μ — step by step. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
MLE for the Bernoulli is where Lesson 8: Maximum Likelihood Estimation puts The coin, finished, Derive the Bernoulli MLE, Bernoulli MLE in code. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

57. Rebuild the recipe: The MLE recipe

Ranking

Put in order

These are the steps of The MLE recipe, scrambled. Put them back in order before the next slide shows you.

  1. Write the model p(x; θ) for one data point
  2. Log & sum: ℓ(θ) = Σ log p(xᵢ; θ) — the log trick turns the product into a sum
  3. Differentiate ℓ with respect to each parameter
  4. Set to zero and solve for the parameter (check it is a max, not a min)
  5. Plug in the data summaries (Σxᵢ, k, Σxᵢ², …) to get a number

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

58. The MLE recipe

Pattern

  1. Write the model p(x; θ) for one data point
  2. Log & sum: ℓ(θ) = Σ log p(xᵢ; θ) — the log trick turns the product into a sum
  3. Differentiate ℓ with respect to each parameter
  4. Set to zero and solve for the parameter (check it is a max, not a min)
  5. Plug in the data summaries (Σxᵢ, k, Σxᵢ², …) to get a number

Every MLE today — Gaussian, Bernoulli, Poisson, Exponential — is this exact five-step loop. Learn the loop, not the formulas.

59. Restore the missing line: Poisson and Exponential fall out too

Fill the middle

Fill in the blanks

From Poisson and Exponential fall out too — one line has had its right-hand side removed. Put it back.

import numpy as np
pois = np.array([2,3,1,4,2,3])
ex = np.array([0.5,1.5,2.0,1.0])
print('Poisson lambda =', pois.mean()) # MLE = sample mean
print('Exp lambda =', 1/ex.mean()) # MLE = 1/sample mean

Why: ex is what everything below it consumes, so the wrong expression here fails later and somewhere else. d/dλ Σ(xᵢ log λ − λ) = Σxᵢ/λ − n = 0 ⇒ λ = Σxᵢ/n.

60. Poisson and Exponential fall out too

Worked example

Apply the same recipe to two more families and confirm the closed forms in code:

import numpy as np
pois = np.array([2,3,1,4,2,3])
ex = np.array([0.5,1.5,2.0,1.0])
print('Poisson lambda =', pois.mean())      # MLE = sample mean
print('Exp     lambda =', 1/ex.mean())       # MLE = 1/sample mean

Poisson: ℓ = Σ(xᵢ log λ − λ) ⇒ λ_MLE = mean = 2.5

Why: d/dλ Σ(xᵢ log λ − λ) = Σxᵢ/λ − n = 0 ⇒ λ = Σxᵢ/n. For [2,3,1,4,2,3], that is 15/6 = 2.5.

Exponential: ℓ = Σ(log λ − λxᵢ) ⇒ λ_MLE = 1/mean = 0.8

Why: d/dλ = n/λ − Σxᵢ = 0 ⇒ λ = n/Σxᵢ = 1/x̄. For mean 1.25, λ = 0.8.

familyMLEvalue
Poissonλ = sample mean2.5
Exponentialλ = 1 / sample mean0.8

61. Inspect it line by line: Poisson and Exponential fall out too

Error analysis

Annotate

Walk the callouts on Poisson and Exponential fall out too. Each one is a place this is easy to get subtly wrong.

  • d/dλ Σ(xᵢ log λ − λ) = Σxᵢ/λ − n = 0 ⇒ λ = Σxᵢ/n. For [2,3,1,4,2,3], that is 15/6 = 2.5.
  • d/dλ = n/λ − Σxᵢ = 0 ⇒ λ = n/Σxᵢ = 1/x̄. For mean 1.25, λ = 0.8.

62. Rule out three: Check yourself — the log trick

Elimination

Eliminate the wrong options

Maximizing ℓ(θ) = Σ log p(xᵢ;θ) gives the same θ as maximizing ∏ p(xᵢ;θ) because…

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. log is strictly increasing, so it preserves the location of the maximum
  • B. log(∏) happens to equal ∏ log for probabilities
  • C. the likelihood and log-likelihood have the same numerical value
  • D. taking log removes the dependence on θ

Survives elimination: A

Why: A monotonically increasing transform changes the height of the peak but never its location, so the argmax is unchanged. As a bonus, log turns the product into a sum, which is why we use it.

63. Check yourself — the log trick

Check

Why is it always safe to maximize the log-likelihood instead of the likelihood?

Check your understanding

Maximizing ℓ(θ) = Σ log p(xᵢ;θ) gives the same θ as maximizing ∏ p(xᵢ;θ) because…

  • A. log is strictly increasing, so it preserves the location of the maximum (correct)
  • B. log(∏) happens to equal ∏ log for probabilities
  • C. the likelihood and log-likelihood have the same numerical value
  • D. taking log removes the dependence on θ

Answer: A

Why: A monotonically increasing transform changes the height of the peak but never its location, so the argmax is unchanged. As a bonus, log turns the product into a sum, which is why we use it.

Why B tempts people
log(∏) = Σ log, not ∏ log. The product becomes a SUM of logs — that's the useful identity, stated wrong here.
Why C tempts people
The values differ — log shrinks them, often making them negative (our peak ℓ was −16.8967). What's preserved is the LOCATION of the max, not the value.
Why D tempts people
The log-likelihood still depends on θ — if it didn't, there would be nothing left to maximize.

64. Answer it before you see the options: Check yourself — the MLE variance

Prediction

Predict first

For Gaussian data, the maximum-likelihood estimate of the variance divides the sum of squared deviations by…

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: n (and is biased low)

Why: Setting ∂ℓ/∂σ² = 0 yields σ²_MLE = (1/n)Σ(xᵢ−μ)². It is biased low because the same data estimates both μ and σ. On our data that is 32/8 = 4.0.

65. Check yourself — the MLE variance

Check

Careful with the denominator.

Check your understanding

For Gaussian data, the maximum-likelihood estimate of the variance divides the sum of squared deviations by…

  • A. n (and is biased low) (correct)
  • B. n−1 (and is unbiased)
  • C. n+1
  • D. √n

Answer: A

Why: Setting ∂ℓ/∂σ² = 0 yields σ²_MLE = (1/n)Σ(xᵢ−μ)². It is biased low because the same data estimates both μ and σ. On our data that is 32/8 = 4.0.

Why B tempts people
1/(n−1) is the unbiased (Bessel-corrected) sample variance — 32/7 ≈ 4.571 here — a deliberately bias-corrected estimator, not the MLE.
Why C tempts people
n+1 minimizes the mean squared error of the variance estimate in some settings, but it is neither the MLE nor the unbiased estimator.
Why D tempts people
√n appears in the standard ERROR of the mean, not in the variance estimator's denominator.

66. Losses are negative log-likelihoods

Section

Part 4 of 6 — MSE & cross-entropy

67. The reveal

Intuition

Here is the payoff. Every loss you have minimized — squared error for regression, cross-entropy for classification — is not an arbitrary choice someone handed you. Each one is the negative log-likelihood of a probabilistic model.

Pick a noise model, write its likelihood, negate the log — and out drops the loss. Change the noise model and the loss changes with it. Let's watch MSE fall out of a Gaussian.

68. Picture it first: A little bell over every point

Picture it

Figure (svg): A rising straight regression line with small vertical bell curves centered on the line at four x positions, and data points scattered near each bell.

Each target is a draw from a Gaussian centered on the line. Maximizing that = minimizing squared distance to the line.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture the regression line, and at each xᵢ a small Gaussian bell standing up off the line. The model says the true yᵢ was drawn from that bell — most likely near the line, occasionally far.

69. A little bell over every point

Intuition

Picture the regression line, and at each xᵢ a small Gaussian bell standing up off the line. The model says the true yᵢ was drawn from that bell — most likely near the line, occasionally far.

Figure (svg): A rising straight regression line with small vertical bell curves centered on the line at four x positions, and data points scattered near each bell.

Each target is a draw from a Gaussian centered on the line. Maximizing that = minimizing squared distance to the line.

The farther a point sits from the line, the lower its bell-height, the smaller its likelihood. Maximizing total likelihood pulls the line to where the points cluster — which is exactly least squares.

70. The Gaussian-noise regression model

Concept

Assume each target is the model's prediction plus independent Gaussian noise:

\[ y_i = f(x_i) + \varepsilon_i, \qquad \varepsilon_i \sim \mathcal{N}(0,\sigma^2) \]

Then yᵢ is Gaussian centered at f(xᵢ) with variance σ², so its density is (2πσ²)^{−1/2} exp(−(yᵢ−f(xᵢ))²/2σ²). Now run the MLE recipe on that.

71. Plan first: From Gaussian likelihood to MSE

Step zero

Discussion prompt

From Gaussian likelihood to MSE — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Negative log-likelihood of the targets

Answer:

  1. Negative log-likelihood of the targets
  2. The log term is constant in the model parameters
  3. Drop the positive scale 1/2σ²
  4. This is the MSE objective — Lesson 7's OLS

72. From Gaussian likelihood to MSE

Worked example

Negative log-likelihood of the targets

Why: Same log as the Gaussian in Part 2, but centered at f(xᵢ) instead of μ, and negated.

\[ -\ell = \frac{n}{2}\log(2\pi\sigma^2) + \frac{1}{2\sigma^2}\sum_i \big(y_i - f(x_i)\big)^2 \]

The log term is constant in the model parameters

Why: (n/2)log(2πσ²) has no f in it, so it does not move the argmin over f's parameters — drop it.

Drop the positive scale 1/2σ²

Why: A positive constant multiplier never moves an argmin. What remains is exactly the sum of squared errors.

\[ \arg\min \;(-\ell) = \arg\min \sum_i \big(y_i - f(x_i)\big)^2 \]

This is the MSE objective — Lesson 7's OLS

Why: Least squares was never a convention. Its normal equations gave w = [2.2, 0.6] because MSE IS the MLE under Gaussian noise.

73. Draw the shape of it: From Gaussian likelihood to MSE

Blank canvas

Draw it

Draw what From Gaussian likelihood to MSE just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

74. Restore the missing line: The tie-back to Lesson 7, in code

Fill the middle

Fill in the blanks

From The tie-back to Lesson 7, in code — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1.,2.,3.,4.,5.])
y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
sse = np.sum((y - X @ w)**2)
print(w, sse, sse/5)

Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Minimizing MSE gave the same w = [2.2, 0.6] as Lesson 7, and the noise-variance MLE is just the mean squared residual, 0.48.

75. The tie-back to Lesson 7, in code

Worked example

Re-fit the 5-student line from Lesson 7 and read off the noise-variance MLE σ² = SSE/n — the least-squares residuals ARE the Gaussian MLE:

import numpy as np
x = np.array([1.,2.,3.,4.,5.])
y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
sse = np.sum((y - X @ w)**2)
print(w, sse, sse/5)

w = [2.2, 0.6], SSE = 2.4, σ²_MLE = 2.4/5 = 0.48

Why: Minimizing MSE gave the same w = [2.2, 0.6] as Lesson 7, and the noise-variance MLE is just the mean squared residual, 0.48.

quantityvalue
w (OLS / MLE)[2.2, 0.6]
SSE = Σ(y − ŷ)²2.4
σ²_MLE = SSE / n0.48

76. Cross-entropy = MLE for classification

Concept

Now the classification side. Model P(yᵢ=1|xᵢ) = pᵢ (e.g. pᵢ = σ(wᵀxᵢ+b)). Each label is Bernoulli, so its log-likelihood is the Bernoulli log from Part 3:

\[ \ell = \sum_i \big[\, y_i \log p_i + (1-y_i)\log(1-p_i)\,\big] \]

The negative of that, averaged, is exactly the binary cross-entropy loss. So minimizing BCE literally IS maximizing the Bernoulli likelihood of the labels.

77. Predict the next row: BCE equals the negative log-likelihood

Pattern

Predict first

The table runs: (1, 0.9) | −log(0.9) | 0.1054 · (0, 0.2) | −log(0.8) | 0.2231 · (1, 0.7) | −log(0.7) | 0.3567

In BCE equals the negative log-likelihood, given the rows so far: what is the next one — the row where point (y, p) is mean of the three?

Correct: mean of the three | = F.binary_cross_entropy | 0.2284

point (y, p)−log-likelihood termvalue
(1, 0.9)−log(0.9)0.1054
(0, 0.2)−log(0.8)0.2231
(1, 0.7)−log(0.7)0.3567
mean of the three= F.binary_cross_entropy0.2284

Why: The relationship between the columns, not the individual numbers, is what generates the next row. y=1 point contributes −log(0.9)=0.1054; y=0 point contributes −log(1−0.2)=−log(0.8)=0.2231; y=1 point contributes −log(0.7)=0.3567.

78. BCE equals the negative log-likelihood

Worked example

Confirm numerically that a hand-computed mean NLL matches torch's BCE for predictions [0.9, 0.2, 0.7], labels [1, 0, 1]:

import numpy as np, torch
import torch.nn.functional as F
y = np.array([1.,0.,1.]); p = np.array([0.9,0.2,0.7])
nll = -np.mean(y*np.log(p) + (1-y)*np.log(1-p))
tb = F.binary_cross_entropy(torch.tensor(p), torch.tensor(y)).item()
print(round(nll,4), round(tb,4))

Per-point −log terms: 0.1054, 0.2231, 0.3567

Why: y=1 point contributes −log(0.9)=0.1054; y=0 point contributes −log(1−0.2)=−log(0.8)=0.2231; y=1 point contributes −log(0.7)=0.3567.

mean NLL = 0.6852/3 = 0.2284 = torch BCE

Why: The three terms sum to 0.6852; divided by n=3 gives 0.2284, exactly what F.binary_cross_entropy returns. BCE is the averaged negative Bernoulli log-likelihood — same formula.

point (y, p)−log-likelihood termvalue
(1, 0.9)−log(0.9)0.1054
(0, 0.2)−log(0.8)0.2231
(1, 0.7)−log(0.7)0.3567
mean of the three= F.binary_cross_entropy0.2284

79. Fill in: −log-likelihood term for BCE equals the negative log-likelihood

Comparison

Comparison matrix

From BCE equals the negative log-likelihood: refill the −log-likelihood term column from what you know. The rest of the table is as it appeared.

point (y, p)−log-likelihood termvalue
(1, 0.9)−log(0.9)0.1054
(0, 0.2)−log(0.8)0.2231
(1, 0.7)−log(0.7)0.3567
mean of the three= F.binary_cross_entropy0.2284

80. Something is wrong here: maximize or minimize?

Anomaly

Predict first

A student writes this, and it looks reasonable:

MLE maximizes likelihood, and 'the loss is the likelihood', so training should maximize the cross-entropy.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This silently drops the negative sign.

The loss is the negative log-likelihood, so maximizing likelihood = minimizing loss.

Why: This silently drops the negative sign. Cross-entropy is the NEGATED log-likelihood, so climbing it drives the likelihood DOWN.

81. Trap: maximize or minimize?

Trap

The trap

MLE maximizes likelihood, and 'the loss is the likelihood', so training should maximize the cross-entropy.

Run gradient ASCENT on the cross-entropy value

Why: This silently drops the negative sign. Cross-entropy is the NEGATED log-likelihood, so climbing it drives the likelihood DOWN.

Result: predictions get worse each step

Why: You are maximizing −ℓ, i.e. minimizing ℓ — the exact opposite of MLE. Loss curves shoot upward and the model diverges.

The fix

The loss is the negative log-likelihood, so maximizing likelihood = minimizing loss.

Minimize L = −ℓ with gradient DESCENT

Why: max ℓ ⇔ min(−ℓ). The single sign flip is what connects the probability story to the descent loop from Lesson 6.

Loss decreases, likelihood of the data increases

Why: Descending BCE raises P(labels | model) step by step — training and MLE are the same motion, read with opposite signs.

82. How sure are you: Check yourself — MSE's origin

Commit first

Predict first

Minimizing mean squared error is equivalent to maximum likelihood under which assumption?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: The target noise is Gaussian with constant variance

Why: With y = f(x) + ε and ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² plus constants — exactly MSE.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

83. Check yourself — MSE's origin

Check

Where does squared-error loss actually come from?

Check your understanding

Minimizing mean squared error is equivalent to maximum likelihood under which assumption?

  • A. The target noise is Gaussian with constant variance (correct)
  • B. The target noise is Laplace-distributed
  • C. The labels are Bernoulli
  • D. No assumption — MSE is always the likelihood

Answer: A

Why: With y = f(x) + ε and ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² plus constants — exactly MSE.

Why B tempts people
A Laplace noise model gives the MEAN ABSOLUTE error (L1 loss), not squared error — the |·| comes from e^{−|·|}.
Why C tempts people
Bernoulli labels give cross-entropy, the classification loss — not MSE for regression.
Why D tempts people
MSE corresponds to a SPECIFIC model (Gaussian noise). Under other noise models the MLE loss is different.

84. Answer it before you see the options: Check yourself — cross-entropy and signs

Prediction

Predict first

Training a classifier by minimizing binary cross-entropy is equivalent to…

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: maximizing the Bernoulli log-likelihood of the labels

Why: BCE is the averaged NEGATIVE Bernoulli log-likelihood, so min BCE = max ℓ. Our numeric check gave the same 0.2284 for the hand NLL and torch's BCE.

85. Check yourself — cross-entropy and signs

Check

Connect the probability story to the training loop.

Check your understanding

Training a classifier by minimizing binary cross-entropy is equivalent to…

  • A. maximizing the Bernoulli log-likelihood of the labels (correct)
  • B. maximizing the cross-entropy of the labels
  • C. minimizing the log-likelihood of the labels
  • D. minimizing the sum of squared probabilities

Answer: A

Why: BCE is the averaged NEGATIVE Bernoulli log-likelihood, so min BCE = max ℓ. Our numeric check gave the same 0.2284 for the hand NLL and torch's BCE.

Why B tempts people
You minimize cross-entropy, not maximize it — maximizing it (gradient ascent) is the Part-4 trap that drives predictions the wrong way.
Why C tempts people
min(−ℓ) is max(ℓ), not min(ℓ). Minimizing the log-likelihood itself would make the data less probable — backwards.
Why D tempts people
Squared probabilities are the Brier score, a different loss; BCE comes from the Bernoulli LOG-likelihood, not squared terms.

86. MAP: MLE plus a prior

Section

Part 5 of 6 — regularization = prior

87. Add a prior belief

Concept

MLE trusts the data alone. MAP (maximum a posteriori) adds a prior p(θ) — a belief about θ before seeing data — and maximizes the posterior p(θ|data) ∝ p(data|θ)·p(θ).

Take logs, and the product becomes a sum: the log-likelihood plus the log-prior. The prior turns into an additive term on the objective.

\[ \theta_{\text{MAP}} = \arg\max_\theta\; \big[\, \ell(\theta) + \log p(\theta) \,\big] \]

88. What has to happen first: A Gaussian prior IS the ridge penalty

Ranking

Put in order

Put the moves of A Gaussian prior IS the ridge penalty into the order they have to happen.

  1. Put a Gaussian prior on the weights: w ~ N(0, τ²I)
  2. MAP objective = log-likelihood + log-prior
  3. Identify λ = σ²/τ² — this is ridge from Lesson 7

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. A belief that weights are small and centered at 0 — the same belief ridge encodes.

89. A Gaussian prior IS the ridge penalty

Worked example

Put a Gaussian prior on the weights: w ~ N(0, τ²I)

Why: A belief that weights are small and centered at 0 — the same belief ridge encodes.

\[ \log p(w) = -\frac{1}{2\tau^2}\lVert w\rVert^2 + \text{const} \]

MAP objective = log-likelihood + log-prior

Why: Add the Gaussian-noise log-likelihood (→ −½σ⁻²·SSE) to the log-prior (→ −½τ⁻²‖w‖²), then negate to get a loss to minimize.

\[ -\big[\ell + \log p(w)\big] \;\propto\; \lVert Xw - y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \]

Identify λ = σ²/τ² — this is ridge from Lesson 7

Why: The log-prior contributed exactly λ‖w‖². Ridge's penalty was never ad hoc — it is a Gaussian prior on the weights. A Laplace prior would give L1 instead.

90. Decode the notation: A Gaussian prior IS the ridge penalty

Notation

Annotate

From A Gaussian prior IS the ridge penalty — read this one piece at a time. What is each part doing?

On: \( -\big[\ell + \log p(w)\big] \;\propto\; \lVert Xw - y\rVert^2 + \frac{\sigma^2}{\tau^2}\lVert w\rVert^2 \)

  • A belief that weights are small and centered at 0 — the same belief ridge encodes.
  • Add the Gaussian-noise log-likelihood (→ −½σ⁻²·SSE) to the log-prior (→ −½τ⁻²‖w‖²), then negate to get a loss to minimize.
  • The log-prior contributed exactly λ‖w‖². Ridge's penalty was never ad hoc — it is a Gaussian prior on the weights. A Laplace prior would give L1 instead.

91. The prior shrinks the estimate

Concept

A prior centered at 0 pulls the estimate toward 0 — the tighter the prior (smaller τ²), the harder the pull. As τ² → ∞ the prior goes flat and MAP relaxes back to plain MLE.

\[ \mu_{\text{MAP}} = \frac{n\,\bar x/\sigma^2}{\,n/\sigma^2 + 1/\tau^2\,} \]

This is the posterior mean for a Gaussian mean with a N(0, τ²) prior — a precision-weighted blend of the data mean x̄ and the prior mean 0.

92. Teach it back: The prior shrinks the estimate

Explain it

Discussion prompt

Explain The prior shrinks the estimate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A prior centered at 0 pulls the estimate toward 0 — the tighter the prior (smaller τ²), the harder the pull. As τ² → ∞ the prior goes flat and MAP relaxes back to plain MLE.

93. What has to be given first: MAP shrinkage, numerically

Missing information

Discussion prompt

Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

A flat prior recovers the sample mean 5. Tightening the prior toward 0 shrinks the estimate — 4.44, then 3.33. This is exactly what ridge does to regression weights.

94. MAP shrinkage, numerically

Worked example

Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
n, xbar, s2 = len(data), data.mean(), 4.0
for tau2 in [np.inf, 4.0, 1.0]:
    prec = n/s2 + (0.0 if np.isinf(tau2) else 1/tau2)
    print(tau2, round((n*xbar/s2)/prec, 4))

τ²=∞ → 5.0 (pure MLE); τ²=4 → 4.4444; τ²=1 → 3.3333

Why: A flat prior recovers the sample mean 5. Tightening the prior toward 0 shrinks the estimate — 4.44, then 3.33. This is exactly what ridge does to regression weights.

prior variance τ²prior strengthμ_MAP
∞none (flat) → MLE5.0000
4.0moderate4.4444
1.0strong3.3333

95. Work backwards from the answer: MAP shrinkage, numerically

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

τ²=∞ → 5.0 (pure MLE); τ²=4 → 4.4444; τ²=1 → 3.3333

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Estimate the mean of our 8-point data (x̄ = 5, σ² = 4) under a N(0, τ²) prior. Watch a tighter prior pull the estimate down from 5 toward 0:

96. Something is wrong here: MAP is a different kind of estimate?

Anomaly

Predict first

A student writes this, and it looks reasonable:

MAP is Bayesian and MLE is frequentist, so MAP must need a totally separate machinery.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.

MAP is MLE plus one extra term: the log-prior. Same optimization, one added regularizer.

Why: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.

97. Trap: MAP is a different kind of estimate?

Trap

The trap

MAP is Bayesian and MLE is frequentist, so MAP must need a totally separate machinery.

Treat the prior as a bolt-on unrelated to the loss

Why: Misses that after taking the log, the prior is just ANOTHER additive term on the very same objective.

Fail to see ridge/L2 as a prior

Why: So regularization looks like a hack you sprinkle on, rather than a modeling choice with a probabilistic meaning.

The fix

MAP is MLE plus one extra term: the log-prior. Same optimization, one added regularizer.

Objective = ℓ(θ) + log p(θ), minimize its negative

Why: The log-prior sits right beside the log-likelihood. Everything you know about maximizing ℓ still applies.

Gaussian prior ⇒ L2 (ridge); Laplace prior ⇒ L1

Why: Regularization strength λ = σ²/τ² comes straight from the prior's variance. Regularizing IS choosing a prior — nothing bolted on.

98. Which of these survive contact with Lesson 8: Maximum Likelihood Estimation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
You flip a coin 10 times and see 7 heads. Nobody told you the coin's bias p. But your gut already leans: this looks like a coin that favors heads.; A subtle point that trips people up: L(θ) is not the probability that θ is correct. It is the probability the model assigns to the data, read as a function of θ.; Maximum likelihood estimation picks the parameter that maximizes that function — the θ that makes exactly this dataset as probable as possible:
Breaks
log transforms the function and shrinks its value, so the maximizing θ must shift too.; Variance 'should' use the unbiased estimator, so divide the sum of squares by n−1.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 8: Maximum Likelihood Estimation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

99. Rebuild the recipe: Deriving a loss from a model

Ranking

Put in order

These are the steps of Deriving a loss from a model, scrambled. Put them back in order before the next slide shows you.

  1. Write the probabilistic model p(x; θ) or p(y | x; θ)
  2. Form the log-likelihood ℓ(θ) = Σ log p
  3. Negate it → that's your loss L = −ℓ (Gaussian noise → MSE, Bernoulli → BCE)
  4. Add −log p(θ) if you have a prior → the regularizer (Gaussian → L2, Laplace → L1)
  5. Minimize L by gradient descent — the Lesson 6 loop

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

100. Deriving a loss from a model

Pattern

  1. Write the probabilistic model p(x; θ) or p(y | x; θ)
  2. Form the log-likelihood ℓ(θ) = Σ log p
  3. Negate it → that's your loss L = −ℓ (Gaussian noise → MSE, Bernoulli → BCE)
  4. Add −log p(θ) if you have a prior → the regularizer (Gaussian → L2, Laplace → L1)
  5. Minimize L by gradient descent — the Lesson 6 loop

This recipe reduces 'which loss?' and 'which regularizer?' to two modeling questions: what is the noise, and what do I believe about the parameters.

101. Where does it stop working: Deriving a loss from a model

Edge cases

Discussion prompt

Deriving a loss from a model works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

This recipe reduces 'which loss?' and 'which regularizer?' to two modeling questions: what is the noise, and what do I believe about the parameters.

102. Rule out three: Check yourself — regularization as a prior

Elimination

Eliminate the wrong options

Adding an L2 (ridge) penalty λ‖w‖² to a squared-error loss is equivalent to MAP estimation with which prior on the weights?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. A zero-mean Gaussian prior, w ~ N(0, τ²I)
  • B. A Laplace prior, p(w) ∝ e^{−|w|/b}
  • C. A uniform prior over all weights
  • D. A Bernoulli prior on each weight

Survives elimination: A

Why: log of a zero-mean Gaussian prior is −‖w‖²/(2τ²) + const, which is exactly the L2 penalty with λ = σ²/τ². A tighter prior (smaller τ²) means larger λ and stronger shrinkage.

103. Check yourself — regularization as a prior

Check

Tie MAP back to Lesson 7's ridge.

Check your understanding

Adding an L2 (ridge) penalty λ‖w‖² to a squared-error loss is equivalent to MAP estimation with which prior on the weights?

  • A. A zero-mean Gaussian prior, w ~ N(0, τ²I) (correct)
  • B. A Laplace prior, p(w) ∝ e^{−|w|/b}
  • C. A uniform prior over all weights
  • D. A Bernoulli prior on each weight

Answer: A

Why: log of a zero-mean Gaussian prior is −‖w‖²/(2τ²) + const, which is exactly the L2 penalty with λ = σ²/τ². A tighter prior (smaller τ²) means larger λ and stronger shrinkage.

Why B tempts people
A Laplace prior gives the L1 penalty (|w|, lasso), not the squared L2 penalty — its log is −|w|/b, not −w²/(2τ²).
Why C tempts people
A uniform (flat) prior has constant log, adds nothing to the objective, and reduces MAP back to plain MLE — no penalty at all.
Why D tempts people
Bernoulli is a distribution over 0/1 outcomes, not a sensible prior over continuous real-valued weights.

104. Your turn: build the MLE

Section

Part 6 of 6 — the project

105. Project: maximum_likelihood_gaussian(data)

Concept

Write one function that returns the MLE μ and σ of a 1-D dataset, then prove it matches NumPy. You have derived every piece — now assemble it yourself.

#requirementtool
1μ = sample meandata.mean()
2σ² = mean squared deviation (÷n)((data-μ)**2).mean()
3Verify against np.var (ddof=0 = MLE)np.var(data)

Build rules: type every line yourself, run after each line, and confirm np.var defaults to ddof=0 (the 1/n MLE) before you trust it.

106. By analogy: Project: maximum_likelihood_gaussian(data)

Analogy

Discussion prompt

Explain Project: maximum_likelihood_gaussian(data) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Write one function that returns the MLE μ and σ of a 1-D dataset, then prove it matches NumPy. You have derived every piece — now assemble it yourself.

107. Milestone 1 — the mean

Worked example

Your turn: compute μ for [2,4,4,4,5,5,7,9]. Say the value out loud from Σ/n before you print it.

Hint: the MLE mean is just data.mean() — no loop needed. Expect 40/8 = 5.0.

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
print(data.sum(), mu)
checkvalue
Σ data40.0
μ = Σ/85.0

108. Milestone 2 — the variance

Worked example

Your turn: compute σ² as the mean squared deviation from μ. Will it be above or below the unbiased 4.571?

Hint: ((data-mu)**2).mean() divides by n (the MLE), so it comes out below the unbiased estimate. Expect 32/8 = 4.0.

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu = data.mean()
var_mle = ((data - mu)**2).mean()
sigma = var_mle**0.5
print(((data-mu)**2).sum(), var_mle, sigma)
quantityvalue
Σ(x−μ)²32.0
σ²_MLE = 32/84.0
σ_MLE = √42.0

109. Watch it run: Milestone 2 — the variance

Pattern

Step through it

Step through Milestone 2 — the variance one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is Σ(x−μ)²
  2. Step 2: quantity is σ²_MLE = 32/8
  3. Step 3: quantity is σ_MLE = √4

110. Milestone 3 — verify against NumPy

Worked example

Your turn: confirm your σ² matches np.var(data) (default ddof=0 = the 1/n MLE), and see how ddof=1 differs.

Hint: np.var(data) uses ddof=0 (÷n); np.var(data, ddof=1) uses ÷(n−1). Only the first should match your var_mle.

import numpy as np
data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
var_mle = ((data - data.mean())**2).mean()
print(var_mle)               # your MLE
print(np.var(data))          # ddof=0 -> MLE
print(np.var(data, ddof=1))  # ddof=1 -> unbiased
callvaluematches your MLE?
your var_mle4.0—
np.var(data)4.0✓ (MLE)
np.var(data, ddof=1)4.571no (unbiased)

111. Guess the shape of the answer: The full function

Estimation

Predict first

Assemble the three milestones into one function and prove it agrees with NumPy in a single boolean:

Commit before you compute: what does The full function come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Returns (5.0, 2.0); np.allclose vs np.var prints True

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Your from-scratch MLE matches NumPy's default variance exactly.

112. The full function

Worked example

Assemble the three milestones into one function and prove it agrees with NumPy in a single boolean:

import numpy as np

def maximum_likelihood_gaussian(data):
    data = np.asarray(data, dtype=float)
    mu = data.mean()                 # MLE mean
    var = ((data - mu)**2).mean()    # MLE variance (1/n)
    return mu, var**0.5

data = np.array([2.,4.,4.,4.,5.,5.,7.,9.])
mu, sigma = maximum_likelihood_gaussian(data)
print(mu, sigma)                              # 5.0 2.0
print(np.allclose(sigma**2, np.var(data)))    # True

Returns (5.0, 2.0); np.allclose vs np.var prints True

Why: Your from-scratch MLE matches NumPy's default variance exactly. If you see (5.0, 2.0) and True, you have implemented maximum likelihood for the Gaussian.

output linevalue
mu, sigma5.0 2.0
allclose vs np.varTrue

113. What each one costs: The full function

Trade off

Comparison matrix

From The full function: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

output linevalue
mu, sigma5.0 2.0
allclose vs np.varTrue

114. Show it off

Concept

Slides closed, out loud: explain (1) why the log trick never moves the argmax, (2) why MSE and cross-entropy are both negative log-likelihoods, and (3) which prior turns MLE into ridge.

Stretch: prove the Poisson MLE is the sample mean (verify on [2,3,1,4,2,3] → 2.5) and the Exponential MLE is 1/mean (on [0.5,1.5,2,1] → 0.8). Both fall straight out of the five-step recipe.

115. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why the log trick never moves the argmax, (2) why MSE and cross-entropy are both negative log-likelihoods, and (3) which prior turns MLE into ridge.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

116. One principle, every loss

Concept

Step back and see how much collapsed into a single idea. Each row below is the same recipe — model, log, negate — run on a different assumption:

assumptionMLE givesthe loss it becomes
Gaussian noise on yleast squaresMSE
Bernoulli labelssuccess probabilitiescross-entropy (BCE)
Gaussian prior on wshrunk weightsL2 / ridge penalty
Laplace prior on wsparse weightsL1 / lasso penalty

You never again memorize a loss. You state a probabilistic model and the loss is forced on you — that is the whole reason this lesson sits under every later one.

117. Fill in: the loss it becomes for One principle, every loss

Comparison

Comparison matrix

From One principle, every loss: refill the the loss it becomes column from what you know. The rest of the table is as it appeared.

assumptionMLE givesthe loss it becomes
Gaussian noise on yleast squaresMSE
Bernoulli labelssuccess probabilitiescross-entropy (BCE)
Gaussian prior on wshrunk weightsL2 / ridge penalty
Laplace prior on wsparse weightsL1 / lasso penalty

118. Connect it up: Lesson 8: Maximum Likelihood Estimation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The likelihood principle · MLE for the Gaussian · MLE for the Bernoulli · Losses are negative log-likelihoods · MAP: MLE plus a prior · Your turn: build the MLE. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

119. What you can do now

Recap

ideathe one thing to remember
MLEargmax of the likelihood = best-explaining θ
log trickmonotone → same argmax; product → sum
lossloss = −log-likelihood (minimize it)
MLE variancedivide by n, not n−1 (biased low)
MSE / BCEGaussian noise → MSE; Bernoulli → cross-entropy
MAPMLE + log-prior; Gaussian prior → ridge

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 8 (Week 3 — MLE Deep Dive) — Barron · USAAIO Round 2 Preparation, 2026
  2. Bishop, Pattern Recognition and Machine Learning, Ch. 1–4 (Maximum Likelihood, Bayesian priors) — Springer, 2006
  3. PyTorch binary_cross_entropy
  4. NumPy var (ddof)
  5. Every likelihood value, log-likelihood, and coefficient produced by real execution — numpy 2.2.6 + torch 2.7.1, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108