Lesson 36: Initialization & Gradient Clipping

USAAIO Lesson 36, from Week 12, fully worked, on why a deep network lives or dies at step zero. It presents exploding and vanishing gradients as a product of per-layer factors, derives gradient norm clipping and contrasts it with per-component clamping, and builds the forward-variance recursion for Var(a_L) one step at a time. It derives Xavier initialization from variance preservation and He initialization from the fact that ReLU zeros half its inputs, with no algebra skipped, then shows that the backward pass needs the SAME scale. It proves the zero-initialization symmetry problem by hand and ends with a 50-layer signal-propagation experiment in which only He survives. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6. The lesson runs to 62 slides.

Subject: Machine Learning · 112 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Initialization & Gradient Clipping

Title

USAAIO · Lesson 36 · Week 12

A deep network lives or dies at step zero. We derive gradient norm clipping move by move, build the variance recursion for a deep net one layer at a time, get Xavier and He from it with no skipped algebra, and run a 50-layer experiment where only He survives.

2. By the end of this lesson you can

Objectives

  1. Explain exploding / vanishing gradients as a product of per-layer factors, and why depth makes it worse
  2. Derive gradient norm clipping g ← g·min(1, τ/‖g‖) and prove it preserves the descent direction while per-component clamping does not
  3. Build the forward recursion Var(a_L) layer by layer and read off when the signal explodes, vanishes, or holds
  4. Derive Xavier Var(W)=1/n from linear variance preservation, then He Var(W)=2/n from ReLU zeroing half the units — every step shown
  5. Prove the backward pass needs the same scale, prove zero init can never learn (broken symmetry), and run a 50-layer experiment where only He keeps std ≈ O(1)

3. What survived from ROC, PR Curves & Calibration?

Warm-up

Discussion prompt

Before we open Lesson 36: Initialization & Gradient Clipping: without looking back, what was the main idea of ROC, PR Curves & Calibration, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the ROC curve and AUC, the precision-recall curve for imbalanced data, why ROC-AUC misleads under heavy imbalance, model calibration, the Expected Calibration Error, and temperature scaling. Build an ROC curve from scratch and fix an overconfident model's calibration.

4. The problem: depth multiplies

Section

Part 1 of 7

5. A deep net is a chain of maps

Concept

A feed-forward net stacks L layers: each takes the previous activation a, multiplies by a weight matrix W, and applies a nonlinearity. The output is a long composition of these maps.

\[ a^{(\ell)} = \phi\!\big(W^{(\ell)} a^{(\ell-1)}\big), \qquad \ell = 1, \dots, L \]

Both the forward signal and the backward gradient pass through every layer. Anything each layer does to the scale gets applied L times — that repetition is the whole story of this lesson.

6. Break it if you can: A deep net is a chain of maps

Counterexample

Discussion prompt

A feed-forward net stacks L layers: each takes the previous activation a, multiplies by a weight matrix W, and applies a nonlinearity. The output is a long composition of these maps.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Both the forward signal and the backward gradient pass through every layer. Anything each layer does to the scale gets applied L times — that repetition is the whole story of this lesson.

7. The gradient is a product

Concept

By the chain rule (Lesson 9), the gradient at an early layer is a product of one Jacobian factor per later layer. Roughly, its magnitude multiplies by some factor c_ℓ at each layer:

\[ \left\lVert \frac{\partial \mathcal{L}}{\partial a^{(1)}} \right\rVert \;\sim\; \Big(\textstyle\prod_{\ell=1}^{L} c_\ell\Big)\, \left\lVert \frac{\partial \mathcal{L}}{\partial a^{(L)}} \right\rVert \]

A product of many numbers is unstable. If the typical factor is above 1 it blows up; if below 1 it dies. Only a factor of exactly 1 survives depth — and hitting that is the job of initialization.

8. Why 'a little off' becomes 'catastrophic'

Intuition

Addition forgives small errors; multiplication does not. A per-layer factor of 1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.

Same story downward: a factor of 0.7 per layer feels close to 1, yet after 30 layers almost nothing is left. The gradient reaches the first layer as noise, and those layers never learn.

So the target is razor-thin: keep the per-layer factor at 1. Next we watch exactly how fast the two failure modes run away.

9. By analogy: Why 'a little off' becomes 'catastrophic'

Analogy

Discussion prompt

Explain Why 'a little off' becomes 'catastrophic' by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Addition forgives small errors; multiplication does not. A per-layer factor of 1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.

10. Guess the shape of the answer: Depth turns 1.5 into 190,000×

Estimation

Predict first

Make the per-layer factor concrete. Take 30 layers, each multiplying the gradient magnitude by a constant c. The overall factor is c^30. Three cases — this snippet is complete and runnable:

Commit before you compute: what does Depth turns 1.5 into 190,000× come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: c = 1.5 → 1.9e+05 (explodes); c = 0.7 → 2.3e-05 (vanishes); c = 1.0 → 1.0 (holds)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A 50% per-layer over-scale becomes a 190,000× blow-up in 30 layers; a 30% under-scale shrinks the gradient to 2.3e-5.

11. Depth turns 1.5 into 190,000×

Worked example

Make the per-layer factor concrete. Take 30 layers, each multiplying the gradient magnitude by a constant c. The overall factor is c^30. Three cases — this snippet is complete and runnable:

import numpy as np
depth = 30
for c in (1.5, 1.0, 0.7):
    total = np.prod(np.full(depth, c))
    print(f'c={c}: c^{depth} = {total:.3e}')

c = 1.5 → 1.9e+05 (explodes); c = 0.7 → 2.3e-05 (vanishes); c = 1.0 → 1.0 (holds)

Why: A 50% per-layer over-scale becomes a 190,000× blow-up in 30 layers; a 30% under-scale shrinks the gradient to 2.3e-5. Only c = 1 is stable. Verified by execution.

per-layer factor cc^30 (verified)verdict
1.51.918e+05explodes
1.01.000e+00stable
0.72.254e-05vanishes

12. Fill in: c^30 (verified) for Depth turns 1.5 into 190,000×

Comparison

Comparison matrix

From Depth turns 1.5 into 190,000×: refill the c^30 (verified) column from what you know. The rest of the table is as it appeared.

per-layer factor cc^30 (verified)verdict
1.51.918e+05explodes
1.01.000e+00stable
0.72.254e-05vanishes

13. Name the two failure modes

Concept

exploding gradients — The per-layer factor is > 1, so the gradient product grows exponentially with depth — huge parameter jumps, inf/NaN losses, divergence. Common in deep and recurrent nets.

vanishing gradients — The per-layer factor is < 1, so the gradient product decays exponentially — early layers receive a gradient near 0 and stop learning. The quieter, more insidious failure.

Both come from the same multiplicative product. Fix the per-layer factor and you fix both at once — which is exactly what initialization does.

14. Take the definitions apart: exploding gradients vs vanishing gradients

Definition probe

Sort into buckets

Every line below is part of the definition of exploding gradients or of vanishing gradients — one or the other, never both. Put each where it belongs.

exploding gradients
The per-layer factor is > 1, so the gradient product grows exponentially with depth; huge parameter jumps, inf/NaN losses, divergence.; Common in deep and recurrent nets.
vanishing gradients
The per-layer factor is < 1, so the gradient product decays exponentially; early layers receive a gradient near 0 and stop learning.; The quieter, more insidious failure.
b1
The per-layer factor is > 1, so the gradient product grows exponentially with depth — huge parameter jumps, inf/NaN losses, divergence. Common in deep and recurrent nets.
b2
The per-layer factor is < 1, so the gradient product decays exponentially — early layers receive a gradient near 0 and stop learning. The quieter, more insidious failure.

15. Two separate fixes

Concept

There are two distinct levers, and this lesson does both:

  1. Clipping — a runtime safety net that caps a gradient the instant it explodes, so one bad step can't wreck training.
  2. Initialization — a design-time choice of Var(W) that keeps the per-layer factor near 1 from the start, so gradients rarely explode or vanish in the first place.

Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.

16. Teach it back: Two separate fixes

Explain it

Discussion prompt

Explain Two separate fixes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.

17. Clipping exploding gradients

Section

Part 2 of 7 — the runtime net

18. The idea: cap the length, keep the aim

Concept

When a gradient g explodes, its direction is usually still fine — it's the length that's dangerous, because the update w ← w − η g takes a giant leap and overshoots.

So we want to shorten g to a safe length τ (the threshold) without turning it — keep the aim, cut the stride. That is exactly what scaling the whole vector by one positive number does.

19. What has to happen first: Derive the clip factor

Ranking

Put in order

Put the moves of Derive the clip factor into the order they have to happen.

  1. Only act when the gradient is too long
  2. Scale by a single positive number s so the new norm equals τ
  3. Fold both cases into one min()

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. If ‖g‖ ≤ τ the step is already safe — leave g untouched.

20. Derive the clip factor

Worked example

Only act when the gradient is too long

Why: If ‖g‖ ≤ τ the step is already safe — leave g untouched. We only rescale when ‖g‖ > τ.

Scale by a single positive number s so the new norm equals τ

Why: Multiplying g by s > 0 keeps its direction; we choose s so ‖s·g‖ = τ.

\[ \lVert s\,g \rVert = s\,\lVert g \rVert = \tau \;\Longrightarrow\; s = \frac{\tau}{\lVert g \rVert} \]

Fold both cases into one min()

Why: When ‖g‖ ≤ τ we want s = 1; when ‖g‖ > τ we want s = τ/‖g‖. Since τ/‖g‖ ≥ 1 exactly when ‖g‖ ≤ τ, the minimum of the two picks the right branch automatically.

\[ \boxed{\; g \leftarrow g \cdot \min\!\Big(1,\; \frac{\tau}{\lVert g \rVert}\Big) \;} \]

21. Decode the notation: Derive the clip factor

Notation

Annotate

From Derive the clip factor — read this one piece at a time. What is each part doing?

On: \( \lVert s\,g \rVert = s\,\lVert g \rVert = \tau \;\Longrightarrow\; s = \frac{\tau}{\lVert g \rVert} \)

  • If ‖g‖ ≤ τ the step is already safe — leave g untouched. We only rescale when ‖g‖ > τ.
  • Multiplying g by s > 0 keeps its direction; we choose s so ‖s·g‖ = τ.
  • When ‖g‖ ≤ τ we want s = 1; when ‖g‖ > τ we want s = τ/‖g‖. Since τ/‖g‖ ≥ 1 exactly when ‖g‖ ≤ τ, the minimum of the two picks the right branch automatically.

22. What has to be given first: Clip by norm — the trace

Missing information

Discussion prompt

Take g = [30, 40], threshold τ = 5. Its norm is √(30² + 40²) = √2500 = 50, so s = 5/50 = 0.1. Every number below is real output:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Every component shrank by the same 0.1, so the direction unit-vector [0.6, 0.8] is untouched — the final print is True. Length capped, aim preserved.

23. Clip by norm — the trace

Worked example

Take g = [30, 40], threshold τ = 5. Its norm is √(30² + 40²) = √2500 = 50, so s = 5/50 = 0.1. Every number below is real output:

import numpy as np
g = np.array([30.0, 40.0])
tau = 5.0
norm = np.linalg.norm(g)          # 50.0
s = min(1.0, tau / norm)          # 0.1
g_clip = g * s                    # [3., 4.]
print(round(norm, 1), round(s, 3))
print(g_clip, round(float(np.linalg.norm(g_clip)), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))

‖g‖ = 50 > 5, so s = 0.1 and g scales to [3, 4] with norm exactly 5

Why: Every component shrank by the same 0.1, so the direction unit-vector [0.6, 0.8] is untouched — the final print is True. Length capped, aim preserved.

quantityvalue (verified)
‖g‖ before50.0
scale s = min(1, τ/‖g‖)0.1
g after clip[3.0, 4.0]
‖g‖ after5.0
unit direction keptTrue

24. What each one costs: Clip by norm — the trace

Trade off

Comparison matrix

From Clip by norm — the trace: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityvalue (verified)
‖g‖ before50.0
scale s = min(1, τ/‖g‖)0.1
g after clip[3.0, 4.0]
‖g‖ after5.0
unit direction keptTrue

25. Why the whole vector, not each entry?

Intuition

There's a tempting alternative: clamp each component into [−τ, τ] separately. It also caps the size — but it silently rotates the gradient, because it shrinks the big components more than the small ones.

Rotating the gradient means you're no longer stepping downhill. Norm clipping multiplies every component by the same s, so the arrow keeps pointing the exact same way. That distinction is the trap on the next slide.

26. Something is wrong here: clamp each component instead of the norm

Anomaly

Predict first

A student writes this, and it looks reasonable:

To cap the gradient, just clamp each entry into [−τ, τ] with np.clip(g, -tau, tau).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Original unit direction was [0.6, 0.8].

Scale the whole vector by one factor s = min(1, τ/‖g‖).

Why: Original unit direction was [0.6, 0.8]. Clamping capped 30 and 40 to the SAME 5, so the ratio changed and the vector turned 8° off. You're no longer descending the loss — verified: keeps-direction = False.

27. Trap: clamp each component instead of the norm

Trap

The trap

To cap the gradient, just clamp each entry into [−τ, τ] with np.clip(g, -tau, tau).

\[ g = [30,\,40] \;\xrightarrow{\text{clamp to } \pm 5}\; [5,\,5] \]

Direction becomes [0.707, 0.707] — rotated

Why: Original unit direction was [0.6, 0.8]. Clamping capped 30 and 40 to the SAME 5, so the ratio changed and the vector turned 8° off. You're no longer descending the loss — verified: keeps-direction = False.

The fix

Scale the whole vector by one factor s = min(1, τ/‖g‖).

\[ g = [30,\,40] \;\xrightarrow{\times\, 0.1}\; [3,\,4] \]

Direction stays [0.6, 0.8] — unchanged

Why: Every component times the same 0.1, so the unit vector is identical and ‖g‖ = 5. Length capped, aim preserved — verified: keeps-direction = True. Clip by NORM, never per component.

28. Predict the next row: Side by side: norm clip vs clamp

Pattern

Predict first

The table runs: original g | [30, 40] | [0.6, 0.8] | — · norm clip (×0.1) | [3, 4] | [0.6, 0.8] | True

In Side by side: norm clip vs clamp, given the rows so far: what is the next one — the row where method is per-component clamp?

Correct: per-component clamp | [5, 5] | [0.707, 0.707] | False

methodresultunit directionkeeps aim?
original g[30, 40][0.6, 0.8]—
norm clip (×0.1)[3, 4][0.6, 0.8]True
per-component clamp[5, 5][0.707, 0.707]False

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two methods produce genuinely different directions.

29. Side by side: norm clip vs clamp

Worked example

Run both on the same g = [30, 40] and compare their unit directions to the original. Complete, runnable:

import numpy as np
g = np.array([30.0, 40.0]); tau = 5.0
u = g / np.linalg.norm(g)                      # [0.6, 0.8]
g_norm = g * min(1.0, tau / np.linalg.norm(g)) # [3., 4.]
g_clamp = np.clip(g, -tau, tau)                # [5., 5.]
print('norm-clip dir:', np.round(g_norm/np.linalg.norm(g_norm), 3))
print('clamp dir:    ', np.round(g_clamp/np.linalg.norm(g_clamp), 3))
print('norm keeps dir:', np.allclose(g_norm/np.linalg.norm(g_norm), u))
print('clamp keeps dir:', np.allclose(g_clamp/np.linalg.norm(g_clamp), u))

norm-clip dir = [0.6, 0.8] (True); clamp dir = [0.707, 0.707] (False)

Why: The two methods produce genuinely different directions. Norm clipping is the only one that leaves the descent direction alone. Verified by execution.

methodresultunit directionkeeps aim?
original g[30, 40][0.6, 0.8]—
norm clip (×0.1)[3, 4][0.6, 0.8]True
per-component clamp[5, 5][0.707, 0.707]False

30. Work backwards from the answer: Side by side: norm clip vs clamp

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

norm-clip dir = [0.6, 0.8] (True); clamp dir = [0.707, 0.707] (False)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Run both on the same g = [30, 40] and compare their unit directions to the original. Complete, runnable:

31. Choosing the threshold τ

Concept

τ is the only knob. Too small and you throttle every step (slow learning); too large and it never triggers (no protection). A common recipe: track the running gradient-norm history and set τ to a high percentile, so only genuine spikes get clipped.

Clipping is a safety net, not a substitute for a good init — a well-initialized net rarely trips it. It matters most in recurrent nets, where the same W is applied every timestep and the product can spike hard (Pascanu et al., 2013).

32. Variance through a layer

Section

Part 3 of 7 — the building block

33. Track the signal's 'loudness'

Intuition

Clipping fought the symptom. To fix the cause, we ask: as the signal passes through a layer, does its typical size grow, shrink, or hold? 'Typical size' is captured by variance — the spread of the activations around zero.

If we can compute how one layer transforms variance, we can chain that rule L times and demand the result stays put. That single per-layer rule is the engine behind Xavier and He.

34. Setup: one linear layer

Concept

Look at a single pre-activation z = W a, where a has n inputs. Each output entry is a sum of n products:

\[ z_i = \sum_{j=1}^{n} W_{ij}\, a_j \]

Assume the weights W_ij are independent, mean zero, with variance Var(W), and independent of the inputs a_j which have mean zero and variance Var(a). These are the standard init assumptions — everything follows from them.

35. Why variance, not the mean

Intuition

Weights are drawn symmetrically around 0, so the mean of z is 0 at every layer — the mean carries no scale information. What actually grows or shrinks through depth is the spread, and spread is variance.

That's why every init rule is a statement about Var(W), never about the mean. Keep the mean at 0 (for symmetry, Part 7) and tune the variance (for scale) — two independent jobs.

36. Where does each piece belong: Lesson 36: Initialization & Gradient Clipping

Sorting

Sort into buckets

These are the pieces of Lesson 36: Initialization & Gradient Clipping, out of order. Put each one back under the part of the lesson it belongs to.

The problem: depth multiplies
A deep net is a chain of maps; The gradient is a product; Why 'a little off' becomes 'catastrophic'
Clipping exploding gradients
The idea: cap the length, keep the aim; Derive the clip factor; Clip by norm — the trace
Variance through a layer
Track the signal's 'loudness'; Setup: one linear layer; Why variance, not the mean
s1
The problem: depth multiplies is where Lesson 36: Initialization & Gradient Clipping puts A deep net is a chain of maps, The gradient is a product, Why 'a little off' becomes 'catastrophic'. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Clipping exploding gradients is where Lesson 36: Initialization & Gradient Clipping puts The idea: cap the length, keep the aim, Derive the clip factor, Clip by norm — the trace. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Variance through a layer is where Lesson 36: Initialization & Gradient Clipping puts Track the signal's 'loudness', Setup: one linear layer, Why variance, not the mean. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

37. Plan first: Variance of one output entry

Step zero

Discussion prompt

Variance of one output entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Variance of a sum of independent terms adds

Answer:

  1. Variance of a sum of independent terms adds
  2. Variance of a product of independent mean-zero factors
  3. Add up n identical terms

38. Variance of one output entry

Worked example

Variance of a sum of independent terms adds

Why: The n products W_ij·a_j are independent and mean-zero, so Var of their sum is the sum of their variances.

\[ \mathrm{Var}(z_i) = \sum_{j=1}^{n} \mathrm{Var}(W_{ij}\, a_j) \]

Variance of a product of independent mean-zero factors

Why: For independent X, Y with mean 0, Var(XY) = Var(X)·Var(Y). So each term is Var(W)·Var(a).

\[ \mathrm{Var}(W_{ij}\,a_j) = \mathrm{Var}(W)\,\mathrm{Var}(a) \]

Add up n identical terms

Why: All n terms are equal, so the sum is n times one of them. This is the master relation for forward signal propagation.

\[ \boxed{\;\mathrm{Var}(z) = n\,\mathrm{Var}(W)\,\mathrm{Var}(a)\;} \]

39. Say it in words: Variance of one output entry

Translation

\( \boxed{\;\mathrm{Var}(z) = n\,\mathrm{Var}(W)\,\mathrm{Var}(a)\;} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

40. Restore the missing line: Confirm Var(z) = n·Var(W)·Var(a)

Fill the middle

Fill in the blanks

From Confirm Var(z) = n·Var(W)·Var(a) — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(0)
n = 512; sigma2 = 1.0 / n
zs = []
for _ in range(4000):
a = rng.normal(size=n) # Var(a) = 1
w = rng.normal(size=n) * np.sqrt(sigma2)
zs.append(w @ a)
print('predicted:', n * sigma2 * 1.0)
print('empirical :', round(float(np.var(zs)), 4))

Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The tiny gap is Monte-Carlo noise from 4000 samples.

41. Confirm Var(z) = n·Var(W)·Var(a)

Worked example

Take n = 512, unit-variance input a, and weights with Var(W) = 1/n. The formula predicts Var(z) = 512·(1/512)·1 = 1. Measure it over 4000 draws — runnable as written:

import numpy as np
rng = np.random.default_rng(0)
n = 512; sigma2 = 1.0 / n
zs = []
for _ in range(4000):
    a = rng.normal(size=n)                 # Var(a) = 1
    w = rng.normal(size=n) * np.sqrt(sigma2)
    zs.append(w @ a)
print('predicted:', n * sigma2 * 1.0)
print('empirical :', round(float(np.var(zs)), 4))

Predicted 1.0, empirical ≈ 1.0195 — they match

Why: The tiny gap is Monte-Carlo noise from 4000 samples. Choosing Var(W) = 1/n makes one linear layer preserve variance exactly. That choice is Xavier — derived next.

quantityvalue (verified)
n · Var(W) · Var(a) (predicted)1.0
empirical Var(z), 4000 draws1.0195
Var(W) used1/512 ≈ 0.001953

42. Chain the layers: a per-layer ratio

Concept

One layer multiplies the variance by a factor r = n·Var(W) (linear) or r = ½·n·Var(W) (after ReLU, which deletes half). Apply it L times and the variances form a geometric sequence:

\[ \mathrm{Var}(a^{(L)}) = r^{L}\,\mathrm{Var}(a^{(0)}), \qquad r = \tfrac{1}{2}\,n\,\mathrm{Var}(W)\ \text{(ReLU)} \]

This is the same product-blows-up problem as the gradient. Signals hold only if r = 1 exactly. r > 1 explodes, r < 1 vanishes. The whole design question is: pick Var(W) so r = 1.

43. Finish it with less help: The ratio predicts the 50-layer std

Faded example

Fill in the blanks

The ratio predicts the 50-layer std, with the scaffolding fading: two lines are gone now — fill both.

n = 256
for name, scale in [('He', 2), ('Xavier', 1), ('too_big', 8)]:
varW = scale / n
r = **0.5 * n * varW** # ReLU per-layer variance ratio
print(f'___: r=___, Var ratio r^50 = ___')

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. He's ½·256·(2/256) = 1 is the ONLY ratio that survives depth.

44. The ratio predicts the 50-layer std

Worked example

Compute the per-layer ReLU variance ratio r = ½·n·Var(W) for three inits at n = 256, then raise to the 50th power. This is a closed-form prediction of the experiment we'll run — runnable:

n = 256
for name, scale in [('He', 2), ('Xavier', 1), ('too_big', 8)]:
    varW = scale / n
    r = 0.5 * n * varW          # ReLU per-layer variance ratio
    print(f'{name}: r={r}, Var ratio r^50 = {r**50:.3e}')

He r = 1 (holds); Xavier r = 0.5 (r⁵⁰ = 8.9e-16); too_big r = 4 (r⁵⁰ = 1.3e+30)

Why: He's ½·256·(2/256) = 1 is the ONLY ratio that survives depth. Xavier's r = 0.5 means the variance halves each layer — 8.9e-16 after 50, whose square root ≈ 3e-8 matches the measured std 1.47e-8. The algebra predicts the experiment.

initr = ½·n·Var(W)Var ratio r⁵⁰ (verified)
He (2/n)1.01.000e+00 (stable)
Xavier (1/n)0.58.882e-16 (vanishes)
too_big (8/n)4.01.268e+30 (explodes)

45. Xavier initialization

Section

Part 4 of 7 — variance preservation

46. Demand: variance holds layer to layer

Concept

For a linear (or tanh-like, symmetric) network, we want each layer's output variance to equal its input variance — so signals neither grow nor shrink through depth. Set Var(z) = Var(a):

\[ \mathrm{Var}(z) = n_{\text{in}}\,\mathrm{Var}(W)\,\mathrm{Var}(a) \;\overset{!}{=}\; \mathrm{Var}(a) \]

47. What has to happen first: Solve for Var(W): the forward condition

Ranking

Put in order

Put the moves of Solve for Var(W): the forward condition into the order they have to happen.

  1. Cancel Var(a) from both sides
  2. Isolate Var(W)
  3. Backward gives the mirror condition

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Var(a) > 0, so divide it out. The activation variance drops away, leaving a pure condition on W and the fan-in.

48. Solve for Var(W): the forward condition

Worked example

Cancel Var(a) from both sides

Why: Var(a) > 0, so divide it out. The activation variance drops away, leaving a pure condition on W and the fan-in.

\[ n_{\text{in}}\,\mathrm{Var}(W) = 1 \]

Isolate Var(W)

Why: Divide by the fan-in n_in. This is the forward-preservation initializer.

\[ \mathrm{Var}(W) = \frac{1}{n_{\text{in}}} \]

Backward gives the mirror condition

Why: Propagating the gradient backward is the same algebra with the transpose, so it wants Var(W) = 1/n_out. Both can't hold unless n_in = n_out.

\[ \mathrm{Var}(W) = \frac{1}{n_{\text{out}}}\ \text{(backward)} \]

49. Draw the shape of it: Solve for Var(W): the forward condition

Blank canvas

Draw it

Draw what Solve for Var(W): the forward condition just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

50. Xavier: compromise between the two

Concept

Forward wants 1/n_in; backward wants 1/n_out. Glorot & Bengio split the difference with the harmonic-style average — average the two denominators:

\[ \boxed{\;\mathrm{Var}(W)_{\text{Xavier}} = \frac{2}{n_{\text{in}} + n_{\text{out}}}\;} \]

When n_in = n_out = n this reduces to 2/(2n) = 1/n — exactly the forward condition. Xavier is built for symmetric activations (tanh, sigmoid) that don't distort variance. ReLU does distort it — that gap is Part 5.

51. Something is wrong here: `1/n` vs `2/(n_in+n_out)`

Anomaly

Predict first

A student writes this, and it looks reasonable:

Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n.

Substitute carefully: the denominator is n_in + n_out = 2n.

Why: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.

52. Trap: `1/n` vs `2/(n_in+n_out)`

Trap

The trap

Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.

Var(W) = 2/n on a square layer

Why: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.

The fix

Substitute carefully: the denominator is n_in + n_out = 2n.

Var(W) = 2/(2n) = 1/n

Why: On a square layer Xavier's per-weight variance is 1/n — the same forward-preservation value. The factor of 2 in Xavier's numerator only offsets the SUM of two fan sizes; it is NOT the ReLU factor of 2 (that's He, coming up).

53. Break it on purpose: `1/n` vs `2/(n_in+n_out)`

Break the constraint

Discussion prompt

The rule this trap just fixed:

Substitute carefully: the denominator is n_in + n_out = 2n.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.

54. He initialization

Section

Part 5 of 7 — the ReLU factor of 2

55. ReLU throws away half the signal

Intuition

ReLU sets every negative pre-activation to 0. For a symmetric, mean-zero z, that's half the values zeroed. Half your variance is deleted at every layer.

So a variance-preserving linear init (Xavier's 1/n) is too weak once ReLU is in the loop: the signal loses half its power per layer and decays. He's fix is a single factor — put the missing half back by doubling Var(W). Let's prove the 'half' claim, then derive the 2.

56. Teach it back: ReLU throws away half the signal

Explain it

Discussion prompt

Explain ReLU throws away half the signal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

ReLU sets every negative pre-activation to 0. For a symmetric, mean-zero z, that's half the values zeroed. Half your variance is deleted at every layer.

57. Guess the shape of the answer: ReLU keeps exactly half the variance

Estimation

Predict first

For symmetric mean-zero z, relu(z) = max(0, z) keeps the positive half untouched and zeros the rest. Claim: E[relu(z)²] = ½·Var(z). Measure it on 2 million draws:

Commit before you compute: what does ReLU keeps exactly half the variance come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: E[relu(z)²] ≈ 0.5006 ≈ ½·Var(z)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Half the z's are zeroed (contribute nothing); the surviving half carry the full second moment, and by symmetry that's exactly half the total.

58. ReLU keeps exactly half the variance

Worked example

For symmetric mean-zero z, relu(z) = max(0, z) keeps the positive half untouched and zeros the rest. Claim: E[relu(z)²] = ½·Var(z). Measure it on 2 million draws:

import numpy as np
rng = np.random.default_rng(0)
z = rng.normal(size=2_000_000)     # mean 0, Var 1
a = np.maximum(0, z)               # ReLU
print('Var(z):', round(float(np.var(z)), 4))
print('E[relu(z)^2]:', round(float(np.mean(a**2)), 4))
print('ratio:', round(float(np.mean(a**2) / np.var(z)), 4))

E[relu(z)²] ≈ 0.5006 ≈ ½·Var(z)

Why: Half the z's are zeroed (contribute nothing); the surviving half carry the full second moment, and by symmetry that's exactly half the total. The ratio 0.5008 confirms E[relu²] = ½·Var(z). Verified.

quantityvalue (verified)
Var(z)0.9997
E[relu(z)²] (variance after ReLU)0.5006
ratio E[relu²] / Var(z)0.5008 (≈ ½)

59. What has to happen first: Derive He: put the half back

Ranking

Put in order

Put the moves of Derive He: put the half back into the order they have to happen.

  1. Forward relation with a ReLU layer
  2. Demand the activation variance holds
  3. Solve for Var(W) — the factor of 2 appears

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The pre-activation still has Var(z) = n·Var(W)·Var(a).

60. Derive He: put the half back

Worked example

Forward relation with a ReLU layer

Why: The pre-activation still has Var(z) = n·Var(W)·Var(a). But the activation a = relu(z) now carries only HALF that variance, from the previous slide.

\[ \mathrm{Var}(a^{(\ell)}) = \tfrac{1}{2}\,\mathrm{Var}(z^{(\ell)}) = \tfrac{1}{2}\, n\,\mathrm{Var}(W)\,\mathrm{Var}(a^{(\ell-1)}) \]

Demand the activation variance holds

Why: Set Var(a^(ℓ)) = Var(a^(ℓ-1)) so the signal is stable through depth, then cancel that common variance.

\[ \tfrac{1}{2}\, n\,\mathrm{Var}(W) = 1 \]

Solve for Var(W) — the factor of 2 appears

Why: Multiply both sides by 2 and divide by n. The 2 is exactly the compensation for ReLU zeroing half the units.

\[ \boxed{\;\mathrm{Var}(W)_{\text{He}} = \frac{2}{n_{\text{in}}}\;} \]

61. Decode the notation: Derive He: put the half back

Notation

Annotate

From Derive He: put the half back — read this one piece at a time. What is each part doing?

On: \( \boxed{\;\mathrm{Var}(W)_{\text{He}} = \frac{2}{n_{\text{in}}}\;} \)

  • The pre-activation still has Var(z) = n·Var(W)·Var(a). But the activation a = relu(z) now carries only HALF that variance, from the previous slide.
  • Set Var(a^(ℓ)) = Var(a^(ℓ-1)) so the signal is stable through depth, then cancel that common variance.
  • Multiply both sides by 2 and divide by n. The 2 is exactly the compensation for ReLU zeroing half the units.

62. Xavier vs He, side by side

Concept

\[ \underbrace{\mathrm{Var}(W) = \frac{2}{n_{\text{in}}+n_{\text{out}}}}_{\text{Xavier — tanh/sigmoid}} \qquad \underbrace{\mathrm{Var}(W) = \frac{2}{n_{\text{in}}}}_{\text{He — ReLU}} \]

The only difference that matters for ReLU: He's numerator-2 sits over a single n_in, so it's about twice Xavier's variance on a square layer (2/n vs 1/n). That extra factor of 2 is the half of the signal ReLU deletes, restored.

initVar(W), square layerstd = √Var, n=256activation to use
Xavier1/n0.0625tanh, sigmoid
He2/n0.088388ReLU

63. Something is wrong here: Xavier init for a ReLU network

Anomaly

Predict first

A student writes this, and it looks reasonable:

Xavier is the classic initializer, so use it everywhere — including deep ReLU nets.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer.

Use He (2/n_in) for ReLU — the factor of 2 refills the deleted half each layer.

Why: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer. Verified: std decays to 1.47e-08 by layer 50 — the signal is gone.

64. Trap: Xavier init for a ReLU network

Trap

The trap

Xavier is the classic initializer, so use it everywhere — including deep ReLU nets.

Var(W) = 1/n on 50 ReLU layers

Why: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer. Verified: std decays to 1.47e-08 by layer 50 — the signal is gone.

The fix

Use He (2/n_in) for ReLU — the factor of 2 refills the deleted half each layer.

Var(W) = 2/n keeps activation std ≈ O(1)

Why: Verified: He holds std ≈ 0.49 at layer 50 while Xavier vanishes. Rule: match the initializer to the activation — ReLU ⇒ He, tanh/sigmoid ⇒ Xavier.

65. Backward & the 50-layer test

Section

Part 6 of 7 — put it under load

66. What 'O(1)' actually buys you

Intuition

'Keep the std O(1)' isn't cosmetic. If activations vanish, the network's output barely responds to its input — nothing to learn from. If they explode, you get inf and NaN on the very first forward pass.

The same holds backward: an O(1) gradient at every layer means every layer gets a signal it can act on. Vanish and the early layers freeze; explode and the update overshoots. O(1) is the only regime where all L layers train together.

67. Backward wants the same scale

Concept

The gradient flows backward through Wᵀ and through the ReLU mask (which zeroes the same half of coordinates). So the gradient variance obeys the mirror of the forward recursion — and it, too, is stabilized by He's 2/n.

That's the point of a good init: it keeps both the forward signal and the backward gradient at O(1) through depth, so early layers get a usable gradient. Let's measure the backward pass directly.

68. Restore the missing line: He stabilizes the backward gradient too

Fill the middle

Fill in the blanks

From He stabilizes the backward gradient too — one line has had its right-hand side removed. Put it back.

import numpy as np
def grad_std(scale, depth=50, w=256, seed=1):
rng = np.random.default_rng(seed)
a = rng.normal(size=w); Ws, masks = [], []
for _ in range(depth): # forward, cache masks
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
z = W @ a; m = (z > 0).astype(float)
a = z * m; Ws.append(W); masks.append(m)
d = rng.normal(size=w) # backward from output
for L in range(depth - 1, -1, -1):
d = (d * masks[L]); d = Ws[L].T @ d
return d.std()
for name, s in [('He', 2), ('Xavier', 1)]:
print(name, f'___')

Why: d is what everything below it consumes, so the wrong expression here fails later and somewhere else. The backward pass sees the same halving from the ReLU mask (line 8) and the same Wᵀ scaling.

69. He stabilizes the backward gradient too

Worked example

Forward-propagate through 50 ReLU layers caching the masks, then back-propagate a unit gradient and read its std at the input. He vs Xavier — complete and runnable:

import numpy as np
def grad_std(scale, depth=50, w=256, seed=1):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w); Ws, masks = [], []
    for _ in range(depth):                 # forward, cache masks
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        z = W @ a; m = (z > 0).astype(float)
        a = z * m; Ws.append(W); masks.append(m)
    d = rng.normal(size=w)                  # backward from output
    for L in range(depth - 1, -1, -1):
        d = (d * masks[L]); d = Ws[L].T @ d
    return d.std()
for name, s in [('He', 2), ('Xavier', 1)]:
    print(name, f'{grad_std(s):.3e}')

He grad std ≈ 1.43 (alive); Xavier ≈ 4.26e-08 (vanished)

Why: The backward pass sees the same halving from the ReLU mask (line 8) and the same Wᵀ scaling. He keeps the gradient O(1) to the input layer; Xavier lets it die — so early layers never learn. Verified.

initgradient std at input layer (verified)
He (2/n)1.429e+00 (stable)
Xavier (1/n)4.260e-08 (vanished)

70. Restore the missing line: 50 ReLU layers, four initializations

Fill the middle

Fill in the blanks

From 50 ReLU layers, four initializations — one line has had its right-hand side removed. Put it back.

import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a) # ReLU
return a.std()
for name, s in [('He', 2), ('Xavier', 1),
('too_big', 8), ('too_small', 0.1)]:
print(name, f'___')

Why: a is what everything below it consumes, so the wrong expression here fails later and somewhere else. Only He's factor-of-2 keeps the std O(1) across 50 ReLU layers.

71. 50 ReLU layers, four initializations

Worked example

The headline experiment. Push a unit-variance signal through 50 ReLU layers under four scales Var(W) = scale/n and read the final activation std. scale = 2 is He, 1 is Xavier, 8 too big, 0.1 too small. Fully runnable:

import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w)
    for _ in range(depth):
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        a = np.maximum(0, W @ a)          # ReLU
    return a.std()
for name, s in [('He', 2), ('Xavier', 1),
                ('too_big', 8), ('too_small', 0.1)]:
    print(name, f'{final_std(s):.2e}')

He → 4.93e-01 (stable); Xavier → 1.47e-08 (vanishes); too_big → 5.55e+14 (explodes); too_small → 1.47e-33 (dead)

Why: Only He's factor-of-2 keeps the std O(1) across 50 ReLU layers. Everything else runs away by depth 50 — over-scale explodes, any under-scale collapses. Verified by execution.

initVar(W) scaleactivation std @ layer 50 (verified)
He2/n4.93e-01 (stable)
Xavier1/n1.47e-08 (vanishes)
too big8/n5.55e+14 (explodes)
too small0.1/n1.47e-33 (dead)

72. Fill in: Var(W) scale for 50 ReLU layers, four initializations

Comparison

Comparison matrix

From 50 ReLU layers, four initializations: refill the Var(W) scale column from what you know. The rest of the table is as it appeared.

initVar(W) scaleactivation std @ layer 50 (verified)
He2/n4.93e-01 (stable)
Xavier1/n1.47e-08 (vanishes)
too big8/n5.55e+14 (explodes)
too small0.1/n1.47e-33 (dead)

73. Predict the next row: Watch it happen layer by layer

Pattern

Predict first

The table runs: 1 | 8.54e-01 | 6.04e-01 · 2 | 7.98e-01 | 3.99e-01 · 5 | 8.97e-01 | 1.59e-01 · 10 | 9.26e-01 | 2.89e-02 · 25 | 7.58e-01 | 1.31e-04

In Watch it happen layer by layer, given the rows so far: what is the next one — the row where layer is 50?

Correct: 50 | 4.93e-01 | 1.47e-08

layerHe stdXavier std
18.54e-016.04e-01
27.98e-013.99e-01
58.97e-011.59e-01
109.26e-012.89e-02
257.58e-011.31e-04
504.93e-011.47e-08

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The full trace shows Xavier is losing roughly half its variance (≈0.7× std) each layer while He holds.

74. Watch it happen layer by layer

Worked example

Read the std after each layer instead of only at the end. He hovers around 1; Xavier halves and halves. This is your diagnostic tool — a per-layer std trace. Runnable:

import numpy as np
def trace(scale, depth=50, w=256, seed=0):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w); out = []
    for _ in range(depth):
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        a = np.maximum(0, W @ a)
        out.append(a.std())
    return out
he, xa = trace(2), trace(1)
for L in (1, 2, 5, 10, 25, 50):
    print(f'layer {L:2d}  He {he[L-1]:.2e}  Xavier {xa[L-1]:.2e}')

He stays near 1 (0.85 → 0.49); Xavier decays (0.60 → 1.47e-08)

Why: The full trace shows Xavier is losing roughly half its variance (≈0.7× std) each layer while He holds. Watching per-layer std is exactly how you diagnose a real init bug. Verified.

layerHe stdXavier std
18.54e-016.04e-01
27.98e-013.99e-01
58.97e-011.59e-01
109.26e-012.89e-02
257.58e-011.31e-04
504.93e-011.47e-08

75. Watch it run: Watch it happen layer by layer

Pattern

Step through it

Step through Watch it happen layer by layer one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: layer is 1
  2. Step 2: layer is 2
  3. Step 3: layer is 5
  4. Step 4: layer is 10
  5. Step 5: layer is 25
  6. Step 6: layer is 50

76. Zero init & symmetry

Section

Part 7 of 7 — the other failure

77. Zero feels safe — and never learns

Intuition

Even with the right variance in mind, a beginner reaches for the 'neutral' choice: set every weight to 0 (or any single constant). It's not a scaling problem — it's a symmetry problem.

If every neuron in a layer starts identical, they compute the same output, receive the same gradient, and take the same step — forever. They stay clones, so the layer can only ever represent one feature. Let's prove that by hand.

78. By analogy: Zero feels safe — and never learns

Analogy

Discussion prompt

Explain Zero feels safe — and never learns by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Even with the right variance in mind, a beginner reaches for the 'neutral' choice: set every weight to 0 (or any single constant). It's not a scaling problem — it's a symmetry problem.

79. What has to be given first: Prove: identical rows, identical gradients

Missing information

Discussion prompt

Tiny net: 2 inputs → 3 hidden (ReLU) → 1 output, all weights zero, MSE loss. Forward is all zeros; backward gives identical rows in dW1. Runnable as written:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

dh = dyhat·W2 is the same for all three neurons because W2 is a single constant, so the three hidden units get the SAME gradient. They will update identically and remain clones. Symmetry is never broken. Verified: rows-all-equal = True.

80. Prove: identical rows, identical gradients

Worked example

Tiny net: 2 inputs → 3 hidden (ReLU) → 1 output, all weights zero, MSE loss. Forward is all zeros; backward gives identical rows in dW1. Runnable as written:

import numpy as np
x = np.array([1.0, 2.0]); y = 1.0
W1 = np.zeros((3, 2))          # hidden weights, all zero
W2 = np.zeros(3)              # output weights, all zero
h_pre = W1 @ x               # [0, 0, 0]
h = np.maximum(0, h_pre)     # [0, 0, 0]
yhat = W2 @ h               # 0.0
dyhat = yhat - y            # -1.0
dh = dyhat * W2             # [0,0,0]  (W2 is zero)
dW1 = np.outer(dh * (h_pre > 0), x)   # 3x2
print('h:', h)
print('dW1 rows all equal:', np.allclose(dW1, dW1[0]))
print('dW1:', dW1.tolist())

Every row of dW1 is identical (here all zero)

Why: dh = dyhat·W2 is the same for all three neurons because W2 is a single constant, so the three hidden units get the SAME gradient. They will update identically and remain clones. Symmetry is never broken. Verified: rows-all-equal = True.

quantityvalue (verified)
hidden activations h[0, 0, 0]
dW1 all rows equalTrue
dW1[[0,0],[0,0],[0,0]]

81. Guess the shape of the answer: Random init breaks the symmetry

Estimation

Predict first

Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:

Commit before you compute: what does Random init breaks the symmetry come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. h_pre = [-0.139, 0.850, 0.188]: the first unit is off (gradient row is zero), the others are on with different gradients.

82. Random init breaks the symmetry

Worked example

Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:

import numpy as np
rng = np.random.default_rng(0)
x = np.array([1.0, 2.0]); y = 1.0
W1 = rng.normal(size=(3, 2)) * np.sqrt(2.0 / 2)   # He
W2 = rng.normal(size=3) * np.sqrt(2.0 / 3)
h_pre = W1 @ x
h = np.maximum(0, h_pre)
yhat = W2 @ h
dh = (yhat - y) * W2
dW1 = np.outer(dh * (h_pre > 0), x)
print('h_pre:', np.round(h_pre, 4))
print('dW1 row0:', np.round(dW1[0], 4))
print('dW1 row1:', np.round(dW1[1], 4))
print('rows identical?', np.allclose(dW1[0], dW1[1]))

Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ

Why: h_pre = [-0.139, 0.850, 0.188]: the first unit is off (gradient row is zero), the others are on with different gradients. rows-identical = False — the neurons will learn DIFFERENT features. Random init is what breaks the tie. Verified.

quantityvalue (verified)
h_pre[-0.139, 0.850, 0.188]
dW1 row 0 (ReLU off)[-0.0, -0.0]
dW1 row 1 (ReLU on)[-0.348, -0.696]
rows identical?False

83. Work backwards from the answer: Random init breaks the symmetry

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:

84. Something is wrong here: initialize weights to zero

Anomaly

Predict first

A student writes this, and it looks reasonable:

Zero is a neutral, unbiased starting point, so set all weights to 0.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever.

Initialize randomly from the He (or Xavier) distribution to break symmetry.

Why: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever. The layer collapses to one effective neuron and never learns distinct features. Bias terms alone can't break it — verified: dW1 rows all equal.

85. Trap: initialize weights to zero

Trap

The trap

Zero is a neutral, unbiased starting point, so set all weights to 0.

W = zeros → every neuron identical

Why: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever. The layer collapses to one effective neuron and never learns distinct features. Bias terms alone can't break it — verified: dW1 rows all equal.

The fix

Initialize randomly from the He (or Xavier) distribution to break symmetry.

W ~ N(0, 2/n_in) → neurons differentiate

Why: Random weights make each neuron start differently, so they receive different gradients and learn different features. Verified: dW1 rows differ. Symmetry-breaking is non-negotiable — and it comes free from any random init.

86. Which of these survive contact with Lesson 36: Initialization & Gradient Clipping?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Addition forgives small errors; multiplication does not. A per-layer factor of 1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.; Both come from the same multiplicative product. Fix the per-layer factor and you fix both at once — which is exactly what initialization does.; Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.
Breaks
To cap the gradient, just clamp each entry into [−τ, τ] with np.clip(g, -tau, tau).; Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 36: Initialization & Gradient Clipping puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

87. Biases: zero is fine

Concept

The zero-init warning is about weights, not biases. Biases are usually initialized to 0 and that's fine — the randomness in W already breaks the neuron symmetry, so identical starting biases don't clone anything.

Only when the weights are also constant do biases matter, and even then a nonzero bias can't rescue the symmetry (every neuron still shares it). So: random weights, zero biases is the standard, safe default.

88. Wrap-up & project

Section

Pattern, checks, build

89. Rebuild the recipe: The init / stability toolkit

Ranking

Put in order

These are the steps of The init / stability toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Exploding grads: clip by norm g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)
  2. ReLU nets: He Var(W) = 2/n_in — the 2 refills the half ReLU deletes
  3. tanh / sigmoid: Xavier Var(W) = 2/(n_in+n_out) — preserves variance for symmetric activations
  4. Never init all weights equal — random draw breaks neuron symmetry
  5. Diagnose by printing per-layer activation and gradient std: it should stay O(1), not drift toward 0 or ∞

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

90. The init / stability toolkit

Pattern

  1. Exploding grads: clip by norm g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)
  2. ReLU nets: He Var(W) = 2/n_in — the 2 refills the half ReLU deletes
  3. tanh / sigmoid: Xavier Var(W) = 2/(n_in+n_out) — preserves variance for symmetric activations
  4. Never init all weights equal — random draw breaks neuron symmetry
  5. Diagnose by printing per-layer activation and gradient std: it should stay O(1), not drift toward 0 or ∞

91. Where does it stop working: The init / stability toolkit

Edge cases

Discussion prompt

The init / stability toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Exploding grads: clip by norm g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)
  2. ReLU nets: He Var(W) = 2/n_in — the 2 refills the half ReLU deletes
  3. tanh / sigmoid: Xavier Var(W) = 2/(n_in+n_out) — preserves variance for symmetric activations
  4. Never init all weights equal — random draw breaks neuron symmetry
  5. Diagnose by printing per-layer activation and gradient std: it should stay O(1), not drift toward 0 or ∞

92. Rule out three: Check yourself — why depth is dangerous

Elimination

Eliminate the wrong options

Why do gradients explode or vanish specifically in DEEP networks?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The gradient is a product of per-layer factors, so a factor slightly off 1 compounds exponentially with depth
  • B. Deep networks use more memory, which rounds the gradients to 0
  • C. ReLU is undefined at 0, and deeper nets hit that point more often
  • D. The learning rate must grow with depth, overshooting the minimum

Survives elimination: A

Why: By the chain rule the gradient at an early layer is a product of one factor per later layer. A product of L numbers near c behaves like c^L: c = 1.5 gives 1.9e5 over 30 layers, c = 0.7 gives 2.3e-5. Only c ≈ 1 survives — which is what initialization targets.

93. Check yourself — why depth is dangerous

Check

Think about what the chain rule does to magnitudes.

Check your understanding

Why do gradients explode or vanish specifically in DEEP networks?

  • A. The gradient is a product of per-layer factors, so a factor slightly off 1 compounds exponentially with depth (correct)
  • B. Deep networks use more memory, which rounds the gradients to 0
  • C. ReLU is undefined at 0, and deeper nets hit that point more often
  • D. The learning rate must grow with depth, overshooting the minimum

Answer: A

Why: By the chain rule the gradient at an early layer is a product of one factor per later layer. A product of L numbers near c behaves like c^L: c = 1.5 gives 1.9e5 over 30 layers, c = 0.7 gives 2.3e-5. Only c ≈ 1 survives — which is what initialization targets.

Why B tempts people
Memory has nothing to do with it — the effect is purely the multiplicative chain-rule product, and it happens in exact arithmetic too.
Why C tempts people
ReLU's kink at 0 is a measure-zero event and isn't the cause; the compounding product of Jacobian factors is.
Why D tempts people
The learning rate is a fixed hyperparameter, not something that grows with depth; the instability comes from the gradient product itself, before the step size is applied.

94. Answer it before you see the options: Check yourself — He vs Xavier

Prediction

Predict first

For a deep network with ReLU activations, the correct initialization variance (square layer, fan-in n) is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: He: Var(W) = 2/n

Why: ReLU zeroes ~half the activations (E[relu²] = ½·Var(z), verified ≈ 0.5006), so He's factor-of-2 variance refills the deleted half and keeps the signal stable — std ≈ 0.49 at layer 50. Xavier's 1/n under-scales for ReLU and the signal decays to 1.47e-8.

95. Check yourself — He vs Xavier

Check

Match the initializer to the activation.

Check your understanding

For a deep network with ReLU activations, the correct initialization variance (square layer, fan-in n) is:

  • A. He: Var(W) = 2/n (correct)
  • B. Xavier: Var(W) = 1/n
  • C. all weights = 0
  • D. all weights = 1

Answer: A

Why: ReLU zeroes ~half the activations (E[relu²] = ½·Var(z), verified ≈ 0.5006), so He's factor-of-2 variance refills the deleted half and keeps the signal stable — std ≈ 0.49 at layer 50. Xavier's 1/n under-scales for ReLU and the signal decays to 1.47e-8.

Why B tempts people
Xavier (1/n on a square layer) preserves variance for symmetric activations; ReLU deletes an extra half each layer, so 1/n under-scales by 2× and activations decay to 1.47e-8 by layer 50.
Why C tempts people
Zero init sets the right variance to nothing AND breaks symmetry-breaking — every neuron stays identical and nothing learns; it's not a variance choice at all.
Why D tempts people
Constant nonzero init has the same broken-symmetry problem as zeros AND blows the forward signal up; a constant is never a valid initializer.

96. Answer it before you see the options: Check yourself — zero init

Prediction

Predict first

Initializing all weights of a layer to 0 fails because:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: every neuron gets identical gradients and never differentiates (broken symmetry)

Why: Identical weights → identical outputs → identical gradients → identical updates, forever. We verified dW1 has three identical rows, so the neurons never become distinct and the layer collapses to a single feature. Random init breaks the tie.

97. Check yourself — zero init

Check

Why exactly do zeros fail? Name the mechanism.

Check your understanding

Initializing all weights of a layer to 0 fails because:

  • A. every neuron gets identical gradients and never differentiates (broken symmetry) (correct)
  • B. the loss immediately becomes NaN
  • C. the gradients explode
  • D. it uses too much memory

Answer: A

Why: Identical weights → identical outputs → identical gradients → identical updates, forever. We verified dW1 has three identical rows, so the neurons never become distinct and the layer collapses to a single feature. Random init breaks the tie.

Why B tempts people
No NaN appears — the forward pass is a finite 0 and the loss is finite; learning is stuck, not numerically broken.
Why C tempts people
Zero weights make the gradient small and identical across neurons, not exploding — the failure is symmetry, not magnitude.
Why D tempts people
Memory is unaffected; the entire problem is the broken symmetry between neurons, not resource use.

98. Rule out three: Check yourself — clipping

Elimination

Eliminate the wrong options

Gradient norm clipping rescales g to g·min(1, τ/‖g‖). Compared with clamping each component to [−τ, τ], it:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. preserves the gradient's direction; per-component clamping does not
  • B. is identical to per-component clamping
  • C. increases the gradient norm
  • D. only works for 1-D gradients

Survives elimination: A

Why: Scaling the whole vector by one factor keeps every ratio between components fixed, so the direction is unchanged (only shorter). Per-component clamping shrinks big components more than small ones and rotates the vector — verified: [30,40]→[5,5] turns [0.6,0.8] into [0.707,0.707].

99. Check yourself — clipping

Check

Which clip keeps you descending the loss?

Check your understanding

Gradient norm clipping rescales g to g·min(1, τ/‖g‖). Compared with clamping each component to [−τ, τ], it:

  • A. preserves the gradient's direction; per-component clamping does not (correct)
  • B. is identical to per-component clamping
  • C. increases the gradient norm
  • D. only works for 1-D gradients

Answer: A

Why: Scaling the whole vector by one factor keeps every ratio between components fixed, so the direction is unchanged (only shorter). Per-component clamping shrinks big components more than small ones and rotates the vector — verified: [30,40]→[5,5] turns [0.6,0.8] into [0.707,0.707].

Why B tempts people
They differ: norm clipping keeps [0.6, 0.8]; value clamping bends it to [0.707, 0.707]. Only norm clipping preserves the descent direction.
Why C tempts people
It caps the norm DOWN to τ when ‖g‖ > τ (and leaves it alone otherwise); it never increases the norm.
Why D tempts people
It works in any dimension — scaling by a scalar is dimension-agnostic, which is precisely its advantage over per-component clamping.

100. Your turn: signal propagation

Section

Build it yourself

101. Project: a 50-layer init lab

Concept

Reproduce the headline result yourself: push a signal through 50 ReLU layers under each init and find the one that keeps the activation std stable — then add a norm-clip. You've derived every piece; now assemble it.

#requirementtool
1He forward pass, read final stdW·√(2/n), then ReLU
2compare He / Xavier / too-big / too-smallvary the scale
3norm-clip a large gradient, keep directiong·min(1, τ/‖g‖)

Build rules: type every line yourself, apply ReLU each layer, seed the RNG so your numbers match, and read the final std — only He should stay O(1).

102. Break it if you can: Project: a 50-layer init lab

Counterexample

Discussion prompt

Build rules: type every line yourself, apply ReLU each layer, seed the RNG so your numbers match, and read the final std — only He should stay O(1).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

103. Milestone 1 — He forward pass

Worked example

Your turn: push a signal through 50 ReLU layers with He init (scale = 2). Predict the order of magnitude of the final std out loud before you print.

Hint: each layer W = randn(w, w)·sqrt(2/w), then a = relu(W @ a). Seed with default_rng(0) so you get the same number.

import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w)
    for _ in range(depth):
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        a = np.maximum(0, W @ a)          # ReLU
    return a.std()
print('He std @ 50:', f'{final_std(2):.3f}')
initstd @ layer 50 (verified)
He (scale 2)0.493 (stable)
targetO(1) — near 0.5, not 0 or ∞

104. Milestone 2 — compare four initializations

Worked example

Your turn: reuse final_std and sweep all four scales. Predict which vanish and which explode before running.

Hint: scales are 2 (He), 1 (Xavier), 8 (too big), 0.1 (too small). Print with f'{x:.2e}' so the tiny/huge numbers are readable.

import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w)
    for _ in range(depth):
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        a = np.maximum(0, W @ a)
    return a.std()
for name, s in [('He', 2), ('Xavier', 1),
                ('too_big', 8), ('too_small', 0.1)]:
    print(name, f'{final_std(s):.2e}')
initstd @ 50 (verified)
He4.93e-01
Xavier1.47e-08
too_big5.55e+14
too_small1.47e-33

105. Watch it run: Milestone 2 — compare four initializations

Pattern

Step through it

Step through Milestone 2 — compare four initializations one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: init is He
  2. Step 2: init is Xavier
  3. Step 3: init is too_big
  4. Step 4: init is too_small

106. Milestone 3 — clip a gradient

Worked example

Your turn: norm-clip a gradient of norm 50 down to τ = 5 and confirm the direction is unchanged. Predict the clipped vector before printing.

Hint: g * min(1, tau/np.linalg.norm(g)), then compare the unit vectors before and after with np.allclose.

import numpy as np
g = np.array([30.0, 40.0]); tau = 5.0
gc = g * min(1.0, tau / np.linalg.norm(g))
print('gc:', gc)
print('norm after:', round(float(np.linalg.norm(gc)), 1))
print('direction kept:',
      np.allclose(gc / np.linalg.norm(gc), g / np.linalg.norm(g)))
quantityvalue (verified)
gc[3.0, 4.0]
‖gc‖ after5.0
direction keptTrue

107. What each one costs: Milestone 3 — clip a gradient

Trade off

Comparison matrix

From Milestone 3 — clip a gradient: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityvalue (verified)
gc[3.0, 4.0]
‖gc‖ after5.0
direction keptTrue

108. The full program

Concept

import numpy as np

def final_std(scale, depth=50, w=256, seed=0):
    rng = np.random.default_rng(seed)
    a = rng.normal(size=w)
    for _ in range(depth):
        W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
        a = np.maximum(0, W @ a)          # ReLU
    return a.std()

for name, s in [('He', 2), ('Xavier', 1),
                ('too_big', 8), ('too_small', 0.1)]:
    print(name, f'{final_std(s):.2e}')

g = np.array([30.0, 40.0])
gc = g * min(1.0, 5.0 / np.linalg.norm(g))
print('clip 50 ->', round(float(np.linalg.norm(gc)), 1),
      'dir kept', np.allclose(gc/np.linalg.norm(gc), g/np.linalg.norm(g)))
printed linevalue (verified)
He / Xavier4.93e-01 / 1.47e-08
too_big / too_small5.55e+14 / 1.47e-33
clip 50 ->5.0 dir kept True

If only He keeps the signal alive through 50 layers and your clip lands norm 5 with dir kept True — you understand why init and clipping decide whether a deep net trains at all.

109. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
He / Xavier4.93e-01 / 1.47e-08
too_big / too_small5.55e+14 / 1.47e-33
clip 50 ->5.0 dir kept True

110. Show it off

Concept

Slides closed, out loud: explain (1) why the gradient is a product so depth compounds errors, (2) where He's factor of 2 comes from over Xavier, and (3) why norm clipping preserves the descent direction but per-component clamping doesn't.

Stretch: re-derive Xavier's 1/n from the forward variance condition on paper, then swap ReLU for tanh in final_std and confirm Xavier now holds O(1) while He slightly over-scales. That's the exact init-to-activation match the exam probes. Init enables ResNets (Week 42) and transformer LayerNorm (Week 29).

111. Connect it up: Lesson 36: Initialization & Gradient Clipping

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The problem: depth multiplies · Clipping exploding gradients · Variance through a layer · Xavier initialization · He initialization · Backward & the 50-layer test. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

112. What you can do now

Recap

ideathe one thing to remember
depthgradient is a PRODUCT — factor must be ≈ 1
clipby NORM: g·min(1, τ/‖g‖), never per component
HeVar(W) = 2/n_in — the 2 is ReLU's deleted half
XavierVar(W) = 2/(n_in+n_out) — for tanh/sigmoid
zero initidentical rows ⇒ symmetry never breaks ⇒ never learns

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 36 (Week 12 — Initialization & Clipping) — Barron · USAAIO Round 2 Preparation, 2026
  2. Glorot & Bengio, Understanding the difficulty of training deep feedforward neural networks (Xavier init) — AISTATS 2010
  3. He, Zhang, Ren & Sun, Delving Deep into Rectifiers (He/Kaiming init) — ICCV 2015
  4. Pascanu, Mikolov & Bengio, On the difficulty of training recurrent neural networks (gradient norm clipping) — ICML 2013
  5. Every std, variance ratio, and clip result produced by real execution — numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108