USAAIO Lesson 36, from Week 12, fully worked, on why a deep network lives or dies at step zero. It presents exploding and vanishing gradients as a product of per-layer factors, derives gradient norm clipping and contrasts it with per-component clamping, and builds the forward-variance recursion for Var(a_L) one step at a time. It derives Xavier initialization from variance preservation and He initialization from the fact that ReLU zeros half its inputs, with no algebra skipped, then shows that the backward pass needs the SAME scale. It proves the zero-initialization symmetry problem by hand and ends with a 50-layer signal-propagation experiment in which only He survives. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6. The lesson runs to 62 slides.
Subject: Machine Learning · 112 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 36 · Week 12
A deep network lives or dies at step zero. We derive gradient norm clipping move by move, build the variance recursion for a deep net one layer at a time, get Xavier and He from it with no skipped algebra, and run a 50-layer experiment where only He survives.
Objectives
g ← g·min(1, τ/‖g‖) and prove it preserves the descent direction while per-component clamping does notVar(a_L) layer by layer and read off when the signal explodes, vanishes, or holdsVar(W)=1/n from linear variance preservation, then He Var(W)=2/n from ReLU zeroing half the units — every step shownstd ≈ O(1)Warm-up
Discussion prompt
Before we open Lesson 36: Initialization & Gradient Clipping: without looking back, what was the main idea of ROC, PR Curves & Calibration, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the ROC curve and AUC, the precision-recall curve for imbalanced data, why ROC-AUC misleads under heavy imbalance, model calibration, the Expected Calibration Error, and temperature scaling. Build an ROC curve from scratch and fix an overconfident model's calibration.
Section
Part 1 of 7
Concept
A feed-forward net stacks L layers: each takes the previous activation a, multiplies by a weight matrix W, and applies a nonlinearity. The output is a long composition of these maps.
\[ a^{(\ell)} = \phi\!\big(W^{(\ell)} a^{(\ell-1)}\big), \qquad \ell = 1, \dots, L \]
Both the forward signal and the backward gradient pass through every layer. Anything each layer does to the scale gets applied L times — that repetition is the whole story of this lesson.
Counterexample
Discussion prompt
A feed-forward net stacks L layers: each takes the previous activation a, multiplies by a weight matrix W, and applies a nonlinearity. The output is a long composition of these maps.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Both the forward signal and the backward gradient pass through every layer. Anything each layer does to the scale gets applied L times — that repetition is the whole story of this lesson.
Concept
By the chain rule (Lesson 9), the gradient at an early layer is a product of one Jacobian factor per later layer. Roughly, its magnitude multiplies by some factor c_ℓ at each layer:
\[ \left\lVert \frac{\partial \mathcal{L}}{\partial a^{(1)}} \right\rVert \;\sim\; \Big(\textstyle\prod_{\ell=1}^{L} c_\ell\Big)\, \left\lVert \frac{\partial \mathcal{L}}{\partial a^{(L)}} \right\rVert \]
A product of many numbers is unstable. If the typical factor is above 1 it blows up; if below 1 it dies. Only a factor of exactly 1 survives depth — and hitting that is the job of initialization.
Intuition
Addition forgives small errors; multiplication does not. A per-layer factor of 1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.
Same story downward: a factor of 0.7 per layer feels close to 1, yet after 30 layers almost nothing is left. The gradient reaches the first layer as noise, and those layers never learn.
So the target is razor-thin: keep the per-layer factor at 1. Next we watch exactly how fast the two failure modes run away.
Analogy
Discussion prompt
Explain Why 'a little off' becomes 'catastrophic' by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Addition forgives small errors; multiplication does not. A per-layer factor of 1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.
Estimation
Predict first
Make the per-layer factor concrete. Take 30 layers, each multiplying the gradient magnitude by a constant c. The overall factor is c^30. Three cases — this snippet is complete and runnable:
Commit before you compute: what does Depth turns 1.5 into 190,000× come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: c = 1.5 → 1.9e+05 (explodes); c = 0.7 → 2.3e-05 (vanishes); c = 1.0 → 1.0 (holds)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A 50% per-layer over-scale becomes a 190,000× blow-up in 30 layers; a 30% under-scale shrinks the gradient to 2.3e-5.
Worked example
Make the per-layer factor concrete. Take 30 layers, each multiplying the gradient magnitude by a constant c. The overall factor is c^30. Three cases — this snippet is complete and runnable:
import numpy as np
depth = 30
for c in (1.5, 1.0, 0.7):
total = np.prod(np.full(depth, c))
print(f'c={c}: c^{depth} = {total:.3e}')c = 1.5 → 1.9e+05 (explodes); c = 0.7 → 2.3e-05 (vanishes); c = 1.0 → 1.0 (holds)
Why: A 50% per-layer over-scale becomes a 190,000× blow-up in 30 layers; a 30% under-scale shrinks the gradient to 2.3e-5. Only c = 1 is stable. Verified by execution.
| per-layer factor c | c^30 (verified) | verdict |
|---|---|---|
| 1.5 | 1.918e+05 | explodes |
| 1.0 | 1.000e+00 | stable |
| 0.7 | 2.254e-05 | vanishes |
Comparison
Comparison matrix
From Depth turns 1.5 into 190,000×: refill the c^30 (verified) column from what you know. The rest of the table is as it appeared.
| per-layer factor c | c^30 (verified) | verdict |
|---|---|---|
| 1.5 | 1.918e+05 | explodes |
| 1.0 | 1.000e+00 | stable |
| 0.7 | 2.254e-05 | vanishes |
Concept
exploding gradients — The per-layer factor is > 1, so the gradient product grows exponentially with depth — huge parameter jumps, inf/NaN losses, divergence. Common in deep and recurrent nets.
vanishing gradients — The per-layer factor is < 1, so the gradient product decays exponentially — early layers receive a gradient near 0 and stop learning. The quieter, more insidious failure.
Both come from the same multiplicative product. Fix the per-layer factor and you fix both at once — which is exactly what initialization does.
Definition probe
Sort into buckets
Every line below is part of the definition of exploding gradients or of vanishing gradients — one or the other, never both. Put each where it belongs.
Concept
There are two distinct levers, and this lesson does both:
Var(W) that keeps the per-layer factor near 1 from the start, so gradients rarely explode or vanish in the first place.Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.
Explain it
Discussion prompt
Explain Two separate fixes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.
Section
Part 2 of 7 — the runtime net
Concept
When a gradient g explodes, its direction is usually still fine — it's the length that's dangerous, because the update w ← w − η g takes a giant leap and overshoots.
So we want to shorten g to a safe length τ (the threshold) without turning it — keep the aim, cut the stride. That is exactly what scaling the whole vector by one positive number does.
Ranking
Put in order
Put the moves of Derive the clip factor into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. If ‖g‖ ≤ τ the step is already safe — leave g untouched.
Worked example
Only act when the gradient is too long
Why: If ‖g‖ ≤ τ the step is already safe — leave g untouched. We only rescale when ‖g‖ > τ.
Scale by a single positive number s so the new norm equals τ
Why: Multiplying g by s > 0 keeps its direction; we choose s so ‖s·g‖ = τ.
\[ \lVert s\,g \rVert = s\,\lVert g \rVert = \tau \;\Longrightarrow\; s = \frac{\tau}{\lVert g \rVert} \]
Fold both cases into one min()
Why: When ‖g‖ ≤ τ we want s = 1; when ‖g‖ > τ we want s = τ/‖g‖. Since τ/‖g‖ ≥ 1 exactly when ‖g‖ ≤ τ, the minimum of the two picks the right branch automatically.
\[ \boxed{\; g \leftarrow g \cdot \min\!\Big(1,\; \frac{\tau}{\lVert g \rVert}\Big) \;} \]
Notation
Annotate
From Derive the clip factor — read this one piece at a time. What is each part doing?
On: \( \lVert s\,g \rVert = s\,\lVert g \rVert = \tau \;\Longrightarrow\; s = \frac{\tau}{\lVert g \rVert} \)
Missing information
Discussion prompt
Take g = [30, 40], threshold τ = 5. Its norm is √(30² + 40²) = √2500 = 50, so s = 5/50 = 0.1. Every number below is real output:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Every component shrank by the same 0.1, so the direction unit-vector [0.6, 0.8] is untouched — the final print is True. Length capped, aim preserved.
Worked example
Take g = [30, 40], threshold τ = 5. Its norm is √(30² + 40²) = √2500 = 50, so s = 5/50 = 0.1. Every number below is real output:
import numpy as np
g = np.array([30.0, 40.0])
tau = 5.0
norm = np.linalg.norm(g) # 50.0
s = min(1.0, tau / norm) # 0.1
g_clip = g * s # [3., 4.]
print(round(norm, 1), round(s, 3))
print(g_clip, round(float(np.linalg.norm(g_clip)), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))‖g‖ = 50 > 5, so s = 0.1 and g scales to [3, 4] with norm exactly 5
Why: Every component shrank by the same 0.1, so the direction unit-vector [0.6, 0.8] is untouched — the final print is True. Length capped, aim preserved.
| quantity | value (verified) |
|---|---|
| ‖g‖ before | 50.0 |
| scale s = min(1, τ/‖g‖) | 0.1 |
| g after clip | [3.0, 4.0] |
| ‖g‖ after | 5.0 |
| unit direction kept | True |
Trade off
Comparison matrix
From Clip by norm — the trace: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | value (verified) |
|---|---|
| ‖g‖ before | 50.0 |
| scale s = min(1, τ/‖g‖) | 0.1 |
| g after clip | [3.0, 4.0] |
| ‖g‖ after | 5.0 |
| unit direction kept | True |
Intuition
There's a tempting alternative: clamp each component into [−τ, τ] separately. It also caps the size — but it silently rotates the gradient, because it shrinks the big components more than the small ones.
Rotating the gradient means you're no longer stepping downhill. Norm clipping multiplies every component by the same s, so the arrow keeps pointing the exact same way. That distinction is the trap on the next slide.
Anomaly
Predict first
A student writes this, and it looks reasonable:
To cap the gradient, just clamp each entry into [−τ, τ] with np.clip(g, -tau, tau).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Original unit direction was [0.6, 0.8].
Scale the whole vector by one factor s = min(1, τ/‖g‖).
Why: Original unit direction was [0.6, 0.8]. Clamping capped 30 and 40 to the SAME 5, so the ratio changed and the vector turned 8° off. You're no longer descending the loss — verified: keeps-direction = False.
Trap
To cap the gradient, just clamp each entry into [−τ, τ] with np.clip(g, -tau, tau).
\[ g = [30,\,40] \;\xrightarrow{\text{clamp to } \pm 5}\; [5,\,5] \]
Direction becomes [0.707, 0.707] — rotated
Why: Original unit direction was [0.6, 0.8]. Clamping capped 30 and 40 to the SAME 5, so the ratio changed and the vector turned 8° off. You're no longer descending the loss — verified: keeps-direction = False.
Scale the whole vector by one factor s = min(1, τ/‖g‖).
\[ g = [30,\,40] \;\xrightarrow{\times\, 0.1}\; [3,\,4] \]
Direction stays [0.6, 0.8] — unchanged
Why: Every component times the same 0.1, so the unit vector is identical and ‖g‖ = 5. Length capped, aim preserved — verified: keeps-direction = True. Clip by NORM, never per component.
Pattern
Predict first
The table runs: original g | [30, 40] | [0.6, 0.8] | — · norm clip (×0.1) | [3, 4] | [0.6, 0.8] | True
In Side by side: norm clip vs clamp, given the rows so far: what is the next one — the row where method is per-component clamp?
Correct: per-component clamp | [5, 5] | [0.707, 0.707] | False
| method | result | unit direction | keeps aim? |
|---|---|---|---|
| original g | [30, 40] | [0.6, 0.8] | — |
| norm clip (×0.1) | [3, 4] | [0.6, 0.8] | True |
| per-component clamp | [5, 5] | [0.707, 0.707] | False |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two methods produce genuinely different directions.
Worked example
Run both on the same g = [30, 40] and compare their unit directions to the original. Complete, runnable:
import numpy as np
g = np.array([30.0, 40.0]); tau = 5.0
u = g / np.linalg.norm(g) # [0.6, 0.8]
g_norm = g * min(1.0, tau / np.linalg.norm(g)) # [3., 4.]
g_clamp = np.clip(g, -tau, tau) # [5., 5.]
print('norm-clip dir:', np.round(g_norm/np.linalg.norm(g_norm), 3))
print('clamp dir: ', np.round(g_clamp/np.linalg.norm(g_clamp), 3))
print('norm keeps dir:', np.allclose(g_norm/np.linalg.norm(g_norm), u))
print('clamp keeps dir:', np.allclose(g_clamp/np.linalg.norm(g_clamp), u))norm-clip dir = [0.6, 0.8] (True); clamp dir = [0.707, 0.707] (False)
Why: The two methods produce genuinely different directions. Norm clipping is the only one that leaves the descent direction alone. Verified by execution.
| method | result | unit direction | keeps aim? |
|---|---|---|---|
| original g | [30, 40] | [0.6, 0.8] | — |
| norm clip (×0.1) | [3, 4] | [0.6, 0.8] | True |
| per-component clamp | [5, 5] | [0.707, 0.707] | False |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
norm-clip dir = [0.6, 0.8] (True); clamp dir = [0.707, 0.707] (False)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Run both on the same g = [30, 40] and compare their unit directions to the original. Complete, runnable:
Concept
τ is the only knob. Too small and you throttle every step (slow learning); too large and it never triggers (no protection). A common recipe: track the running gradient-norm history and set τ to a high percentile, so only genuine spikes get clipped.
Clipping is a safety net, not a substitute for a good init — a well-initialized net rarely trips it. It matters most in recurrent nets, where the same W is applied every timestep and the product can spike hard (Pascanu et al., 2013).
Section
Part 3 of 7 — the building block
Intuition
Clipping fought the symptom. To fix the cause, we ask: as the signal passes through a layer, does its typical size grow, shrink, or hold? 'Typical size' is captured by variance — the spread of the activations around zero.
If we can compute how one layer transforms variance, we can chain that rule L times and demand the result stays put. That single per-layer rule is the engine behind Xavier and He.
Concept
Look at a single pre-activation z = W a, where a has n inputs. Each output entry is a sum of n products:
\[ z_i = \sum_{j=1}^{n} W_{ij}\, a_j \]
Assume the weights W_ij are independent, mean zero, with variance Var(W), and independent of the inputs a_j which have mean zero and variance Var(a). These are the standard init assumptions — everything follows from them.
Intuition
Weights are drawn symmetrically around 0, so the mean of z is 0 at every layer — the mean carries no scale information. What actually grows or shrinks through depth is the spread, and spread is variance.
That's why every init rule is a statement about Var(W), never about the mean. Keep the mean at 0 (for symmetry, Part 7) and tune the variance (for scale) — two independent jobs.
Sorting
Sort into buckets
These are the pieces of Lesson 36: Initialization & Gradient Clipping, out of order. Put each one back under the part of the lesson it belongs to.
Step zero
Discussion prompt
Variance of one output entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Variance of a sum of independent terms adds
Answer:
Worked example
Variance of a sum of independent terms adds
Why: The n products W_ij·a_j are independent and mean-zero, so Var of their sum is the sum of their variances.
\[ \mathrm{Var}(z_i) = \sum_{j=1}^{n} \mathrm{Var}(W_{ij}\, a_j) \]
Variance of a product of independent mean-zero factors
Why: For independent X, Y with mean 0, Var(XY) = Var(X)·Var(Y). So each term is Var(W)·Var(a).
\[ \mathrm{Var}(W_{ij}\,a_j) = \mathrm{Var}(W)\,\mathrm{Var}(a) \]
Add up n identical terms
Why: All n terms are equal, so the sum is n times one of them. This is the master relation for forward signal propagation.
\[ \boxed{\;\mathrm{Var}(z) = n\,\mathrm{Var}(W)\,\mathrm{Var}(a)\;} \]
Translation
\( \boxed{\;\mathrm{Var}(z) = n\,\mathrm{Var}(W)\,\mathrm{Var}(a)\;} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Fill the middle
Fill in the blanks
From Confirm Var(z) = n·Var(W)·Var(a) — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(0)
n = 512; sigma2 = 1.0 / n
zs = []
for _ in range(4000):
a = rng.normal(size=n) # Var(a) = 1
w = rng.normal(size=n) * np.sqrt(sigma2)
zs.append(w @ a)
print('predicted:', n * sigma2 * 1.0)
print('empirical :', round(float(np.var(zs)), 4))
Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The tiny gap is Monte-Carlo noise from 4000 samples.
Worked example
Take n = 512, unit-variance input a, and weights with Var(W) = 1/n. The formula predicts Var(z) = 512·(1/512)·1 = 1. Measure it over 4000 draws — runnable as written:
import numpy as np
rng = np.random.default_rng(0)
n = 512; sigma2 = 1.0 / n
zs = []
for _ in range(4000):
a = rng.normal(size=n) # Var(a) = 1
w = rng.normal(size=n) * np.sqrt(sigma2)
zs.append(w @ a)
print('predicted:', n * sigma2 * 1.0)
print('empirical :', round(float(np.var(zs)), 4))Predicted 1.0, empirical ≈ 1.0195 — they match
Why: The tiny gap is Monte-Carlo noise from 4000 samples. Choosing Var(W) = 1/n makes one linear layer preserve variance exactly. That choice is Xavier — derived next.
| quantity | value (verified) |
|---|---|
| n · Var(W) · Var(a) (predicted) | 1.0 |
| empirical Var(z), 4000 draws | 1.0195 |
| Var(W) used | 1/512 ≈ 0.001953 |
Concept
One layer multiplies the variance by a factor r = n·Var(W) (linear) or r = ½·n·Var(W) (after ReLU, which deletes half). Apply it L times and the variances form a geometric sequence:
\[ \mathrm{Var}(a^{(L)}) = r^{L}\,\mathrm{Var}(a^{(0)}), \qquad r = \tfrac{1}{2}\,n\,\mathrm{Var}(W)\ \text{(ReLU)} \]
This is the same product-blows-up problem as the gradient. Signals hold only if r = 1 exactly. r > 1 explodes, r < 1 vanishes. The whole design question is: pick Var(W) so r = 1.
Faded example
Fill in the blanks
The ratio predicts the 50-layer std, with the scaffolding fading: two lines are gone now — fill both.
n = 256
for name, scale in [('He', 2), ('Xavier', 1), ('too_big', 8)]:
varW = scale / n
r = **0.5 * n * varW** # ReLU per-layer variance ratio
print(f'___: r=___, Var ratio r^50 = ___')
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. He's ½·256·(2/256) = 1 is the ONLY ratio that survives depth.
Worked example
Compute the per-layer ReLU variance ratio r = ½·n·Var(W) for three inits at n = 256, then raise to the 50th power. This is a closed-form prediction of the experiment we'll run — runnable:
n = 256
for name, scale in [('He', 2), ('Xavier', 1), ('too_big', 8)]:
varW = scale / n
r = 0.5 * n * varW # ReLU per-layer variance ratio
print(f'{name}: r={r}, Var ratio r^50 = {r**50:.3e}')He r = 1 (holds); Xavier r = 0.5 (r⁵⁰ = 8.9e-16); too_big r = 4 (r⁵⁰ = 1.3e+30)
Why: He's ½·256·(2/256) = 1 is the ONLY ratio that survives depth. Xavier's r = 0.5 means the variance halves each layer — 8.9e-16 after 50, whose square root ≈ 3e-8 matches the measured std 1.47e-8. The algebra predicts the experiment.
| init | r = ½·n·Var(W) | Var ratio r⁵⁰ (verified) |
|---|---|---|
| He (2/n) | 1.0 | 1.000e+00 (stable) |
| Xavier (1/n) | 0.5 | 8.882e-16 (vanishes) |
| too_big (8/n) | 4.0 | 1.268e+30 (explodes) |
Section
Part 4 of 7 — variance preservation
Concept
For a linear (or tanh-like, symmetric) network, we want each layer's output variance to equal its input variance — so signals neither grow nor shrink through depth. Set Var(z) = Var(a):
\[ \mathrm{Var}(z) = n_{\text{in}}\,\mathrm{Var}(W)\,\mathrm{Var}(a) \;\overset{!}{=}\; \mathrm{Var}(a) \]
Ranking
Put in order
Put the moves of Solve for Var(W): the forward condition into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Var(a) > 0, so divide it out. The activation variance drops away, leaving a pure condition on W and the fan-in.
Worked example
Cancel Var(a) from both sides
Why: Var(a) > 0, so divide it out. The activation variance drops away, leaving a pure condition on W and the fan-in.
\[ n_{\text{in}}\,\mathrm{Var}(W) = 1 \]
Isolate Var(W)
Why: Divide by the fan-in n_in. This is the forward-preservation initializer.
\[ \mathrm{Var}(W) = \frac{1}{n_{\text{in}}} \]
Backward gives the mirror condition
Why: Propagating the gradient backward is the same algebra with the transpose, so it wants Var(W) = 1/n_out. Both can't hold unless n_in = n_out.
\[ \mathrm{Var}(W) = \frac{1}{n_{\text{out}}}\ \text{(backward)} \]
Blank canvas
Draw it
Draw what Solve for Var(W): the forward condition just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
Forward wants 1/n_in; backward wants 1/n_out. Glorot & Bengio split the difference with the harmonic-style average — average the two denominators:
\[ \boxed{\;\mathrm{Var}(W)_{\text{Xavier}} = \frac{2}{n_{\text{in}} + n_{\text{out}}}\;} \]
When n_in = n_out = n this reduces to 2/(2n) = 1/n — exactly the forward condition. Xavier is built for symmetric activations (tanh, sigmoid) that don't distort variance. ReLU does distort it — that gap is Part 5.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n.
Substitute carefully: the denominator is n_in + n_out = 2n.
Why: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.
Trap
Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.
Var(W) = 2/n on a square layer
Why: Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.
Substitute carefully: the denominator is n_in + n_out = 2n.
Var(W) = 2/(2n) = 1/n
Why: On a square layer Xavier's per-weight variance is 1/n — the same forward-preservation value. The factor of 2 in Xavier's numerator only offsets the SUM of two fan sizes; it is NOT the ReLU factor of 2 (that's He, coming up).
Break the constraint
Discussion prompt
The rule this trap just fixed:
Substitute carefully: the denominator is n_in + n_out = 2n.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Plugging n_in = n_out = n into 2/(n_in+n_out) gives 2/(n+n) = 2/(2n) = 1/n, NOT 2/n. Forgetting to add the two n's in the denominator doubles the variance and starts the signal growing.
Section
Part 5 of 7 — the ReLU factor of 2
Intuition
ReLU sets every negative pre-activation to 0. For a symmetric, mean-zero z, that's half the values zeroed. Half your variance is deleted at every layer.
So a variance-preserving linear init (Xavier's 1/n) is too weak once ReLU is in the loop: the signal loses half its power per layer and decays. He's fix is a single factor — put the missing half back by doubling Var(W). Let's prove the 'half' claim, then derive the 2.
Explain it
Discussion prompt
Explain ReLU throws away half the signal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
ReLU sets every negative pre-activation to 0. For a symmetric, mean-zero z, that's half the values zeroed. Half your variance is deleted at every layer.
Estimation
Predict first
For symmetric mean-zero z, relu(z) = max(0, z) keeps the positive half untouched and zeros the rest. Claim: E[relu(z)²] = ½·Var(z). Measure it on 2 million draws:
Commit before you compute: what does ReLU keeps exactly half the variance come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: E[relu(z)²] ≈ 0.5006 ≈ ½·Var(z)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Half the z's are zeroed (contribute nothing); the surviving half carry the full second moment, and by symmetry that's exactly half the total.
Worked example
For symmetric mean-zero z, relu(z) = max(0, z) keeps the positive half untouched and zeros the rest. Claim: E[relu(z)²] = ½·Var(z). Measure it on 2 million draws:
import numpy as np
rng = np.random.default_rng(0)
z = rng.normal(size=2_000_000) # mean 0, Var 1
a = np.maximum(0, z) # ReLU
print('Var(z):', round(float(np.var(z)), 4))
print('E[relu(z)^2]:', round(float(np.mean(a**2)), 4))
print('ratio:', round(float(np.mean(a**2) / np.var(z)), 4))E[relu(z)²] ≈ 0.5006 ≈ ½·Var(z)
Why: Half the z's are zeroed (contribute nothing); the surviving half carry the full second moment, and by symmetry that's exactly half the total. The ratio 0.5008 confirms E[relu²] = ½·Var(z). Verified.
| quantity | value (verified) |
|---|---|
| Var(z) | 0.9997 |
| E[relu(z)²] (variance after ReLU) | 0.5006 |
| ratio E[relu²] / Var(z) | 0.5008 (≈ ½) |
Ranking
Put in order
Put the moves of Derive He: put the half back into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The pre-activation still has Var(z) = n·Var(W)·Var(a).
Worked example
Forward relation with a ReLU layer
Why: The pre-activation still has Var(z) = n·Var(W)·Var(a). But the activation a = relu(z) now carries only HALF that variance, from the previous slide.
\[ \mathrm{Var}(a^{(\ell)}) = \tfrac{1}{2}\,\mathrm{Var}(z^{(\ell)}) = \tfrac{1}{2}\, n\,\mathrm{Var}(W)\,\mathrm{Var}(a^{(\ell-1)}) \]
Demand the activation variance holds
Why: Set Var(a^(ℓ)) = Var(a^(ℓ-1)) so the signal is stable through depth, then cancel that common variance.
\[ \tfrac{1}{2}\, n\,\mathrm{Var}(W) = 1 \]
Solve for Var(W) — the factor of 2 appears
Why: Multiply both sides by 2 and divide by n. The 2 is exactly the compensation for ReLU zeroing half the units.
\[ \boxed{\;\mathrm{Var}(W)_{\text{He}} = \frac{2}{n_{\text{in}}}\;} \]
Notation
Annotate
From Derive He: put the half back — read this one piece at a time. What is each part doing?
On: \( \boxed{\;\mathrm{Var}(W)_{\text{He}} = \frac{2}{n_{\text{in}}}\;} \)
Concept
\[ \underbrace{\mathrm{Var}(W) = \frac{2}{n_{\text{in}}+n_{\text{out}}}}_{\text{Xavier — tanh/sigmoid}} \qquad \underbrace{\mathrm{Var}(W) = \frac{2}{n_{\text{in}}}}_{\text{He — ReLU}} \]
The only difference that matters for ReLU: He's numerator-2 sits over a single n_in, so it's about twice Xavier's variance on a square layer (2/n vs 1/n). That extra factor of 2 is the half of the signal ReLU deletes, restored.
| init | Var(W), square layer | std = √Var, n=256 | activation to use |
|---|---|---|---|
| Xavier | 1/n | 0.0625 | tanh, sigmoid |
| He | 2/n | 0.088388 | ReLU |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Xavier is the classic initializer, so use it everywhere — including deep ReLU nets.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer.
Use He (2/n_in) for ReLU — the factor of 2 refills the deleted half each layer.
Why: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer. Verified: std decays to 1.47e-08 by layer 50 — the signal is gone.
Trap
Xavier is the classic initializer, so use it everywhere — including deep ReLU nets.
Var(W) = 1/n on 50 ReLU layers
Why: Each ReLU halves the variance and Xavier's 1/n only preserves the LINEAR part, so the activation loses ~half its power per layer. Verified: std decays to 1.47e-08 by layer 50 — the signal is gone.
Use He (2/n_in) for ReLU — the factor of 2 refills the deleted half each layer.
Var(W) = 2/n keeps activation std ≈ O(1)
Why: Verified: He holds std ≈ 0.49 at layer 50 while Xavier vanishes. Rule: match the initializer to the activation — ReLU ⇒ He, tanh/sigmoid ⇒ Xavier.
Section
Part 6 of 7 — put it under load
Intuition
'Keep the std O(1)' isn't cosmetic. If activations vanish, the network's output barely responds to its input — nothing to learn from. If they explode, you get inf and NaN on the very first forward pass.
The same holds backward: an O(1) gradient at every layer means every layer gets a signal it can act on. Vanish and the early layers freeze; explode and the update overshoots. O(1) is the only regime where all L layers train together.
Concept
The gradient flows backward through Wᵀ and through the ReLU mask (which zeroes the same half of coordinates). So the gradient variance obeys the mirror of the forward recursion — and it, too, is stabilized by He's 2/n.
That's the point of a good init: it keeps both the forward signal and the backward gradient at O(1) through depth, so early layers get a usable gradient. Let's measure the backward pass directly.
Fill the middle
Fill in the blanks
From He stabilizes the backward gradient too — one line has had its right-hand side removed. Put it back.
import numpy as np
def grad_std(scale, depth=50, w=256, seed=1):
rng = np.random.default_rng(seed)
a = rng.normal(size=w); Ws, masks = [], []
for _ in range(depth): # forward, cache masks
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
z = W @ a; m = (z > 0).astype(float)
a = z * m; Ws.append(W); masks.append(m)
d = rng.normal(size=w) # backward from output
for L in range(depth - 1, -1, -1):
d = (d * masks[L]); d = Ws[L].T @ d
return d.std()
for name, s in [('He', 2), ('Xavier', 1)]:
print(name, f'___')
Why: d is what everything below it consumes, so the wrong expression here fails later and somewhere else. The backward pass sees the same halving from the ReLU mask (line 8) and the same Wᵀ scaling.
Worked example
Forward-propagate through 50 ReLU layers caching the masks, then back-propagate a unit gradient and read its std at the input. He vs Xavier — complete and runnable:
import numpy as np
def grad_std(scale, depth=50, w=256, seed=1):
rng = np.random.default_rng(seed)
a = rng.normal(size=w); Ws, masks = [], []
for _ in range(depth): # forward, cache masks
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
z = W @ a; m = (z > 0).astype(float)
a = z * m; Ws.append(W); masks.append(m)
d = rng.normal(size=w) # backward from output
for L in range(depth - 1, -1, -1):
d = (d * masks[L]); d = Ws[L].T @ d
return d.std()
for name, s in [('He', 2), ('Xavier', 1)]:
print(name, f'{grad_std(s):.3e}')He grad std ≈ 1.43 (alive); Xavier ≈ 4.26e-08 (vanished)
Why: The backward pass sees the same halving from the ReLU mask (line 8) and the same Wᵀ scaling. He keeps the gradient O(1) to the input layer; Xavier lets it die — so early layers never learn. Verified.
| init | gradient std at input layer (verified) |
|---|---|
| He (2/n) | 1.429e+00 (stable) |
| Xavier (1/n) | 4.260e-08 (vanished) |
Fill the middle
Fill in the blanks
From 50 ReLU layers, four initializations — one line has had its right-hand side removed. Put it back.
import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a) # ReLU
return a.std()
for name, s in [('He', 2), ('Xavier', 1),
('too_big', 8), ('too_small', 0.1)]:
print(name, f'___')
Why: a is what everything below it consumes, so the wrong expression here fails later and somewhere else. Only He's factor-of-2 keeps the std O(1) across 50 ReLU layers.
Worked example
The headline experiment. Push a unit-variance signal through 50 ReLU layers under four scales Var(W) = scale/n and read the final activation std. scale = 2 is He, 1 is Xavier, 8 too big, 0.1 too small. Fully runnable:
import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a) # ReLU
return a.std()
for name, s in [('He', 2), ('Xavier', 1),
('too_big', 8), ('too_small', 0.1)]:
print(name, f'{final_std(s):.2e}')He → 4.93e-01 (stable); Xavier → 1.47e-08 (vanishes); too_big → 5.55e+14 (explodes); too_small → 1.47e-33 (dead)
Why: Only He's factor-of-2 keeps the std O(1) across 50 ReLU layers. Everything else runs away by depth 50 — over-scale explodes, any under-scale collapses. Verified by execution.
| init | Var(W) scale | activation std @ layer 50 (verified) |
|---|---|---|
| He | 2/n | 4.93e-01 (stable) |
| Xavier | 1/n | 1.47e-08 (vanishes) |
| too big | 8/n | 5.55e+14 (explodes) |
| too small | 0.1/n | 1.47e-33 (dead) |
Comparison
Comparison matrix
From 50 ReLU layers, four initializations: refill the Var(W) scale column from what you know. The rest of the table is as it appeared.
| init | Var(W) scale | activation std @ layer 50 (verified) |
|---|---|---|
| He | 2/n | 4.93e-01 (stable) |
| Xavier | 1/n | 1.47e-08 (vanishes) |
| too big | 8/n | 5.55e+14 (explodes) |
| too small | 0.1/n | 1.47e-33 (dead) |
Pattern
Predict first
The table runs: 1 | 8.54e-01 | 6.04e-01 · 2 | 7.98e-01 | 3.99e-01 · 5 | 8.97e-01 | 1.59e-01 · 10 | 9.26e-01 | 2.89e-02 · 25 | 7.58e-01 | 1.31e-04
In Watch it happen layer by layer, given the rows so far: what is the next one — the row where layer is 50?
Correct: 50 | 4.93e-01 | 1.47e-08
| layer | He std | Xavier std |
|---|---|---|
| 1 | 8.54e-01 | 6.04e-01 |
| 2 | 7.98e-01 | 3.99e-01 |
| 5 | 8.97e-01 | 1.59e-01 |
| 10 | 9.26e-01 | 2.89e-02 |
| 25 | 7.58e-01 | 1.31e-04 |
| 50 | 4.93e-01 | 1.47e-08 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The full trace shows Xavier is losing roughly half its variance (≈0.7× std) each layer while He holds.
Worked example
Read the std after each layer instead of only at the end. He hovers around 1; Xavier halves and halves. This is your diagnostic tool — a per-layer std trace. Runnable:
import numpy as np
def trace(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w); out = []
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a)
out.append(a.std())
return out
he, xa = trace(2), trace(1)
for L in (1, 2, 5, 10, 25, 50):
print(f'layer {L:2d} He {he[L-1]:.2e} Xavier {xa[L-1]:.2e}')He stays near 1 (0.85 → 0.49); Xavier decays (0.60 → 1.47e-08)
Why: The full trace shows Xavier is losing roughly half its variance (≈0.7× std) each layer while He holds. Watching per-layer std is exactly how you diagnose a real init bug. Verified.
| layer | He std | Xavier std |
|---|---|---|
| 1 | 8.54e-01 | 6.04e-01 |
| 2 | 7.98e-01 | 3.99e-01 |
| 5 | 8.97e-01 | 1.59e-01 |
| 10 | 9.26e-01 | 2.89e-02 |
| 25 | 7.58e-01 | 1.31e-04 |
| 50 | 4.93e-01 | 1.47e-08 |
Pattern
Step through it
Step through Watch it happen layer by layer one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 7 of 7 — the other failure
Intuition
Even with the right variance in mind, a beginner reaches for the 'neutral' choice: set every weight to 0 (or any single constant). It's not a scaling problem — it's a symmetry problem.
If every neuron in a layer starts identical, they compute the same output, receive the same gradient, and take the same step — forever. They stay clones, so the layer can only ever represent one feature. Let's prove that by hand.
Analogy
Discussion prompt
Explain Zero feels safe — and never learns by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Even with the right variance in mind, a beginner reaches for the 'neutral' choice: set every weight to 0 (or any single constant). It's not a scaling problem — it's a symmetry problem.
Missing information
Discussion prompt
Tiny net: 2 inputs → 3 hidden (ReLU) → 1 output, all weights zero, MSE loss. Forward is all zeros; backward gives identical rows in dW1. Runnable as written:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
dh = dyhat·W2 is the same for all three neurons because W2 is a single constant, so the three hidden units get the SAME gradient. They will update identically and remain clones. Symmetry is never broken. Verified: rows-all-equal = True.
Worked example
Tiny net: 2 inputs → 3 hidden (ReLU) → 1 output, all weights zero, MSE loss. Forward is all zeros; backward gives identical rows in dW1. Runnable as written:
import numpy as np
x = np.array([1.0, 2.0]); y = 1.0
W1 = np.zeros((3, 2)) # hidden weights, all zero
W2 = np.zeros(3) # output weights, all zero
h_pre = W1 @ x # [0, 0, 0]
h = np.maximum(0, h_pre) # [0, 0, 0]
yhat = W2 @ h # 0.0
dyhat = yhat - y # -1.0
dh = dyhat * W2 # [0,0,0] (W2 is zero)
dW1 = np.outer(dh * (h_pre > 0), x) # 3x2
print('h:', h)
print('dW1 rows all equal:', np.allclose(dW1, dW1[0]))
print('dW1:', dW1.tolist())Every row of dW1 is identical (here all zero)
Why: dh = dyhat·W2 is the same for all three neurons because W2 is a single constant, so the three hidden units get the SAME gradient. They will update identically and remain clones. Symmetry is never broken. Verified: rows-all-equal = True.
| quantity | value (verified) |
|---|---|
| hidden activations h | [0, 0, 0] |
| dW1 all rows equal | True |
| dW1 | [[0,0],[0,0],[0,0]] |
Estimation
Predict first
Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:
Commit before you compute: what does Random init breaks the symmetry come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. h_pre = [-0.139, 0.850, 0.188]: the first unit is off (gradient row is zero), the others are on with different gradients.
Worked example
Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:
import numpy as np
rng = np.random.default_rng(0)
x = np.array([1.0, 2.0]); y = 1.0
W1 = rng.normal(size=(3, 2)) * np.sqrt(2.0 / 2) # He
W2 = rng.normal(size=3) * np.sqrt(2.0 / 3)
h_pre = W1 @ x
h = np.maximum(0, h_pre)
yhat = W2 @ h
dh = (yhat - y) * W2
dW1 = np.outer(dh * (h_pre > 0), x)
print('h_pre:', np.round(h_pre, 4))
print('dW1 row0:', np.round(dW1[0], 4))
print('dW1 row1:', np.round(dW1[1], 4))
print('rows identical?', np.allclose(dW1[0], dW1[1]))Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ
Why: h_pre = [-0.139, 0.850, 0.188]: the first unit is off (gradient row is zero), the others are on with different gradients. rows-identical = False — the neurons will learn DIFFERENT features. Random init is what breaks the tie. Verified.
| quantity | value (verified) |
|---|---|
| h_pre | [-0.139, 0.850, 0.188] |
| dW1 row 0 (ReLU off) | [-0.0, -0.0] |
| dW1 row 1 (ReLU on) | [-0.348, -0.696] |
| rows identical? | False |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Neuron 0's pre-activation is negative → ReLU kills it; rows of dW1 differ
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same net, but draw W from a He normal. Now the three neurons see different pre-activations, ReLU keeps a different subset alive, and the gradient rows differ. Runnable:
Anomaly
Predict first
A student writes this, and it looks reasonable:
Zero is a neutral, unbiased starting point, so set all weights to 0.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever.
Initialize randomly from the He (or Xavier) distribution to break symmetry.
Why: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever. The layer collapses to one effective neuron and never learns distinct features. Bias terms alone can't break it — verified: dW1 rows all equal.
Trap
Zero is a neutral, unbiased starting point, so set all weights to 0.
W = zeros → every neuron identical
Why: Identical weights ⇒ identical outputs ⇒ identical gradients ⇒ identical updates, forever. The layer collapses to one effective neuron and never learns distinct features. Bias terms alone can't break it — verified: dW1 rows all equal.
Initialize randomly from the He (or Xavier) distribution to break symmetry.
W ~ N(0, 2/n_in) → neurons differentiate
Why: Random weights make each neuron start differently, so they receive different gradients and learn different features. Verified: dW1 rows differ. Symmetry-breaking is non-negotiable — and it comes free from any random init.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
1.5 sounds harmless, but you apply it 30 times, and 1.5 compounds like interest.; Both come from the same multiplicative product. Fix the per-layer factor and you fix both at once — which is exactly what initialization does.; Clipping treats the symptom; initialization treats the cause. Part 1 does clipping — it's the quick win — then Parts 2–6 build initialization from variance.[−τ, τ] with np.clip(g, -tau, tau).; Xavier is 'the 2/(n_in+n_out) one', so on a square layer with n_in = n_out = n its variance is 2/n.Concept
The zero-init warning is about weights, not biases. Biases are usually initialized to 0 and that's fine — the randomness in W already breaks the neuron symmetry, so identical starting biases don't clone anything.
Only when the weights are also constant do biases matter, and even then a nonzero bias can't rescue the symmetry (every neuron still shares it). So: random weights, zero biases is the standard, safe default.
Section
Pattern, checks, build
Ranking
Put in order
These are the steps of The init / stability toolkit, scrambled. Put them back in order before the next slide shows you.
g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)Var(W) = 2/n_in — the 2 refills the half ReLU deletesVar(W) = 2/(n_in+n_out) — preserves variance for symmetric activationsO(1), not drift toward 0 or ∞Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)Var(W) = 2/n_in — the 2 refills the half ReLU deletesVar(W) = 2/(n_in+n_out) — preserves variance for symmetric activationsO(1), not drift toward 0 or ∞Edge cases
Discussion prompt
The init / stability toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
g·min(1, τ/‖g‖) — caps length, keeps direction (never clamp per component)Var(W) = 2/n_in — the 2 refills the half ReLU deletesVar(W) = 2/(n_in+n_out) — preserves variance for symmetric activationsO(1), not drift toward 0 or ∞Elimination
Eliminate the wrong options
Why do gradients explode or vanish specifically in DEEP networks?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: By the chain rule the gradient at an early layer is a product of one factor per later layer. A product of L numbers near c behaves like c^L: c = 1.5 gives 1.9e5 over 30 layers, c = 0.7 gives 2.3e-5. Only c ≈ 1 survives — which is what initialization targets.
Check
Think about what the chain rule does to magnitudes.
Check your understanding
Why do gradients explode or vanish specifically in DEEP networks?
Answer: A
Why: By the chain rule the gradient at an early layer is a product of one factor per later layer. A product of L numbers near c behaves like c^L: c = 1.5 gives 1.9e5 over 30 layers, c = 0.7 gives 2.3e-5. Only c ≈ 1 survives — which is what initialization targets.
Prediction
Predict first
For a deep network with ReLU activations, the correct initialization variance (square layer, fan-in n) is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: He: Var(W) = 2/n
Why: ReLU zeroes ~half the activations (E[relu²] = ½·Var(z), verified ≈ 0.5006), so He's factor-of-2 variance refills the deleted half and keeps the signal stable — std ≈ 0.49 at layer 50. Xavier's 1/n under-scales for ReLU and the signal decays to 1.47e-8.
Check
Match the initializer to the activation.
Check your understanding
For a deep network with ReLU activations, the correct initialization variance (square layer, fan-in n) is:
Answer: A
Why: ReLU zeroes ~half the activations (E[relu²] = ½·Var(z), verified ≈ 0.5006), so He's factor-of-2 variance refills the deleted half and keeps the signal stable — std ≈ 0.49 at layer 50. Xavier's 1/n under-scales for ReLU and the signal decays to 1.47e-8.
Prediction
Predict first
Initializing all weights of a layer to 0 fails because:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: every neuron gets identical gradients and never differentiates (broken symmetry)
Why: Identical weights → identical outputs → identical gradients → identical updates, forever. We verified dW1 has three identical rows, so the neurons never become distinct and the layer collapses to a single feature. Random init breaks the tie.
Check
Why exactly do zeros fail? Name the mechanism.
Check your understanding
Initializing all weights of a layer to 0 fails because:
Answer: A
Why: Identical weights → identical outputs → identical gradients → identical updates, forever. We verified dW1 has three identical rows, so the neurons never become distinct and the layer collapses to a single feature. Random init breaks the tie.
Elimination
Eliminate the wrong options
Gradient norm clipping rescales g to g·min(1, τ/‖g‖). Compared with clamping each component to [−τ, τ], it:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Scaling the whole vector by one factor keeps every ratio between components fixed, so the direction is unchanged (only shorter). Per-component clamping shrinks big components more than small ones and rotates the vector — verified: [30,40]→[5,5] turns [0.6,0.8] into [0.707,0.707].
Check
Which clip keeps you descending the loss?
Check your understanding
Gradient norm clipping rescales g to g·min(1, τ/‖g‖). Compared with clamping each component to [−τ, τ], it:
Answer: A
Why: Scaling the whole vector by one factor keeps every ratio between components fixed, so the direction is unchanged (only shorter). Per-component clamping shrinks big components more than small ones and rotates the vector — verified: [30,40]→[5,5] turns [0.6,0.8] into [0.707,0.707].
Section
Build it yourself
Concept
Reproduce the headline result yourself: push a signal through 50 ReLU layers under each init and find the one that keeps the activation std stable — then add a norm-clip. You've derived every piece; now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | He forward pass, read final std | W·√(2/n), then ReLU |
| 2 | compare He / Xavier / too-big / too-small | vary the scale |
| 3 | norm-clip a large gradient, keep direction | g·min(1, τ/‖g‖) |
Build rules: type every line yourself, apply ReLU each layer, seed the RNG so your numbers match, and read the final std — only He should stay O(1).
Counterexample
Discussion prompt
Build rules: type every line yourself, apply ReLU each layer, seed the RNG so your numbers match, and read the final std — only He should stay O(1).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: push a signal through 50 ReLU layers with He init (scale = 2). Predict the order of magnitude of the final std out loud before you print.
Hint: each layer W = randn(w, w)·sqrt(2/w), then a = relu(W @ a). Seed with default_rng(0) so you get the same number.
import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a) # ReLU
return a.std()
print('He std @ 50:', f'{final_std(2):.3f}')| init | std @ layer 50 (verified) |
|---|---|
| He (scale 2) | 0.493 (stable) |
| target | O(1) — near 0.5, not 0 or ∞ |
Worked example
Your turn: reuse final_std and sweep all four scales. Predict which vanish and which explode before running.
Hint: scales are 2 (He), 1 (Xavier), 8 (too big), 0.1 (too small). Print with f'{x:.2e}' so the tiny/huge numbers are readable.
import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a)
return a.std()
for name, s in [('He', 2), ('Xavier', 1),
('too_big', 8), ('too_small', 0.1)]:
print(name, f'{final_std(s):.2e}')| init | std @ 50 (verified) |
|---|---|
| He | 4.93e-01 |
| Xavier | 1.47e-08 |
| too_big | 5.55e+14 |
| too_small | 1.47e-33 |
Pattern
Step through it
Step through Milestone 2 — compare four initializations one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: norm-clip a gradient of norm 50 down to τ = 5 and confirm the direction is unchanged. Predict the clipped vector before printing.
Hint: g * min(1, tau/np.linalg.norm(g)), then compare the unit vectors before and after with np.allclose.
import numpy as np
g = np.array([30.0, 40.0]); tau = 5.0
gc = g * min(1.0, tau / np.linalg.norm(g))
print('gc:', gc)
print('norm after:', round(float(np.linalg.norm(gc)), 1))
print('direction kept:',
np.allclose(gc / np.linalg.norm(gc), g / np.linalg.norm(g)))| quantity | value (verified) |
|---|---|
| gc | [3.0, 4.0] |
| ‖gc‖ after | 5.0 |
| direction kept | True |
Trade off
Comparison matrix
From Milestone 3 — clip a gradient: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | value (verified) |
|---|---|
| gc | [3.0, 4.0] |
| ‖gc‖ after | 5.0 |
| direction kept | True |
Concept
import numpy as np
def final_std(scale, depth=50, w=256, seed=0):
rng = np.random.default_rng(seed)
a = rng.normal(size=w)
for _ in range(depth):
W = rng.normal(size=(w, w)) * np.sqrt(scale / w)
a = np.maximum(0, W @ a) # ReLU
return a.std()
for name, s in [('He', 2), ('Xavier', 1),
('too_big', 8), ('too_small', 0.1)]:
print(name, f'{final_std(s):.2e}')
g = np.array([30.0, 40.0])
gc = g * min(1.0, 5.0 / np.linalg.norm(g))
print('clip 50 ->', round(float(np.linalg.norm(gc)), 1),
'dir kept', np.allclose(gc/np.linalg.norm(gc), g/np.linalg.norm(g)))| printed line | value (verified) |
|---|---|
| He / Xavier | 4.93e-01 / 1.47e-08 |
| too_big / too_small | 5.55e+14 / 1.47e-33 |
| clip 50 -> | 5.0 dir kept True |
If only He keeps the signal alive through 50 layers and your clip lands norm 5 with dir kept True — you understand why init and clipping decide whether a deep net trains at all.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| He / Xavier | 4.93e-01 / 1.47e-08 |
| too_big / too_small | 5.55e+14 / 1.47e-33 |
| clip 50 -> | 5.0 dir kept True |
Concept
Slides closed, out loud: explain (1) why the gradient is a product so depth compounds errors, (2) where He's factor of 2 comes from over Xavier, and (3) why norm clipping preserves the descent direction but per-component clamping doesn't.
Stretch: re-derive Xavier's 1/n from the forward variance condition on paper, then swap ReLU for tanh in final_std and confirm Xavier now holds O(1) while He slightly over-scales. That's the exact init-to-activation match the exam probes. Init enables ResNets (Week 42) and transformer LayerNorm (Week 29).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The problem: depth multiplies · Clipping exploding gradients · Variance through a layer · Xavier initialization · He initialization · Backward & the 50-layer test. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
1.5³⁰ ≈ 1.9e5, 0.7³⁰ ≈ 2.3e-5, only 1³⁰ = 1 survivesg·min(1, τ/‖g‖) — caps length, keeps direction; clamping per component rotates itVar(z) = n·Var(W)·Var(a), and get Xavier 1/n from preservation and He 2/n from ReLU's deleted halfdW1 — broken symmetry)std ≈ 0.49, and diagnose init bugs from a per-layer std trace| idea | the one thing to remember |
|---|---|
| depth | gradient is a PRODUCT — factor must be ≈ 1 |
| clip | by NORM: g·min(1, τ/‖g‖), never per component |
| He | Var(W) = 2/n_in — the 2 is ReLU's deleted half |
| Xavier | Var(W) = 2/(n_in+n_out) — for tanh/sigmoid |
| zero init | identical rows ⇒ symmetry never breaks ⇒ never learns |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.