Lesson 20: Hypothesis Testing

USAAIO Lesson 20, from Week 7 on statistics, fully worked. It covers H₀ and H₁ and Type I and Type II errors, derives the p-value as P(data|H₀), and dismantles the misreading of it as P(H₀|data). It then builds a one-sample t-test from the sample mean and the ddof=1 sample standard deviation, through the t₉ tail, to p = 0.0282, which rejects, and verifies it against scipy. From there it covers critical values, the t, chi-squared, and F family with a worked chi-squared test, power against sample size, the 85-against-83 two-proportion test, and multiple testing with Bonferroni on seeded noise. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 111 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Hypothesis Testing

Title

USAAIO · Lesson 20 · Week 7 (Statistics)

Is that result real, or noise? We build the one-sample t-test from x̄ to p-value with no skipped step, verify it against scipy, then meet the whole test family, power, and why 20 tests find 'significance' in pure noise.

2. By the end of this lesson you can

Objectives

  1. State H₀ vs H₁ precisely and name the two error types — Type I (α) and Type II (β)
  2. Define a p-value as P(data this extreme | H₀) and reject the P(H₀|data) misreading
  3. Compute a one-sample t-statistic t=(x̄−μ₀)/(s/√n) by hand, entry by entry, and its two-sided p-value from scipy
  4. Pick the right test — t (means), χ² (categorical counts), F (variances) — and read a critical value
  5. Relate power = 1−β to sample size, and control multiple testing with Bonferroni α/m

3. What survived from PCA Implementation?

Warm-up

Discussion prompt

Before we open Lesson 20: Hypothesis Testing: without looking back, what was the main idea of PCA Implementation, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the PCA algorithm via eigendecomposition or SVD, the explained-variance ratio, choosing k with a scree plot / 95% threshold, PCA for 2D visualization, and its linear-only limitation. Build a PCA class from scratch (both routes) and verify against sklearn on the digits dataset.

4. The framework

Section

Part 1 of 6 — H₀, errors, p-values

5. The running example

Concept

A teammate ships a model and claims its true accuracy is exactly μ₀ = 0.80. You re-measure it on 10 independent held-out folds and want to know: is the claim consistent with what you saw?

foldaccuracy
1–50.82, 0.79, 0.85, 0.83, 0.78
6–100.86, 0.81, 0.84, 0.80, 0.87

The sample average will land a little above 0.80. The whole lesson is one question: is that gap real, or just the luck of which 10 folds we drew?

6. Fill in: accuracy for The running example

Comparison

Comparison matrix

From The running example: refill the accuracy column from what you know. The rest of the table is as it appeared.

foldaccuracy
1–50.82, 0.79, 0.85, 0.83, 0.78
6–100.86, 0.81, 0.84, 0.80, 0.87

7. The courtroom mindset

Intuition

Hypothesis testing works like a trial. The defendant is presumed innocent — that's H₀ (no effect). You never prove innocence; you either gather enough evidence to convict, or you don't.

Convicting an innocent defendant is a Type I error (false positive). Letting a guilty one walk is a Type II error (false negative). Setting α = 0.05 is choosing how much reasonable doubt you demand before convicting.

Key consequence: 'fail to reject H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.

8. Break it if you can: The courtroom mindset

Counterexample

Discussion prompt

Hypothesis testing works like a trial. The defendant is presumed innocent — that's H₀ (no effect). You never prove innocence; you either gather enough evidence to convict, or you don't.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Key consequence: 'fail to reject H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.

9. Null, alternative, and the two errors

Concept

The null H₀ is the boring default — here μ = 0.80, the claim is exactly right. The alternative H₁ is μ ≠ 0.80. We assume H₀, then ask if the data is too surprising to keep it.

H₀ trueH₀ false
reject H₀Type I error (α)correct ✓
keep H₀correct ✓Type II error (β)

Type I = false positive (cry wolf), rate α. Type II = false negative (miss a real effect), rate β. You fix α (usually 0.05) before seeing data.

10. What each one costs: Null, alternative, and the two errors

Trade off

Comparison matrix

From Null, alternative, and the two errors: every row here is a choice with a cost. Fill the H₀ true column, then say which row you would actually pick and what you give up for it.

H₀ trueH₀ false
reject H₀Type I error (α)correct ✓
keep H₀correct ✓Type II error (β)

11. What 'surprising' means

Intuition

Pretend H₀ is true: the model really is at 0.80 and every wobble is pure sampling noise. Then ask: how often would noise alone produce a sample this far from 0.80, or farther?

If the answer is 'rarely' — say 3% of the time — then either we witnessed a rare fluke, or H₀ is wrong. That tail probability is the p-value; small means the data fights the null.

Crucially this is a statement about the data, computed under the assumption H₀ holds. It is not the probability that H₀ is true.

12. By analogy: What 'surprising' means

Analogy

Discussion prompt

Explain What 'surprising' means by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Pretend H₀ is true: the model really is at 0.80 and every wobble is pure sampling noise. Then ask: how often would noise alone produce a sample this far from 0.80, or farther?

13. α is a pre-commitment, not a dial

Concept

α is the false-positive rate you're willing to tolerate, chosen before the data. It defines the rejection region: reject when the evidence would arise less than α of the time under H₀.

Fixing α up front is what makes the guarantee honest. Picking α after seeing p — nudging 0.05 to 0.10 so your result 'passes' — is moving the goalposts, and it voids the 5% false-positive promise.

14. Teach it back: α is a pre-commitment, not a dial

Explain it

Discussion prompt

Explain α is a pre-commitment, not a dial to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

α is the false-positive rate you're willing to tolerate, chosen before the data. It defines the rejection region: reject when the evidence would arise less than α of the time under H₀.

15. The p-value, precisely

Concept

The p-value is the probability, assuming H₀, of a test statistic at least as extreme as the one observed:

\[ p = P\big(\,|T| \ge |t_{\text{obs}}| \;\big|\; H_0 \text{ true}\,\big) \]

Small p ⇒ the observed data is unlikely under H₀ ⇒ reject H₀ at level α when p < α. It is P(data | H₀), never P(H₀ | data).

16. Something is wrong here: is the p-value P(H₀ is true)?

Anomaly

Predict first

A student writes this, and it looks reasonable:

We got p = 0.03, so there is a 3% chance the null hypothesis is true.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).

p = 0.03 means: if H₀ were true, data this extreme would occur only 3% of the time.

Why: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).

17. Trap: is the p-value P(H₀ is true)?

Trap

The trap

We got p = 0.03, so there is a 3% chance the null hypothesis is true.

Read p as P(H₀ | data)

Why: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).

Conclude 'H₀ is 97% likely false'

Why: A posterior like P(H₀|data) needs a PRIOR on H₀ and Bayes' rule. The p-value alone can never deliver it.

The fix

p = 0.03 means: if H₀ were true, data this extreme would occur only 3% of the time.

Read p as P(data this extreme | H₀)

Why: The direction of conditioning is fixed by construction: we hold H₀ and integrate the tail of the statistic.

p < α ⇒ reject H₀ (a decision, not a probability of H₀)

Why: Rejecting controls the false-positive RATE at α over many experiments. It says nothing about the probability this particular H₀ is true. The single most-tested stats misconception.

18. Build the t-test

Section

Part 2 of 6 — x̄ to p, no skipped step

19. Guess the shape of the answer: Step 1 — the sample mean, by hand

Estimation

Predict first

The 10 accuracies sum, then divide by n = 10. Add them in order — no shortcut:

Commit before you compute: what does Step 1 — the sample mean, by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: x̄ = 8.25 / 10 = 0.825

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The sample mean sits 0.025 above the claimed μ₀ = 0.80.

20. Step 1 — the sample mean, by hand

Worked example

The 10 accuracies sum, then divide by n = 10. Add them in order — no shortcut:

\[ \sum_i x_i = 0.82+0.79+0.85+0.83+0.78+0.86+0.81+0.84+0.80+0.87 = 8.25 \]

x̄ = 8.25 / 10 = 0.825

Why: The sample mean sits 0.025 above the claimed μ₀ = 0.80. That gap is the raw signal — but we can't judge it until we scale it by the noise.

quantityvalue (verified)
Σ xᵢ8.25
n10
x̄ = Σxᵢ / n0.825

21. Watch it run: Step 1 — the sample mean, by hand

Pattern

Step through it

Step through Step 1 — the sample mean, by hand one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is Σ xᵢ
  2. Step 2: quantity is n
  3. Step 3: quantity is x̄ = Σxᵢ / n

22. What has to be given first: Step 2 — the sample std, deviation by…

Missing information

Discussion prompt

Spread uses the sample standard deviation with ddof = 1 (divide by n−1, not n). First every deviation dᵢ = xᵢ − x̄ and its square:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Sum the last column. This is the total squared spread around the mean.

23. Step 2 — the sample std, deviation by deviation

Worked example

Spread uses the sample standard deviation with ddof = 1 (divide by n−1, not n). First every deviation dᵢ = xᵢ − x̄ and its square:

xᵢdᵢ = xᵢ − 0.825dᵢ²
0.82−0.0050.000025
0.79−0.0350.001225
0.85+0.0250.000625
0.83+0.0050.000025
0.78−0.0450.002025
0.86+0.0350.001225
0.81−0.0150.000225
0.84+0.0150.000225
0.80−0.0250.000625
0.87+0.0450.002025

Σ dᵢ² = 0.00825

Why: Sum the last column. This is the total squared spread around the mean.

s² = 0.00825 / (10−1) = 0.0009167, s = √0.0009167 = 0.030277

Why: Dividing by n−1 (not n) makes s² an UNBIASED estimate of the population variance — that is exactly what ddof=1 does. Matches np.std(ddof=1) = 0.030277.

24. What happens as it grows: Step 2 — the sample std, deviation by…

Scale up

Step through it

Step through Step 2 — the sample std, deviation by deviation and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: xᵢ is 0.82
  2. Step 2: xᵢ is 0.79
  3. Step 3: xᵢ is 0.85
  4. Step 4: xᵢ is 0.83
  5. Step 5: xᵢ is 0.78
  6. Step 6: xᵢ is 0.86
  7. Step 7: xᵢ is 0.81
  8. Step 8: xᵢ is 0.84
  9. Step 9: xᵢ is 0.80
  10. Step 10: xᵢ is 0.87

25. Predict the next row: Step 3 — standard error, then the t-statistic

Pattern

Predict first

The table runs: x̄ − μ₀ | 0.025 · SE = s/√n | 0.0095743

In Step 3 — standard error, then the t-statistic, given the rows so far: what is the next one — the row where quantity is t = (x̄−μ₀)/SE?

Correct: t = (x̄−μ₀)/SE | 2.6112

quantityvalue (verified)
x̄ − μ₀0.025
SE = s/√n0.0095743
t = (x̄−μ₀)/SE2.6112

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Averaging 10 folds cuts the noise by √10 — that √n is why bigger samples resolve smaller effects.

26. Step 3 — standard error, then the t-statistic

Worked example

The standard error is how much x̄ itself wobbles: the sample sd shrunk by √n. Then t counts how many standard errors the gap x̄ − μ₀ spans.

SE = s/√n = 0.030277/√10 = 0.030277/3.1623 = 0.0095743

Why: Averaging 10 folds cuts the noise by √10 — that √n is why bigger samples resolve smaller effects.

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n}} = \frac{0.825 - 0.80}{0.0095743} = \frac{0.025}{0.0095743} \]

t = 2.6112

Why: The sample mean is 2.61 standard errors above the claimed value. Now we ask how often noise alone reaches |t| ≥ 2.61 — that's the p-value.

quantityvalue (verified)
x̄ − μ₀0.025
SE = s/√n0.0095743
t = (x̄−μ₀)/SE2.6112

27. Decode the notation: Step 3 — standard error, then the t-statistic

Notation

Annotate

From Step 3 — standard error, then the t-statistic — read this one piece at a time. What is each part doing?

On: \( t = \frac{\bar x - \mu_0}{s/\sqrt{n}} = \frac{0.825 - 0.80}{0.0095743} = \frac{0.025}{0.0095743} \)

  • Averaging 10 folds cuts the noise by √10 — that √n is why bigger samples resolve smaller effects.
  • The sample mean is 2.61 standard errors above the claimed value. Now we ask how often noise alone reaches |t| ≥ 2.61 — that's the p-value.

28. Restore the missing line: Step 3 in code — the t-statistic

Fill the middle

Fill in the blanks

From Step 3 in code — the t-statistic — one line has had its right-hand side removed. Put it back.

import numpy as np
# 10 held-out accuracy scores of a model claimed to hit 0.80
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
xbar = data.mean() # sample mean
s = data.std(ddof=1) # SAMPLE sd (ddof=1)
se = s / np.sqrt(n) # standard error
t = (xbar - mu0) / se # t statistic
print(round(xbar,4), round(s,4), round(t,4))

Why: s is what everything below it consumes, so the wrong expression here fails later and somewhere else. Drop ddof=1 and std divides by n, shrinking s, inflating t and faking significance.

29. Step 3 in code — the t-statistic

Worked example

Every hand number, reproduced in NumPy. This block is self-contained — it re-imports and re-defines the data:

import numpy as np
# 10 held-out accuracy scores of a model claimed to hit 0.80
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
xbar = data.mean()                       # sample mean
s = data.std(ddof=1)                     # SAMPLE sd (ddof=1)
se = s / np.sqrt(n)                       # standard error
t = (xbar - mu0) / se                     # t statistic
print(round(xbar,4), round(s,4), round(t,4))

ddof=1 is the whole game on line 7

Why: Drop ddof=1 and std divides by n, shrinking s, inflating t and faking significance. The sample sd MUST use n−1.

printedvalue
xbar0.825
s0.0303
t2.6112

30. Watch it run: Step 3 in code — the t-statistic

Pattern

Step through it

Step through Step 3 in code — the t-statistic one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: printed is xbar
  2. Step 2: printed is s
  3. Step 3: printed is t

31. What the t-test assumes

Concept

The one-sample t is valid when the observations are independent and the sample mean is approximately normal — true if the data itself is roughly normal, or (by the Central Limit Theorem) if n is large enough.

With n = 10 we lean on approximate normality of the folds. Heavy skew or outliers at small n break it — then reach for a nonparametric test (e.g. Wilcoxon) instead.

Independence is the assumption people violate most: reusing overlapping folds or correlated samples silently inflates significance.

32. Why the t distribution, not the normal

Concept

We estimated the sd from the same 10 points, so t carries extra uncertainty. Under H₀ it follows a Student-t distribution with df = n−1 = 9 — like a normal but with heavier tails that shrink toward the normal as n grows.

\[ \text{under } H_0:\quad T = \frac{\bar X - \mu_0}{S/\sqrt{n}} \sim t_{\,df=n-1} \]

To get the p-value we measure how much of that t₉ distribution lies beyond ±|t_obs| — the two tails, because H₁ is ≠ (either direction counts).

33. Heavier tails, and why they matter

Intuition

The t₉ curve looks like a bell but with fatter tails than the normal — it admits that with only 10 points our sd estimate could be off, so extreme values are a bit more likely.

That is why the two-sided 5% cutoff is ±2.262 for t₉, wider than the normal's ±1.96. Fewer data → fatter tails → a higher bar to clear before you reject.

As n → ∞ the sd estimate sharpens and t collapses onto the standard normal. Small samples are exactly where using t instead of z matters.

34. Finish it with less help: Step 4 — tail area to p-value

Faded example

Fill in the blanks

Step 4 — tail area to p-value, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
df = n - 1
one_tail = stats.t.sf(abs(t), df=df) # upper-tail area
p = **2 * one_tail** # two-sided
print(round(t,4), df, round(one_tail,4), round(p,4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. By symmetry of t₉, the lower tail beyond −|t| is also 0.0141.

35. Step 4 — tail area to p-value

Worked example

stats.t.sf(|t|, df) is the upper-tail area beyond |t|. Double it for the two-sided p (both tails are equal by symmetry):

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
df = n - 1
one_tail = stats.t.sf(abs(t), df=df)      # upper-tail area
p = 2 * one_tail                          # two-sided
print(round(t,4), df, round(one_tail,4), round(p,4))

one tail = 0.0141, so p = 2 × 0.0141 = 0.0282

Why: By symmetry of t₉, the lower tail beyond −|t| is also 0.0141. Two-sided p adds both: 0.0282.

quantityvalue (verified)
t2.6112
df9
upper tail P(T ≥ |t|)0.0141
p = 2 × upper tail0.0282

36. Work backwards from the answer: Step 4 — tail area to p-value

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

one tail = 0.0141, so p = 2 × 0.0141 = 0.0282

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

stats.t.sf(|t|, df) is the upper-tail area beyond |t|. Double it for the two-sided p (both tails are equal by symmetry):

37. One-sided or two-sided?

Concept

Our H₁ is μ ≠ 0.80 — an effect in either direction — so we count both tails: p = 2 × 0.0141 = 0.0282. This is the honest default when you didn't pre-commit to a direction.

A one-sided H₁: μ > 0.80 counts only the upper tail: p = 0.0141. It's more powerful, but only legitimate if you fixed the direction before seeing the data — otherwise it's a p-halving trick.

Rule of thumb for the exam: default to two-sided unless the problem explicitly justifies a direction.

38. The decision: reject H₀

Concept

We fixed α = 0.05 in advance. The data gave p = 0.0282.

\[ p = 0.0282 \;<\; 0.05 = \alpha \;\Longrightarrow\; \boxed{\text{reject } H_0} \]

So the evidence is inconsistent with μ = 0.80: the model's true accuracy is very likely above the claimed value. (Reject controls the false-positive rate at 5% — it doesn't make us 97% certain.)

39. The same decision by critical value

Worked example

Equivalent to comparing p to α: compare |t| to the critical value — the cutoff that leaves α/2 = 2.5% in each tail of t₉:

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
t = (data.mean()-0.80)/(data.std(ddof=1)/np.sqrt(len(data)))
tcrit = stats.t.ppf(0.975, df=9)          # two-sided 5% cutoff
print(round(t,4), round(tcrit,4))
print('reject H0:', abs(t) > tcrit)

t_crit = t.ppf(0.975, df=9) = 2.2622

Why: ppf is the inverse CDF: the value with 97.5% of t₉ below it, so 2.5% above. Two tails → 5% total = α.

|t| = 2.6112 > 2.2622 = t_crit ⇒ reject

Why: |t| lands in the rejection region beyond the cutoff, so p < α automatically. The p-value route and the critical-value route ALWAYS agree.

quantityvalue (verified)
|t| observed2.6112
t_crit (0.975, df=9)2.2622
|t| > t_crit ?True → reject H₀

40. Inspect it line by line: The same decision by critical value

Error analysis

Annotate

Walk the callouts on The same decision by critical value. Each one is a place this is easy to get subtly wrong.

  • ppf is the inverse CDF: the value with 97.5% of t₉ below it, so 2.5% above. Two tails → 5% total = α.
  • |t| lands in the rejection region beyond the cutoff, so p < α automatically. The p-value route and the critical-value route ALWAYS agree.

41. The confidence interval says the same thing

Concept

A 95% confidence interval for the true mean is x̄ ± t_crit · SE — the set of μ₀ values this data would not reject at α = 0.05. It carries the same information as the p-value, plus a range.

\[ \bar x \pm t_{0.975,\,9}\cdot \frac{s}{\sqrt n} = 0.825 \pm 2.2622\cdot 0.0095743 \]

The margin is 2.2622 × 0.0095743 = 0.02166. If μ₀ = 0.80 falls outside the interval, that's exactly equivalent to rejecting H₀.

42. Restore the missing line: Compute the 95% CI

Fill the middle

Fill in the blanks

From Compute the 95% CI — one line has had its right-hand side removed. Put it back.

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
xbar = data.mean()
se = data.std(ddof=1) / np.sqrt(n)
tc = stats.t.ppf(0.975, df=n-1) # 2.5% in each tail
lo, hi = xbar - tcse, xbar + tcse
print(round(lo,4), round(hi,4))
print('0.80 inside?', lo <= 0.80 <= hi)

Why: se is what everything below it consumes, so the wrong expression here fails later and somewhere else. The plausible range for the true accuracy.

43. Compute the 95% CI

Worked example

Same ingredients as the test — x̄, SE, and the t₉ critical value — assembled into an interval. Self-contained:

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
xbar = data.mean()
se = data.std(ddof=1) / np.sqrt(n)
tc = stats.t.ppf(0.975, df=n-1)           # 2.5% in each tail
lo, hi = xbar - tc*se, xbar + tc*se
print(round(lo,4), round(hi,4))
print('0.80 inside?', lo <= 0.80 <= hi)

CI = 0.825 ± 0.02166 = [0.8033, 0.8467]

Why: The plausible range for the true accuracy. Its lower end 0.8033 sits ABOVE 0.80.

0.80 is NOT in [0.8033, 0.8467] ⇒ reject H₀

Why: The interval excludes the null value, the mirror image of p = 0.0282 < 0.05. CI and test always agree — the CI just also tells you the effect is roughly 0.003–0.047 above 0.80.

quantityvalue (verified)
x̄0.825
t_crit · SE (margin)0.02166
95% CI[0.8033, 0.8467]
contains μ₀ = 0.80 ?No → reject H₀

44. Draw the shape of it: Compute the 95% CI

Blank canvas

Draw it

Draw what Compute the 95% CI just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

45. Guess the shape of the answer: Verify the whole test against scipy

Estimation

Predict first

Our from-scratch t and p must match the library's ttest_1samp exactly. Self-contained:

Commit before you compute: what does Verify the whole test against scipy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: ttest_1samp returns (statistic, pvalue) — both match

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. np.isclose confirms agreement to floating-point tolerance.

46. Verify the whole test against scipy

Worked example

Our from-scratch t and p must match the library's ttest_1samp exactly. Self-contained:

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
res = stats.ttest_1samp(data, mu0)        # library reference
print(round(t,4), round(p,4))
print(round(res.statistic,4), round(res.pvalue,4))
print(np.isclose(t,res.statistic) and np.isclose(p,res.pvalue))

ttest_1samp returns (statistic, pvalue) — both match

Why: np.isclose confirms agreement to floating-point tolerance. If your hand pipeline matches scipy, you understand every gear inside the black box.

sourcetp
from scratch2.61120.0282
scipy.ttest_1samp2.61120.0282
np.isclose matchTrueTrue

47. Something is wrong here: population sd (ddof=0) in a t-test

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use data.std() — NumPy's default — for the standard deviation in t.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: NumPy's default divides by n, giving the POPULATION sd.

Use data.std(ddof=1) — the sample sd — because we estimated the mean from the same data.

Why: NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.

48. Trap: population sd (ddof=0) in a t-test

Trap

The trap

Use data.std() — NumPy's default — for the standard deviation in t.

s = data.std() (ddof=0, divides by n)

Why: NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.

t = 0.025 / (0.02872/√10) = 2.752 → overstated

Why: A smaller denominator inflates t and shrinks p — you'd report a stronger effect than the data supports. Wrong distribution, wrong number.

The fix

Use data.std(ddof=1) — the sample sd — because we estimated the mean from the same data.

s = data.std(ddof=1) (divides by n−1)

Why: Dividing by n−1 gives an unbiased variance estimate: s = 0.030277, matching the t₉ theory.

t = 0.025 / (0.030277/√10) = 2.6112 ✓

Why: This is the value that agrees with scipy and with the t distribution. In a t-test, ALWAYS ddof=1.

49. Break it on purpose: population sd (ddof=0) in a t-test

Break the constraint

Discussion prompt

The rule this trap just fixed:

Dividing by n−1 gives an unbiased variance estimate: s = 0.030277, matching the t₉ theory.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.

50. The test family

Section

Part 3 of 6 — t, χ², F

51. A test statistic is a signal-to-noise ratio

Intuition

Look back at t = (x̄ − μ₀)/(s/√n): the numerator is the signal (how far you are from the null) and the denominator is the noise (how much that estimate wobbles). Every test statistic has this shape.

That's why the same effect gets more 'significant' as n grows: the noise s/√n shrinks, the ratio blows up. And why a noisy metric needs a bigger effect to clear the bar.

So building any test = decide the signal, decide how to measure its noise, form the ratio, look up its tail. χ² and F just measure signal and noise differently.

52. One recipe, three statistics

Concept

Every test is the same loop: form a statistic, then read its tail under H₀'s known distribution. What changes is what you compare.

testcomparesH₀ distributionuse when
t-testmeansStudent-tcontinuous data, small n
χ² testcategorical countschi-squaredfrequencies / contingency tables
F-testvariancesFANOVA, comparing spread

We built the t. Next a concrete χ² on categorical counts, so the family isn't just a table.

53. Finish it with less help: A chi-squared goodness-of-fit test

Faded example

Fill in the blanks

A chi-squared goodness-of-fit test, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from scipy import stats
observed = np.array([22,17,20,18,12,11]) # 100 die rolls, faces 1..6
expected = np.array([100/6]*6) # fair-die expectation
chi2 = np.sum((observed-expected)2/expected)
p =
stats.chi2.sf(chi2, df=5)** # df = k-1 = 5
print(round(chi2,4), round(p,4))
print(stats.chisquare(observed).pvalue.round(4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Six faces → 5 degrees of freedom (the counts must sum to 100, costing one).

54. A chi-squared goodness-of-fit test

Worked example

Roll a die 100 times; is it fair? H₀: each face has expected count 100/6 ≈ 16.67. The statistic sums (observed−expected)²/expected over faces:

import numpy as np
from scipy import stats
observed = np.array([22,17,20,18,12,11])  # 100 die rolls, faces 1..6
expected = np.array([100/6]*6)             # fair-die expectation
chi2 = np.sum((observed-expected)**2/expected)
p = stats.chi2.sf(chi2, df=5)              # df = k-1 = 5
print(round(chi2,4), round(p,4))
print(stats.chisquare(observed).pvalue.round(4))

χ² = Σ (O−E)²/E = 5.72, df = k−1 = 5

Why: Six faces → 5 degrees of freedom (the counts must sum to 100, costing one). Each term measures how far one face's count strayed from 16.67.

p = chi2.sf(5.72, df=5) = 0.3344 → fail to reject

Why: p ≫ 0.05: this much deviation is ordinary for a fair die over 100 rolls. No evidence of bias. scipy.chisquare agrees.

faceobservedexpected(O−E)²/E
12216.671.7067
21716.670.0067
3–620,18,12,1116.67 each0.6667, 0.1067, 1.3067, 1.9267
Σ100100χ² = 5.72

55. Draw the shape of it: A chi-squared goodness-of-fit test

Blank canvas

Draw it

Draw what A chi-squared goodness-of-fit test just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

56. A two-sample t-test (Welch)

Worked example

Compare two models on 5 folds each. H₀: equal means. Welch's t divides the mean gap by the combined standard error √(s_A²/n_A + s_B²/n_B) — no equal-variance assumption:

import numpy as np
from scipy import stats
A = np.array([0.82,0.79,0.85,0.83,0.78])  # model A folds
B = np.array([0.80,0.75,0.79,0.77,0.76])  # model B folds
sew = np.sqrt(A.var(ddof=1)/len(A) + B.var(ddof=1)/len(B))
t = (A.mean() - B.mean()) / sew           # Welch statistic
ref = stats.ttest_ind(A, B, equal_var=False)
print(round(t,4), round(ref.statistic,4), round(ref.pvalue,4))

means 0.814 vs 0.774, gap 0.040; combined SE = √(0.00083/5 + 0.00043/5) = 0.01587

Why: Each group contributes its own variance over its own n. The gap spans 0.040/0.01587 = 2.52 standard errors.

t = 2.5198, p = 0.0386 (df ≈ 7.27) ⇒ reject H₀

Why: p < 0.05: the two models differ significantly on this data. Matches scipy.ttest_ind(equal_var=False) exactly.

quantityvalue (verified)
mean A / mean B0.814 / 0.774
combined SE0.01587
Welch t2.5198
p-value0.0386

57. Fill in: value (verified) for A two-sample t-test (Welch)

Comparison

Comparison matrix

From A two-sample t-test (Welch): refill the value (verified) column from what you know. The rest of the table is as it appeared.

quantityvalue (verified)
mean A / mean B0.814 / 0.774
combined SE0.01587
Welch t2.5198
p-value0.0386

58. Restore the missing line: An F-test: three models at once (ANOVA)

Fill the middle

Fill in the blanks

From An F-test: three models at once (ANOVA) — one line has had its right-hand side removed. Put it back.

import numpy as np
from scipy import stats
g1 = np.array([0.82,0.79,0.85]) # three models,
g2 = np.array([0.80,0.75,0.79]) # three folds each
g3 = np.array([0.88,0.90,0.86])
F, p = stats.f_oneway(g1, g2, g3) # one-way ANOVA
print(round(F,4), round(p,4))

Why: g2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Between-group spread is 11.4× the within-group noise — far more than chance would produce.

59. An F-test: three models at once (ANOVA)

Worked example

With three models you can't just run pairwise t-tests (that's multiple testing!). One-way ANOVA uses an F statistic — the ratio of between-group spread to within-group spread — to test H₀: all means equal.

import numpy as np
from scipy import stats
g1 = np.array([0.82,0.79,0.85])           # three models,
g2 = np.array([0.80,0.75,0.79])           # three folds each
g3 = np.array([0.88,0.90,0.86])
F, p = stats.f_oneway(g1, g2, g3)         # one-way ANOVA
print(round(F,4), round(p,4))

F = 11.4, p = 0.009 ⇒ reject H₀

Why: Between-group spread is 11.4× the within-group noise — far more than chance would produce. At least one model's mean differs. A large F lives in the upper tail of the F distribution, giving a small p.

groupfoldsmean
g10.82, 0.79, 0.850.820
g20.80, 0.75, 0.790.780
g30.88, 0.90, 0.860.880
F / p—11.4 / 0.009

60. Power & sample size

Section

Part 4 of 6 — what you can resolve

61. Significance is not 'which number is bigger'

Intuition

Model A scores 85%, model B 83%. A is higher — but is the 2-point gap real, or the coin-flips of which test examples each happened to get right?

The answer depends entirely on how many examples you tested on. The exact same 2-point gap can be pure noise at n=100 and rock-solid at n=10000. The gap alone is not a test.

62. Power = 1 − β

Concept

Power is the probability of correctly rejecting H₀ when H₁ is actually true — the chance of catching a real effect.

\[ \text{power} = 1 - \beta = P(\text{reject } H_0 \mid H_1 \text{ true}) \]

Power rises with the effect size and with the sample size (via the √n in the SE). Too-small n → low power → you miss real effects; huge n → even trivial effects become 'significant'.

63. Effect size: significance without magnitude is empty

Concept

A p-value tells you an effect is detectable; effect size tells you if it's big enough to care about. For a one-sample mean, Cohen's d reports the gap in sd units:

\[ d = \frac{\bar x - \mu_0}{s} = \frac{0.825 - 0.80}{0.030277} = 0.826 \]

d ≈ 0.83 is a large effect (rough guide: 0.2 small, 0.5 medium, 0.8 large). Always report the effect size next to p — with huge n, a trivial d can still be 'significant'.

64. Type II error β, made concrete

Concept

Suppose the true accuracy really is 0.825 (a 0.025 effect) with sd 0.030. Even so, with only n = 10 we miss it — fail to reject — some fraction of the time. That miss rate is β.

\[ \text{power} = 1 - \beta \approx 0.74 \quad\Longrightarrow\quad \beta \approx 0.26 \]

So even a real, large effect goes undetected about 26% of the time at n = 10. Underpowered studies quietly produce false negatives — the fix is more data, which shrinks β.

65. Predict the next row: Power climbs with n

Pattern

Predict first

The table runs: 10 | 2.108 | 0.5589 · 30 | 3.651 | 0.9546

In Power climbs with n, given the rows so far: what is the next one — the row where n is 100?

Correct: 100 | 6.667 | 1.0000

neffect (z-units)power = 1−β
102.1080.5589
303.6510.9546
1006.6671.0000

Why: The relationship between the columns, not the individual numbers, is what generates the next row. With only 10 folds we'd miss this real effect 44% of the time.

66. Power climbs with n

Worked example

Suppose the true accuracy really is 0.02 above the null, with sd 0.03. How likely are we to detect it at α = 0.05 for different n?

import numpy as np
from scipy import stats
delta, sigma, alpha = 0.02, 0.03, 0.05     # true shift, sd, level
zc = stats.norm.ppf(1 - alpha/2)           # 1.96
for n in [10, 30, 100]:
    se = sigma / np.sqrt(n)
    ncp = delta / se                       # effect in z-units
    power = stats.norm.sf(zc-ncp) + stats.norm.cdf(-zc-ncp)
    print(n, round(ncp,3), round(power,4))

n=10 → power 0.56; n=30 → 0.95; n=100 → 1.00

Why: With only 10 folds we'd miss this real effect 44% of the time. At n=30 we catch it 95% of the time. The effect never changed — our resolution did.

neffect (z-units)power = 1−β
102.1080.5589
303.6510.9546
1006.6671.0000

67. What has to be given first: 85% vs 83%: same gap, opposite verdicts

Missing information

Discussion prompt

A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Hand-check for n=100: 0.84·0.16 = 0.1344, times 2/100 = 0.002688, root = 0.051846.

68. 85% vs 83%: same gap, opposite verdicts

Worked example

A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:

import numpy as np
from scipy import stats
p1, p2 = 0.85, 0.83                        # two model accuracies
for n in [100, 1000, 10000]:               # n per group
    pp = (p1 + p2) / 2                      # pooled proportion
    se = np.sqrt(pp*(1-pp)*(2/n))
    z = (p1 - p2) / se
    print(n, round(z,4), round(2*stats.norm.sf(abs(z)),6))

SE(100) = √(0.84·0.16·2/100) = √0.002688 = 0.051846

Why: Hand-check for n=100: 0.84·0.16 = 0.1344, times 2/100 = 0.002688, root = 0.051846.

n=100: z=0.386, p=0.70 (noise). n=10000: z=3.858, p=0.0001 (decisive)

Why: The identical 2-point gap is indistinguishable from chance at n=100 and overwhelming at n=10000. 'A > B' is a ranking, not a test.

n per groupzp-valueverdict
1000.38580.6997not significant
10001.21990.2225not significant
100003.85760.0001significant

69. Work backwards from the answer: 85% vs 83%: same gap, opposite verdicts

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

n=100: z=0.386, p=0.70 (noise). n=10000: z=3.858, p=0.0001 (decisive)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:

70. Something is wrong here: 85 > 83, so ship A

Anomaly

Predict first

A student writes this, and it looks reasonable:

Model A scored higher on the test set, so A is the better model — ship it.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.

Test the difference; report significance (or a confidence interval), not just the raw numbers.

Why: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.

71. Trap: 85 > 83, so ship A

Trap

The trap

Model A scored higher on the test set, so A is the better model — ship it.

Pick A from the point estimates alone

Why: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.

Assume the ranking is stable

Why: On a fresh test set the ranking could easily flip. You reported a coin-flip as a finding.

The fix

Test the difference; report significance (or a confidence interval), not just the raw numbers.

Run a two-proportion test at your fixed α

Why: Turns 'A looks higher' into a decision that accounts for how many examples backed it.

Significant only with enough n; else say 'no detectable difference'

Why: The same gap is meaningless at n=100 and decisive at n=10000. Effect size AND n both matter.

72. Multiple testing

Section

Part 5 of 6 — the p-hacking trap

73. False positives pile up

Concept

Each test at α = 0.05 has a 5% false-positive rate. Run m independent tests and the chance of at least one false alarm is:

\[ P(\ge 1\text{ false positive}) = 1 - (1-\alpha)^m \]

m tests1 − 0.95ᵐP(≥ 1 false positive)
10.950.0500
50.95⁵0.2262
100.95¹⁰0.4013
200.95²⁰0.6415

74. Watch it run: False positives pile up

Pattern

Step through it

Step through False positives pile up one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: m tests is 1
  2. Step 2: m tests is 5
  3. Step 3: m tests is 10
  4. Step 4: m tests is 20

75. Bonferroni: tighten the threshold

Concept

To keep the family-wise false-positive rate at α across m tests, compare each p-value to a stricter cutoff:

\[ \alpha_{\text{corrected}} = \frac{\alpha}{m} \qquad(\text{Bonferroni}) \]

Simple and conservative. For large m, FDR (Benjamini–Hochberg) controls the false-discovery rate with more power — but Bonferroni is the exam default.

76. Guess the shape of the answer: p-hacking on pure noise

Estimation

Predict first

Run 20 one-sample t-tests on completely random data — no real effect anywhere — and count the 'significant' ones. Seeded so you see these exact results:

Commit before you compute: what does p-hacking on pure noise come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 0 survive Bonferroni's 0.05/20 = 0.0025

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The smallest p (0.0111) exceeds 0.0025, so the corrected test correctly reports nothing significant — which is the truth.

77. p-hacking on pure noise

Worked example

Run 20 one-sample t-tests on completely random data — no real effect anywhere — and count the 'significant' ones. Seeded so you see these exact results:

import numpy as np
from scipy import stats
rng = np.random.default_rng(7)            # fixed seed -> reproducible
pvals = [stats.ttest_1samp(rng.normal(size=30), 0).pvalue
         for _ in range(20)]              # 20 tests on PURE noise
print(sum(p < 0.05      for p in pvals))  # naive count
print(sum(p < 0.05/20   for p in pvals))  # Bonferroni count

4 of 20 'significant' at 0.05 — every one a false positive

Why: Seed 7 produced 4 p-values below 0.05 (min p = 0.0111) on noise. You EXPECT ~1 by chance (20×0.05); this draw gave 4.

0 survive Bonferroni's 0.05/20 = 0.0025

Why: The smallest p (0.0111) exceeds 0.0025, so the corrected test correctly reports nothing significant — which is the truth.

threshold'significant' tests
α = 0.05 (naive)4 of 20 (all false)
α = 0.05/20 (Bonferroni)0 of 20 (correct)
min p-value observed0.0111

78. Inspect it line by line: p-hacking on pure noise

Error analysis

Annotate

Walk the callouts on p-hacking on pure noise. Each one is a place this is easy to get subtly wrong.

  • Seed 7 produced 4 p-values below 0.05 (min p = 0.0111) on noise. You EXPECT ~1 by chance (20×0.05); this draw gave 4.
  • The smallest p (0.0111) exceeds 0.0025, so the corrected test correctly reports nothing significant — which is the truth.

79. When Bonferroni is too strict: FDR

Concept

Bonferroni controls the chance of any false positive, so with thousands of tests (genomics, feature screens) it becomes so strict it misses real effects — low power.

Benjamini–Hochberg (FDR) instead controls the expected fraction of discoveries that are false. Sort the m p-values ascending and keep test k while p₍ₖ₎ ≤ (k/m)·α.

On our 20 noise tests, BH's rank-1 threshold is (1/20)·0.05 = 0.0025; the smallest p 0.0111 exceeds it, so BH also reports nothing — correct. FDR simply gains power when there are many true effects mixed in.

80. Teach it back: When Bonferroni is too strict: FDR

Explain it

Discussion prompt

Explain When Bonferroni is too strict: FDR to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Bonferroni controls the chance of any false positive, so with thousands of tests (genomics, feature screens) it becomes so strict it misses real effects — low power.

81. Something is wrong here: report the one that worked

Anomaly

Predict first

A student writes this, and it looks reasonable:

Tried 20 features; feature 7 came out significant (p = 0.011), so report that discovery.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64.

Correct for the number of tests before claiming significance.

Why: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64. The 'winner' is very likely noise.

82. Trap: report the one that worked

Trap

The trap

Tried 20 features; feature 7 came out significant (p = 0.011), so report that discovery.

Cherry-pick the significant test, drop the other 19

Why: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64. The 'winner' is very likely noise.

Publish without mentioning the 19

Why: Hiding the tests you ran is p-hacking. The result won't replicate on fresh data.

The fix

Correct for the number of tests before claiming significance.

Compare each p to α/m (Bonferroni), or control FDR

Why: p = 0.011 vs 0.05/20 = 0.0025 → not significant. Honest correction dissolves the fake finding.

Report all tests run, corrected

Why: Disclosing m and applying the correction is what separates a real discovery from noise mining.

83. Which of these survive contact with Lesson 20: Hypothesis Testing?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Key consequence: 'fail to reject H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.; Crucially this is a statement about the data, computed under the assumption H₀ holds. It is not the probability that H₀ is true.; The p-value is the probability, assuming H₀, of a test statistic at least as extreme as the one observed:
Breaks
We got p = 0.03, so there is a 3% chance the null hypothesis is true.; Use data.std() — NumPy's default — for the standard deviation in t.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 20: Hypothesis Testing puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

84. Recipe & checks

Section

Part 6 of 6

85. Rebuild the recipe: The hypothesis-testing recipe

Ranking

Put in order

These are the steps of The hypothesis-testing recipe, scrambled. Put them back in order before the next slide shows you.

  1. State H₀ and H₁; fix α before looking at the data
  2. Pick the test: t (means), χ² (categorical counts), F (variance)
  3. Compute the statistic — e.g. t=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀
  4. Decide: reject if p < α (or |t| > t_crit); report the effect size, not just 'significant'
  5. Correct for multiple tests with α/m (Bonferroni) or FDR if you ran many

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

86. The hypothesis-testing recipe

Pattern

  1. State H₀ and H₁; fix α before looking at the data
  2. Pick the test: t (means), χ² (categorical counts), F (variance)
  3. Compute the statistic — e.g. t=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀
  4. Decide: reject if p < α (or |t| > t_crit); report the effect size, not just 'significant'
  5. Correct for multiple tests with α/m (Bonferroni) or FDR if you ran many

87. Where does it stop working: The hypothesis-testing recipe

Edge cases

Discussion prompt

The hypothesis-testing recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. State H₀ and H₁; fix α before looking at the data
  2. Pick the test: t (means), χ² (categorical counts), F (variance)
  3. Compute the statistic — e.g. t=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀
  4. Decide: reject if p < α (or |t| > t_crit); report the effect size, not just 'significant'
  5. Correct for multiple tests with α/m (Bonferroni) or FDR if you ran many

88. Rule out three: Check yourself — the p-value

Elimination

Eliminate the wrong options

Our t-test gave p = 0.0282. Which statement is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. If μ were 0.80, a sample this far from 0.80 would occur about 2.8% of the time
  • B. There is a 2.8% probability that H₀ (μ = 0.80) is true
  • C. There is a 97.2% probability the model's accuracy exceeds 0.80
  • D. The effect size is 2.8%

Survives elimination: A

Why: A p-value is P(data this extreme | H₀). Assuming μ = 0.80, a sample mean at least 2.61 SE away arises ~2.8% of the time — rare, so we reject H₀. It never states P(H₀).

89. Check yourself — the p-value

Check

Read the conditioning precisely.

Check your understanding

Our t-test gave p = 0.0282. Which statement is correct?

  • A. If μ were 0.80, a sample this far from 0.80 would occur about 2.8% of the time (correct)
  • B. There is a 2.8% probability that H₀ (μ = 0.80) is true
  • C. There is a 97.2% probability the model's accuracy exceeds 0.80
  • D. The effect size is 2.8%

Answer: A

Why: A p-value is P(data this extreme | H₀). Assuming μ = 0.80, a sample mean at least 2.61 SE away arises ~2.8% of the time — rare, so we reject H₀. It never states P(H₀).

Why B tempts people
That's P(H₀ | data), the reversed conditioning. Getting a posterior needs a prior and Bayes' rule — the p-value is not a posterior.
Why C tempts people
Also a posterior claim (about H₁). 1 − p is not P(H₁ true); the p-value can't deliver the probability an effect is real.
Why D tempts people
The p-value measures surprise under H₀, not effect size. Here the effect is x̄−μ₀ = 0.025, and a large n could make a tiny effect have p = 0.028.

90. Answer it before you see the options: Check yourself — ddof

Prediction

Predict first

Computing t = (x̄ − μ₀)/(s/√n) in NumPy, what should s be?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: data.std(ddof=1) — divide by n−1 (sample sd)

Why: The mean was estimated from the same data, so the unbiased variance divides by n−1: ddof=1. That gives s = 0.030277, t = 2.6112, matching scipy and the t₉ theory.

91. Check yourself — ddof

Check

Which standard deviation belongs in a one-sample t?

Check your understanding

Computing t = (x̄ − μ₀)/(s/√n) in NumPy, what should s be?

  • A. data.std(ddof=1) — divide by n−1 (sample sd) (correct)
  • B. data.std() — NumPy's default (divide by n)
  • C. data.var() — the variance, not the sd
  • D. data.std(ddof=1) / n — divide by n once more

Answer: A

Why: The mean was estimated from the same data, so the unbiased variance divides by n−1: ddof=1. That gives s = 0.030277, t = 2.6112, matching scipy and the t₉ theory.

Why B tempts people
ddof=0 is the population sd (divide by n). It underestimates spread here (0.02872), inflating t to 2.752 and faking a stronger result.
Why C tempts people
var is s², not s. The t formula needs the standard deviation; using the variance is dimensionally wrong.
Why D tempts people
The √n is already in the SE denominator (s/√n). Dividing s by n as well double-counts the sample size and collapses t toward 0.

92. How sure are you: Check yourself — sample size

Commit first

Predict first

Model A 85% vs B 83% is statistically significant when…

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: the sample is large enough (e.g. ~10000/group), not at ~100/group

Why: The two-proportion test gives p = 0.70 at n=100 (noise) but p = 0.0001 at n=10000. Significance of a fixed gap depends on the sample size behind it.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

93. Check yourself — sample size

Check

When is a fixed gap meaningful?

Check your understanding

Model A 85% vs B 83% is statistically significant when…

  • A. the sample is large enough (e.g. ~10000/group), not at ~100/group (correct)
  • B. always — 85 > 83
  • C. never — 2 points is too small to matter
  • D. only if both models use the same architecture

Answer: A

Why: The two-proportion test gives p = 0.70 at n=100 (noise) but p = 0.0001 at n=10000. Significance of a fixed gap depends on the sample size behind it.

Why B tempts people
Raw ordering ignores sampling noise; at n=100 the gap (z=0.386) is indistinguishable from chance.
Why C tempts people
A 2-point gap CAN be significant — with enough data it's decisive (p=0.0001 at n=10000).
Why D tempts people
Architecture is irrelevant to whether the accuracy difference is statistically distinguishable; the test needs only the counts and n.

94. Answer it before you see the options: Check yourself — multiple testing

Prediction

Predict first

You run 20 independent tests at α = 0.05 on data with no real effects. Roughly how many come out 'significant', and how do you control it?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: About 1 by chance; Bonferroni's α/20 = 0.0025 suppresses it

Why: Each test has a 5% false-positive rate, so the expected count is 20 × 0.05 = 1 (our seed happened to give 4). Bonferroni lowers the threshold to 0.05/20 = 0.0025 to hold the family-wise rate at 5%.

95. Check yourself — multiple testing

Check

Account for how many tests you ran.

Check your understanding

You run 20 independent tests at α = 0.05 on data with no real effects. Roughly how many come out 'significant', and how do you control it?

  • A. About 1 by chance; Bonferroni's α/20 = 0.0025 suppresses it (correct)
  • B. 0 — there are no real effects
  • C. All 20
  • D. Exactly 5

Answer: A

Why: Each test has a 5% false-positive rate, so the expected count is 20 × 0.05 = 1 (our seed happened to give 4). Bonferroni lowers the threshold to 0.05/20 = 0.0025 to hold the family-wise rate at 5%.

Why B tempts people
Even with no real effect, the 5% Type I rate produces false positives — that IS the multiple-testing problem. P(≥1) = 0.64 at m=20.
Why C tempts people
Each test independently has only a 5% false-positive chance; all 20 firing by chance is astronomically unlikely (0.05²⁰).
Why D tempts people
5 would be a 25% rate. The expected count is 20 × 0.05 = 1, not 5.

96. Rule out three: Check yourself — confidence intervals

Elimination

Eliminate the wrong options

Our 95% CI for the true accuracy is [0.8033, 0.8467]. What does that tell you about testing H₀: μ = 0.80 at α = 0.05?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Reject H₀ — 0.80 lies outside the interval
  • B. Fail to reject H₀ — the interval is wide
  • C. There's a 95% probability the true mean is exactly 0.825
  • D. The p-value must be greater than 0.05

Survives elimination: A

Why: A 95% CI is exactly the set of μ₀ values a two-sided test at α=0.05 would NOT reject. Since 0.80 is outside [0.8033, 0.8467], the test rejects — matching p = 0.0282 < 0.05.

97. Check yourself — confidence intervals

Check

Connect the interval to the test.

Check your understanding

Our 95% CI for the true accuracy is [0.8033, 0.8467]. What does that tell you about testing H₀: μ = 0.80 at α = 0.05?

  • A. Reject H₀ — 0.80 lies outside the interval (correct)
  • B. Fail to reject H₀ — the interval is wide
  • C. There's a 95% probability the true mean is exactly 0.825
  • D. The p-value must be greater than 0.05

Answer: A

Why: A 95% CI is exactly the set of μ₀ values a two-sided test at α=0.05 would NOT reject. Since 0.80 is outside [0.8033, 0.8467], the test rejects — matching p = 0.0282 < 0.05.

Why B tempts people
Interval width doesn't decide the test; what matters is whether μ₀ is inside. 0.80 is outside, so we reject.
Why C tempts people
The CI is a range of plausible means, not a probability spike on 0.825. The interval either contains the fixed true μ or not; 95% refers to the procedure's long-run coverage.
Why D tempts people
The opposite: because 0.80 is outside the 95% CI, p must be BELOW 0.05. CI-exclusion and p < α are equivalent.

98. Your turn: build a t-test

Section

Project

99. Project: one-sample t-test from scratch

Concept

Rebuild the whole test on the 10-fold accuracy data using only NumPy for the math, then prove it matches scipy.stats.ttest_1samp. You derived every piece — now assemble it yourself.

#requirementtool
1t = (x̄ − μ₀)/(s/√n)np.mean, np.std(ddof=1)
2two-sided p from the t distributionscipy.stats.t.sf
3verify vs scipystats.ttest_1samp

Build rules: type every line yourself, use ddof=1 for the sample sd, and double the one-tailed area for a two-sided p. Predict each output before you run it.

100. By analogy: Project: one-sample t-test from scratch

Analogy

Discussion prompt

Explain Project: one-sample t-test from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Rebuild the whole test on the 10-fold accuracy data using only NumPy for the math, then prove it matches scipy.stats.ttest_1samp. You derived every piece — now assemble it yourself.

101. Milestone 1 — the t-statistic

Worked example

Your turn: compute t for the 10-fold sample against μ₀ = 0.80. Predict its sign first (is x̄ above or below 0.80?).

Hint: t = (data.mean()-mu0)/(data.std(ddof=1)/np.sqrt(n)). Expect t > 0 because x̄ = 0.825 > 0.80.

import numpy as np
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
print(round(t, 4))
quantityvalue
x̄0.825
s (ddof=1)0.0303
t2.6112

102. Milestone 2 — the p-value

Worked example

Your turn: turn t into a two-sided p-value with n−1 = 9 degrees of freedom. Predict: above or below 0.05?

Hint: p = 2*stats.t.sf(abs(t), df=n-1). The survival function sf is the upper-tail area; double it for two sides.

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
t = (data.mean() - 0.80) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
print(round(p, 4))
quantityvalue
df9
p-value0.0282
reject H₀?yes (p < 0.05)

103. Watch it run: Milestone 2 — the p-value

Pattern

Step through it

Step through Milestone 2 — the p-value one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is df
  2. Step 2: quantity is p-value
  3. Step 3: quantity is reject H₀?

104. Milestone 3 — verify vs scipy

Worked example

Your turn: confirm your t and p match scipy.stats.ttest_1samp. Predict: exact match or just close?

Hint: stats.ttest_1samp(data, 0.80) returns (statistic, pvalue); compare with np.isclose.

import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
t = (data.mean() - 0.80) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
res = stats.ttest_1samp(data, 0.80)
print(np.isclose(t,res.statistic) and np.isclose(p,res.pvalue))
sourcetp
your code2.61120.0282
scipy2.61120.0282
np.isclose matchTrueTrue

105. What each one costs: Milestone 3 — verify vs scipy

Trade off

Comparison matrix

From Milestone 3 — verify vs scipy: every row here is a choice with a cost. Fill the t column, then say which row you would actually pick and what you give up for it.

sourcetp
your code2.61120.0282
scipy2.61120.0282
np.isclose matchTrueTrue

106. The full program

Concept

import numpy as np
from scipy import stats

def one_sample_t(data, mu0):
    n = len(data)
    t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
    p = 2 * stats.t.sf(abs(t), df=n-1)
    return t, p

data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
t, p = one_sample_t(data, 0.80)
ref = stats.ttest_1samp(data, 0.80)
print('t =', round(t,4), ' p =', round(p,4))
print('matches scipy:',
      np.isclose(t, ref.statistic) and np.isclose(p, ref.pvalue))
printed linevalue
t = 2.6112 p =0.0282
matches scipy:True

If your t reads 2.6112, p reads 0.0282, and scipy agrees — you built a hypothesis test from the definition up.

107. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
t = 2.6112 p =0.0282
matches scipy:True

108. Show it off

Concept

Slides closed, out loud: explain (1) why a p-value is P(data | H₀) and not P(H₀), (2) why ddof=1 and not ddof=0, and (3) why the same 85-vs-83 gap flips significance with n.

Stretch: extend one_sample_t to a two-sample t-test (Welch's), then run 20 tests on noise and watch Bonferroni erase the false positives — the exact robustness gap the exam likes to probe.

109. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why a p-value is P(data | H₀) and not P(H₀), (2) why ddof=1 and not ddof=0, and (3) why the same 85-vs-83 gap flips significance with n.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Stretch: extend one_sample_t to a two-sample t-test (Welch's), then run 20 tests on noise and watch Bonferroni erase the false positives — the exact robustness gap the exam likes to probe.

110. Connect it up: Lesson 20: Hypothesis Testing

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The framework · Build the t-test · The test family · Power & sample size · Multiple testing · Recipe & checks. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

111. What you can do now

Recap

ideathe one thing to remember
p-valueP(data this extreme | H₀), not P(H₀)
t-statistic(x̄−μ₀)/(s/√n) with ddof=1
decisionreject if p < α ⇔ |t| > t_crit
significancedepends on sample size, not just the gap
power1 − β; grows with n and effect size
multiple testscorrect with α/m (Bonferroni) or FDR

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 20 (Week 7 — Hypothesis Testing) — Barron · USAAIO Round 2 Preparation, 2026
  2. scipy.stats.ttest_1samp / t.sf / t.ppf / chi2.sf / norm.sf
  3. Every t, p, z, χ², power, and multiple-testing count produced by real execution — numpy 2.2.6 + scipy 1.16, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108