USAAIO Lesson 20, from Week 7 on statistics, fully worked. It covers H₀ and H₁ and Type I and Type II errors, derives the p-value as P(data|H₀), and dismantles the misreading of it as P(H₀|data). It then builds a one-sample t-test from the sample mean and the ddof=1 sample standard deviation, through the t₉ tail, to p = 0.0282, which rejects, and verifies it against scipy. From there it covers critical values, the t, chi-squared, and F family with a worked chi-squared test, power against sample size, the 85-against-83 two-proportion test, and multiple testing with Bonferroni on seeded noise. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 111 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 20 · Week 7 (Statistics)
Is that result real, or noise? We build the one-sample t-test from x̄ to p-value with no skipped step, verify it against scipy, then meet the whole test family, power, and why 20 tests find 'significance' in pure noise.
Objectives
H₀ vs H₁ precisely and name the two error types — Type I (α) and Type II (β)P(data this extreme | H₀) and reject the P(H₀|data) misreadingt=(x̄−μ₀)/(s/√n) by hand, entry by entry, and its two-sided p-value from scipyα/mWarm-up
Discussion prompt
Before we open Lesson 20: Hypothesis Testing: without looking back, what was the main idea of PCA Implementation, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the PCA algorithm via eigendecomposition or SVD, the explained-variance ratio, choosing k with a scree plot / 95% threshold, PCA for 2D visualization, and its linear-only limitation. Build a PCA class from scratch (both routes) and verify against sklearn on the digits dataset.
Section
Part 1 of 6 — H₀, errors, p-values
Concept
A teammate ships a model and claims its true accuracy is exactly μ₀ = 0.80. You re-measure it on 10 independent held-out folds and want to know: is the claim consistent with what you saw?
| fold | accuracy |
|---|---|
| 1–5 | 0.82, 0.79, 0.85, 0.83, 0.78 |
| 6–10 | 0.86, 0.81, 0.84, 0.80, 0.87 |
The sample average will land a little above 0.80. The whole lesson is one question: is that gap real, or just the luck of which 10 folds we drew?
Comparison
Comparison matrix
From The running example: refill the accuracy column from what you know. The rest of the table is as it appeared.
| fold | accuracy |
|---|---|
| 1–5 | 0.82, 0.79, 0.85, 0.83, 0.78 |
| 6–10 | 0.86, 0.81, 0.84, 0.80, 0.87 |
Intuition
Hypothesis testing works like a trial. The defendant is presumed innocent — that's H₀ (no effect). You never prove innocence; you either gather enough evidence to convict, or you don't.
Convicting an innocent defendant is a Type I error (false positive). Letting a guilty one walk is a Type II error (false negative). Setting α = 0.05 is choosing how much reasonable doubt you demand before convicting.
Key consequence: 'fail to reject H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.
Counterexample
Discussion prompt
Hypothesis testing works like a trial. The defendant is presumed innocent — that's H₀ (no effect). You never prove innocence; you either gather enough evidence to convict, or you don't.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Key consequence: 'fail to reject H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.
Concept
The null H₀ is the boring default — here μ = 0.80, the claim is exactly right. The alternative H₁ is μ ≠ 0.80. We assume H₀, then ask if the data is too surprising to keep it.
| H₀ true | H₀ false | |
|---|---|---|
| reject H₀ | Type I error (α) | correct ✓ |
| keep H₀ | correct ✓ | Type II error (β) |
Type I = false positive (cry wolf), rate α. Type II = false negative (miss a real effect), rate β. You fix α (usually 0.05) before seeing data.
Trade off
Comparison matrix
From Null, alternative, and the two errors: every row here is a choice with a cost. Fill the H₀ true column, then say which row you would actually pick and what you give up for it.
| H₀ true | H₀ false | |
|---|---|---|
| reject H₀ | Type I error (α) | correct ✓ |
| keep H₀ | correct ✓ | Type II error (β) |
Intuition
Pretend H₀ is true: the model really is at 0.80 and every wobble is pure sampling noise. Then ask: how often would noise alone produce a sample this far from 0.80, or farther?
If the answer is 'rarely' — say 3% of the time — then either we witnessed a rare fluke, or H₀ is wrong. That tail probability is the p-value; small means the data fights the null.
Crucially this is a statement about the data, computed under the assumption H₀ holds. It is not the probability that H₀ is true.
Analogy
Discussion prompt
Explain What 'surprising' means by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Pretend H₀ is true: the model really is at 0.80 and every wobble is pure sampling noise. Then ask: how often would noise alone produce a sample this far from 0.80, or farther?
Concept
α is the false-positive rate you're willing to tolerate, chosen before the data. It defines the rejection region: reject when the evidence would arise less than α of the time under H₀.
Fixing α up front is what makes the guarantee honest. Picking α after seeing p — nudging 0.05 to 0.10 so your result 'passes' — is moving the goalposts, and it voids the 5% false-positive promise.
Explain it
Discussion prompt
Explain α is a pre-commitment, not a dial to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
α is the false-positive rate you're willing to tolerate, chosen before the data. It defines the rejection region: reject when the evidence would arise less than α of the time under H₀.
Concept
The p-value is the probability, assuming H₀, of a test statistic at least as extreme as the one observed:
\[ p = P\big(\,|T| \ge |t_{\text{obs}}| \;\big|\; H_0 \text{ true}\,\big) \]
Small p ⇒ the observed data is unlikely under H₀ ⇒ reject H₀ at level α when p < α. It is P(data | H₀), never P(H₀ | data).
Anomaly
Predict first
A student writes this, and it looks reasonable:
We got p = 0.03, so there is a 3% chance the null hypothesis is true.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).
p = 0.03 means: if H₀ were true, data this extreme would occur only 3% of the time.
Why: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).
Trap
We got p = 0.03, so there is a 3% chance the null hypothesis is true.
Read p as P(H₀ | data)
Why: This flips the conditioning. The p-value is computed by ASSUMING H₀ and measuring the data's tail — it is P(data | H₀).
Conclude 'H₀ is 97% likely false'
Why: A posterior like P(H₀|data) needs a PRIOR on H₀ and Bayes' rule. The p-value alone can never deliver it.
p = 0.03 means: if H₀ were true, data this extreme would occur only 3% of the time.
Read p as P(data this extreme | H₀)
Why: The direction of conditioning is fixed by construction: we hold H₀ and integrate the tail of the statistic.
p < α ⇒ reject H₀ (a decision, not a probability of H₀)
Why: Rejecting controls the false-positive RATE at α over many experiments. It says nothing about the probability this particular H₀ is true. The single most-tested stats misconception.
Section
Part 2 of 6 — x̄ to p, no skipped step
Estimation
Predict first
The 10 accuracies sum, then divide by n = 10. Add them in order — no shortcut:
Commit before you compute: what does Step 1 — the sample mean, by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: x̄ = 8.25 / 10 = 0.825
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The sample mean sits 0.025 above the claimed μ₀ = 0.80.
Worked example
The 10 accuracies sum, then divide by n = 10. Add them in order — no shortcut:
\[ \sum_i x_i = 0.82+0.79+0.85+0.83+0.78+0.86+0.81+0.84+0.80+0.87 = 8.25 \]
x̄ = 8.25 / 10 = 0.825
Why: The sample mean sits 0.025 above the claimed μ₀ = 0.80. That gap is the raw signal — but we can't judge it until we scale it by the noise.
| quantity | value (verified) |
|---|---|
| Σ xᵢ | 8.25 |
| n | 10 |
| x̄ = Σxᵢ / n | 0.825 |
Pattern
Step through it
Step through Step 1 — the sample mean, by hand one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
Spread uses the sample standard deviation with ddof = 1 (divide by n−1, not n). First every deviation dᵢ = xᵢ − x̄ and its square:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Sum the last column. This is the total squared spread around the mean.
Worked example
Spread uses the sample standard deviation with ddof = 1 (divide by n−1, not n). First every deviation dᵢ = xᵢ − x̄ and its square:
| xᵢ | dᵢ = xᵢ − 0.825 | dᵢ² |
|---|---|---|
| 0.82 | −0.005 | 0.000025 |
| 0.79 | −0.035 | 0.001225 |
| 0.85 | +0.025 | 0.000625 |
| 0.83 | +0.005 | 0.000025 |
| 0.78 | −0.045 | 0.002025 |
| 0.86 | +0.035 | 0.001225 |
| 0.81 | −0.015 | 0.000225 |
| 0.84 | +0.015 | 0.000225 |
| 0.80 | −0.025 | 0.000625 |
| 0.87 | +0.045 | 0.002025 |
Σ dᵢ² = 0.00825
Why: Sum the last column. This is the total squared spread around the mean.
s² = 0.00825 / (10−1) = 0.0009167, s = √0.0009167 = 0.030277
Why: Dividing by n−1 (not n) makes s² an UNBIASED estimate of the population variance — that is exactly what ddof=1 does. Matches np.std(ddof=1) = 0.030277.
Scale up
Step through it
Step through Step 2 — the sample std, deviation by deviation and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Pattern
Predict first
The table runs: x̄ − μ₀ | 0.025 · SE = s/√n | 0.0095743
In Step 3 — standard error, then the t-statistic, given the rows so far: what is the next one — the row where quantity is t = (x̄−μ₀)/SE?
Correct: t = (x̄−μ₀)/SE | 2.6112
| quantity | value (verified) |
|---|---|
| x̄ − μ₀ | 0.025 |
| SE = s/√n | 0.0095743 |
| t = (x̄−μ₀)/SE | 2.6112 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Averaging 10 folds cuts the noise by √10 — that √n is why bigger samples resolve smaller effects.
Worked example
The standard error is how much x̄ itself wobbles: the sample sd shrunk by √n. Then t counts how many standard errors the gap x̄ − μ₀ spans.
SE = s/√n = 0.030277/√10 = 0.030277/3.1623 = 0.0095743
Why: Averaging 10 folds cuts the noise by √10 — that √n is why bigger samples resolve smaller effects.
\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n}} = \frac{0.825 - 0.80}{0.0095743} = \frac{0.025}{0.0095743} \]
t = 2.6112
Why: The sample mean is 2.61 standard errors above the claimed value. Now we ask how often noise alone reaches |t| ≥ 2.61 — that's the p-value.
| quantity | value (verified) |
|---|---|
| x̄ − μ₀ | 0.025 |
| SE = s/√n | 0.0095743 |
| t = (x̄−μ₀)/SE | 2.6112 |
Notation
Annotate
From Step 3 — standard error, then the t-statistic — read this one piece at a time. What is each part doing?
On: \( t = \frac{\bar x - \mu_0}{s/\sqrt{n}} = \frac{0.825 - 0.80}{0.0095743} = \frac{0.025}{0.0095743} \)
Fill the middle
Fill in the blanks
From Step 3 in code — the t-statistic — one line has had its right-hand side removed. Put it back.
import numpy as np
# 10 held-out accuracy scores of a model claimed to hit 0.80
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
xbar = data.mean() # sample mean
s = data.std(ddof=1) # SAMPLE sd (ddof=1)
se = s / np.sqrt(n) # standard error
t = (xbar - mu0) / se # t statistic
print(round(xbar,4), round(s,4), round(t,4))
Why: s is what everything below it consumes, so the wrong expression here fails later and somewhere else. Drop ddof=1 and std divides by n, shrinking s, inflating t and faking significance.
Worked example
Every hand number, reproduced in NumPy. This block is self-contained — it re-imports and re-defines the data:
import numpy as np
# 10 held-out accuracy scores of a model claimed to hit 0.80
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
xbar = data.mean() # sample mean
s = data.std(ddof=1) # SAMPLE sd (ddof=1)
se = s / np.sqrt(n) # standard error
t = (xbar - mu0) / se # t statistic
print(round(xbar,4), round(s,4), round(t,4))ddof=1 is the whole game on line 7
Why: Drop ddof=1 and std divides by n, shrinking s, inflating t and faking significance. The sample sd MUST use n−1.
| printed | value |
|---|---|
| xbar | 0.825 |
| s | 0.0303 |
| t | 2.6112 |
Pattern
Step through it
Step through Step 3 in code — the t-statistic one row at a time. What is driving the change, and what would the row after the last one be?
Concept
The one-sample t is valid when the observations are independent and the sample mean is approximately normal — true if the data itself is roughly normal, or (by the Central Limit Theorem) if n is large enough.
With n = 10 we lean on approximate normality of the folds. Heavy skew or outliers at small n break it — then reach for a nonparametric test (e.g. Wilcoxon) instead.
Independence is the assumption people violate most: reusing overlapping folds or correlated samples silently inflates significance.
Concept
We estimated the sd from the same 10 points, so t carries extra uncertainty. Under H₀ it follows a Student-t distribution with df = n−1 = 9 — like a normal but with heavier tails that shrink toward the normal as n grows.
\[ \text{under } H_0:\quad T = \frac{\bar X - \mu_0}{S/\sqrt{n}} \sim t_{\,df=n-1} \]
To get the p-value we measure how much of that t₉ distribution lies beyond ±|t_obs| — the two tails, because H₁ is ≠ (either direction counts).
Intuition
The t₉ curve looks like a bell but with fatter tails than the normal — it admits that with only 10 points our sd estimate could be off, so extreme values are a bit more likely.
That is why the two-sided 5% cutoff is ±2.262 for t₉, wider than the normal's ±1.96. Fewer data → fatter tails → a higher bar to clear before you reject.
As n → ∞ the sd estimate sharpens and t collapses onto the standard normal. Small samples are exactly where using t instead of z matters.
Faded example
Fill in the blanks
Step 4 — tail area to p-value, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
df = n - 1
one_tail = stats.t.sf(abs(t), df=df) # upper-tail area
p = **2 * one_tail** # two-sided
print(round(t,4), df, round(one_tail,4), round(p,4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. By symmetry of t₉, the lower tail beyond −|t| is also 0.0141.
Worked example
stats.t.sf(|t|, df) is the upper-tail area beyond |t|. Double it for the two-sided p (both tails are equal by symmetry):
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
df = n - 1
one_tail = stats.t.sf(abs(t), df=df) # upper-tail area
p = 2 * one_tail # two-sided
print(round(t,4), df, round(one_tail,4), round(p,4))one tail = 0.0141, so p = 2 × 0.0141 = 0.0282
Why: By symmetry of t₉, the lower tail beyond −|t| is also 0.0141. Two-sided p adds both: 0.0282.
| quantity | value (verified) |
|---|---|
| t | 2.6112 |
| df | 9 |
| upper tail P(T ≥ |t|) | 0.0141 |
| p = 2 × upper tail | 0.0282 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
one tail = 0.0141, so p = 2 × 0.0141 = 0.0282
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
stats.t.sf(|t|, df) is the upper-tail area beyond |t|. Double it for the two-sided p (both tails are equal by symmetry):
Concept
Our H₁ is μ ≠ 0.80 — an effect in either direction — so we count both tails: p = 2 × 0.0141 = 0.0282. This is the honest default when you didn't pre-commit to a direction.
A one-sided H₁: μ > 0.80 counts only the upper tail: p = 0.0141. It's more powerful, but only legitimate if you fixed the direction before seeing the data — otherwise it's a p-halving trick.
Rule of thumb for the exam: default to two-sided unless the problem explicitly justifies a direction.
Concept
We fixed α = 0.05 in advance. The data gave p = 0.0282.
\[ p = 0.0282 \;<\; 0.05 = \alpha \;\Longrightarrow\; \boxed{\text{reject } H_0} \]
So the evidence is inconsistent with μ = 0.80: the model's true accuracy is very likely above the claimed value. (Reject controls the false-positive rate at 5% — it doesn't make us 97% certain.)
Worked example
Equivalent to comparing p to α: compare |t| to the critical value — the cutoff that leaves α/2 = 2.5% in each tail of t₉:
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
t = (data.mean()-0.80)/(data.std(ddof=1)/np.sqrt(len(data)))
tcrit = stats.t.ppf(0.975, df=9) # two-sided 5% cutoff
print(round(t,4), round(tcrit,4))
print('reject H0:', abs(t) > tcrit)t_crit = t.ppf(0.975, df=9) = 2.2622
Why: ppf is the inverse CDF: the value with 97.5% of t₉ below it, so 2.5% above. Two tails → 5% total = α.
|t| = 2.6112 > 2.2622 = t_crit ⇒ reject
Why: |t| lands in the rejection region beyond the cutoff, so p < α automatically. The p-value route and the critical-value route ALWAYS agree.
| quantity | value (verified) |
|---|---|
| |t| observed | 2.6112 |
| t_crit (0.975, df=9) | 2.2622 |
| |t| > t_crit ? | True → reject H₀ |
Error analysis
Annotate
Walk the callouts on The same decision by critical value. Each one is a place this is easy to get subtly wrong.
Concept
A 95% confidence interval for the true mean is x̄ ± t_crit · SE — the set of μ₀ values this data would not reject at α = 0.05. It carries the same information as the p-value, plus a range.
\[ \bar x \pm t_{0.975,\,9}\cdot \frac{s}{\sqrt n} = 0.825 \pm 2.2622\cdot 0.0095743 \]
The margin is 2.2622 × 0.0095743 = 0.02166. If μ₀ = 0.80 falls outside the interval, that's exactly equivalent to rejecting H₀.
Fill the middle
Fill in the blanks
From Compute the 95% CI — one line has had its right-hand side removed. Put it back.
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
xbar = data.mean()
se = data.std(ddof=1) / np.sqrt(n)
tc = stats.t.ppf(0.975, df=n-1) # 2.5% in each tail
lo, hi = xbar - tcse, xbar + tcse
print(round(lo,4), round(hi,4))
print('0.80 inside?', lo <= 0.80 <= hi)
Why: se is what everything below it consumes, so the wrong expression here fails later and somewhere else. The plausible range for the true accuracy.
Worked example
Same ingredients as the test — x̄, SE, and the t₉ critical value — assembled into an interval. Self-contained:
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
xbar = data.mean()
se = data.std(ddof=1) / np.sqrt(n)
tc = stats.t.ppf(0.975, df=n-1) # 2.5% in each tail
lo, hi = xbar - tc*se, xbar + tc*se
print(round(lo,4), round(hi,4))
print('0.80 inside?', lo <= 0.80 <= hi)CI = 0.825 ± 0.02166 = [0.8033, 0.8467]
Why: The plausible range for the true accuracy. Its lower end 0.8033 sits ABOVE 0.80.
0.80 is NOT in [0.8033, 0.8467] ⇒ reject H₀
Why: The interval excludes the null value, the mirror image of p = 0.0282 < 0.05. CI and test always agree — the CI just also tells you the effect is roughly 0.003–0.047 above 0.80.
| quantity | value (verified) |
|---|---|
| x̄ | 0.825 |
| t_crit · SE (margin) | 0.02166 |
| 95% CI | [0.8033, 0.8467] |
| contains μ₀ = 0.80 ? | No → reject H₀ |
Blank canvas
Draw it
Draw what Compute the 95% CI just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Estimation
Predict first
Our from-scratch t and p must match the library's ttest_1samp exactly. Self-contained:
Commit before you compute: what does Verify the whole test against scipy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: ttest_1samp returns (statistic, pvalue) — both match
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. np.isclose confirms agreement to floating-point tolerance.
Worked example
Our from-scratch t and p must match the library's ttest_1samp exactly. Self-contained:
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
res = stats.ttest_1samp(data, mu0) # library reference
print(round(t,4), round(p,4))
print(round(res.statistic,4), round(res.pvalue,4))
print(np.isclose(t,res.statistic) and np.isclose(p,res.pvalue))ttest_1samp returns (statistic, pvalue) — both match
Why: np.isclose confirms agreement to floating-point tolerance. If your hand pipeline matches scipy, you understand every gear inside the black box.
| source | t | p |
|---|---|---|
| from scratch | 2.6112 | 0.0282 |
| scipy.ttest_1samp | 2.6112 | 0.0282 |
| np.isclose match | True | True |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use data.std() — NumPy's default — for the standard deviation in t.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: NumPy's default divides by n, giving the POPULATION sd.
Use data.std(ddof=1) — the sample sd — because we estimated the mean from the same data.
Why: NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.
Trap
Use data.std() — NumPy's default — for the standard deviation in t.
s = data.std() (ddof=0, divides by n)
Why: NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.
t = 0.025 / (0.02872/√10) = 2.752 → overstated
Why: A smaller denominator inflates t and shrinks p — you'd report a stronger effect than the data supports. Wrong distribution, wrong number.
Use data.std(ddof=1) — the sample sd — because we estimated the mean from the same data.
s = data.std(ddof=1) (divides by n−1)
Why: Dividing by n−1 gives an unbiased variance estimate: s = 0.030277, matching the t₉ theory.
t = 0.025 / (0.030277/√10) = 2.6112 ✓
Why: This is the value that agrees with scipy and with the t distribution. In a t-test, ALWAYS ddof=1.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Dividing by n−1 gives an unbiased variance estimate: s = 0.030277, matching the t₉ theory.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
NumPy's default divides by n, giving the POPULATION sd. On this data that's 0.02872, smaller than the sample sd 0.030277.
Section
Part 3 of 6 — t, χ², F
Intuition
Look back at t = (x̄ − μ₀)/(s/√n): the numerator is the signal (how far you are from the null) and the denominator is the noise (how much that estimate wobbles). Every test statistic has this shape.
That's why the same effect gets more 'significant' as n grows: the noise s/√n shrinks, the ratio blows up. And why a noisy metric needs a bigger effect to clear the bar.
So building any test = decide the signal, decide how to measure its noise, form the ratio, look up its tail. χ² and F just measure signal and noise differently.
Concept
Every test is the same loop: form a statistic, then read its tail under H₀'s known distribution. What changes is what you compare.
| test | compares | H₀ distribution | use when |
|---|---|---|---|
| t-test | means | Student-t | continuous data, small n |
| χ² test | categorical counts | chi-squared | frequencies / contingency tables |
| F-test | variances | F | ANOVA, comparing spread |
We built the t. Next a concrete χ² on categorical counts, so the family isn't just a table.
Faded example
Fill in the blanks
A chi-squared goodness-of-fit test, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from scipy import stats
observed = np.array([22,17,20,18,12,11]) # 100 die rolls, faces 1..6
expected = np.array([100/6]*6) # fair-die expectation
chi2 = np.sum((observed-expected)2/expected)
p = stats.chi2.sf(chi2, df=5)** # df = k-1 = 5
print(round(chi2,4), round(p,4))
print(stats.chisquare(observed).pvalue.round(4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Six faces → 5 degrees of freedom (the counts must sum to 100, costing one).
Worked example
Roll a die 100 times; is it fair? H₀: each face has expected count 100/6 ≈ 16.67. The statistic sums (observed−expected)²/expected over faces:
import numpy as np
from scipy import stats
observed = np.array([22,17,20,18,12,11]) # 100 die rolls, faces 1..6
expected = np.array([100/6]*6) # fair-die expectation
chi2 = np.sum((observed-expected)**2/expected)
p = stats.chi2.sf(chi2, df=5) # df = k-1 = 5
print(round(chi2,4), round(p,4))
print(stats.chisquare(observed).pvalue.round(4))χ² = Σ (O−E)²/E = 5.72, df = k−1 = 5
Why: Six faces → 5 degrees of freedom (the counts must sum to 100, costing one). Each term measures how far one face's count strayed from 16.67.
p = chi2.sf(5.72, df=5) = 0.3344 → fail to reject
Why: p ≫ 0.05: this much deviation is ordinary for a fair die over 100 rolls. No evidence of bias. scipy.chisquare agrees.
| face | observed | expected | (O−E)²/E |
|---|---|---|---|
| 1 | 22 | 16.67 | 1.7067 |
| 2 | 17 | 16.67 | 0.0067 |
| 3–6 | 20,18,12,11 | 16.67 each | 0.6667, 0.1067, 1.3067, 1.9267 |
| Σ | 100 | 100 | χ² = 5.72 |
Blank canvas
Draw it
Draw what A chi-squared goodness-of-fit test just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Worked example
Compare two models on 5 folds each. H₀: equal means. Welch's t divides the mean gap by the combined standard error √(s_A²/n_A + s_B²/n_B) — no equal-variance assumption:
import numpy as np
from scipy import stats
A = np.array([0.82,0.79,0.85,0.83,0.78]) # model A folds
B = np.array([0.80,0.75,0.79,0.77,0.76]) # model B folds
sew = np.sqrt(A.var(ddof=1)/len(A) + B.var(ddof=1)/len(B))
t = (A.mean() - B.mean()) / sew # Welch statistic
ref = stats.ttest_ind(A, B, equal_var=False)
print(round(t,4), round(ref.statistic,4), round(ref.pvalue,4))means 0.814 vs 0.774, gap 0.040; combined SE = √(0.00083/5 + 0.00043/5) = 0.01587
Why: Each group contributes its own variance over its own n. The gap spans 0.040/0.01587 = 2.52 standard errors.
t = 2.5198, p = 0.0386 (df ≈ 7.27) ⇒ reject H₀
Why: p < 0.05: the two models differ significantly on this data. Matches scipy.ttest_ind(equal_var=False) exactly.
| quantity | value (verified) |
|---|---|
| mean A / mean B | 0.814 / 0.774 |
| combined SE | 0.01587 |
| Welch t | 2.5198 |
| p-value | 0.0386 |
Comparison
Comparison matrix
From A two-sample t-test (Welch): refill the value (verified) column from what you know. The rest of the table is as it appeared.
| quantity | value (verified) |
|---|---|
| mean A / mean B | 0.814 / 0.774 |
| combined SE | 0.01587 |
| Welch t | 2.5198 |
| p-value | 0.0386 |
Fill the middle
Fill in the blanks
From An F-test: three models at once (ANOVA) — one line has had its right-hand side removed. Put it back.
import numpy as np
from scipy import stats
g1 = np.array([0.82,0.79,0.85]) # three models,
g2 = np.array([0.80,0.75,0.79]) # three folds each
g3 = np.array([0.88,0.90,0.86])
F, p = stats.f_oneway(g1, g2, g3) # one-way ANOVA
print(round(F,4), round(p,4))
Why: g2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Between-group spread is 11.4× the within-group noise — far more than chance would produce.
Worked example
With three models you can't just run pairwise t-tests (that's multiple testing!). One-way ANOVA uses an F statistic — the ratio of between-group spread to within-group spread — to test H₀: all means equal.
import numpy as np
from scipy import stats
g1 = np.array([0.82,0.79,0.85]) # three models,
g2 = np.array([0.80,0.75,0.79]) # three folds each
g3 = np.array([0.88,0.90,0.86])
F, p = stats.f_oneway(g1, g2, g3) # one-way ANOVA
print(round(F,4), round(p,4))F = 11.4, p = 0.009 ⇒ reject H₀
Why: Between-group spread is 11.4× the within-group noise — far more than chance would produce. At least one model's mean differs. A large F lives in the upper tail of the F distribution, giving a small p.
| group | folds | mean |
|---|---|---|
| g1 | 0.82, 0.79, 0.85 | 0.820 |
| g2 | 0.80, 0.75, 0.79 | 0.780 |
| g3 | 0.88, 0.90, 0.86 | 0.880 |
| F / p | — | 11.4 / 0.009 |
Section
Part 4 of 6 — what you can resolve
Intuition
Model A scores 85%, model B 83%. A is higher — but is the 2-point gap real, or the coin-flips of which test examples each happened to get right?
The answer depends entirely on how many examples you tested on. The exact same 2-point gap can be pure noise at n=100 and rock-solid at n=10000. The gap alone is not a test.
Concept
Power is the probability of correctly rejecting H₀ when H₁ is actually true — the chance of catching a real effect.
\[ \text{power} = 1 - \beta = P(\text{reject } H_0 \mid H_1 \text{ true}) \]
Power rises with the effect size and with the sample size (via the √n in the SE). Too-small n → low power → you miss real effects; huge n → even trivial effects become 'significant'.
Concept
A p-value tells you an effect is detectable; effect size tells you if it's big enough to care about. For a one-sample mean, Cohen's d reports the gap in sd units:
\[ d = \frac{\bar x - \mu_0}{s} = \frac{0.825 - 0.80}{0.030277} = 0.826 \]
d ≈ 0.83 is a large effect (rough guide: 0.2 small, 0.5 medium, 0.8 large). Always report the effect size next to p — with huge n, a trivial d can still be 'significant'.
Concept
Suppose the true accuracy really is 0.825 (a 0.025 effect) with sd 0.030. Even so, with only n = 10 we miss it — fail to reject — some fraction of the time. That miss rate is β.
\[ \text{power} = 1 - \beta \approx 0.74 \quad\Longrightarrow\quad \beta \approx 0.26 \]
So even a real, large effect goes undetected about 26% of the time at n = 10. Underpowered studies quietly produce false negatives — the fix is more data, which shrinks β.
Pattern
Predict first
The table runs: 10 | 2.108 | 0.5589 · 30 | 3.651 | 0.9546
In Power climbs with n, given the rows so far: what is the next one — the row where n is 100?
Correct: 100 | 6.667 | 1.0000
| n | effect (z-units) | power = 1−β |
|---|---|---|
| 10 | 2.108 | 0.5589 |
| 30 | 3.651 | 0.9546 |
| 100 | 6.667 | 1.0000 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. With only 10 folds we'd miss this real effect 44% of the time.
Worked example
Suppose the true accuracy really is 0.02 above the null, with sd 0.03. How likely are we to detect it at α = 0.05 for different n?
import numpy as np
from scipy import stats
delta, sigma, alpha = 0.02, 0.03, 0.05 # true shift, sd, level
zc = stats.norm.ppf(1 - alpha/2) # 1.96
for n in [10, 30, 100]:
se = sigma / np.sqrt(n)
ncp = delta / se # effect in z-units
power = stats.norm.sf(zc-ncp) + stats.norm.cdf(-zc-ncp)
print(n, round(ncp,3), round(power,4))n=10 → power 0.56; n=30 → 0.95; n=100 → 1.00
Why: With only 10 folds we'd miss this real effect 44% of the time. At n=30 we catch it 95% of the time. The effect never changed — our resolution did.
| n | effect (z-units) | power = 1−β |
|---|---|---|
| 10 | 2.108 | 0.5589 |
| 30 | 3.651 | 0.9546 |
| 100 | 6.667 | 1.0000 |
Missing information
Discussion prompt
A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Hand-check for n=100: 0.84·0.16 = 0.1344, times 2/100 = 0.002688, root = 0.051846.
Worked example
A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:
import numpy as np
from scipy import stats
p1, p2 = 0.85, 0.83 # two model accuracies
for n in [100, 1000, 10000]: # n per group
pp = (p1 + p2) / 2 # pooled proportion
se = np.sqrt(pp*(1-pp)*(2/n))
z = (p1 - p2) / se
print(n, round(z,4), round(2*stats.norm.sf(abs(z)),6))SE(100) = √(0.84·0.16·2/100) = √0.002688 = 0.051846
Why: Hand-check for n=100: 0.84·0.16 = 0.1344, times 2/100 = 0.002688, root = 0.051846.
n=100: z=0.386, p=0.70 (noise). n=10000: z=3.858, p=0.0001 (decisive)
Why: The identical 2-point gap is indistinguishable from chance at n=100 and overwhelming at n=10000. 'A > B' is a ranking, not a test.
| n per group | z | p-value | verdict |
|---|---|---|---|
| 100 | 0.3858 | 0.6997 | not significant |
| 1000 | 1.2199 | 0.2225 | not significant |
| 10000 | 3.8576 | 0.0001 | significant |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
n=100: z=0.386, p=0.70 (noise). n=10000: z=3.858, p=0.0001 (decisive)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
A two-proportion z-test on the 2-point gap, swept over sample size per group. Pooled proportion p̂ = (0.85+0.83)/2 = 0.84:
Anomaly
Predict first
A student writes this, and it looks reasonable:
Model A scored higher on the test set, so A is the better model — ship it.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.
Test the difference; report significance (or a confidence interval), not just the raw numbers.
Why: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.
Trap
Model A scored higher on the test set, so A is the better model — ship it.
Pick A from the point estimates alone
Why: Ignores sampling noise. At n=100 the 2-point gap has p=0.70 — squarely inside what chance produces.
Assume the ranking is stable
Why: On a fresh test set the ranking could easily flip. You reported a coin-flip as a finding.
Test the difference; report significance (or a confidence interval), not just the raw numbers.
Run a two-proportion test at your fixed α
Why: Turns 'A looks higher' into a decision that accounts for how many examples backed it.
Significant only with enough n; else say 'no detectable difference'
Why: The same gap is meaningless at n=100 and decisive at n=10000. Effect size AND n both matter.
Section
Part 5 of 6 — the p-hacking trap
Concept
Each test at α = 0.05 has a 5% false-positive rate. Run m independent tests and the chance of at least one false alarm is:
\[ P(\ge 1\text{ false positive}) = 1 - (1-\alpha)^m \]
| m tests | 1 − 0.95ᵐ | P(≥ 1 false positive) |
|---|---|---|
| 1 | 0.95 | 0.0500 |
| 5 | 0.95⁵ | 0.2262 |
| 10 | 0.95¹⁰ | 0.4013 |
| 20 | 0.95²⁰ | 0.6415 |
Pattern
Step through it
Step through False positives pile up one row at a time. What is driving the change, and what would the row after the last one be?
Concept
To keep the family-wise false-positive rate at α across m tests, compare each p-value to a stricter cutoff:
\[ \alpha_{\text{corrected}} = \frac{\alpha}{m} \qquad(\text{Bonferroni}) \]
Simple and conservative. For large m, FDR (Benjamini–Hochberg) controls the false-discovery rate with more power — but Bonferroni is the exam default.
Estimation
Predict first
Run 20 one-sample t-tests on completely random data — no real effect anywhere — and count the 'significant' ones. Seeded so you see these exact results:
Commit before you compute: what does p-hacking on pure noise come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 0 survive Bonferroni's 0.05/20 = 0.0025
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The smallest p (0.0111) exceeds 0.0025, so the corrected test correctly reports nothing significant — which is the truth.
Worked example
Run 20 one-sample t-tests on completely random data — no real effect anywhere — and count the 'significant' ones. Seeded so you see these exact results:
import numpy as np
from scipy import stats
rng = np.random.default_rng(7) # fixed seed -> reproducible
pvals = [stats.ttest_1samp(rng.normal(size=30), 0).pvalue
for _ in range(20)] # 20 tests on PURE noise
print(sum(p < 0.05 for p in pvals)) # naive count
print(sum(p < 0.05/20 for p in pvals)) # Bonferroni count4 of 20 'significant' at 0.05 — every one a false positive
Why: Seed 7 produced 4 p-values below 0.05 (min p = 0.0111) on noise. You EXPECT ~1 by chance (20×0.05); this draw gave 4.
0 survive Bonferroni's 0.05/20 = 0.0025
Why: The smallest p (0.0111) exceeds 0.0025, so the corrected test correctly reports nothing significant — which is the truth.
| threshold | 'significant' tests |
|---|---|
| α = 0.05 (naive) | 4 of 20 (all false) |
| α = 0.05/20 (Bonferroni) | 0 of 20 (correct) |
| min p-value observed | 0.0111 |
Error analysis
Annotate
Walk the callouts on p-hacking on pure noise. Each one is a place this is easy to get subtly wrong.
Concept
Bonferroni controls the chance of any false positive, so with thousands of tests (genomics, feature screens) it becomes so strict it misses real effects — low power.
Benjamini–Hochberg (FDR) instead controls the expected fraction of discoveries that are false. Sort the m p-values ascending and keep test k while p₍ₖ₎ ≤ (k/m)·α.
On our 20 noise tests, BH's rank-1 threshold is (1/20)·0.05 = 0.0025; the smallest p 0.0111 exceeds it, so BH also reports nothing — correct. FDR simply gains power when there are many true effects mixed in.
Explain it
Discussion prompt
Explain When Bonferroni is too strict: FDR to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Bonferroni controls the chance of any false positive, so with thousands of tests (genomics, feature screens) it becomes so strict it misses real effects — low power.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Tried 20 features; feature 7 came out significant (p = 0.011), so report that discovery.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64.
Correct for the number of tests before claiming significance.
Why: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64. The 'winner' is very likely noise.
Trap
Tried 20 features; feature 7 came out significant (p = 0.011), so report that discovery.
Cherry-pick the significant test, drop the other 19
Why: With 20 tests at α=0.05 you EXPECT a false positive — P(≥1) = 0.64. The 'winner' is very likely noise.
Publish without mentioning the 19
Why: Hiding the tests you ran is p-hacking. The result won't replicate on fresh data.
Correct for the number of tests before claiming significance.
Compare each p to α/m (Bonferroni), or control FDR
Why: p = 0.011 vs 0.05/20 = 0.0025 → not significant. Honest correction dissolves the fake finding.
Report all tests run, corrected
Why: Disclosing m and applying the correction is what separates a real discovery from noise mining.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
H₀' means 'not enough evidence to convict' — not 'proven innocent'. Absence of evidence isn't evidence of absence.; Crucially this is a statement about the data, computed under the assumption H₀ holds. It is not the probability that H₀ is true.; The p-value is the probability, assuming H₀, of a test statistic at least as extreme as the one observed:p = 0.03, so there is a 3% chance the null hypothesis is true.; Use data.std() — NumPy's default — for the standard deviation in t.Section
Part 6 of 6
Ranking
Put in order
These are the steps of The hypothesis-testing recipe, scrambled. Put them back in order before the next slide shows you.
H₀ and H₁; fix α before looking at the datat=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀p < α (or |t| > t_crit); report the effect size, not just 'significant'α/m (Bonferroni) or FDR if you ran manyWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
H₀ and H₁; fix α before looking at the datat=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀p < α (or |t| > t_crit); report the effect size, not just 'significant'α/m (Bonferroni) or FDR if you ran manyEdge cases
Discussion prompt
The hypothesis-testing recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
H₀ and H₁; fix α before looking at the datat=(x̄−μ₀)/(s/√n) with ddof=1 — and its p-value under H₀p < α (or |t| > t_crit); report the effect size, not just 'significant'α/m (Bonferroni) or FDR if you ran manyElimination
Eliminate the wrong options
Our t-test gave p = 0.0282. Which statement is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A p-value is P(data this extreme | H₀). Assuming μ = 0.80, a sample mean at least 2.61 SE away arises ~2.8% of the time — rare, so we reject H₀. It never states P(H₀).
Check
Read the conditioning precisely.
Check your understanding
Our t-test gave p = 0.0282. Which statement is correct?
Answer: A
Why: A p-value is P(data this extreme | H₀). Assuming μ = 0.80, a sample mean at least 2.61 SE away arises ~2.8% of the time — rare, so we reject H₀. It never states P(H₀).
Prediction
Predict first
Computing t = (x̄ − μ₀)/(s/√n) in NumPy, what should s be?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: data.std(ddof=1) — divide by n−1 (sample sd)
Why: The mean was estimated from the same data, so the unbiased variance divides by n−1: ddof=1. That gives s = 0.030277, t = 2.6112, matching scipy and the t₉ theory.
Check
Which standard deviation belongs in a one-sample t?
Check your understanding
Computing t = (x̄ − μ₀)/(s/√n) in NumPy, what should s be?
Answer: A
Why: The mean was estimated from the same data, so the unbiased variance divides by n−1: ddof=1. That gives s = 0.030277, t = 2.6112, matching scipy and the t₉ theory.
Commit first
Predict first
Model A 85% vs B 83% is statistically significant when…
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: the sample is large enough (e.g. ~10000/group), not at ~100/group
Why: The two-proportion test gives p = 0.70 at n=100 (noise) but p = 0.0001 at n=10000. Significance of a fixed gap depends on the sample size behind it.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
When is a fixed gap meaningful?
Check your understanding
Model A 85% vs B 83% is statistically significant when…
Answer: A
Why: The two-proportion test gives p = 0.70 at n=100 (noise) but p = 0.0001 at n=10000. Significance of a fixed gap depends on the sample size behind it.
Prediction
Predict first
You run 20 independent tests at α = 0.05 on data with no real effects. Roughly how many come out 'significant', and how do you control it?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: About 1 by chance; Bonferroni's α/20 = 0.0025 suppresses it
Why: Each test has a 5% false-positive rate, so the expected count is 20 × 0.05 = 1 (our seed happened to give 4). Bonferroni lowers the threshold to 0.05/20 = 0.0025 to hold the family-wise rate at 5%.
Check
Account for how many tests you ran.
Check your understanding
You run 20 independent tests at α = 0.05 on data with no real effects. Roughly how many come out 'significant', and how do you control it?
Answer: A
Why: Each test has a 5% false-positive rate, so the expected count is 20 × 0.05 = 1 (our seed happened to give 4). Bonferroni lowers the threshold to 0.05/20 = 0.0025 to hold the family-wise rate at 5%.
Elimination
Eliminate the wrong options
Our 95% CI for the true accuracy is [0.8033, 0.8467]. What does that tell you about testing H₀: μ = 0.80 at α = 0.05?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A 95% CI is exactly the set of μ₀ values a two-sided test at α=0.05 would NOT reject. Since 0.80 is outside [0.8033, 0.8467], the test rejects — matching p = 0.0282 < 0.05.
Check
Connect the interval to the test.
Check your understanding
Our 95% CI for the true accuracy is [0.8033, 0.8467]. What does that tell you about testing H₀: μ = 0.80 at α = 0.05?
Answer: A
Why: A 95% CI is exactly the set of μ₀ values a two-sided test at α=0.05 would NOT reject. Since 0.80 is outside [0.8033, 0.8467], the test rejects — matching p = 0.0282 < 0.05.
Section
Project
Concept
Rebuild the whole test on the 10-fold accuracy data using only NumPy for the math, then prove it matches scipy.stats.ttest_1samp. You derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | t = (x̄ − μ₀)/(s/√n) | np.mean, np.std(ddof=1) |
| 2 | two-sided p from the t distribution | scipy.stats.t.sf |
| 3 | verify vs scipy | stats.ttest_1samp |
Build rules: type every line yourself, use ddof=1 for the sample sd, and double the one-tailed area for a two-sided p. Predict each output before you run it.
Analogy
Discussion prompt
Explain Project: one-sample t-test from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Rebuild the whole test on the 10-fold accuracy data using only NumPy for the math, then prove it matches scipy.stats.ttest_1samp. You derived every piece — now assemble it yourself.
Worked example
Your turn: compute t for the 10-fold sample against μ₀ = 0.80. Predict its sign first (is x̄ above or below 0.80?).
Hint: t = (data.mean()-mu0)/(data.std(ddof=1)/np.sqrt(n)). Expect t > 0 because x̄ = 0.825 > 0.80.
import numpy as np
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
mu0 = 0.80
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
print(round(t, 4))| quantity | value |
|---|---|
| x̄ | 0.825 |
| s (ddof=1) | 0.0303 |
| t | 2.6112 |
Worked example
Your turn: turn t into a two-sided p-value with n−1 = 9 degrees of freedom. Predict: above or below 0.05?
Hint: p = 2*stats.t.sf(abs(t), df=n-1). The survival function sf is the upper-tail area; double it for two sides.
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
t = (data.mean() - 0.80) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
print(round(p, 4))| quantity | value |
|---|---|
| df | 9 |
| p-value | 0.0282 |
| reject H₀? | yes (p < 0.05) |
Pattern
Step through it
Step through Milestone 2 — the p-value one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: confirm your t and p match scipy.stats.ttest_1samp. Predict: exact match or just close?
Hint: stats.ttest_1samp(data, 0.80) returns (statistic, pvalue); compare with np.isclose.
import numpy as np
from scipy import stats
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
n = len(data)
t = (data.mean() - 0.80) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
res = stats.ttest_1samp(data, 0.80)
print(np.isclose(t,res.statistic) and np.isclose(p,res.pvalue))| source | t | p |
|---|---|---|
| your code | 2.6112 | 0.0282 |
| scipy | 2.6112 | 0.0282 |
| np.isclose match | True | True |
Trade off
Comparison matrix
From Milestone 3 — verify vs scipy: every row here is a choice with a cost. Fill the t column, then say which row you would actually pick and what you give up for it.
| source | t | p |
|---|---|---|
| your code | 2.6112 | 0.0282 |
| scipy | 2.6112 | 0.0282 |
| np.isclose match | True | True |
Concept
import numpy as np
from scipy import stats
def one_sample_t(data, mu0):
n = len(data)
t = (data.mean() - mu0) / (data.std(ddof=1)/np.sqrt(n))
p = 2 * stats.t.sf(abs(t), df=n-1)
return t, p
data = np.array([0.82,0.79,0.85,0.83,0.78,0.86,0.81,0.84,0.80,0.87])
t, p = one_sample_t(data, 0.80)
ref = stats.ttest_1samp(data, 0.80)
print('t =', round(t,4), ' p =', round(p,4))
print('matches scipy:',
np.isclose(t, ref.statistic) and np.isclose(p, ref.pvalue))| printed line | value |
|---|---|
| t = 2.6112 p = | 0.0282 |
| matches scipy: | True |
If your t reads 2.6112, p reads 0.0282, and scipy agrees — you built a hypothesis test from the definition up.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| t = 2.6112 p = | 0.0282 |
| matches scipy: | True |
Concept
Slides closed, out loud: explain (1) why a p-value is P(data | H₀) and not P(H₀), (2) why ddof=1 and not ddof=0, and (3) why the same 85-vs-83 gap flips significance with n.
Stretch: extend one_sample_t to a two-sample t-test (Welch's), then run 20 tests on noise and watch Bonferroni erase the false positives — the exact robustness gap the exam likes to probe.
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why a p-value is P(data | H₀) and not P(H₀), (2) why ddof=1 and not ddof=0, and (3) why the same 85-vs-83 gap flips significance with n.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Stretch: extend one_sample_t to a two-sample t-test (Welch's), then run 20 tests on noise and watch Bonferroni erase the false positives — the exact robustness gap the exam likes to probe.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The framework · Build the t-test · The test family · Power & sample size · Multiple testing · Recipe & checks. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
H₀/H₁ and name Type I (α) and Type II (β) errorsP(data this extreme | H₀), never P(H₀)t=(x̄−μ₀)/(s/√n) with ddof=1 — here t = 2.6112, p = 0.0282, reject H₀ — and verify vs scipyα/m (or FDR)| idea | the one thing to remember |
|---|---|
| p-value | P(data this extreme | H₀), not P(H₀) |
| t-statistic | (x̄−μ₀)/(s/√n) with ddof=1 |
| decision | reject if p < α ⇔ |t| > t_crit |
| significance | depends on sample size, not just the gap |
| power | 1 − β; grows with n and effect size |
| multiple tests | correct with α/m (Bonferroni) or FDR |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.