Chapter 9 tested a parameter against a fixed number; this chapter tests two populations against each other, and the six steps of a hypothesis test carry over unchanged. What changes is the parameter, which is now the difference of two population means, its point estimate, which is the difference of the two sample means, and the standard error, which combines the variability of both samples by adding their variances. Because the population standard deviations are unknown and estimated by the two sample standard deviations, the test statistic follows a Student t distribution, and its degrees of freedom come from the Aspin-Welch formula — a calculation the book calls complicated, which is rarely a whole number and is done by machine. The book is emphatic that the sample variances are not pooled, since pooling would assume the two populations share a standard deviation that this test deliberately does not assume.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 10 — Hypothesis Testing with Two Samples
Two Population Means with Unknown Standard Deviations
Objectives
Five outcomes, and the first is a change of parameter rather than of method.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-519 — the section these objectives are drawn from
Warm-up
Chapter 9 ran complete tests in six steps: hypotheses, tail, distribution, p-value, decision, conclusion.
Discussion prompt
A study compares the mean hours girls and boys spend on sport. Nothing is claimed about either group's mean on its own — only about whether they differ. What is the parameter being tested?
Hint: Chapter 9 required the null to name a single value. What single value does this claim name?
Answer:
Neither mean alone. The claim is about their DIFFERENCE, and the null names a single value for it: zero. Writing the null as mu-girls minus mu-boys equals 0 makes that explicit.
So the parameter has changed from a mean to a difference of means, and the point estimate changes to match — the difference of the two sample means, which for this study is 2 minus 3.2, or minus 1.2 hours.
Everything else carries over. The alternative still fixes the tail, a t distribution still supplies the p-value, and the decision still compares that p-value with alpha. What needs working out is the standard error of a difference, and how many degrees of freedom it has.
Concept
A difference between two samples depends on both the means and the standard deviations, and very different means can occur by chance if there is great variation among the individual samples. To account for that variation we take the difference of the sample means and divide by the standard error, which standardises the difference and gives a t-score test statistic.
the two-sample t test — Also called the Aspin-Welch t-test, after the developers of its degrees-of-freedom formula. It compares two independent population means when both population standard deviations are unknown.
\[ t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \]
The reason the standard error matters so much here is worth stating plainly, and the book states it: very different means can occur by chance if there is great variation among the individual samples. A gap of 1.2 hours means nothing until it is measured against how much the two sample means would vary by chance alone — which is exactly what dividing by the standard error does.
Figure (svg): Two columns contrasting a one-sample test with a two-sample test of means
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512
Section
Section 1
Concept
The hypotheses now concern two population means. The null states they are equal — equivalently that their difference is zero — and the alternative states they differ, or that one exceeds the other. The random variable is the difference of the sample means.
the two notations — H-nought written as mu-1 equals mu-2, or as mu-1 minus mu-2 equals zero. The second makes the parameter explicit and shows that the null still names a single value, as chapter 9 required.
\[ H_0: \mu_1 = \mu_2 \;\Longleftrightarrow\; H_0: \mu_1 - \mu_2 = 0 \]
The book's remark on wording is worth carrying: the words the same tell you the null has an equals, and since there are no other words to indicate the alternative, assume it says is different. So a comparison with no stated direction is two-tailed by default, and a direction has to be supplied by the problem before a one-tailed test is legitimate.
Figure (svg): A four-column table pairing three wordings with their hypotheses in both notations and the resulting tail
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513 — Example 10.1's hypotheses and the wording rule
Picture it
Six rows, of which three change.
Figure (svg): Two columns contrasting a one-sample test with a two-sample test of means
The unchanged rows are the ones worth noticing. The meaning of a p-value, the two error types, the decision rule and the conclusion wording all carry over exactly — so learning chapter 9 well makes this chapter mostly a matter of new formulas in a familiar frame.
Worked example
Girls and boys aged seven to eleven, hours of sport per day.
\[ \text{the average time is believed to be the SAME} \]
Read the claim
Why: The same.
Place it
Why: Equalities go in the null.
\[ H 0: \mu _{g} = \mu _{b} \]
Read the direction
Why: None given.
Write the alternative
Why: Two-sided.
\[ H a: \mu _{g} \ne \mu _{b} \]
Figure (svg): The solution to Worked example setting up Example 10.1 shown as a ladder of expressions, one row per legal move
\[ H_0: \mu_g = \mu_b, \qquad H_a: \mu_g \ne \mu_b \]
Verify: confirm the parameter named is a single number
Why: Written as mu-g minus mu-b equals zero, the null names one value — zero — exactly as chapter 9 required so that a distribution can be built around it. A null saying merely that two unknown means are equal without fixing their difference would supply no value to centre a sampling distribution on, and no p-value could be computed.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513
Sorting
Ask whether the two-sample setting changes it.
Sort into buckets
Sort each element.
Only the two computational rows change, which is why this chapter is shorter than chapter 9 despite covering four tests. Everything conceptual was settled there.
Worked example
Example 10.2, where the wording supplies a direction.
\[ \text{a graduate of college A has taken MORE maths classes, on average} \]
Read the claim
Why: A has taken more.
Place it
Why: Claims go in the alternative.
\[ H a: \mu _{A} > \mu _{B} \]
Write the null
Why: Its complement.
\[ H 0: \mu _{A} \le \mu _{B} \]
Read the tail
Why: Greater than.
Figure (svg): The solution to Worked example a directional claim shown as a ladder of expressions, one row per legal move
\[ H_0: \mu_A \le \mu_B, \qquad H_a: \mu_A > \mu_B \]
Verify: confirm which group is subtracted from which, and why it matters
Why: Writing the alternative as mu-A greater than mu-B means the difference mu-A minus mu-B is positive, so the right tail is the one wanted. Had the subscripts been reversed the same claim would need a LEFT-tailed test on mu-B minus mu-A. Fixing the order of subtraction before computing anything, and keeping it, is what prevents a sign error from inverting the conclusion.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 514-515
Trap
\[ H_a: \mu_A > \mu_B, \text{ then computing } \bar{x}_B - \bar{x}_A \]
Subtract in whichever order the data are listed
Why: Both differences describe the same comparison.
\[ t \text{ changes sign, so the right tail becomes the wrong one} \]
The p-value comes out as one minus the correct value, turning a rejection into a comfortable non-rejection.
\[ \text{fix group 1 and group 2 in the hypotheses, and keep that order} \]
Subtract in the same order throughout
Why: The hypotheses define which is group one.
For a two-tailed test the order does not matter, since both tails are counted — which is exactly what makes the habit easy to lose on two-tailed problems and costly on one-tailed ones. Writing group 1 and group 2 explicitly beside the hypotheses, before any arithmetic, is the cheapest guard.
Fill the middle
The equivalent way of writing the null.
Fill in the blanks
H_0: \mu_1 = \mu_2 \text0 H_0: \mu_1 - \mu_2 = ___
Why: Zero. Writing it this way shows that the null still names a single specific value for the parameter, which is what allows a sampling distribution to be centred on it.
Two truths and a lie
All three concern the setup.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Hypotheses always concern population parameters — mu-1 and mu-2 — while the sample means are the evidence. Section 9.1's rule is unchanged by there being two groups.
Prediction
Commit before reasoning.
Predict first
For a two-tailed test, does it matter which group is subtracted from which?
Correct: No for two-tailed, yes for one-tailed.
Why: Reversing the order flips the sign of the statistic. A two-tailed test counts both tails so the p-value is unchanged, but a one-tailed test would then be looking at the wrong tail and would report one minus the correct p-value. That asymmetry is why the habit of fixing the order is easy to lose and expensive to lose.
Section
Section 2
Concept
Because the population standard deviations are unknown, they are estimated by the two sample standard deviations. The estimated standard error of the difference in sample means is the square root of the first sample's variance over its size plus the second's over its size.
the standard error of a difference — The square root of s-one-squared over n-one plus s-two-squared over n-two. The variances are added because the two samples are independent, so their errors do not cancel.
\[ \text{SE} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}} \]
Adding rather than subtracting is the part that surprises people, and section 7.2's account of sums explains it. Two independent quantities each carry uncertainty, and combining them — whether by adding or by subtracting — combines those uncertainties rather than cancelling them. The difference of two means is therefore more variable than either mean alone, which is why comparing two groups needs more data than describing one.
Figure (svg): A card giving the standard error of a difference of means and the resulting t statistic
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512 — the standard error and the test statistic
Picture it
The standard error and the statistic built from it.
Figure (svg): A card giving the standard error of a difference of means and the resulting t statistic
Notice that each sample contributes its own variance divided by its own size, so a small noisy group can dominate the standard error even when the other group is large and tidy. That is why the degrees of freedom also depend on both, and why they are not simply the total sample size.
Worked example
Girls with n of 9 and boys with n of 16.
\[ s_g = 0.866, \; n_g = 9; \quad s_b = 1.00, \; n_b = 16 \]
Girls' contribution
Why: 0.866 squared over 9.
\[ 0.0833 \]
Boys' contribution
Why: One over 16.
\[ 0.0625 \]
Add them
Why: The variance of the difference.
\[ 0.1458 \]
Take the square root
Why: The standard error.
\[ 0.3819 \]
Figure (svg): The solution to Worked example the standard error for Example 10.1 shown as a ladder of expressions, one row per legal move
\[ \text{SE} = \sqrt{\frac{0.866^2}{9} + \frac{1^2}{16}} \approx 0.3819 \]
Verify: confirm the standard error exceeds each group's own
Why: The girls' mean has a standard error of 0.866 over 3, about 0.289, and the boys' is 1 over 4, or 0.25. The difference's standard error of 0.382 exceeds both, as it must — combining two uncertain quantities cannot produce something more certain than either. A standard error smaller than one of the two would signal that the variances had been subtracted.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513
Faded example
College A: n = 11, s = 1.5. College B: n = 9, s = 1.0.
Fill in the blanks
\text0.3157 = \sqrt0.5618___ + \frac______} = \sqrt___} \approx ___
Why: 0.2045 plus 0.1111 gives about 0.3157, whose square root is about 0.5618 maths classes — Example 10.2's standard error.
Worked example
Comparing the two contributions.
\[ 0.0833 \text{ against } 0.0625 \]
Girls' share
Why: 0.0833 of 0.1458.
\[ \text{about } 57 \% \]
Boys' share
Why: 0.0625 of 0.1458.
\[ \text{about } 43 \% \]
Note the sample sizes
Why: Nine against sixteen.
Draw the lesson
Why: The smaller sample dominates.
Figure (svg): The solution to Worked example which group dominates shown as a ladder of expressions, one row per legal move
\[ \frac{0.866^2}{9} > \frac{1^2}{16} \]
Verify: confirm what this implies for study design
Why: Adding observations to the smaller group reduces the standard error more than adding them to the larger one, so a fixed budget of extra subjects is best spent on whichever group is currently smaller or noisier. Balanced designs are common precisely because they are close to optimal when the two populations have similar spreads.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512
Error analysis
The correct value is about 0.3819.
Annotate
On: \( \begin{aligned} &(1)\; \sqrt{\tfrac{0.866^2}{9} - \tfrac{1^2}{16}} \approx 0.144 \\ &(2)\; \tfrac{0.866}{\sqrt{9}} + \tfrac{1}{\sqrt{16}} = 0.539 \\ &(3)\; \sqrt{\tfrac{0.866^2 + 1^2}{9 + 16}} \approx 0.264 \\ &(4)\; \sqrt{\tfrac{0.866^2}{9} + \tfrac{1^2}{16}} \approx 0.382 \end{aligned} \)
Error (1) is caught by the check that the difference's standard error must exceed each group's own. Errors (2) and (3) both produce plausible numbers, and only computing the formula as written catches them.
Two truths and a lie
All three concern the standard error.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is the section's most tempting error, since the statistic itself is a difference. But independence means the two errors do not cancel — a difference of two uncertain quantities is more uncertain than either, not less.
Prediction
Commit before reasoning.
Predict first
Two groups have n = 9 and n = 16 with similar spreads. Twenty more subjects are available. Where do they reduce the standard error most?
Correct: Mostly to the smaller group.
Why: The smaller group contributes the larger term to the variance, so reducing it has the bigger effect. Balancing the two sizes is close to optimal when the spreads are similar, which is why balanced designs are the default — though a genuinely noisier group deserves more even when it is already larger.
Estimation
Two sample means differ by 1.2 hours, with a standard error of 0.38.
Predict first
Roughly how many standard errors apart are they?
Correct: About 3.
Why: 1.2 divided by 0.38 is about 3.14, which is a substantial departure — three standard errors would be unusual under a null of no difference. Making this division before computing a p-value gives an immediate sense of whether the result will be significant.
Section
Section 3
Concept
The number of degrees of freedom requires a somewhat complicated calculation, which a computer or calculator does easily. They are not always a whole number. The book adds a direct instruction: the sample variances are not pooled — if the question comes up, do not pool the variances.
the Aspin-Welch degrees of freedom — A formula combining both samples' variances and sizes. It gives a non-integer value in general, and the Student t approximation is very good when both sample sizes are five or larger.
\[ \text{df} = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{(s_1^2/n_1)^2}{n_1-1} + \frac{(s_2^2/n_2)^2}{n_2-1}} \]
The rule against pooling is a statement about assumptions rather than about arithmetic. A pooled test assumes the two populations share one standard deviation and combines the samples to estimate it, which gives more degrees of freedom and a narrower test — but only if the assumption holds. The Aspin-Welch approach assumes nothing about the two spreads, which is why the book prefers it and why its degrees of freedom come out fractional.
Figure (svg): The conditions and the rule against pooling the variances
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513 — the degrees of freedom formula and the warning against pooling
Picture it
What the test assumes, and what it deliberately does not.
Figure (svg): The conditions and the rule against pooling the variances
The second and third lines are chapter 9's normality assumption in two-sample form, and they behave the same way: small samples need approximately normal populations, and large ones are rescued by the central limit theorem. The fifth line is the only genuinely new instruction.
Worked example
The book reports 18.8462.
\[ s_g = 0.866, \; n_g = 9; \quad s_b = 1.00, \; n_b = 16 \]
Form the two terms
Why: Each variance over its n.
\[ 0.0833\text{ and } 0.0625 \]
Square their sum
Why: 0.1458 squared.
\[ 0.02127 \]
Form the denominator
Why: Each term squared over n minus one.
\[ 0.001128 \]
Divide
Why: The degrees of freedom.
\[ 18.85 \]
Figure (svg): The solution to Worked example the degrees of freedom for Example 10.1 shown as a ladder of expressions, one row per legal move
\[ \text{df} \approx 18.85 \]
Verify: confirm the value sits between what the two samples alone would give
Why: A one-sample t on the girls would have 8 degrees of freedom and on the boys 15; the combined figure of 18.85 exceeds both but falls well short of the 23 that pooling would give. That bracketing is a useful check: an Aspin-Welch df should always be less than n-one plus n-two minus two and at least the smaller sample's own.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513
Two truths and a lie
All three concern the degrees of freedom.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false: that is the POOLED degrees of freedom, which the book forbids. The Aspin-Welch value is generally smaller and fractional, because it makes no assumption that the two populations share a standard deviation.
Worked example
Why the book forbids the shortcut.
\[ s_g = 0.866 \text{ against } s_b = 1.00 \]
State the pooling assumption
Why: One common sigma.
\[ \sigma _{1} = \sigma _{2} \]
Note what is observed
Why: 0.866 against 1.00.
Ask what pooling buys
Why: More degrees of freedom.
\[ 23\text{ rather than } 18.85 \]
Ask what it costs
Why: An untested assumption.
Figure (svg): The solution to Worked example what pooling would have assumed shown as a ladder of expressions, one row per legal move
\[ \text{pooled df} = n_1 + n_2 - 2 = 23 \]
Verify: confirm the practical size of the difference here
Why: With 23 degrees of freedom rather than 18.85 the two-tailed p-value for this statistic would be about 0.0046 rather than 0.0054 — a small change that would not alter the decision. The reason to follow the rule anyway is that the difference is not always small: when the sample sizes and spreads are both very unequal, pooling can be badly wrong, and the Aspin-Welch approach never is.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513
Trap
\[ \text{df} = n_1 + n_2 - 2 = 23, \text{ using a pooled } s \]
Combine the samples to estimate one common spread
Why: It gives a whole number and more degrees of freedom.
\[ \text{but the two populations may not share a spread} \]
The book's instruction is explicit: if the question comes up, do not pool the variances.
\[ \text{df} \approx 18.85, \text{ from the Aspin-Welch formula} \]
Use each sample's own variance and let the df come out fractional
Why: The test assumes nothing about the two spreads.
A non-integer degrees of freedom looks wrong to anyone used to n minus one, and that is often what prompts the pooling. It is worth accepting: the fractional value is the honest one, and every calculator and statistics package computes it without complaint.
Faded example
Two samples of 11 and 9, using the Aspin-Welch formula.
Fill in the blanks
\text18 11 + 9 - 2 = 17.4, \text___ ___, \text___
Why: Example 10.2's Aspin-Welch degrees of freedom are about 17.40, below the pooled 18 — as the formula's output always is when the two variances differ.
Prediction
Commit before reasoning.
Predict first
What does pooling the variances assume that the Aspin-Welch approach does not?
Correct: That the two populations share a standard deviation.
Why: Pooling combines both samples into one estimate of a single common spread, which only makes sense if such a spread exists. Normality is assumed by both approaches, equal sample sizes by neither, and equal means is the null being tested rather than an assumption.
Sorting
From the book's list.
Sort into buckets
Sort each statement.
Item (b) is the one to remember, since it is exactly what the Aspin-Welch approach avoids assuming — and the reason its degrees of freedom are more complicated than a pooled test's.
Section
Section 4
Concept
With the hypotheses written, the standard error computed and the degrees of freedom obtained, the remaining steps are chapter 9's unchanged: find the p-value on the t distribution, compare it with a preset alpha, decide, and write the conclusion in context.
the test statistic — The difference of the sample means, minus the difference the null claims, divided by the standard error. Under a null of no difference the middle term is zero, so the statistic is simply the observed difference over its standard error.
\[ t = \frac{\bar{x}_1 - \bar{x}_2}{\text{SE}}, \quad \text{df from Aspin-Welch} \]
The book draws the picture on the scale of the difference rather than of t, which is worth imitating once: it says half the p-value is below minus 1.2 and half is above 1.2, in hours of sport. Working on the original scale keeps the reader's attention on the quantity that matters, and the t scale is only a device for looking up an area.
Figure (svg): A t distribution with both tails beyond plus and minus three point one four shaded
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 513-516 — Examples 10.1 and 10.2, worked
Picture it
Example 10.1: do girls and boys differ?
Figure (svg): A t distribution with both tails beyond plus and minus three point one four shaded
The two shaded tails together give 0.0054, well below the 5 percent level, so the null is rejected — the means differ. On the original scale the boys average 1.2 hours more per day, and the test says a gap that size is very unlikely to have arisen by chance from equal populations.
Worked example
A two-tailed test at the 5 percent level.
\[ \bar{x}_g = 2, \; \bar{x}_b = 3.2; \; \alpha = 0.05 \]
Compute the difference
Why: Two minus 3.2.
\[ -1.2\text{ hours} \]
Divide by the standard error
Why: Minus 1.2 over 0.3819.
\[ t = -3.14 \]
Find the p-value
Why: Two tails on 18.85 df.
\[ 0.0054 \]
Decide
Why: Alpha exceeds it.
\[ \text{reject } H 0 \]
Figure (svg): The solution to Worked example Example 10.1 in full shown as a ladder of expressions, one row per legal move
\[ t \approx -3.14, \quad p = 0.0054 < 0.05 \]
Verify: confirm the direction of the difference is reported, not just its significance
Why: Rejecting says the means differ; it does not by itself say which is larger. The sample means make that clear — boys averaged 3.2 hours against girls' 2 — and a conclusion should state it, since a reader wants to know the direction as well as that a difference exists. The two-tailed test establishes the difference; the data establish its sign.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 513-514
Faded example
Sample means 2 and 3.2, standard error 0.3819.
Fill in the blanks
t = \frac-3.140.0054 \approx ___, \quad \text___ p = ___
Why: The difference is about 3.14 standard errors below zero, and both tails beyond that on 18.85 degrees of freedom give 0.0054.
Worked example
A right-tailed test at the 1 percent level.
\[ \bar{x}_A = 4, \; \bar{x}_B = 3.5; \; \alpha = 0.01 \]
Compute the difference
Why: Four minus 3.5.
\[ 0.5\text{ classes} \]
Divide by the standard error
Why: 0.5 over 0.5618.
\[ t = 0.89 \]
Find the p-value
Why: Right tail on 17.40 df.
\[ 0.1928 \]
Decide
Why: Alpha is far below it.
Figure (svg): The solution to Worked example Example 10.2 in full shown as a ladder of expressions, one row per legal move
\[ t \approx 0.89, \quad p = 0.1928 > 0.01 \]
Verify: confirm why half a class is not enough evidence
Why: The gap of 0.5 classes is less than one standard error of 0.5618, so a difference this size arises easily by chance from populations with equal means. Comparing the raw difference against the standard error before computing anything would have predicted a large p-value — a gap under one standard error is never significant at any conventional level.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 514-516
Trap
\[ \bar{x}_A - \bar{x}_B = 0.5 \text{ classes, so college A is ahead} \]
Read the observed gap as the answer
Why: Half a class is a real difference in the data.
\[ t = 0.89, \; p = 0.1928 \]
A gap smaller than one standard error is exactly what equal populations produce routinely.
\[ \text{compare the gap with the standard error: } \frac{0.5}{0.5618} = 0.89 \]
Standardise before judging
Why: The book: very different means can occur by chance if there is great variation.
This is the whole reason the test exists. Two samples always differ somewhat, so a difference on its own establishes nothing — the question is whether it is large relative to how much the difference would vary by chance, and only the standard error answers that.
Estimation
A difference of 0.5 with a standard error of 0.56.
Predict first
Without computing, will this be significant at any conventional level?
Correct: No: the gap is under one standard error.
Why: A statistic under 1 leaves a right-tail area above 0.15 on any t distribution, so no conventional alpha would reject. Dividing the gap by the standard error before computing gives an immediate read on the outcome, and a ratio under about 1.7 rarely reaches significance one-tailed.
Two truths and a lie
All three concern running the test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false as stated, because the p-value depends on the gap relative to the standard error. A gap of 5 with a standard error of 10 is far less significant than a gap of 1 with a standard error of 0.1 — which is exactly the point the book makes about variation among the samples.
Explain it
A classmate says college A's graduates plainly take more maths, since 4 exceeds 3.5.
Discussion prompt
In two sentences or fewer, explain what is missing.
Hint: Ask how much two sample means would differ by chance.
Answer:
Two samples from identical populations would still differ, and here the typical size of that chance difference is about 0.56 classes — larger than the 0.5 observed.
So the gap is smaller than the noise, which is why the p-value is 0.19 and the evidence is insufficient.
Section
Section 5
Concept
The two independent samples must be simple random samples from two distinct populations. If the sample sizes are small the distributions matter and should be normal; if they are large the distributions need not be. The book adds that when the sum of the sample sizes exceeds 30, a normal distribution can approximate the Student t — but to use the t whenever possible.
the normal approximation — Permitted when n-one plus n-two exceeds 30, since the t approaches the normal as degrees of freedom grow. The book's own advice is to use the t regardless, because a calculator makes it no harder.
\[ n_1 + n_2 > 30 \;\Rightarrow\; \text{normal is adequate, but } t \text{ is better} \]
The advice to use the t whenever possible is the same judgement section 8.2 reached about one-sample intervals: the correction costs nothing now that calculators compute any t, and skipping it always errs in the direction of overstating the evidence. The approximation note is best read as reassurance about older tables rather than as a recommendation.
Figure (svg): The conditions and the rule against pooling the variances
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-514 — the conditions and the note on the normal approximation
Picture it
What has to hold for the test to mean anything.
Figure (svg): The conditions and the rule against pooling the variances
The first condition is the one no sample size repairs, exactly as in chapters 8 and 9. Two independently drawn random samples is a statement about how the data were collected, and a study that compares self-selected groups cannot be rescued by the arithmetic that follows.
Worked example
Example 10.1's setting.
\[ n_g = 9, \; n_b = 16; \; \text{both populations normal} \]
Independence
Why: Two distinct groups.
Random sampling
Why: Stated as a study.
Sample sizes
Why: Nine and sixteen.
Normality
Why: Stated in the problem.
Figure (svg): The solution to Worked example checking the conditions shown as a ladder of expressions, one row per legal move
\[ \text{small } n \;\Rightarrow\; \text{normality matters, and is given} \]
Verify: confirm why the problem bothers to state normality
Why: With samples of 9 and 16 the central limit theorem cannot be relied on, so the populations' own shapes carry the assumption — which is why the book says each population has a normal distribution rather than leaving it implicit. For samples of a hundred each the sentence would be unnecessary, and its presence is a signal that the sample sizes are small enough to need it.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513
Discrimination
Check independence, randomness and the sample sizes.
Sort into buckets
Sort each study.
Worked example
The book's note about the sum of the sample sizes.
\[ n_1 + n_2 = 25 \text{ against } n_1 + n_2 = 60 \]
At a total of 25
Why: Below 30.
At a total of 60
Why: Above 30.
Recall the advice
Why: Use t whenever possible.
Say why
Why: It costs nothing.
Figure (svg): The solution to Worked example when a normal would serve shown as a ladder of expressions, one row per legal move
\[ n_1 + n_2 > 30 \;\Rightarrow\; \text{normal permitted} \]
Verify: confirm the direction of the error a normal would make
Why: The t has heavier tails, so a normal gives a smaller p-value on the same statistic — overstating the evidence, exactly as section 8.2 found. At 18.85 degrees of freedom the difference is small; at 5 it would not be. Since the error always runs the same way, the safe default is the t.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 514
Trap
\[ n_1 + n_2 = 60 > 30 \;\Rightarrow\; \text{use a normal} \]
Apply the approximation because it is permitted
Why: The book says it is adequate.
\[ \text{but the book also says to use the t whenever possible} \]
Permitted and preferable are different things, and the approximation always errs toward significance.
\[ \text{use the t, with Aspin-Welch degrees of freedom} \]
Take the approximation as reassurance rather than instruction
Why: It cost something when tables were the only tool; it costs nothing now.
The whole note exists because a t table needs a degrees-of-freedom row and an Aspin-Welch value is fractional — awkward on paper and trivial on a calculator. Reading a historical convenience as current advice is how outdated shortcuts survive.
Two truths and a lie
All three concern the conditions.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it is the same error chapters 8 and 9 warned about. A larger non-random sample estimates a biased quantity more precisely, which makes the problem harder to detect rather than easier.
Prediction
Commit before reasoning.
Predict first
Using a normal instead of the t on the same statistic gives what?
Correct: A smaller p-value.
Why: The t has heavier tails than the normal, so it leaves more area beyond any statistic — meaning the correct p-value is the larger one. The approximation therefore always flatters the result, which is why the book's advice to prefer the t is worth following even when the shortcut is permitted.
Faded example
Two studies, with samples of 9 and 16, and of 200 and 250.
Fill in the blanks
\textfirst central limit theorem \text___ ___ \text___
Why: With samples of 9 and 16 the populations' own shapes carry the assumption; with 200 and 250 the central limit theorem makes the sample means normal whatever the populations look like.
Comparison
Fill the blanks. Three rows change and the rest carry over.
Comparison matrix
| One sample (ch. 9) | Two samples (10.1) | |
|---|---|---|
| Parameter | mu | mu-1 minus mu-2 |
| Point estimate | x-bar | x-bar-1 minus x-bar-2 |
| Standard error | s over root n | root of s1-squared over n1 plus s2-squared over n2 |
| Degrees of freedom | n minus one | the Aspin-Welch formula, usually fractional |
Everything not in this table — the meaning of a p-value, the two error types, the decision rule, the conclusion wording — is unchanged from chapter 9, which is why this chapter can cover four tests in the space chapter 9 gave to one.
Pattern
Six steps, and only the fourth is new.
Check the standard error exceeds each group's own before proceeding, and divide the observed gap by it to predict roughly what the p-value will be.
OpenStax Introductory Business Statistics 2e, §10.1 Comparing Two Independent Population Means §10.1 Comparing Two Independent Population Means
Check
The standard error.
Check your understanding
Two samples have s = 0.866 with n = 9, and s = 1.00 with n = 16. What is the standard error of the difference?
Answer: A
Why: Add each variance over its own sample size — 0.0833 plus 0.0625 — and take the square root.
Check
Degrees of freedom.
Check your understanding
For samples of 9 and 16, what are the degrees of freedom?
Answer: A
Why: The Aspin-Welch formula combines both variances and sizes, and gives a value that is rarely a whole number.
Check
Pooling.
Check your understanding
Why does the book instruct that the variances not be pooled?
Answer: A
Why: The Aspin-Welch approach makes no assumption about the two spreads, which is why its degrees of freedom are more complicated.
Real world
A hospital compares recovery times under two surgical techniques. Technique A gives a mean of 5.1 days from 12 patients with a standard deviation of 1.9; technique B gives 6.3 days from 40 patients with a standard deviation of 3.8. A surgeon reports that A is clearly better, since its mean is over a day shorter and its variability is half as large.
Discussion prompt
Test the claim properly, and say what the comparison of variabilities does and does not establish.
Hint: The smaller, tidier sample is not necessarily the more informative one.
Answer:
The standard error is dominated by the larger, noisier group. Technique A contributes 1.9 squared over 12, about 0.301; technique B contributes 3.8 squared over 40, about 0.361. Their sum is 0.662, so the standard error is about 0.814 days.
\[ t = \frac{5.1 - 6.3}{0.814} \approx -1.47, \qquad \text{df} \approx 27.5, \qquad p_{\text{two}} \approx 0.15 \]
The observed gap of 1.2 days is under one and a half standard errors, so the two-tailed p-value is around 0.15 and the null is not rejected at any conventional level. The surgeon's clearly is not supported: a difference this size arises readily by chance from populations with equal means.
The comparison of variabilities establishes nothing about the means and may not even establish a difference in spread. With only 12 patients, technique A's standard deviation of 1.9 is itself very imprecisely estimated — a sample that small routinely produces a spread half or double the population's. Testing whether two variances differ is a separate procedure, and it is section 13.4's subject rather than this one's.
Two design points follow. The imbalance is costly: 12 against 40 puts most of the uncertainty in the smaller group despite its tidier data, and a dozen more patients on technique A would sharpen the comparison more than a dozen more on B. And a difference of 1.2 days may well be clinically important, so the honest report is not that the techniques are equivalent but that this study was too small to settle it — which is a call for a larger trial rather than a conclusion.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why are the two variances added rather than subtracted when forming the standard error of a difference?
Correct: Because the uncertainties combine rather than cancel.
\[ \text{SE} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}} > \max\left(\frac{s_1}{\sqrt{n_1}}, \frac{s_2}{\sqrt{n_2}}\right) \]
Why: Two independent quantities each carry error, and combining them by any arithmetic — adding or subtracting — combines those errors. The difference of two sample means is therefore more variable than either mean alone, which is why the standard error must exceed each group's own. That a negative would be impossible is a symptom of the error rather than its reason.
Explain it
They used degrees of freedom of n-one plus n-two minus two, saying the fractional Aspin-Welch value looked wrong.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask what their formula assumes about the two populations.
Answer:
Their formula is the POOLED one, and it only applies if the two populations share a standard deviation — an assumption this test deliberately avoids making.
A fractional degrees of freedom is the honest output of the Aspin-Welch formula, and the book says explicitly not to pool the variances.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, name group 1 and group 2 in writing and keep that order. For the second, each variance over its own n, added, then rooted — and check the result exceeds both. For the third, use Aspin-Welch and accept a fractional value. For the fourth, report which mean was larger as well as whether the difference was significant. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw two columns headed one sample and two samples, and fill in four rows: parameter, point estimate, standard error and degrees of freedom — circling the three that change. Beneath, write the standard error formula and beside it the two things that could go wrong with it: subtracting the variances, and pooling them. In the middle of the page, work Example 10.1 completely: write the hypotheses in both notations, compute each variance over its own sample size, add them, take the root to get 0.3819, compute the statistic of about minus 3.14, and note the degrees of freedom as 18.85 with a remark that it is not a whole number. Draw the t curve with both tails shaded and label the p-value 0.0054, then write the conclusion naming the level and the direction. Beside it, work Example 10.2 the same way and note that the gap of 0.5 is smaller than the standard error of 0.56 — then write one sentence on why that alone predicts a large p-value. At the bottom, list the five conditions from the book, marking which one no sample size repairs.
Check your standard error in each example by confirming it exceeds both individual standard errors — 0.289 and 0.25 for the first example, so 0.382 is right. Check the degrees of freedom by confirming each falls below the pooled n-one plus n-two minus two, which is what the Aspin-Welch formula always gives when the variances differ.
Recap
Five things, and the first three are the only genuinely new material.
| If you see | Then |
|---|---|
| Two independent groups compared | The parameter is a difference of means |
| No stated direction | A two-tailed test by default |
| Two sample standard deviations | Add each variance over its own n |
| A standard error below one group's own | The variances were subtracted |
| A whole-number df from two samples | The variances were probably pooled |
| A gap smaller than the standard error | It will not be significant |
| Paired or before-and-after data | Not this test: see section 10.4 |
Section 10.2 takes the same comparison with the population standard deviations known, which replaces the t with a normal and removes the degrees-of-freedom problem entirely. The book concedes at the outset that the situation is not likely, and the section is short accordingly.
OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-519 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.