10.1 Two Population Means with Unknown Standard Deviations

Chapter 9 tested a parameter against a fixed number; this chapter tests two populations against each other, and the six steps of a hypothesis test carry over unchanged. What changes is the parameter, which is now the difference of two population means, its point estimate, which is the difference of the two sample means, and the standard error, which combines the variability of both samples by adding their variances. Because the population standard deviations are unknown and estimated by the two sample standard deviations, the test statistic follows a Student t distribution, and its degrees of freedom come from the Aspin-Welch formula — a calculation the book calls complicated, which is rarely a whole number and is done by machine. The book is emphatic that the sample variances are not pooled, since pooling would assume the two populations share a standard deviation that this test deliberately does not assume.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 10.1 Two Population Means with Unknown Standard Deviations

Title

Statistics · Chapter 10 — Hypothesis Testing with Two Samples

Two Population Means with Unknown Standard Deviations

2. By the end of this lesson you can

Objectives

Five outcomes, and the first is a change of parameter rather than of method.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-519 — the section these objectives are drawn from

3. What you already have

Warm-up

Chapter 9 ran complete tests in six steps: hypotheses, tail, distribution, p-value, decision, conclusion.

Discussion prompt

A study compares the mean hours girls and boys spend on sport. Nothing is claimed about either group's mean on its own — only about whether they differ. What is the parameter being tested?

Hint: Chapter 9 required the null to name a single value. What single value does this claim name?

Answer:

Neither mean alone. The claim is about their DIFFERENCE, and the null names a single value for it: zero. Writing the null as mu-girls minus mu-boys equals 0 makes that explicit.

So the parameter has changed from a mean to a difference of means, and the point estimate changes to match — the difference of the two sample means, which for this study is 2 minus 3.2, or minus 1.2 hours.

Everything else carries over. The alternative still fixes the tail, a t distribution still supplies the p-value, and the decision still compares that p-value with alpha. What needs working out is the standard error of a difference, and how many degrees of freedom it has.

4. The parameter is a difference; the method is unchanged

Concept

A difference between two samples depends on both the means and the standard deviations, and very different means can occur by chance if there is great variation among the individual samples. To account for that variation we take the difference of the sample means and divide by the standard error, which standardises the difference and gives a t-score test statistic.

the two-sample t test — Also called the Aspin-Welch t-test, after the developers of its degrees-of-freedom formula. It compares two independent population means when both population standard deviations are unknown.

\[ t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \]

The reason the standard error matters so much here is worth stating plainly, and the book states it: very different means can occur by chance if there is great variation among the individual samples. A gap of 1.2 hours means nothing until it is measured against how much the two sample means would vary by chance alone — which is exactly what dividing by the standard error does.

Figure (svg): Two columns contrasting a one-sample test with a two-sample test of means

The six steps of chapter 9 carry over untouched. Only the three rows in the middle change, and they change in the same way for every test in this chapter.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512

5. What changes, and what does not

Section

Section 1

6. A new parameter in an old procedure

Concept

The hypotheses now concern two population means. The null states they are equal — equivalently that their difference is zero — and the alternative states they differ, or that one exceeds the other. The random variable is the difference of the sample means.

the two notations — H-nought written as mu-1 equals mu-2, or as mu-1 minus mu-2 equals zero. The second makes the parameter explicit and shows that the null still names a single value, as chapter 9 required.

\[ H_0: \mu_1 = \mu_2 \;\Longleftrightarrow\; H_0: \mu_1 - \mu_2 = 0 \]

The book's remark on wording is worth carrying: the words the same tell you the null has an equals, and since there are no other words to indicate the alternative, assume it says is different. So a comparison with no stated direction is two-tailed by default, and a direction has to be supplied by the problem before a one-tailed test is legitimate.

Figure (svg): A four-column table pairing three wordings with their hypotheses in both notations and the resulting tail

Writing the null as a difference equal to zero is worth doing once: it shows that the parameter really is a single number, as chapter 9 required.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513 — Example 10.1's hypotheses and the wording rule

7. One sample against two

Picture it

Six rows, of which three change.

Figure (svg): Two columns contrasting a one-sample test with a two-sample test of means

The six steps of chapter 9 carry over untouched. Only the three rows in the middle change, and they change in the same way for every test in this chapter.

The unchanged rows are the ones worth noticing. The meaning of a p-value, the two error types, the decision rule and the conclusion wording all carry over exactly — so learning chapter 9 well makes this chapter mostly a matter of new formulas in a familiar frame.

8. Worked example: setting up Example 10.1

Worked example

Girls and boys aged seven to eleven, hours of sport per day.

\[ \text{the average time is believed to be the SAME} \]

Read the claim

Why: The same.

Place it

Why: Equalities go in the null.

\[ H 0: \mu _{g} = \mu _{b} \]

Read the direction

Why: None given.

Write the alternative

Why: Two-sided.

\[ H a: \mu _{g} \ne \mu _{b} \]

Figure (svg): The solution to Worked example setting up Example 10.1 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \mu_g = \mu_b, \qquad H_a: \mu_g \ne \mu_b \]

Verify: confirm the parameter named is a single number

Why: Written as mu-g minus mu-b equals zero, the null names one value — zero — exactly as chapter 9 required so that a distribution can be built around it. A null saying merely that two unknown means are equal without fixing their difference would supply no value to centre a sampling distribution on, and no p-value could be computed.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513

9. What carries over from chapter 9?

Sorting

Ask whether the two-sample setting changes it.

Sort into buckets

Sort each element.

Unchanged
the meaning of a p-value; the decision rule against alpha; the two error types
Changes
the formula for the standard error; the degrees of freedom
same
It belongs to the logic of testing, which does not depend on how many samples there are.
new
It depends on the sampling distribution of the statistic, which now involves two samples.

Only the two computational rows change, which is why this chapter is shorter than chapter 9 despite covering four tests. Everything conceptual was settled there.

10. Worked example: a directional claim

Worked example

Example 10.2, where the wording supplies a direction.

\[ \text{a graduate of college A has taken MORE maths classes, on average} \]

Read the claim

Why: A has taken more.

Place it

Why: Claims go in the alternative.

\[ H a: \mu _{A} > \mu _{B} \]

Write the null

Why: Its complement.

\[ H 0: \mu _{A} \le \mu _{B} \]

Read the tail

Why: Greater than.

Figure (svg): The solution to Worked example a directional claim shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \mu_A \le \mu_B, \qquad H_a: \mu_A > \mu_B \]

Verify: confirm which group is subtracted from which, and why it matters

Why: Writing the alternative as mu-A greater than mu-B means the difference mu-A minus mu-B is positive, so the right tail is the one wanted. Had the subscripts been reversed the same claim would need a LEFT-tailed test on mu-B minus mu-A. Fixing the order of subtraction before computing anything, and keeping it, is what prevents a sign error from inverting the conclusion.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 514-515

11. Trap: letting the order of subtraction drift

Trap

The trap

\[ H_a: \mu_A > \mu_B, \text{ then computing } \bar{x}_B - \bar{x}_A \]

Subtract in whichever order the data are listed

Why: Both differences describe the same comparison.

\[ t \text{ changes sign, so the right tail becomes the wrong one} \]

The p-value comes out as one minus the correct value, turning a rejection into a comfortable non-rejection.

The fix

\[ \text{fix group 1 and group 2 in the hypotheses, and keep that order} \]

Subtract in the same order throughout

Why: The hypotheses define which is group one.

For a two-tailed test the order does not matter, since both tails are counted — which is exactly what makes the habit easy to lose on two-tailed problems and costly on one-tailed ones. Writing group 1 and group 2 explicitly beside the hypotheses, before any arithmetic, is the cheapest guard.

12. The null as a difference

Fill the middle

The equivalent way of writing the null.

Fill in the blanks

H_0: \mu_1 = \mu_2 \text0 H_0: \mu_1 - \mu_2 = ___

Why: Zero. Writing it this way shows that the null still names a single specific value for the parameter, which is what allows a sampling distribution to be centred on it.

13. One of these is false

Two truths and a lie

All three concern the setup.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A comparison with no stated direction is two-tailed
  • C. The point estimate is the difference of the sample means
  • B. The hypotheses name the two sample means

Survives elimination: B

Why: The survivor is false. Hypotheses always concern population parameters — mu-1 and mu-2 — while the sample means are the evidence. Section 9.1's rule is unchanged by there being two groups.

14. Does the order matter?

Prediction

Commit before reasoning.

Predict first

For a two-tailed test, does it matter which group is subtracted from which?

  • No for a two-tailed test, but yes for a one-tailed one
  • Yes, always
  • No, never
  • Only if the sample sizes differ

Correct: No for two-tailed, yes for one-tailed.

Why: Reversing the order flips the sign of the statistic. A two-tailed test counts both tails so the p-value is unchanged, but a one-tailed test would then be looking at the wrong tail and would report one minus the correct p-value. That asymmetry is why the habit of fixing the order is easy to lose and expensive to lose.

15. The standard error of a difference

Section

Section 2

16. Add the variances, never subtract them

Concept

Because the population standard deviations are unknown, they are estimated by the two sample standard deviations. The estimated standard error of the difference in sample means is the square root of the first sample's variance over its size plus the second's over its size.

the standard error of a difference — The square root of s-one-squared over n-one plus s-two-squared over n-two. The variances are added because the two samples are independent, so their errors do not cancel.

\[ \text{SE} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}} \]

Adding rather than subtracting is the part that surprises people, and section 7.2's account of sums explains it. Two independent quantities each carry uncertainty, and combining them — whether by adding or by subtracting — combines those uncertainties rather than cancelling them. The difference of two means is therefore more variable than either mean alone, which is why comparing two groups needs more data than describing one.

Figure (svg): A card giving the standard error of a difference of means and the resulting t statistic

The difference of two means is more variable than either mean on its own, which is why comparing groups needs more data than describing one.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512 — the standard error and the test statistic

17. Two variances, added

Picture it

The standard error and the statistic built from it.

Figure (svg): A card giving the standard error of a difference of means and the resulting t statistic

The difference of two means is more variable than either mean on its own, which is why comparing groups needs more data than describing one.

Notice that each sample contributes its own variance divided by its own size, so a small noisy group can dominate the standard error even when the other group is large and tidy. That is why the degrees of freedom also depend on both, and why they are not simply the total sample size.

18. Worked example: the standard error for Example 10.1

Worked example

Girls with n of 9 and boys with n of 16.

\[ s_g = 0.866, \; n_g = 9; \quad s_b = 1.00, \; n_b = 16 \]

Girls' contribution

Why: 0.866 squared over 9.

\[ 0.0833 \]

Boys' contribution

Why: One over 16.

\[ 0.0625 \]

Add them

Why: The variance of the difference.

\[ 0.1458 \]

Take the square root

Why: The standard error.

\[ 0.3819 \]

Figure (svg): The solution to Worked example the standard error for Example 10.1 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{SE} = \sqrt{\frac{0.866^2}{9} + \frac{1^2}{16}} \approx 0.3819 \]

Verify: confirm the standard error exceeds each group's own

Why: The girls' mean has a standard error of 0.866 over 3, about 0.289, and the boys' is 1 over 4, or 0.25. The difference's standard error of 0.382 exceeds both, as it must — combining two uncertain quantities cannot produce something more certain than either. A standard error smaller than one of the two would signal that the variances had been subtracted.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513

19. Compute a standard error

Faded example

College A: n = 11, s = 1.5. College B: n = 9, s = 1.0.

Fill in the blanks

\text0.3157 = \sqrt0.5618___ + \frac______} = \sqrt___} \approx ___

Why: 0.2045 plus 0.1111 gives about 0.3157, whose square root is about 0.5618 maths classes — Example 10.2's standard error.

20. Worked example: which group dominates

Worked example

Comparing the two contributions.

\[ 0.0833 \text{ against } 0.0625 \]

Girls' share

Why: 0.0833 of 0.1458.

\[ \text{about } 57 \% \]

Boys' share

Why: 0.0625 of 0.1458.

\[ \text{about } 43 \% \]

Note the sample sizes

Why: Nine against sixteen.

Draw the lesson

Why: The smaller sample dominates.

Figure (svg): The solution to Worked example which group dominates shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{0.866^2}{9} > \frac{1^2}{16} \]

Verify: confirm what this implies for study design

Why: Adding observations to the smaller group reduces the standard error more than adding them to the larger one, so a fixed budget of extra subjects is best spent on whichever group is currently smaller or noisier. Balanced designs are common precisely because they are close to optimal when the two populations have similar spreads.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 512

21. Error analysis: four standard errors for Example 10.1

Error analysis

The correct value is about 0.3819.

Annotate

On: \( \begin{aligned} &(1)\; \sqrt{\tfrac{0.866^2}{9} - \tfrac{1^2}{16}} \approx 0.144 \\ &(2)\; \tfrac{0.866}{\sqrt{9}} + \tfrac{1}{\sqrt{16}} = 0.539 \\ &(3)\; \sqrt{\tfrac{0.866^2 + 1^2}{9 + 16}} \approx 0.264 \\ &(4)\; \sqrt{\tfrac{0.866^2}{9} + \tfrac{1^2}{16}} \approx 0.382 \end{aligned} \)

  • (1) subtracts the variances, which would make a difference more certain than its parts — impossible for independent samples.
  • (2) adds the two standard errors rather than the variances, overstating the result.
  • (3) pools everything over the combined sample size, which is the pooling the book forbids.
  • (4) is correct: each variance over its own sample size, added, then rooted.

Error (1) is caught by the check that the difference's standard error must exceed each group's own. Errors (2) and (3) both produce plausible numbers, and only computing the formula as written catches them.

22. One of these is false

Two truths and a lie

All three concern the standard error.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The variances are added, not the standard deviations
  • C. The result exceeds either group's own standard error
  • B. The variances are subtracted for a difference of means

Survives elimination: B

Why: The survivor is false and is the section's most tempting error, since the statistic itself is a difference. But independence means the two errors do not cancel — a difference of two uncertain quantities is more uncertain than either, not less.

23. Where should extra data go?

Prediction

Commit before reasoning.

Predict first

Two groups have n = 9 and n = 16 with similar spreads. Twenty more subjects are available. Where do they reduce the standard error most?

  • Mostly to the smaller group
  • Mostly to the larger group
  • It makes no difference
  • Split exactly evenly

Correct: Mostly to the smaller group.

Why: The smaller group contributes the larger term to the variance, so reducing it has the bigger effect. Balancing the two sizes is close to optimal when the spreads are similar, which is why balanced designs are the default — though a genuinely noisier group deserves more even when it is already larger.

24. Is the gap meaningful?

Estimation

Two sample means differ by 1.2 hours, with a standard error of 0.38.

Predict first

Roughly how many standard errors apart are they?

  • About 3
  • About 1
  • About 0.3
  • About 12

Correct: About 3.

Why: 1.2 divided by 0.38 is about 3.14, which is a substantial departure — three standard errors would be unusual under a null of no difference. Making this division before computing a p-value gives an immediate sense of whether the result will be significant.

25. Degrees of freedom, and not pooling

Section

Section 3

26. Aspin-Welch, and a rule against a shortcut

Concept

The number of degrees of freedom requires a somewhat complicated calculation, which a computer or calculator does easily. They are not always a whole number. The book adds a direct instruction: the sample variances are not pooled — if the question comes up, do not pool the variances.

the Aspin-Welch degrees of freedom — A formula combining both samples' variances and sizes. It gives a non-integer value in general, and the Student t approximation is very good when both sample sizes are five or larger.

\[ \text{df} = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{(s_1^2/n_1)^2}{n_1-1} + \frac{(s_2^2/n_2)^2}{n_2-1}} \]

The rule against pooling is a statement about assumptions rather than about arithmetic. A pooled test assumes the two populations share one standard deviation and combines the samples to estimate it, which gives more degrees of freedom and a narrower test — but only if the assumption holds. The Aspin-Welch approach assumes nothing about the two spreads, which is why the book prefers it and why its degrees of freedom come out fractional.

Figure (svg): The conditions and the rule against pooling the variances

Pooling assumes the two populations share a standard deviation, which this test deliberately does not assume.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513 — the degrees of freedom formula and the warning against pooling

27. Five conditions and one warning

Picture it

What the test assumes, and what it deliberately does not.

Figure (svg): The conditions and the rule against pooling the variances

Pooling assumes the two populations share a standard deviation, which this test deliberately does not assume.

The second and third lines are chapter 9's normality assumption in two-sample form, and they behave the same way: small samples need approximately normal populations, and large ones are rescued by the central limit theorem. The fifth line is the only genuinely new instruction.

28. Worked example: the degrees of freedom for Example 10.1

Worked example

The book reports 18.8462.

\[ s_g = 0.866, \; n_g = 9; \quad s_b = 1.00, \; n_b = 16 \]

Form the two terms

Why: Each variance over its n.

\[ 0.0833\text{ and } 0.0625 \]

Square their sum

Why: 0.1458 squared.

\[ 0.02127 \]

Form the denominator

Why: Each term squared over n minus one.

\[ 0.001128 \]

Divide

Why: The degrees of freedom.

\[ 18.85 \]

Figure (svg): The solution to Worked example the degrees of freedom for Example 10.1 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{df} \approx 18.85 \]

Verify: confirm the value sits between what the two samples alone would give

Why: A one-sample t on the girls would have 8 degrees of freedom and on the boys 15; the combined figure of 18.85 exceeds both but falls well short of the 23 that pooling would give. That bracketing is a useful check: an Aspin-Welch df should always be less than n-one plus n-two minus two and at least the smaller sample's own.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513

29. One of these is false

Two truths and a lie

All three concern the degrees of freedom.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. They are rarely a whole number
  • C. The approximation is very good when both samples are at least five
  • B. They equal n-one plus n-two minus two

Survives elimination: B

Why: The survivor is false: that is the POOLED degrees of freedom, which the book forbids. The Aspin-Welch value is generally smaller and fractional, because it makes no assumption that the two populations share a standard deviation.

30. Worked example: what pooling would have assumed

Worked example

Why the book forbids the shortcut.

\[ s_g = 0.866 \text{ against } s_b = 1.00 \]

State the pooling assumption

Why: One common sigma.

\[ \sigma _{1} = \sigma _{2} \]

Note what is observed

Why: 0.866 against 1.00.

Ask what pooling buys

Why: More degrees of freedom.

\[ 23\text{ rather than } 18.85 \]

Ask what it costs

Why: An untested assumption.

Figure (svg): The solution to Worked example what pooling would have assumed shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{pooled df} = n_1 + n_2 - 2 = 23 \]

Verify: confirm the practical size of the difference here

Why: With 23 degrees of freedom rather than 18.85 the two-tailed p-value for this statistic would be about 0.0046 rather than 0.0054 — a small change that would not alter the decision. The reason to follow the rule anyway is that the difference is not always small: when the sample sizes and spreads are both very unequal, pooling can be badly wrong, and the Aspin-Welch approach never is.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 513

31. Trap: pooling the variances to simplify the degrees of freedom

Trap

The trap

\[ \text{df} = n_1 + n_2 - 2 = 23, \text{ using a pooled } s \]

Combine the samples to estimate one common spread

Why: It gives a whole number and more degrees of freedom.

\[ \text{but the two populations may not share a spread} \]

The book's instruction is explicit: if the question comes up, do not pool the variances.

The fix

\[ \text{df} \approx 18.85, \text{ from the Aspin-Welch formula} \]

Use each sample's own variance and let the df come out fractional

Why: The test assumes nothing about the two spreads.

A non-integer degrees of freedom looks wrong to anyone used to n minus one, and that is often what prompts the pooling. It is worth accepting: the fractional value is the honest one, and every calculator and statistics package computes it without complaint.

32. Bracket the degrees of freedom

Faded example

Two samples of 11 and 9, using the Aspin-Welch formula.

Fill in the blanks

\text18 11 + 9 - 2 = 17.4, \text___ ___, \text___

Why: Example 10.2's Aspin-Welch degrees of freedom are about 17.40, below the pooled 18 — as the formula's output always is when the two variances differ.

33. Why not pool?

Prediction

Commit before reasoning.

Predict first

What does pooling the variances assume that the Aspin-Welch approach does not?

  • That the two populations have the same standard deviation
  • That the two samples are the same size
  • That both populations are normal
  • That the means are equal

Correct: That the two populations share a standard deviation.

Why: Pooling combines both samples into one estimate of a single common spread, which only makes sense if such a spread exists. Normality is assumed by both approaches, equal sample sizes by neither, and equal means is the null being tested rather than an assumption.

34. Which conditions does the test need?

Sorting

From the book's list.

Sort into buckets

Sort each statement.

Required
the two samples are independent simple random samples; small samples should come from normal populations; large samples need not come from normal populations
Not required
the two populations have equal standard deviations; the two sample sizes must be equal
yes
The book lists it among the test's characteristics.
no
It is either the assumption pooling would make, or a convenience that the formulas do not need.

Item (b) is the one to remember, since it is exactly what the Aspin-Welch approach avoids assuming — and the reason its degrees of freedom are more complicated than a pooled test's.

35. Running the test

Section

Section 4

36. The six steps, with the new statistic

Concept

With the hypotheses written, the standard error computed and the degrees of freedom obtained, the remaining steps are chapter 9's unchanged: find the p-value on the t distribution, compare it with a preset alpha, decide, and write the conclusion in context.

the test statistic — The difference of the sample means, minus the difference the null claims, divided by the standard error. Under a null of no difference the middle term is zero, so the statistic is simply the observed difference over its standard error.

\[ t = \frac{\bar{x}_1 - \bar{x}_2}{\text{SE}}, \quad \text{df from Aspin-Welch} \]

The book draws the picture on the scale of the difference rather than of t, which is worth imitating once: it says half the p-value is below minus 1.2 and half is above 1.2, in hours of sport. Working on the original scale keeps the reader's attention on the quantity that matters, and the t scale is only a device for looking up an area.

Figure (svg): A t distribution with both tails beyond plus and minus three point one four shaded

The book: half the p-value is below minus 1.2 and half is above 1.2, on the scale of the difference itself.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 513-516 — Examples 10.1 and 10.2, worked

37. A two-tailed test on 18.85 degrees of freedom

Picture it

Example 10.1: do girls and boys differ?

Figure (svg): A t distribution with both tails beyond plus and minus three point one four shaded

The book: half the p-value is below minus 1.2 and half is above 1.2, on the scale of the difference itself.

The two shaded tails together give 0.0054, well below the 5 percent level, so the null is rejected — the means differ. On the original scale the boys average 1.2 hours more per day, and the test says a gap that size is very unlikely to have arisen by chance from equal populations.

38. Worked example: Example 10.1 in full

Worked example

A two-tailed test at the 5 percent level.

\[ \bar{x}_g = 2, \; \bar{x}_b = 3.2; \; \alpha = 0.05 \]

Compute the difference

Why: Two minus 3.2.

\[ -1.2\text{ hours} \]

Divide by the standard error

Why: Minus 1.2 over 0.3819.

\[ t = -3.14 \]

Find the p-value

Why: Two tails on 18.85 df.

\[ 0.0054 \]

Decide

Why: Alpha exceeds it.

\[ \text{reject } H 0 \]

Figure (svg): The solution to Worked example Example 10.1 in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ t \approx -3.14, \quad p = 0.0054 < 0.05 \]

Verify: confirm the direction of the difference is reported, not just its significance

Why: Rejecting says the means differ; it does not by itself say which is larger. The sample means make that clear — boys averaged 3.2 hours against girls' 2 — and a conclusion should state it, since a reader wants to know the direction as well as that a difference exists. The two-tailed test establishes the difference; the data establish its sign.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 513-514

39. Compute the statistic

Faded example

Sample means 2 and 3.2, standard error 0.3819.

Fill in the blanks

t = \frac-3.140.0054 \approx ___, \quad \text___ p = ___

Why: The difference is about 3.14 standard errors below zero, and both tails beyond that on 18.85 degrees of freedom give 0.0054.

40. Worked example: Example 10.2 in full

Worked example

A right-tailed test at the 1 percent level.

\[ \bar{x}_A = 4, \; \bar{x}_B = 3.5; \; \alpha = 0.01 \]

Compute the difference

Why: Four minus 3.5.

\[ 0.5\text{ classes} \]

Divide by the standard error

Why: 0.5 over 0.5618.

\[ t = 0.89 \]

Find the p-value

Why: Right tail on 17.40 df.

\[ 0.1928 \]

Decide

Why: Alpha is far below it.

Figure (svg): The solution to Worked example Example 10.2 in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ t \approx 0.89, \quad p = 0.1928 > 0.01 \]

Verify: confirm why half a class is not enough evidence

Why: The gap of 0.5 classes is less than one standard error of 0.5618, so a difference this size arises easily by chance from populations with equal means. Comparing the raw difference against the standard error before computing anything would have predicted a large p-value — a gap under one standard error is never significant at any conventional level.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 514-516

41. Trap: judging a difference without its standard error

Trap

The trap

\[ \bar{x}_A - \bar{x}_B = 0.5 \text{ classes, so college A is ahead} \]

Read the observed gap as the answer

Why: Half a class is a real difference in the data.

\[ t = 0.89, \; p = 0.1928 \]

A gap smaller than one standard error is exactly what equal populations produce routinely.

The fix

\[ \text{compare the gap with the standard error: } \frac{0.5}{0.5618} = 0.89 \]

Standardise before judging

Why: The book: very different means can occur by chance if there is great variation.

This is the whole reason the test exists. Two samples always differ somewhat, so a difference on its own establishes nothing — the question is whether it is large relative to how much the difference would vary by chance, and only the standard error answers that.

42. Predict the outcome

Estimation

A difference of 0.5 with a standard error of 0.56.

Predict first

Without computing, will this be significant at any conventional level?

  • No: the gap is under one standard error
  • Yes: any difference is significant with enough data
  • Yes: 0.5 is a large gap
  • It cannot be judged without the p-value

Correct: No: the gap is under one standard error.

Why: A statistic under 1 leaves a right-tail area above 0.15 on any t distribution, so no conventional alpha would reject. Dividing the gap by the standard error before computing gives an immediate read on the outcome, and a ratio under about 1.7 rarely reaches significance one-tailed.

43. One of these is false

Two truths and a lie

All three concern running the test.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Under a null of no difference the statistic is the gap over its standard error
  • C. A two-tailed test establishes a difference but not its direction
  • B. A larger observed gap always gives a smaller p-value

Survives elimination: B

Why: The survivor is false as stated, because the p-value depends on the gap relative to the standard error. A gap of 5 with a standard error of 10 is far less significant than a gap of 1 with a standard error of 0.1 — which is exactly the point the book makes about variation among the samples.

44. Explain the standardising

Explain it

A classmate says college A's graduates plainly take more maths, since 4 exceeds 3.5.

Discussion prompt

In two sentences or fewer, explain what is missing.

Hint: Ask how much two sample means would differ by chance.

Answer:

Two samples from identical populations would still differ, and here the typical size of that chance difference is about 0.56 classes — larger than the 0.5 observed.

So the gap is smaller than the noise, which is why the p-value is 0.19 and the evidence is insufficient.

45. Conditions and approximations

Section

Section 5

46. What must hold, and when a normal will serve

Concept

The two independent samples must be simple random samples from two distinct populations. If the sample sizes are small the distributions matter and should be normal; if they are large the distributions need not be. The book adds that when the sum of the sample sizes exceeds 30, a normal distribution can approximate the Student t — but to use the t whenever possible.

the normal approximation — Permitted when n-one plus n-two exceeds 30, since the t approaches the normal as degrees of freedom grow. The book's own advice is to use the t regardless, because a calculator makes it no harder.

\[ n_1 + n_2 > 30 \;\Rightarrow\; \text{normal is adequate, but } t \text{ is better} \]

The advice to use the t whenever possible is the same judgement section 8.2 reached about one-sample intervals: the correction costs nothing now that calculators compute any t, and skipping it always errs in the direction of overstating the evidence. The approximation note is best read as reassurance about older tables rather than as a recommendation.

Figure (svg): The conditions and the rule against pooling the variances

Pooling assumes the two populations share a standard deviation, which this test deliberately does not assume.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-514 — the conditions and the note on the normal approximation

47. The conditions again

Picture it

What has to hold for the test to mean anything.

Figure (svg): The conditions and the rule against pooling the variances

Pooling assumes the two populations share a standard deviation, which this test deliberately does not assume.

The first condition is the one no sample size repairs, exactly as in chapters 8 and 9. Two independently drawn random samples is a statement about how the data were collected, and a study that compares self-selected groups cannot be rescued by the arithmetic that follows.

48. Worked example: checking the conditions

Worked example

Example 10.1's setting.

\[ n_g = 9, \; n_b = 16; \; \text{both populations normal} \]

Independence

Why: Two distinct groups.

Random sampling

Why: Stated as a study.

Sample sizes

Why: Nine and sixteen.

Normality

Why: Stated in the problem.

Figure (svg): The solution to Worked example checking the conditions shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{small } n \;\Rightarrow\; \text{normality matters, and is given} \]

Verify: confirm why the problem bothers to state normality

Why: With samples of 9 and 16 the central limit theorem cannot be relied on, so the populations' own shapes carry the assumption — which is why the book says each population has a normal distribution rather than leaving it implicit. For samples of a hundred each the sentence would be unnecessary, and its presence is a signal that the sample sizes are small enough to need it.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-513

49. Does the test apply?

Discrimination

Check independence, randomness and the sample sizes.

Sort into buckets

Sort each study.

This test applies
two random samples of 9 and 16 from normal populations; two random samples of 30 each, populations skewed; random samples of 11 and 9 from normal populations
It does not
before and after measurements on the same 8 subjects; volunteers for treatment against those who declined
ok
Two independent random samples, with either normal populations or samples large enough not to need them.
no
Either the samples are not independent, or the groups were not randomly formed.

50. Worked example: when a normal would serve

Worked example

The book's note about the sum of the sample sizes.

\[ n_1 + n_2 = 25 \text{ against } n_1 + n_2 = 60 \]

At a total of 25

Why: Below 30.

At a total of 60

Why: Above 30.

Recall the advice

Why: Use t whenever possible.

Say why

Why: It costs nothing.

Figure (svg): The solution to Worked example when a normal would serve shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n_1 + n_2 > 30 \;\Rightarrow\; \text{normal permitted} \]

Verify: confirm the direction of the error a normal would make

Why: The t has heavier tails, so a normal gives a smaller p-value on the same statistic — overstating the evidence, exactly as section 8.2 found. At 18.85 degrees of freedom the difference is small; at 5 it would not be. Since the error always runs the same way, the safe default is the t.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, p. 514

51. Trap: treating a large total as licence to skip the t

Trap

The trap

\[ n_1 + n_2 = 60 > 30 \;\Rightarrow\; \text{use a normal} \]

Apply the approximation because it is permitted

Why: The book says it is adequate.

\[ \text{but the book also says to use the t whenever possible} \]

Permitted and preferable are different things, and the approximation always errs toward significance.

The fix

\[ \text{use the t, with Aspin-Welch degrees of freedom} \]

Take the approximation as reassurance rather than instruction

Why: It cost something when tables were the only tool; it costs nothing now.

The whole note exists because a t table needs a degrees-of-freedom row and an Aspin-Welch value is fractional — awkward on paper and trivial on a calculator. Reading a historical convenience as current advice is how outdated shortcuts survive.

52. One of these is false

Two truths and a lie

All three concern the conditions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Normality matters more when the samples are small
  • C. The samples must be independent of each other
  • B. A large total sample size makes random sampling unnecessary

Survives elimination: B

Why: The survivor is false, and it is the same error chapters 8 and 9 warned about. A larger non-random sample estimates a biased quantity more precisely, which makes the problem harder to detect rather than easier.

53. Which way does the approximation err?

Prediction

Commit before reasoning.

Predict first

Using a normal instead of the t on the same statistic gives what?

  • A smaller p-value, overstating the evidence
  • A larger p-value
  • The same p-value
  • An error that could go either way

Correct: A smaller p-value.

Why: The t has heavier tails than the normal, so it leaves more area beyond any statistic — meaning the correct p-value is the larger one. The approximation therefore always flatters the result, which is why the book's advice to prefer the t is worth following even when the shortcut is permitted.

54. When is normality critical?

Faded example

Two studies, with samples of 9 and 16, and of 200 and 250.

Fill in the blanks

\textfirst central limit theorem \text___ ___ \text___

Why: With samples of 9 and 16 the populations' own shapes carry the assumption; with 200 and 250 the central limit theorem makes the sample means normal whatever the populations look like.

55. One-sample against two-sample

Comparison

Fill the blanks. Three rows change and the rest carry over.

Comparison matrix

One sample (ch. 9)Two samples (10.1)
Parametermumu-1 minus mu-2
Point estimatex-barx-bar-1 minus x-bar-2
Standard errors over root nroot of s1-squared over n1 plus s2-squared over n2
Degrees of freedomn minus onethe Aspin-Welch formula, usually fractional

Everything not in this table — the meaning of a p-value, the two error types, the decision rule, the conclusion wording — is unchanged from chapter 9, which is why this chapter can cover four tests in the space chapter 9 gave to one.

56. A two-sample t test, in order

Pattern

Six steps, and only the fourth is new.

  1. Confirm the two samples are independent random samples, and that small samples come from approximately normal populations.
  2. Write the hypotheses, naming which group is 1 and which is 2, and fix that order for the rest of the problem.
  3. Read the tail from the alternative, defaulting to two-tailed when the wording gives no direction.
  4. Compute the standard error by adding each variance over its own sample size, and get the degrees of freedom from the Aspin-Welch formula — without pooling.
  5. Compute the statistic as the difference of sample means over the standard error, and find the p-value on that t.
  6. Compare with alpha, decide, and write the conclusion naming the level, the direction and the context.

Check the standard error exceeds each group's own before proceeding, and divide the observed gap by it to predict roughly what the p-value will be.

OpenStax Introductory Business Statistics 2e, §10.1 Comparing Two Independent Population Means §10.1 Comparing Two Independent Population Means

57. Check yourself 1 of 3

Check

The standard error.

Check your understanding

Two samples have s = 0.866 with n = 9, and s = 1.00 with n = 16. What is the standard error of the difference?

  • A. About 0.382 (correct)
  • B. About 0.144
  • C. About 0.539
  • D. About 0.264

Answer: A

Why: Add each variance over its own sample size — 0.0833 plus 0.0625 — and take the square root.

Why B tempts people
That subtracts the variances, which would make a difference more certain than its parts.
Why C tempts people
That adds the two standard errors rather than the variances.
Why D tempts people
That pools everything over the combined sample size, which the book forbids.

58. Check yourself 2 of 3

Check

Degrees of freedom.

Check your understanding

For samples of 9 and 16, what are the degrees of freedom?

  • A. About 18.85, from the Aspin-Welch formula (correct)
  • B. 23, from n1 plus n2 minus 2
  • C. 24, from n1 plus n2 minus 1
  • D. 8, from the smaller sample

Answer: A

Why: The Aspin-Welch formula combines both variances and sizes, and gives a value that is rarely a whole number.

Why B tempts people
That is the pooled value, which assumes the populations share a standard deviation.
Why C tempts people
That formula corresponds to no test in this chapter.
Why D tempts people
That would be the degrees of freedom for a one-sample test on the smaller group alone.

59. Check yourself 3 of 3

Check

Pooling.

Check your understanding

Why does the book instruct that the variances not be pooled?

  • A. Pooling assumes the two populations share a standard deviation (correct)
  • B. Pooling is arithmetically harder
  • C. Pooling gives fewer degrees of freedom
  • D. Pooling only works for equal sample sizes

Answer: A

Why: The Aspin-Welch approach makes no assumption about the two spreads, which is why its degrees of freedom are more complicated.

Why B tempts people
Pooling is easier, which is part of why it is tempting.
Why C tempts people
Pooling gives MORE degrees of freedom, which is the other part of its appeal.
Why D tempts people
Pooling can be applied at any sample sizes; the objection is to its assumption.

60. Where this shows up outside the textbook

Real world

A hospital compares recovery times under two surgical techniques. Technique A gives a mean of 5.1 days from 12 patients with a standard deviation of 1.9; technique B gives 6.3 days from 40 patients with a standard deviation of 3.8. A surgeon reports that A is clearly better, since its mean is over a day shorter and its variability is half as large.

Discussion prompt

Test the claim properly, and say what the comparison of variabilities does and does not establish.

Hint: The smaller, tidier sample is not necessarily the more informative one.

Answer:

The standard error is dominated by the larger, noisier group. Technique A contributes 1.9 squared over 12, about 0.301; technique B contributes 3.8 squared over 40, about 0.361. Their sum is 0.662, so the standard error is about 0.814 days.

\[ t = \frac{5.1 - 6.3}{0.814} \approx -1.47, \qquad \text{df} \approx 27.5, \qquad p_{\text{two}} \approx 0.15 \]

The observed gap of 1.2 days is under one and a half standard errors, so the two-tailed p-value is around 0.15 and the null is not rejected at any conventional level. The surgeon's clearly is not supported: a difference this size arises readily by chance from populations with equal means.

The comparison of variabilities establishes nothing about the means and may not even establish a difference in spread. With only 12 patients, technique A's standard deviation of 1.9 is itself very imprecisely estimated — a sample that small routinely produces a spread half or double the population's. Testing whether two variances differ is a separate procedure, and it is section 13.4's subject rather than this one's.

Two design points follow. The imbalance is costly: 12 against 40 puts most of the uncertainty in the smaller group despite its tidier data, and a dozen more patients on technique A would sharpen the comparison more than a dozen more on B. And a difference of 1.2 days may well be clinically important, so the honest report is not that the techniques are equivalent but that this study was too small to settle it — which is a call for a larger trial rather than a conclusion.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why are the two variances added rather than subtracted when forming the standard error of a difference?

  • Because subtraction could give a negative number
  • Because the samples are independent, so their uncertainties combine rather than cancel
  • Because the means were subtracted
  • Because variances are always positive

Correct: Because the uncertainties combine rather than cancel.

\[ \text{SE} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}} > \max\left(\frac{s_1}{\sqrt{n_1}}, \frac{s_2}{\sqrt{n_2}}\right) \]

Why: Two independent quantities each carry error, and combining them by any arithmetic — adding or subtracting — combines those errors. The difference of two sample means is therefore more variable than either mean alone, which is why the standard error must exceed each group's own. That a negative would be impossible is a symptom of the error rather than its reason.

62. Explain it to someone a year behind you

Explain it

They used degrees of freedom of n-one plus n-two minus two, saying the fractional Aspin-Welch value looked wrong.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what their formula assumes about the two populations.

Answer:

Their formula is the POOLED one, and it only applies if the two populations share a standard deviation — an assumption this test deliberately avoids making.

A fractional degrees of freedom is the honest output of the Aspin-Welch formula, and the book says explicitly not to pool the variances.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Writing the hypotheses and fixing which group is 1
  • Computing the standard error by adding the two variances
  • Getting the degrees of freedom without pooling
  • Running the test and stating the direction in the conclusion

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, name group 1 and group 2 in writing and keep that order. For the second, each variance over its own n, added, then rooted — and check the result exceeds both. For the third, use Aspin-Welch and accept a fractional value. For the fourth, report which mean was larger as well as whether the difference was significant. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, draw two columns headed one sample and two samples, and fill in four rows: parameter, point estimate, standard error and degrees of freedom — circling the three that change. Beneath, write the standard error formula and beside it the two things that could go wrong with it: subtracting the variances, and pooling them. In the middle of the page, work Example 10.1 completely: write the hypotheses in both notations, compute each variance over its own sample size, add them, take the root to get 0.3819, compute the statistic of about minus 3.14, and note the degrees of freedom as 18.85 with a remark that it is not a whole number. Draw the t curve with both tails shaded and label the p-value 0.0054, then write the conclusion naming the level and the direction. Beside it, work Example 10.2 the same way and note that the gap of 0.5 is smaller than the standard error of 0.56 — then write one sentence on why that alone predicts a large p-value. At the bottom, list the five conditions from the book, marking which one no sample size repairs.

Check your standard error in each example by confirming it exceeds both individual standard errors — 0.289 and 0.25 for the first example, so 0.382 is right. Check the degrees of freedom by confirming each falls below the pooled n-one plus n-two minus two, which is what the Aspin-Welch formula always gives when the variances differ.

65. What you can do now

Recap

Five things, and the first three are the only genuinely new material.

If you seeThen
Two independent groups comparedThe parameter is a difference of means
No stated directionA two-tailed test by default
Two sample standard deviationsAdd each variance over its own n
A standard error below one group's ownThe variances were subtracted
A whole-number df from two samplesThe variances were probably pooled
A gap smaller than the standard errorIt will not be significant
Paired or before-and-after dataNot this test: see section 10.4

Section 10.2 takes the same comparison with the population standard deviations known, which replaces the t with a normal and removes the degrees-of-freedom problem entirely. The book concedes at the outset that the situation is not likely, and the section is short accordingly.

OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations §10.1, pp. 512-519 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §10.1 Two Population Means with Unknown Standard Deviations — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 512-519
  2. OpenStax Introductory Business Statistics 2e, §10.1 Comparing Two Independent Population Means — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0
  3. OpenStax Introductory Business Statistics 2e, §10.3 Test for Differences in Means: Assuming Equal Population Variances — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108