11.4 Test for Homogeneity

A new question answered by an existing procedure. A goodness-of-fit test can decide whether a population fits a given distribution, but it cannot decide whether two populations follow the same unknown distribution — and that second question is often the one that matters, since a researcher comparing men with women or before with after usually has no external distribution to test against. The test for homogeneity fills the gap, and the book is explicit that its statistic is computed exactly as the test of independence's: expected counts from the margins, the same sum of squared departures, the same right tail, the same requirement of at least five per cell. The hypotheses are worded about two populations having the same distribution, the degrees of freedom are the number of columns minus one for the usual two-population case, and the conclusion is only that the distributions differ — never how.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 11.4 Test for Homogeneity

Title

Statistics · Chapter 11 — The Chi-Square Distribution

Test for Homogeneity

2. By the end of this lesson you can

Objectives

Six outcomes, and none of them is a new calculation.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-581 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 11.2 tested a distribution against a claim; section 11.3 tested two factors for independence.

Discussion prompt

You have 250 men and 300 women, each reporting one of four living arrangements. You want to know whether the two groups distribute themselves the same way. Can a goodness-of-fit test do this?

Hint: Ask what a goodness-of-fit test needs before it can start.

Answer:

No. A goodness-of-fit test needs a distribution supplied from outside — the national percentages, the uniform assumption, the fair-coin probabilities. Here there is no such claim: the question is whether two groups agree with EACH OTHER, and neither one is a known standard.

The book states the gap plainly. The goodness-of-fit test can decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution.

What is available instead is each group's own data, which is exactly the situation section 11.3 handles: expected counts built from the table's own margins. So the arithmetic is already in hand, and only the question and its wording are new.

4. The same procedure, a different question

Concept

A test for homogeneity draws a conclusion about whether two populations have the same distribution. To calculate its test statistic, follow the same procedure as with the test of independence. The null is that the two distributions are the same; the alternative is that they are not.

a test for homogeneity — A chi-square test comparing two populations across a categorical variable with more than two response values. Common uses: men against women, before against after, east against west.

\[ H_0: \text{the distributions are the same} \quad\text{against}\quad H_a: \text{they are not} \]

The condition on the variable is worth noting: the book says it is categorical with more than two possible response values. With exactly two values there would be one proportion per population, and section 10.4's two-proportion test would answer the same question more directly — and would additionally say which population's proportion is larger, which this test never does.

Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can

The book's own framing: a new question answered by an existing procedure.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579

5. The gap this test fills

Section

Section 1

6. Two populations, no given distribution

Concept

The goodness-of-fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to draw that conclusion.

the unknown distribution — Neither population's distribution is claimed in advance. The test asks only whether they agree with each other, without ever estimating what they agree on.

\[ \text{unknown } D_1, D_2: \quad H_0: D_1 = D_2 \]

The word unknown carries the weight. If the national distribution of living arrangements were published, a goodness-of-fit test could check each sex against it separately — two tests, two questions. The homogeneity test compares the two groups directly, needs no external figure, and answers in one verdict. That is a genuinely different capability rather than a convenience.

Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can

The book's own framing: a new question answered by an existing procedure.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 578 — the opening paragraph of section 11.4

7. What each test can decide

Picture it

The question each one answers, and where the arithmetic comes from.

Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can

The book's own framing: a new question answered by an existing procedure.

The lower panel is the point of the section. A new question is being asked, but nothing new has to be computed — which is why the book's whole treatment fits on a page and a half.

8. Worked example: why goodness of fit will not do

Worked example

For Example 11.8's living arrangements.

\[ 250 \text{ men}, \; 300 \text{ women} \]

Recall what it requires

Why: A claimed distribution.

Check what is available

Why: Only two samples.

Consider using one group as the standard

Why: Women's proportions.

Conclude

Why: The tool does not fit.

Figure (svg): The solution to Worked example why goodness of fit will not do shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{no given } p_i \;\Longrightarrow\; \text{goodness of fit unavailable} \]

Verify: confirm why treating one sample as the standard would be wrong

Why: The women's proportions carry sampling error of their own, so comparing the men against them would ignore half the uncertainty in the comparison and produce a p-value that is too small. The homogeneity test accounts for variability in both samples at once, which is exactly what pooling the margins accomplishes — the same idea as section 10.4's pooled proportion.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 578

9. Which test applies?

Sorting

Each describes a situation.

Sort into buckets

Sort by which chi-square test fits.

Goodness of fit
blood types in one sample against known population percentages; leading digits against Benford's predicted percentages
Independence or homogeneity
living arrangements for 250 men against 300 women; voter preference before and after an event; smoking status by exercise level in one sample of 400
gof
An external distribution is quoted, and one group is compared against it.
other
No external distribution; expected counts must come from the table's own margins.

Items (b) and (c) are homogeneity — two separately drawn samples — while (d) is independence, since one sample of 400 was classified two ways. Telling those two apart is section 11.5's subject.

10. Worked example: Example 11.9's question

Worked example

Voter preferences among three candidates, surveyed before and after an earthquake.

\[ \text{before against after} \]

Count the populations

Why: Two surveys.

Count the response values

Why: Three candidates.

Check for a claimed distribution

Why: None given.

Name the test

Why: All three signals agree.

Figure (svg): The solution to Worked example Example 11.9's question shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \text{preferences were the same before and after} \]

Verify: confirm this is one of the book's own common uses

Why: The book lists before against after as a standard application, alongside men against women and east against west. All three share the same shape: two groups that are compared with each other rather than against any outside standard, which is the situation the test was built for.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581

11. Trap: using one sample's proportions as the expected distribution

Trap

The trap

\[ E_{\text{men}} = 250 \cdot \frac{91}{300}, \; 250 \cdot \frac{86}{300}, \; \ldots \]

Treat the women's observed proportions as a known distribution

Why: It looks like a goodness-of-fit test then.

\[ \text{but those proportions are estimates, not facts} \]

The women's sample has its own sampling error, and this treats it as exact — so the resulting p-value is too small.

The fix

\[ E = \frac{(\text{row total})(\text{column total})}{n}, \quad n = 550 \]

Pool both samples to build the expected counts

Why: The margins use all 550 people.

This mirrors section 10.4 exactly. There, testing two proportions used a POOLED estimate under the null rather than either sample's own value, for the same reason: if the null says the two agree, the best estimate of what they agree on uses both samples. Here the column totals do that pooling automatically.

12. One of these is false

Two truths and a lie

All three concern what this test is for.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It compares two populations with each other
  • C. No external distribution is needed
  • B. It requires the common distribution to be known

Survives elimination: B

Why: The survivor is false, and it reverses the section's opening point. The distribution is unknown; the test asks only whether the two populations share it, and it never estimates what it is.

13. What if the variable had two values?

Prediction

Commit before reasoning.

Predict first

Two populations compared on a categorical variable with only two response values. What applies?

  • A two-proportion test from section 10.4, which also says which is larger
  • Only a homogeneity test
  • A goodness-of-fit test
  • Nothing available

Correct: A two-proportion test.

Why: With two categories each population is described by one proportion, so section 10.4's test answers the question directly — and gives a direction, which a chi-square never does. The book's own note says the variable here is categorical with more than two possible response values.

14. The hypotheses

Faded example

Example 11.9's, in the book's wording.

Fill in the blanks

H_0: \textsame not the same; \; H_a: \text___ ___

Why: Both are statements about whole distributions rather than about any parameter, which is why they are written as sentences — the same practice as in sections 11.2 and 11.3.

15. The arithmetic is unchanged

Section

Section 2

16. Computed the same way as the test for independence

Concept

Use a chi-square test statistic, computed in the same way as the test for independence. The expected count in each cell is its row total times its column total over the grand total, and every value in the table must be at least five.

what carries over — The expected-count formula, the statistic, the right tail, and the condition. All four come from section 11.3 without modification.

\[ E = \frac{(\text{row})(\text{column})}{n}, \qquad \chi^2 = \sum \frac{(O-E)^2}{E} \]

There is a reason the arithmetic can be reused. If the two populations share a distribution, then knowing which population a person came from tells nothing about which category they fall in — which is precisely independence between the row variable and the column variable. So the null hypotheses of the two tests, differently worded, impose the same numerical structure on the table.

Figure (svg): The observed table of living arrangements for men and women college students

The statistic is 10.1287 on three degrees of freedom, giving a p-value of 0.0175.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579 — follow the same procedure as with the test of independence

17. Example 11.8's table

Picture it

Living arrangements for 250 men and 300 women.

Figure (svg): The observed table of living arrangements for men and women college students

The statistic is 10.1287 on three degrees of freedom, giving a p-value of 0.0175.

The two row totals were fixed by the researcher before any data were collected, which is the design signature of a homogeneity test. The column totals emerged from the data, and the expected counts use both.

18. Worked example: one expected count

Worked example

Men in dormitories, from Example 11.8.

\[ \text{row } 250, \; \text{column } 163, \; n = 550 \]

Total the two samples

Why: 250 plus 300.

\[ 550 \]

Total the dormitory column

Why: 72 plus 91.

\[ 163 \]

Apply the formula

Why: 250 times 163 over 550.

\[ 74.09 \]

Compare with observed

Why: 72 were seen.

Figure (svg): The solution to Worked example one expected count shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ E = \frac{(250)(163)}{550} = 74.09 \]

Verify: confirm the reading of that number

Why: It says that if living arrangements were distributed identically in both groups, the pooled dormitory rate of 163 out of 550 — about 29.6 percent — should apply to the men too, giving 29.6 percent of 250, or 74.09. The pooled rate is the null hypothesis's best estimate of the shared distribution, exactly as section 10.4's pooled proportion was.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 579

19. Another expected count

Faded example

Women with parents: row 300, column 137, total 550.

Fill in the blanks

E = \frac74.7388 = ___, \quad \text___ ___ \text___

Why: Women lived with parents considerably more than the pooled rate predicts — 88 against 74.73 — and this cell contributes 2.36 to the statistic of 10.13, the largest single share.

20. Worked example: Example 11.8 in full

Worked example

Do men and women college students have the same distribution of living arrangements? Test at 5 percent.

\[ \alpha = 0.05 \]

Hypotheses

Why: Same against not the same.

Degrees of freedom

Why: Four columns minus one.

\[ 3 \]

Check the condition

Why: Smallest expected.

\[ 36.36,\text{ above five} \]

Statistic

Why: Summed over eight cells.

\[ 10.1287 \]

Right tail on df 3

Why: Beyond 10.1287.

\[ 0.0175 \]

Figure (svg): The solution to Worked example Example 11.8 in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 10.1287, \quad \text{df} = 3, \quad p = 0.0175 < 0.05 \]

Verify: confirm the statistic against its degrees of freedom

Why: The mean of a chi-square with three degrees of freedom is 3, and the standard deviation is the root of six, about 2.45 — so 10.13 sits nearly three standard deviations above the centre, consistent with a p-value under 2 percent. Note the decision would reverse at 1 percent, since 0.0175 exceeds 0.01, which is why the book's assumption of 5 percent when none is given matters to the outcome.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-580

21. Error analysis: four claims about the arithmetic

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; \text{the expected counts come from the pooled margins} \\ &(2)\; \text{the statistic differs from the independence statistic} \\ &(3)\; \text{the test is right-tailed} \\ &(4)\; \text{the condition is at least five per cell} \end{aligned} \)

  • (1) is correct: row total times column total over the grand total, using both samples.
  • (2) is false. The book says the statistic is computed in the same way, and it is identical.
  • (3) is correct, for the same reason as in the previous two sections: every departure is squared.
  • (4) is correct, and the book states it twice — once as a note about expected values and once in its requirements list.

The only false claim is (2), and noticing that it is false is most of the lesson. Everything computational transfers, which is why the section is so short and why section 11.5 is needed to keep the tests distinct.

22. One of these is false

Two truths and a lie

All three concern what carries over from section 11.3.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The expected-count formula is identical
  • C. The right-tail reasoning is identical
  • B. The expected counts use only one sample

Survives elimination: B

Why: The survivor is false. The column totals pool both samples, which is what makes the expected counts reflect the shared distribution the null asserts. Using one sample alone would be the trap this lesson warns about.

23. Why can the arithmetic be reused?

Prediction

Commit before reasoning.

Predict first

Why does a homogeneity test compute exactly what an independence test computes?

  • If two populations share a distribution, group membership tells nothing about category — which is independence
  • Because both use chi-square tables
  • Because the tables look the same
  • It is a coincidence

Correct: Shared distributions make group and category independent.

Why: The two nulls are worded differently but impose the same structure: in both cases the cell probabilities factor into a row part and a column part. That is why one set of expected counts serves both, and why only the design and the question distinguish the tests.

24. The pooled rate

Estimation

Example 11.8: 163 of 550 lived in dormitories.

Predict first

Roughly what share, and what does it predict for the 300 women?

  • About 30 percent, so about 89
  • About 30 percent, so about 74
  • About 50 percent, so about 150
  • About 20 percent, so about 60

Correct: About 30 percent, so about 89.

Why: The pooled dormitory rate is 29.6 percent, and applying it to 300 women gives 88.91 — the book's expected count. Reading expected counts as a pooled rate applied to each sample is often faster than the formula and makes the null hypothesis's content visible.

25. Degrees of freedom

Section

Section 3

26. Columns minus one, for two populations

Concept

The book gives the degrees of freedom as the number of columns minus one. Both of its examples compare exactly two populations, so the table has two rows — and the general product formula reduces to exactly that.

the two-row case — With r equal to 2, rows minus one is 1, so the product collapses to columns minus one. The rule is section 11.3's, not a third thing to memorise.

\[ (r-1)(c-1) \;\overset{r=2}{\longrightarrow}\; (1)(c-1) = c - 1 \]

The reduction matters when more than two populations are compared, which is a common extension. Three regions across four categories give two times three, which is six — not three. Reading the book's rule as a special case rather than as the definition keeps that generalisation available and avoids an error that would otherwise be invisible.

Figure (svg): The book's own summary of the test for homogeneity

The book's five-part summary, reproduced. Every computational line points back to section 11.3.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581 — df = number of columns minus one, in both examples

27. The book's summary

Picture it

Five headings: hypotheses, statistic, degrees of freedom, requirements, common uses.

Figure (svg): The book's own summary of the test for homogeneity

The book's five-part summary, reproduced. Every computational line points back to section 11.3.

Every computational line points back to section 11.3. What the summary adds is the wording of the hypotheses and the description of when the test is used — which is genuinely the whole of what is new.

28. Worked example: both of the book's examples

Worked example

Checking the rule against the product formula.

\[ 2 \times 4 \text{ and } 2 \times 3 \]

Example 11.8

Why: Four columns minus one.

\[ 3 \]

By the product rule

Why: One times three.

\[ 3 \]

Example 11.9

Why: Three columns minus one.

\[ 2 \]

By the product rule

Why: One times two.

\[ 2 \]

Figure (svg): The solution to Worked example both of the book's examples shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (2-1)(4-1) = 3, \qquad (2-1)(3-1) = 2 \]

Verify: confirm the rules diverge as soon as there are more than two populations

Why: Comparing three regions across four categories gives columns minus one equal to 3 by the book's shortcut, but the product rule gives two times three, which is 6. The product rule is the correct one; the shortcut is safe only for the two-population case the book's examples use, and applying it more widely would halve the degrees of freedom.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581

29. Three populations

Faded example

Three regions compared across six categories.

Fill in the blanks

\text2 = (5)(___) = 10

Why: Ten, not five. The book's columns-minus-one shortcut applies only when there are exactly two populations, which is the case in both of its examples.

30. Worked example: Example 11.9 in full

Worked example

Voter preferences among three candidates, before and after an earthquake, at 5 percent.

\[ \alpha = 0.05 \]

Hypotheses

Why: Same against not the same.

Degrees of freedom

Why: Three columns minus one.

\[ 2 \]

Statistic

Why: Summed over six cells.

\[ 3.2603 \]

Compare with the mean

Why: Mean is 2.

Right tail

Why: Beyond 3.2603 on df 2.

\[ 0.1959 \]

Figure (svg): The solution to Worked example Example 11.9 in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 3.2603, \quad p = 0.1959 > 0.05 \]

Verify: confirm the conclusion is stated with the book's care

Why: The book's wording is that there is insufficient evidence to conclude that the distribution was not the same — not that it was the same. That distinction matters here because the after-survey had 636 respondents against 430 before, and the observed shares differ by a couple of percentage points; the test simply cannot separate that from sampling variation.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 580-581

31. Trap: applying columns minus one to more than two populations

Trap

The trap

\[ 3 \text{ regions} \times 5 \text{ categories} \;\Rightarrow\; \text{df} = 4 \]

Use the book's stated rule directly

Why: It says columns minus one.

\[ \text{but that rule assumed two rows} \]

With three populations the table has three rows, and rows minus one is 2 rather than 1.

The fix

\[ \text{df} = (3-1)(5-1) = 8 \]

Use the product rule, which the shortcut is a case of

Why: It holds for any number of populations.

The error halves the degrees of freedom, which shifts the reference curve well to the left and makes every statistic look more extreme than it is. Since a chi-square with 4 degrees of freedom has mean 4 while one with 8 has mean 8, a statistic near 10 would seem significant under the wrong count and unremarkable under the right one.

32. One of these is false

Two truths and a lie

All three concern the count.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. For two populations, columns minus one is correct
  • C. The general rule is rows minus one times columns minus one
  • B. Columns minus one holds for any number of populations

Survives elimination: B

Why: The survivor is false. Three populations across four categories give two times three, which is six rather than three, and using the shortcut there would halve the degrees of freedom.

33. Table to degrees of freedom

Matching

Match each comparison to its count.

Match the pairs

  • l1. 2 populations, 4 categories
  • l2. 2 populations, 3 categories
  • l3. 3 populations, 4 categories
  • l4. 4 populations, 5 categories
  • r1. 3
  • r2. 2
  • r3. 6
  • r4. 12

Why: The first two are the book's own examples. The last two show how quickly the count grows once more than two populations are compared, and why the product rule has to be the one held on to.

34. What does the shortcut assume?

Prediction

Commit before reasoning.

Predict first

The rule columns minus one is correct under what condition?

  • Exactly two populations are compared
  • The categories number more than two
  • The sample sizes are equal
  • Always

Correct: Exactly two populations.

Why: With two rows, rows minus one is 1 and the product formula collapses. Both of the book's examples compare two groups, so the shortcut is correct throughout section 11.4 — but it is a consequence of that design rather than a general rule.

35. What a rejection does not say

Section

Section 4

36. Only that the distributions differ

Concept

The book adds an explicit warning after Example 11.8: notice that the conclusion is only that the distributions are not the same. The test for homogeneity cannot be used to draw any conclusions about how they differ.

an omnibus test — One that detects any departure without identifying its kind. The price of a single verdict on a whole distribution is that the verdict carries no direction.

\[ \text{reject } H_0 \;\Longrightarrow\; D_1 \ne D_2 \text{ only} \]

The contrast with section 10.4 is sharp. A two-proportion test rejects and reports which proportion is larger, because its statistic keeps the sign of the difference. A chi-square squares every departure, so the sign is destroyed before the cells are added — and no amount of examining the statistic afterwards can recover it.

Figure (svg): A chi-square curve on three degrees of freedom with the right tail beyond 10.13 shaded thin

Example 11.8 decided: living arrangements differ between men and women college students.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 580 — we cannot use the test for homogeneity to draw conclusions about how they differ

37. Example 11.8 decided

Picture it

The statistic against its distribution.

Figure (svg): A chi-square curve on three degrees of freedom with the right tail beyond 10.13 shaded thin

Example 11.8 decided: living arrangements differ between men and women college students.

The p-value of 0.0175 supports one claim only: the two distributions differ. Which categories differ, and in which direction, is not something the test reports — though the cells can of course be described.

38. Worked example: what may and may not be said

Worked example

After rejecting in Example 11.8.

\[ p = 0.0175 \]

The test's verdict

Why: The alternative.

Adding a direction

Why: Not from this test.

Naming a category

Why: Not from this test.

Describing the cells

Why: A description.

Figure (svg): The solution to Worked example what may and may not be said shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{sufficient evidence that } D_{\text{men}} \ne D_{\text{women}} \]

Verify: confirm what the cells show, as description rather than inference

Why: Women lived with parents more than the pooled rate predicts, at 88 against 74.73, and men chose the other category more, at 45 against 36.36. Both are accurate descriptions of these 550 students and worth reporting — but each is a claim the test did not separately examine, and the book's warning is aimed exactly at presenting them as though it had.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-580

39. One of these is false

Two truths and a lie

All three concern the conclusion.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The conclusion is only that the distributions differ
  • C. The statistic destroys the direction of each departure
  • B. The test reports which population has the larger proportion in each category

Survives elimination: B

Why: The survivor is false and is exactly what the book warns against. A chi-square gives one unsigned verdict; recovering directions requires either separate comparisons or a test designed to give them.

40. Worked example: what a two-proportion test would add

Worked example

If living arrangements had only two categories.

\[ \text{dormitory or not} \]

The chi-square

Why: One verdict.

The two-proportion z

Why: Keeps the sign.

Its confidence interval

Why: Section 8.4's.

Note the trade

Why: Fewer categories, more detail.

Figure (svg): The solution to Worked example what a two-proportion test would add shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ z \text{ signed} \;\text{against}\; \chi^2 \ge 0 \text{ unsigned} \]

Verify: confirm this is a genuine trade rather than a defect

Why: A chi-square handles four categories at once, which no two-proportion test can. What it gives up is the direction, because summing squared departures across categories leaves nothing to attach a sign to. The right response is to choose the test that matches the question — an omnibus verdict across many categories, or a directed comparison of one.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 580

41. Error analysis: four conclusions from Example 11.8

Error analysis

The p-value was 0.0175 at a 5 percent level. Which are defensible?

Annotate

On: \( \begin{aligned} &(1)\; \text{the distributions are not the same} \\ &(2)\; \text{women live with parents more often than men} \\ &(3)\; \text{sex determines living arrangement} \\ &(4)\; \text{at 1 percent, the evidence is insufficient} \end{aligned} \)

  • (1) is the test's conclusion and the book's own wording.
  • (2) describes these data correctly, but the book warns the test draws no conclusions about how the distributions differ.
  • (3) is causal, and the data are observational.
  • (4) is correct: 0.0175 exceeds 0.01, so the same data would not reject at that level.

Statement (4) is worth noticing because the book assumes 5 percent when no level is given. The evidence here is moderate rather than overwhelming, and the decision genuinely depends on the level chosen in advance.

42. Decide at two levels

Faded example

Example 11.8's p-value is 0.0175.

Fill in the blanks

\textreject \alpha = 0.05: \; do not reject; \qquad \text___ \alpha = 0.01: \; ___

Why: The same data, two decisions. This is why the level has to be fixed before seeing the p-value, a point chapter 9 made and one that this example illustrates cleanly.

43. Why no direction?

Prediction

Commit before reasoning.

Predict first

Why can a chi-square test not say which way two distributions differ?

  • Every departure is squared, so the sign is gone before the cells are summed
  • Because the distributions are unknown
  • Because the sample sizes differ
  • Because it uses expected counts

Correct: Squaring destroys the sign.

Why: A cell running 20 above expectation and one running 20 below contribute identically. Once summed, there is nothing left in the statistic to indicate which cells ran which way — which is the price of combining every category into one number.

44. Explain the limit

Explain it

A classmate writes that Example 11.8 shows women are more likely to live with their parents.

Discussion prompt

In two sentences or fewer, tell them what the test supports.

Hint: Ask what the alternative hypothesis actually says.

Answer:

The test's alternative is only that the two distributions are not the same, and the book adds a warning that it cannot be used to draw conclusions about how they differ.

That women lived with parents more often is a true description of these 550 students, and worth reporting as such — but it is not a claim the test separately examined.

45. Homogeneity against independence

Section

Section 5

46. Told apart by the design, not the arithmetic

Concept

The two tests compute identical numbers from a two-way table. What separates them is what was fixed before the data were collected: one sample classified two ways gives a test of independence, while two separately drawn samples classified one way gives a test for homogeneity.

fixed margins — In a homogeneity study the row totals are chosen by the researcher — 250 men, 300 women. In an independence study only the grand total is, and both sets of margins emerge from the data.

\[ \text{one sample, two questions} \quad\text{against}\quad \text{two samples, one question} \]

This is why the tests can share a procedure without being the same test. The arithmetic examines whether the interior of the table matches its margins; what that comparison MEANS depends on where the margins came from. Section 11.5 turns this into a practical rule for choosing between them, based on how the hypotheses are worded.

Figure (svg): A card contrasting one sample cross-classified against two samples separately drawn

Two designs, one calculation: what was fixed before the data were collected decides which test it is.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579 — common uses: comparing two populations

47. Two designs, one calculation

Picture it

What was fixed in advance, in each case.

Figure (svg): A card contrasting one sample cross-classified against two samples separately drawn

Two designs, one calculation: what was fixed before the data were collected decides which test it is.

A table of numbers alone cannot say which study produced it. That is why the choice between the two tests is made from the description of how the data were collected, before any arithmetic begins.

48. Worked example: reading two designs

Worked example

Example 11.6 against Example 11.8.

\[ \text{which is which?} \]

Example 11.6

Why: 839 volunteers, one sample.

What was fixed

Why: Only the total.

Example 11.8

Why: 250 men and 300 women.

What was fixed

Why: Both row totals.

Figure (svg): The solution to Worked example reading two designs shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 839 \text{ classified twice} \quad\text{against}\quad 250 + 300 \text{ classified once} \]

Verify: confirm the arithmetic really is indistinguishable

Why: Both tests take the same table, build the same expected counts from row times column over n, sum the same statistic and read the same right tail. If the numbers of Example 11.8 had arisen from a single sample of 550 people classified by sex and living arrangement, every calculation would be identical and only the wording of the hypotheses and conclusion would change.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579

49. Independence or homogeneity?

Sorting

Each describes how data were collected.

Sort into buckets

Sort by which test the design calls for.

Test of independence
839 volunteers classified by type and hours; 400 students classified by anxiety and need to succeed
Test for homogeneity
250 men and 300 women asked about living arrangements; 430 voters before and 636 after an earthquake; 100 families and 200 singles asked what car they drive
ind
One sample, classified by two variables at once.
hom
Two samples drawn separately, each classified by one variable.

Items (a) and (d) are Examples 11.6 and 11.7; (b), (c) and (e) are Examples 11.8, 11.9 and Try It 11.8. The signal every time is how many samples were drawn, not what the table looks like.

50. Worked example: the sample sizes as a tell

Worked example

Round row totals often signal a homogeneity design.

\[ 250 \text{ and } 300 \]

Note the roundness

Why: Both exact hundreds or halves.

Compare Example 11.6

Why: 255, 290, 294.

Draw the inference

Why: A researcher picked the first.

Note the caution

Why: Not conclusive.

Figure (svg): The solution to Worked example the sample sizes as a tell shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 250, 300 \text{ chosen} \quad\text{against}\quad 255, 290, 294 \text{ observed} \]

Verify: confirm this is a hint rather than a rule

Why: Round numbers can arise by chance, and a homogeneity study might well end up with 247 men after non-responses. The reliable signal is always the description of how the data were collected — the book's examples say plainly that 250 men and 300 women were randomly selected, and that sentence is what settles it.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 579

51. Trap: choosing the test from the table alone

Trap

The trap

\[ \text{a } 2 \times 4 \text{ table} \;\Rightarrow\; \text{a test for homogeneity} \]

Read the design off the table's shape

Why: Homogeneity examples all have two rows.

\[ \text{but a } 2 \times 4 \text{ independence table is equally possible} \]

One sample of 550 people classified by sex and living arrangement would produce the same shape.

The fix

\[ \text{read how the data were collected, then choose} \]

Ask what was fixed before sampling began

Why: Two samples means homogeneity; one means independence.

Since the arithmetic is identical, choosing wrongly costs nothing numerically — the statistic, degrees of freedom and p-value all come out the same. What it costs is the interpretation: a homogeneity conclusion about two populations is a different claim from an independence conclusion about two factors, and only one of them fits the study that was run.

52. One of these is false

Two truths and a lie

All three concern the two designs.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The arithmetic is the same for both
  • C. A homogeneity design fixes the row totals in advance
  • B. The two can be told apart from the table of counts

Survives elimination: B

Why: The survivor is false. The same table could arise from either design, so the choice depends entirely on the description of how the data were collected — which is why the book introduces each example with a sentence about the sampling.

53. What does choosing wrongly cost?

Prediction

Commit before reasoning.

Predict first

A homogeneity study is analysed as a test of independence. What changes?

  • Nothing numerically; only the interpretation is wrong
  • The statistic changes
  • The degrees of freedom change
  • The test becomes invalid

Correct: Only the interpretation.

Why: Every number comes out the same, since the procedures are identical. What differs is the claim being made: independence between two factors within one population, against agreement between two populations. Reporting the wrong one misdescribes the study even though the arithmetic is sound.

54. How many samples?

Estimation

Try It 11.8 surveys 100 families and 200 singles about car type.

Predict first

Which test, and how many degrees of freedom across five car types?

  • Homogeneity, 4
  • Independence, 4
  • Homogeneity, 8
  • Goodness of fit, 4

Correct: Homogeneity, with 4 degrees of freedom.

Why: Two separately drawn samples of chosen sizes make it a homogeneity test, and five categories give columns minus one, which is four — the same as the product rule's one times four. The statistic there is 62.91, far beyond the mean of 4, so the distributions differ decisively.

55. The three chi-square tests so far

Comparison

Fill the blanks. Two of the rows are identical across all three.

Comparison matrix

Goodness of fitIndependenceHomogeneity
Samples drawnoneonetwo or more
Expected counts froman outside claimthe marginsthe margins
Degrees of freedomcells minus one(r-1)(c-1)columns minus one, for two groups
Tailrightrightright

The last two columns agree on everything computational, which is the whole point of this section. What separates them is the design behind the table and the wording of the hypotheses, and section 11.5 makes that separation into a rule.

56. Running a test for homogeneity, in order

Pattern

Six steps, and five of them are section 11.3's.

  1. Confirm the design: two or more separately drawn samples, each classified by one categorical variable.
  2. State the hypotheses: the distributions of the populations are the same, against not the same.
  3. Build expected counts as row total times column total over the grand total, pooling all samples.
  4. Check every value is at least five.
  5. Count degrees of freedom as rows minus one times columns minus one, which is columns minus one for two groups.
  6. Compute the statistic, take the right tail, decide, and conclude that the distributions do or do not differ — without saying how.

Read expected counts as the pooled rate applied to each sample; it is faster than the formula and makes the null's content visible.

OpenStax Introductory Business Statistics 2e, §11.5 Test for Homogeneity §11.5 Test for Homogeneity

57. Check yourself 1 of 3

Check

Choosing the test.

Check your understanding

Two hundred urban and 200 rural residents each name their preferred transport from five options. Which test applies?

  • A. A test for homogeneity (correct)
  • B. A goodness-of-fit test
  • C. A test of independence
  • D. A two-proportion test

Answer: A

Why: Two separately drawn samples, each classified by one categorical variable with more than two values, is exactly the homogeneity design.

Why B tempts people
No external distribution is given to test against.
Why C tempts people
That would require one sample classified two ways, not two samples of chosen sizes.
Why D tempts people
There are five categories, not two, so a single proportion does not describe either group.

58. Check yourself 2 of 3

Check

Degrees of freedom.

Check your understanding

Three regions are compared across four categories. What are the degrees of freedom?

  • A. 6 (correct)
  • B. 3
  • C. 11
  • D. 2

Answer: A

Why: Two times three is six. The book's columns-minus-one rule is the two-population case of this, and does not apply with three groups.

Why B tempts people
That applies columns minus one, which assumes exactly two populations.
Why C tempts people
That is cells minus one, the goodness-of-fit rule.
Why D tempts people
That is rows minus one alone, without multiplying by columns minus one.

59. Check yourself 3 of 3

Check

The conclusion.

Check your understanding

A test for homogeneity rejects at 5 percent. What may be concluded?

  • A. The two distributions are not the same (correct)
  • B. The first population has more in category one
  • C. Group membership causes the difference
  • D. The distributions are the same

Answer: A

Why: The book is explicit that the conclusion is only that the distributions are not the same, and that no conclusion about how they differ may be drawn.

Why B tempts people
That is a direction, and squaring destroyed every direction before the cells were summed.
Why C tempts people
Causal claims need assignment, which these observational designs do not have.
Why D tempts people
That is the null, which was rejected.

60. Where this shows up outside the textbook

Real world

A company surveys 400 employees before a policy change and 400 after, asking each to pick one of six reasons for job satisfaction. A test for homogeneity gives a p-value of 0.31, and the report concludes that the policy change had no effect on why employees are satisfied.

Discussion prompt

Assess that conclusion, and say what else would need checking.

Hint: Ask what failing to reject establishes, and what the test could have detected.

Answer:

The conclusion overstates a non-rejection, in exactly the way chapter 9 warned against. A p-value of 0.31 says the data are consistent with no change; it does not establish that no change occurred. The honest wording is that there is insufficient evidence that the distribution of reasons differed before and after.

The power question is the substantive one. With 400 per group and six categories, expected counts run around 130 per cell if the categories are roughly balanced — enough to detect a shift of several percentage points, but not a shift of one or two. Before treating the non-result as informative, someone should ask how large a change the study could have found, and whether a change smaller than that would still have mattered to the company.

\[ \text{df} = (2-1)(6-1) = 5, \qquad \chi^2 \text{ needed to reject at } 5\% \approx 11.07 \]

Two design points deserve attention too. If the same employees were surveyed twice, the two samples are not independent, and a chi-square test for homogeneity assumes they are — paired categorical data need a different treatment, the same way section 10.5 needed a paired t rather than a two-sample one. And if the surveys drew different employees, turnover between them could change the population itself, so the two samples might differ in composition rather than in attitude.

What the study supports: no detectable change in the distribution of reasons, at this sample size. What it does not support: that the policy had no effect. The difference between those two statements is the whole content of chapter 9's asymmetry between rejecting and failing to reject, and it matters most exactly when the result is the one the reporter was hoping for.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

What distinguishes a test for homogeneity from a test of independence?

  • A different test statistic
  • The design behind the table and the wording of the hypotheses; the arithmetic is identical
  • Different degrees of freedom rules
  • One is right-tailed and one is not

Correct: The design and the wording; the arithmetic is identical.

\[ E = \frac{RC}{n}, \; \chi^2 = \sum \frac{(O-E)^2}{E}, \; \text{df} = (r-1)(c-1) \quad \text{in both} \]

Why: The book says to follow the same procedure as with the test of independence, and it means it — the same expected counts, the same statistic, the same right tail, the same condition. A homogeneity study draws two samples of chosen sizes and asks whether two populations agree; an independence study draws one sample and asks whether two factors are related. Nothing in the table itself distinguishes them.

62. Explain it to someone a year behind you

Explain it

They used columns minus one for a table comparing four regions across three categories, getting df = 2.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask how many rows the book's examples had.

Answer:

Columns minus one is the book's shortcut for comparing exactly two populations, where rows minus one happens to equal 1 — but with four regions there are four rows, so rows minus one is 3.

The degrees of freedom are three times two, which is six rather than two.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Saying which question a goodness-of-fit test cannot answer
  • Telling a homogeneity design from an independence design
  • Counting degrees of freedom for more than two populations
  • Stating exactly what a rejection does and does not establish

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, whether two populations share an unknown distribution. For the second, count the samples drawn, not the rows in the table. For the third, use the product rule, of which columns minus one is a special case. For the fourth, the distributions differ, never how. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twelve minutes.

Draw it

At the top, write the question a goodness-of-fit test cannot answer — whether two populations follow the same unknown distribution — and beside it the hypotheses for a test of homogeneity in words. Underneath, write in capitals that the arithmetic is the test of independence's, unchanged, and list the four things that carry over: the expected-count formula, the statistic, the right tail, and the at-least-five condition. In the middle, work Example 11.8: the two-by-four table of 72, 84, 49, 45 over 91, 86, 88, 35 with row totals 250 and 300 and column totals 163, 170, 137, 80; the expected count for men in dormitories as 250 times 163 over 550, which is 74.09; the degrees of freedom as three; the statistic 10.1287 and p-value 0.0175. Beside it write the two decisions — reject at 5 percent, do not reject at 1 percent. At the bottom left, write the degrees-of-freedom rule as a product and show columns minus one falling out of it when there are two rows. At the bottom right, write the book's warning in a box: the conclusion is only that the distributions are not the same, never how they differ.

Check your expected count by reading it the other way: 163 of 550 is 29.6 percent, and 29.6 percent of 250 is 74.09. Check your degrees-of-freedom box by testing it on three populations across four categories, where the product rule gives 6 and the shortcut would wrongly give 3.

65. What you can do now

Recap

Six things, and only two of them are new since section 11.3.

If you seeThen
Two separately drawn samplesA test for homogeneity
One sample classified two waysA test of independence instead
An external distribution quotedA goodness-of-fit test instead
Two populations, two categoriesA two-proportion test from section 10.4, which gives a direction
Expected counts neededRow times column over n, pooling every sample
Two populations comparedDegrees of freedom are columns minus one
Three or more populationsUse the product rule; the shortcut no longer holds
A rejectionThe distributions differ; describe the cells separately if at all

Section 11.5 is the book's own comparison of the three tests. Since so little separates them computationally, it turns the distinction into a practical rule based on how the hypotheses are worded — which is the reliable signal when a problem does not describe its sampling clearly.

OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-581 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 578-581
  2. OpenStax Introductory Business Statistics 2e, §11.5 Test for Homogeneity — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108