A new question answered by an existing procedure. A goodness-of-fit test can decide whether a population fits a given distribution, but it cannot decide whether two populations follow the same unknown distribution — and that second question is often the one that matters, since a researcher comparing men with women or before with after usually has no external distribution to test against. The test for homogeneity fills the gap, and the book is explicit that its statistic is computed exactly as the test of independence's: expected counts from the margins, the same sum of squared departures, the same right tail, the same requirement of at least five per cell. The hypotheses are worded about two populations having the same distribution, the degrees of freedom are the number of columns minus one for the usual two-population case, and the conclusion is only that the distributions differ — never how.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 11 — The Chi-Square Distribution
Test for Homogeneity
Objectives
Six outcomes, and none of them is a new calculation.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-581 — the section these objectives are drawn from
Warm-up
Section 11.2 tested a distribution against a claim; section 11.3 tested two factors for independence.
Discussion prompt
You have 250 men and 300 women, each reporting one of four living arrangements. You want to know whether the two groups distribute themselves the same way. Can a goodness-of-fit test do this?
Hint: Ask what a goodness-of-fit test needs before it can start.
Answer:
No. A goodness-of-fit test needs a distribution supplied from outside — the national percentages, the uniform assumption, the fair-coin probabilities. Here there is no such claim: the question is whether two groups agree with EACH OTHER, and neither one is a known standard.
The book states the gap plainly. The goodness-of-fit test can decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution.
What is available instead is each group's own data, which is exactly the situation section 11.3 handles: expected counts built from the table's own margins. So the arithmetic is already in hand, and only the question and its wording are new.
Concept
A test for homogeneity draws a conclusion about whether two populations have the same distribution. To calculate its test statistic, follow the same procedure as with the test of independence. The null is that the two distributions are the same; the alternative is that they are not.
a test for homogeneity — A chi-square test comparing two populations across a categorical variable with more than two response values. Common uses: men against women, before against after, east against west.
\[ H_0: \text{the distributions are the same} \quad\text{against}\quad H_a: \text{they are not} \]
The condition on the variable is worth noting: the book says it is categorical with more than two possible response values. With exactly two values there would be one proportion per population, and section 10.4's two-proportion test would answer the same question more directly — and would additionally say which population's proportion is larger, which this test never does.
Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579
Section
Section 1
Concept
The goodness-of-fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to draw that conclusion.
the unknown distribution — Neither population's distribution is claimed in advance. The test asks only whether they agree with each other, without ever estimating what they agree on.
\[ \text{unknown } D_1, D_2: \quad H_0: D_1 = D_2 \]
The word unknown carries the weight. If the national distribution of living arrangements were published, a goodness-of-fit test could check each sex against it separately — two tests, two questions. The homogeneity test compares the two groups directly, needs no external figure, and answers in one verdict. That is a genuinely different capability rather than a convenience.
Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 578 — the opening paragraph of section 11.4
Picture it
The question each one answers, and where the arithmetic comes from.
Figure (svg): A card showing the question a goodness-of-fit test cannot answer and this test can
The lower panel is the point of the section. A new question is being asked, but nothing new has to be computed — which is why the book's whole treatment fits on a page and a half.
Worked example
For Example 11.8's living arrangements.
\[ 250 \text{ men}, \; 300 \text{ women} \]
Recall what it requires
Why: A claimed distribution.
Check what is available
Why: Only two samples.
Consider using one group as the standard
Why: Women's proportions.
Conclude
Why: The tool does not fit.
Figure (svg): The solution to Worked example why goodness of fit will not do shown as a ladder of expressions, one row per legal move
\[ \text{no given } p_i \;\Longrightarrow\; \text{goodness of fit unavailable} \]
Verify: confirm why treating one sample as the standard would be wrong
Why: The women's proportions carry sampling error of their own, so comparing the men against them would ignore half the uncertainty in the comparison and produce a p-value that is too small. The homogeneity test accounts for variability in both samples at once, which is exactly what pooling the margins accomplishes — the same idea as section 10.4's pooled proportion.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 578
Sorting
Each describes a situation.
Sort into buckets
Sort by which chi-square test fits.
Items (b) and (c) are homogeneity — two separately drawn samples — while (d) is independence, since one sample of 400 was classified two ways. Telling those two apart is section 11.5's subject.
Worked example
Voter preferences among three candidates, surveyed before and after an earthquake.
\[ \text{before against after} \]
Count the populations
Why: Two surveys.
Count the response values
Why: Three candidates.
Check for a claimed distribution
Why: None given.
Name the test
Why: All three signals agree.
Figure (svg): The solution to Worked example Example 11.9's question shown as a ladder of expressions, one row per legal move
\[ H_0: \text{preferences were the same before and after} \]
Verify: confirm this is one of the book's own common uses
Why: The book lists before against after as a standard application, alongside men against women and east against west. All three share the same shape: two groups that are compared with each other rather than against any outside standard, which is the situation the test was built for.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581
Trap
\[ E_{\text{men}} = 250 \cdot \frac{91}{300}, \; 250 \cdot \frac{86}{300}, \; \ldots \]
Treat the women's observed proportions as a known distribution
Why: It looks like a goodness-of-fit test then.
\[ \text{but those proportions are estimates, not facts} \]
The women's sample has its own sampling error, and this treats it as exact — so the resulting p-value is too small.
\[ E = \frac{(\text{row total})(\text{column total})}{n}, \quad n = 550 \]
Pool both samples to build the expected counts
Why: The margins use all 550 people.
This mirrors section 10.4 exactly. There, testing two proportions used a POOLED estimate under the null rather than either sample's own value, for the same reason: if the null says the two agree, the best estimate of what they agree on uses both samples. Here the column totals do that pooling automatically.
Two truths and a lie
All three concern what this test is for.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and it reverses the section's opening point. The distribution is unknown; the test asks only whether the two populations share it, and it never estimates what it is.
Prediction
Commit before reasoning.
Predict first
Two populations compared on a categorical variable with only two response values. What applies?
Correct: A two-proportion test.
Why: With two categories each population is described by one proportion, so section 10.4's test answers the question directly — and gives a direction, which a chi-square never does. The book's own note says the variable here is categorical with more than two possible response values.
Faded example
Example 11.9's, in the book's wording.
Fill in the blanks
H_0: \textsame not the same; \; H_a: \text___ ___
Why: Both are statements about whole distributions rather than about any parameter, which is why they are written as sentences — the same practice as in sections 11.2 and 11.3.
Section
Section 2
Concept
Use a chi-square test statistic, computed in the same way as the test for independence. The expected count in each cell is its row total times its column total over the grand total, and every value in the table must be at least five.
what carries over — The expected-count formula, the statistic, the right tail, and the condition. All four come from section 11.3 without modification.
\[ E = \frac{(\text{row})(\text{column})}{n}, \qquad \chi^2 = \sum \frac{(O-E)^2}{E} \]
There is a reason the arithmetic can be reused. If the two populations share a distribution, then knowing which population a person came from tells nothing about which category they fall in — which is precisely independence between the row variable and the column variable. So the null hypotheses of the two tests, differently worded, impose the same numerical structure on the table.
Figure (svg): The observed table of living arrangements for men and women college students
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579 — follow the same procedure as with the test of independence
Picture it
Living arrangements for 250 men and 300 women.
Figure (svg): The observed table of living arrangements for men and women college students
The two row totals were fixed by the researcher before any data were collected, which is the design signature of a homogeneity test. The column totals emerged from the data, and the expected counts use both.
Worked example
Men in dormitories, from Example 11.8.
\[ \text{row } 250, \; \text{column } 163, \; n = 550 \]
Total the two samples
Why: 250 plus 300.
\[ 550 \]
Total the dormitory column
Why: 72 plus 91.
\[ 163 \]
Apply the formula
Why: 250 times 163 over 550.
\[ 74.09 \]
Compare with observed
Why: 72 were seen.
Figure (svg): The solution to Worked example one expected count shown as a ladder of expressions, one row per legal move
\[ E = \frac{(250)(163)}{550} = 74.09 \]
Verify: confirm the reading of that number
Why: It says that if living arrangements were distributed identically in both groups, the pooled dormitory rate of 163 out of 550 — about 29.6 percent — should apply to the men too, giving 29.6 percent of 250, or 74.09. The pooled rate is the null hypothesis's best estimate of the shared distribution, exactly as section 10.4's pooled proportion was.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 579
Faded example
Women with parents: row 300, column 137, total 550.
Fill in the blanks
E = \frac74.7388 = ___, \quad \text___ ___ \text___
Why: Women lived with parents considerably more than the pooled rate predicts — 88 against 74.73 — and this cell contributes 2.36 to the statistic of 10.13, the largest single share.
Worked example
Do men and women college students have the same distribution of living arrangements? Test at 5 percent.
\[ \alpha = 0.05 \]
Hypotheses
Why: Same against not the same.
Degrees of freedom
Why: Four columns minus one.
\[ 3 \]
Check the condition
Why: Smallest expected.
\[ 36.36,\text{ above five} \]
Statistic
Why: Summed over eight cells.
\[ 10.1287 \]
Right tail on df 3
Why: Beyond 10.1287.
\[ 0.0175 \]
Figure (svg): The solution to Worked example Example 11.8 in full shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 10.1287, \quad \text{df} = 3, \quad p = 0.0175 < 0.05 \]
Verify: confirm the statistic against its degrees of freedom
Why: The mean of a chi-square with three degrees of freedom is 3, and the standard deviation is the root of six, about 2.45 — so 10.13 sits nearly three standard deviations above the centre, consistent with a p-value under 2 percent. Note the decision would reverse at 1 percent, since 0.0175 exceeds 0.01, which is why the book's assumption of 5 percent when none is given matters to the outcome.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-580
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; \text{the expected counts come from the pooled margins} \\ &(2)\; \text{the statistic differs from the independence statistic} \\ &(3)\; \text{the test is right-tailed} \\ &(4)\; \text{the condition is at least five per cell} \end{aligned} \)
The only false claim is (2), and noticing that it is false is most of the lesson. Everything computational transfers, which is why the section is so short and why section 11.5 is needed to keep the tests distinct.
Two truths and a lie
All three concern what carries over from section 11.3.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The column totals pool both samples, which is what makes the expected counts reflect the shared distribution the null asserts. Using one sample alone would be the trap this lesson warns about.
Prediction
Commit before reasoning.
Predict first
Why does a homogeneity test compute exactly what an independence test computes?
Correct: Shared distributions make group and category independent.
Why: The two nulls are worded differently but impose the same structure: in both cases the cell probabilities factor into a row part and a column part. That is why one set of expected counts serves both, and why only the design and the question distinguish the tests.
Estimation
Example 11.8: 163 of 550 lived in dormitories.
Predict first
Roughly what share, and what does it predict for the 300 women?
Correct: About 30 percent, so about 89.
Why: The pooled dormitory rate is 29.6 percent, and applying it to 300 women gives 88.91 — the book's expected count. Reading expected counts as a pooled rate applied to each sample is often faster than the formula and makes the null hypothesis's content visible.
Section
Section 3
Concept
The book gives the degrees of freedom as the number of columns minus one. Both of its examples compare exactly two populations, so the table has two rows — and the general product formula reduces to exactly that.
the two-row case — With r equal to 2, rows minus one is 1, so the product collapses to columns minus one. The rule is section 11.3's, not a third thing to memorise.
\[ (r-1)(c-1) \;\overset{r=2}{\longrightarrow}\; (1)(c-1) = c - 1 \]
The reduction matters when more than two populations are compared, which is a common extension. Three regions across four categories give two times three, which is six — not three. Reading the book's rule as a special case rather than as the definition keeps that generalisation available and avoids an error that would otherwise be invisible.
Figure (svg): The book's own summary of the test for homogeneity
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581 — df = number of columns minus one, in both examples
Picture it
Five headings: hypotheses, statistic, degrees of freedom, requirements, common uses.
Figure (svg): The book's own summary of the test for homogeneity
Every computational line points back to section 11.3. What the summary adds is the wording of the hypotheses and the description of when the test is used — which is genuinely the whole of what is new.
Worked example
Checking the rule against the product formula.
\[ 2 \times 4 \text{ and } 2 \times 3 \]
Example 11.8
Why: Four columns minus one.
\[ 3 \]
By the product rule
Why: One times three.
\[ 3 \]
Example 11.9
Why: Three columns minus one.
\[ 2 \]
By the product rule
Why: One times two.
\[ 2 \]
Figure (svg): The solution to Worked example both of the book's examples shown as a ladder of expressions, one row per legal move
\[ (2-1)(4-1) = 3, \qquad (2-1)(3-1) = 2 \]
Verify: confirm the rules diverge as soon as there are more than two populations
Why: Comparing three regions across four categories gives columns minus one equal to 3 by the book's shortcut, but the product rule gives two times three, which is 6. The product rule is the correct one; the shortcut is safe only for the two-population case the book's examples use, and applying it more widely would halve the degrees of freedom.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-581
Faded example
Three regions compared across six categories.
Fill in the blanks
\text2 = (5)(___) = 10
Why: Ten, not five. The book's columns-minus-one shortcut applies only when there are exactly two populations, which is the case in both of its examples.
Worked example
Voter preferences among three candidates, before and after an earthquake, at 5 percent.
\[ \alpha = 0.05 \]
Hypotheses
Why: Same against not the same.
Degrees of freedom
Why: Three columns minus one.
\[ 2 \]
Statistic
Why: Summed over six cells.
\[ 3.2603 \]
Compare with the mean
Why: Mean is 2.
Right tail
Why: Beyond 3.2603 on df 2.
\[ 0.1959 \]
Figure (svg): The solution to Worked example Example 11.9 in full shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 3.2603, \quad p = 0.1959 > 0.05 \]
Verify: confirm the conclusion is stated with the book's care
Why: The book's wording is that there is insufficient evidence to conclude that the distribution was not the same — not that it was the same. That distinction matters here because the after-survey had 636 respondents against 430 before, and the observed shares differ by a couple of percentage points; the test simply cannot separate that from sampling variation.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 580-581
Trap
\[ 3 \text{ regions} \times 5 \text{ categories} \;\Rightarrow\; \text{df} = 4 \]
Use the book's stated rule directly
Why: It says columns minus one.
\[ \text{but that rule assumed two rows} \]
With three populations the table has three rows, and rows minus one is 2 rather than 1.
\[ \text{df} = (3-1)(5-1) = 8 \]
Use the product rule, which the shortcut is a case of
Why: It holds for any number of populations.
The error halves the degrees of freedom, which shifts the reference curve well to the left and makes every statistic look more extreme than it is. Since a chi-square with 4 degrees of freedom has mean 4 while one with 8 has mean 8, a statistic near 10 would seem significant under the wrong count and unremarkable under the right one.
Two truths and a lie
All three concern the count.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Three populations across four categories give two times three, which is six rather than three, and using the shortcut there would halve the degrees of freedom.
Matching
Match each comparison to its count.
Match the pairs
Why: The first two are the book's own examples. The last two show how quickly the count grows once more than two populations are compared, and why the product rule has to be the one held on to.
Prediction
Commit before reasoning.
Predict first
The rule columns minus one is correct under what condition?
Correct: Exactly two populations.
Why: With two rows, rows minus one is 1 and the product formula collapses. Both of the book's examples compare two groups, so the shortcut is correct throughout section 11.4 — but it is a consequence of that design rather than a general rule.
Section
Section 4
Concept
The book adds an explicit warning after Example 11.8: notice that the conclusion is only that the distributions are not the same. The test for homogeneity cannot be used to draw any conclusions about how they differ.
an omnibus test — One that detects any departure without identifying its kind. The price of a single verdict on a whole distribution is that the verdict carries no direction.
\[ \text{reject } H_0 \;\Longrightarrow\; D_1 \ne D_2 \text{ only} \]
The contrast with section 10.4 is sharp. A two-proportion test rejects and reports which proportion is larger, because its statistic keeps the sign of the difference. A chi-square squares every departure, so the sign is destroyed before the cells are added — and no amount of examining the statistic afterwards can recover it.
Figure (svg): A chi-square curve on three degrees of freedom with the right tail beyond 10.13 shaded thin
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 580 — we cannot use the test for homogeneity to draw conclusions about how they differ
Picture it
The statistic against its distribution.
Figure (svg): A chi-square curve on three degrees of freedom with the right tail beyond 10.13 shaded thin
The p-value of 0.0175 supports one claim only: the two distributions differ. Which categories differ, and in which direction, is not something the test reports — though the cells can of course be described.
Worked example
After rejecting in Example 11.8.
\[ p = 0.0175 \]
The test's verdict
Why: The alternative.
Adding a direction
Why: Not from this test.
Naming a category
Why: Not from this test.
Describing the cells
Why: A description.
Figure (svg): The solution to Worked example what may and may not be said shown as a ladder of expressions, one row per legal move
\[ \text{sufficient evidence that } D_{\text{men}} \ne D_{\text{women}} \]
Verify: confirm what the cells show, as description rather than inference
Why: Women lived with parents more than the pooled rate predicts, at 88 against 74.73, and men chose the other category more, at 45 against 36.36. Both are accurate descriptions of these 550 students and worth reporting — but each is a claim the test did not separately examine, and the book's warning is aimed exactly at presenting them as though it had.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 579-580
Two truths and a lie
All three concern the conclusion.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is exactly what the book warns against. A chi-square gives one unsigned verdict; recovering directions requires either separate comparisons or a test designed to give them.
Worked example
If living arrangements had only two categories.
\[ \text{dormitory or not} \]
The chi-square
Why: One verdict.
The two-proportion z
Why: Keeps the sign.
Its confidence interval
Why: Section 8.4's.
Note the trade
Why: Fewer categories, more detail.
Figure (svg): The solution to Worked example what a two-proportion test would add shown as a ladder of expressions, one row per legal move
\[ z \text{ signed} \;\text{against}\; \chi^2 \ge 0 \text{ unsigned} \]
Verify: confirm this is a genuine trade rather than a defect
Why: A chi-square handles four categories at once, which no two-proportion test can. What it gives up is the direction, because summing squared departures across categories leaves nothing to attach a sign to. The right response is to choose the test that matches the question — an omnibus verdict across many categories, or a directed comparison of one.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 580
Error analysis
The p-value was 0.0175 at a 5 percent level. Which are defensible?
Annotate
On: \( \begin{aligned} &(1)\; \text{the distributions are not the same} \\ &(2)\; \text{women live with parents more often than men} \\ &(3)\; \text{sex determines living arrangement} \\ &(4)\; \text{at 1 percent, the evidence is insufficient} \end{aligned} \)
Statement (4) is worth noticing because the book assumes 5 percent when no level is given. The evidence here is moderate rather than overwhelming, and the decision genuinely depends on the level chosen in advance.
Faded example
Example 11.8's p-value is 0.0175.
Fill in the blanks
\textreject \alpha = 0.05: \; do not reject; \qquad \text___ \alpha = 0.01: \; ___
Why: The same data, two decisions. This is why the level has to be fixed before seeing the p-value, a point chapter 9 made and one that this example illustrates cleanly.
Prediction
Commit before reasoning.
Predict first
Why can a chi-square test not say which way two distributions differ?
Correct: Squaring destroys the sign.
Why: A cell running 20 above expectation and one running 20 below contribute identically. Once summed, there is nothing left in the statistic to indicate which cells ran which way — which is the price of combining every category into one number.
Explain it
A classmate writes that Example 11.8 shows women are more likely to live with their parents.
Discussion prompt
In two sentences or fewer, tell them what the test supports.
Hint: Ask what the alternative hypothesis actually says.
Answer:
The test's alternative is only that the two distributions are not the same, and the book adds a warning that it cannot be used to draw conclusions about how they differ.
That women lived with parents more often is a true description of these 550 students, and worth reporting as such — but it is not a claim the test separately examined.
Section
Section 5
Concept
The two tests compute identical numbers from a two-way table. What separates them is what was fixed before the data were collected: one sample classified two ways gives a test of independence, while two separately drawn samples classified one way gives a test for homogeneity.
fixed margins — In a homogeneity study the row totals are chosen by the researcher — 250 men, 300 women. In an independence study only the grand total is, and both sets of margins emerge from the data.
\[ \text{one sample, two questions} \quad\text{against}\quad \text{two samples, one question} \]
This is why the tests can share a procedure without being the same test. The arithmetic examines whether the interior of the table matches its margins; what that comparison MEANS depends on where the margins came from. Section 11.5 turns this into a practical rule for choosing between them, based on how the hypotheses are worded.
Figure (svg): A card contrasting one sample cross-classified against two samples separately drawn
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579 — common uses: comparing two populations
Picture it
What was fixed in advance, in each case.
Figure (svg): A card contrasting one sample cross-classified against two samples separately drawn
A table of numbers alone cannot say which study produced it. That is why the choice between the two tests is made from the description of how the data were collected, before any arithmetic begins.
Worked example
Example 11.6 against Example 11.8.
\[ \text{which is which?} \]
Example 11.6
Why: 839 volunteers, one sample.
What was fixed
Why: Only the total.
Example 11.8
Why: 250 men and 300 women.
What was fixed
Why: Both row totals.
Figure (svg): The solution to Worked example reading two designs shown as a ladder of expressions, one row per legal move
\[ 839 \text{ classified twice} \quad\text{against}\quad 250 + 300 \text{ classified once} \]
Verify: confirm the arithmetic really is indistinguishable
Why: Both tests take the same table, build the same expected counts from row times column over n, sum the same statistic and read the same right tail. If the numbers of Example 11.8 had arisen from a single sample of 550 people classified by sex and living arrangement, every calculation would be identical and only the wording of the hypotheses and conclusion would change.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-579
Sorting
Each describes how data were collected.
Sort into buckets
Sort by which test the design calls for.
Items (a) and (d) are Examples 11.6 and 11.7; (b), (c) and (e) are Examples 11.8, 11.9 and Try It 11.8. The signal every time is how many samples were drawn, not what the table looks like.
Worked example
Round row totals often signal a homogeneity design.
\[ 250 \text{ and } 300 \]
Note the roundness
Why: Both exact hundreds or halves.
Compare Example 11.6
Why: 255, 290, 294.
Draw the inference
Why: A researcher picked the first.
Note the caution
Why: Not conclusive.
Figure (svg): The solution to Worked example the sample sizes as a tell shown as a ladder of expressions, one row per legal move
\[ 250, 300 \text{ chosen} \quad\text{against}\quad 255, 290, 294 \text{ observed} \]
Verify: confirm this is a hint rather than a rule
Why: Round numbers can arise by chance, and a homogeneity study might well end up with 247 men after non-responses. The reliable signal is always the description of how the data were collected — the book's examples say plainly that 250 men and 300 women were randomly selected, and that sentence is what settles it.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, p. 579
Trap
\[ \text{a } 2 \times 4 \text{ table} \;\Rightarrow\; \text{a test for homogeneity} \]
Read the design off the table's shape
Why: Homogeneity examples all have two rows.
\[ \text{but a } 2 \times 4 \text{ independence table is equally possible} \]
One sample of 550 people classified by sex and living arrangement would produce the same shape.
\[ \text{read how the data were collected, then choose} \]
Ask what was fixed before sampling began
Why: Two samples means homogeneity; one means independence.
Since the arithmetic is identical, choosing wrongly costs nothing numerically — the statistic, degrees of freedom and p-value all come out the same. What it costs is the interpretation: a homogeneity conclusion about two populations is a different claim from an independence conclusion about two factors, and only one of them fits the study that was run.
Two truths and a lie
All three concern the two designs.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The same table could arise from either design, so the choice depends entirely on the description of how the data were collected — which is why the book introduces each example with a sentence about the sampling.
Prediction
Commit before reasoning.
Predict first
A homogeneity study is analysed as a test of independence. What changes?
Correct: Only the interpretation.
Why: Every number comes out the same, since the procedures are identical. What differs is the claim being made: independence between two factors within one population, against agreement between two populations. Reporting the wrong one misdescribes the study even though the arithmetic is sound.
Estimation
Try It 11.8 surveys 100 families and 200 singles about car type.
Predict first
Which test, and how many degrees of freedom across five car types?
Correct: Homogeneity, with 4 degrees of freedom.
Why: Two separately drawn samples of chosen sizes make it a homogeneity test, and five categories give columns minus one, which is four — the same as the product rule's one times four. The statistic there is 62.91, far beyond the mean of 4, so the distributions differ decisively.
Comparison
Fill the blanks. Two of the rows are identical across all three.
Comparison matrix
| Goodness of fit | Independence | Homogeneity | |
|---|---|---|---|
| Samples drawn | one | one | two or more |
| Expected counts from | an outside claim | the margins | the margins |
| Degrees of freedom | cells minus one | (r-1)(c-1) | columns minus one, for two groups |
| Tail | right | right | right |
The last two columns agree on everything computational, which is the whole point of this section. What separates them is the design behind the table and the wording of the hypotheses, and section 11.5 makes that separation into a rule.
Pattern
Six steps, and five of them are section 11.3's.
Read expected counts as the pooled rate applied to each sample; it is faster than the formula and makes the null's content visible.
OpenStax Introductory Business Statistics 2e, §11.5 Test for Homogeneity §11.5 Test for Homogeneity
Check
Choosing the test.
Check your understanding
Two hundred urban and 200 rural residents each name their preferred transport from five options. Which test applies?
Answer: A
Why: Two separately drawn samples, each classified by one categorical variable with more than two values, is exactly the homogeneity design.
Check
Degrees of freedom.
Check your understanding
Three regions are compared across four categories. What are the degrees of freedom?
Answer: A
Why: Two times three is six. The book's columns-minus-one rule is the two-population case of this, and does not apply with three groups.
Check
The conclusion.
Check your understanding
A test for homogeneity rejects at 5 percent. What may be concluded?
Answer: A
Why: The book is explicit that the conclusion is only that the distributions are not the same, and that no conclusion about how they differ may be drawn.
Real world
A company surveys 400 employees before a policy change and 400 after, asking each to pick one of six reasons for job satisfaction. A test for homogeneity gives a p-value of 0.31, and the report concludes that the policy change had no effect on why employees are satisfied.
Discussion prompt
Assess that conclusion, and say what else would need checking.
Hint: Ask what failing to reject establishes, and what the test could have detected.
Answer:
The conclusion overstates a non-rejection, in exactly the way chapter 9 warned against. A p-value of 0.31 says the data are consistent with no change; it does not establish that no change occurred. The honest wording is that there is insufficient evidence that the distribution of reasons differed before and after.
The power question is the substantive one. With 400 per group and six categories, expected counts run around 130 per cell if the categories are roughly balanced — enough to detect a shift of several percentage points, but not a shift of one or two. Before treating the non-result as informative, someone should ask how large a change the study could have found, and whether a change smaller than that would still have mattered to the company.
\[ \text{df} = (2-1)(6-1) = 5, \qquad \chi^2 \text{ needed to reject at } 5\% \approx 11.07 \]
Two design points deserve attention too. If the same employees were surveyed twice, the two samples are not independent, and a chi-square test for homogeneity assumes they are — paired categorical data need a different treatment, the same way section 10.5 needed a paired t rather than a two-sample one. And if the surveys drew different employees, turnover between them could change the population itself, so the two samples might differ in composition rather than in attitude.
What the study supports: no detectable change in the distribution of reasons, at this sample size. What it does not support: that the policy had no effect. The difference between those two statements is the whole content of chapter 9's asymmetry between rejecting and failing to reject, and it matters most exactly when the result is the one the reporter was hoping for.
Commit first
Answer, then rate your confidence honestly.
Predict first
What distinguishes a test for homogeneity from a test of independence?
Correct: The design and the wording; the arithmetic is identical.
\[ E = \frac{RC}{n}, \; \chi^2 = \sum \frac{(O-E)^2}{E}, \; \text{df} = (r-1)(c-1) \quad \text{in both} \]
Why: The book says to follow the same procedure as with the test of independence, and it means it — the same expected counts, the same statistic, the same right tail, the same condition. A homogeneity study draws two samples of chosen sizes and asks whether two populations agree; an independence study draws one sample and asks whether two factors are related. Nothing in the table itself distinguishes them.
Explain it
They used columns minus one for a table comparing four regions across three categories, getting df = 2.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask how many rows the book's examples had.
Answer:
Columns minus one is the book's shortcut for comparing exactly two populations, where rows minus one happens to equal 1 — but with four regions there are four rows, so rows minus one is 3.
The degrees of freedom are three times two, which is six rather than two.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, whether two populations share an unknown distribution. For the second, count the samples drawn, not the rows in the table. For the third, use the product rule, of which columns minus one is a special case. For the fourth, the distributions differ, never how. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Twelve minutes.
Draw it
At the top, write the question a goodness-of-fit test cannot answer — whether two populations follow the same unknown distribution — and beside it the hypotheses for a test of homogeneity in words. Underneath, write in capitals that the arithmetic is the test of independence's, unchanged, and list the four things that carry over: the expected-count formula, the statistic, the right tail, and the at-least-five condition. In the middle, work Example 11.8: the two-by-four table of 72, 84, 49, 45 over 91, 86, 88, 35 with row totals 250 and 300 and column totals 163, 170, 137, 80; the expected count for men in dormitories as 250 times 163 over 550, which is 74.09; the degrees of freedom as three; the statistic 10.1287 and p-value 0.0175. Beside it write the two decisions — reject at 5 percent, do not reject at 1 percent. At the bottom left, write the degrees-of-freedom rule as a product and show columns minus one falling out of it when there are two rows. At the bottom right, write the book's warning in a box: the conclusion is only that the distributions are not the same, never how they differ.
Check your expected count by reading it the other way: 163 of 550 is 29.6 percent, and 29.6 percent of 250 is 74.09. Check your degrees-of-freedom box by testing it on three populations across four categories, where the product rule gives 6 and the shortcut would wrongly give 3.
Recap
Six things, and only two of them are new since section 11.3.
| If you see | Then |
|---|---|
| Two separately drawn samples | A test for homogeneity |
| One sample classified two ways | A test of independence instead |
| An external distribution quoted | A goodness-of-fit test instead |
| Two populations, two categories | A two-proportion test from section 10.4, which gives a direction |
| Expected counts needed | Row times column over n, pooling every sample |
| Two populations compared | Degrees of freedom are columns minus one |
| Three or more populations | Use the product rule; the shortcut no longer holds |
| A rejection | The distributions differ; describe the cells separately if at all |
Section 11.5 is the book's own comparison of the three tests. Since so little separates them computationally, it turns the distinction into a practical rule based on how the hypotheses are worded — which is the reliable signal when a problem does not describe its sampling clearly.
OpenStax Introductory Statistics 2e, §11.4 Test for Homogeneity §11.4, pp. 578-581 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.