11.5 Comparison of the Chi-Square Tests

A summary section, and a necessary one. The chi-square statistic has now been used in three different circumstances, and since all three share the statistic, the right tail and the at-least-five condition, nothing in the arithmetic distinguishes them. The book's list separates them by the study instead, and two counts settle nearly every case: how many populations were sampled, and how many questions were asked. A goodness-of-fit test has one population, one question, and a known distribution to test against; a test of independence has one population and two questions, arranged in a contingency table; a test for homogeneity has two populations, one question, and no known distribution at all. The surest signal of all is how the hypotheses are worded, and this lesson works through the book's three pairs.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 11.5 Comparison of the Chi-Square Tests

Title

Statistics · Chapter 11 — The Chi-Square Distribution

Comparison of the Chi-Square Tests

2. By the end of this lesson you can

Objectives

Five outcomes. None involves a calculation.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the section these objectives are drawn from

3. What you already have

Warm-up

Three tests, three sections, one statistic.

Discussion prompt

Given only a two-way table of counts and its chi-square statistic, which of the three tests produced it?

Hint: Ask what the numbers themselves could possibly reveal about the study.

Answer:

There is no way to tell. A two-way table could come from one sample cross-classified — a test of independence — or from two separately drawn samples each classified once, which is a test for homogeneity. The expected counts, the statistic, the degrees of freedom and the p-value are identical either way.

Even a single row of counts is ambiguous until a claimed distribution appears: with one, it is goodness of fit; without one, there is nothing to test against at all.

So the choice of test is made before the arithmetic, from the description of how the data were collected and what is being asked. This section is the book's own procedure for making it.

4. Three circumstances, one statistic

Concept

You have seen the chi-square test statistic used in three different circumstances. The tests are told apart by how many populations were sampled, how many questions were asked, and whether a distribution is known in advance — not by anything in the calculation.

choosing the test — A goodness-of-fit test has one population and one question with a known distribution; independence has one population and two questions; homogeneity has two populations and one question.

\[ \text{gof}: 1\text{ pop}, 1\text{ q} \quad \text{ind}: 1\text{ pop}, 2\text{ q} \quad \text{hom}: 2\text{ pop}, 1\text{ q} \]

No two rows of that grid agree on both counts, which is what makes the pair decisive. Goodness of fit and independence share a population count but differ on questions; goodness of fit and homogeneity share a question count but differ on populations; independence and homogeneity differ on both. Two numbers therefore identify the test uniquely.

Figure (svg): A table distinguishing the three chi-square tests by populations, questions and whether a distribution is known

The book's three-bullet summary, arranged so the distinguishing features line up.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

5. Counting populations and questions

Section

Section 1

6. Two numbers identify the test

Concept

A goodness-of-fit test has a single qualitative survey question or a single outcome of an experiment from a single population. A test of independence has two qualitative survey questions or experiments, arranged in a contingency table. A test for homogeneity has a single question given to two different populations.

a question — One categorical variable measured on each subject. Two questions means each subject is classified twice, which is what produces a contingency table from a single sample.

\[ (\text{populations}, \text{questions}): \; (1,1), \; (1,2), \; (2,1) \]

The counts are read from the study's description rather than from the table. Example 11.6 asked 839 volunteers two things — what type of volunteer they were and how many hours they gave — which is one population and two questions. Example 11.8 asked one thing, about living arrangements, of two populations drawn separately. Both produced a two-way table, and only the descriptions distinguish them.

Figure (svg): A table distinguishing the three chi-square tests by populations, questions and whether a distribution is known

The book's three-bullet summary, arranged so the distinguishing features line up.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the three bullets

7. The three tests, lined up

Picture it

Populations, questions, and whether a distribution is given.

Figure (svg): A table distinguishing the three chi-square tests by populations, questions and whether a distribution is known

The book's three-bullet summary, arranged so the distinguishing features line up.

The third column resolves the remaining ambiguity. When one population is compared against another whose distribution is KNOWN, it is a goodness-of-fit test; when the other distribution has to be estimated from a sample, it is homogeneity.

8. Worked example: routing four studies

Worked example

Counting populations and questions for each.

\[ \text{which test?} \]

Absences across five weekdays

Why: One sample, one question.

Volunteer type by hours

Why: One sample, two questions.

Men against women on living

Why: Two samples, one question.

Voters before and after

Why: Two samples, one question.

Figure (svg): The solution to Worked example routing four studies shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (1,1), \; (1,2), \; (2,1), \; (2,1) \]

Verify: confirm the first one really has a known distribution

Why: Example 11.2 tested whether absences occur with equal frequencies, and a uniform distribution over five weekdays is completely specified — one fifth each — without estimating anything. That is what makes it a known distribution and the test a goodness-of-fit test. Had the question been whether two workplaces had the same pattern of absences, nothing would be known and it would be homogeneity.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

9. How many populations?

Sorting

Each describes a study.

Sort into buckets

Sort by how many populations were sampled.

One population
400 students classified by anxiety and by need to succeed; 600 far-western families reporting streaming services; 839 volunteers classified by type and by hours
Two populations
250 men and 300 women asked about living arrangements; 100 families and 200 singles asked what car they drive
one
A single sample was drawn, then classified.
two
Two separate samples were drawn, of sizes the researcher chose.

Items (a) and (e) each produce a table with three or more rows from ONE population, which is exactly why rows cannot be counted as populations.

10. Worked example: the ambiguous case

Worked example

Comparing one population against another population's distribution.

\[ \text{one sample against another population} \]

If the other is known

Why: Percentages published.

Build expected counts

Why: n times each percentage.

If the other is a sample

Why: Its distribution estimated.

Build expected counts

Why: Pooled margins.

Figure (svg): The solution to Worked example the ambiguous case shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{known } p_i \;\to\; \text{gof}; \qquad \text{estimated} \;\to\; \text{hom} \]

Verify: confirm why the distinction changes the arithmetic

Why: A known distribution contributes no uncertainty, so the expected counts are exact and the degrees of freedom are cells minus one. An estimated one carries sampling error, which the homogeneity test accounts for by pooling both samples into the margins — and by spending degrees of freedom on the estimation, which is why its count is a product rather than a simple subtraction.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

11. Trap: counting rows instead of populations

Trap

The trap

\[ \text{a table with 3 rows} \;\Rightarrow\; \text{three populations} \;\Rightarrow\; \text{homogeneity} \]

Read the population count off the table

Why: Homogeneity examples do have one row per population.

\[ \text{but Example 11.6's three rows are one sample} \]

Those 839 volunteers were a single sample, and the rows are levels of a variable rather than separate populations.

The fix

\[ \text{read the sampling description, then count} \]

Ask how many separate samples were drawn

Why: Rows are a display choice; samples are a design fact.

The tell in the text is usually a sentence naming two sample sizes — 250 men and 300 women were randomly selected — against one naming a single total, as in a random sample of 400 students took a test. The book's examples always say which, and that sentence is what the choice rests on.

12. The two counts

Faded example

A study asks 500 shoppers one question about preferred payment method, and compares the result against published national percentages.

Fill in the blanks

\text1 = 1, \quad \text___ = ___ \;\Rightarrow\; \text___

Why: One and one, with a known distribution supplied by the national figures — the goodness-of-fit signature exactly.

13. One of these is false

Two truths and a lie

All three concern the counts.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. No two tests share both counts
  • C. The counts come from the study description
  • B. The number of table rows gives the number of populations

Survives elimination: B

Why: The survivor is false. Example 11.6's table has three rows from one sample of 839 volunteers, and Example 11.7's has three rows from one sample of 400 students. Rows are levels of a variable unless the description says separate samples were drawn.

14. Two questions of one sample

Prediction

Commit before reasoning.

Predict first

A single sample is asked two categorical questions. Which test?

  • A test of independence
  • A goodness-of-fit test
  • A test for homogeneity
  • Any of the three

Correct: A test of independence.

Why: One population and two questions is that test's signature, and the two questions are what produce a contingency table. The goal, in the book's words, is to see whether the two variables are unrelated or related.

15. The wording of the hypotheses

Section

Section 2

16. The surest signal

Concept

Each test has its own null and alternative wording. Goodness of fit asks whether the population FITS the given distribution; independence asks whether two variables are INDEPENDENT; homogeneity asks whether two populations FOLLOW THE SAME distribution.

reading the null — A null naming a given distribution means goodness of fit; one naming two variables or factors means independence; one naming two populations means homogeneity.

\[ \text{fits} \quad \text{against} \quad \text{independent} \quad \text{against} \quad \text{the same as each other} \]

This is the most reliable route when a problem's description of its sampling is thin, which happens often in practice. If the null can be written at all, its subject identifies the test: a distribution supplied from outside, a pair of variables, or a pair of populations. Those three subjects are mutually exclusive.

Figure (svg): Three pairs of hypotheses, one for each chi-square test

The book's own wordings. A null about fitting names a given distribution; one about factors names two variables; one about following names two populations.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the three hypothesis pairs

17. The book's three pairs

Picture it

Reproduced exactly as stated.

Figure (svg): Three pairs of hypotheses, one for each chi-square test

The book's own wordings. A null about fitting names a given distribution; one about factors names two variables; one about following names two populations.

Notice that every alternative is simply the negation of its null, with no direction anywhere. That is consistent with all three tests being right-tailed: there is no direction for the hypotheses to express, because the statistic could not detect one.

18. Worked example: identifying tests from nulls alone

Worked example

Three hypotheses, with no other information.

\[ \text{which test?} \]

Blood types fit the national distribution

Why: A given distribution.

Smoking and exercise are independent

Why: Two variables.

Urban and rural follow the same distribution

Why: Two populations.

Note the pattern

Why: The subject decides.

Figure (svg): The solution to Worked example identifying tests from nulls alone shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{distribution} \;\to\; \text{gof}; \; \text{variables} \;\to\; \text{ind}; \; \text{populations} \;\to\; \text{hom} \]

Verify: confirm the three subjects really cannot overlap

Why: A given distribution comes from outside the study; two variables are two measurements on one sample; two populations are two separate samples. A single study cannot be described by more than one of those, so the subject of a correctly written null always identifies the test uniquely — and if it seems to fit two, the null has been written loosely.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

19. Null to test

Matching

Match each null hypothesis to its test.

Match the pairs

  • l1. the population fits the given distribution
  • l2. the two variables are independent
  • l3. the two populations follow the same distribution
  • l4. the population variance equals 25
  • r1. goodness of fit
  • r2. test of independence
  • r3. test for homogeneity
  • r4. test of a single variance

Why: The fourth is section 11.6's, and it is the only chi-square test in the chapter whose null names a parameter — which is also why it is the only one that can be left-tailed.

20. Worked example: rewriting a vague null

Worked example

Someone writes the null as the two groups are the same.

\[ H_0: \text{the two groups are the same} \]

Ask what same means

Why: Same on what?

Ask what the groups are

Why: Populations or levels?

If separate samples

Why: Two populations.

If one sample split

Why: Two levels of a variable.

Figure (svg): The solution to Worked example rewriting a vague null shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \text{same distribution} \quad\text{or}\quad H_0: \text{independent} \]

Verify: confirm the numbers would be identical either way

Why: Both readings produce the same expected counts, the same statistic and the same p-value, so nothing computational forces a choice. That is precisely why the wording matters: it is the only place where the difference between the two studies is recorded, and a vague null loses it.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

21. Error analysis: four hypothesis statements

Error analysis

Which are correctly worded for their tests?

Annotate

On: \( \begin{aligned} &(1)\; H_0: \text{the population fits the given distribution} \\ &(2)\; H_0: \mu_1 = \mu_2 \\ &(3)\; H_0: \text{the two factors are dependent} \\ &(4)\; H_0: \text{the two populations follow the same distribution} \end{aligned} \)

  • (1) is the book's goodness-of-fit null, correctly worded.
  • (2) is a two-sample means null from section 10.1, and belongs to no chi-square test — these tests have no parameter.
  • (3) has the hypotheses swapped. The null asserts independence; dependence is the alternative.
  • (4) is the book's homogeneity null, correctly worded.

Error (3) is worth guarding against because it reverses the entire logic of the test. Independence is the null because it is the statement that generates the expected counts; dependence supplies no numbers at all, so it could not be the hypothesis being tested against.

22. One of these is false

Two truths and a lie

All three concern the hypotheses.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Each alternative is the plain negation of its null
  • C. None of the three nulls names a parameter
  • B. The alternative for independence states which factor causes the other

Survives elimination: B

Why: The survivor is false. The alternative states only that the two factors are dependent. Causation is not something any of these tests addresses, and with observational data it would not follow from any p-value.

23. Complete the pair

Faded example

For a test for homogeneity, in the book's wording.

Fill in the blanks

H_0: \textsame different \text___; \; H_a: \text___ ___ \text___

Why: Both are statements about whole distributions with no direction attached, which is what makes the test right-tailed and its conclusion silent about how the populations differ.

24. Why is independence the null?

Prediction

Commit before reasoning.

Predict first

Why does the null assert independence rather than dependence?

  • Independence supplies the expected counts; dependence supplies no numbers
  • Because independence is more likely
  • Because it is more conservative
  • By convention

Correct: Independence supplies the expected counts.

Why: A hypothesis has to be specific enough to predict what the data should look like, and independence does exactly that through row times column over n. Dependence describes infinitely many possible tables and predicts none of them, so it cannot serve as the hypothesis being tested against — the same reason a null is always an equality.

25. Known against unknown distributions

Section

Section 3

26. The word that separates two of the tests

Concept

Goodness of fit is typically used to see if the population is uniform, if the population is normal, or if the population is the same as another population with a KNOWN distribution. Homogeneity decides whether two populations with UNKNOWN distributions have the same distribution as each other.

a known distribution — One fully specified without reference to the data: a uniform over five categories, a set of published national percentages, a fair-coin distribution. No estimation is involved.

\[ \text{known } D \;\to\; \text{gof}; \qquad \text{both unknown} \;\to\; \text{hom} \]

The overlap is real and the book resolves it explicitly. Both tests can compare a population against another population — the difference is entirely whether that other population's distribution is available as fact or has to be estimated from a sample. Example 11.3 compared far-western families against the national distribution, which was published; had it compared them against a sample of eastern families, it would have been a homogeneity test.

Figure (svg): What a goodness-of-fit test is typically used for

The book's own list. Example 11.2 is the first use, and Example 11.3 the third.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the three typical uses of goodness of fit

27. Three typical uses

Picture it

What a goodness-of-fit test is normally reached for.

Figure (svg): What a goodness-of-fit test is typically used for

The book's own list. Example 11.2 is the first use, and Example 11.3 the third.

The first two uses have no counterpart among the other tests: only goodness of fit can check whether a population is uniform or normal, because only it accepts a distribution specified from outside. The third use is where the overlap lives.

28. Worked example: the same comparison, two ways

Worked example

Far-western families' streaming services, compared two different ways.

\[ \text{against national figures, or against an eastern sample} \]

Against published percentages

Why: Known.

Its degrees of freedom

Why: Five cells minus one.

\[ 4 \]

Against an eastern sample

Why: Unknown, estimated.

Its degrees of freedom

Why: Two rows, five columns.

\[ 4 \]

Figure (svg): The solution to Worked example the same comparison, two ways shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k-1 = 4 \qquad\text{against}\qquad (2-1)(5-1) = 4 \]

Verify: confirm why the two are not interchangeable despite matching here

Why: The expected counts differ substantially. Against published percentages they are 600 times each national figure, fixed and exact. Against an eastern sample they come from pooled margins and reflect both samples' variability. The matching degrees of freedom are an artefact of comparing exactly two groups across five categories, and would diverge with any other shape.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

29. Known or estimated?

Sorting

Each is a distribution a study might compare against.

Sort into buckets

Sort by whether it counts as known.

Known: goodness of fit
a uniform distribution over five weekdays; published national census percentages; the fair-coin distribution for two flips
Estimated: homogeneity
percentages from a separate sample of 300 people; proportions computed from this year's other survey
known
Fully specified without reference to sampled data.
est
Computed from a sample, so it carries sampling error of its own.

Items (a) and (d) come from theory, and (c) from a census that measures the whole population. Items (b) and (e) come from samples, so the comparison has to account for their variability too.

30. Worked example: uniform and normal

Worked example

Two uses only goodness of fit can serve.

\[ \text{uniform or normal?} \]

A uniform claim

Why: Equal frequency.

Build expected counts

Why: n over k.

A normal claim

Why: A shape.

Compare with the others

Why: They need two groups.

Figure (svg): The solution to Worked example uniform and normal shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ E_i = \frac{n}{k} \text{ for uniform} \]

Verify: confirm this explains the chapter's lab

Why: Section 11.7 is the book's own lab, in which students collect thirty grocery receipt totals and test whether they fit either the uniform or the exponential distribution. Both are goodness-of-fit questions about a single population against a specified shape, and the lab's own note warns that two categories may need combining so every expected value reaches five — the condition from section 11.2, appearing in practice.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

31. Trap: treating a sample's distribution as known

Trap

The trap

\[ \text{the eastern sample gave } 21\%, 24\%, 32\%, 14\%, 9\% \]

Use those as the given distribution for a goodness-of-fit test

Why: They are percentages, so they look like a claim.

\[ \text{but they carry sampling error} \]

Percentages computed from a sample are estimates, and treating them as exact ignores half the uncertainty in the comparison.

The fix

\[ \text{pool both samples: } E = \frac{RC}{n} \]

Run a test for homogeneity instead

Why: It accounts for variability in both samples.

The consequence of getting this wrong is a p-value that is too small — the test appears more confident than the data justify, because one sample's noise has been treated as fact. The signal to watch for is where the percentages came from: published figures and theoretical models are known, and anything computed from data is not.

32. One of these is false

Two truths and a lie

All three concern known distributions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Only goodness of fit can test whether a population is uniform
  • C. A census gives a known distribution
  • B. Any set of percentages counts as a known distribution

Survives elimination: B

Why: The survivor is false. Percentages computed from a sample are estimates, and using them as though they were exact understates the uncertainty and produces a p-value that is too small.

33. What if both are samples?

Prediction

Commit before reasoning.

Predict first

Two samples are compared and neither distribution is known. Which test?

  • Homogeneity, which pools both samples for its expected counts
  • Goodness of fit, using one sample as the standard
  • Independence
  • None applies

Correct: Homogeneity.

Why: That is exactly the case the book says a goodness-of-fit test will not suffice for. Pooling both samples into the margins accounts for variability in each, which is the same principle as section 10.4's pooled proportion under a null of equality.

34. How many of the three uses overlap?

Estimation

Goodness of fit's three typical uses: uniform, normal, and same as a known population.

Predict first

How many of them could a homogeneity test also serve?

  • One, and only with an unknown distribution
  • All three
  • None
  • Two

Correct: One, and only with an unknown distribution.

Why: Only the third use involves comparing two populations, and homogeneity applies to it precisely when the second distribution is unknown. Testing for uniformity or normality compares one population against a specified shape, and no two-population test can pose that question at all.

35. What all three share

Section

Section 4

36. Everything computational

Concept

The three tests use the same statistic, the same right tail, the same requirement that every expected count reach five, and hypotheses stated in words rather than as equations. What differs is where the expected counts come from and how the degrees of freedom are counted.

why a comparison section is needed — Because the shared features are exactly the ones a student notices while calculating. Nothing in the arithmetic prompts the question of which test is being run.

\[ \chi^2 = \sum \frac{(O-E)^2}{E} \quad \text{in all three} \]

There is one more shared feature worth naming: none of the three has a parameter. Every test from chapter 9 through chapter 10 concerned a mean, a proportion or a difference of them, and could be reported with a confidence interval. These tests concern whole distributions, so there is nothing to build an interval around — which is why the chapter offers none.

Figure (svg): A card listing what the three tests share and what separates them

Five things are identical across the three tests, which is exactly why a separate comparison section is needed.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the same statistic used in three circumstances

37. Shared and distinct

Picture it

Five things in common, five that differ.

Figure (svg): A card listing what the three tests share and what separates them

Five things are identical across the three tests, which is exactly why a separate comparison section is needed.

The left column is why this section exists. A student working through the arithmetic sees only shared features, so the distinction has to be drawn from the study description before any calculation begins.

38. Worked example: the degrees-of-freedom rules side by side

Worked example

The one computational difference that changes answers.

\[ \text{three rules} \]

Goodness of fit

Why: Cells minus one.

\[ k - 1 \]

Independence

Why: A product.

\[ (r - 1) (c - 1) \]

Homogeneity

Why: The same product.

\[ (r - 1) (c - 1) \]

Note

Why: Two of three agree.

Figure (svg): The solution to Worked example the degrees-of-freedom rules side by side shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k-1 \quad\text{against}\quad (r-1)(c-1) \]

Verify: confirm the shared rule reflects a shared structure

Why: Independence and homogeneity both build expected counts from margins, so both are constrained by every row and column total — hence the same product. Goodness of fit is constrained only by the grand total, hence the simple subtraction. The rules differ exactly where the constraints do, which makes them derivable rather than memorisable.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

39. One of these is false

Two truths and a lie

All three concern what the tests share.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. All three require expected counts of at least five
  • C. All three state hypotheses in words
  • B. All three use the same degrees-of-freedom rule

Survives elimination: B

Why: The survivor is false, and it is the single computational difference. Goodness of fit uses cells minus one, while independence and homogeneity both use the product of the reduced dimensions.

40. Worked example: what changes if the test is misidentified

Worked example

An independence study analysed as homogeneity, or the reverse.

\[ \text{cost of the error} \]

The expected counts

Why: Same formula.

The statistic

Why: Same sum.

The degrees of freedom

Why: Same product.

The conclusion

Why: Different claim.

Figure (svg): The solution to Worked example what changes if the test is misidentified shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{same } \chi^2, \text{ df}, p; \; \text{different conclusion} \]

Verify: confirm the same is not true of confusing either with goodness of fit

Why: Mistaking a contingency table for a goodness-of-fit problem changes the degrees of freedom — nine cells would give 8 rather than 4 — and with Example 11.6's statistic that moves the p-value from 0.0113 to 0.1124, reversing the decision at 5 percent. So the two errors are not equally harmless: one costs only the interpretation, the other costs the answer.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

41. Error analysis: four claims about what the tests share

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; \text{all three use the same test statistic} \\ &(2)\; \text{all three are right-tailed} \\ &(3)\; \text{all three count degrees of freedom the same way} \\ &(4)\; \text{none of the three has a parameter} \end{aligned} \)

  • (1) is correct: the sum of squared departures over expected counts, in every case.
  • (2) is correct, and for the same reason each time — squaring destroys direction.
  • (3) is false. Goodness of fit uses cells minus one; the other two use a product.
  • (4) is correct, which is why all three state their hypotheses in words.

Statement (3) is the only false one, and it is the only computational difference among the three. That is worth holding on to: identify the test to get the degrees of freedom right, then everything else follows identically.

42. Two rules, three tests

Faded example

Fill in the degrees-of-freedom rules.

Fill in the blanks

\textk - 1: (r-1)(c-1); \qquad \text___: ___

Why: Two rules across three tests. The shared rule reflects a shared structure — both build expected counts from margins and are constrained by every row and column total.

43. Why no confidence intervals?

Prediction

Commit before reasoning.

Predict first

Why does this chapter offer no confidence intervals for its first three tests?

  • There is no parameter to build an interval around
  • The distribution is skewed
  • The statistic cannot be negative
  • The sample sizes are too small

Correct: There is no parameter.

Why: An interval estimates a number, and these tests concern whole distributions rather than any single quantity. Chapters 8 through 10 could always pair a test with an interval because each concerned a mean or a proportion; here there is nothing for an interval to be about.

44. Explain the necessity

Explain it

A classmate asks why the book needs a whole section just to compare three tests it has already explained.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what the arithmetic reveals about which test is being run.

Answer:

The three share their statistic, their right tail, their condition and — for two of them — their degrees of freedom, so nothing that appears during the calculation indicates which test is being performed.

The choice has to be made from the study's description before any arithmetic starts, and this section is the procedure for making it.

45. Routing a study in practice

Section

Section 5

46. Two questions, in order

Concept

How many populations were sampled, and how many questions were asked. One population and two questions is independence; two populations and one question is homogeneity; one population and one question with a known distribution is goodness of fit.

the routing procedure — Read the sampling description first, count populations, then count questions. Fall back on the wording of the hypotheses when the description is unclear.

\[ (1,2) \to \text{ind}; \quad (2,1) \to \text{hom}; \quad (1,1) \to \text{gof} \]

The fourth possibility, two populations and two questions, is outside this chapter. It would require comparing whole contingency tables across groups, which needs methods beyond an introductory course — so in practice the two counts always land on one of the three tests covered here.

Figure (svg): A decision tree routing a study to one of the three chi-square tests

Two questions route every study. The arithmetic that follows is nearly identical whichever branch is taken.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — the summary list

47. The decision tree

Picture it

Two branches, three destinations.

Figure (svg): A decision tree routing a study to one of the three chi-square tests

Two questions route every study. The arithmetic that follows is nearly identical whichever branch is taken.

Everything below the destinations is identical: the same statistic, the same right tail, the same conclusion structure. Only the expected counts and the degrees of freedom depend on which box was reached.

48. Worked example: five studies routed

Worked example

Applying the two counts.

\[ \text{route each} \]

Blood types against known percentages

Why: One, one, known.

Anxiety by need to succeed, 400 students

Why: One, two.

Urban against rural on transport

Why: Two, one.

Grocery totals against a uniform shape

Why: One, one, known.

Degree subject by employment sector

Why: One, two.

Figure (svg): The solution to Worked example five studies routed shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (1,1), (1,2), (2,1), (1,1), (1,2) \]

Verify: confirm the fourth against the book's own lab

Why: Section 11.7 has students collect thirty grocery receipt totals from one supermarket and test whether they fit the uniform or exponential distribution — one population, one measurement, and a specified shape. That is the goodness-of-fit signature exactly, and the lab's warning about combining categories to reach an expected five confirms which test is intended.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

49. Route each study

Sorting

Each names its sampling.

Sort into buckets

Sort by which test applies.

Goodness of fit
one sample of 600 against published percentages; one sample of 60 tested for uniformity
Independence
one sample of 400 classified by two variables
Homogeneity
two samples of 250 and 300, one question each; two samples surveyed before and after an event
gof
One population, one question, and a known distribution to test against.
ind
One population, two questions, giving a contingency table.
hom
Two populations, one question, with no known distribution.

Every routing here came from the sampling description alone. None required looking at a single count.

50. Worked example: when the description is thin

Worked example

A problem gives a table and asks whether the groups differ, with no sampling detail.

\[ \text{no design stated} \]

Try the counts

Why: Cannot: no description.

Read the question asked

Why: Do the groups differ?

Write the null it implies

Why: Same distribution.

Note the arithmetic

Why: Identical either way.

Figure (svg): The solution to Worked example when the description is thin shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \text{the groups follow the same distribution} \]

Verify: confirm nothing is at risk when the two readings are confused

Why: Independence and homogeneity produce identical expected counts, statistics, degrees of freedom and p-values, so an ambiguous case cannot yield a wrong number. Only the sentence describing the conclusion changes, and stating it in the terms the question used — do the groups differ — keeps the report faithful to what was asked.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582

51. Trap: routing from the table's shape

Trap

The trap

\[ \text{a two-row table} \;\Rightarrow\; \text{homogeneity} \]

Match the shape against remembered examples

Why: Both homogeneity examples had two rows.

\[ \text{but a two-row table can be either} \]

One sample of 550 classified by sex and living arrangement produces the same two-row table from an independence design.

The fix

\[ \text{count the samples described, then the questions asked} \]

Route from the study, not the display

Why: The table is the same either way.

The shape does carry one reliable signal, though: a single ROW of counts cannot be a contingency test at all, since neither independence nor homogeneity has anything to cross-classify. So a one-row table means goodness of fit, and only two-way tables are ambiguous.

52. Route from counts

Faded example

Two samples, one question each.

Fill in the blanks

(\text2, \text1) = (___, ___) \;\Rightarrow\; \text___

Why: Two and one is the homogeneity signature, and no other test shares it. Reversing the two counts to one and two would give a test of independence instead.

53. One of these is false

Two truths and a lie

All three concern routing.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. A one-row table cannot be a contingency test
  • C. Confusing independence with homogeneity changes no numbers
  • B. A two-row table is always a homogeneity test

Survives elimination: B

Why: The survivor is false. One sample classified by two variables, one of which has two levels, gives a two-row table from an independence design. Only the sampling description resolves it.

54. Two populations, two questions?

Prediction

Commit before reasoning.

Predict first

Where does a study with two populations AND two questions fall?

  • Outside this chapter's three tests
  • Independence
  • Homogeneity
  • Goodness of fit

Correct: Outside this chapter.

Why: Comparing whole contingency tables across populations requires methods beyond an introductory course. The three tests here cover the three combinations that remain, which is why the two counts always route successfully within the material as presented.

55. The three tests, side by side

Comparison

Fill the blanks. Two rows separate them; the rest do not.

Comparison matrix

Goodness of fitIndependenceHomogeneity
Populationsoneonetwo
Questionsonetwoone
Null namesa given distributiontwo variablestwo populations
Degrees of freedomcells minus one(r-1)(c-1)(r-1)(c-1)

The first two rows never repeat a pair, which is what makes them decisive. The last row shows that two of the three tests are computationally identical, so the distinction between them lives entirely in the third row.

56. Choosing a chi-square test, in order

Pattern

Four steps, all before any arithmetic.

  1. Read the sampling description and count how many populations were sampled.
  2. Count how many questions each subject was asked.
  3. If one and one, check that a distribution is given; if it must be estimated from a sample, the study is really a two-population comparison.
  4. Write the null in the book's wording for that test — fits, independent, or the same distribution — and only then begin computing.

If the description is too thin to count, read the question being asked and write the null it implies; between independence and homogeneity the arithmetic is identical either way.

OpenStax Introductory Business Statistics 2e, §11.6 Comparison of the Chi-Square Tests §11.6 Comparison of the Chi-Square Tests

57. Check yourself 1 of 3

Check

Routing.

Check your understanding

A single sample of 500 people is asked both their region and their preferred news source. Which test?

  • A. A test of independence (correct)
  • B. A test for homogeneity
  • C. A goodness-of-fit test
  • D. A test of a single variance

Answer: A

Why: One population and two questions is the independence signature, and the two questions produce the contingency table.

Why B tempts people
That would need two separately drawn samples, not one classified twice.
Why C tempts people
No external distribution is given, and each subject answered two questions rather than one.
Why D tempts people
That concerns a variance, and the data here are categorical.

58. Check yourself 2 of 3

Check

Known against unknown.

Check your understanding

A sample of 300 is compared against percentages computed from a separate sample of 400. Which test?

  • A. A test for homogeneity (correct)
  • B. A goodness-of-fit test
  • C. A test of independence
  • D. Either, since the arithmetic is the same

Answer: A

Why: Percentages from a sample are estimates, not a known distribution, so both distributions are unknown and the comparison is between two populations.

Why B tempts people
That requires a known distribution; treating a sample's percentages as exact would understate the uncertainty.
Why C tempts people
That needs one sample classified two ways, not two samples.
Why D tempts people
The arithmetic differs here: goodness of fit would use fixed expected counts and cells minus one.

59. Check yourself 3 of 3

Check

The hypotheses.

Check your understanding

A null reads: the two variables are independent. Which test is it?

  • A. A test of independence (correct)
  • B. A test for homogeneity
  • C. A goodness-of-fit test
  • D. Cannot be determined

Answer: A

Why: A null naming two variables or factors belongs to the test of independence; homogeneity's names two populations and goodness of fit's names a given distribution.

Why B tempts people
Its null reads that the two populations follow the same distribution.
Why C tempts people
Its null reads that the population fits the given distribution.
Why D tempts people
The subject of the null identifies the test uniquely, since the three subjects cannot overlap.

60. Where this shows up outside the textbook

Real world

A market research report states: we surveyed 1,200 customers about their preferred delivery option and cross-tabulated the results by membership tier, finding a chi-square statistic of 34.7 with p below 0.001, which shows that our three membership tiers have significantly different delivery preferences.

Discussion prompt

Which test was actually run, which does the conclusion describe, and does the difference matter here?

Hint: Count the samples described, then read the conclusion's subject.

Answer:

The study is a test of independence. One sample of 1,200 customers was drawn and then classified two ways — delivery option and membership tier — which is one population and two questions. The tiers are levels of a variable within that single sample, not three populations that were separately sampled.

The conclusion is worded as homogeneity, describing three tiers as though each had been sampled in its own right and their distributions compared. That is the wrong description of what was done, and it matters for a practical reason rather than an arithmetic one: it implies the tier sizes were chosen by the researcher, when in fact they emerged from the data and reflect however the customer base happens to be distributed.

\[ \text{one sample, two questions} \;\Rightarrow\; \text{independence}; \qquad \text{df} = (3-1)(c-1) \]

The numbers are unaffected, since the two tests compute identically — the statistic, degrees of freedom and p-value would be the same under either description. So nothing needs recalculating, and the finding that tier and delivery preference are associated stands.

Two things would improve the report. Stating it as an association between tier and preference rather than as a difference between populations describes the study accurately. And since the tiers were observed rather than assigned, the association could easily run through order size, geography or tenure — so the sentence about what to do next should not assume that moving a customer between tiers would change their delivery preference.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

What distinguishes the three chi-square tests?

  • Their test statistics
  • How many populations were sampled and how many questions were asked
  • Which tail is used
  • The significance level

Correct: The number of populations and the number of questions.

\[ (1,1) \to \text{gof}; \quad (1,2) \to \text{ind}; \quad (2,1) \to \text{hom} \]

Why: All three use the same statistic and the same right tail, and all three require expected counts of at least five. Goodness of fit has one population and one question with a known distribution; independence has one population and two questions; homogeneity has two populations and one question. No two share both counts, so the pair identifies the test — and the wording of the hypotheses confirms it.

62. Explain it to someone a year behind you

Explain it

They ask how to tell a test of independence from a test for homogeneity when both give the same numbers.

Discussion prompt

In two sentences or fewer, tell them what to look at.

Hint: Ask what was decided before the data were collected.

Answer:

Count the samples the study describes: one sample classified by two variables is independence, while two separately drawn samples each answering one question is homogeneity.

The arithmetic is identical either way, so the choice only affects how the conclusion is worded — an association between two variables, or a difference between two populations.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Counting populations and questions from a study description
  • Recognising each test from the wording of its null
  • Deciding whether a comparison distribution counts as known
  • Saying what all three tests share, and why that matters

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, count samples drawn and things each subject was asked. For the second, the null's subject: a given distribution, two variables, or two populations. For the third, known means specified without using sampled data. For the fourth, everything computational is shared, which is why the choice comes first. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twelve minutes — this is a summary section, so the page is the point.

Draw it

Draw a three-column table across the top with the headings goodness of fit, independence and homogeneity, and four rows: populations, questions, whether a distribution is known, and the degrees-of-freedom rule. Fill it in — one, one, yes, cells minus one; one, two, not applicable, the product; two, one, no, the product. Underneath, write out all three pairs of hypotheses in the book's wording, and underline the subject of each null: a given distribution, two variables, two populations. To the right, draw the decision tree: how many populations, branching to homogeneity on two, and on one branching again by how many questions to goodness of fit or independence. At the bottom, draw two boxes side by side. In the left write the five things all three share — the statistic, the right tail, the at-least-five condition, the chi-square distribution and hypotheses in words. In the right write the one computational thing that differs: the degrees-of-freedom rule, with goodness of fit alone against the other two.

Check your table by confirming that no two columns agree on BOTH of the first two rows, which is what makes those counts decisive. Check your bottom boxes by asking whether anything in the left box could help you choose a test — nothing in it can, which is why the choice has to come first.

65. What you can do now

Recap

Five things, and none of them is a calculation.

If you seeThen
One sample, one question, a known distributionGoodness of fit, with df of cells minus one
One sample, two questionsIndependence, with df a product
Two samples, one questionHomogeneity, with df the same product
Percentages computed from a sampleNot a known distribution: use homogeneity
A null naming a given distributionGoodness of fit
A null naming two factorsIndependence
A null naming two populationsHomogeneity
A one-row table of countsGoodness of fit; there is nothing to cross-classify

Section 11.6 breaks every pattern in this comparison. Its null names a parameter, its data are quantitative rather than categorical, its degrees of freedom are n minus one, and it is the only chi-square test in the chapter that can be left-tailed or two-tailed.

OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests §11.5, p. 582 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §11.5 Comparison of the Chi-Square Tests — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 582-582
  2. OpenStax Introductory Business Statistics 2e, §11.6 Comparison of the Chi-Square Tests — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108