11.2 Goodness-of-Fit Test

The chi-square distribution's first use, and the chapter's first genuinely new question: instead of testing a single mean or proportion, a goodness-of-fit test asks whether a whole set of observed counts is consistent with a claimed distribution. The statistic sums observed minus expected, squared, over expected across the cells, so a departure in any cell and in either direction makes it larger — which is why the test is almost always right-tailed and why small statistics mean good agreement. The degrees of freedom are the number of categories minus one, never the number of observations minus one, and every expected count must be at least five, a condition that in the book's first example forces two categories to be combined before the test can be run at all. Worked through the book's four examples: absenteeism, absences by weekday, streaming services and a pair of coins.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 11.2 Goodness-of-Fit Test

Title

Statistics · Chapter 11 — The Chi-Square Distribution

Goodness-of-Fit Test

2. By the end of this lesson you can

Objectives

Six outcomes. Two of them are conditions to check rather than calculations to do.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-573 — the section these objectives are drawn from

3. What you already have

Warm-up

Chapter 9 tested one parameter at a time; section 11.1 gave a distribution built from squares.

Discussion prompt

A die is rolled 60 times and lands on each face 12, 8, 11, 9, 13 and 7 times. How would you test whether it is fair, using only chapter 9's tools?

Hint: Chapter 9 could test one proportion. How many would you need here?

Answer:

Chapter 9 could test each face separately — six tests, one per proportion, each asking whether that face's rate differs from one sixth. But six tests answer six questions, not one, and running all of them inflates the chance that at least one comes out significant by accident.

What is actually wanted is a single verdict on the whole set of counts at once. That needs a statistic that combines every cell's departure into one number, and section 11.1 supplies the shape: squaring each departure makes them all positive so they accumulate rather than cancel, and the sum of squares has a chi-square distribution.

This section builds exactly that. Six counts against six expectations become one statistic and one p-value, answering the question that was actually asked.

4. Comparing a whole set of counts against a claim

Concept

In this type of hypothesis test you determine whether the data fit a particular distribution or not. The test statistic sums, across the cells, the squared difference between observed and expected divided by expected. The degrees of freedom are the number of categories minus one.

a goodness-of-fit test — A chi-square test of whether observed counts are consistent with the distribution a hypothesis claims. The null hypothesis is that the data fit; the alternative is that they do not.

\[ \chi^2 = \sum_{i=1}^{k} \frac{(O - E)^2}{E}, \qquad \text{df} = k - 1 \]

The book's phrasing of the hypotheses is worth copying: they may be written in sentences or as equations or inequalities, and in practice sentences are clearer. Example 11.1's are simply that student absenteeism fits faculty perception, against that it does not — no parameter is named, because the claim is about a whole distribution rather than any single number.

Figure (svg): A card breaking the goodness-of-fit statistic into its parts

The book's own summary of the test: the statistic, what its letters mean, the degrees of freedom, and the one condition.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 565

5. The statistic and its degrees of freedom

Section

Section 1

6. One term per cell

Concept

The observed values are the data and the expected values are what you would expect to get if the null hypothesis were true. Each cell contributes a squared difference scaled by its expected count, and the degrees of freedom are the number of categories minus one.

scaling by E — Dividing by the expected count makes a discrepancy of ten matter more when only twelve were expected than when three hundred were. Without it, large cells would dominate every test.

\[ \frac{(O-E)^2}{E} \quad \text{per cell}, \qquad \text{df} = (\text{number of categories}) - 1 \]

The book adds a note that is easy to skip past: df is not 600 minus 1 in Example 11.3, even though 600 families were surveyed. The degrees of freedom count the cells, and with five categories of streaming service the answer is four however many families were asked. Nothing about the sample size enters the count.

Figure (svg): A card breaking the goodness-of-fit statistic into its parts

The book's own summary of the test: the statistic, what its letters mean, the degrees of freedom, and the one condition.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-571 — the statistic, the df rule, and the note that df is not n minus one

7. The statistic in pieces

Picture it

What each symbol means, and the one condition attached.

Figure (svg): A card breaking the goodness-of-fit statistic into its parts

The book's own summary of the test: the statistic, what its letters mean, the degrees of freedom, and the one condition.

Every term is a square over a positive count, so no term can be negative and none can cancel another. That is why the statistic only grows as the data depart from expectation, whatever direction the departures run in.

8. Worked example: Example 11.2, absences by weekday

Worked example

Sixty managers named the day with the most employee absences: 15, 12, 9, 9 and 15 across Monday to Friday. Do the days occur with equal frequencies? Test at 5 percent.

\[ O = 15, 12, 9, 9, 15 \]

Total the observations

Why: Fifteen plus twelve plus nine plus nine plus fifteen.

\[ 60 \]

Split it equally

Why: Sixty over five days.

\[ E = 12\text{ each} \]

Sum the terms

Why: Nine, zero, nine, nine and nine, each over twelve.

\[ 3.00 \]

Count the cells

Why: Five days minus one.

\[ d f = 4 \]

Find the right tail

Why: Beyond 3 on four df.

\[ 0.5578 \]

Figure (svg): The solution to Worked example Example 11.2, absences by weekday shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 3, \quad \text{df} = 4, \quad p = 0.5578 \]

Verify: confirm the statistic is plausible before consulting any table

Why: The mean of a chi-square with four degrees of freedom is 4, and the statistic here is 3 — below its own mean, so more than half the distribution lies above it and the p-value must exceed one half. It does, at 0.5578. The book's conclusion follows: at a 5 percent level there is not sufficient evidence to conclude that the absent days do not occur with equal frequencies.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 568-569

9. One cell's contribution

Faded example

Example 11.3's last cell observed 15 against an expected 48.

Fill in the blanks

\frac108922.69 = \frac___}___ = ___

Why: That single cell contributes 22.69 of the total 29.65 — over three quarters of the statistic. Looking at which cell dominates is how a rejected goodness-of-fit test gets interpreted.

10. Worked example: Example 11.4, are the coins fair?

Worked example

Two coins are flipped 100 times, giving 20 HH, 27 HT, 30 TH and 23 TT. Test at 5 percent.

\[ 20\;HH,\; 27\;HT,\; 30\;TH,\; 23\;TT \]

Define the variable

Why: X counts heads in one flip of the two.

\[ X = 0, 1, 2 \]

Collapse to three cells

Why: HT and TH both give one head.

\[ 20, 57, 23 \]

Expect on fairness

Why: A quarter, a half, a quarter.

\[ 25, 50, 25 \]

Sum the terms

Why: One plus 0.98 plus 0.16.

\[ 2.14 \]

Right tail on df 2

Why: Three cells minus one.

\[ 0.3430 \]

Figure (svg): The solution to Worked example Example 11.4, are the coins fair shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 2.14, \quad \text{df} = 2, \quad p = 0.3430 \]

Verify: confirm the cell count, which is where this example is easy to get wrong

Why: The sample space has four outcomes but the random variable takes only three values, since HT and TH are both one head. So there are three cells and two degrees of freedom, not four cells and three. Using four cells would put 27 and 30 against expectations of 25 each — a different statistic on different degrees of freedom, and the wrong answer to the question actually asked about the number of heads.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 573

11. Trap: using n minus one for the degrees of freedom

Trap

The trap

\[ 600 \text{ families in 5 categories} \;\Rightarrow\; \text{df} = 599 \]

Reach for the familiar n minus one

Why: It was right for every t test in chapters 8 through 10.

\[ \text{but the cells are what is being counted} \]

The book flags this directly with a note reading df is not 600 minus 1.

The fix

\[ \text{df} = k - 1 = 5 - 1 = 4 \]

Count categories, not observations

Why: The sample size never enters the degrees of freedom here.

Sample size does matter, but through the expected counts rather than the degrees of freedom. Surveying 6,000 families instead of 600 would multiply every expected count by ten and make real departures far easier to detect — while leaving the degrees of freedom at four.

12. One of these is false

Two truths and a lie

All three concern the statistic.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Each term is divided by its expected count
  • C. The degrees of freedom are categories minus one
  • B. Terms can cancel if some cells are above and others below

Survives elimination: B

Why: The survivor is false. Each term is squared before being added, so every departure contributes positively whatever its direction. That is exactly why the test detects any kind of misfit and why the statistic can only be large when the fit is poor.

13. Why divide by E?

Prediction

Commit before reasoning.

Predict first

What does dividing each squared difference by its expected count accomplish?

  • It judges each discrepancy against how much was expected there
  • It keeps the statistic below one
  • It makes the terms negative when observed is small
  • It converts counts into percentages

Correct: It judges each discrepancy relative to its expected count.

Why: Missing by ten is serious when twelve were expected and trivial when three hundred were. Dividing by E puts every cell on a comparable footing, which is what lets a small cell and a large one contribute meaningfully to the same total.

14. Which cell dominates?

Estimation

Example 11.3: observed 66, 119, 340, 60, 15 against expected 60, 96, 330, 66, 48.

Predict first

Which cell contributes most to the statistic of 29.65?

  • The last, 15 against 48
  • The third, 340 against 330
  • The second, 119 against 96
  • They contribute equally

Correct: The last: 15 against 48.

Why: Its contribution is 22.69, against 5.51 for the second cell and only 0.30 for the third. The third cell has the largest counts but the smallest relative miss, which is exactly what dividing by E is designed to capture.

15. Why the test is right-tailed

Section

Section 2

16. Only large values are evidence

Concept

The goodness-of-fit test is almost always right-tailed. If the observed values and the corresponding expected values are not close to each other, then the test statistic can get very large and will be way out in the right tail of the chi-square curve.

right-tailed by construction — The statistic measures disagreement, so poor fit can only make it large. A small statistic means the observations sit close to expectation, which is never evidence against the null.

\[ p\text{-value} = P(\chi^2 > \chi^2_{\text{obs}}) \]

This is a real break from chapter 9, where the direction of the tail was a decision that followed from how the alternative hypothesis was worded. Here it follows from the statistic's construction instead: because every departure is squared, disagreement in any direction pushes the statistic the same way, so there is no left-tailed version of the question to ask.

Figure (svg): A chi-square curve on four degrees of freedom with the area to the right of three shaded, covering more than half the distribution

Example 11.2 drawn: the observed absences agree with the uniform expectation, so the shaded right tail is large.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-570 — almost always right-tailed; this test is always right-tailed

17. Example 11.2's picture

Picture it

A statistic of 3 against a distribution with mean 4.

Figure (svg): A chi-square curve on four degrees of freedom with the area to the right of three shaded, covering more than half the distribution

Example 11.2 drawn: the observed absences agree with the uniform expectation, so the shaded right tail is large.

The shaded region is over half the curve, which is what a p-value of 0.5578 means. The book asks for exactly this picture, properly labelled with the right tail shaded, as part of the solution.

18. Worked example: Example 11.3, streaming services

Worked example

Six hundred far-western families reported 66, 119, 340, 60 and 15 across five categories, against a national distribution of 10, 16, 55, 11 and 8 percent. Test at 1 percent.

\[ \alpha = 0.01 \]

Build expected counts

Why: Each percent times 600.

\[ 60, 96, 330, 66, 48 \]

Check the condition

Why: Smallest expected is 48.

Sum the terms

Why: By calculator.

\[ 29.65 \]

Degrees of freedom

Why: Five cells minus one.

\[ 4 \]

Right tail

Why: Beyond 29.65.

\[ 0.000006 \]

Figure (svg): The solution to Worked example Example 11.3, streaming services shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 29.65, \quad p = 0.000006 < 0.01 \]

Verify: confirm the statistic is extreme relative to its degrees of freedom

Why: The mean of the distribution is 4 and its standard deviation is the root of 8, about 2.83 — so a statistic of 29.65 sits some nine standard deviations above the mean. A p-value in the millionths is exactly what that predicts, and the conclusion is that the far western distribution genuinely differs, driven overwhelmingly by the 4-or-more cell where 15 appeared against 48 expected.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571

19. Which statistics are evidence?

Sorting

Each is paired with its degrees of freedom.

Sort into buckets

Sort by whether the result is evidence against the null.

Evidence against the null
chi-square 29.65 on df 4; chi-square 14.29 on df 3
Not evidence
chi-square 3 on df 4; chi-square 2.14 on df 2; chi-square 0 on df 5
yes
The statistic is far above its degrees of freedom, so the right tail beyond it is tiny.
no
The statistic is at or below its degrees of freedom, so the right tail is large.

Items (a), (b), (c) and (d) are Examples 11.3, 11.2, 11.1 and 11.4, with p-values of 0.000006, 0.5578, 0.0025 and 0.3430. Item (e) is a perfect fit.

20. Worked example: what a small statistic would mean

Worked example

Suppose the far-western counts had been 60, 96, 330, 66, 48 exactly.

\[ O = E \text{ in every cell} \]

Each term

Why: Zero over its expected count.

\[ 0 \]

The sum

Why: Five zeros.

\[ 0 \]

The right tail beyond 0

Why: The whole distribution.

\[ p = 1 \]

Interpret

Why: Perfect agreement.

Figure (svg): The solution to Worked example what a small statistic would mean shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 0 \;\Longrightarrow\; p = 1 \]

Verify: confirm this is why a left tail would be meaningless

Why: The statistic's smallest possible value is zero, and it means the observations equal the expectations exactly. There is no way for data to be worse than fitting perfectly, so the lower end of the scale carries no evidence against the null and no left-tailed test exists. That is a structural fact about the statistic rather than a choice about how to word the alternative.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 565

21. Error analysis: four statements about the tail

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; \text{a small statistic is evidence against the null} \\ &(2)\; \text{the p-value is the area to the right of the statistic} \\ &(3)\; \text{the direction is chosen from the alternative's wording} \\ &(4)\; \text{a large statistic means observed and expected are far apart} \end{aligned} \)

  • (1) is backwards. A small statistic means observed and expected are close, which supports the null rather than contradicting it.
  • (2) is correct, and is what the book's own graphs shade.
  • (3) was true in chapter 9 but is false here. The direction follows from the statistic's construction, not from the hypothesis's wording.
  • (4) is correct, and it is the book's own justification for the right tail.

Error (3) is the one worth naming, because chapter 9 spent a whole section on choosing the tail. That habit has to be dropped: every test in sections 11.2 to 11.4 is right-tailed regardless of how the alternative is phrased.

22. One of these is false

Two truths and a lie

All three concern the direction of the test.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The p-value is a right-tail area
  • C. The statistic cannot go below zero
  • B. A left-tailed goodness-of-fit test is used when observed counts are too small

Survives elimination: B

Why: The survivor is false. Cells with observed below expected contribute positively to the statistic just as cells above do, since the difference is squared. There is no left-tailed version because there is no way for data to disagree in a way that makes the statistic small.

23. What changed from chapter 9?

Prediction

Commit before reasoning.

Predict first

In chapter 9 the tail was chosen from the alternative hypothesis. Why not here?

  • The statistic squares every departure, so all disagreement lands in the upper tail
  • Because chi-square has no left tail
  • Because the alternative is always two-sided
  • Because the significance level is fixed

Correct: Every departure is squared, so all disagreement lands on the right.

Why: A z or t statistic keeps the sign of the departure, so the tail had to match the alternative's direction. A chi-square statistic discards the sign by squaring, so both directions of departure push it the same way and only the upper tail carries evidence.

24. Explain the picture

Explain it

A classmate has drawn Example 11.2's graph and shaded the area to the LEFT of 3.

Discussion prompt

In two sentences or fewer, tell them what is wrong.

Hint: Ask what the p-value is meant to measure.

Answer:

The p-value is the chance of a statistic at least as extreme as the one observed, and extreme means large here — so the shading belongs to the right of 3, not the left.

Shading the left would give about 0.44 instead of 0.5578, and would treat a good fit as though it were evidence against the null.

25. Building the expected counts

Section

Section 3

26. What the data would look like if the null were true

Concept

The expected values are the values you would expect to get if the null hypothesis were true. When the claim is a set of percentages, multiply each by the sample total; when the claim is that categories are equally likely, divide the total by the number of categories.

expected counts — Not the observed data, and not necessarily whole numbers. They are what the hypothesised distribution predicts for a sample of this size, and they always sum to the observed total.

\[ E_i = n \cdot p_i, \qquad \sum E_i = n \]

The check that the expected counts total the sample size is the fastest way to catch an arithmetic slip. In Example 11.3 they total 60 plus 96 plus 330 plus 66 plus 48, which is 600 — the number of families surveyed. If they do not add to n, either a percentage was misread or the multiplication went wrong, and the statistic will be meaningless.

Figure (svg): A table converting the claimed percentages into expected frequencies for six hundred families

A claimed distribution given as percentages has to be turned into counts before the statistic can be formed.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571 — converting percentages to expected frequencies

27. Percentages into counts

Picture it

Example 11.3's conversion, alongside what was actually observed.

Figure (svg): A table converting the claimed percentages into expected frequencies for six hundred families

A claimed distribution given as percentages has to be turned into counts before the statistic can be formed.

The book's advice about letting the calculator do the arithmetic — entering 0.10 times 600 rather than 60 — matters when the percentages produce awkward decimals. Rounding expected counts before summing is a common source of small discrepancies.

28. Worked example: equal frequencies

Worked example

Example 11.2's expected counts, from the claim of a uniform distribution.

\[ 60 \text{ absences over } 5 \text{ days} \]

Total the observations

Why: 15 + 12 + 9 + 9 + 15.

\[ 60 \]

State the claim

Why: Equal frequencies.

Multiply

Why: Sixty times one fifth.

\[ 12 \]

Repeat for each cell

Why: All the same here.

\[ 12, 12, 12, 12, 12 \]

Figure (svg): The solution to Worked example equal frequencies shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ E = 12 \text{ for each of the five days} \]

Verify: confirm the expected counts sum to the observed total

Why: Five twelves are 60, which matches the total number of absences reported. That check applies to every goodness-of-fit test regardless of the claimed distribution, since the expected counts are the total split according to the hypothesised proportions and those proportions sum to one.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 568

29. Expected from a percentage

Faded example

A claimed 24 percent, in a sample of 250.

Fill in the blanks

E = 0.24 \times 250 = 60

Why: Expected counts need not be whole numbers, though this one is. A claimed 23 percent of 250 would give 57.5, and that value is used as it stands rather than rounded.

30. Worked example: percentages

Worked example

Example 11.3's national distribution applied to 600 families.

\[ 10\%, 16\%, 55\%, 11\%, 8\% \]

Convert to decimals

Why: Each percentage over 100.

\[ 0.10\text{ to } 0.08 \]

Multiply by 600

Why: One cell at a time.

\[ 60, 96, 330, 66, 48 \]

Check the total

Why: Sum them.

\[ 600 \]

Check the condition

Why: The smallest.

\[ 48,\text{ comfortably above } 5 \]

Figure (svg): The solution to Worked example percentages shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ E = 60, 96, 330, 66, 48 \]

Verify: confirm the percentages themselves were complete

Why: Ten plus sixteen plus fifty-five plus eleven plus eight is one hundred, so the claimed distribution accounts for every family and nothing is missing. A set of percentages that fails to total 100 means a category has been left out, and the test cannot be run until it is found — the expected counts would not sum to n.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 570

31. Trap: comparing percentages instead of counts

Trap

The trap

\[ \frac{(11 - 10)^2}{10} + \frac{(19.8 - 16)^2}{16} + \cdots \]

Convert the observed data to percentages and compare those

Why: It seems to put the two on the same footing.

\[ \text{the statistic no longer depends on the sample size} \]

Six hundred families and six families would give the same statistic, which cannot be right — a departure seen in 600 observations is far stronger evidence.

The fix

\[ \frac{(66 - 60)^2}{60} + \frac{(119 - 96)^2}{96} + \cdots = 29.65 \]

Convert the claim to counts and compare counts

Why: The sample size then enters through the expected values.

This is the same principle as the standard error's dependence on n throughout chapters 8 to 10: evidence should strengthen with more data. Working in percentages throws that away, and a test built on percentages would reach the same verdict from six families as from six hundred.

32. Does the condition hold?

Sorting

Each is an expected count in some cell.

Sort into buckets

Sort by whether the cell may be used as it stands.

Usable: at least five
expected 48; expected 5; expected 12
Must be combined with a neighbour
expected 2; expected 4.8
ok
The expected count is five or more, which is what the book requires.
combine
The expected count is below five, so this cell must be merged with an adjacent one.

Item (c) is exactly at the boundary and passes, since the requirement is at least five. Item (d) at 4.8 fails despite rounding to five — the condition applies to the expected count itself, not to a rounded version of it.

33. One of these is false

Two truths and a lie

All three concern expected counts.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. They always sum to the sample total
  • C. They need not be whole numbers
  • B. They are computed from the observed data

Survives elimination: B

Why: The survivor is false. Expected counts come from the hypothesised distribution — the theory — and the only thing they take from the data is the total. That separation is what makes the comparison meaningful; if both sides came from the data they would always agree.

34. What if the percentages do not total 100?

Prediction

Commit before reasoning.

Predict first

A claimed distribution's percentages sum to 92. What has gone wrong?

  • A category has been omitted, so the expected counts will not sum to n
  • Nothing; the test rescales automatically
  • The sample is too small
  • The test becomes two-tailed

Correct: A category is missing.

Why: Every observation must fall in exactly one cell, so the claimed proportions must account for all of them. Percentages totalling 92 mean 8 percent of the population is unclassified, and the expected counts would fall short of the sample total — a discrepancy that would show up in the sum check.

35. The condition, and combining cells

Section

Section 4

36. Every expected count at least five

Concept

The expected value for each cell needs to be at least five in order for you to use this test. When a cell falls short, it is combined with an adjacent one — which reduces the number of categories and so reduces the degrees of freedom.

combining cells — Merging adjacent categories so that each expected count reaches five. The observed counts are added together too, and the category labels are rewritten to describe the merged range.

\[ E_i \ge 5 \text{ for every } i \]

The reason behind the condition is the same one that sat behind chapter 7's normal approximation to the binomial. A chi-square distribution describes the statistic only approximately, and the approximation relies on each cell's count behaving roughly normally. With an expected count of two that is far from true, and the resulting p-value cannot be trusted.

Figure (svg): Two tables showing the absence categories before and after the last two are combined

The condition has to be checked before the test is set up, because meeting it changes the degrees of freedom.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-567 — the note on expected counts, and Example 11.1's combining

37. Example 11.1's tables

Picture it

The absence categories, and why two of them merge.

Figure (svg): Two tables showing the absence categories before and after the last two are combined

The condition has to be checked before the test is set up, because meeting it changes the degrees of freedom.

The book asks first whether the test can be run on the charts as they appear, and the answer is no. Combining the last two categories gives a 9-or-more cell expecting 8 and observing 5, and reduces the cells from five to four so the degrees of freedom become 3.

38. Worked example: Example 11.1, absenteeism

Worked example

Faculty expected 50, 30, 12, 6 and 2 students in five absence bands; the survey observed 35, 40, 20, 1 and 4.

\[ E = 50, 30, 12, 6, 2 \]

Check every expected count

Why: The last is two.

Combine the last two

Why: Six plus two, one plus four.

\[ E = 8, O = 5 \]

Recheck the condition

Why: New smallest is eight.

Recount the cells

Why: Four now.

\[ d f = 4 - 1 = 3 \]

Figure (svg): The solution to Worked example Example 11.1, absenteeism shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{df} = \text{number of cells} - 1 = 4 - 1 = 3 \]

Verify: confirm the totals survive the combining

Why: The expected counts still total 100 — 50 plus 30 plus 12 plus 8 — and the observed still total 100 as well, at 35 plus 40 plus 20 plus 5. Merging cells redistributes counts without creating or destroying any, so both totals must be unchanged; if either has moved, the merge was done wrongly.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 566-567

39. After combining

Faded example

Six categories, two of which are merged into one.

Fill in the blanks

\text5 = 4, \qquad \text___ = ___

Why: Merging two cells into one leaves five, so the degrees of freedom are four rather than the five that the original six categories would have given.

40. Worked example: completing Example 11.1

Worked example

The book stops at the degrees of freedom, so here is the test it sets up.

\[ O = 35, 40, 20, 5; \; E = 50, 30, 12, 8 \]

First cell

Why: 225 over 50.

\[ 4.500 \]

Second cell

Why: 100 over 30.

\[ 3.333 \]

Third cell

Why: 64 over 12.

\[ 5.333 \]

Fourth cell

Why: 9 over 8.

\[ 1.125 \]

Sum, and take the right tail on df 3

Why: Add them.

\[ 14.29, p = 0.0025 \]

Figure (svg): The solution to Worked example completing Example 11.1 shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 = 14.29, \quad \text{df} = 3, \quad p = 0.0025 \]

Verify: confirm the statistic is extreme for three degrees of freedom

Why: The mean is 3 and the standard deviation is the root of 6, about 2.45, so 14.29 sits about 4.6 standard deviations above the centre — comfortably consistent with a p-value of 0.0025. Reading the cells shows where the misfit lies: faculty expected half the students in the lowest band and only 35 were there, while the 3-to-5 and 6-to-8 bands both ran well above expectation. Students were absent more than faculty believed.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 566-567

41. Error analysis: four responses to a small expected count

Error analysis

A cell has an expected count of 3. Which responses are legitimate?

Annotate

On: \( \begin{aligned} &(1)\; \text{combine it with an adjacent cell} \\ &(2)\; \text{drop the cell from the analysis} \\ &(3)\; \text{run the test anyway and note the caveat} \\ &(4)\; \text{combine, but keep the original degrees of freedom} \end{aligned} \)

  • (1) is the book's own remedy, used in Example 11.1.
  • (2) is wrong: dropping a cell discards its observations, so the expected counts no longer sum to the sample total.
  • (3) violates the stated condition, and the p-value from a chi-square approximation that does not hold is not interpretable.
  • (4) is the subtle error. Combining reduces the number of cells, and the degrees of freedom must fall with it.

Error (4) is worth dwelling on because it produces a plausible-looking answer. Example 11.1 with four degrees of freedom instead of three gives a p-value of 0.0064 rather than 0.0025 — same decision here, but a different number, and the difference grows as more cells are merged.

42. One of these is false

Two truths and a lie

All three concern the condition.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It applies to expected counts, not observed ones
  • C. Combining changes the degrees of freedom
  • B. Cells with small OBSERVED counts must be combined

Survives elimination: B

Why: The survivor is false, and Example 11.1 shows why the distinction matters: its 9-to-11 cell observed only 1, but its expected count was 6, so that cell was not the problem. The 12-or-more cell was combined because its EXPECTED count was 2.

43. Which cells can be merged?

Prediction

Commit before reasoning.

Predict first

Categories run 0 to 2, 3 to 5, 6 to 8, 9 to 11 and 12 or more. Which merge is sensible?

  • 9 to 11 with 12 or more, since they are adjacent
  • 0 to 2 with 12 or more, since they are the extremes
  • Any two cells, since order does not matter
  • None; the categories are fixed

Correct: 9 to 11 with 12 or more.

Why: The categories are ordered, so merging adjacent ones produces a meaningful new category — 9 or more. Merging the two extremes would create a category describing students who were absent either very little or very much, which answers no sensible question and would be impossible to interpret if the test rejected.

44. How much does the df error cost?

Estimation

Example 11.1's statistic of 14.29 on the correct df of 3 gives p = 0.0025.

Predict first

Keeping df at 4 by mistake would give roughly what p-value?

  • About 0.006
  • About 0.0025
  • About 0.05
  • About 0.0001

Correct: About 0.006.

Why: More degrees of freedom shift the curve right, so the same statistic sits less far out and the tail area grows — here from 0.0025 to 0.0064. The decision is unchanged at 5 percent, but the reported evidence is overstated by more than a factor of two, and with more merged cells the gap widens.

45. Running and reporting the test

Section

Section 5

46. The chapter 9 rhythm, with a new middle

Concept

The overall shape is the familiar one: hypotheses, a statistic, a p-value, a comparison with alpha, a decision and a conclusion in context. What is new sits in the middle — building the expected counts, checking the condition, and counting the cells.

the conclusion in context — Stated in the language of the problem rather than in symbols, and about fit rather than about a parameter. The book's conclusions name the distribution being tested every time.

\[ H_0: \text{the data fit} \quad\text{against}\quad H_a: \text{the data do not fit} \]

One feature of the hypotheses is worth noticing. There is no parameter, so there is nothing to write an equation about — which is why the book says the hypotheses may be written in sentences. A goodness-of-fit test's null is a statement about a whole distribution, and sentences state that more honestly than any equation would.

Figure (svg): The order of operations for a goodness-of-fit test

Steps one, five and six are the familiar chapter 9 rhythm; steps two through four are what is new here.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-570 — the hypotheses, and Example 11.3's full write-up

47. The order of operations

Picture it

Six steps, and the reason step three precedes step four.

Figure (svg): The order of operations for a goodness-of-fit test

Steps one, five and six are the familiar chapter 9 rhythm; steps two through four are what is new here.

The ordering is not cosmetic. Combining cells to satisfy the condition changes how many cells there are, so counting degrees of freedom before the check has been done gives the wrong answer — which is precisely the mistake Example 11.1 is designed to prevent.

48. Worked example: writing the hypotheses

Worked example

For Example 11.3's streaming services.

\[ H_0 \text{ and } H_a \]

Name the two distributions

Why: Far western families, and the nation.

The null

Why: They agree.

The alternative

Why: The opposite.

Note what is absent

Why: No parameter.

Figure (svg): The solution to Worked example writing the hypotheses shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \text{same distribution}; \quad H_a: \text{different} \]

Verify: confirm the alternative really is the negation of the null

Why: Chapter 9 established that the two hypotheses must partition every possibility, and they do here: either the distributions agree in every category or they differ in at least one. Note that the alternative does not say HOW they differ, which is consistent with the test being right-tailed — the statistic detects any departure without identifying its direction.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 571

49. Stage to content

Matching

Match each stage of Example 11.3's write-up to what it contains.

Match the pairs

  • l1. hypotheses
  • l2. distribution for the test
  • l3. probability statement
  • l4. conclusion
  • r1. same distribution against different
  • r2. chi-square with df = 4
  • r3. p-value = 0.000006
  • r4. sufficient evidence of a difference, in context

Why: The sequence is identical to chapter 9's. What changed is the content of the second stage — a chi-square with cell-based degrees of freedom rather than a t with n minus one.

50. Worked example: the full write-up

Worked example

Example 11.3, in the book's own sequence.

\[ \alpha = 0.01 \]

Hypotheses

Why: In sentences.

Distribution for the test

Why: Chi-square.

\[ d f = 5 - 1 = 4 \]

Test statistic

Why: By calculator.

\[ 29.65 \]

Probability statement

Why: A right tail.

\[ p = 0.000006 \]

Compare and decide

Why: Alpha exceeds the p-value.

Figure (svg): The solution to Worked example the full write-up shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p = 0.000006 < 0.01 = \alpha \]

Verify: confirm the conclusion is stated in the problem's own terms

Why: The book's wording names the variable and both populations rather than saying only that the null was rejected. That practice matters more here than in chapter 9, since a goodness-of-fit test's null has no parameter — a reader who is only told the null was rejected has no way to reconstruct what claim was being tested.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571

51. Trap: concluding that a particular cell is different

Trap

The trap

\[ p = 0.000006 \;\Rightarrow\; \text{far western families have fewer with 4+ services} \]

Read a specific cell's story out of a rejected test

Why: The 4-or-more cell did dominate the statistic.

\[ \text{but the test's alternative names no cell} \]

The test rejects the whole distribution, and the alternative hypothesis is only that the distributions differ.

The fix

\[ \text{the distributions differ; the 4+ cell contributed } 22.69 \text{ of } 29.65 \]

Report the verdict the test gives, then describe the cells separately

Why: The contributions are informative but are not what was tested.

The distinction is the same one chapter 9 drew between a test's conclusion and the description that motivated it. Naming the cell that dominated is good practice and helps a reader understand the result; presenting it as the test's own conclusion overstates what a single p-value licenses.

52. Decide

Faded example

Example 11.2, at a 5 percent level.

Fill in the blanks

p = 0.5578, \quad \alpha = 0.05: \quad p > \alpha, \textdo not reject ___ H_0

Why: The p-value far exceeds alpha, so the decision is not to reject. The book's conclusion is that there is not sufficient evidence to conclude that the absent days do not occur with equal frequencies — a double negative that is worth reading twice.

53. One of these is false

Two truths and a lie

All three concern reporting.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The hypotheses may be written in sentences
  • C. The conclusion should name the distribution being tested
  • B. Rejecting identifies which category caused the misfit

Survives elimination: B

Why: The survivor is false. The test delivers one verdict on the whole distribution. Examining individual cell contributions is a sensible follow-up and often illuminating, but it is a description of the data rather than a conclusion the test itself supports.

54. What does a large p-value license?

Prediction

Commit before reasoning.

Predict first

Example 11.2's p-value is 0.5578. What may be concluded?

  • There is not sufficient evidence that the days differ in frequency
  • The days definitely occur with equal frequencies
  • The sample was too small
  • The test was run incorrectly

Correct: There is not sufficient evidence that the days differ.

Why: This is chapter 9's discipline applied here. Failing to reject never establishes the null; it only reports that the data are consistent with it. The book's own phrasing is careful about this, and with sixty observations spread over five days the test would only have detected a fairly large departure anyway.

55. A goodness-of-fit test against a one-proportion test

Comparison

Fill the blanks. Both use counts; only one gives a single verdict on a whole distribution.

Comparison matrix

One-proportion testGoodness of fit
What is claimedone proportion's valuea whole distribution
Distribution usedthe standard normalchi-square
Tails availableleft, right or bothright only
Conditionnp and nq at least fiveevery expected count at least five

The last row shows the two conditions are the same idea. Both ensure enough is expected in each category for a continuous approximation to a count to be reasonable — chapter 7's requirement, appearing again in a new setting.

56. Running a goodness-of-fit test, in order

Pattern

Six steps, and the third must precede the fourth.

  1. State the hypotheses in sentences: the data fit the distribution, against they do not.
  2. Build the expected counts by multiplying each claimed proportion by the sample total, and check they sum to n.
  3. Check that every expected count is at least five; combine adjacent cells where it is not.
  4. Count the cells after combining and subtract one for the degrees of freedom.
  5. Sum observed minus expected, squared, over expected across the cells.
  6. Take the right-tail area beyond the statistic, compare with alpha, decide, and conclude in context.

Compare the statistic with the degrees of freedom before looking up anything: well above means a small p-value, at or below means a large one.

OpenStax Introductory Business Statistics 2e, §11.3 Goodness-of-Fit Test §11.3 Goodness-of-Fit Test

57. Check yourself 1 of 3

Check

Degrees of freedom.

Check your understanding

A goodness-of-fit test uses six categories and 400 observations. After checking the condition, no cells needed combining. What are the degrees of freedom?

  • A. 5 (correct)
  • B. 399
  • C. 6
  • D. 394

Answer: A

Why: The degrees of freedom are the number of categories minus one, so six categories give five — whatever the sample size.

Why B tempts people
That is n minus one, which the book explicitly warns against here.
Why C tempts people
That is the number of categories, before subtracting the one constraint.
Why D tempts people
That subtracts the categories from the observations, which matches no rule.

58. Check yourself 2 of 3

Check

The condition.

Check your understanding

A cell has an expected count of 3 and an observed count of 9. What must be done?

  • A. Combine it with an adjacent cell and reduce the df (correct)
  • B. Nothing: the observed count exceeds five
  • C. Drop the cell from the analysis
  • D. Increase the sample size

Answer: A

Why: The condition applies to expected counts, and three falls short. Combining with a neighbour fixes it and reduces the number of cells by one.

Why B tempts people
The condition is stated for expected values, not observed ones.
Why C tempts people
Dropping a cell discards its observations, so the expected counts would no longer sum to n.
Why D tempts people
That would be a design change, not something available once the data are collected.

59. Check yourself 3 of 3

Check

Interpretation.

Check your understanding

A goodness-of-fit test yields a statistic of 2 on 8 degrees of freedom. What follows?

  • A. The observed counts agree closely with expectation (correct)
  • B. The data strongly contradict the claimed distribution
  • C. The test was run incorrectly
  • D. The p-value is below 0.05

Answer: A

Why: The statistic is well below its mean of 8, so the departures are small and the right-tail p-value is large — about 0.98 here.

Why B tempts people
That would require a statistic well ABOVE the degrees of freedom.
Why C tempts people
A small statistic is a perfectly ordinary result; it means good agreement.
Why D tempts people
The p-value is near 0.98, since almost the whole distribution lies above 2.

60. Where this shows up outside the textbook

Real world

An auditor checks whether the leading digits of 8,000 expense claims follow Benford's law, which predicts 30.1 percent ones down to 4.6 percent nines. The chi-square statistic is 62.4 on eight degrees of freedom, giving a p-value below 0.000001, and the auditor concludes the claims were fabricated.

Discussion prompt

Assess the statistical work and the conclusion separately.

Hint: Ask what the test establishes, and what it does not.

Answer:

The statistical work is sound. Nine leading digits give nine cells and eight degrees of freedom, and the smallest expected count is 4.6 percent of 8,000, which is 368 — far above the required five. A statistic of 62.4 against a mean of 8 is about 13 standard deviations out, so the tiny p-value is exactly what the numbers predict.

\[ \text{df} = 9 - 1 = 8, \qquad E_9 = 0.046(8000) = 368, \qquad \chi^2 = 62.4 \]

But the conclusion goes far beyond what the test supports. What was rejected is that the digits follow Benford's law. Fabrication is one explanation; so are a spending policy with reimbursement caps that cluster claims near round thresholds, a currency conversion applied to some claims, a category of fixed-price items, or simply a population whose values span too narrow a range for Benford's law to apply at all — the law describes quantities ranging over several orders of magnitude, which expense claims often do not.

The sample size deserves attention too. With 8,000 observations the test detects departures far too small to matter practically. A distribution differing from Benford's by a percentage point or two in each cell would be rejected decisively while being consistent with entirely ordinary spending. That is chapter 9's distinction between statistical and practical significance, and it bites hard here because n is large.

The right report says the digits depart significantly from Benford's law, names the cells driving the departure, notes how large that departure is in percentage-point terms, and treats the finding as a reason to look further rather than as a conclusion. This is close to how digit analysis is actually used in auditing — as a screen that directs attention, never as evidence of wrongdoing on its own.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why is a goodness-of-fit test right-tailed rather than left- or two-tailed?

  • Because the alternative is always worded as a difference
  • Because every departure is squared, so all disagreement makes the statistic large
  • Because chi-square has no left tail
  • By convention, to simplify tables

Correct: Because every departure is squared.

\[ \frac{(O-E)^2}{E} \ge 0 \text{ in every cell} \;\Longrightarrow\; \text{misfit only ever increases } \chi^2 \]

Why: A chi-square statistic discards the sign of each difference, so a cell running above expectation and one running below both push it upward. Poor fit can only make it large, and a small statistic means the data agree with the claim — which is never evidence against the null. The chi-square distribution does have a lower tail; it simply carries no evidence for this question.

62. Explain it to someone a year behind you

Explain it

They ran Example 11.1 without combining cells and got df = 4.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask them to check the smallest expected count first.

Answer:

The 12-or-more cell expects only two students, below the required five, so it has to be combined with the 9-to-11 cell before the test can be run at all.

That leaves four cells rather than five, so the degrees of freedom are three — the check has to happen before the count.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Building expected counts from a set of claimed percentages
  • Checking the condition and combining cells correctly
  • Counting degrees of freedom after a merge
  • Explaining why the test is right-tailed

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, multiply each proportion by n and check the total is n. For the second, look at expected counts only, and merge adjacent cells. For the third, count the cells you actually ended with and subtract one. For the fourth, every term is squared, so misfit only pushes the statistic up. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, write the test statistic as a sum over cells of observed minus expected, squared, over expected, and label O, E and k beside it. Under it write the degrees of freedom rule — categories minus one — and beside that write in capitals that it is NOT n minus one, with Example 11.3's 600 families as the reminder. In the middle of the page, work Example 11.1 in full: the original five categories with their expected and observed counts, an arrow showing the last two merging because 2 is below 5, the new four-cell table, the four contributions 4.500, 3.333, 5.333 and 1.125, their total of 14.29, df of 3 and p of 0.0025. To its right, draw a chi-square curve on four degrees of freedom with a vertical line at 3 and the area to its RIGHT shaded, labelled 0.5578, and write one sentence saying why a statistic below its mean gives a p-value above one half. At the bottom, list the six steps in order and circle the arrow from step three to step four, noting that combining changes the cell count.

Check your Example 11.1 work by confirming both the expected and observed columns still total 100 after the merge. Check your curve by confirming the shaded region covers more than half the area, which is what a p-value above 0.5 has to look like.

65. What you can do now

Recap

Six things, and two of them are checks rather than calculations.

If you seeThen
A claimed distribution as percentagesMultiply each by n to get expected counts
A claim of equal likelihoodDivide n by the number of categories
An expected count below fiveCombine that cell with an adjacent one
Cells combinedRecount them before finding the degrees of freedom
A statistic well above its dfA small p-value: poor fit
A statistic at or below its dfA large p-value: good agreement
A rejected testReport the distribution, then describe the dominant cells separately

Section 11.3 keeps the same statistic and changes only where the expected counts come from. Instead of a claimed distribution supplying them, they are computed from a two-way table's own row and column totals — which tests whether two variables are independent.

OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-573 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 565-573
  2. OpenStax Introductory Business Statistics 2e, §11.3 Goodness-of-Fit Test — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108