The chi-square distribution's first use, and the chapter's first genuinely new question: instead of testing a single mean or proportion, a goodness-of-fit test asks whether a whole set of observed counts is consistent with a claimed distribution. The statistic sums observed minus expected, squared, over expected across the cells, so a departure in any cell and in either direction makes it larger — which is why the test is almost always right-tailed and why small statistics mean good agreement. The degrees of freedom are the number of categories minus one, never the number of observations minus one, and every expected count must be at least five, a condition that in the book's first example forces two categories to be combined before the test can be run at all. Worked through the book's four examples: absenteeism, absences by weekday, streaming services and a pair of coins.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 11 — The Chi-Square Distribution
Goodness-of-Fit Test
Objectives
Six outcomes. Two of them are conditions to check rather than calculations to do.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-573 — the section these objectives are drawn from
Warm-up
Chapter 9 tested one parameter at a time; section 11.1 gave a distribution built from squares.
Discussion prompt
A die is rolled 60 times and lands on each face 12, 8, 11, 9, 13 and 7 times. How would you test whether it is fair, using only chapter 9's tools?
Hint: Chapter 9 could test one proportion. How many would you need here?
Answer:
Chapter 9 could test each face separately — six tests, one per proportion, each asking whether that face's rate differs from one sixth. But six tests answer six questions, not one, and running all of them inflates the chance that at least one comes out significant by accident.
What is actually wanted is a single verdict on the whole set of counts at once. That needs a statistic that combines every cell's departure into one number, and section 11.1 supplies the shape: squaring each departure makes them all positive so they accumulate rather than cancel, and the sum of squares has a chi-square distribution.
This section builds exactly that. Six counts against six expectations become one statistic and one p-value, answering the question that was actually asked.
Concept
In this type of hypothesis test you determine whether the data fit a particular distribution or not. The test statistic sums, across the cells, the squared difference between observed and expected divided by expected. The degrees of freedom are the number of categories minus one.
a goodness-of-fit test — A chi-square test of whether observed counts are consistent with the distribution a hypothesis claims. The null hypothesis is that the data fit; the alternative is that they do not.
\[ \chi^2 = \sum_{i=1}^{k} \frac{(O - E)^2}{E}, \qquad \text{df} = k - 1 \]
The book's phrasing of the hypotheses is worth copying: they may be written in sentences or as equations or inequalities, and in practice sentences are clearer. Example 11.1's are simply that student absenteeism fits faculty perception, against that it does not — no parameter is named, because the claim is about a whole distribution rather than any single number.
Figure (svg): A card breaking the goodness-of-fit statistic into its parts
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 565
Section
Section 1
Concept
The observed values are the data and the expected values are what you would expect to get if the null hypothesis were true. Each cell contributes a squared difference scaled by its expected count, and the degrees of freedom are the number of categories minus one.
scaling by E — Dividing by the expected count makes a discrepancy of ten matter more when only twelve were expected than when three hundred were. Without it, large cells would dominate every test.
\[ \frac{(O-E)^2}{E} \quad \text{per cell}, \qquad \text{df} = (\text{number of categories}) - 1 \]
The book adds a note that is easy to skip past: df is not 600 minus 1 in Example 11.3, even though 600 families were surveyed. The degrees of freedom count the cells, and with five categories of streaming service the answer is four however many families were asked. Nothing about the sample size enters the count.
Figure (svg): A card breaking the goodness-of-fit statistic into its parts
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-571 — the statistic, the df rule, and the note that df is not n minus one
Picture it
What each symbol means, and the one condition attached.
Figure (svg): A card breaking the goodness-of-fit statistic into its parts
Every term is a square over a positive count, so no term can be negative and none can cancel another. That is why the statistic only grows as the data depart from expectation, whatever direction the departures run in.
Worked example
Sixty managers named the day with the most employee absences: 15, 12, 9, 9 and 15 across Monday to Friday. Do the days occur with equal frequencies? Test at 5 percent.
\[ O = 15, 12, 9, 9, 15 \]
Total the observations
Why: Fifteen plus twelve plus nine plus nine plus fifteen.
\[ 60 \]
Split it equally
Why: Sixty over five days.
\[ E = 12\text{ each} \]
Sum the terms
Why: Nine, zero, nine, nine and nine, each over twelve.
\[ 3.00 \]
Count the cells
Why: Five days minus one.
\[ d f = 4 \]
Find the right tail
Why: Beyond 3 on four df.
\[ 0.5578 \]
Figure (svg): The solution to Worked example Example 11.2, absences by weekday shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 3, \quad \text{df} = 4, \quad p = 0.5578 \]
Verify: confirm the statistic is plausible before consulting any table
Why: The mean of a chi-square with four degrees of freedom is 4, and the statistic here is 3 — below its own mean, so more than half the distribution lies above it and the p-value must exceed one half. It does, at 0.5578. The book's conclusion follows: at a 5 percent level there is not sufficient evidence to conclude that the absent days do not occur with equal frequencies.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 568-569
Faded example
Example 11.3's last cell observed 15 against an expected 48.
Fill in the blanks
\frac108922.69 = \frac___}___ = ___
Why: That single cell contributes 22.69 of the total 29.65 — over three quarters of the statistic. Looking at which cell dominates is how a rejected goodness-of-fit test gets interpreted.
Worked example
Two coins are flipped 100 times, giving 20 HH, 27 HT, 30 TH and 23 TT. Test at 5 percent.
\[ 20\;HH,\; 27\;HT,\; 30\;TH,\; 23\;TT \]
Define the variable
Why: X counts heads in one flip of the two.
\[ X = 0, 1, 2 \]
Collapse to three cells
Why: HT and TH both give one head.
\[ 20, 57, 23 \]
Expect on fairness
Why: A quarter, a half, a quarter.
\[ 25, 50, 25 \]
Sum the terms
Why: One plus 0.98 plus 0.16.
\[ 2.14 \]
Right tail on df 2
Why: Three cells minus one.
\[ 0.3430 \]
Figure (svg): The solution to Worked example Example 11.4, are the coins fair shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 2.14, \quad \text{df} = 2, \quad p = 0.3430 \]
Verify: confirm the cell count, which is where this example is easy to get wrong
Why: The sample space has four outcomes but the random variable takes only three values, since HT and TH are both one head. So there are three cells and two degrees of freedom, not four cells and three. Using four cells would put 27 and 30 against expectations of 25 each — a different statistic on different degrees of freedom, and the wrong answer to the question actually asked about the number of heads.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 573
Trap
\[ 600 \text{ families in 5 categories} \;\Rightarrow\; \text{df} = 599 \]
Reach for the familiar n minus one
Why: It was right for every t test in chapters 8 through 10.
\[ \text{but the cells are what is being counted} \]
The book flags this directly with a note reading df is not 600 minus 1.
\[ \text{df} = k - 1 = 5 - 1 = 4 \]
Count categories, not observations
Why: The sample size never enters the degrees of freedom here.
Sample size does matter, but through the expected counts rather than the degrees of freedom. Surveying 6,000 families instead of 600 would multiply every expected count by ten and make real departures far easier to detect — while leaving the degrees of freedom at four.
Two truths and a lie
All three concern the statistic.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Each term is squared before being added, so every departure contributes positively whatever its direction. That is exactly why the test detects any kind of misfit and why the statistic can only be large when the fit is poor.
Prediction
Commit before reasoning.
Predict first
What does dividing each squared difference by its expected count accomplish?
Correct: It judges each discrepancy relative to its expected count.
Why: Missing by ten is serious when twelve were expected and trivial when three hundred were. Dividing by E puts every cell on a comparable footing, which is what lets a small cell and a large one contribute meaningfully to the same total.
Estimation
Example 11.3: observed 66, 119, 340, 60, 15 against expected 60, 96, 330, 66, 48.
Predict first
Which cell contributes most to the statistic of 29.65?
Correct: The last: 15 against 48.
Why: Its contribution is 22.69, against 5.51 for the second cell and only 0.30 for the third. The third cell has the largest counts but the smallest relative miss, which is exactly what dividing by E is designed to capture.
Section
Section 2
Concept
The goodness-of-fit test is almost always right-tailed. If the observed values and the corresponding expected values are not close to each other, then the test statistic can get very large and will be way out in the right tail of the chi-square curve.
right-tailed by construction — The statistic measures disagreement, so poor fit can only make it large. A small statistic means the observations sit close to expectation, which is never evidence against the null.
\[ p\text{-value} = P(\chi^2 > \chi^2_{\text{obs}}) \]
This is a real break from chapter 9, where the direction of the tail was a decision that followed from how the alternative hypothesis was worded. Here it follows from the statistic's construction instead: because every departure is squared, disagreement in any direction pushes the statistic the same way, so there is no left-tailed version of the question to ask.
Figure (svg): A chi-square curve on four degrees of freedom with the area to the right of three shaded, covering more than half the distribution
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-570 — almost always right-tailed; this test is always right-tailed
Picture it
A statistic of 3 against a distribution with mean 4.
Figure (svg): A chi-square curve on four degrees of freedom with the area to the right of three shaded, covering more than half the distribution
The shaded region is over half the curve, which is what a p-value of 0.5578 means. The book asks for exactly this picture, properly labelled with the right tail shaded, as part of the solution.
Worked example
Six hundred far-western families reported 66, 119, 340, 60 and 15 across five categories, against a national distribution of 10, 16, 55, 11 and 8 percent. Test at 1 percent.
\[ \alpha = 0.01 \]
Build expected counts
Why: Each percent times 600.
\[ 60, 96, 330, 66, 48 \]
Check the condition
Why: Smallest expected is 48.
Sum the terms
Why: By calculator.
\[ 29.65 \]
Degrees of freedom
Why: Five cells minus one.
\[ 4 \]
Right tail
Why: Beyond 29.65.
\[ 0.000006 \]
Figure (svg): The solution to Worked example Example 11.3, streaming services shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 29.65, \quad p = 0.000006 < 0.01 \]
Verify: confirm the statistic is extreme relative to its degrees of freedom
Why: The mean of the distribution is 4 and its standard deviation is the root of 8, about 2.83 — so a statistic of 29.65 sits some nine standard deviations above the mean. A p-value in the millionths is exactly what that predicts, and the conclusion is that the far western distribution genuinely differs, driven overwhelmingly by the 4-or-more cell where 15 appeared against 48 expected.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571
Sorting
Each is paired with its degrees of freedom.
Sort into buckets
Sort by whether the result is evidence against the null.
Items (a), (b), (c) and (d) are Examples 11.3, 11.2, 11.1 and 11.4, with p-values of 0.000006, 0.5578, 0.0025 and 0.3430. Item (e) is a perfect fit.
Worked example
Suppose the far-western counts had been 60, 96, 330, 66, 48 exactly.
\[ O = E \text{ in every cell} \]
Each term
Why: Zero over its expected count.
\[ 0 \]
The sum
Why: Five zeros.
\[ 0 \]
The right tail beyond 0
Why: The whole distribution.
\[ p = 1 \]
Interpret
Why: Perfect agreement.
Figure (svg): The solution to Worked example what a small statistic would mean shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 0 \;\Longrightarrow\; p = 1 \]
Verify: confirm this is why a left tail would be meaningless
Why: The statistic's smallest possible value is zero, and it means the observations equal the expectations exactly. There is no way for data to be worse than fitting perfectly, so the lower end of the scale carries no evidence against the null and no left-tailed test exists. That is a structural fact about the statistic rather than a choice about how to word the alternative.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 565
Error analysis
Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; \text{a small statistic is evidence against the null} \\ &(2)\; \text{the p-value is the area to the right of the statistic} \\ &(3)\; \text{the direction is chosen from the alternative's wording} \\ &(4)\; \text{a large statistic means observed and expected are far apart} \end{aligned} \)
Error (3) is the one worth naming, because chapter 9 spent a whole section on choosing the tail. That habit has to be dropped: every test in sections 11.2 to 11.4 is right-tailed regardless of how the alternative is phrased.
Two truths and a lie
All three concern the direction of the test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Cells with observed below expected contribute positively to the statistic just as cells above do, since the difference is squared. There is no left-tailed version because there is no way for data to disagree in a way that makes the statistic small.
Prediction
Commit before reasoning.
Predict first
In chapter 9 the tail was chosen from the alternative hypothesis. Why not here?
Correct: Every departure is squared, so all disagreement lands on the right.
Why: A z or t statistic keeps the sign of the departure, so the tail had to match the alternative's direction. A chi-square statistic discards the sign by squaring, so both directions of departure push it the same way and only the upper tail carries evidence.
Explain it
A classmate has drawn Example 11.2's graph and shaded the area to the LEFT of 3.
Discussion prompt
In two sentences or fewer, tell them what is wrong.
Hint: Ask what the p-value is meant to measure.
Answer:
The p-value is the chance of a statistic at least as extreme as the one observed, and extreme means large here — so the shading belongs to the right of 3, not the left.
Shading the left would give about 0.44 instead of 0.5578, and would treat a good fit as though it were evidence against the null.
Section
Section 3
Concept
The expected values are the values you would expect to get if the null hypothesis were true. When the claim is a set of percentages, multiply each by the sample total; when the claim is that categories are equally likely, divide the total by the number of categories.
expected counts — Not the observed data, and not necessarily whole numbers. They are what the hypothesised distribution predicts for a sample of this size, and they always sum to the observed total.
\[ E_i = n \cdot p_i, \qquad \sum E_i = n \]
The check that the expected counts total the sample size is the fastest way to catch an arithmetic slip. In Example 11.3 they total 60 plus 96 plus 330 plus 66 plus 48, which is 600 — the number of families surveyed. If they do not add to n, either a percentage was misread or the multiplication went wrong, and the statistic will be meaningless.
Figure (svg): A table converting the claimed percentages into expected frequencies for six hundred families
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571 — converting percentages to expected frequencies
Picture it
Example 11.3's conversion, alongside what was actually observed.
Figure (svg): A table converting the claimed percentages into expected frequencies for six hundred families
The book's advice about letting the calculator do the arithmetic — entering 0.10 times 600 rather than 60 — matters when the percentages produce awkward decimals. Rounding expected counts before summing is a common source of small discrepancies.
Worked example
Example 11.2's expected counts, from the claim of a uniform distribution.
\[ 60 \text{ absences over } 5 \text{ days} \]
Total the observations
Why: 15 + 12 + 9 + 9 + 15.
\[ 60 \]
State the claim
Why: Equal frequencies.
Multiply
Why: Sixty times one fifth.
\[ 12 \]
Repeat for each cell
Why: All the same here.
\[ 12, 12, 12, 12, 12 \]
Figure (svg): The solution to Worked example equal frequencies shown as a ladder of expressions, one row per legal move
\[ E = 12 \text{ for each of the five days} \]
Verify: confirm the expected counts sum to the observed total
Why: Five twelves are 60, which matches the total number of absences reported. That check applies to every goodness-of-fit test regardless of the claimed distribution, since the expected counts are the total split according to the hypothesised proportions and those proportions sum to one.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 568
Faded example
A claimed 24 percent, in a sample of 250.
Fill in the blanks
E = 0.24 \times 250 = 60
Why: Expected counts need not be whole numbers, though this one is. A claimed 23 percent of 250 would give 57.5, and that value is used as it stands rather than rounded.
Worked example
Example 11.3's national distribution applied to 600 families.
\[ 10\%, 16\%, 55\%, 11\%, 8\% \]
Convert to decimals
Why: Each percentage over 100.
\[ 0.10\text{ to } 0.08 \]
Multiply by 600
Why: One cell at a time.
\[ 60, 96, 330, 66, 48 \]
Check the total
Why: Sum them.
\[ 600 \]
Check the condition
Why: The smallest.
\[ 48,\text{ comfortably above } 5 \]
Figure (svg): The solution to Worked example percentages shown as a ladder of expressions, one row per legal move
\[ E = 60, 96, 330, 66, 48 \]
Verify: confirm the percentages themselves were complete
Why: Ten plus sixteen plus fifty-five plus eleven plus eight is one hundred, so the claimed distribution accounts for every family and nothing is missing. A set of percentages that fails to total 100 means a category has been left out, and the test cannot be run until it is found — the expected counts would not sum to n.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 570
Trap
\[ \frac{(11 - 10)^2}{10} + \frac{(19.8 - 16)^2}{16} + \cdots \]
Convert the observed data to percentages and compare those
Why: It seems to put the two on the same footing.
\[ \text{the statistic no longer depends on the sample size} \]
Six hundred families and six families would give the same statistic, which cannot be right — a departure seen in 600 observations is far stronger evidence.
\[ \frac{(66 - 60)^2}{60} + \frac{(119 - 96)^2}{96} + \cdots = 29.65 \]
Convert the claim to counts and compare counts
Why: The sample size then enters through the expected values.
This is the same principle as the standard error's dependence on n throughout chapters 8 to 10: evidence should strengthen with more data. Working in percentages throws that away, and a test built on percentages would reach the same verdict from six families as from six hundred.
Sorting
Each is an expected count in some cell.
Sort into buckets
Sort by whether the cell may be used as it stands.
Item (c) is exactly at the boundary and passes, since the requirement is at least five. Item (d) at 4.8 fails despite rounding to five — the condition applies to the expected count itself, not to a rounded version of it.
Two truths and a lie
All three concern expected counts.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. Expected counts come from the hypothesised distribution — the theory — and the only thing they take from the data is the total. That separation is what makes the comparison meaningful; if both sides came from the data they would always agree.
Prediction
Commit before reasoning.
Predict first
A claimed distribution's percentages sum to 92. What has gone wrong?
Correct: A category is missing.
Why: Every observation must fall in exactly one cell, so the claimed proportions must account for all of them. Percentages totalling 92 mean 8 percent of the population is unclassified, and the expected counts would fall short of the sample total — a discrepancy that would show up in the sum check.
Section
Section 4
Concept
The expected value for each cell needs to be at least five in order for you to use this test. When a cell falls short, it is combined with an adjacent one — which reduces the number of categories and so reduces the degrees of freedom.
combining cells — Merging adjacent categories so that each expected count reaches five. The observed counts are added together too, and the category labels are rewritten to describe the merged range.
\[ E_i \ge 5 \text{ for every } i \]
The reason behind the condition is the same one that sat behind chapter 7's normal approximation to the binomial. A chi-square distribution describes the statistic only approximately, and the approximation relies on each cell's count behaving roughly normally. With an expected count of two that is far from true, and the resulting p-value cannot be trusted.
Figure (svg): Two tables showing the absence categories before and after the last two are combined
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-567 — the note on expected counts, and Example 11.1's combining
Picture it
The absence categories, and why two of them merge.
Figure (svg): Two tables showing the absence categories before and after the last two are combined
The book asks first whether the test can be run on the charts as they appear, and the answer is no. Combining the last two categories gives a 9-or-more cell expecting 8 and observing 5, and reduces the cells from five to four so the degrees of freedom become 3.
Worked example
Faculty expected 50, 30, 12, 6 and 2 students in five absence bands; the survey observed 35, 40, 20, 1 and 4.
\[ E = 50, 30, 12, 6, 2 \]
Check every expected count
Why: The last is two.
Combine the last two
Why: Six plus two, one plus four.
\[ E = 8, O = 5 \]
Recheck the condition
Why: New smallest is eight.
Recount the cells
Why: Four now.
\[ d f = 4 - 1 = 3 \]
Figure (svg): The solution to Worked example Example 11.1, absenteeism shown as a ladder of expressions, one row per legal move
\[ \text{df} = \text{number of cells} - 1 = 4 - 1 = 3 \]
Verify: confirm the totals survive the combining
Why: The expected counts still total 100 — 50 plus 30 plus 12 plus 8 — and the observed still total 100 as well, at 35 plus 40 plus 20 plus 5. Merging cells redistributes counts without creating or destroying any, so both totals must be unchanged; if either has moved, the merge was done wrongly.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 566-567
Faded example
Six categories, two of which are merged into one.
Fill in the blanks
\text5 = 4, \qquad \text___ = ___
Why: Merging two cells into one leaves five, so the degrees of freedom are four rather than the five that the original six categories would have given.
Worked example
The book stops at the degrees of freedom, so here is the test it sets up.
\[ O = 35, 40, 20, 5; \; E = 50, 30, 12, 8 \]
First cell
Why: 225 over 50.
\[ 4.500 \]
Second cell
Why: 100 over 30.
\[ 3.333 \]
Third cell
Why: 64 over 12.
\[ 5.333 \]
Fourth cell
Why: 9 over 8.
\[ 1.125 \]
Sum, and take the right tail on df 3
Why: Add them.
\[ 14.29, p = 0.0025 \]
Figure (svg): The solution to Worked example completing Example 11.1 shown as a ladder of expressions, one row per legal move
\[ \chi^2 = 14.29, \quad \text{df} = 3, \quad p = 0.0025 \]
Verify: confirm the statistic is extreme for three degrees of freedom
Why: The mean is 3 and the standard deviation is the root of 6, about 2.45, so 14.29 sits about 4.6 standard deviations above the centre — comfortably consistent with a p-value of 0.0025. Reading the cells shows where the misfit lies: faculty expected half the students in the lowest band and only 35 were there, while the 3-to-5 and 6-to-8 bands both ran well above expectation. Students were absent more than faculty believed.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 566-567
Error analysis
A cell has an expected count of 3. Which responses are legitimate?
Annotate
On: \( \begin{aligned} &(1)\; \text{combine it with an adjacent cell} \\ &(2)\; \text{drop the cell from the analysis} \\ &(3)\; \text{run the test anyway and note the caveat} \\ &(4)\; \text{combine, but keep the original degrees of freedom} \end{aligned} \)
Error (4) is worth dwelling on because it produces a plausible-looking answer. Example 11.1 with four degrees of freedom instead of three gives a p-value of 0.0064 rather than 0.0025 — same decision here, but a different number, and the difference grows as more cells are merged.
Two truths and a lie
All three concern the condition.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false, and Example 11.1 shows why the distinction matters: its 9-to-11 cell observed only 1, but its expected count was 6, so that cell was not the problem. The 12-or-more cell was combined because its EXPECTED count was 2.
Prediction
Commit before reasoning.
Predict first
Categories run 0 to 2, 3 to 5, 6 to 8, 9 to 11 and 12 or more. Which merge is sensible?
Correct: 9 to 11 with 12 or more.
Why: The categories are ordered, so merging adjacent ones produces a meaningful new category — 9 or more. Merging the two extremes would create a category describing students who were absent either very little or very much, which answers no sensible question and would be impossible to interpret if the test rejected.
Estimation
Example 11.1's statistic of 14.29 on the correct df of 3 gives p = 0.0025.
Predict first
Keeping df at 4 by mistake would give roughly what p-value?
Correct: About 0.006.
Why: More degrees of freedom shift the curve right, so the same statistic sits less far out and the tail area grows — here from 0.0025 to 0.0064. The decision is unchanged at 5 percent, but the reported evidence is overstated by more than a factor of two, and with more merged cells the gap widens.
Section
Section 5
Concept
The overall shape is the familiar one: hypotheses, a statistic, a p-value, a comparison with alpha, a decision and a conclusion in context. What is new sits in the middle — building the expected counts, checking the condition, and counting the cells.
the conclusion in context — Stated in the language of the problem rather than in symbols, and about fit rather than about a parameter. The book's conclusions name the distribution being tested every time.
\[ H_0: \text{the data fit} \quad\text{against}\quad H_a: \text{the data do not fit} \]
One feature of the hypotheses is worth noticing. There is no parameter, so there is nothing to write an equation about — which is why the book says the hypotheses may be written in sentences. A goodness-of-fit test's null is a statement about a whole distribution, and sentences state that more honestly than any equation would.
Figure (svg): The order of operations for a goodness-of-fit test
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-570 — the hypotheses, and Example 11.3's full write-up
Picture it
Six steps, and the reason step three precedes step four.
Figure (svg): The order of operations for a goodness-of-fit test
The ordering is not cosmetic. Combining cells to satisfy the condition changes how many cells there are, so counting degrees of freedom before the check has been done gives the wrong answer — which is precisely the mistake Example 11.1 is designed to prevent.
Worked example
For Example 11.3's streaming services.
\[ H_0 \text{ and } H_a \]
Name the two distributions
Why: Far western families, and the nation.
The null
Why: They agree.
The alternative
Why: The opposite.
Note what is absent
Why: No parameter.
Figure (svg): The solution to Worked example writing the hypotheses shown as a ladder of expressions, one row per legal move
\[ H_0: \text{same distribution}; \quad H_a: \text{different} \]
Verify: confirm the alternative really is the negation of the null
Why: Chapter 9 established that the two hypotheses must partition every possibility, and they do here: either the distributions agree in every category or they differ in at least one. Note that the alternative does not say HOW they differ, which is consistent with the test being right-tailed — the statistic detects any departure without identifying its direction.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, p. 571
Matching
Match each stage of Example 11.3's write-up to what it contains.
Match the pairs
Why: The sequence is identical to chapter 9's. What changed is the content of the second stage — a chi-square with cell-based degrees of freedom rather than a t with n minus one.
Worked example
Example 11.3, in the book's own sequence.
\[ \alpha = 0.01 \]
Hypotheses
Why: In sentences.
Distribution for the test
Why: Chi-square.
\[ d f = 5 - 1 = 4 \]
Test statistic
Why: By calculator.
\[ 29.65 \]
Probability statement
Why: A right tail.
\[ p = 0.000006 \]
Compare and decide
Why: Alpha exceeds the p-value.
Figure (svg): The solution to Worked example the full write-up shown as a ladder of expressions, one row per legal move
\[ p = 0.000006 < 0.01 = \alpha \]
Verify: confirm the conclusion is stated in the problem's own terms
Why: The book's wording names the variable and both populations rather than saying only that the null was rejected. That practice matters more here than in chapter 9, since a goodness-of-fit test's null has no parameter — a reader who is only told the null was rejected has no way to reconstruct what claim was being tested.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 570-571
Trap
\[ p = 0.000006 \;\Rightarrow\; \text{far western families have fewer with 4+ services} \]
Read a specific cell's story out of a rejected test
Why: The 4-or-more cell did dominate the statistic.
\[ \text{but the test's alternative names no cell} \]
The test rejects the whole distribution, and the alternative hypothesis is only that the distributions differ.
\[ \text{the distributions differ; the 4+ cell contributed } 22.69 \text{ of } 29.65 \]
Report the verdict the test gives, then describe the cells separately
Why: The contributions are informative but are not what was tested.
The distinction is the same one chapter 9 drew between a test's conclusion and the description that motivated it. Naming the cell that dominated is good practice and helps a reader understand the result; presenting it as the test's own conclusion overstates what a single p-value licenses.
Faded example
Example 11.2, at a 5 percent level.
Fill in the blanks
p = 0.5578, \quad \alpha = 0.05: \quad p > \alpha, \textdo not reject ___ H_0
Why: The p-value far exceeds alpha, so the decision is not to reject. The book's conclusion is that there is not sufficient evidence to conclude that the absent days do not occur with equal frequencies — a double negative that is worth reading twice.
Two truths and a lie
All three concern reporting.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The test delivers one verdict on the whole distribution. Examining individual cell contributions is a sensible follow-up and often illuminating, but it is a description of the data rather than a conclusion the test itself supports.
Prediction
Commit before reasoning.
Predict first
Example 11.2's p-value is 0.5578. What may be concluded?
Correct: There is not sufficient evidence that the days differ.
Why: This is chapter 9's discipline applied here. Failing to reject never establishes the null; it only reports that the data are consistent with it. The book's own phrasing is careful about this, and with sixty observations spread over five days the test would only have detected a fairly large departure anyway.
Comparison
Fill the blanks. Both use counts; only one gives a single verdict on a whole distribution.
Comparison matrix
| One-proportion test | Goodness of fit | |
|---|---|---|
| What is claimed | one proportion's value | a whole distribution |
| Distribution used | the standard normal | chi-square |
| Tails available | left, right or both | right only |
| Condition | np and nq at least five | every expected count at least five |
The last row shows the two conditions are the same idea. Both ensure enough is expected in each category for a continuous approximation to a count to be reasonable — chapter 7's requirement, appearing again in a new setting.
Pattern
Six steps, and the third must precede the fourth.
Compare the statistic with the degrees of freedom before looking up anything: well above means a small p-value, at or below means a large one.
OpenStax Introductory Business Statistics 2e, §11.3 Goodness-of-Fit Test §11.3 Goodness-of-Fit Test
Check
Degrees of freedom.
Check your understanding
A goodness-of-fit test uses six categories and 400 observations. After checking the condition, no cells needed combining. What are the degrees of freedom?
Answer: A
Why: The degrees of freedom are the number of categories minus one, so six categories give five — whatever the sample size.
Check
The condition.
Check your understanding
A cell has an expected count of 3 and an observed count of 9. What must be done?
Answer: A
Why: The condition applies to expected counts, and three falls short. Combining with a neighbour fixes it and reduces the number of cells by one.
Check
Interpretation.
Check your understanding
A goodness-of-fit test yields a statistic of 2 on 8 degrees of freedom. What follows?
Answer: A
Why: The statistic is well below its mean of 8, so the departures are small and the right-tail p-value is large — about 0.98 here.
Real world
An auditor checks whether the leading digits of 8,000 expense claims follow Benford's law, which predicts 30.1 percent ones down to 4.6 percent nines. The chi-square statistic is 62.4 on eight degrees of freedom, giving a p-value below 0.000001, and the auditor concludes the claims were fabricated.
Discussion prompt
Assess the statistical work and the conclusion separately.
Hint: Ask what the test establishes, and what it does not.
Answer:
The statistical work is sound. Nine leading digits give nine cells and eight degrees of freedom, and the smallest expected count is 4.6 percent of 8,000, which is 368 — far above the required five. A statistic of 62.4 against a mean of 8 is about 13 standard deviations out, so the tiny p-value is exactly what the numbers predict.
\[ \text{df} = 9 - 1 = 8, \qquad E_9 = 0.046(8000) = 368, \qquad \chi^2 = 62.4 \]
But the conclusion goes far beyond what the test supports. What was rejected is that the digits follow Benford's law. Fabrication is one explanation; so are a spending policy with reimbursement caps that cluster claims near round thresholds, a currency conversion applied to some claims, a category of fixed-price items, or simply a population whose values span too narrow a range for Benford's law to apply at all — the law describes quantities ranging over several orders of magnitude, which expense claims often do not.
The sample size deserves attention too. With 8,000 observations the test detects departures far too small to matter practically. A distribution differing from Benford's by a percentage point or two in each cell would be rejected decisively while being consistent with entirely ordinary spending. That is chapter 9's distinction between statistical and practical significance, and it bites hard here because n is large.
The right report says the digits depart significantly from Benford's law, names the cells driving the departure, notes how large that departure is in percentage-point terms, and treats the finding as a reason to look further rather than as a conclusion. This is close to how digit analysis is actually used in auditing — as a screen that directs attention, never as evidence of wrongdoing on its own.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why is a goodness-of-fit test right-tailed rather than left- or two-tailed?
Correct: Because every departure is squared.
\[ \frac{(O-E)^2}{E} \ge 0 \text{ in every cell} \;\Longrightarrow\; \text{misfit only ever increases } \chi^2 \]
Why: A chi-square statistic discards the sign of each difference, so a cell running above expectation and one running below both push it upward. Poor fit can only make it large, and a small statistic means the data agree with the claim — which is never evidence against the null. The chi-square distribution does have a lower tail; it simply carries no evidence for this question.
Explain it
They ran Example 11.1 without combining cells and got df = 4.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask them to check the smallest expected count first.
Answer:
The 12-or-more cell expects only two students, below the required five, so it has to be combined with the 9-to-11 cell before the test can be run at all.
That leaves four cells rather than five, so the degrees of freedom are three — the check has to happen before the count.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, multiply each proportion by n and check the total is n. For the second, look at expected counts only, and merge adjacent cells. For the third, count the cells you actually ended with and subtract one. For the fourth, every term is squared, so misfit only pushes the statistic up. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, write the test statistic as a sum over cells of observed minus expected, squared, over expected, and label O, E and k beside it. Under it write the degrees of freedom rule — categories minus one — and beside that write in capitals that it is NOT n minus one, with Example 11.3's 600 families as the reminder. In the middle of the page, work Example 11.1 in full: the original five categories with their expected and observed counts, an arrow showing the last two merging because 2 is below 5, the new four-cell table, the four contributions 4.500, 3.333, 5.333 and 1.125, their total of 14.29, df of 3 and p of 0.0025. To its right, draw a chi-square curve on four degrees of freedom with a vertical line at 3 and the area to its RIGHT shaded, labelled 0.5578, and write one sentence saying why a statistic below its mean gives a p-value above one half. At the bottom, list the six steps in order and circle the arrow from step three to step four, noting that combining changes the cell count.
Check your Example 11.1 work by confirming both the expected and observed columns still total 100 after the merge. Check your curve by confirming the shaded region covers more than half the area, which is what a p-value above 0.5 has to look like.
Recap
Six things, and two of them are checks rather than calculations.
| If you see | Then |
|---|---|
| A claimed distribution as percentages | Multiply each by n to get expected counts |
| A claim of equal likelihood | Divide n by the number of categories |
| An expected count below five | Combine that cell with an adjacent one |
| Cells combined | Recount them before finding the degrees of freedom |
| A statistic well above its df | A small p-value: poor fit |
| A statistic at or below its df | A large p-value: good agreement |
| A rejected test | Report the distribution, then describe the dominant cells separately |
Section 11.3 keeps the same statistic and changes only where the expected counts come from. Instead of a claimed distribution supplying them, they are computed from a two-way table's own row and column totals — which tests whether two variables are independent.
OpenStax Introductory Statistics 2e, §11.2 Goodness-of-Fit Test §11.2, pp. 565-573 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.