11.1 Facts About the Chi-Square Distribution

A short section introducing the distribution the whole chapter runs on. A chi-square random variable with k degrees of freedom is the sum of k independent squared standard normal variables, and every property the book lists follows from that definition. Because the terms are squares the statistic is never negative; because there are k of them the mean is k and the standard deviation is the square root of twice k; and because it is a sum of many independent quantities, the central limit theorem makes the curve approximately normal once the degrees of freedom exceed about ninety. The curve is nonsymmetrical and skewed to the right, there is a different curve for every degrees of freedom, and the mean sits just to the right of the peak. How the degrees of freedom are counted depends on which of the chapter's three tests is being run.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 11.1 Facts About the Chi-Square Distribution

Title

Statistics · Chapter 11 — The Chi-Square Distribution

Facts About the Chi-Square Distribution

2. By the end of this lesson you can

Objectives

Five outcomes, and all five are consequences of one definition.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565 — the section these objectives are drawn from

3. What you already have

Warm-up

Chapter 6 gave the standard normal; chapter 7 gave the behaviour of sums of independent quantities.

Discussion prompt

Take a standard normal variable and square it. What values can the result take, and what is its average?

Hint: A square is never negative, and the average of a squared standard normal is its variance.

Answer:

Only values at or above zero, since squaring destroys the sign. So the resulting distribution is bounded on the left at zero and can extend indefinitely to the right — which is already a description of a skewed shape.

Its average is 1, because the variance of a standard normal is 1 and the variance of a mean-zero variable is exactly the average of its square. So one squared standard normal averages 1, and k of them average k.

That is the chi-square distribution with k degrees of freedom, and its whole character — non-negative, right-skewed, with mean k — has just been derived without anything new. This section states five facts, and every one of them follows from that construction.

4. A sum of squared standard normals

Concept

The random variable for a chi-square distribution with k degrees of freedom is the sum of k independent squared standard normal variables. The mean of the distribution is the degrees of freedom, and the standard deviation is the square root of twice the degrees of freedom.

the chi-square distribution — Written with the Greek letter chi squared, and indexed by degrees of freedom that depend on how it is being used. Its random variable may be written with any upper-case letter.

\[ \chi^2 = Z_1^2 + Z_2^2 + \cdots + Z_k^2, \qquad \mu = \text{df}, \quad \sigma = \sqrt{2\,\text{df}} \]

The book notes that the degrees of freedom depend on how chi-square is being used, and that the three major uses each calculate them differently — a goodness-of-fit test counts categories, a test of independence multiplies two dimensions, and a test of a single variance uses n minus one. That variety is why the definition is worth holding on to: it is the one thing common to all three.

Figure (svg): A card giving the definition of a chi-square variable as a sum of squared standard normals

Nothing about the chi-square distribution has to be memorised separately once its definition is understood.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

5. The definition and what follows

Section

Section 1

6. Squares, and k of them

Concept

Two features of the definition do all the work. The terms are squares, so nothing can be negative; and there are k independent terms, so the mean is k and the sum behaves like any sum of independent quantities.

what follows from squaring — The statistic is never below zero, and the distribution is bounded on the left while unbounded on the right — which is what makes it skewed.

\[ Z_i^2 \ge 0 \;\Longrightarrow\; \chi^2 \ge 0 \]

The mean deserves a moment. A standard normal has mean zero and variance one, and variance is the average squared deviation from the mean — so the average of Z squared is exactly 1. Adding k independent such terms gives a mean of k, which is why the mean equals the degrees of freedom rather than being some separate fact to remember.

Figure (svg): A card giving the definition of a chi-square variable as a sum of squared standard normals

Nothing about the chi-square distribution has to be memorised separately once its definition is understood.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564 — the definition and the mean and standard deviation

7. One definition, three consequences

Picture it

What squaring gives, and what summing gives.

Figure (svg): A card giving the definition of a chi-square variable as a sum of squared standard normals

Nothing about the chi-square distribution has to be memorised separately once its definition is understood.

The third box is chapter 7's central limit theorem applied here: a chi-square variable is a sum of independent quantities, so for a large enough number of them it is approximately normal. That is the book's fourth fact, and it needs no separate justification.

8. Worked example: the mean and standard deviation

Worked example

For a chi-square with 15 degrees of freedom.

\[ \text{df} = 15 \]

The mean

Why: Equal to df.

\[ 15 \]

Twice the df

Why: Thirty.

\[ 30 \]

Take the root

Why: The standard deviation.

\[ 5.48 \]

Compare them

Why: Spread against centre.

Figure (svg): The solution to Worked example the mean and standard deviation shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \mu = 15, \qquad \sigma = \sqrt{30} \approx 5.48 \]

Verify: confirm the spread is a plausible fraction of the mean

Why: The standard deviation is about 37 percent of the mean here, which is large enough for the distribution to be visibly lopsided — it cannot extend more than about 2.7 standard deviations below its mean before hitting zero. That constraint on the left with none on the right is exactly what right skew means, and it is why the skew is severe at small df and mild at large.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

9. Mean and spread

Faded example

A chi-square distribution with 50 degrees of freedom.

Fill in the blanks

\mu = 50, \qquad \sigma = \sqrt10 = ___

Why: The mean is the degrees of freedom, 50, and the standard deviation is the square root of twice that, which is 10.

10. Worked example: why the statistic cannot be negative

Worked example

Tracing it back to the definition.

\[ \text{can } \chi^2 \text{ be } -3? \]

Recall the definition

Why: A sum of squares.

\[ \text{each term at least } 0 \]

Sum non-negative terms

Why: The total.

\[ \text{also at least } 0 \]

Check the test statistics

Why: Each is a sum of squares over positives.

Conclude

Why: Never negative.

Figure (svg): The solution to Worked example why the statistic cannot be negative shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \chi^2 \ge 0 \text{ always} \]

Verify: confirm this holds for the chapter's actual test statistics too

Why: The goodness-of-fit statistic is a sum of terms of the form observed minus expected, squared, divided by expected — a square over a positive count, so every term is non-negative and so is the total. A negative chi-square statistic is therefore always an arithmetic error, which makes the sign a free check on every calculation in this chapter.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565

11. Trap: expecting a symmetric distribution

Trap

The trap

\[ \text{find the value with } 2.5 \text{ percent in each tail, symmetrically} \]

Treat chi-square like the normal or the t

Why: Both earlier distributions were symmetric.

\[ \text{but the curve is skewed and bounded at zero} \]

The two tails are not mirror images, so a symmetric interval around the mean is not the right construction.

The fix

\[ \text{look up each tail's cut-off separately} \]

Treat the two tails as unrelated, since the curve is asymmetric

Why: The book's first fact.

This is why section 8.1's note about non-symmetrical confidence intervals mattered: an interval for a variance is built from two separately looked-up chi-square values rather than from a point estimate plus and minus a margin. Every symmetric habit from chapters 6 to 10 has to be set aside here.

12. One of these is false

Two truths and a lie

All three concern the definition.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The statistic is never negative
  • C. The mean equals the degrees of freedom
  • B. The distribution is symmetric about its mean

Survives elimination: B

Why: The survivor is false and is the book's first fact reversed. The curve is nonsymmetrical and skewed to the right, because it is bounded at zero on the left and unbounded on the right.

13. Where does the skew come from?

Prediction

Commit before reasoning.

Predict first

Why is the chi-square distribution skewed to the right?

  • It is bounded at zero on the left but unbounded on the right
  • Because the degrees of freedom are small
  • Because it is a sum
  • Because the normal is skewed

Correct: It is bounded at zero but unbounded above.

Why: A distribution cannot be symmetric when it has a hard floor close to its centre and no ceiling. As the degrees of freedom grow the mean moves further from that floor relative to the spread, which is why the skew fades — but it never disappears entirely.

14. How far can it reach below the mean?

Estimation

A chi-square with 4 degrees of freedom has mean 4 and standard deviation about 2.83.

Predict first

How many standard deviations below the mean is zero?

  • About 1.4
  • About 4
  • About 2.8
  • About 0.7

Correct: About 1.4.

Why: Four divided by 2.83 is about 1.41, so the distribution runs out of room barely one and a half standard deviations below its centre while extending indefinitely above. That asymmetry in available room is precisely the skew, and it is severe at small degrees of freedom.

15. A curve for every df

Section

Section 2

16. The shape changes with the degrees of freedom

Concept

There is a different chi-square curve for each degrees of freedom, and the mean is located just to the right of the peak. Small degrees of freedom give a curve crowded against zero; larger ones give a curve that is flatter, further along the axis and closer to symmetric.

the family of curves — Indexed by degrees of freedom, which unlike the t's are not always n minus one — each of the chapter's three tests counts them differently.

\[ \chi^2_{\text{df}}: \quad \text{one curve per df} \]

The book's fifth fact — that the mean sits just to the right of the peak — is a direct consequence of right skew, and it is the same relationship section 2.6 established for any right-skewed distribution: the long tail pulls the mean above the mode. Chapter 5's exponential showed the same pattern, and for the same reason.

Figure (svg): Four chi-square curves of increasing degrees of freedom, each skewed right and each flatter and further along the axis than the last

The book's second and fifth facts, drawn: a curve per df, and the mean located just to the right of the peak.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564 — facts two and five

17. Four curves

Picture it

Degrees of freedom of 2, 4, 8 and 15.

Figure (svg): Four chi-square curves of increasing degrees of freedom, each skewed right and each flatter and further along the axis than the last

The book's second and fifth facts, drawn: a curve per df, and the mean located just to the right of the peak.

Each dashed mark sits at that curve's mean, and each is a little to the right of its peak. The curves also flatten as they move right, because the total area stays one while the spread grows — the same trade-off every density in the book obeys.

18. Worked example: comparing two curves

Worked example

Degrees of freedom of 2 and 15.

\[ \text{df} = 2 \text{ against } \text{df} = 15 \]

Means

Why: Equal to df.

\[ 2\text{ and } 15 \]

Standard deviations

Why: Root of twice df.

\[ 2\text{ and } 5.48 \]

Relative spread

Why: sd over mean.

\[ 1.00\text{ and } 0.37 \]

Say what that means

Why: Less lopsided.

Figure (svg): The solution to Worked example comparing two curves shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{\sigma}{\mu} = 1.00 \quad\text{against}\quad 0.37 \]

Verify: confirm the relative spread is what governs the skew

Why: At two degrees of freedom the standard deviation equals the mean, so zero is only one standard deviation away and the curve is severely lopsided. At fifteen it is 2.7 standard deviations away and the shape is much more even. The ratio of standard deviation to mean is root two over root df, so it falls steadily and the skew fades with it.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

19. Which curve is further right?

Sorting

The mean is the degrees of freedom.

Sort into buckets

Sort each pair by whether the first curve sits further right.

The first is further right
df = 15 against df = 4; df = 90 against df = 10; df = 50 against df = 20
The second is
df = 2 against df = 8; df = 3 against df = 30
yes
The first has the larger degrees of freedom, so the larger mean.
no
The second has the larger degrees of freedom.

Since the mean IS the degrees of freedom, the comparison needs no calculation at all — which is one of the more convenient consequences of the definition.

20. Worked example: locating the peak

Worked example

Why the mean sits to its right.

\[ \text{df} = 8 \]

Recall the mean

Why: Equal to df.

\[ 8 \]

Note the skew direction

Why: Right.

Recall section 2.6

Why: The tail pulls the mean.

Conclude

Why: The peak is left of 8.

\[ \text{at df minus } 2,\text{ which is } 6 \]

Figure (svg): The solution to Worked example locating the peak shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{mode} = \text{df} - 2 = 6 \;<\; \mu = 8 \]

Verify: confirm this matches every right-skewed distribution met so far

Why: Section 2.6 established the ordering for right-skewed data, and section 5.3's exponential showed it numerically with a median at 69 percent of the mean. The chi-square is another instance of the same phenomenon, which is why the book can state it as a fact without deriving it — the reader has met the pattern three times already.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

21. Error analysis: four claims about the family of curves

Error analysis

Which correctly describe the chi-square distribution?

Annotate

On: \( \begin{aligned} &(1)\; \text{the degrees of freedom are always } n - 1 \\ &(2)\; \text{the curve can take negative values} \\ &(3)\; \text{the peak is to the right of the mean} \\ &(4)\; \text{the mean is the degrees of freedom} \end{aligned} \)

  • (1) is false in general. It holds for a test of a single variance, but a goodness-of-fit test uses categories minus one and a contingency test multiplies two dimensions.
  • (2) is false: a sum of squares is never negative, and the book states it as its third fact.
  • (3) has the ordering backwards. Right skew puts the mean to the RIGHT of the peak.
  • (4) is correct, and it follows from each squared standard normal averaging one.

Error (1) is the one worth guarding against, because it is true in one of the chapter's cases and false in the other two. The book flags it directly: the degrees of freedom for the three major uses are each calculated differently.

22. One of these is false

Two truths and a lie

All three concern the family.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The mean lies just to the right of the peak
  • C. There is a different curve for each df
  • B. The degrees of freedom are always n minus one

Survives elimination: B

Why: The survivor is false. Only the test of a single variance uses n minus one; a goodness-of-fit test uses categories minus one, and a contingency test multiplies rows minus one by columns minus one. The book warns that the three uses count them differently.

23. Compare two curves

Faded example

Chi-square with 4 and with 25 degrees of freedom.

Fill in the blanks

\frac0.710.28 = \frac___}}___}: \quad ___ \text___ = 4, \text___ ___ \text___ = 25

Why: The ratio falls from about 0.71 to about 0.28, which is why the second curve is far less skewed — it has proportionally much more room below its mean.

24. What happens as df grows?

Prediction

Commit before reasoning.

Predict first

As the degrees of freedom increase, the chi-square curve becomes what?

  • Wider, further right, and closer to symmetric
  • Narrower and more skewed
  • Unchanged in shape
  • Bounded above

Correct: Wider, further right, and closer to symmetric.

Why: The mean is df so the curve moves right, the standard deviation is the root of twice df so it widens, and the ratio of the two falls so the skew fades. All three follow from the two formulas, which is why the family can be described completely by one parameter.

25. The normal approximation

Section

Section 3

26. Above about ninety degrees of freedom

Concept

When the degrees of freedom exceed 90, the chi-square curve approximates the normal distribution. The book gives its own instance: for a chi-square with 1,000 degrees of freedom the mean is 1,000 and the standard deviation is the square root of 2,000, about 44.7, so the distribution is approximately normal with those parameters.

the large-df approximation — A chi-square with many degrees of freedom is approximately N(df, root of twice df). It follows from the central limit theorem, since the statistic is a sum of independent terms.

\[ \text{df} > 90 \;\Longrightarrow\; \chi^2 \approx N\left(\text{df}, \sqrt{2\,\text{df}}\right) \]

This is chapter 7 doing exactly what it always does. A chi-square variable is a sum of df independent squared normals, so as df grows the central limit theorem takes hold and the sum becomes normal — with the mean and standard deviation the definition already supplies. Nothing new is being asserted; the theorem is simply being applied to a sum that happens to have a name.

Figure (svg): A chi-square curve with a thousand degrees of freedom drawn over a normal curve of matching mean and standard deviation, the two nearly indistinguishable

The book's fourth fact, and its own worked instance: chi-square with df of 1,000 is approximately N(1000, 44.7).

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564 — the fourth fact and its worked instance

27. A thousand degrees of freedom

Picture it

The chi-square curve against its normal approximation.

Figure (svg): A chi-square curve with a thousand degrees of freedom drawn over a normal curve of matching mean and standard deviation, the two nearly indistinguishable

The book's fourth fact, and its own worked instance: chi-square with df of 1,000 is approximately N(1000, 44.7).

The two curves are nearly indistinguishable at this scale, which is what the fourth fact claims. A trace of right skew survives — the chi-square is very slightly higher in the upper tail — but nothing that would affect a practical calculation.

28. Worked example: the book's own instance

Worked example

A chi-square with 1,000 degrees of freedom.

\[ \text{df} = 1000 \]

The mean

Why: Equal to df.

\[ 1, 000 \]

Twice the df

Why: Two thousand.

\[ 2, 000 \]

Take the root

Why: The standard deviation.

\[ 44.7 \]

State the approximation

Why: By the fourth fact.

\[ N(1000, 44.7) \]

Figure (svg): The solution to Worked example the book's own instance shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ X \approx N(1000, 44.7) \]

Verify: confirm the relative spread is small enough for symmetry

Why: The standard deviation is 44.7 against a mean of 1,000, so zero is over twenty standard deviations below the centre — the boundary that causes the skew is now far out of reach, and the distribution has ample room on both sides. That ratio of about 0.045 is what makes the approximation work, and it is why the threshold is stated in degrees of freedom.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

29. The approximating normal

Faded example

A chi-square with 200 degrees of freedom.

Fill in the blanks

\approx N(200, \sqrt400}) = N(200, 20)

Why: Twice 200 is 400, whose root is 20 — so the distribution is approximately normal with mean 200 and standard deviation 20.

30. Worked example: why ninety

Worked example

Checking the threshold against the relative spread.

\[ \text{df} = 90 \]

The mean

Why: Ninety.

\[ 90 \]

The standard deviation

Why: Root of 180.

\[ 13.42 \]

The ratio

Why: 13.42 over 90.

\[ 0.149 \]

Distance to zero

Why: One over the ratio.

\[ 6.7 s d \]

Figure (svg): The solution to Worked example why ninety shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{90}{13.42} \approx 6.7 \]

Verify: confirm this explains why the threshold falls where it does

Why: A normal distribution has essentially no probability more than about four standard deviations from its centre, so once the floor at zero sits nearly seven standard deviations away it stops mattering. That is the substance behind the book's ninety: it is roughly where the bound ceases to distort the shape, rather than an arbitrary round number.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

31. Trap: using the approximation at small degrees of freedom

Trap

The trap

\[ \text{df} = 4 \;\Rightarrow\; \text{use } N(4, 2.83) \]

Apply the fourth fact without checking the threshold

Why: The formulas for the mean and spread always hold.

\[ \text{the normal would assign probability below zero} \]

At four degrees of freedom zero is only 1.4 standard deviations below the mean, so a normal would put substantial probability on impossible values.

The fix

\[ \text{use the chi-square itself, and reserve the approximation for df} > 90 \]

Check the degrees of freedom before approximating

Why: The mean and spread formulas hold at every df; the normal SHAPE does not.

The distinction is worth being precise about: the mean is df and the standard deviation is the root of twice df at every degrees of freedom, and only the claim about the SHAPE requires df above ninety. Every test in this chapter uses small degrees of freedom, so the approximation never arises in practice here.

32. One of these is false

Two truths and a lie

All three concern the approximation.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It follows from the central limit theorem
  • C. The mean and spread formulas hold at every df
  • B. The approximation applies at any degrees of freedom

Survives elimination: B

Why: The survivor is false. The book gives a threshold of 90, and below it the boundary at zero distorts the shape too much for a normal to describe. At four degrees of freedom a normal would assign real probability to negative values.

33. Where does the bound stop mattering?

Estimation

The floor at zero sits mean over standard deviation below the centre.

Predict first

At roughly what degrees of freedom is zero four standard deviations below the mean?

  • About 32
  • About 4
  • About 90
  • About 400

Correct: About 32.

Why: The distance is df over the root of twice df, which is the root of df over two — and that equals 4 when df is 32. So the bound is already fairly remote by then, and the book's threshold of 90 is comfortably conservative. The approximation improves steadily rather than switching on at a point.

34. Which chapter supplies the reason?

Prediction

Commit before reasoning.

Predict first

Which earlier result explains why a chi-square becomes normal for large df?

  • The central limit theorem, since chi-square is a sum of independent terms
  • The empirical rule
  • The law of large numbers
  • Bayes' rule

Correct: The central limit theorem.

Why: Chapter 7 established that sums of independent quantities tend toward normality as their number grows, whatever the individual terms look like. A chi-square with df degrees of freedom is a sum of df independent squared normals, so it is exactly the kind of quantity the theorem covers.

35. Degrees of freedom in the three tests

Section

Section 4

36. Counted differently each time

Concept

The degrees of freedom depend on how chi-square is being used, and the book warns that the three major uses each calculate them differently. Only for a test of a single variance are they n minus one.

the three counts — Categories minus one for goodness of fit; rows minus one times columns minus one for a contingency table; and n minus one for a test of a single variance.

\[ k - 1; \quad (r-1)(c-1); \quad n - 1 \]

The underlying principle is the one section 8.2 established: each constraint imposed on the data costs a degree of freedom. A goodness-of-fit test's expected counts must total the observed total, which is one constraint. A contingency table's expected counts must reproduce every row and column total, which is a larger set of constraints and gives the product formula. The counting differs because the constraints differ.

Figure (svg): The book's five facts about the chi-square distribution

Facts one, three and five follow from squaring; facts two and four follow from summing k of them.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564 — the note that the three uses count df differently

37. The five facts, with the df note

Picture it

What the book states about the distribution.

Figure (svg): The book's five facts about the chi-square distribution

Facts one, three and five follow from squaring; facts two and four follow from summing k of them.

The note beneath the list is the one that matters most for the rest of the chapter. Everything else about the distribution is settled once the degrees of freedom are known, so getting them right is the only genuinely test-specific step.

38. Worked example: counting for three tests

Worked example

The same distribution, three different counts.

\[ \text{4 categories; a } 3\times 3 \text{ table; a sample of } 20 \]

Goodness of fit

Why: Categories minus one.

\[ 3 \]

Contingency table

Why: Two times two.

\[ 4 \]

Single variance

Why: n minus one.

\[ 19 \]

Note the pattern

Why: All different.

Figure (svg): The solution to Worked example counting for three tests shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k-1 = 3; \quad (3-1)(3-1) = 4; \quad n-1 = 19 \]

Verify: confirm why the contingency count is a product

Why: A three-by-three table has nine cells, but once two entries in each row and two in each column are known the rest are forced by the totals — leaving a two-by-two block of free values, which is four. Counting the genuinely free cells rather than memorising the formula makes it clear why the answer is a product rather than a sum.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

39. Test to degrees of freedom

Matching

Match each use to its count.

Match the pairs

  • l1. goodness of fit, k categories
  • l2. test of independence, r by c table
  • l3. test for homogeneity, r by c table
  • l4. test of a single variance, n data
  • r1. k minus 1
  • r2. (r minus 1)(c minus 1)
  • r3. (r minus 1)(c minus 1)
  • r4. n minus 1

Why: Independence and homogeneity share both their formula and their arithmetic entirely — section 11.4 makes that explicit — so only three distinct counts appear across four tests.

40. Worked example: the constraint that costs the degree

Worked example

Why a goodness-of-fit test loses exactly one.

\[ k \text{ categories with expected counts} \]

Note the total

Why: Expected must match observed.

Count the cells

Why: k of them.

Subtract the constraint

Why: One is determined.

\[ k\text{ minus } 1 \]

Compare with section 8.2

Why: The same principle.

Figure (svg): The solution to Worked example the constraint that costs the degree shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{df} = k - 1 \]

Verify: confirm this is the same rule as section 8.2's

Why: There the n deviations from a sample mean had to sum to zero, so one was determined and the degrees of freedom were n minus one. Here the k expected counts must sum to the sample total, so one is determined and the degrees of freedom are k minus one. One principle — a constraint costs a degree of freedom — producing both results.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565

41. Trap: using n minus one for every chi-square test

Trap

The trap

\[ 60 \text{ observations in } 5 \text{ categories} \;\Rightarrow\; \text{df} = 59 \]

Apply the familiar n minus one

Why: It worked for every t test.

\[ \text{but there are only five categories} \]

The degrees of freedom count the free CELLS, not the observations, so the answer is four.

The fix

\[ \text{df} = k - 1 = 5 - 1 = 4 \]

Count what the test's constraints leave free

Why: For goodness of fit that is categories minus one.

The error is large and always in the direction of a smaller p-value, since a chi-square with 59 degrees of freedom sits far to the right of one with 4. It is also invisible in the output, which is why the book's warning about the three counts is worth reading carefully before the first test is attempted.

42. Count for a table

Faded example

A contingency table with 3 rows and 5 columns.

Fill in the blanks

\text2 = (3 - 1)(5 - 1) = 4 \times ___ = 8

Why: Two times four is eight — the number of cells that can vary freely once every row and column total is fixed.

43. One of these is false

Two truths and a lie

All three concern degrees of freedom.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The three uses count them differently
  • C. A constraint on the data costs a degree of freedom
  • B. The count is always based on the sample size

Survives elimination: B

Why: The survivor is false. Only the test of a single variance uses the sample size; the other two count categories or table cells. A goodness-of-fit test on 600 observations in five categories has four degrees of freedom, not 599.

44. Which way does the error run?

Prediction

Commit before reasoning.

Predict first

Using n minus one instead of categories minus one gives what kind of error?

  • A much larger df, so a much smaller p-value
  • A smaller df, so a larger p-value
  • No change to the p-value
  • An impossible statistic

Correct: A much larger df and a much smaller p-value.

Why: A chi-square curve with many degrees of freedom sits far to the right, so a statistic that would be extreme on four degrees of freedom looks unremarkable on 59 — actually giving a LARGER p-value. The direction depends on the statistic's size, which is precisely what makes the error unpredictable and dangerous rather than conservative.

45. Reading a chi-square picture

Section

Section 5

46. Right-tailed, and bounded at zero

Concept

Because the statistic measures how far observed values fall from expected ones, large values mean poor agreement and small values mean good agreement. Every test in the next three sections is therefore right-tailed, and the picture always shows a shaded upper tail.

why right-tailed — A chi-square statistic grows when observations depart from expectation in either direction, so all disagreement lands in the upper tail. There is no lower tail of disagreement to look at.

\[ \text{large } \chi^2 \;\Longrightarrow\; \text{poor fit} \;\Longrightarrow\; \text{small } p \]

The exception is section 11.6's test of a single variance, which the book notes may be right-tailed, left-tailed or two-tailed. That is because there the statistic measures a sample variance against a claimed one, and a variance can be too small as well as too large — so both directions carry meaning, unlike a measure of disagreement.

Figure (svg): A four-column table showing the chi-square mean, standard deviation and their ratio at five degrees of freedom values

The last column is why the skew fades: a distribution whose spread is small relative to its centre has little room to be lopsided.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565 — the facts, and the goodness-of-fit tail

47. Mean, spread and their ratio

Picture it

How the distribution's shape changes across degrees of freedom.

Figure (svg): A four-column table showing the chi-square mean, standard deviation and their ratio at five degrees of freedom values

The last column is why the skew fades: a distribution whose spread is small relative to its centre has little room to be lopsided.

The last column falls steadily, and it is the single number that predicts everything about a chi-square curve's appearance. A ratio near one gives a curve crushed against zero; a ratio near zero gives one indistinguishable from a normal.

48. Worked example: reading a statistic against its df

Worked example

The same value on two different curves.

\[ \chi^2 = 12 \text{ on df } = 4 \text{ and on df } = 20 \]

On df = 4

Why: Mean 4, sd 2.83.

\[ 2.8\text{ sd above} \]

Its right-tail area

Why: Small.

\[ \text{about } 0.017 \]

On df = 20

Why: Mean 20, sd 6.32.

\[ 1.3\text{ sd BELOW} \]

Its right-tail area

Why: Large.

\[ \text{about } 0.916 \]

Figure (svg): The solution to Worked example reading a statistic against its df shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p \approx 0.017 \quad\text{against}\quad p \approx 0.916 \]

Verify: confirm the comparison always runs against the mean

Why: Since the mean IS the degrees of freedom, comparing the statistic with df gives an instant read: well above means a small p-value and below means a large one. A statistic of 12 against a mean of 4 is far out; the same statistic against a mean of 20 is on the low side of ordinary. That one comparison predicts the outcome before any lookup.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, p. 564

49. Good fit or poor?

Sorting

Compare each statistic with its degrees of freedom.

Sort into buckets

Sort each result.

Good agreement: a large p-value
chi-square 3 on df 4; chi-square 2 on df 10
Poor agreement: a small p-value
chi-square 14.3 on df 3; chi-square 12.99 on df 4; chi-square 25 on df 5
good
The statistic is at or below its degrees of freedom, so the observed counts are close to expectation.
poor
The statistic is well above its degrees of freedom, so the departure is large.

Comparing the statistic with df is the whole of the intuition, since df IS the mean. Item (a) is Example 11.2's result, whose p-value was 0.5578; item (c) is Example 11.6's, at 0.0113.

50. Worked example: a statistic below its mean

Worked example

What an unusually small chi-square means.

\[ \chi^2 = 1 \text{ on df } = 10 \]

Compare with the mean

Why: One against ten.

Find the right-tail area

Why: Nearly all of it.

\[ \text{about } 0.9998 \]

Say what it means for fit

Why: Observed near expected.

Note the decision

Why: A huge p-value.

Figure (svg): The solution to Worked example a statistic below its mean shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p \approx 0.9998 \]

Verify: confirm what an implausibly small statistic can indicate

Why: Agreement this close is itself unusual — real sampling produces some departure from expectation, so a statistic far below its degrees of freedom happens only about once in five thousand samples here. Historically such results have been treated as evidence of data that were adjusted rather than observed, which is a use of the LOWER tail that the goodness-of-fit test's right-tailed convention never examines.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565

51. Trap: reading a small statistic as evidence against the null

Trap

The trap

\[ \chi^2 = 1 \text{ is far from } 10, \text{ so the fit is poor} \]

Treat any departure from the mean as evidence

Why: Both directions are unusual.

\[ \text{but small means observed CLOSE to expected} \]

The statistic measures disagreement, so a small value means good agreement rather than bad.

The fix

\[ \text{small } \chi^2 \;\Rightarrow\; \text{good fit} \;\Rightarrow\; \text{large } p \]

Read the statistic as a measure of disagreement

Why: Only large values indicate a problem.

This is why the tests are right-tailed and why the direction is not a choice the way it was in chapter 9. The alternative hypothesis — that the data do not fit — can only show up as a large statistic, so the p-value is always an upper tail.

52. One of these is false

Two truths and a lie

All three concern reading the statistic.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Larger statistics mean poorer agreement
  • C. Goodness-of-fit and contingency tests are right-tailed
  • B. Every chi-square test is right-tailed

Survives elimination: B

Why: The survivor is false. Section 11.6's test of a single variance may be right-tailed, left-tailed or two-tailed, because a variance can be smaller than claimed as well as larger — and both departures are meaningful.

53. Predict the p-value

Estimation

A chi-square statistic of 4 on 10 degrees of freedom.

Predict first

Roughly what p-value should be expected?

  • Large, well above 0.5
  • Small, below 0.05
  • Exactly 0.5
  • Impossible to say

Correct: Large, well above 0.5.

Why: The statistic is well below its mean of 10, so most of the distribution lies above it and the right-tail area is large — about 0.95 here. Comparing the statistic against the degrees of freedom gives this read instantly, and it catches a badly miscounted df as well.

54. Explain the tail

Explain it

A classmate asks why chi-square tests do not have a left-tailed version like chapter 9's.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what a small statistic would mean.

Answer:

The statistic measures how far the observations fall from what the null predicts, so any kind of disagreement — in any direction, in any cell — makes it larger.

A small statistic means the data agree closely with the null, which is never evidence against it, so there is nothing for a left tail to detect.

55. Chi-square against the earlier distributions

Comparison

Fill the blanks. Every symmetric habit from chapters 6 to 10 has to be set aside.

Comparison matrix

Normal and tChi-square
Shapesymmetricskewed right
Rangethe whole linezero and above
Meanzero, after standardisingthe degrees of freedom
Tails usedleft, right or bothright, except for a test of a variance

The second row is the source of everything else. A distribution with a hard floor near its centre cannot be symmetric, and once the floor becomes remote — which is what large degrees of freedom achieve — the shape becomes normal again.

56. Working with a chi-square distribution, in order

Pattern

Four steps, and the second is the only test-specific one.

  1. Identify which of the chapter's uses applies: goodness of fit, independence, homogeneity, or a single variance.
  2. Count the degrees of freedom by that test's own rule — categories minus one, the product of two dimensions, or n minus one.
  3. Recall that the mean is the degrees of freedom and the spread is the square root of twice them.
  4. Compare the statistic with the degrees of freedom for a quick read: well above means a small p-value, at or below means a large one.

The statistic can never be negative, so a negative value is always an arithmetic error rather than an unusual result.

OpenStax Introductory Business Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1 Facts About the Chi-Square Distribution

57. Check yourself 1 of 3

Check

The parameters.

Check your understanding

A chi-square distribution has 18 degrees of freedom. What are its mean and standard deviation?

  • A. Mean 18, standard deviation 6 (correct)
  • B. Mean 18, standard deviation 36
  • C. Mean 0, standard deviation 1
  • D. Mean 17, standard deviation 6

Answer: A

Why: The mean is the degrees of freedom, and the standard deviation is the square root of twice them: the root of 36 is 6.

Why B tempts people
That is twice the degrees of freedom, before taking the square root.
Why C tempts people
Those are the standard normal's parameters, not chi-square's.
Why D tempts people
The mean is df itself, not df minus one.

58. Check yourself 2 of 3

Check

The shape.

Check your understanding

Why can a chi-square statistic never be negative?

  • A. It is a sum of squared quantities (correct)
  • B. The degrees of freedom are positive
  • C. Probabilities cannot be negative
  • D. It is standardised

Answer: A

Why: The definition is a sum of squared standard normals, and every test statistic in the chapter is likewise a sum of squares over positive quantities.

Why B tempts people
The degrees of freedom index the curve but do not bound the variable.
Why C tempts people
That is true of probabilities but says nothing about the statistic's range.
Why D tempts people
Standardising does not prevent negative values; the standard normal takes plenty of them.

59. Check yourself 3 of 3

Check

Degrees of freedom.

Check your understanding

A goodness-of-fit test has 600 observations in five categories. What are the degrees of freedom?

  • A. 4 (correct)
  • B. 599
  • C. 5
  • D. 595

Answer: A

Why: A goodness-of-fit test uses categories minus one, so five categories give four degrees of freedom whatever the sample size.

Why B tempts people
That applies n minus one, which is the rule for a test of a single variance rather than for goodness of fit.
Why C tempts people
That is the number of categories, before subtracting the one constraint.
Why D tempts people
That subtracts the categories from the observations, which corresponds to no rule in the chapter.

60. Where this shows up outside the textbook

Real world

A researcher runs a goodness-of-fit test on 400 observations across six categories and reports a chi-square statistic of 4.1 with a p-value of 0.9999, concluding the data fit the hypothesised distribution extremely well. A colleague suggests the fit is suspiciously good.

Discussion prompt

Assess both readings, and say what the colleague might be getting at.

Hint: Compare the statistic with what the distribution predicts for it.

Answer:

The first reading is arithmetically wrong. With six categories the degrees of freedom are five, so the mean of the statistic is 5 and a value of 4.1 is entirely ordinary — slightly below average. The right-tail area beyond 4.1 on five degrees of freedom is about 0.535, not 0.9999, so the reported p-value does not match the reported statistic.

\[ \text{df} = 6 - 1 = 5, \qquad \mu = 5, \qquad P(\chi^2 > 4.1) \approx 0.535 \]

A p-value of 0.9999 would require a statistic of about 0.4, not 4.1 — so either the statistic or the p-value has been miscomputed, most likely by using the wrong degrees of freedom. Checking the statistic against its mean would have caught this immediately, which is why that comparison is worth making on every chi-square result.

The colleague's instinct is sound in general, even if it does not apply here. A statistic far BELOW its degrees of freedom means the observed counts match expectation more closely than random sampling would normally produce, and historically that pattern has been used as evidence that data were adjusted rather than collected — most famously in re-examinations of some early genetics results. A p-value of 0.9999 genuinely would be suspicious; 0.535 is not.

Two points follow. A chi-square result should always be reported with its degrees of freedom, since neither the statistic nor the p-value is interpretable without them — and the mismatch here is only visible because both were given. And the lower tail, which the right-tailed convention never examines, carries real information about whether data are too good; looking at it is a standard forensic check even though no test in this chapter performs it.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why is the mean of a chi-square distribution equal to its degrees of freedom?

  • By convention, to make the tables simpler
  • Because it is a sum of k squared standard normals, and each averages one
  • Because the degrees of freedom are always n minus one
  • Because the distribution is skewed

Correct: Because each squared standard normal averages one.

\[ E[Z^2] = \text{Var}(Z) = 1 \;\Longrightarrow\; E[\chi^2_k] = k \]

Why: A standard normal has mean zero and variance one, and variance is the average squared deviation from the mean — so the average of Z squared is exactly 1. Adding k independent such terms gives a mean of k. The result is a consequence of the definition rather than a convention, and it is what makes comparing a statistic against its degrees of freedom such a useful quick check.

62. Explain it to someone a year behind you

Explain it

They used n minus one as the degrees of freedom for a goodness-of-fit test on 60 observations in five categories.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what the constraint on the expected counts actually is.

Answer:

The degrees of freedom count the free cells, not the observations — and with five categories whose expected counts must total the sample size, only four can vary freely.

So the answer is 4 rather than 59, and the book warns that the chapter's three uses each count degrees of freedom differently.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Stating the definition and deriving the five facts from it
  • Computing the mean and standard deviation from the degrees of freedom
  • Saying why the curve becomes normal for large degrees of freedom
  • Knowing that the three tests count degrees of freedom differently

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, a sum of k squared standard normals. For the second, the mean is df and the spread is the root of twice df. For the third, it is a sum of independent terms, so chapter 7 applies. For the fourth, categories minus one, the product of two dimensions, or n minus one. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twelve minutes — this is a short section.

Draw it

At the top, write the definition of a chi-square variable as a sum of k squared standard normals, and beneath it draw three arrows to the three things it implies: never negative, mean equal to k, and normal for large k. Below that, list the book's five facts and mark beside each which part of the definition produces it. In the middle of the page, draw four chi-square curves on one axis for degrees of freedom 2, 4, 8 and 15, marking each curve's mean with a dashed line and noting that it falls just right of the peak. Beside them, make a table of df, mean, standard deviation and their ratio for df of 2, 4, 10, 90 and 1000, and write one sentence on why the last column explains the fading skew. At the bottom, write the three degrees-of-freedom rules — categories minus one, rows minus one times columns minus one, and n minus one — and beside each name the test it belongs to.

Check your table by confirming the ratio column falls steadily and equals the square root of two over the square root of df at every row. Check your curves by confirming each mean mark sits to the right of that curve's highest point, which is the book's fifth fact and the one easiest to draw backwards.

65. What you can do now

Recap

Five things, and all of them come from one definition.

If you seeThen
A chi-square with df degrees of freedomMean df, standard deviation root of twice df
A negative chi-square statisticAn arithmetic error: it cannot happen
A statistic well above its dfPoor agreement, so a small p-value
A statistic well below its dfGood agreement, so a large p-value
df above about 90The normal approximation applies
A goodness-of-fit testdf is categories minus one, not n minus one
A contingency tabledf is rows minus one times columns minus one

Section 11.2 puts the distribution to its first use. A goodness-of-fit test compares a whole set of observed counts against what a hypothesised distribution predicts, and answers a question no single mean or proportion could pose: whether the data fit a distribution at all.

OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution §11.1, pp. 564-565 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §11.1 Facts About the Chi-Square Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 564-565
  2. OpenStax Introductory Business Statistics 2e, §11.1 Facts About the Chi-Square Distribution — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108