13.4 Test of Two Variances

The chapter's second use of the F distribution, and the course's last section. It is often desirable to compare two variances rather than two averages — college administrators would like two professors to have the same variation in their grading, a lid and a container must vary alike to fit, a supermarket may care about the variability of two checkers' times. The statistic is the ratio of the two sample variances, since the population variances cancel under a null of equality, and its degrees of freedom are one less than each sample size. Because a variance can be larger or smaller than another, the test may be left-tailed, right-tailed or two-tailed. The section also carries a warning unlike anything else in the book: this test is very sensitive to deviations from normality, can give p-values too high or too low in unpredictable ways, and many texts suggest students not use it at all — a caveat this lesson treats as the section's most important content.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 13.4 Test of Two Variances

Title

Statistics · Chapter 13 — F Distribution and One-Way ANOVA

Test of Two Variances

2. By the end of this lesson you can

Objectives

Six outcomes, and the last is the one the book most wants remembered.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 11.6 tested one variance against a claimed value; section 13.2 built a ratio of two variance estimates.

Discussion prompt

Two instructors grade the same 30 exams. The first's grades have variance 52.3 and the second's 89.9. How would you test whether the two populations have the same variance?

Hint: You have two sample variances and a distribution built for ratios.

Answer:

Take their ratio. If the two population variances are genuinely equal, then two samples of the same size should give sample variances close in value and their ratio close to one.

\[ F = \frac{s_1^2}{s_2^2} = \frac{52.3}{89.9} = 0.5818 \]

A ratio of 0.58 means the first instructor's grades varied only about 58 percent as much as the second's. Whether that is more than sampling variation would produce is what the F distribution answers, with 29 degrees of freedom for each sample.

Note the direction. Here the ratio came out BELOW one, so the evidence points toward the first variance being smaller — and the test that examines it is left-tailed. That option did not exist for the one-way ANOVA, and this section is where it returns.

4. A ratio of two sample variances

Concept

Another of the uses of the F distribution is testing two variances. Since we are interested in comparing the two sample variances, we use the F ratio, which has the distribution F with n1 minus one and n2 minus one degrees of freedom. If the null hypothesis is that the two population variances are equal, the F ratio becomes the ratio of the two sample variances.

why the ratio works — If the two populations have equal variances, the two sample variances are close in value and their ratio is close to one. If the population variances are very different, the sample variances tend to be very different too.

\[ F = \frac{s_1^2}{s_2^2}, \qquad F \sim F_{n_1-1,\,n_2-1} \]

The full statistic divides each sample variance by its own population variance before taking the ratio. Under a null of equality those two population variances are the same and cancel, leaving the sample variances alone — which is why the test is computable at all, since the population variances are unknown.

Figure (svg): A card giving the F ratio for a test of two variances

The population variances cancel under the null, leaving a ratio of the two sample variances.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691

5. The ratio and its degrees of freedom

Section

Section 1

6. The population variances cancel

Concept

The F ratio divides each sample variance by its population variance and takes the quotient. If the null hypothesis is that the two population variances are equal, they cancel and the F ratio becomes the ratio of the two sample variances alone.

two degrees of freedom — n1 minus one for the numerator and n2 minus one for the denominator — one less than each sample size, since each sample variance is computed about its own sample mean.

\[ F = \frac{s_1^2/\sigma_1^2}{s_2^2/\sigma_2^2} \;\overset{H_0}{=}\; \frac{s_1^2}{s_2^2} \]

The book adds a note about which variance goes on top: the F ratio could also be written the other way up, and it depends on the alternative hypothesis and on which sample variance is larger. Putting the larger one on top forces the ratio above one, which is convenient for a two-tailed test but reverses the tail for a directional one — so the choice has to be made deliberately.

Figure (svg): A card giving the F ratio for a test of two variances

The population variances cancel under the null, leaving a ratio of the two sample variances.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691 — the F ratio, the cancellation, and the note on which variance goes on top

7. The statistic

Picture it

Before and after the cancellation.

Figure (svg): A card giving the F ratio for a test of two variances

The population variances cancel under the null, leaving a ratio of the two sample variances.

The lower panels give the two degrees of freedom, and the yellow note is the book's own: which sample variance sits on top is a choice, and it interacts with the direction of the alternative.

8. Worked example: Example 13.5's statistic

Worked example

Two instructors each grade the same 30 exams. The first's grades have variance 52.3 and the second's 89.9.

\[ s_1^2 = 52.3, \; s_2^2 = 89.9, \; n_1 = n_2 = 30 \]

The ratio

Why: First over second.

\[ 52.3\text{ over } 89.9 \]

Evaluate

Why: Divide.

\[ 0.5818 \]

Numerator df

Why: n1 minus one.

\[ 29 \]

Denominator df

Why: n2 minus one.

\[ 29 \]

Figure (svg): The solution to Worked example Example 13.5's statistic shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F_{29,29} = \frac{52.3}{89.9} = 0.5818 \]

Verify: confirm what the value says before any lookup

Why: A ratio of 0.58 means the first instructor's grade variance is about 58 percent of the second's, so their standard deviations are about 7.2 and 9.5 points. Whether that gap is more than two samples of thirty would produce by chance is the question, and with equal degrees of freedom the curve is centred near one — so 0.58 sits noticeably below where the null predicts.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691

9. Compute the statistic

Faded example

Samples of 16 and 21 give variances 40 and 25.

Fill in the blanks

F = \frac1.615 = ___ \text___ ___ \text___ 20 \text___

Why: A ratio above one puts the evidence in the upper tail, so a right-tailed alternative would be the one this statistic supports.

10. Worked example: why the population variances cancel

Worked example

Tracing the algebra under the null.

\[ H_0: \sigma_1^2 = \sigma_2^2 \]

The full ratio

Why: Each sample over its population.

Under the null

Why: The two are equal.

Substitute

Why: Both denominators the same.

What remains

Why: The sample variances.

\[ s 1\text{ squared over } s 2\text{ squared} \]

Figure (svg): The solution to Worked example why the population variances cancel shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{s_1^2/\sigma^2}{s_2^2/\sigma^2} = \frac{s_1^2}{s_2^2} \]

Verify: confirm why this makes the test possible at all

Why: The population variances are unknown — that is the whole reason for testing. The full statistic could not be computed from data, and it is only the null's assumption of equality that removes them. That is the same device section 10.4 used with a pooled proportion: assuming the null lets an unknown parameter be eliminated, and the resulting statistic is what the data can supply.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691

11. Trap: using n1 plus n2 minus two for the degrees of freedom

Trap

The trap

\[ \text{df} = n_1 + n_2 - 2 = 58 \]

Carry the two-sample t test's rule across

Why: Two samples of thirty were drawn.

\[ \text{but an F has TWO degrees of freedom} \]

Each sample contributes its own count, so the answer is a pair — 29 and 29 — rather than a single 58.

The fix

\[ F \sim F_{n_1-1,\,n_2-1} = F_{29,29} \]

Give one degrees of freedom per sample

Why: The numerator's from sample one, the denominator's from sample two.

The habit comes from chapter 10, where a pooled two-sample t test genuinely used n1 plus n2 minus two. The F distribution keeps the two counts separate because the numerator and denominator come from different samples and are not pooled — which is also why the order in which they are written matters.

12. One of these is false

Two truths and a lie

All three concern the statistic.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The population variances cancel under the null
  • C. Each sample contributes its own degrees of freedom
  • B. The degrees of freedom are n1 plus n2 minus two

Survives elimination: B

Why: The survivor is false and imports the pooled two-sample t rule. An F distribution has two separate degrees-of-freedom parameters, one per sample, and they are not added.

13. Which variance goes on top?

Prediction

Commit before reasoning.

Predict first

The book notes the ratio could be written either way up. What decides it?

  • The alternative hypothesis, and which sample variance is larger
  • The larger sample size
  • Alphabetical order
  • It never matters

Correct: The alternative, and which sample variance is larger.

Why: Putting the larger sample variance on top forces the ratio above one, which suits a two-tailed test. For a directional alternative the numbering has to match the hypothesis — Example 13.5 puts the first instructor's variance on top precisely because the claim is about the FIRST being smaller.

14. What ratio would equal variances give?

Estimation

Two samples from populations with the same variance.

Predict first

What F value would you expect?

  • Near one
  • Near zero
  • Near the sample size
  • Exactly one

Correct: Near one.

Why: The book says it directly: if the two populations have equal variances, the two sample variances are close in value and their ratio is close to one. Exactly one would be a coincidence, since sample variances vary from sample to sample even when their populations agree.

15. Choosing the tail

Section

Section 2

16. Left, right or two

Concept

If F is close to one, the evidence favours the null hypothesis that the two population variances are equal. If F is much larger than one, the evidence is against it. A test of two variances may be left-tailed, right-tailed or two-tailed.

reading the direction — A claim that the first variance is larger gives a right-tailed test; that it is smaller gives a left-tailed one; that it merely differs gives a two-tailed test.

\[ \sigma_1^2 > \sigma_2^2 \to \text{right}; \quad < \to \text{left}; \quad \ne \to \text{two} \]

This is the second time in the course a chi-square-family test has offered a choice of tail, and both times for the same reason. Section 11.6's test of a single variance could be left-tailed because a variance can be smaller than claimed as well as larger; the same holds when comparing two. Every other test in chapters 11 and 13 measures disagreement and can only be right-tailed.

Figure (svg): Three cards showing how the alternative selects the tail for a test of two variances

The same choice section 11.6's test of a single variance had, and for the same reason: a variance can be too small as well as too large.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691 — a test of two variances may be left, right, or two-tailed

17. Three alternatives, three tails

Picture it

With the wording that signals each.

Figure (svg): Three cards showing how the alternative selects the tail for a test of two variances

The same choice section 11.6's test of a single variance had, and for the same reason: a variance can be too small as well as too large.

Example 13.5 uses the middle card. The claim is that the first instructor's variance is SMALLER, so the alternative points below and the test is left-tailed — the chapter's only departure from the right tail.

18. Worked example: Example 13.5's hypotheses

Worked example

Test the claim that the first instructor's variance is smaller, at a 10 percent level.

\[ \text{the first variance is smaller} \]

The parameters

Why: Two population variances.

\[ \sigma 1\text{ squared and } \sigma 2\text{ squared} \]

The null

Why: An equality.

The claim

Why: First is smaller.

\[ \sigma 1\text{ squared below } \sigma 2\text{ squared} \]

The tail

Why: The alternative points down.

Figure (svg): The solution to Worked example Example 13.5's hypotheses shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \sigma_1^2 = \sigma_2^2; \quad H_a: \sigma_1^2 < \sigma_2^2 \]

Verify: confirm the tail matches which variance was placed on top

Why: Because the first instructor's variance sits in the numerator, a smaller first variance produces a smaller ratio — so the alternative's downward direction corresponds to the lower tail of the F curve. Had the second variance been placed on top, the same claim would have produced a right-tailed test with a statistic of 1.719. Both are correct; mixing them is not.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691

19. Which tail?

Sorting

Each is a claim, with the first sample's variance in the numerator.

Sort into buckets

Sort by the tail the test uses.

Left-tailed
the first instructor's variance is smaller; the lid varies less than the container
Right-tailed
the new process is more variable than the old
Two-tailed
the two checkers' times differ in variability; the two groups' variances are not the same
left
The alternative claims the numerator's variance is smaller.
right
The alternative claims the numerator's variance is larger.
two
The alternative claims only a difference, with no direction.

Item (a) is Example 13.5 and item (c) is the supermarket motivation the book gives in its opening paragraph. The two-tailed cases are the ones where either instructor being more variable would matter.

20. Worked example: the same data, both ways up

Worked example

Example 13.5's variances, with the ratio inverted.

\[ \frac{89.9}{52.3} \]

The inverted ratio

Why: Second over first.

\[ 1.719 \]

The hypotheses

Why: Now about the second.

The tail

Why: The alternative points up.

The p-value

Why: The upper tail beyond 1.719.

\[ 0.0753 \]

Figure (svg): The solution to Worked example the same data, both ways up shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ P(F_{29,29} > 1.719) = P(F_{29,29} < 0.5818) = 0.0753 \]

Verify: confirm the two arrangements must agree

Why: Inverting an F statistic and reversing the tail gives the same probability, since one over an F on df1 and df2 has an F distribution on df2 and df1. So the choice of which variance goes on top cannot change the conclusion, provided the tail is reversed with it — and a p-value that differs between the two arrangements means the tail was not reversed.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692

21. Error analysis: four setups for a claim that variance one is smaller

Error analysis

With the first variance in the numerator. Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; H_a: \sigma_1^2 < \sigma_2^2, \text{ left-tailed} \\ &(2)\; H_a: \sigma_1^2 < \sigma_2^2, \text{ right-tailed} \\ &(3)\; H_a: s_1^2 < s_2^2, \text{ left-tailed} \\ &(4)\; H_0: \sigma_1^2 < \sigma_2^2, \text{ left-tailed} \end{aligned} \)

  • (1) is correct and is Example 13.5's setup.
  • (2) has the right hypotheses and the wrong tail. With the first variance on top, a smaller first variance makes F small, so the evidence is in the lower tail.
  • (3) uses SAMPLE variances. Hypotheses concern population parameters; the sample variances are the evidence.
  • (4) makes the alternative into the null. The null must be the equality, since only that generates a distribution for the statistic.

Error (2) is the one to guard against, because the arithmetic is identical and only the tail differs. Taking the right tail here would give a p-value of 0.9247 and reverse the conclusion entirely.

22. Example 13.5's p-value

Faded example

The statistic is 0.5818 and the test is left-tailed.

Fill in the blanks

p\text< = P(F 0.0753 0.5818) = ___

Why: Taking the right tail instead would give 0.9247 and reverse the conclusion — the same error section 11.6 warned about for a single variance.

23. One of these is false

Two truths and a lie

All three concern the tail.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. This test may be left, right or two-tailed
  • C. An F near one favours the null
  • B. Like the ANOVA, this test is always right-tailed

Survives elimination: B

Why: The survivor is false and is the section's main structural difference from the rest of the chapter. A variance can be smaller than another as well as larger, so both directions carry meaning — exactly as in section 11.6.

24. Why does this test have a choice of tail?

Prediction

Commit before reasoning.

Predict first

The one-way ANOVA is always right-tailed and this test is not. Why?

  • Its statistic measures a ratio of two variances, and either can be the larger
  • Because it uses two samples
  • Because the F distribution is skewed
  • Because alpha is 10 percent

Correct: Either variance can be the larger.

Why: An ANOVA's numerator can only be inflated by differing means, so its alternative pushes in one direction. Here the alternative may claim the first variance is larger or smaller, and both are meaningful — the same reason section 11.6's single-variance test could be left-tailed while sections 11.2 to 11.4 could not.

25. The complete test

Section

Section 3

26. Example 13.5, end to end

Concept

Two college instructors each grade the same 30 exams. The first's grades have a variance of 52.3 and the second's 89.9. Testing the claim that the first instructor's variance is smaller, at a 10 percent level, gives F = 0.5818 on 29 and 29 degrees of freedom with a left-tail p-value of 0.0753.

the conclusion — With a 10 percent level of significance, from the data, there is sufficient evidence to conclude that the variance in grades for the first instructor is smaller.

\[ F_{29,29} = 0.5818, \quad p = 0.0753 < 0.10 \]

The book's motivating remark is worth keeping: in most colleges it is desirable for the variances of exam grades to be nearly the same among instructors. So a rejection here is not good news about the first instructor — it says the two are grading with different consistency, which is the thing the administration wanted to avoid.

Figure (svg): An F curve on twenty-nine and twenty-nine degrees of freedom with a dashed line at one and the area to the left of 0.5818 shaded

The book's Figure 13.6. With equal degrees of freedom the curve is centred near one, and the statistic falls well below it.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692 — Example 13.5 in full

27. Example 13.5 decided

Picture it

The book's Figure 13.6.

Figure (svg): An F curve on twenty-nine and twenty-nine degrees of freedom with a dashed line at one and the area to the left of 0.5818 shaded

The book's Figure 13.6. With equal degrees of freedom the curve is centred near one, and the statistic falls well below it.

The shaded region is the LEFT tail, at 0.0753. That is below a 10 percent level and above a 5 percent one, so this conclusion depends on an unusually generous alpha having been fixed in advance.

28. Worked example: the full write-up

Worked example

Example 13.5, in the book's own sequence.

\[ \alpha = 0.10 \]

Hypotheses

Why: First smaller.

Statistic

Why: 52.3 over 89.9.

\[ 0.5818 \]

Distribution

Why: Both samples of 30.

\[ F s u b 29, 29 \]

Probability statement

Why: The lower tail.

\[ 0.0753 \]

Compare and decide

Why: 0.0753 below 0.10.

Figure (svg): The solution to Worked example the full write-up shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ p = P(F < 0.5818) = 0.0753 < 0.10 \]

Verify: confirm how sensitive this conclusion is to the level

Why: At 5 percent the same p-value of 0.0753 would not reject, so the conclusion rests entirely on the 10 percent level having been chosen in advance. Ten percent is a generous level by the standards of the rest of the course, where 5 percent was the default and 1 percent appeared several times — which makes it worth stating prominently rather than in passing.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692

29. Stage to content

Matching

Match each stage of Example 13.5.

Match the pairs

  • l1. hypotheses
  • l2. test statistic
  • l3. distribution for the test
  • l4. probability statement
  • r1. variances equal, against the first smaller
  • r2. F = 0.5818
  • r3. F sub 29, 29
  • r4. P(F < 0.5818) = 0.0753

Why: The same four stages as every test since chapter 9. Only the fourth carries anything unusual — the inequality points the other way from every other test in this chapter.

30. Worked example: what would change at 5 percent

Worked example

The same data, at the course's more usual level.

\[ \alpha = 0.05 \]

Compare

Why: 0.0753 against 0.05.

The decision

Why: Not below alpha.

The conclusion

Why: Insufficient evidence.

The lesson

Why: Fix alpha first.

Figure (svg): The solution to Worked example what would change at 5 percent shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 0.05 < 0.0753 < 0.10 \]

Verify: confirm this is the third time the chapter has raised the point

Why: Example 13.2's p-value of 0.0248 fell between 1 and 5 percent, Example 13.3's choice of 1 percent turned out not to matter, and this one falls between 5 and 10. Chapter 9 required the level to be fixed before the data are seen, and the chapter's own examples show how often a result lands close enough to a boundary for that requirement to bite.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 692

31. Trap: taking the right tail out of habit

Trap

The trap

\[ p = P(F > 0.5818) = 0.9247 \]

Use the upper tail, as the ANOVA does

Why: Every other test in the chapter is right-tailed.

\[ \text{but the alternative says SMALLER} \]

A p-value of 0.92 would mean failing to reject, and the study's actual finding would be missed.

The fix

\[ p = P(F < 0.5818) = 0.0753 \]

Read the tail from the alternative

Why: This is the chapter's only test where the choice exists.

The error is dangerous because 0.9247 is a perfectly ordinary-looking p-value leading to a perfectly ordinary do-not-reject. Nothing in the arithmetic flags it, and the only defence is noticing before computing that the statistic came out below one and the alternative points downward — which are the same observation.

32. The degrees of freedom

Faded example

Both instructors graded 30 exams.

Fill in the blanks

F \sim F_29,\,29}

Why: Equal sample sizes give equal degrees of freedom, which makes the curve nearly centred on one — so the statistic's distance below one is easy to read.

33. One of these is false

Two truths and a lie

All three concern the test.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It rejects at 10 percent but would not at 5
  • C. The conclusion is that the first instructor's variance is smaller
  • B. The conclusion is that the first instructor grades more accurately

Survives elimination: B

Why: The survivor is false. A smaller variance means more consistent grades, not more accurate ones — the first instructor could be consistently generous or consistently harsh. Variance describes spread, and nothing in this test concerns the average grade at all.

34. Is a rejection here good news?

Prediction

Commit before reasoning.

Predict first

The book notes that colleges want instructors' grade variances to be nearly the same. What does rejecting mean for them?

  • That the two instructors grade with different consistency, which is what they wanted to avoid
  • That the first instructor is better
  • That the grades are accurate
  • Nothing of practical interest

Correct: The two grade with different consistency.

Why: The null of equal variances is the desirable state here, so rejecting it identifies a problem rather than a finding. That reverses the usual framing — in most of the course rejecting was the interesting outcome — and it is a useful reminder that which hypothesis is desirable depends entirely on the context.

35. The warning

Section

Section 4

36. Very sensitive to non-normality

Concept

Unlike most other tests in this book, the F test for equality of two variances is very sensitive to deviations from normality. If the two distributions are not normal, the test can give higher p-values than it should, or lower ones, in ways that are unpredictable. Many texts suggest that students not use this test at all, but in the interest of completeness we include it here.

unpredictable — The key word. A test whose error runs in a known direction can be interpreted with caution; one whose error direction depends on the unknown shape of the populations cannot.

\[ \text{non-normal} \;\Rightarrow\; p \text{ too high OR too low} \]

No other test in the course carries a caveat of this strength, and the contrast with the t procedures is the point. Chapter 8's t interval and chapter 10's two-sample tests become robust as samples grow, because the central limit theorem describes sample MEANS. Nothing plays that role for sample variances, which is the same limitation section 11.6 noted for a single variance — and here it applies to two at once.

Figure (svg): A card reproducing the book's warning about this test's sensitivity to non-normality

No other test in the course carries a caveat this strong, and it is worth reading before the arithmetic.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690 — the paragraph of warning after the requirements

37. The warning, as the book states it

Picture it

Reproduced rather than summarised.

Figure (svg): A card reproducing the book's warning about this test's sensitivity to non-normality

No other test in the course carries a caveat this strong, and it is worth reading before the arithmetic.

The phrase many texts suggest that students not use this test at all is remarkable in a textbook that is teaching the test. It is included for completeness, and the honest way to learn it is with the caveat attached.

38. Worked example: reading the two requirements

Worked example

The book lists two conditions before the warning.

\[ \text{the requirements} \]

First

Why: Both populations normal.

Second

Why: The two independent.

Which is fragile

Why: The first.

What secures each

Why: Design against shape.

Figure (svg): The solution to Worked example reading the two requirements shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{normal populations}; \quad \text{independent samples} \]

Verify: confirm why the asymmetry matters practically

Why: Independence is a property of the sampling procedure, so a well-designed study guarantees it. Normality is a property of the populations themselves, which no design can impose — it can only be examined after the fact, and with small samples not examined well. So the fragile requirement is also the one hardest to check, which is the substance of the book's recommendation.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690

39. One of these is false

Two truths and a lie

All three concern the warning.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The error can run in either direction
  • C. Many texts suggest not using this test at all
  • B. A large sample makes the test robust

Survives elimination: B

Why: The survivor is false. The central limit theorem normalises sample means, not sample variances, so the distribution of this statistic keeps depending on the populations' shape however large the samples are.

40. Worked example: comparing with the t procedures

Worked example

Why other tests tolerate non-normality and this one does not.

\[ \bar{x} \text{ against } s^2 \]

A sample mean

Why: A scaled sum.

\[ \text{chapter } 7\text{ applies} \]

Its behaviour for large n

Why: Approximately normal.

A sample variance

Why: Not a sum of the data.

Its behaviour for large n

Why: Depends on shape.

Figure (svg): The solution to Worked example comparing with the t procedures shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \bar{X} \approx N \text{ for large } n; \quad s^2 \text{ has no such guarantee} \]

Verify: confirm this is the same point section 11.6 made

Why: There, the test of a single variance assumed normality and the assumption did not soften with sample size. Here the same limitation applies to two sample variances simultaneously, and the departures can compound — which is presumably why the warning appears in this section and not that one, and why it is worded so much more strongly.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690

41. Error analysis: four responses to the warning

Error analysis

Which are consistent with what the book says?

Annotate

On: \( \begin{aligned} &(1)\; \text{check both populations for normality first} \\ &(2)\; \text{use a large sample, which removes the concern} \\ &(3)\; \text{treat a } p \text{-value near alpha with particular caution} \\ &(4)\; \text{ignore it, since the book teaches the test} \end{aligned} \)

  • (1) is the right first response, and the requirements list it before anything else.
  • (2) is false. The central limit theorem covers means, not variances, so no sample size repairs it.
  • (3) is sound. If the p-value could be too high or too low, a result near the decision boundary is the least trustworthy kind.
  • (4) misreads the inclusion. The book says it is included in the interest of completeness while noting many texts suggest not using it — that is an argument for caution, not permission to skip the caveat.

Response (3) is the practical one. Example 13.5's own p-value of 0.0753 against a 10 percent level is exactly the kind of near-boundary result the warning makes hardest to trust.

42. Robust to non-normality?

Sorting

Each is a procedure from the course.

Sort into buckets

Sort by whether large samples repair non-normality.

Large samples help
a t interval for a single mean; a two-sample t test for means; a z test for a proportion
They do not
an F test of two variances; a chi-square test of a single variance
yes
The statistic is built from a mean or a proportion, which the central limit theorem normalises.
no
The statistic is built from variances, which the theorem says nothing about.

The division is clean and follows one principle: chapter 7's theorem covers sums and means. Every procedure it does not cover carries a normality assumption that no sample size removes.

43. Why does unpredictable matter so much?

Prediction

Commit before reasoning.

Predict first

The book says the error can go either way, unpredictably. Why is that worse than a known bias?

  • A known direction could be allowed for; an unknown one cannot
  • It makes the arithmetic harder
  • It only affects large samples
  • It is not worse

Correct: A known direction could be allowed for.

Why: If a test were known to give p-values that are too small, a reader could demand a stricter threshold and proceed. When the direction depends on the unknown shape of the populations, no adjustment is available — which is what turns a caveat into a recommendation against use.

44. Explain the caveat

Explain it

A classmate says the warning cannot be serious, since the book teaches the test and works an example.

Discussion prompt

In two sentences or fewer, explain what the warning means in practice.

Hint: Ask what they would do with a p-value of 0.06 from non-normal data.

Answer:

The book includes it for completeness while noting that many texts suggest students not use it, so the right response is to check both populations for normality before trusting any p-value it produces.

A result near the decision boundary is the least trustworthy kind here, because the p-value could be too high or too low and nothing tells you which.

45. The chapter's two uses, side by side

Section

Section 5

46. One distribution, two questions

Concept

The fifth fact about the F distribution named its other uses, and this is one of them. A one-way ANOVA compares several means by way of variances; a test of two variances compares the variances themselves.

the sharpest contrast — One-way ANOVA ASSUMES the group variances are equal in order to conclude about means. This test makes that same equality its hypothesis.

\[ \text{ANOVA}: \; \sigma^2 \text{ assumed equal}; \qquad \text{here}: \; \sigma_1^2 = \sigma_2^2 \text{ tested} \]

The relationship between them is practical as well as conceptual. Since a one-way ANOVA requires equal variances, a test of two variances is one way of examining that requirement — though the warning in this section means it should be used cautiously for that purpose, and a look at side-by-side box plots is often more informative.

Figure (svg): A table comparing the chapter's two uses of the F distribution

One distribution, two questions — which is the book's fifth fact about the F distribution, made concrete.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — the chapter's two uses of the F distribution

47. The two uses

Picture it

Six rows of comparison.

Figure (svg): A table comparing the chapter's two uses of the F distribution

One distribution, two questions — which is the book's fifth fact about the F distribution, made concrete.

Every row differs, which is a fair summary of how much a single distribution can be made to do. Only the distribution itself is shared, and even the two degrees of freedom are counted differently.

48. Worked example: choosing between them

Worked example

Three situations.

\[ \text{which test?} \]

Five mulches, mean yields

Why: Several means.

Two instructors, grade spread

Why: Two variances.

Four sororities, mean grades

Why: Several means.

The signal

Why: Means or spreads.

Figure (svg): The solution to Worked example choosing between them shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{means} \to \text{ANOVA}; \quad \text{two variances} \to \text{this test} \]

Verify: confirm the signal is the question rather than the data

Why: The same two columns of numbers could support either test — the instructors' grades have means as well as variances, and the mulch yields have variances as well as means. What decides is which quantity the question asks about, exactly as section 11.5 found for the chi-square tests. The data alone never settle it.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692

49. Which use of the F distribution?

Sorting

Each describes a question.

Sort into buckets

Sort by which test applies.

One-way ANOVA
do five mulches give the same mean yield?; do four sororities have the same mean grade?; do three growing media give the same mean height?
Test of two variances
do two instructors grade with the same consistency?; does a lid vary as much as its container?
an
Several means are being compared.
va
Two variances are being compared directly.

Items (b) and (d) are two of the three motivations the book gives in its opening paragraph, alongside the variability of two supermarket checkers' times.

50. Worked example: the degrees of freedom in each

Worked example

Two groups of thirty, compared two ways.

\[ n_1 = n_2 = 30 \]

As an ANOVA

Why: k = 2, n = 60.

\[ 1\text{ and } 58 \]

As two variances

Why: Each sample separately.

\[ 29\text{ and } 29 \]

Note the difference

Why: Same data.

Why

Why: Different questions.

Figure (svg): The solution to Worked example the degrees of freedom in each shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F_{1,58} \quad\text{against}\quad F_{29,29} \]

Verify: confirm why the same data give different counts

Why: An ANOVA pools all sixty observations to estimate one within-group variance, spending two degrees of freedom on the two group means; the variance test keeps the samples separate and each variance carries its own count. The two statistics answer different questions from the same data, so there is no reason for their degrees of freedom to agree.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692

51. Trap: using this test to justify the ANOVA assumption uncritically

Trap

The trap

\[ \text{variance test does not reject} \;\Rightarrow\; \text{ANOVA's equal-variance assumption holds} \]

Take a non-rejection as confirmation

Why: It is the relevant hypothesis.

\[ \text{but failing to reject never establishes a null} \]

And this particular test is the one the book warns is least trustworthy.

The fix

\[ \text{examine the group spreads directly, e.g. with box plots} \]

Treat the assumption as something to inspect, not to test away

Why: A non-rejection from a fragile test is weak evidence.

Two problems compound here. Failing to reject never proves a null, as chapter 9 established; and this test's p-value may be too high or too low unpredictably when the populations are not normal — which is exactly the situation in which the equal-variance assumption is most likely to be in doubt. Looking at the data remains the better check.

52. One of these is false

Two truths and a lie

All three compare the two uses.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. ANOVA assumes equal variances; this test hypothesises them
  • C. They count degrees of freedom differently
  • B. Both are always right-tailed

Survives elimination: B

Why: The survivor is false. The ANOVA is always right-tailed, but a test of two variances may be left, right or two-tailed — and Example 13.5 is left-tailed.

53. Two counts from one dataset

Faded example

Two samples of 25 observations each.

Fill in the blanks

\text48 F_24}}; \qquad \text___ F____,24}

Why: The same fifty observations give 1 and 48 under one test and 24 and 24 under the other, because the two statistics are built from different quantities.

54. What does the fifth fact point to?

Prediction

Commit before reasoning.

Predict first

Section 13.3's fifth fact named other uses for the F distribution. Which are they?

  • Comparing two variances, and two-way analysis of variance
  • Comparing two means
  • Testing a correlation
  • Testing a proportion

Correct: Comparing two variances and two-way ANOVA.

Why: The first is this section; the second is named and set aside as beyond the scope of the chapter. Between them they explain why the distribution earns a chapter of its own rather than appearing only as an ANOVA detail.

55. This test against section 11.6's

Comparison

Fill the blanks. Both concern variances, and both can go either way.

Comparison matrix

One variance (11.6)Two variances (13.4)
Compared againsta claimed valueanother sample's variance
Distributionchi-squareF
Degrees of freedomn - 1n1 - 1 and n2 - 1
Tails availableleft, right or twoleft, right or two

The last row is what sets both apart from every other test in chapters 11 and 13. A statistic that measures a ratio of variances can fall meaningfully on either side of its null value, while one that measures disagreement cannot.

56. Running a test of two variances, in order

Pattern

Six steps, and the first is the one the book warns about most.

  1. Check that both populations are normal — this test is very sensitive to departures from it.
  2. Check that the two samples are independent.
  3. State the hypotheses about the two population variances, with the null as equality.
  4. Decide which sample variance goes in the numerator, consistently with the alternative.
  5. Compute F as that ratio, with degrees of freedom one less than each sample size.
  6. Take the area in the tail the alternative points to, compare with alpha, decide, and conclude in context.

If the statistic comes out below one, the evidence is in the lower tail — do not take the upper tail out of habit.

OpenStax Introductory Business Statistics 2e, §12.1 Test of Two Variances §12.1 Test of Two Variances

57. Check yourself 1 of 3

Check

The statistic.

Check your understanding

Samples of 20 and 25 give variances 30 and 12. What is F, and on what degrees of freedom?

  • A. 2.5 on 19 and 24 (correct)
  • B. 2.5 on 20 and 25
  • C. 0.4 on 19 and 24
  • D. 2.5 on 43

Answer: A

Why: Thirty over twelve is 2.5, and the degrees of freedom are one less than each sample size.

Why B tempts people
The degrees of freedom are n minus one for each sample, not n itself.
Why C tempts people
That inverts the ratio; with the first variance named first it goes on top.
Why D tempts people
An F distribution has two degrees-of-freedom parameters, not one combined count.

58. Check yourself 2 of 3

Check

The tail.

Check your understanding

With the first sample's variance on top, the claim is that the first population is LESS variable. Which tail?

  • A. Left-tailed (correct)
  • B. Right-tailed
  • C. Two-tailed
  • D. Right-tailed, as all F tests are

Answer: A

Why: A smaller numerator variance gives a smaller ratio, so the alternative points below one and the p-value is a left-tail area.

Why B tempts people
That would test whether the first is MORE variable.
Why C tempts people
The claim has a direction, so a two-tailed test would not match it.
Why D tempts people
A one-way ANOVA is always right-tailed; a test of two variances is not.

59. Check yourself 3 of 3

Check

The warning.

Check your understanding

The two populations are clearly skewed. What does the book's warning imply?

  • A. The p-value may be too high or too low, unpredictably (correct)
  • B. The p-value will be too small
  • C. A larger sample will fix it
  • D. The test becomes two-tailed

Answer: A

Why: The book says the test can give higher p-values than it should, or lower ones, in ways that are unpredictable — which is why many texts suggest not using it at all.

Why B tempts people
The direction of the error is precisely what cannot be predicted.
Why C tempts people
The central limit theorem covers means, not variances, so sample size does not repair it.
Why D tempts people
The tail follows from the alternative hypothesis and has nothing to do with normality.

60. Where this shows up outside the textbook

Real world

A manufacturer compares two machines' fill volumes, taking 25 bottles from each. The sample variances are 4.1 and 1.9, giving F = 2.16 with a two-tailed p-value of 0.058, and the report concludes that the two machines are equally consistent and no maintenance is needed.

Discussion prompt

Assess the arithmetic, the conclusion and the assumptions.

Hint: Check the decision, then ask what a non-rejection establishes.

Answer:

The arithmetic is right. Twenty-five bottles from each machine give 24 and 24 degrees of freedom, and 4.1 over 1.9 is 2.16 — whose two-tailed p-value on that curve is close to the reported 0.058. At a 5 percent level this does not reject.

\[ F_{24,24} = \frac{4.1}{1.9} = 2.16, \qquad p \approx 0.058 \]

The conclusion over-claims twice. Failing to reject never establishes that the variances are equal, and this p-value is barely above 5 percent — the data are close to indicating a difference rather than supporting equality. The sample variances differ by more than a factor of two, which in a filling operation is a substantial practical gap whatever the test says.

And the assumption is the fragile one. Fill volumes from a drifting or occasionally double-filling machine are not normal, and those are exactly the failure modes that would inflate a variance. The book's warning applies with full force: the p-value could be too high or too low, unpredictably — and at 0.058, a result this close to the boundary is the least trustworthy kind.

What should happen: plot the fifty measurements by machine, as histograms and in production order, to check the shape and look for drift. Report that machine one's sample variance is more than double machine two's, that the difference is borderline at conventional levels, and that the normality the test requires is doubtful. The honest recommendation is to collect more data, not to close the question — which is close to what the book means by suggesting many texts avoid this test entirely.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why does the book warn so strongly about this test?

  • Because the arithmetic is difficult
  • Because it is very sensitive to non-normality, and the resulting error can go either way unpredictably
  • Because it requires large samples
  • Because the F distribution is skewed

Correct: It is very sensitive to non-normality, unpredictably.

\[ \text{non-normal} \;\Rightarrow\; p \text{ biased in an unknown direction} \]

Why: Unlike most other tests in the book, this one's p-value can be too high or too low when the populations are not normal, and nothing indicates which. Sample size does not help, because the central limit theorem normalises sample means and says nothing about sample variances. The book includes the test in the interest of completeness while noting that many texts suggest students not use it at all.

62. Explain it to someone a year behind you

Explain it

They computed F = 0.58 for a claim that the first variance is smaller, then took the right tail and got 0.92.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask which direction the alternative points.

Answer:

The alternative says the first variance is SMALLER, and since it sits in the numerator that makes the ratio small — so the evidence is in the LEFT tail, at 0.0753 rather than 0.92.

This is the chapter's only test where the tail is a choice, so it is the only place the habit of taking the upper tail goes wrong.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck — and the course.

Predict first

Which of these would you least want handed to you cold?

  • Computing the statistic and its two degrees of freedom
  • Choosing the tail from the alternative hypothesis
  • Running the complete test and concluding in context
  • Stating the warning and what it means in practice

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, the ratio of sample variances on one less than each sample size. For the second, the direction the alternative points, given which variance is on top. For the third, the same six stages as every test since chapter 9. For the fourth, very sensitive to non-normality, with the error running either way. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes, and it closes the course.

Draw it

At the top, write the full F ratio with each sample variance over its own population variance, then show the population variances cancelling under the null to leave the ratio of sample variances alone. Beside it write the degrees of freedom as n1 minus one and n2 minus one, and note that they are not added. Under that, draw three boxes for the three alternatives — first variance greater, less, and not equal — with the tail each selects, and circle the middle one as Example 13.5's. In the middle, work Example 13.5 completely: variances 52.3 and 89.9 from samples of 30, the ratio 0.5818, the distribution F sub 29 comma 29, and the LEFT-tail p-value of 0.0753 against a 10 percent level, giving a rejection. Draw the curve beside it with a dashed line at one and the area to the LEFT of 0.5818 shaded, and write that taking the right tail would give 0.9247 and the opposite conclusion. At the bottom left, box the book's warning in full: very sensitive to deviations from normality, p-values too high or too low unpredictably, and many texts suggest not using it at all. At the bottom right, make a six-row table comparing this test with one-way ANOVA, and underline the last row — ANOVA assumes equal variances, this test tests them.

Check your Example 13.5 work by confirming the inverted ratio 89.9 over 52.3 gives 1.719 and its RIGHT tail gives the same 0.0753 — the two arrangements must agree. Check your warning box by asking whether you could say, in one sentence, what you would do differently on encountering non-normal data.

65. What you can do now

Recap

Six things, and the last one is what the book most wants carried away.

If you seeThen
Two sample variances to compareF is their ratio, on n1-1 and n2-1
A claim that one variance is largerRight-tailed, with that variance on top
A claim that one is smallerLeft-tailed, with that variance on top
A claim they merely differTwo-tailed
F near oneThe evidence favours equal variances
F below one with a 'smaller' alternativeThe lower tail, not the upper
Non-normal populationsThe p-value is unreliable in an unknown direction
A near-boundary p-value from this testThe least trustworthy kind of result

That completes chapter 13 and the course. Thirteen chapters have moved from describing a single sample, through the distributions that describe chance, to inference about one parameter, two, and finally several at once — and the last two chapters showed the same F and chi-square distributions doing several different jobs. Section 13.5 is the book's own lab, which puts one-way ANOVA to work on data you collect yourself.

OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 690-692
  2. OpenStax Introductory Business Statistics 2e, §12.1 Test of Two Variances — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108