The chapter's second use of the F distribution, and the course's last section. It is often desirable to compare two variances rather than two averages — college administrators would like two professors to have the same variation in their grading, a lid and a container must vary alike to fit, a supermarket may care about the variability of two checkers' times. The statistic is the ratio of the two sample variances, since the population variances cancel under a null of equality, and its degrees of freedom are one less than each sample size. Because a variance can be larger or smaller than another, the test may be left-tailed, right-tailed or two-tailed. The section also carries a warning unlike anything else in the book: this test is very sensitive to deviations from normality, can give p-values too high or too low in unpredictable ways, and many texts suggest students not use it at all — a caveat this lesson treats as the section's most important content.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 13 — F Distribution and One-Way ANOVA
Test of Two Variances
Objectives
Six outcomes, and the last is the one the book most wants remembered.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — the section these objectives are drawn from
Warm-up
Section 11.6 tested one variance against a claimed value; section 13.2 built a ratio of two variance estimates.
Discussion prompt
Two instructors grade the same 30 exams. The first's grades have variance 52.3 and the second's 89.9. How would you test whether the two populations have the same variance?
Hint: You have two sample variances and a distribution built for ratios.
Answer:
Take their ratio. If the two population variances are genuinely equal, then two samples of the same size should give sample variances close in value and their ratio close to one.
\[ F = \frac{s_1^2}{s_2^2} = \frac{52.3}{89.9} = 0.5818 \]
A ratio of 0.58 means the first instructor's grades varied only about 58 percent as much as the second's. Whether that is more than sampling variation would produce is what the F distribution answers, with 29 degrees of freedom for each sample.
Note the direction. Here the ratio came out BELOW one, so the evidence points toward the first variance being smaller — and the test that examines it is left-tailed. That option did not exist for the one-way ANOVA, and this section is where it returns.
Concept
Another of the uses of the F distribution is testing two variances. Since we are interested in comparing the two sample variances, we use the F ratio, which has the distribution F with n1 minus one and n2 minus one degrees of freedom. If the null hypothesis is that the two population variances are equal, the F ratio becomes the ratio of the two sample variances.
why the ratio works — If the two populations have equal variances, the two sample variances are close in value and their ratio is close to one. If the population variances are very different, the sample variances tend to be very different too.
\[ F = \frac{s_1^2}{s_2^2}, \qquad F \sim F_{n_1-1,\,n_2-1} \]
The full statistic divides each sample variance by its own population variance before taking the ratio. Under a null of equality those two population variances are the same and cancel, leaving the sample variances alone — which is why the test is computable at all, since the population variances are unknown.
Figure (svg): A card giving the F ratio for a test of two variances
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691
Section
Section 1
Concept
The F ratio divides each sample variance by its population variance and takes the quotient. If the null hypothesis is that the two population variances are equal, they cancel and the F ratio becomes the ratio of the two sample variances alone.
two degrees of freedom — n1 minus one for the numerator and n2 minus one for the denominator — one less than each sample size, since each sample variance is computed about its own sample mean.
\[ F = \frac{s_1^2/\sigma_1^2}{s_2^2/\sigma_2^2} \;\overset{H_0}{=}\; \frac{s_1^2}{s_2^2} \]
The book adds a note about which variance goes on top: the F ratio could also be written the other way up, and it depends on the alternative hypothesis and on which sample variance is larger. Putting the larger one on top forces the ratio above one, which is convenient for a two-tailed test but reverses the tail for a directional one — so the choice has to be made deliberately.
Figure (svg): A card giving the F ratio for a test of two variances
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691 — the F ratio, the cancellation, and the note on which variance goes on top
Picture it
Before and after the cancellation.
Figure (svg): A card giving the F ratio for a test of two variances
The lower panels give the two degrees of freedom, and the yellow note is the book's own: which sample variance sits on top is a choice, and it interacts with the direction of the alternative.
Worked example
Two instructors each grade the same 30 exams. The first's grades have variance 52.3 and the second's 89.9.
\[ s_1^2 = 52.3, \; s_2^2 = 89.9, \; n_1 = n_2 = 30 \]
The ratio
Why: First over second.
\[ 52.3\text{ over } 89.9 \]
Evaluate
Why: Divide.
\[ 0.5818 \]
Numerator df
Why: n1 minus one.
\[ 29 \]
Denominator df
Why: n2 minus one.
\[ 29 \]
Figure (svg): The solution to Worked example Example 13.5's statistic shown as a ladder of expressions, one row per legal move
\[ F_{29,29} = \frac{52.3}{89.9} = 0.5818 \]
Verify: confirm what the value says before any lookup
Why: A ratio of 0.58 means the first instructor's grade variance is about 58 percent of the second's, so their standard deviations are about 7.2 and 9.5 points. Whether that gap is more than two samples of thirty would produce by chance is the question, and with equal degrees of freedom the curve is centred near one — so 0.58 sits noticeably below where the null predicts.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691
Faded example
Samples of 16 and 21 give variances 40 and 25.
Fill in the blanks
F = \frac1.615 = ___ \text___ ___ \text___ 20 \text___
Why: A ratio above one puts the evidence in the upper tail, so a right-tailed alternative would be the one this statistic supports.
Worked example
Tracing the algebra under the null.
\[ H_0: \sigma_1^2 = \sigma_2^2 \]
The full ratio
Why: Each sample over its population.
Under the null
Why: The two are equal.
Substitute
Why: Both denominators the same.
What remains
Why: The sample variances.
\[ s 1\text{ squared over } s 2\text{ squared} \]
Figure (svg): The solution to Worked example why the population variances cancel shown as a ladder of expressions, one row per legal move
\[ \frac{s_1^2/\sigma^2}{s_2^2/\sigma^2} = \frac{s_1^2}{s_2^2} \]
Verify: confirm why this makes the test possible at all
Why: The population variances are unknown — that is the whole reason for testing. The full statistic could not be computed from data, and it is only the null's assumption of equality that removes them. That is the same device section 10.4 used with a pooled proportion: assuming the null lets an unknown parameter be eliminated, and the resulting statistic is what the data can supply.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-691
Trap
\[ \text{df} = n_1 + n_2 - 2 = 58 \]
Carry the two-sample t test's rule across
Why: Two samples of thirty were drawn.
\[ \text{but an F has TWO degrees of freedom} \]
Each sample contributes its own count, so the answer is a pair — 29 and 29 — rather than a single 58.
\[ F \sim F_{n_1-1,\,n_2-1} = F_{29,29} \]
Give one degrees of freedom per sample
Why: The numerator's from sample one, the denominator's from sample two.
The habit comes from chapter 10, where a pooled two-sample t test genuinely used n1 plus n2 minus two. The F distribution keeps the two counts separate because the numerator and denominator come from different samples and are not pooled — which is also why the order in which they are written matters.
Two truths and a lie
All three concern the statistic.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and imports the pooled two-sample t rule. An F distribution has two separate degrees-of-freedom parameters, one per sample, and they are not added.
Prediction
Commit before reasoning.
Predict first
The book notes the ratio could be written either way up. What decides it?
Correct: The alternative, and which sample variance is larger.
Why: Putting the larger sample variance on top forces the ratio above one, which suits a two-tailed test. For a directional alternative the numbering has to match the hypothesis — Example 13.5 puts the first instructor's variance on top precisely because the claim is about the FIRST being smaller.
Estimation
Two samples from populations with the same variance.
Predict first
What F value would you expect?
Correct: Near one.
Why: The book says it directly: if the two populations have equal variances, the two sample variances are close in value and their ratio is close to one. Exactly one would be a coincidence, since sample variances vary from sample to sample even when their populations agree.
Section
Section 2
Concept
If F is close to one, the evidence favours the null hypothesis that the two population variances are equal. If F is much larger than one, the evidence is against it. A test of two variances may be left-tailed, right-tailed or two-tailed.
reading the direction — A claim that the first variance is larger gives a right-tailed test; that it is smaller gives a left-tailed one; that it merely differs gives a two-tailed test.
\[ \sigma_1^2 > \sigma_2^2 \to \text{right}; \quad < \to \text{left}; \quad \ne \to \text{two} \]
This is the second time in the course a chi-square-family test has offered a choice of tail, and both times for the same reason. Section 11.6's test of a single variance could be left-tailed because a variance can be smaller than claimed as well as larger; the same holds when comparing two. Every other test in chapters 11 and 13 measures disagreement and can only be right-tailed.
Figure (svg): Three cards showing how the alternative selects the tail for a test of two variances
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691 — a test of two variances may be left, right, or two-tailed
Picture it
With the wording that signals each.
Figure (svg): Three cards showing how the alternative selects the tail for a test of two variances
Example 13.5 uses the middle card. The claim is that the first instructor's variance is SMALLER, so the alternative points below and the test is left-tailed — the chapter's only departure from the right tail.
Worked example
Test the claim that the first instructor's variance is smaller, at a 10 percent level.
\[ \text{the first variance is smaller} \]
The parameters
Why: Two population variances.
\[ \sigma 1\text{ squared and } \sigma 2\text{ squared} \]
The null
Why: An equality.
The claim
Why: First is smaller.
\[ \sigma 1\text{ squared below } \sigma 2\text{ squared} \]
The tail
Why: The alternative points down.
Figure (svg): The solution to Worked example Example 13.5's hypotheses shown as a ladder of expressions, one row per legal move
\[ H_0: \sigma_1^2 = \sigma_2^2; \quad H_a: \sigma_1^2 < \sigma_2^2 \]
Verify: confirm the tail matches which variance was placed on top
Why: Because the first instructor's variance sits in the numerator, a smaller first variance produces a smaller ratio — so the alternative's downward direction corresponds to the lower tail of the F curve. Had the second variance been placed on top, the same claim would have produced a right-tailed test with a statistic of 1.719. Both are correct; mixing them is not.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 691
Sorting
Each is a claim, with the first sample's variance in the numerator.
Sort into buckets
Sort by the tail the test uses.
Item (a) is Example 13.5 and item (c) is the supermarket motivation the book gives in its opening paragraph. The two-tailed cases are the ones where either instructor being more variable would matter.
Worked example
Example 13.5's variances, with the ratio inverted.
\[ \frac{89.9}{52.3} \]
The inverted ratio
Why: Second over first.
\[ 1.719 \]
The hypotheses
Why: Now about the second.
The tail
Why: The alternative points up.
The p-value
Why: The upper tail beyond 1.719.
\[ 0.0753 \]
Figure (svg): The solution to Worked example the same data, both ways up shown as a ladder of expressions, one row per legal move
\[ P(F_{29,29} > 1.719) = P(F_{29,29} < 0.5818) = 0.0753 \]
Verify: confirm the two arrangements must agree
Why: Inverting an F statistic and reversing the tail gives the same probability, since one over an F on df1 and df2 has an F distribution on df2 and df1. So the choice of which variance goes on top cannot change the conclusion, provided the tail is reversed with it — and a p-value that differs between the two arrangements means the tail was not reversed.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692
Error analysis
With the first variance in the numerator. Which are correct?
Annotate
On: \( \begin{aligned} &(1)\; H_a: \sigma_1^2 < \sigma_2^2, \text{ left-tailed} \\ &(2)\; H_a: \sigma_1^2 < \sigma_2^2, \text{ right-tailed} \\ &(3)\; H_a: s_1^2 < s_2^2, \text{ left-tailed} \\ &(4)\; H_0: \sigma_1^2 < \sigma_2^2, \text{ left-tailed} \end{aligned} \)
Error (2) is the one to guard against, because the arithmetic is identical and only the tail differs. Taking the right tail here would give a p-value of 0.9247 and reverse the conclusion entirely.
Faded example
The statistic is 0.5818 and the test is left-tailed.
Fill in the blanks
p\text< = P(F 0.0753 0.5818) = ___
Why: Taking the right tail instead would give 0.9247 and reverse the conclusion — the same error section 11.6 warned about for a single variance.
Two truths and a lie
All three concern the tail.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and is the section's main structural difference from the rest of the chapter. A variance can be smaller than another as well as larger, so both directions carry meaning — exactly as in section 11.6.
Prediction
Commit before reasoning.
Predict first
The one-way ANOVA is always right-tailed and this test is not. Why?
Correct: Either variance can be the larger.
Why: An ANOVA's numerator can only be inflated by differing means, so its alternative pushes in one direction. Here the alternative may claim the first variance is larger or smaller, and both are meaningful — the same reason section 11.6's single-variance test could be left-tailed while sections 11.2 to 11.4 could not.
Section
Section 3
Concept
Two college instructors each grade the same 30 exams. The first's grades have a variance of 52.3 and the second's 89.9. Testing the claim that the first instructor's variance is smaller, at a 10 percent level, gives F = 0.5818 on 29 and 29 degrees of freedom with a left-tail p-value of 0.0753.
the conclusion — With a 10 percent level of significance, from the data, there is sufficient evidence to conclude that the variance in grades for the first instructor is smaller.
\[ F_{29,29} = 0.5818, \quad p = 0.0753 < 0.10 \]
The book's motivating remark is worth keeping: in most colleges it is desirable for the variances of exam grades to be nearly the same among instructors. So a rejection here is not good news about the first instructor — it says the two are grading with different consistency, which is the thing the administration wanted to avoid.
Figure (svg): An F curve on twenty-nine and twenty-nine degrees of freedom with a dashed line at one and the area to the left of 0.5818 shaded
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692 — Example 13.5 in full
Picture it
The book's Figure 13.6.
Figure (svg): An F curve on twenty-nine and twenty-nine degrees of freedom with a dashed line at one and the area to the left of 0.5818 shaded
The shaded region is the LEFT tail, at 0.0753. That is below a 10 percent level and above a 5 percent one, so this conclusion depends on an unusually generous alpha having been fixed in advance.
Worked example
Example 13.5, in the book's own sequence.
\[ \alpha = 0.10 \]
Hypotheses
Why: First smaller.
Statistic
Why: 52.3 over 89.9.
\[ 0.5818 \]
Distribution
Why: Both samples of 30.
\[ F s u b 29, 29 \]
Probability statement
Why: The lower tail.
\[ 0.0753 \]
Compare and decide
Why: 0.0753 below 0.10.
Figure (svg): The solution to Worked example the full write-up shown as a ladder of expressions, one row per legal move
\[ p = P(F < 0.5818) = 0.0753 < 0.10 \]
Verify: confirm how sensitive this conclusion is to the level
Why: At 5 percent the same p-value of 0.0753 would not reject, so the conclusion rests entirely on the 10 percent level having been chosen in advance. Ten percent is a generous level by the standards of the rest of the course, where 5 percent was the default and 1 percent appeared several times — which makes it worth stating prominently rather than in passing.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 691-692
Matching
Match each stage of Example 13.5.
Match the pairs
Why: The same four stages as every test since chapter 9. Only the fourth carries anything unusual — the inequality points the other way from every other test in this chapter.
Worked example
The same data, at the course's more usual level.
\[ \alpha = 0.05 \]
Compare
Why: 0.0753 against 0.05.
The decision
Why: Not below alpha.
The conclusion
Why: Insufficient evidence.
The lesson
Why: Fix alpha first.
Figure (svg): The solution to Worked example what would change at 5 percent shown as a ladder of expressions, one row per legal move
\[ 0.05 < 0.0753 < 0.10 \]
Verify: confirm this is the third time the chapter has raised the point
Why: Example 13.2's p-value of 0.0248 fell between 1 and 5 percent, Example 13.3's choice of 1 percent turned out not to matter, and this one falls between 5 and 10. Chapter 9 required the level to be fixed before the data are seen, and the chapter's own examples show how often a result lands close enough to a boundary for that requirement to bite.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 692
Trap
\[ p = P(F > 0.5818) = 0.9247 \]
Use the upper tail, as the ANOVA does
Why: Every other test in the chapter is right-tailed.
\[ \text{but the alternative says SMALLER} \]
A p-value of 0.92 would mean failing to reject, and the study's actual finding would be missed.
\[ p = P(F < 0.5818) = 0.0753 \]
Read the tail from the alternative
Why: This is the chapter's only test where the choice exists.
The error is dangerous because 0.9247 is a perfectly ordinary-looking p-value leading to a perfectly ordinary do-not-reject. Nothing in the arithmetic flags it, and the only defence is noticing before computing that the statistic came out below one and the alternative points downward — which are the same observation.
Faded example
Both instructors graded 30 exams.
Fill in the blanks
F \sim F_29,\,29}
Why: Equal sample sizes give equal degrees of freedom, which makes the curve nearly centred on one — so the statistic's distance below one is easy to read.
Two truths and a lie
All three concern the test.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A smaller variance means more consistent grades, not more accurate ones — the first instructor could be consistently generous or consistently harsh. Variance describes spread, and nothing in this test concerns the average grade at all.
Prediction
Commit before reasoning.
Predict first
The book notes that colleges want instructors' grade variances to be nearly the same. What does rejecting mean for them?
Correct: The two grade with different consistency.
Why: The null of equal variances is the desirable state here, so rejecting it identifies a problem rather than a finding. That reverses the usual framing — in most of the course rejecting was the interesting outcome — and it is a useful reminder that which hypothesis is desirable depends entirely on the context.
Section
Section 4
Concept
Unlike most other tests in this book, the F test for equality of two variances is very sensitive to deviations from normality. If the two distributions are not normal, the test can give higher p-values than it should, or lower ones, in ways that are unpredictable. Many texts suggest that students not use this test at all, but in the interest of completeness we include it here.
unpredictable — The key word. A test whose error runs in a known direction can be interpreted with caution; one whose error direction depends on the unknown shape of the populations cannot.
\[ \text{non-normal} \;\Rightarrow\; p \text{ too high OR too low} \]
No other test in the course carries a caveat of this strength, and the contrast with the t procedures is the point. Chapter 8's t interval and chapter 10's two-sample tests become robust as samples grow, because the central limit theorem describes sample MEANS. Nothing plays that role for sample variances, which is the same limitation section 11.6 noted for a single variance — and here it applies to two at once.
Figure (svg): A card reproducing the book's warning about this test's sensitivity to non-normality
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690 — the paragraph of warning after the requirements
Picture it
Reproduced rather than summarised.
Figure (svg): A card reproducing the book's warning about this test's sensitivity to non-normality
The phrase many texts suggest that students not use this test at all is remarkable in a textbook that is teaching the test. It is included for completeness, and the honest way to learn it is with the caveat attached.
Worked example
The book lists two conditions before the warning.
\[ \text{the requirements} \]
First
Why: Both populations normal.
Second
Why: The two independent.
Which is fragile
Why: The first.
What secures each
Why: Design against shape.
Figure (svg): The solution to Worked example reading the two requirements shown as a ladder of expressions, one row per legal move
\[ \text{normal populations}; \quad \text{independent samples} \]
Verify: confirm why the asymmetry matters practically
Why: Independence is a property of the sampling procedure, so a well-designed study guarantees it. Normality is a property of the populations themselves, which no design can impose — it can only be examined after the fact, and with small samples not examined well. So the fragile requirement is also the one hardest to check, which is the substance of the book's recommendation.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690
Two truths and a lie
All three concern the warning.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The central limit theorem normalises sample means, not sample variances, so the distribution of this statistic keeps depending on the populations' shape however large the samples are.
Worked example
Why other tests tolerate non-normality and this one does not.
\[ \bar{x} \text{ against } s^2 \]
A sample mean
Why: A scaled sum.
\[ \text{chapter } 7\text{ applies} \]
Its behaviour for large n
Why: Approximately normal.
A sample variance
Why: Not a sum of the data.
Its behaviour for large n
Why: Depends on shape.
Figure (svg): The solution to Worked example comparing with the t procedures shown as a ladder of expressions, one row per legal move
\[ \bar{X} \approx N \text{ for large } n; \quad s^2 \text{ has no such guarantee} \]
Verify: confirm this is the same point section 11.6 made
Why: There, the test of a single variance assumed normality and the assumption did not soften with sample size. Here the same limitation applies to two sample variances simultaneously, and the departures can compound — which is presumably why the warning appears in this section and not that one, and why it is worded so much more strongly.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, p. 690
Error analysis
Which are consistent with what the book says?
Annotate
On: \( \begin{aligned} &(1)\; \text{check both populations for normality first} \\ &(2)\; \text{use a large sample, which removes the concern} \\ &(3)\; \text{treat a } p \text{-value near alpha with particular caution} \\ &(4)\; \text{ignore it, since the book teaches the test} \end{aligned} \)
Response (3) is the practical one. Example 13.5's own p-value of 0.0753 against a 10 percent level is exactly the kind of near-boundary result the warning makes hardest to trust.
Sorting
Each is a procedure from the course.
Sort into buckets
Sort by whether large samples repair non-normality.
The division is clean and follows one principle: chapter 7's theorem covers sums and means. Every procedure it does not cover carries a normality assumption that no sample size removes.
Prediction
Commit before reasoning.
Predict first
The book says the error can go either way, unpredictably. Why is that worse than a known bias?
Correct: A known direction could be allowed for.
Why: If a test were known to give p-values that are too small, a reader could demand a stricter threshold and proceed. When the direction depends on the unknown shape of the populations, no adjustment is available — which is what turns a caveat into a recommendation against use.
Explain it
A classmate says the warning cannot be serious, since the book teaches the test and works an example.
Discussion prompt
In two sentences or fewer, explain what the warning means in practice.
Hint: Ask what they would do with a p-value of 0.06 from non-normal data.
Answer:
The book includes it for completeness while noting that many texts suggest students not use it, so the right response is to check both populations for normality before trusting any p-value it produces.
A result near the decision boundary is the least trustworthy kind here, because the p-value could be too high or too low and nothing tells you which.
Section
Section 5
Concept
The fifth fact about the F distribution named its other uses, and this is one of them. A one-way ANOVA compares several means by way of variances; a test of two variances compares the variances themselves.
the sharpest contrast — One-way ANOVA ASSUMES the group variances are equal in order to conclude about means. This test makes that same equality its hypothesis.
\[ \text{ANOVA}: \; \sigma^2 \text{ assumed equal}; \qquad \text{here}: \; \sigma_1^2 = \sigma_2^2 \text{ tested} \]
The relationship between them is practical as well as conceptual. Since a one-way ANOVA requires equal variances, a test of two variances is one way of examining that requirement — though the warning in this section means it should be used cautiously for that purpose, and a look at side-by-side box plots is often more informative.
Figure (svg): A table comparing the chapter's two uses of the F distribution
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — the chapter's two uses of the F distribution
Picture it
Six rows of comparison.
Figure (svg): A table comparing the chapter's two uses of the F distribution
Every row differs, which is a fair summary of how much a single distribution can be made to do. Only the distribution itself is shared, and even the two degrees of freedom are counted differently.
Worked example
Three situations.
\[ \text{which test?} \]
Five mulches, mean yields
Why: Several means.
Two instructors, grade spread
Why: Two variances.
Four sororities, mean grades
Why: Several means.
The signal
Why: Means or spreads.
Figure (svg): The solution to Worked example choosing between them shown as a ladder of expressions, one row per legal move
\[ \text{means} \to \text{ANOVA}; \quad \text{two variances} \to \text{this test} \]
Verify: confirm the signal is the question rather than the data
Why: The same two columns of numbers could support either test — the instructors' grades have means as well as variances, and the mulch yields have variances as well as means. What decides is which quantity the question asks about, exactly as section 11.5 found for the chi-square tests. The data alone never settle it.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692
Sorting
Each describes a question.
Sort into buckets
Sort by which test applies.
Items (b) and (d) are two of the three motivations the book gives in its opening paragraph, alongside the variability of two supermarket checkers' times.
Worked example
Two groups of thirty, compared two ways.
\[ n_1 = n_2 = 30 \]
As an ANOVA
Why: k = 2, n = 60.
\[ 1\text{ and } 58 \]
As two variances
Why: Each sample separately.
\[ 29\text{ and } 29 \]
Note the difference
Why: Same data.
Why
Why: Different questions.
Figure (svg): The solution to Worked example the degrees of freedom in each shown as a ladder of expressions, one row per legal move
\[ F_{1,58} \quad\text{against}\quad F_{29,29} \]
Verify: confirm why the same data give different counts
Why: An ANOVA pools all sixty observations to estimate one within-group variance, spending two degrees of freedom on the two group means; the variance test keeps the samples separate and each variance carries its own count. The two statistics answer different questions from the same data, so there is no reason for their degrees of freedom to agree.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692
Trap
\[ \text{variance test does not reject} \;\Rightarrow\; \text{ANOVA's equal-variance assumption holds} \]
Take a non-rejection as confirmation
Why: It is the relevant hypothesis.
\[ \text{but failing to reject never establishes a null} \]
And this particular test is the one the book warns is least trustworthy.
\[ \text{examine the group spreads directly, e.g. with box plots} \]
Treat the assumption as something to inspect, not to test away
Why: A non-rejection from a fragile test is weak evidence.
Two problems compound here. Failing to reject never proves a null, as chapter 9 established; and this test's p-value may be too high or too low unpredictably when the populations are not normal — which is exactly the situation in which the equal-variance assumption is most likely to be in doubt. Looking at the data remains the better check.
Two truths and a lie
All three compare the two uses.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The ANOVA is always right-tailed, but a test of two variances may be left, right or two-tailed — and Example 13.5 is left-tailed.
Faded example
Two samples of 25 observations each.
Fill in the blanks
\text48 F_24}}; \qquad \text___ F____,24}
Why: The same fifty observations give 1 and 48 under one test and 24 and 24 under the other, because the two statistics are built from different quantities.
Prediction
Commit before reasoning.
Predict first
Section 13.3's fifth fact named other uses for the F distribution. Which are they?
Correct: Comparing two variances and two-way ANOVA.
Why: The first is this section; the second is named and set aside as beyond the scope of the chapter. Between them they explain why the distribution earns a chapter of its own rather than appearing only as an ANOVA detail.
Comparison
Fill the blanks. Both concern variances, and both can go either way.
Comparison matrix
| One variance (11.6) | Two variances (13.4) | |
|---|---|---|
| Compared against | a claimed value | another sample's variance |
| Distribution | chi-square | F |
| Degrees of freedom | n - 1 | n1 - 1 and n2 - 1 |
| Tails available | left, right or two | left, right or two |
The last row is what sets both apart from every other test in chapters 11 and 13. A statistic that measures a ratio of variances can fall meaningfully on either side of its null value, while one that measures disagreement cannot.
Pattern
Six steps, and the first is the one the book warns about most.
If the statistic comes out below one, the evidence is in the lower tail — do not take the upper tail out of habit.
OpenStax Introductory Business Statistics 2e, §12.1 Test of Two Variances §12.1 Test of Two Variances
Check
The statistic.
Check your understanding
Samples of 20 and 25 give variances 30 and 12. What is F, and on what degrees of freedom?
Answer: A
Why: Thirty over twelve is 2.5, and the degrees of freedom are one less than each sample size.
Check
The tail.
Check your understanding
With the first sample's variance on top, the claim is that the first population is LESS variable. Which tail?
Answer: A
Why: A smaller numerator variance gives a smaller ratio, so the alternative points below one and the p-value is a left-tail area.
Check
The warning.
Check your understanding
The two populations are clearly skewed. What does the book's warning imply?
Answer: A
Why: The book says the test can give higher p-values than it should, or lower ones, in ways that are unpredictable — which is why many texts suggest not using it at all.
Real world
A manufacturer compares two machines' fill volumes, taking 25 bottles from each. The sample variances are 4.1 and 1.9, giving F = 2.16 with a two-tailed p-value of 0.058, and the report concludes that the two machines are equally consistent and no maintenance is needed.
Discussion prompt
Assess the arithmetic, the conclusion and the assumptions.
Hint: Check the decision, then ask what a non-rejection establishes.
Answer:
The arithmetic is right. Twenty-five bottles from each machine give 24 and 24 degrees of freedom, and 4.1 over 1.9 is 2.16 — whose two-tailed p-value on that curve is close to the reported 0.058. At a 5 percent level this does not reject.
\[ F_{24,24} = \frac{4.1}{1.9} = 2.16, \qquad p \approx 0.058 \]
The conclusion over-claims twice. Failing to reject never establishes that the variances are equal, and this p-value is barely above 5 percent — the data are close to indicating a difference rather than supporting equality. The sample variances differ by more than a factor of two, which in a filling operation is a substantial practical gap whatever the test says.
And the assumption is the fragile one. Fill volumes from a drifting or occasionally double-filling machine are not normal, and those are exactly the failure modes that would inflate a variance. The book's warning applies with full force: the p-value could be too high or too low, unpredictably — and at 0.058, a result this close to the boundary is the least trustworthy kind.
What should happen: plot the fifty measurements by machine, as histograms and in production order, to check the shape and look for drift. Report that machine one's sample variance is more than double machine two's, that the difference is borderline at conventional levels, and that the normality the test requires is doubtful. The honest recommendation is to collect more data, not to close the question — which is close to what the book means by suggesting many texts avoid this test entirely.
Commit first
Answer, then rate your confidence honestly.
Predict first
Why does the book warn so strongly about this test?
Correct: It is very sensitive to non-normality, unpredictably.
\[ \text{non-normal} \;\Rightarrow\; p \text{ biased in an unknown direction} \]
Why: Unlike most other tests in the book, this one's p-value can be too high or too low when the populations are not normal, and nothing indicates which. Sample size does not help, because the central limit theorem normalises sample means and says nothing about sample variances. The book includes the test in the interest of completeness while noting that many texts suggest students not use it at all.
Explain it
They computed F = 0.58 for a claim that the first variance is smaller, then took the right tail and got 0.92.
Discussion prompt
In two sentences or fewer, correct them.
Hint: Ask which direction the alternative points.
Answer:
The alternative says the first variance is SMALLER, and since it sits in the numerator that makes the ratio small — so the evidence is in the LEFT tail, at 0.0753 rather than 0.92.
This is the chapter's only test where the tail is a choice, so it is the only place the habit of taking the upper tail goes wrong.
Exit ticket
Name the weakest spot before you close the deck — and the course.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the first, the ratio of sample variances on one less than each sample size. For the second, the direction the alternative points, given which variance is on top. For the third, the same six stages as every test since chapter 9. For the fourth, very sensitive to non-normality, with the error running either way. Do five problems of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes, and it closes the course.
Draw it
At the top, write the full F ratio with each sample variance over its own population variance, then show the population variances cancelling under the null to leave the ratio of sample variances alone. Beside it write the degrees of freedom as n1 minus one and n2 minus one, and note that they are not added. Under that, draw three boxes for the three alternatives — first variance greater, less, and not equal — with the tail each selects, and circle the middle one as Example 13.5's. In the middle, work Example 13.5 completely: variances 52.3 and 89.9 from samples of 30, the ratio 0.5818, the distribution F sub 29 comma 29, and the LEFT-tail p-value of 0.0753 against a 10 percent level, giving a rejection. Draw the curve beside it with a dashed line at one and the area to the LEFT of 0.5818 shaded, and write that taking the right tail would give 0.9247 and the opposite conclusion. At the bottom left, box the book's warning in full: very sensitive to deviations from normality, p-values too high or too low unpredictably, and many texts suggest not using it at all. At the bottom right, make a six-row table comparing this test with one-way ANOVA, and underline the last row — ANOVA assumes equal variances, this test tests them.
Check your Example 13.5 work by confirming the inverted ratio 89.9 over 52.3 gives 1.719 and its RIGHT tail gives the same 0.0753 — the two arrangements must agree. Check your warning box by asking whether you could say, in one sentence, what you would do differently on encountering non-normal data.
Recap
Six things, and the last one is what the book most wants carried away.
| If you see | Then |
|---|---|
| Two sample variances to compare | F is their ratio, on n1-1 and n2-1 |
| A claim that one variance is larger | Right-tailed, with that variance on top |
| A claim that one is smaller | Left-tailed, with that variance on top |
| A claim they merely differ | Two-tailed |
| F near one | The evidence favours equal variances |
| F below one with a 'smaller' alternative | The lower tail, not the upper |
| Non-normal populations | The p-value is unreliable in an unknown direction |
| A near-boundary p-value from this test | The least trustworthy kind of result |
That completes chapter 13 and the course. Thirteen chapters have moved from describing a single sample, through the distributions that describe chance, to inference about one parameter, two, and finally several at once — and the last two chapters showed the same F and chi-square distributions doing several different jobs. Section 13.5 is the book's own lab, which puts one-way ANOVA to work on data you collect yourself.
OpenStax Introductory Statistics 2e, §13.4 Test of Two Variances §13.4, pp. 690-692 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.