13.2 The F Distribution and the F-Ratio

Section 13.1's box-plot picture turned into arithmetic. Two estimates of the population variance are computed: the variance between samples, which is the variance of the sample means multiplied by n and is called the explained variation, and the variance within samples, which is the average of the sample variances and is called the unexplained variation. The F statistic is their ratio. Its logic is stated plainly by the book: MSwithin is an estimate of the population variance, while MSbetween consists of the population variance plus a variance produced from the differences between the samples — so under the null both estimate the same value and F should be approximately one, while differing means add a term to the numerator and push the ratio above it. The distribution is named after Sir Ronald Fisher and has two sets of degrees of freedom, k minus one for the numerator and n minus k for the denominator.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 13.2 The F Distribution and the F-Ratio

Title

Statistics · Chapter 13 — F Distribution and One-Way ANOVA

The F Distribution and the F-Ratio

2. By the end of this lesson you can

Objectives

Six outcomes, and the second is the one that makes the ratio interpretable.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-684 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 13.1 established that differing means inflate the combined spread while leaving within-group spread unchanged.

Discussion prompt

How would you turn that observation into a single number that measures whether the means differ?

Hint: You have two quantities that behave differently. What do you do with two quantities you want to compare?

Answer:

Compute both and take their ratio. One estimate of the population variance is built from how far the GROUP MEANS sit from each other; the other is built from how far individual values sit from their own group's mean.

The second is blind to differences among the means, since shifting a whole group up or down leaves its internal spread untouched. The first is not.

\[ F = \frac{\text{MS}_{\text{between}}}{\text{MS}_{\text{within}}} \]

If the means are all equal, both estimate the same population variance and the ratio should be near one. If the means differ, the numerator picks up an extra term and the ratio grows. That single number is the F statistic, and the rest of this section is how to compute it.

4. A ratio of two variance estimates

Concept

To calculate the F ratio, two estimates of the variance are made. The variance between samples is the variance of the sample means multiplied by n and is also called the variation due to treatment or explained variation; the variance within samples is the average of the sample variances and is also called the variation due to error or unexplained variation.

the F distribution — Named after Sir Ronald Fisher, an English statistician. The F statistic is a ratio, and there are two sets of degrees of freedom — one for the numerator and one for the denominator.

\[ F = \frac{\text{MS}_{\text{between}}}{\text{MS}_{\text{within}}}, \qquad F \sim F_{k-1,\,n-k} \]

The explained-and-unexplained language is the same distinction chapter 12 made with r-squared. There, the regression line explained part of y's variation and left the rest as scatter; here the grouping explains part of the response's variation and leaves the rest as within-group noise. Both split a total into a part a model accounts for and a part it does not.

Figure (svg): A card contrasting the between-samples and within-samples variance estimates

The book's own framing: an extra term in the numerator is what pushes the ratio above one.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-681

5. The two variance estimates

Section

Section 1

6. One responds to the means, one does not

Concept

The one-way ANOVA test depends on the fact that MSbetween can be influenced by population differences among means of the several groups. Since MSwithin compares values of each group to its own group mean, the fact that group means might be different does not affect MSwithin.

explained and unexplained — Between-samples variation is due to treatment and is explained; within-samples variation is due to error and is unexplained. The names describe what each part of the total variation is attributed to.

\[ \text{MS}_{\text{between}} \text{ sees the means}; \quad \text{MS}_{\text{within}} \text{ does not} \]

That asymmetry is the whole design. A statistic built only from within-group spread could never detect differing means, and one built only from the means could not tell a real difference from ordinary sampling noise. Taking the ratio measures the first against the second, which is exactly the comparison section 13.1's box plots invited.

Figure (svg): A card contrasting the between-samples and within-samples variance estimates

The book's own framing: an extra term in the numerator is what pushes the ratio above one.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-681 — the two estimates, and why MSwithin is unaffected

7. The two estimates

Picture it

What each is, and what each is called.

Figure (svg): A card contrasting the between-samples and within-samples variance estimates

The book's own framing: an extra term in the numerator is what pushes the ratio above one.

The lower panels state the logic in the book's own terms. Because the numerator carries an extra term whenever the means differ, and that term is never negative, the ratio can only be pushed upward — which is why the test is right-tailed.

8. Worked example: why F should be near one under the null

Worked example

The book's argument, step by step.

\[ H_0 \text{ true} \]

MSwithin

Why: Average of group variances.

MSbetween under H0

Why: Means differ only by chance.

Their ratio

Why: Two estimates of one thing.

What moves it

Why: Sampling error only.

Figure (svg): The solution to Worked example why F should be near one under the null shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F \approx 1 \text{ when } H_0 \text{ holds} \]

Verify: confirm what happens when the null is false

Why: The book says MSbetween consists of the population variance PLUS a variance produced from the differences between the samples. Since variances are always positive, adding that term makes the numerator generally larger than the denominator and the ratio larger than one. It adds one honest caveat: if the population effect is small, it is not unlikely that MSwithin will be larger in a given sample — so an F below one does happen, as Example 13.1's 0.3769 shows.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, p. 681

9. Which estimate does it describe?

Sorting

Each phrase comes from the book.

Sort into buckets

Sort by which variance estimate it names.

Between samples
the variance of the sample means multiplied by n; variation due to treatment; explained variation
Within samples
the average of the sample variances; variation due to error
bt
It measures how far the group means sit from each other.
wt
It measures how far values sit from their own group's mean.

The unexplained half is also called the pooled variance, which connects it to section 10.4's pooled proportion — both average information across groups under an assumption that the groups share a parameter.

10. Worked example: Example 13.4's two estimates

Worked example

Three children's bean plants, five each. Sample means 24.2, 25.4 and 24.4; sample variances 11.7, 18.3 and 16.3.

\[ n = 5, \; k = 3 \]

Variance of the three means

Why: Of 24.2, 25.4, 24.4.

\[ 0.413 \]

MSbetween

Why: n times that.

\[ 5(0.413) = 2.067 \]

Mean of the three variances

Why: Of 11.7, 18.3, 16.3.

\[ 15.433 \]

MSwithin

Why: That mean.

\[ 15.433 \]

Figure (svg): The solution to Worked example Example 13.4's two estimates shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F = \frac{2.067}{15.433} = 0.134 \]

Verify: confirm the reading of a ratio well below one

Why: The group means differ by barely a single inch while individual plants within a group vary by several, so the between-group signal is far smaller than the within-group noise. An F of 0.134 says exactly that, and the right-tail p-value is 0.876 — the data are entirely consistent with all three media producing the same mean height.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 682-683

11. Trap: reading a small F as evidence of a difference

Trap

The trap

\[ F = 0.134 \text{ is far from } 1 \;\Rightarrow\; \text{the means differ} \]

Treat any departure from one as evidence

Why: The null predicted a ratio near one.

\[ \text{but only LARGE values indicate differing means} \]

A small F means the between-group variation is even smaller than the within-group variation — the opposite of what differing means produce.

The fix

\[ F = 0.134 \;\Rightarrow\; p = 0.876, \text{ no evidence at all} \]

Read only the upper tail

Why: The extra term in the numerator can only push F up.

This is the same asymmetry chapter 11's goodness-of-fit statistic had. There, disagreement could only make the statistic larger; here, differing means can only make the numerator larger. In both cases the lower tail contains no alternative to test, which is why section 13.3 says the ANOVA test is always right-tailed.

12. One of these is false

Two truths and a lie

All three concern the two estimates.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. MSwithin is unaffected by differences among the means
  • C. Under the null both estimate the same population variance
  • B. MSbetween is unaffected by differences among the means

Survives elimination: B

Why: The survivor is false and reverses the design. MSbetween is built from the variance of the sample means, so it responds directly to their differences — that responsiveness is the entire mechanism of the test.

13. What is in the numerator?

Prediction

Commit before reasoning.

Predict first

According to the book, what does MSbetween consist of?

  • The population variance plus a variance produced by the differences between samples
  • Only the differences between samples
  • The population variance alone
  • The sum of the group means

Correct: The population variance plus an extra term.

Why: That decomposition is what makes the ratio interpretable. The denominator estimates the population variance alone, so the ratio is one plus something non-negative — near one when the means agree, larger when they do not.

14. The balanced shortcut

Faded example

Three groups of five, with group means varying by 0.413 and group variances averaging 15.433.

Fill in the blanks

F = \frac0.413}}0.134 = ___

Why: That is Example 13.4 exactly. The shortcut needs equal group sizes; with unequal sizes the sums of squares must be weighted, as in Example 13.1.

15. The sums of squares

Section

Section 2

16. One correction term, used twice

Concept

The total sum of squares is the sum of squares of all values from every group combined, less the grand total squared divided by n. SS between adds each group sum squared over its group size and subtracts the same correction; SS within is the difference between the two.

s-sub-j — The book's symbol for the SUM of the values in the jth group — not a standard deviation, despite the letter. Example 13.1's s1 = 16.5 is a total of four weight losses.

\[ \text{SS}_{\text{total}} = \sum x^2 - \frac{(\sum x)^2}{n} \]

The notation clash deserves the warning. Every earlier chapter used s for a sample standard deviation, and this section uses s-sub-j for a group total. Reading Example 13.1's line s1 = 16.5, s2 = 15, s3 = 15.5 as three standard deviations would make nonsense of everything that follows, since those are sums of three or four small weight losses.

Figure (svg): A table of the symbols used in the sum-of-squares calculations

Two of these six quantities — the group sums and the sum of squares — are all the sums of squares need.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 681-682 — the notation list and the sum-of-squares formulas

17. The notation

Picture it

Six symbols, with their values in Example 13.1.

Figure (svg): A table of the symbols used in the sum-of-squares calculations

Two of these six quantities — the group sums and the sum of squares — are all the sums of squares need.

Only two quantities are actually needed: the group sums and the sum of every squared value. Everything else in the table is derived from those two plus the group sizes.

18. Worked example: Example 13.1's sums of squares

Worked example

Three diet plans with weight losses summing to 16.5, 15 and 15.5 across groups of 4, 3 and 3. Every value squared and added gives 244.

\[ \sum x = 47, \; \sum x^2 = 244, \; n = 10 \]

The correction term

Why: 47 squared over 10.

\[ 220.9 \]

SS total

Why: 244 minus 220.9.

\[ 23.1 \]

The group terms

Why: 272.25/4 + 225/3 + 240.25/3.

\[ 223.1458 \]

SS between

Why: Less the correction.

\[ 2.2458 \]

SS within

Why: 23.1 minus 2.2458.

\[ 20.8542 \]

Figure (svg): The solution to Worked example Example 13.1's sums of squares shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 23.1 = 2.2458 + 20.8542 \]

Verify: confirm the decomposition adds up

Why: The two parts must sum to the total exactly, since SS within was defined as the difference. That makes the addition a check on the two computed pieces rather than a third calculation, and it is the same explained-plus-unexplained split chapter 12 made with r-squared — there the shares summed to one, here the sums of squares sum to the total.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 682-683

19. The correction term

Faded example

Ten values totalling 47.

Fill in the blanks

\frac2209220.9 = \frac___}___ = ___

Why: The same 220.9 is subtracted in both the total and the between-group sums of squares, which is why SS within can be obtained as a difference.

20. Worked example: the degrees of freedom

Worked example

For the same three diet plans, with ten observations in total.

\[ k = 3, \; n = 10 \]

Numerator

Why: k minus one.

\[ 2 \]

Denominator

Why: n minus k.

\[ 7 \]

Total

Why: n minus one.

\[ 9 \]

Check

Why: Two plus seven.

\[ 9 \]

Figure (svg): The solution to Worked example the degrees of freedom shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ (k-1) + (n-k) = n - 1 \]

Verify: confirm the additivity is exact rather than coincidental

Why: The identity holds for any k and n, since the k cancels: k minus one plus n minus k is n minus one. So the degrees of freedom decompose exactly as the sums of squares do, which is what makes the ANOVA table's rows consistent. Only the mean squares fail to add, since each is a different quotient.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 681-683

21. Error analysis: four readings of the notation

Error analysis

In Example 13.1 the book writes s1 = 16.5, s2 = 15, s3 = 15.5. Which reading is right?

Annotate

On: \( \begin{aligned} &(1)\; \text{the three group standard deviations} \\ &(2)\; \text{the three group sums} \\ &(3)\; \text{the three group means} \\ &(4)\; \text{the three group variances} \end{aligned} \)

  • (1) is the natural reading and is wrong. Everywhere else s is a standard deviation; here the book defines s-sub-j as a group sum.
  • (2) is correct: 16.5 is the total of plan 1's four weight losses.
  • (3) would give means near 4, not 16.5 — dividing 16.5 by 4 gives 4.125.
  • (4) would be even further off, since these are small numbers whose squared deviations are fractions.

The check that settles it is plausibility: three or four weight losses of about four pounds each cannot have a standard deviation of 16.5, and cannot have a mean of 16.5 either. Only a sum makes arithmetic sense.

22. One of these is false

Two truths and a lie

All three concern the sums of squares.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. SS between plus SS within equals SS total
  • C. The two degrees of freedom add to n minus one
  • B. MS between plus MS within equals MS total

Survives elimination: B

Why: The survivor is false, and the ANOVA table shows it by leaving the Total row's MS cell empty. Each mean square is a different sum divided by a different divisor, so they have no reason to add — and the table does not ask them to.

23. Symbol to meaning

Matching

Match each of the book's symbols.

Match the pairs

  • l1. k
  • l2. n-sub-j
  • l3. s-sub-j
  • l4. n
  • r1. the number of groups
  • r2. the size of the jth group
  • r3. the sum of the values in the jth group
  • r4. the total number of values combined

Why: The third is the one to fix in memory, since the letter s means something else in every other chapter of the book.

24. Why is SS within computed as a difference?

Prediction

Commit before reasoning.

Predict first

The book gets SS within by subtracting SS between from SS total. Why not compute it directly?

  • Both other sums use the same correction term, so subtracting is less work and less error-prone
  • It cannot be computed directly
  • Because the groups differ in size
  • To make the table add up

Correct: Subtracting is less work.

Why: Computing it directly would mean finding each group's mean and squaring every deviation from it, group by group. Since the total and between-group sums share the same correction term and use quantities already in hand, the difference gives the same number with far less arithmetic.

25. The ANOVA table

Section

Section 3

26. Three rows, and the F in the corner

Concept

Data are typically put into a table for easy viewing, and one-way ANOVA results are often displayed this way by computer software. The rows are Factor between, Error within and Total; the columns are the sum of squares, the degrees of freedom, the mean square and F.

the ANOVA table — A standard layout. The Factor row's MS divided by the Error row's MS gives F, and the Total row carries only a sum of squares and its degrees of freedom.

\[ F = \frac{\text{MS(Factor)}}{\text{MS(Error)}} \]

Reading the table is a skill in itself, because every statistical package produces one and they all share this shape. The two things to locate are the F in the top right and the degrees of freedom in the middle column, since those three numbers together determine the p-value — everything else in the table is intermediate working.

Figure (svg): The one-way ANOVA table for Example 13.1's diet plans

An F of 0.3769 is below one, so the between-group variation is smaller than the within-group variation.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 682-683 — the table layout and Example 13.1's completed version

27. Example 13.1's table

Picture it

Filled in completely.

Figure (svg): The one-way ANOVA table for Example 13.1's diet plans

An F of 0.3769 is below one, so the between-group variation is smaller than the within-group variation.

The Total row has no mean square and no F, which is not an omission. There is nothing for a total mean square to be compared against, so the cell is left blank in every ANOVA table.

28. Worked example: completing the table

Worked example

With SS between of 2.2458, SS within of 20.8542, k = 3 and n = 10.

\[ \text{fill the table} \]

MS Factor

Why: 2.2458 over 2.

\[ 1.1229 \]

MS Error

Why: 20.8542 over 7.

\[ 2.9792 \]

F

Why: 1.1229 over 2.9792.

\[ 0.3769 \]

Read it

Why: Below one.

Figure (svg): The solution to Worked example completing the table shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F = \frac{1.1229}{2.9792} = 0.3769 \]

Verify: confirm the F against the two mean squares directly

Why: Within-group variation per degree of freedom is 2.9792 while between-group variation per degree of freedom is only 1.1229 — so the three diet plans' means sit closer together than random variation within a plan would typically produce. The right-tail p-value is 0.699, which is a long way from any level of significance.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, p. 683

29. Complete a row

Faded example

A Factor row has SS = 36 and df = 3.

Fill in the blanks

\text3 = \frac12___} = ___

Why: Each row's mean square uses that row's own degrees of freedom, which is why the mean squares do not add even though the sums of squares do.

30. Worked example: reading a calculator's output

Worked example

Example 13.2's calculator display for the tomato data.

\[ F = 4.4810, \; p = 0.0248 \]

Factor row

Why: df 4, SS 36,648,561.

\[ MS 9, 162, 140 \]

Error row

Why: df 10, SS 20,446,726.

\[ MS 2, 044, 673 \]

Check the F

Why: Divide the mean squares.

\[ 4.481 \]

Check the total

Why: Add the sums of squares.

\[ 57, 095, 287 \]

Figure (svg): The solution to Worked example reading a calculator's output shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \frac{9{,}162{,}140}{2{,}044{,}673} = 4.481 \]

Verify: confirm both internal checks the output permits

Why: The two mean squares divide to give the reported F, and the two sums of squares add to the reported total — two independent confirmations that the output is internally consistent. Doing them takes seconds and catches a mistyped data value, which would otherwise be invisible in a display of five numbers.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 682-683

31. Trap: dividing the sums of squares instead of the mean squares

Trap

The trap

\[ F = \frac{2.2458}{20.8542} = 0.108 \]

Take the ratio of the two sums of squares

Why: They are the two quantities in the SS column.

\[ \text{but they rest on different degrees of freedom} \]

Two degrees of freedom went into the first and seven into the second, so the raw sums are not comparable.

The fix

\[ F = \frac{2.2458/2}{20.8542/7} = \frac{1.1229}{2.9792} = 0.3769 \]

Divide each sum by its own degrees of freedom first

Why: That is what a mean square is.

The error runs in a predictable direction here, since the denominator's degrees of freedom usually exceed the numerator's — so dividing the raw sums understates F, often by a factor of several. In Example 13.1 the wrong route gives 0.108 against the correct 0.3769.

32. One of these is false

Two truths and a lie

All three concern the table.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. F is MS Factor divided by MS Error
  • C. The Total row has no mean square
  • B. F is SS Factor divided by SS Error

Survives elimination: B

Why: The survivor is false. The sums of squares rest on different degrees of freedom, so each must be divided by its own before they can be compared. Using the raw sums understates F substantially.

33. Which column?

Sorting

Each is a quantity in an ANOVA table.

Sort into buckets

Sort by which column it appears in.

Sum of squares
2.2458; 20.8542
Degrees of freedom
k minus 1; n minus k
Mean square
1.1229
ss
A sum of squared quantities, before any division.
df
A count derived from k and n.
ms
A sum of squares divided by its own degrees of freedom.

Every mean square is a sum of squares from the same row divided by the degrees of freedom from the same row, which is the table's organising principle.

34. Predict the p-value

Estimation

An F of 0.3769 on 2 and 7 degrees of freedom.

Predict first

Roughly what right-tail p-value should be expected?

  • Large, around 0.7
  • Small, below 0.05
  • About 0.5
  • About 0.05

Correct: Large, around 0.7.

Why: The statistic is well below one, so the between-group variation is smaller than the within-group variation and most of the distribution lies above it. The exact value is 0.699, and no level of significance would reject.

35. The distribution and its notation

Section

Section 4

36. Two degrees of freedom, in order

Concept

There are two sets of degrees of freedom, one for the numerator and one for the denominator. If F follows an F distribution with four degrees of freedom for the numerator and ten for the denominator, then F is distributed as F-sub-four-comma-ten.

the notation — F-sub-df-numerator-comma-df-denominator, where the numerator's is k minus one and the denominator's is n minus k. The order matters: F-sub-4-comma-10 is a different distribution from F-sub-10-comma-4.

\[ F \sim F_{\text{df(num)},\,\text{df(denom)}} = F_{k-1,\,n-k} \]

The book adds a derivation note worth knowing even though its details are beyond the course: the F distribution is derived from Student's t, and the values of F are squares of the corresponding values of the t distribution. That is why one-way ANOVA on exactly two groups gives the same answer as a two-sample t test — the F statistic is the t statistic squared.

Figure (svg): An F curve on four and ten degrees of freedom, skewed right, with a dashed line at one and the area beyond 4.481 shaded

Example 13.2's picture: the null predicts a ratio near one, and 4.481 sits far into the right tail.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-684 — the notation, the two degrees of freedom, and the derivation note

37. An F curve

Picture it

Four and ten degrees of freedom, with Example 13.2's statistic marked.

Figure (svg): An F curve on four and ten degrees of freedom, skewed right, with a dashed line at one and the area beyond 4.481 shaded

Example 13.2's picture: the null predicts a ratio near one, and 4.481 sits far into the right tail.

The dashed line at one is what the null predicts. Example 13.2's statistic of 4.481 sits far to its right, and the shaded tail beyond it is the p-value of 0.0248.

38. Worked example: naming the distribution

Worked example

Example 13.2's tomato study has five conditions and fifteen plants.

\[ k = 5, \; n = 15 \]

Numerator df

Why: k minus one.

\[ 4 \]

Denominator df

Why: n minus k.

\[ 10 \]

Write the notation

Why: Numerator first.

\[ F s u b 4, 10 \]

Note the order

Why: Not the reverse.

\[ F\text{ sub } 10, 4\text{ differs} \]

Figure (svg): The solution to Worked example naming the distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F \sim F_{4,10} \]

Verify: confirm the order matters by comparing the two

Why: The 5 percent critical value of F-sub-4-comma-10 is 3.48, while for F-sub-10-comma-4 it is 5.96 — so swapping the two would raise the bar substantially and could reverse a decision. The numerator's degrees of freedom always comes first, and it is always k minus one for a one-way ANOVA.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-684

39. Name the distribution

Faded example

Four sororities with five sisters each.

Fill in the blanks

F \sim F_3,\,16}

Why: That is Example 13.3's F-sub-3-comma-16, whose statistic is 2.2303 and p-value 0.1241.

40. Worked example: the connection to the t distribution

Worked example

Comparing exactly two groups by both methods.

\[ k = 2 \]

Numerator df

Why: Two minus one.

\[ 1 \]

Denominator df

Why: n minus two.

The book's note

Why: F values are squares of t.

\[ F = t\text{ squared} \]

So the tests agree

Why: Same p-value.

Figure (svg): The solution to Worked example the connection to the t distribution shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F_{1,\,n-2} = (t_{n-2})^2 \]

Verify: confirm this explains why ANOVA is a generalisation rather than a rival

Why: The book says one-way ANOVA expands the t-test for comparing more than two groups, and the squaring relationship is what that expansion means concretely. Squaring also explains why the F test has no direction: a t of 2 and a t of minus 2 both square to 4, so the sign is destroyed — the same loss chapter 11's chi-square statistics suffered for the same reason.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, p. 680

41. Error analysis: four statements about the distribution

Error analysis

Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; \text{the numerator df is } k - 1 \\ &(2)\; \text{the denominator df is } n - 1 \\ &(3)\; F_{4,10} \text{ and } F_{10,4} \text{ are the same} \\ &(4)\; F \text{ values are squares of } t \text{ values} \end{aligned} \)

  • (1) is correct: one degree of freedom is spent on the overall mean across k groups.
  • (2) is wrong. The denominator df is n minus k; n minus one is the TOTAL degrees of freedom.
  • (3) is wrong. The two differ, and their critical values differ substantially — 3.48 against 5.96 at 5 percent.
  • (4) is correct, and it is the book's own derivation note.

Error (2) is the one to watch, because n minus one does appear in the ANOVA table — as the Total row's degrees of freedom. Taking it as the denominator's would use the wrong distribution entirely.

42. One of these is false

Two truths and a lie

All three concern the notation.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The numerator degrees of freedom is written first
  • C. F is derived from the t distribution
  • B. Swapping the two degrees of freedom gives the same distribution

Survives elimination: B

Why: The survivor is false. F-sub-4-comma-10 and F-sub-10-comma-4 are genuinely different, with 5 percent critical values of 3.48 and 5.96 — so the order carries real information and reversing it can change a decision.

43. What does the squaring relationship explain?

Prediction

Commit before reasoning.

Predict first

F values are squares of the corresponding t values. What follows for a two-group comparison?

  • ANOVA and the two-sample t test give the same p-value, and ANOVA loses the direction
  • ANOVA is more powerful
  • The t test is invalid
  • The two give different answers

Correct: Same p-value, and ANOVA loses the direction.

Why: Squaring makes a t of 2 and a t of minus 2 both give F equal to 4, so the two-tailed t p-value and the right-tailed F p-value coincide. What is lost is the sign, which is why an ANOVA never says which group's mean is larger.

44. Where should F land under the null?

Estimation

The mean of an F distribution is the denominator df over that df minus two.

Predict first

For F-sub-4-comma-10, roughly what is the mean?

  • 1.25
  • 1.00
  • 4.00
  • 10.0

Correct: 1.25.

Why: Ten over eight is 1.25 — near one, as the logic predicts, but slightly above it because the distribution is skewed right. The mean approaches one as the denominator degrees of freedom grow, which is the book's formula for the F mean at work.

45. The balanced-design shortcut

Section

Section 5

46. When every group is the same size

Concept

The foregoing calculations were done with groups of different sizes. If the groups are the same size, the calculations simplify somewhat: the F ratio becomes n times the variance of the sample means, divided by the pooled variance.

the shortcut — MSbetween is the sample size times the variance of the k group means; MSwithin is the mean of the k sample variances. Both require equal group sizes.

\[ F = \frac{n \cdot s^2_{\bar{x}}}{s^2_{\text{pooled}}} \]

The shortcut makes the mechanism visible in a way the sums of squares do not. The numerator is literally the spread of the group means, scaled by how many observations each rests on; the denominator is literally the typical spread within a group. Section 13.1's box-plot argument is written directly into the formula.

Figure (svg): A card giving the balanced-design shortcut for the F ratio

The same F, computed from group means and group variances directly rather than through sums of squares.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 681-682 — the F-ratio formula when the groups are the same size

47. The shortcut

Picture it

With Example 13.4's numbers.

Figure (svg): A card giving the balanced-design shortcut for the F ratio

The same F, computed from group means and group variances directly rather than through sums of squares.

The degrees of freedom are unchanged — still k minus one and n minus k — so only the route to the F statistic is shorter, not the test itself.

48. Worked example: Example 13.4 in full

Worked example

Three children's bean plants, five each. Does it appear the three media produce the same mean height? Test at 3 percent.

\[ \alpha = 0.03 \]

Group means

Why: Computed first.

\[ 24.2, 25.4, 24.4 \]

Variance of those three

Why: The numerator's core.

\[ 0.413 \]

MSbetween

Why: Times n = 5.

\[ 2.067 \]

Group variances, averaged

Why: 11.7, 18.3, 16.3.

\[ 15.433 \]

F, and its right tail on 2 and 12

Why: 2.067 over 15.433.

\[ 0.134, p = 0.876 \]

Figure (svg): The solution to Worked example Example 13.4 in full shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ F = \frac{2.067}{15.433} = 0.134 \]

Verify: confirm the shortcut agrees with the sums-of-squares route

Why: Computing the full table for these fifteen plants gives SS between of 4.1333 on 2 degrees of freedom, so MS between is 2.0667 — the shortcut's answer exactly. SS within is 185.2 on 12 degrees of freedom, giving 15.4333. The two routes must agree whenever the groups are equal in size, and checking one against the other is a useful confirmation on a first attempt.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 682-683

49. Apply the shortcut

Faded example

Four groups of six, with group means varying by 2.5 and group variances averaging 10.

Fill in the blanks

F = \frac6 \times 2.5}1.5 = ___

Why: An F of 1.5 on 3 and 20 degrees of freedom gives a right-tail p-value of about 0.246 — above any usual level, so this would not reject.

50. Worked example: why the shortcut needs equal sizes

Worked example

Trying it on Example 13.1's unequal groups of 4, 3 and 3.

\[ n_j = 4, 3, 3 \]

Which n to use

Why: The groups differ.

The book's remedy

Why: Weight the variance.

So use the sums of squares

Why: The general route.

\[ \text{as Example } 13.1\text{ does} \]

The result

Why: Either way.

\[ F = 0.3769 \]

Figure (svg): The solution to Worked example why the shortcut needs equal sizes shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{unequal } n_j \;\Rightarrow\; \text{use SS} \]

Verify: confirm the book states the weighting explicitly

Why: It says that if the samples are different sizes, the variance between samples is weighted to account for the different sample sizes, and that when the sample sizes are different the variance within samples is weighted too. The sums-of-squares formulas do that weighting automatically — each group sum is divided by its own group size — which is why they work in both cases while the shortcut works in only one.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-683

51. Trap: using the shortcut on unequal groups

Trap

The trap

\[ F = \frac{\bar{n} \cdot s^2_{\bar{x}}}{s^2_{\text{pooled}}} \text{ with } \bar{n} = 3.33 \]

Substitute an average group size

Why: It seems a reasonable compromise.

\[ \text{but the weighting is not an average} \]

Each group's contribution must be weighted by its own size, which an average n cannot reproduce.

The fix

\[ \text{use the sum-of-squares formulas, which weight automatically} \]

Reserve the shortcut for balanced designs

Why: The book introduces it that way explicitly.

The sums of squares divide each group's squared total by that group's own size, so a larger group's mean contributes more — exactly the weighting the shortcut cannot express. The general route is barely longer once the group sums and the sum of squared values are in hand, and it always works.

52. One of these is false

Two truths and a lie

All three concern the shortcut.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. It requires every group to be the same size
  • C. It gives the same F as the sums-of-squares route
  • B. It uses different degrees of freedom

Survives elimination: B

Why: The survivor is false. The degrees of freedom are k minus one and n minus k whichever route computes the statistic — only the arithmetic differs, not the test.

53. What does multiplying by n accomplish?

Prediction

Commit before reasoning.

Predict first

Why is the variance of the sample means multiplied by n?

  • Each sample mean rests on n observations, so it varies less than an individual value does
  • To make the units match
  • To make F larger
  • Because there are n groups

Correct: Each mean rests on n observations.

Why: Chapter 7 established that a sample mean's variance is the population variance divided by n. So the variance OF the sample means estimates the population variance divided by n, and multiplying by n recovers an estimate of the population variance itself — which is what makes it comparable with MSwithin.

54. Which route applies?

Sorting

Each describes a study's group sizes.

Sort into buckets

Sort by whether the balanced shortcut may be used.

Balanced: the shortcut applies
three groups of five bean plants; four groups of five sorority members; five groups of three tomato plants
Unbalanced: use the sums of squares
groups of 4, 3 and 3 weight losses; groups of 12, 15 and 9 patients
yes
Every group has the same number of observations.
no
The group sizes differ, so there is no single n to multiply by.

Items (a), (c) and (d) are Examples 13.4, 13.3 and 13.2 — three of the chapter's four worked studies are balanced, and only Example 13.1 is not.

55. The two mean squares

Comparison

Fill the blanks. Both estimate the same thing under the null, and only one responds to differing means.

Comparison matrix

MS betweenMS within
Built fromthe variance of the group meansthe average of the group variances
Also calledexplained variationunexplained variation
Responds to differing meansyesno
Degrees of freedomk - 1n - k

The third row is the entire design. If both responded to the means, the ratio would be uninformative; if neither did, it would be blind. One doing so and the other not is what makes the comparison a test.

56. Computing an F ratio, in order

Pattern

Six steps, or three if the design is balanced.

  1. Record k, each group size, each group sum, the grand total and the sum of every squared value.
  2. Compute the correction term as the grand total squared over n.
  3. SS total is the sum of squared values less the correction; SS between adds each group sum squared over its size, less the same correction.
  4. SS within is the difference between them.
  5. Divide each sum of squares by its own degrees of freedom — k minus one and n minus k — to get the mean squares.
  6. Divide MS between by MS within to get F, and name the distribution as F with those two degrees of freedom in that order.

For equal group sizes, F is n times the variance of the group means over the mean of the group variances — a three-step route to the same number.

OpenStax Introductory Business Statistics 2e, §12.3 The F Distribution and the F-Ratio §12.3 The F Distribution and the F-Ratio

57. Check yourself 1 of 3

Check

The ratio.

Check your understanding

Under the null hypothesis, what value should the F ratio be near?

  • A. One (correct)
  • B. Zero
  • C. The number of groups
  • D. The sample size

Answer: A

Why: Both mean squares estimate the same population variance when the means are equal, so their ratio should be approximately one, with sampling error causing small departures.

Why B tempts people
An F near zero would mean the group means sit far closer together than within-group variation predicts.
Why C tempts people
The number of groups sets the numerator degrees of freedom, not the expected ratio.
Why D tempts people
The sample size affects the degrees of freedom, not the ratio's expected value.

58. Check yourself 2 of 3

Check

Degrees of freedom.

Check your understanding

Six groups with four observations each. What are the two degrees of freedom?

  • A. 5 and 18 (correct)
  • B. 6 and 24
  • C. 5 and 23
  • D. 3 and 20

Answer: A

Why: The numerator is k minus one, which is 5, and the denominator is n minus k, which is 24 minus 6, or 18.

Why B tempts people
Those are k and n themselves, before subtracting.
Why C tempts people
Twenty-three is n minus one, the TOTAL degrees of freedom rather than the denominator's.
Why D tempts people
That reverses the roles of the group count and the group size.

59. Check yourself 3 of 3

Check

The table.

Check your understanding

An ANOVA table shows SS Factor = 24 with df 3, and SS Error = 60 with df 20. What is F?

  • A. 2.67 (correct)
  • B. 0.40
  • C. 8.00
  • D. 0.375

Answer: A

Why: The mean squares are 24 over 3, which is 8, and 60 over 20, which is 3; their ratio is 2.67.

Why B tempts people
That divides the sums of squares directly, ignoring the different degrees of freedom.
Why C tempts people
That is MS Factor alone, before dividing by MS Error.
Why D tempts people
That inverts the correct ratio, putting MS Error on top.

60. Where this shows up outside the textbook

Real world

A software report gives an ANOVA table with SS Factor = 480 on 3 degrees of freedom and SS Error = 1,200 on 4 degrees of freedom, and reports F = 0.40 with a note that the groups do not differ.

Discussion prompt

Find the errors and say what the data actually show.

Hint: Check the F against the mean squares, and check the degrees of freedom against each other.

Answer:

The reported F is the ratio of the SUMS of squares, not the mean squares. Dividing 480 by 1,200 gives exactly 0.40, which is what was reported. The correct calculation divides each sum by its own degrees of freedom first.

\[ F = \frac{480/3}{1200/4} = \frac{160}{300} = 0.533 \]

The degrees of freedom are also suspicious. With 3 for the factor there are four groups, so k is 4 — and a denominator of 4 means n minus k is 4, giving n = 8, or two observations per group. Two observations per group is enough to compute the table and far too few for the normality assumption to be checkable, so the whole analysis rests on an assumption nothing can verify.

The conclusion happens to survive both errors. The correct F of 0.533 on 3 and 4 degrees of freedom gives a right-tail p-value near 0.69, so there is genuinely no evidence that the group means differ. But that is luck: had the sums of squares been reversed, the wrong route would have understated a large F and hidden a real difference.

What should be reported: F = 0.533 on 3 and 4 degrees of freedom with a p-value near 0.69, and an explicit note that with two observations per group the test has almost no power — it would fail to detect even a substantial difference. Failing to reject here says considerably less than the phrase the groups do not differ implies, which is chapter 9's asymmetry showing up one last time.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

Why should the F ratio be near one when the null hypothesis is true?

  • Because F is always near one
  • Because both mean squares then estimate the same population variance
  • Because the degrees of freedom are equal
  • Because the groups are the same size

Correct: Both estimate the same population variance.

\[ \text{MS}_{\text{between}} = \sigma^2 + \text{(term from differing means)}; \quad \text{MS}_{\text{within}} = \sigma^2 \]

Why: MSwithin is an estimate of the population variance, and MSbetween consists of that variance plus a term produced by differences between the samples. When the means are all equal that extra term vanishes and both quantities estimate one thing, so their ratio should be near one — with only sampling error causing variations away from it.

62. Explain it to someone a year behind you

Explain it

They computed F by dividing SS between by SS within.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask how many degrees of freedom went into each sum.

Answer:

The two sums of squares rest on different degrees of freedom — k minus one and n minus k — so each has to be divided by its own before they can be compared.

Those quotients are the mean squares, and F is their ratio; using the raw sums usually understates F by a factor of several.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Explaining why F should be near one under the null
  • Computing the three sums of squares from group sums
  • Filling in an ANOVA table and reading F from it
  • Using the balanced shortcut and knowing when it applies

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, both mean squares then estimate the same population variance. For the second, subtract the same correction term from both. For the third, each row's MS uses that row's own df. For the fourth, n times the variance of the means over the mean of the variances, and only for equal group sizes. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Fifteen minutes.

Draw it

At the top, draw two boxes side by side. In the left write BETWEEN samples with the variance of the sample means times n, and the labels variation due to treatment and explained variation. In the right write WITHIN samples with the average of the sample variances, and the labels variation due to error and unexplained variation. Beneath them write the sentence that makes the ratio interpretable: MSwithin estimates the population variance, MSbetween estimates it plus a term from the differing means, so F is near one under the null. In the middle, work Example 13.1 completely: the group sums 16.5, 15 and 15.5 over sizes 4, 3 and 3; the correction term 47 squared over 10, which is 220.9; SS total 244 minus 220.9, which is 23.1; SS between 223.1458 minus 220.9, which is 2.2458; SS within by subtraction, 20.8542. Then draw the ANOVA table with all three rows filled — df 2, 7 and 9; MS 1.1229 and 2.9792; F equal to 0.3769. Note beside it that the sums of squares add and the degrees of freedom add, but the mean squares do not. At the bottom left, sketch an F curve skewed right with a dashed line at one and note the notation F sub numerator-df comma denominator-df. At the bottom right, write the balanced shortcut and work Example 13.4 through it: 5 times 0.413 over 15.433 equals 0.134.

Check your Example 13.1 work by confirming 2.2458 plus 20.8542 gives 23.1 exactly. Check your table by confirming 2 plus 7 gives 9, and check that you did NOT add the two mean squares — the Total row's MS cell should be empty.

65. What you can do now

Recap

Six things, and the second is what makes the number mean something.

If you seeThen
k groups and n observationsDegrees of freedom k minus 1 and n minus k
Two sums of squaresDivide each by its own df before taking the ratio
s-sub-j in this chapterA group SUM, not a standard deviation
Equal group sizesThe shortcut applies: n times the variance of the means over the pooled variance
Unequal group sizesUse the sums of squares, which weight automatically
F near oneConsistent with all the means being equal
F well above oneThe between-group variation exceeds the within-group variation
F below oneEven less between-group variation than within; no evidence at all

Section 13.3 lists the distribution's properties and works three complete tests with it — including the tomato study whose F of 4.481 gives a p-value of 0.0248, the chapter's one rejection.

OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio §13.2, pp. 680-684 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §13.2 The F Distribution and the F-Ratio — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 680-684
  2. OpenStax Introductory Business Statistics 2e, §12.3 The F Distribution and the F-Ratio — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108