13.1 One-Way ANOVA

The setup section for the chapter, containing one genuinely surprising sentence: the purpose of a one-way ANOVA test is to determine the existence of a statistically significant difference among several group means, and the test actually uses variances to help determine whether those means are equal. The reason is visible in a picture. If every group has the same population mean, pooling the groups changes nothing and the combined data has about the variance each group has; if the means differ, pooling spreads the data out and the combined variance is larger than any single group's. So a variance computed between groups, compared against one computed within groups, detects a difference in means. The null hypothesis is that all k group means are equal and the alternative is that at least two are not, which means a rejection identifies no particular pair. Five assumptions have to hold, and the third — equal population variances — is the one quietly carrying the argument.

Subject: Statistics · 65 slides · symbolic lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Section 13.1 One-Way ANOVA

Title

Statistics · Chapter 13 — F Distribution and One-Way ANOVA

One-Way ANOVA

2. By the end of this lesson you can

Objectives

Five outcomes, and the first is the one that makes the chapter make sense.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680 — the section these objectives are drawn from

3. What you already have

Warm-up

Section 10.1 compared two population means with a t test.

Discussion prompt

An environmentalist wants to know whether the mean pollution level differs among four bodies of water. Why not run a t test on every pair?

Hint: Count the pairs, then think about what each test's error rate contributes.

Answer:

Four groups give six pairs, so six tests. Each carries its own 5 percent chance of a false alarm when nothing is going on, and running six of them makes the chance that at least one fires roughly 26 percent rather than 5.

\[ 1 - (0.95)^6 \approx 0.26 \]

The book states this directly: it is preferable to use ANOVA when there are more than two groups instead of performing pairwise t-tests, because performing multiple tests introduces the likelihood of making a Type 1 error.

What is wanted instead is one test asking a single question — are all four means equal? — with one 5 percent error rate attached to it. That is what analysis of variance provides, and the surprise is that it answers a question about means by comparing variances.

4. A test about means, built from variances

Concept

The purpose of a one-way ANOVA test is to determine the existence of a statistically significant difference among several group means. The test actually uses variances to help determine if the means are equal or not.

analysis of variance — A method for comparing averages between more than two groups. One-way, or single factor, ANOVA is the simplest form, and the one this chapter covers.

\[ H_0: \mu_1 = \mu_2 = \cdots = \mu_k \]

The mechanism is worth stating before any formula appears. If the null hypothesis is true the variance of the combined data is approximately the same as the variance of each of the populations; if it is false, the variance of the combined data is larger, and that increase is caused by the different means. Everything in section 13.2 is a way of measuring that increase.

Figure (svg): Two sets of three box plots, the first with the three medians level and the second with them stepping upward

The book's Figure 13.2: in (a) the differences are due to random variation; in (b) they are too large to be.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680

5. How variances reveal means

Section

Section 1

6. Pooling spreads data out when the means differ

Concept

In the first graph the null hypothesis holds and the three populations have the same distribution, so the variance of the combined data is approximately the same as the variance of each of the populations. If the null hypothesis is false, the variance of the combined data is larger, which is caused by the different means.

combined variance — The variance of all the observations pooled together, ignoring which group they came from. It exceeds the typical within-group variance exactly when the group means differ.

\[ \text{means equal} \;\Rightarrow\; \text{combined variance} \approx \text{group variance} \]

The asymmetry is what makes the test work. Spreading the group means apart always increases the combined variance and never decreases it, so the comparison has a direction — which is why section 13.3 can say the one-way ANOVA test is always right-tailed. It is the same reason chapter 11's chi-square tests were right-tailed, arrived at differently.

Figure (svg): Two sets of three box plots, the first with the three medians level and the second with them stepping upward

The book's Figure 13.2: in (a) the differences are due to random variation; in (b) they are too large to be.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680 — the two box-plot graphs and their interpretation

7. The book's Figure 13.2

Picture it

Two possibilities, drawn as box plots.

Figure (svg): Two sets of three box plots, the first with the three medians level and the second with them stepping upward

The book's Figure 13.2: in (a) the differences are due to random variation; in (b) they are too large to be.

In the left panel the differences among the group medians are due to random variation; in the right they are too large to be. The test's job is to decide which picture the data resemble, and it does so by measuring the combined spread.

8. Worked example: what happens to the combined variance

Worked example

Three groups of five with within-group variance 4. First all means are 20; then they are 12, 20 and 28.

\[ \text{same means, then spread means} \]

Case one: all means 20

Why: Pooling adds nothing.

\[ \text{combined variance about } 4 \]

Case two: means 12, 20, 28

Why: The means themselves vary.

\[ \text{variance of means is } 64 \]

Combined spread

Why: Within plus between.

\[ \text{about } 4 + 64 \]

Compare

Why: Seventeen times larger.

Figure (svg): The solution to Worked example what happens to the combined variance shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 4 \;\longrightarrow\; \approx 68 \]

Verify: confirm the within-group variance did not change

Why: Each group still has variance 4 — shifting a whole group up or down moves its values together and leaves its internal spread untouched. That is the key: the WITHIN-group variance is blind to differences in means, while the combined variance is not. Comparing the two isolates exactly the quantity of interest, which is what the F ratio does.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

9. One of these is false

Two truths and a lie

All three concern the mechanism.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Different group means inflate the combined variance
  • C. The within-group variance is unaffected by differing means
  • B. The test concludes that the group variances are different

Survives elimination: B

Why: The survivor is false. Equal variances is one of the five ASSUMPTIONS; the conclusion is about means. Section 13.4's test of two variances is the one whose hypotheses concern variances.

10. Worked example: reading the two pictures

Worked example

The book's box plots, with three group medians in each.

\[ \text{(a) level, (b) stepping upward} \]

Panel (a)

Why: Medians level.

\[ H 0\text{ plausible} \]

Its combined spread

Why: Same as each box.

Panel (b)

Why: Medians step up.

\[ H 0\text{ doubtful} \]

Its combined spread

Why: Wider than any box.

Figure (svg): The solution to Worked example reading the two pictures shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{(a) H}_0 \text{ true}; \quad \text{(b) H}_0 \text{ false} \]

Verify: confirm what makes the second panel convincing rather than merely different

Why: The book's caption is careful: in (a) the differences are due to random variation, and in (b) the differences are too large to be due to random variation. Some difference among sample means always occurs, so the question is never whether they differ but whether they differ by more than within-group variation would readily produce. The F ratio is the measurement of exactly that.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

11. Trap: expecting the test to compare variances as its conclusion

Trap

The trap

\[ \text{ANOVA} \;\Rightarrow\; \text{the group VARIANCES differ} \]

Read the name as the conclusion

Why: It is called analysis of variance.

\[ \text{but the hypotheses are about MEANS} \]

Equal variances is an ASSUMPTION of the test, not what it concludes about.

The fix

\[ H_0: \mu_1 = \cdots = \mu_k, \; \text{tested by comparing variances} \]

Read the name as the method, and the hypotheses as the question

Why: Variances are the instrument.

The confusion is easy because section 13.4 does test two variances, using the same F distribution. Keeping the two apart matters: one-way ANOVA assumes equal variances in order to conclude about means, while a test of two variances concludes about the variances themselves.

12. Why does this direction work?

Prediction

Commit before reasoning.

Predict first

Why can differing means only INCREASE the combined variance, never decrease it?

  • Separating group centres adds spread to the pooled data on top of the within-group spread
  • Because variances are positive
  • Because the samples are random
  • It can decrease it, in some cases

Correct: Separating the centres adds spread.

Why: Pooling groups whose centres coincide gives the same spread as one group; moving the centres apart pushes observations further from the overall mean without changing any group's internal spread. That one-directional effect is why section 13.3 can say the ANOVA test is always right-tailed.

13. Which picture does it match?

Sorting

Each describes three groups.

Sort into buckets

Sort by which of the book's two panels it resembles.

Panel (a): differences plausibly random
sample means 24.2, 25.4 and 24.4, each group's variance about 15; sample means 3.5, 5.0 and 5.2, each group's variance about 3; sample means 20.1, 19.9 and 20.0, each group's variance about 6
Panel (b): differences too large for that
sample means 12, 20 and 28, each group's variance about 4; sample means 3,512, 5,504 and 7,804, each variance about 1,000,000
a1
The means differ by little relative to the within-group spread.
b1
The means differ by much more than the within-group spread would explain.

Item (a) is Example 13.4's bean plants, whose F is 0.134 and p is 0.876; item (d) is Example 13.2's tomatoes, whose F is 4.481 and p is 0.0248. The comparison that matters is always means against within-group spread, never means alone.

14. The comparison in words

Faded example

Complete the sentence describing what the test measures.

Fill in the blanks

\textwithin against \text___ ___ H_0

Why: Between-group variance responds to differences in means; within-group variance does not. Their ratio is the F statistic, and it is large exactly when the means are far apart relative to the noise.

15. The hypotheses

Section

Section 2

16. All equal, against at least two not

Concept

The null hypothesis is simply that all the group population means are the same. The alternative hypothesis is that at least one pair of means is different — that is, mu-i is not equal to mu-j for some i not equal to j.

an omnibus alternative — It asserts only that the means are not all equal. Rejecting the null therefore identifies no particular pair as different, and identifying one needs methods beyond this chapter.

\[ H_0: \mu_1 = \cdots = \mu_k; \qquad H_a: \mu_i \ne \mu_j \text{ for some } i \ne j \]

The null is one claim rather than several. Writing it as a chain of equalities asserts that all k means coincide, and its negation is satisfied by any single departure — one group differing from the rest, or all k differing from each other. That breadth is exactly what makes a single test possible, and it is what the test gives up in precision.

Figure (svg): A card giving the null and alternative hypotheses for one-way ANOVA

The alternative is the plain negation of the null, which is why it can say only that some pair differs.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680 — the null and alternative hypotheses

17. The two hypotheses

Picture it

For k groups.

Figure (svg): A card giving the null and alternative hypotheses for one-way ANOVA

The alternative is the plain negation of the null, which is why it can say only that some pair differs.

The lower panels are the price and the reason. One test carries one error rate because it asks one question — and the answer, when it comes, is correspondingly unspecific.

18. Worked example: writing the hypotheses for five groups

Worked example

Example 13.2 compares tomato yields under five mulching conditions.

\[ k = 5 \]

Name the parameters

Why: One mean per condition.

\[ \mu 1\text{ through } \mu 5 \]

The null

Why: A chain of equalities.

The alternative

Why: The negation.

Note what it omits

Why: No pair named.

Figure (svg): The solution to Worked example writing the hypotheses for five groups shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ H_0: \mu_1 = \mu_2 = \mu_3 = \mu_4 = \mu_5 \]

Verify: confirm the alternative covers every way the null can fail

Why: The null asserts four equalities at once, and its negation holds if any one of them fails — whether one condition stands apart or all five differ. The alternative has to be that broad to be the null's negation, which is why the book writes it as some i not equal to j rather than naming a pattern.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

19. Write the hypotheses

Faded example

For three diet plans, as in Example 13.1.

Fill in the blanks

H_0: \mu_1 = \mu_2 = μ3; \qquad H_a: \texttwo ___ \text___

Why: The alternative requires only two means to differ, not all three — it is the negation of the null, not its opposite extreme.

20. Worked example: what the book's conclusion says

Worked example

Example 13.2 rejects at 5 percent with a p-value of 0.0248.

\[ \text{reject } H_0 \]

What was rejected

Why: All means equal.

What follows

Why: Some differ.

Which two

Why: Not identified.

The book's wording

Why: Some of the mulches.

Figure (svg): The solution to Worked example what the book's conclusion says shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{some } \mu_i \ne \mu_j, \text{ pair unspecified} \]

Verify: confirm the wording avoids over-claiming

Why: The book writes that we may conclude that at least SOME of the mulches led to different mean yields, and describes the evidence as making the differences unlikely to be due to chance alone. Neither phrase names a condition or a direction, which matches exactly what the alternative hypothesis asserts.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

21. Error analysis: four statements of the hypotheses

Error analysis

For four groups. Which are correct?

Annotate

On: \( \begin{aligned} &(1)\; H_0: \mu_1 = \mu_2 = \mu_3 = \mu_4 \\ &(2)\; H_a: \mu_1 \ne \mu_2 \ne \mu_3 \ne \mu_4 \\ &(3)\; H_a: \text{at least two of the means are unequal} \\ &(4)\; H_0: \bar{x}_1 = \bar{x}_2 = \bar{x}_3 = \bar{x}_4 \end{aligned} \)

  • (1) is the book's own null: a single chain of equalities among the population means.
  • (2) asserts that ALL four differ, which is stronger than the null's negation. Three equal means and one different would satisfy the correct alternative and not this one.
  • (3) is correct, and is the book's wording.
  • (4) uses sample means, which are known from the data. Hypotheses are always about population parameters.

Error (2) is the substantive one. The alternative must be the exact negation of the null, and the negation of all four being equal is that they are not all equal — a much weaker claim than all four being different.

22. One of these is false

Two truths and a lie

All three concern the hypotheses.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. The null asserts all k means are equal
  • C. Rejecting identifies no particular pair
  • B. The alternative asserts that all k means differ

Survives elimination: B

Why: The survivor is false. The alternative is the null's negation, satisfied by a single departure — one group standing apart while the others coincide would satisfy it, and would not satisfy a claim that all k differ.

23. What does rejecting license?

Prediction

Commit before reasoning.

Predict first

A one-way ANOVA on four groups rejects. What may be concluded?

  • At least two of the four means differ, with no pair identified
  • All four means differ
  • The first and last means differ
  • The group with the largest sample mean differs from the rest

Correct: At least two differ, with no pair identified.

Why: The alternative hypothesis says only that some pair is unequal. Reading a specific pair out of the result — even the pair with the most extreme sample means — is a separate claim the test did not examine, and the book's own conclusions are worded to avoid it.

24. How many pairwise tests would it take?

Estimation

Five groups compared pairwise.

Predict first

How many pairs, and roughly what chance of at least one false alarm at 5 percent?

  • 10 pairs, about 40 percent
  • 5 pairs, about 25 percent
  • 20 pairs, about 60 percent
  • 10 pairs, about 5 percent

Correct: 10 pairs, about 40 percent.

Why: Five groups give ten pairs, and one minus 0.95 to the tenth is about 0.40. That is eight times the intended error rate, and it is the precise reason the book prefers a single ANOVA to pairwise tests.

25. The five assumptions

Section

Section 3

26. What has to hold before the test means anything

Concept

Each population from which a sample is taken is assumed to be normal; all samples are randomly selected and independent; the populations are assumed to have equal standard deviations or variances; the factor is a categorical variable; and the response is a numerical variable.

the equal-variance assumption — The third, and the one the argument rests on. If groups may have different spreads, an inflated combined variance no longer implies different means.

\[ \sigma_1 = \sigma_2 = \cdots = \sigma_k \]

The book makes the connection explicit later: the hypothesis of equal means implies that the populations have the same normal distribution, BECAUSE it is assumed that the populations are normal and that they have equal variances. So the first and third assumptions together turn a claim about means into a claim about whole distributions — which is what lets the test proceed by comparing spreads.

Figure (svg): The five assumptions for a one-way ANOVA test

The last two describe the data's shape: one categorical grouping variable, one numerical measurement.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680 — the five assumptions, and the note on equal means implying the same distribution

27. The five assumptions

Picture it

The book's own list.

Figure (svg): The five assumptions for a one-way ANOVA test

The last two describe the data's shape: one categorical grouping variable, one numerical measurement.

The last two are structural rather than statistical: they describe what kind of data the method applies to. The first three are conditions on the populations, and only the second can be secured by design.

28. Worked example: why equal variances matters

Worked example

Suppose three groups have equal means but variances 1, 1 and 100.

\[ \mu \text{ equal}, \; \sigma^2 = 1, 1, 100 \]

The means

Why: All the same.

The combined variance

Why: Dominated by group three.

\[ \text{about } 34 \]

Against a typical group

Why: Two have variance 1.

The wrong inference

Why: Looks like differing means.

Figure (svg): The solution to Worked example why equal variances matters shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{large combined variance} \;\nRightarrow\; \text{unequal means} \]

Verify: confirm what the test would actually do here

Why: MSwithin averages the group variances and would come out near 34, which is large — so the F ratio would not be inflated, and the test would probably not reject. The real damage is subtler: with unequal variances the F statistic no longer follows the F distribution the p-value is read from, so the stated error rate is wrong in a direction that depends on which groups are largest. That is why the assumption is stated rather than checked casually.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 679

29. Assumption or conclusion?

Sorting

Each is a statement about a one-way ANOVA.

Sort into buckets

Sort each one.

An assumption of the test
the populations have equal variances; each population is normal; the samples are independent
A hypothesis being tested
at least two group means differ; the group means are all equal
as
It must hold for the test to be valid, and is not what the test examines.
hy
It is the null or the alternative — the claim the test decides between.

Item (a) is the one most often misplaced, because the method is called analysis of VARIANCE. Equal variances is assumed; equal means is tested.

30. Worked example: identifying factor and response

Worked example

Example 13.3 compares grade means across four sororities.

\[ \text{which variable is which?} \]

What sorts the observations

Why: Which sorority.

Its type

Why: Names, not numbers.

What is measured

Why: Grade mean.

Its type

Why: A number.

Figure (svg): The solution to Worked example identifying factor and response shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ \text{factor: sorority}; \quad \text{response: grade mean} \]

Verify: confirm the response must be numerical for the question to make sense

Why: The hypotheses compare population MEANS, and a mean can only be taken of numbers. A categorical response — pass or fail, say — has no mean to compare, and a different method entirely would be needed. So assumption five is not a technicality but a restatement of what the test is asking.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680

31. Trap: treating the assumptions as formalities

Trap

The trap

\[ \text{the data are grouped and numeric, so run the test} \]

Check only that the arithmetic can be done

Why: Every example has this shape.

\[ \text{but normality and equal variances are separate claims} \]

A strongly skewed response or wildly unequal group spreads makes the p-value unreliable in ways that are not visible in the output.

The fix

\[ \text{check the shape and spread of each group first} \]

Look at the groups' distributions before testing

Why: Box plots per group show both at once.

The book's own Figure 13.2 is a set of box plots, which is not a coincidence: side-by-side box plots show the centres, the spreads and the shapes together, so all three of the statistical assumptions can be inspected in one picture before any F ratio is computed.

32. Assumption to what it secures

Matching

Match each assumption to what it makes possible.

Match the pairs

  • l1. populations are normal
  • l2. populations have equal variances
  • l3. the factor is categorical
  • l4. the response is numerical
  • r1. the F statistic follows an F distribution
  • r2. an inflated combined variance implies unequal means
  • r3. the observations sort into groups
  • r4. there are means to compare at all

Why: The first two are statistical and the last two structural. The second is the one carrying the chapter's central argument, which is why it is worth naming separately.

33. One of these is false

Two truths and a lie

All three concern the assumptions.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. There are five of them
  • C. Equal variances is assumed, not tested
  • B. The samples may be dependent if the groups are large

Survives elimination: B

Why: The survivor is false. The second assumption requires all samples to be randomly selected and independent, and no sample size relaxes it. Dependent groups — the same subjects measured under several conditions — need a different method entirely.

34. What does the equal-variance assumption buy?

Prediction

Commit before reasoning.

Predict first

Combined with normality, what does assuming equal variances let the null hypothesis say?

  • That all the populations have the same normal distribution, not merely the same mean
  • That the samples are the same size
  • That the response is numerical
  • That the test is right-tailed

Correct: That all the populations have the same distribution.

Why: The book states it: the hypothesis of equal means implies that the populations have the same normal distribution, because it is assumed that the populations are normal and have equal variances. That is what makes a within-group variance a legitimate estimate of one common population variance.

35. Why one test rather than many

Section

Section 4

36. Multiple tests accumulate error

Concept

It is preferable to use ANOVA when there are more than two groups instead of performing pairwise t-tests, because performing multiple tests introduces the likelihood of making a Type 1 error.

a Type 1 error — Rejecting a true null hypothesis. Each test at 5 percent carries that risk, and running many tests makes the chance of at least one false alarm far exceed 5 percent.

\[ k \text{ groups} \;\Rightarrow\; \tfrac{k(k-1)}{2} \text{ pairs} \]

The arithmetic is stark. Three groups give three pairs and a combined false-alarm risk near 14 percent; ten groups give 45 pairs and a risk near 90 percent. So with ten genuinely identical populations, pairwise testing would almost certainly report a difference — which makes the reported 5 percent meaningless.

Figure (svg): A card showing the number of pairwise comparisons among k groups and the resulting error risk

One omnibus test at 5 percent replaces many pairwise tests whose combined error rate is far above 5 percent.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680 — the note preferring ANOVA to pairwise tests

37. The cost of many tests

Picture it

Pairs and accumulated risk, at three group counts.

Figure (svg): A card showing the number of pairwise comparisons among k groups and the resulting error risk

One omnibus test at 5 percent replaces many pairwise tests whose combined error rate is far above 5 percent.

One omnibus test at 5 percent carries a 5 percent risk. That is the trade: a single controlled error rate, in exchange for an answer that names no particular pair.

38. Worked example: computing the accumulated risk

Worked example

Five groups, all with identical population means, compared by pairwise t tests at 5 percent.

\[ k = 5 \]

Count the pairs

Why: Five choose two.

\[ 10 \]

Each test's non-rejection chance

Why: One minus alpha.

\[ 0.95 \]

All ten correct

Why: If independent.

\[ 0.95\text{ to the tenth} \]

At least one wrong

Why: The complement.

\[ \text{about } 0.40 \]

Figure (svg): The solution to Worked example computing the accumulated risk shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 1 - (0.95)^{10} \approx 0.40 \]

Verify: confirm the caveat about independence

Why: The ten pairwise tests share data — the same group appears in four of them — so they are not independent and the exact figure differs somewhat from 0.40. The direction and rough size are right, and that is what matters: the accumulated risk is far above 5 percent and grows quickly with k, which is the whole argument for a single test.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

39. Count the pairs

Faded example

Six groups compared pairwise.

Fill in the blanks

\frac5}}15 = ___ \text___

Why: Fifteen tests at 5 percent give a combined false-alarm chance near 54 percent — more than a coin flip, with all six populations identical.

40. Worked example: what the single test costs

Worked example

Comparing what each approach delivers.

\[ \text{ANOVA against pairwise} \]

ANOVA's error rate

Why: One test.

\[ 5 \% \]

ANOVA's answer

Why: Some pair differs.

Pairwise error rate

Why: Accumulates.

\[ 40 \%\text{ at } k = 5 \]

Pairwise answer

Why: Names pairs.

Figure (svg): The solution to Worked example what the single test costs shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ 5\% \text{ and vague} \quad\text{against}\quad 40\% \text{ and specific} \]

Verify: confirm this is a real trade rather than one option dominating

Why: It is a genuine trade, and later courses resolve it with methods that make pairwise comparisons while controlling the overall error rate. The book flags that this is a very brief overview and that the topic is studied in much greater detail in future statistics courses — the follow-up comparison is exactly what those courses supply.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680

41. Error analysis: four responses to comparing several groups

Error analysis

Which are sound?

Annotate

On: \( \begin{aligned} &(1)\; \text{run one ANOVA at } 5\% \\ &(2)\; \text{run all pairwise } t \text{ tests at } 5\% \\ &(3)\; \text{run only the pair with the most extreme sample means} \\ &(4)\; \text{run an ANOVA, then investigate further if it rejects} \end{aligned} \)

  • (1) is the book's recommendation, and carries one controlled error rate.
  • (2) accumulates the error rate, reaching about 40 percent at five groups.
  • (3) is worse than (2). Choosing the pair after seeing the data guarantees the most extreme comparison, so its true error rate is far above 5 percent even though only one test was run.
  • (4) is the practical approach, and it is what later courses formalise.

Error (3) is the subtle one, and it is worth naming because it looks like the disciplined choice. Running one test sounds conservative, but selecting WHICH test to run by inspecting the data reintroduces every comparison that was not run.

42. One of these is false

Two truths and a lie

All three concern multiple testing.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. Each pairwise test carries its own error risk
  • C. One ANOVA carries a single controlled error rate
  • B. Running one carefully chosen pairwise test avoids the problem

Survives elimination: B

Why: The survivor is false when the choice is made after seeing the data. Selecting the most extreme-looking pair guarantees the largest difference, so the effective error rate reflects all the comparisons that could have been chosen, not the one that was.

43. At what k does the risk exceed half?

Estimation

The risk is one minus 0.95 raised to the number of pairs.

Predict first

Roughly how many groups before the chance of a false alarm passes 50 percent?

  • About 6
  • About 3
  • About 12
  • About 20

Correct: About 6.

Why: Six groups give fifteen pairs, and one minus 0.95 to the fifteenth is about 0.54. So with only six identical populations, pairwise testing is more likely than not to report a spurious difference — which is a small number of groups for so complete a failure.

44. Explain the preference

Explain it

A classmate asks why they cannot just run a t test on each pair, since they know how to do those already.

Discussion prompt

In two sentences or fewer, explain.

Hint: Ask what happens to the error rate.

Answer:

Each t test carries its own 5 percent chance of a false alarm, and with five groups there are ten of them, so the chance that at least one fires when nothing is happening is about 40 percent.

One ANOVA asks a single question and carries a single 5 percent rate, at the cost of not naming which pair differs.

45. Reading a study's structure

Section

Section 5

46. One factor, one response, k groups

Concept

The factor is a categorical variable that sorts the observations into groups, and the response is a numerical variable measured on each observation. One-way ANOVA is also called single factor ANOVA, because exactly one factor does the sorting.

one-way — One factor. Studies with two grouping variables need two-way analysis of variance, which the book says is beyond the scope of this chapter.

\[ k \text{ groups}, \; n = \textstyle\sum n_j \text{ observations} \]

The word single carries a real restriction. A study comparing tomato yields under five mulches AND three watering schedules has two factors, and one-way ANOVA cannot handle it — the interaction between them is exactly what two-way analysis exists to examine. Every example in this chapter has one factor, and identifying it is the first step of any ANOVA.

Figure (svg): A table of four studies showing the factor and the response in each

Every example in the chapter has this form, and identifying the two variables is the first step of any ANOVA.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680 — single factor or one-way ANOVA, and the chapter's examples

47. Four studies

Picture it

The factor and the response in each.

Figure (svg): A table of four studies showing the factor and the response in each

Every example in the chapter has this form, and identifying the two variables is the first step of any ANOVA.

Each row has one categorical sorting variable and one numerical measurement. Recognising that shape is what identifies a problem as a one-way ANOVA rather than something else.

48. Worked example: identifying k and n

Worked example

Example 13.2's tomatoes: five mulching conditions with three plants each.

\[ 5 \text{ groups of } 3 \]

Count the groups

Why: Five conditions.

\[ k = 5 \]

Count per group

Why: Three plants each.

\[ n _{j} = 3 \]

Total observations

Why: Five times three.

\[ n = 15 \]

Note the design

Why: Equal sizes.

Figure (svg): The solution to Worked example identifying k and n shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ k = 5, \quad n = 15 \]

Verify: confirm why both counts are needed

Why: Section 13.2 will use k minus one and n minus k as the two degrees of freedom, so both counts enter the test directly. For these data that gives 4 and 10 — the book's F-sub-4-comma-10 — and neither number can be recovered from the other without knowing the group sizes.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

49. Count k and n

Faded example

Four sororities with five sisters sampled from each.

Fill in the blanks

k = 4, \qquad n = 20

Why: That is Example 13.3, whose degrees of freedom are therefore 3 and 16 — the book's F-sub-3-comma-16.

50. Worked example: balanced against unbalanced

Worked example

Example 13.1 has group sizes 4, 3 and 3; Example 13.3 has four groups of five.

\[ \text{which is balanced?} \]

Example 13.1

Why: Sizes 4, 3, 3.

Example 13.3

Why: All five.

What changes

Why: The arithmetic.

What does not

Why: The hypotheses.

Figure (svg): The solution to Worked example balanced against unbalanced shown as a ladder of expressions, one row per legal move

The whole solution at once: each drop is one legal move.

\[ n_1 = \cdots = n_k \;\Rightarrow\; \text{balanced} \]

Verify: confirm the shortcut and where it appears

Why: Section 13.2 gives an F-ratio formula for when the groups are the same size, using the variance of the sample means and the mean of the sample variances — and Example 13.4 works its bean-plant data that way. Unbalanced designs need the weighted sum-of-squares computation instead, which is what Example 13.1 demonstrates. Both give the same F when both apply.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, p. 680

51. Trap: treating two grouping variables as one factor

Trap

The trap

\[ \text{5 mulches} \times 3 \text{ waterings} \;\Rightarrow\; \text{15 groups, one-way ANOVA} \]

Collapse two factors into a single list of combinations

Why: It produces groups that can be compared.

\[ \text{but the two factors may interact} \]

A mulch might help under heavy watering and hurt under light, and a one-way test cannot detect or describe that.

The fix

\[ \text{two factors} \;\Rightarrow\; \text{two-way analysis of variance} \]

Count the grouping variables before choosing the method

Why: One-way means exactly one factor.

The book names the boundary directly: two-way analysis is beyond the scope of this chapter, and is listed among the other uses of the F distribution. Running a fifteen-group one-way test would produce a valid p-value for a question nobody asked — whether all fifteen combinations have the same mean — while leaving the actual question about mulch and water untouched.

52. Factor or response?

Sorting

Each is a variable from one of the chapter's studies.

Sort into buckets

Sort each one.

The factor: categorical
which of three diet plans; which of five soil covers; which sorority
The response: numerical
weight lost in pounds; tomato yield in grams
f
It sorts observations into groups and takes category values.
r
It is measured on each observation and takes numerical values.

The response must be numerical because the hypotheses compare means, and a mean can only be taken of numbers. That is the content of assumption five.

53. One of these is false

Two truths and a lie

All three concern the structure.

Eliminate the wrong options

Two are true. Knock those out and keep the false one.

  • A. One-way means exactly one factor
  • C. The groups need not be the same size
  • B. Two grouping variables can be combined into one factor

Survives elimination: B

Why: The survivor is false. Combining them tests whether all the combinations have equal means, which answers neither question about the individual factors and cannot detect an interaction between them. Two factors need two-way analysis, which the book places beyond this chapter.

54. What does 'balanced' mean?

Prediction

Commit before reasoning.

Predict first

The book calls Example 13.3 a balanced design. What does that mean?

  • Each group has the same number of observations
  • The group means are equal
  • The group variances are equal
  • The samples are random

Correct: Each group has the same number of observations.

Why: The book's note says a balanced design is one where each factor has the same number of observations — four sororities with five sisters each. It permits a simpler F-ratio formula but changes neither the hypotheses nor the conclusion.

55. A two-sample t test against one-way ANOVA

Comparison

Fill the blanks. ANOVA generalises the t test to more than two groups.

Comparison matrix

Two-sample t testOne-way ANOVA
Groups comparedexactly twotwo or more
The nullthe two means are equalall k means are equal
What rejecting nameswhich mean is largerno particular pair
Equal variancesnot required by Welch's versionassumed

The third row is the cost of the generalisation. A two-sample test keeps the sign of the difference; ANOVA compares variances and so discards direction entirely — the same trade chapter 11's chi-square tests made.

56. Setting up a one-way ANOVA, in order

Pattern

Five steps, all before any arithmetic.

  1. Identify the categorical factor and the numerical response.
  2. Count the groups k and the total observations n, noting whether the design is balanced.
  3. Check the five assumptions, looking especially at each group's shape and spread.
  4. State the null that all k means are equal and the alternative that at least two differ.
  5. Recall that a rejection will name no particular pair, and plan any follow-up accordingly.

Side-by-side box plots show centres, spreads and shapes at once, which lets the three statistical assumptions be inspected in one picture.

OpenStax Introductory Business Statistics 2e, §12.2 One-Way ANOVA §12.2 One-Way ANOVA

57. Check yourself 1 of 3

Check

The mechanism.

Check your understanding

Why does a test about means compare variances?

  • A. Because different means inflate the combined variance while leaving within-group variance unchanged (correct)
  • B. Because means are hard to compute
  • C. Because the test concludes about variances
  • D. Because the F distribution requires it

Answer: A

Why: Pooling groups whose centres differ spreads the data out, so an inflated combined variance is evidence the means are not all equal.

Why B tempts people
Group means are trivial to compute and are used in the calculation.
Why C tempts people
The hypotheses concern means; equal variances is an assumption.
Why D tempts people
The distribution follows from the method rather than dictating it.

58. Check yourself 2 of 3

Check

The alternative.

Check your understanding

For four groups, what does the alternative hypothesis assert?

  • A. At least two of the four means are unequal (correct)
  • B. All four means are unequal
  • C. The largest mean exceeds the smallest
  • D. The four variances are unequal

Answer: A

Why: The alternative is the null's negation, so a single departure from equality satisfies it.

Why B tempts people
That is stronger than the negation: three equal means and one different would satisfy the correct alternative.
Why C tempts people
That names a direction, which the alternative does not.
Why D tempts people
Equal variances is an assumption of the test, not a hypothesis.

59. Check yourself 3 of 3

Check

The assumptions.

Check your understanding

Which of these is an assumption rather than a conclusion?

  • A. The populations have equal variances (correct)
  • B. At least two means differ
  • C. The group means are all equal
  • D. The F statistic exceeds one

Answer: A

Why: Equal variances is the third of the five assumptions, and the argument connecting variances to means depends on it.

Why B tempts people
That is the alternative hypothesis.
Why C tempts people
That is the null hypothesis.
Why D tempts people
That is an outcome of the calculation, neither assumed nor hypothesised.

60. Where this shows up outside the textbook

Real world

A clinic compares recovery times across four treatment protocols using six patients each, runs all six pairwise t tests at 5 percent, finds one pair significant, and reports that protocol B produces faster recovery than protocol D.

Discussion prompt

Assess the analysis and say what should have been done.

Hint: Count the tests, and ask how the reported pair was chosen.

Answer:

Six tests were run at 5 percent each, so the chance of at least one firing when all four protocols are identical is about 26 percent — five times the rate the report implies. Finding one significant pair among six is close to what pure chance produces.

\[ 1 - (0.95)^6 \approx 0.26 \]

Reporting only the significant pair makes it worse. The pair was selected by looking at which test came out significant, so the reported comparison is the most extreme of six — and its true error rate reflects all six, not one. This is the same problem as choosing which test to run after seeing the data.

The right approach is one ANOVA on all four protocols, carrying a single 5 percent error rate and testing whether the four mean recovery times are all equal. If it rejects, a follow-up comparison that controls the overall error rate identifies which protocols differ — that is what later courses supply, and it is exactly the gap the book flags when it says this is a very brief overview.

Two further things need checking before any of it. Six patients per group is small, so the normality assumption matters and cannot be verified well from six observations; and recovery times are often right-skewed, which would violate it. Side-by-side box plots of the four groups would show the shapes and spreads together — and with groups this small, the honest report may be that the study cannot distinguish four protocols at all.

61. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

How can comparing variances test a claim about means?

  • It cannot; the name is historical
  • Because separating the group means inflates the combined variance while leaving within-group variance unchanged
  • Because variances and means are proportional
  • Because the F distribution converts one into the other

Correct: Separating the means inflates the combined variance.

\[ \text{within-group variance: blind to the means}; \quad \text{between-group: not} \]

Why: Shifting a whole group up or down moves its values together, so its internal spread is unchanged — but pooling groups whose centres differ pushes observations further from the overall mean and widens the combined spread. Comparing a variance measured between groups against one measured within groups therefore isolates the differences among means, which is what the book's two box-plot panels show.

62. Explain it to someone a year behind you

Explain it

They read that ANOVA rejected for four groups and wrote that all four means are different.

Discussion prompt

In two sentences or fewer, correct them.

Hint: Ask what the alternative hypothesis actually claims.

Answer:

The alternative says only that at least two of the four means differ, so one group standing apart while the other three coincide would produce exactly this result.

The test names no particular pair, and identifying which groups differ needs a follow-up method beyond this chapter.

63. Exit ticket

Exit ticket

Name the weakest spot before you close the deck.

Predict first

Which of these would you least want handed to you cold?

  • Explaining how variances can test a claim about means
  • Stating the hypotheses and what a rejection does not identify
  • Listing the five assumptions and why equal variances matters
  • Explaining why one test beats many pairwise ones

Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.

Why: For the first, separating means inflates the combined spread and not the within-group spread. For the second, all equal against at least two unequal, with no pair named. For the third, without equal variances an inflated combined variance no longer implies unequal means. For the fourth, each test carries its own error risk and they accumulate. Do five problems of your chosen kind rather than twenty mixed ones.

64. Draw the lesson on one page

Connect it up

Paper. Twelve minutes.

Draw it

At the top, draw two sets of three box plots side by side. In the left set put the three medians level with each other; in the right set step them upward. Label the left H0 true and the right H0 false, and beneath each write what happens to the variance of the combined data — approximately the same as each group's on the left, larger on the right. Under the pair, write the sentence that makes the chapter work: the test uses VARIANCES to decide whether the MEANS are equal. In the middle of the page, write the two hypotheses — the null as a chain of equalities among all k means, and the alternative that at least two differ — and box a note saying that rejecting names no particular pair. Below that, list the five assumptions, and circle the third, writing beside it that without equal variances an inflated combined spread no longer implies unequal means. At the bottom, draw three cells for k equal to 3, 5 and 10, giving the pair counts 3, 10 and 45 and the accumulated false-alarm risks of about 14, 40 and 90 percent, with a note that one ANOVA carries one 5 percent rate instead.

Check your box plots by confirming the individual boxes are the same height in both panels — the within-group spread must not change, or the picture makes the wrong point. Check your bottom row by computing one minus 0.95 to the power of the pair count for each cell.

65. What you can do now

Recap

Five things, and the first is what makes the rest coherent.

If you seeThen
More than two group means to compareOne-way ANOVA, not pairwise t tests
One categorical factor and one numerical responseThe one-way ANOVA structure
Two grouping variablesTwo-way analysis, beyond this chapter
Equal group sizesA balanced design, with a simpler F formula available
A rejectionAt least two means differ; no pair is named
Groups with very unequal spreadsThe third assumption is in doubt
A categorical responseNo means to compare; a different method is needed

Section 13.2 turns the box-plot picture into arithmetic. The variance between samples and the variance within samples each become a mean square, their ratio becomes the F statistic, and the distribution that statistic follows is the one the chapter is named after.

OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA §13.1, pp. 679-680 — everything on these slides traces back here

Sources

  1. OpenStax Introductory Statistics 2e, §13.1 One-Way ANOVA — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 679-680
  2. OpenStax Introductory Business Statistics 2e, §12.2 One-Way ANOVA — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108