2.7 Measures of the Spread of the Data

Build variance and standard deviation from squares you can see, prove why a sample divides by n − 1, derive Chebyshev's inequality, and compare data sets with z-scores.

Subject: Statistics · 176 slides · applied lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Measures of the Spread of the Data

Title

Statistics · §2.7

Variance, SD, the sample divisor and z-scores, each shown to be true

2. You will leave able to do these things with real data

Objectives

  1. Show why deviations total zero
  2. Compute variance and SD from data
  3. Choose the right divisor, and say why
  4. Place values several SDs out; cap far data
  5. Compare data sets with z-scores

3. Six rules for sums and means carry the proofs

Concept

\[ {\textstyle\sum} x = x_1 + \cdots + x_n \]

Σ means add up

Why: one term per value

\[ \bar{x} = \frac{{\textstyle\sum} x}{n},\ \ \mu = \frac{{\textstyle\sum} x}{N} \]

Means (§2.5)

Why: n or N values share the total

\[ {\textstyle\sum} (u - v) = {\textstyle\sum} u - {\textstyle\sum} v \]

Sums of differences

Why: addition can be regrouped in any order

\[ \text{avg}(u \pm v) = \text{avg}\,u \pm \text{avg}\,v \]

Averages add, subtract

Why: divide both split sums by the same n

\[ {\textstyle\sum} c\,u = c{\textstyle\sum} u,\ \ \text{avg}(c\,u) = c\,\text{avg}\,u \]

Factor out c

Why: every term shares c

\[ {\textstyle\sum}_{1}^{n} c = n\,c \]

Add c n times

Why: n equal copies

4. Five rules for roots and squares carry the later proofs

Concept

\[ \sqrt{a^2} = a\ \ (a \ge 0) \]

A root undoes a square

Why: area back to side

\[ \sqrt{ab} = \sqrt{a}\sqrt{b}\ \ (a, b \ge 0) \]

A root splits over a product

Why: √a·√b squared gives back ab

\[ (a + b)^2 = a^2 + 2ab + b^2 \]

Square of a sum

Why: four pieces of a square

\[ u^2 \ge 0,\ \ |u|^2 = u^2 \]

Squares ignore sign

Why: signs cancel in pairs

\[ (c\,d)^2 = c^2 d^2,\ \Big(\frac{u}{v}\Big)^2 = \frac{u^2}{v^2} \]

Square each factor

Why: factors can regroup

5. Five order rules say when an inequality survives

Concept

\[ c > 0:\ u \le v \Rightarrow c\,u \le c\,v,\ \frac{u}{c} \le \frac{v}{c} \]

Scale by c > 0

Why: order is kept

\[ 0 \le u \le v \Rightarrow u^2 \le v^2 \]

Square non-negatives

Why: bigger side, bigger square

\[ u \le v \Rightarrow 1 - u \ge 1 - v \]

Subtract from 1

Why: the order flips

\[ u_i \ge c \text{ for } m \text{ terms} \Rightarrow {\textstyle\sum} u_i \ge m\,c \]

Add m inequalities

Why: sums of ≥ keep ≥

\[ u \le v \Rightarrow u + c \le v + c \]

Add to both sides

Why: equal additions leave the gap unchanged

6. Order, sign and median rules finish the toolkit

Concept

\[ 0 < u \le v \Rightarrow \tfrac{1}{u} \ge \tfrac{1}{v} \]

Flip reciprocals of positives

Why: divide 1 by more, get less

\[ 0 \le u \le v \Rightarrow \sqrt{u} \le \sqrt{v} \]

Roots keep order

Why: a bigger area has a longer side

\[ b \ge 0 \Rightarrow a \le a + b \]

Adding a non-negative can only grow

Why: nothing is taken away

\[ u \le v,\ v \le w \Rightarrow u \le w \]

Chain inequalities

Why: a bound passes along

\[ (-u)(-v) = uv,\ \ (-u)v = -uv \]

Sign rule

Why: equal signs multiply to positive

\[ 1,\ 3,\ \mathbf{7},\ 8,\ 9 \]

Median: middle of the ordered list

Why: half lie on each side

7. Both checkout lines average 5 minutes, yet one is riskier with a train to catch

Prediction

Figure (svg): Dot plots of line A's waits 2, 4, 5, 6, 8 and line B's waits 0, 2, 4, 8, 11 minutes on axes from 0 to 12, both with the mean marked at 5

Predict first

Line A: 2, 4, 5, 6, 8 min. Line B: 0, 2, 4, 8, 11 min.

Your train leaves in 7 minutes. Which line?

  • Line A
  • Line B
  • Either: same mean

Correct: Line A

Why: Equal means hide different spreads. B's waits stray further from 5, which suggests more risk; the evidence today is that two of B's five waits pass 7, against one of A's.

8. Today's count hints line B is riskier, but five waits could be luck

Worked example

Figure (svg): Dot plots of line A's waits 2, 4, 5, 6, 8 and line B's waits 0, 2, 4, 8, 11 minutes on axes from 0 to 12, both with the mean marked at 5

\[ \text{A: } \textcolor{#1f5fbf}{8} > 7 \]

Count A's waits over 7

Why: evidence for the riskier line

\[ \text{B: } \textcolor{#1f5fbf}{8},\ \textcolor{#1f5fbf}{11} > 7 \]

Count B's waits over 7

Why: to compare with A's one

\[ 2 + 4 + 5 + 6 + 8 = 0 + 2 + 4 + 8 + 11 = 25 \]

Total each line's waits

Why: equal totals give equal means

\[ 25 \div 5 = 5 \]

Check: share each total among 5 waits

Why: both means are 5

Needed: a number for each line's spread

9. Measure each wait against the mean

Section

Idea 1 of 5

10. Line A's five waits sit at different distances from its 5-minute mean

Concept

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 0 deviation arrows from the mean

Goal: one number for how far these waits sit from 5.

Discussion prompt

What would you measure for each of the five waits?

Answer:

Its distance from the mean: 2 is 3 below, 8 is 3 above.

11. Line A's first three waits sit 3 below, 1 below and exactly at the mean

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 0 deviation arrows from the mean

deviation = value − mean (signed)

\[ \textcolor{#1f5fbf}{2} - 5 = \textcolor{#1f5fbf}{-3} \]

Subtract the mean from 2

Why: below 5, so negative

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 1 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{4} - 5 = \textcolor{#1f5fbf}{-1} \]

Subtract the mean from 4

Why: the sign marks the side

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 2 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{5} - 5 = \textcolor{#1f5fbf}{0} \]

Subtract the mean from 5

Why: a wait at the mean has no distance

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 3 deviation arrows from the mean

\[ 5 + (-3) = 2,\ 5 + (-1) = 4,\ 5 + 0 = 5 \]

Check: hop back from the mean

Why: each lands on its wait

12. Line A's signed distances total 0, yet its waits vary

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 3 deviation arrows from the mean

\( {{2 - 5 = -3}},\quad\allowbreak \allowbreak {{4 - 5 = -1}},\quad\allowbreak \allowbreak {{5 - 5 = 0}} \)

\[ \textcolor{#1f5fbf}{6} - 5 = \textcolor{#1f5fbf}{+1} \]

Subtract the mean from 6

Why: above, so positive

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 4 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{8} - 5 = \textcolor{#1f5fbf}{+3} \]

Subtract the mean from 8

Why: largest wait, longest arrow

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-3 - 1 + 0 + 1 + 3} = 0 \]

Add the distances

Why: spread should show up

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ 0 \div 5 = 0 \]

Divide by 5

Why: should give a typical distance

\[ (-3 + 3) + (-1 + 1) + 0 = 0 \]

Check: pair the arrows

Why: each one cancels

13. Line B's distances are bigger, so compare its shortfall with its overshoot

Prediction

Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 0 deviation arrows from the mean

Predict first

Line B: 0, 2, 4, 8, 11 minutes. Mean 5.

Which is bigger: how far B's low waits fall short of 5 in total, or how far its high waits overshoot?

  • The shortfall
  • The overshoot
  • They are equal

Correct: They are equal

Why: Below: 5 + 3 + 1 = 9 minutes short. Above: 3 + 6 = 9 minutes over. Equal, so B's signed distances total zero as line A's did.

14. Line B's distances cancel exactly as line A's did

Worked example

Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 0 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-5,\ -3,\ -1,\ +3,\ +6} \]

Subtract 5 from each wait

Why: to test whether B's arrows cancel too

Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean

\[ -5 - 3 - 1 = -9 \]

Add the three below

Why: the shortfall to compare

\[ +3 + 6 = +9 \]

Add the two above

Why: the overshoot to compare

\[ -9 + 9 = 0 \]

Combine the sides

Why: opposite amounts of equal size cancel

Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean

\[ (-5 + 6) + (-3 + 3) - 1 = 0 \]

Check: regroup the arrows

Why: any grouping gives one total

15. Deviations from the mean total zero for any data

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ {\textstyle\sum} (x - \bar{x}) = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{{\textstyle\sum} \bar{x}} \]

Split the sum

Why: differences add as totals

Figure (svg): Line A's waits 2, 4, 5, 6, 8 laid end to end reach 25; five copies of the mean 5 laid end to end also reach 25

\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{n\bar{x}} \]

Count the x̄'s

Why: one per value: n of them

Figure (svg): Line A's waits 2, 4, 5, 6, 8 laid end to end reach 25; five copies of the mean 5 laid end to end also reach 25

\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{n \cdot \frac{{\textstyle\sum} x}{n}} \]

Substitute x̄'s definition

Why: the prerequisite mean

Figure (svg): Line A's waits laid end to end reach 25; below them five copies of the mean, each now labelled Σx/n, also reach 25

\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#1f5fbf}{{\textstyle\sum} x} \]

Cancel n against n

Why: multiplying undoes dividing, leaving Σx

Figure (svg): Line A's waits laid end to end reach 25; below them the five copies of Σx/n have merged into one bar Σx = 25

\[ = 0 \]

Cancel the two Σx

Why: equal totals leave nothing over

Figure (svg): Line A's waits laid end to end and the merged bar Σx both end at 25, marked difference 0

\[ \textcolor{#1f5fbf}{25} - \textcolor{#6b7280}{5 \cdot 5} = 0 \]

Check with line A

Why: total 25, five 5s

16. Signed distances cancel, so they can't show spread

Concept

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-3 - 1} = \textcolor{#1f5fbf}{-4},\ \ \textcolor{#1f5fbf}{+1 + 3} = \textcolor{#1f5fbf}{+4} \]

Total each side's arrows

Why: to test the balance

\[ \text{waits } 5, 5, 5, 5, 5:\ \mu = 5 \]

Try no spread at all

Why: spread should read zero

Figure (svg): Five waits all equal to 5 stacked on a number line from 0 to 10, every one at the mean

\[ 5 - 5 = \textcolor{#1f5fbf}{0} \text{ for each wait} \]

Subtract the mean

Why: each value sits on it

\[ 0 + 0 + 0 + 0 + 0 = 0 \]

Add them

Why: signed sums miss spread

\[ (5 + 5 + 5 + 5 + 5) - 5 \times 5 = 0 \]

Check with Σx − Nμ

Why: a true mean leaves zero

17. Four of line A's signed distances force the fifth

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ -3 - 1 + 0 + 1 + \textcolor{#1f5fbf}{d} = 0 \]

Hide the last arrow as d

Why: the total must still be zero

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean, the last one grey and marked forced

\[ -3 + \textcolor{#1f5fbf}{d} = 0 \]

Add the four known distances

Why: leaves d as the only unknown

\[ \textcolor{#1f5fbf}{d} = +3 \]

Add 3 to both sides

Why: isolates the hidden distance

\[ 8 - 5 = +3 \]

Check against the hidden wait 8

Why: its arrow is 3 long

18. Averaging signed deviations reports zero spread for line B, whose waits run 0 to 11

Trap

The trap

\[ -5 - 3 - 1 + 3 + 6 = 0 \]

Add the signed distances

Why: hoping for a total spread

\[ 0 \div 5 = 0 \]

Divide by 5

Why: treated as the typical distance

\[ \text{spread} = 0 \]

Report zero spread

Why: fails here: the waits differ

The fix

\[ 5 + 3 + 1 + 3 + 6 = 18 \]

Add the distances, ignoring sign

Why: unsigned sizes cannot cancel

\[ 18 \div 5 = 3.6 \]

Divide by 5

Why: a typical distance, not zero

\[ 3.6 \times 5 = 18 \]

Check: undo the divide

Why: rebuilds the total distance

Idea 2 compares this with squaring.

19. One quiz score is lost, and the zero total pins down its deviation

Prediction

Predict first

Four quiz scores: three sit −2, +5, +1 from the mean. One is lost.

How far from the mean is the lost score?

  • 4 points above
  • 4 points below
  • At the mean
  • Cannot be found

Correct: 4 points below

Why: All four deviations must total zero. The three known ones total +4, so the lost one is −4: four points below.

20. The lost score sits 4 points below the mean

Worked example

\[ -2 + 5 + 1 + d = 0 \]

Call the lost deviation d

Why: all four must total zero

\[ 4 + d = 0 \]

Add the three known deviations

Why: what is still unbalanced

\[ d = -4 \]

Subtract 4 from both sides

Why: leaves the lost deviation alone

\[ -2 + 5 + 1 - 4 = 0 \]

Check the four balance

Why: zero, as it must be

21. Four runners' lap times give deviations you can check without a calculator

Prediction

Predict first

Four runners' lap times: 61, 66, 70, 83 seconds. Their mean is 70 seconds.

Which list is their deviations from the mean?

  • −9, −4, 0, +13
  • 9, 4, 0, 13
  • +9, +4, 0, −13
  • −9, −4, 0, +4

Correct: −9, −4, 0, +13

Why: Time minus mean gives negatives below 70, and the list must total zero. Only the first list does both.

22. The lap-time deviations are −9, −4, 0 and +13 seconds, and they balance

Worked example

Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 0 deviation arrows from the mean 70

\[ \textcolor{#1f5fbf}{61} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{-9} \]

Subtract the mean from 61

Why: below, so negative

Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 1 deviation arrows from the mean 70

\[ \textcolor{#1f5fbf}{66} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{-4} \]

Subtract it from 66

Why: below, so negative

Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 2 deviation arrows from the mean 70

\[ \textcolor{#1f5fbf}{70} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{0},\ \ \textcolor{#1f5fbf}{83} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{+13} \]

Subtract it from 70 and 83

Why: the sign marks the side

Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 4 deviation arrows from the mean 70

\[ \textcolor{#1f5fbf}{-9 - 4 + 0 + 13} = 0 \]

Check the total

Why: any list that fails is wrong

Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 4 deviation arrows from the mean 70

23. Square the distances so nothing cancels

Section

Idea 2 of 5

24. Line B needs a spread number that can't cancel

Concept

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with arrows from the mean 5 labelled by unsigned distance 5, 3, 1, 3, 6, distances total 18

Unsigned distances from line B's mean 5 total 18: an average of 3.6 minutes.

Discussion prompt

Measured from 4 instead of the mean 5, is line B's unsigned-distance total bigger or smaller?

Answer:

Smaller: 4 + 2 + 0 + 4 + 7 = 17. The next slides test squares the same way.

25. Squaring line B's distances gives a second candidate, 4 minutes

Concept

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\( {{5 + 3 + 1 + 3 + 6 = 18}},\quad\allowbreak \allowbreak {{18 \div 5 = 3.6}} \)

\[ (\textcolor{#1f5fbf}{-5})^2, (\textcolor{#1f5fbf}{-3})^2, (\textcolor{#1f5fbf}{-1})^2, \textcolor{#1f5fbf}{3}^2, \textcolor{#1f5fbf}{6}^2 = \textcolor{#b54708}{25,\ 9,\ 1,\ 9,\ 36} \]

Square each deviation

Why: no negatives survive

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#b54708}{25 + 9 + 1 + 9 + 36} = \textcolor{#b54708}{80} \]

Add the squares

Why: none can cancel either

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#b54708}{80} \div 5 = \textcolor{#b54708}{16} \]

Divide by 5

Why: a typical square

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \sqrt{\textcolor{#b54708}{16}} = \textcolor{#1f5fbf}{4} \]

Take the root

Why: so it compares with the waits

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#1f5fbf}{4}^2 \times 5 = \textcolor{#b54708}{80} \]

Check: undo both

Why: rebuilds the total area

26. Squared distances total least about the mean 5

Worked example

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\( {\text{about } 5\text{: } 25 + 9 + 1 + 9 + 36 = 80} \)

\[ \textcolor{#1f5fbf}{4,\ 2,\ 0,\ 4,\ 7;\ \ 6,\ 4,\ 2,\ 2,\ 5} \]

Measure from 4 and 6

Why: rival centres near 5

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ \textcolor{#b54708}{16,\ 4,\ 0,\ 16,\ 49;\ \ 36,\ 16,\ 4,\ 4,\ 25} \]

Square each distance

Why: as candidate 2 did

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ \textcolor{#b54708}{16 + 4 + 0 + 16 + 49} = \textcolor{#b54708}{85} \]

Add squares from 4

Why: rival total beside 80

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ \textcolor{#b54708}{36 + 16 + 4 + 4 + 25} = \textcolor{#b54708}{85} \]

Add squares from 6

Why: tests the centre above 5

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ (36 + 4) + (16 + 4) + 25 = \textcolor{#b54708}{85} \]

Check by regrouping

Why: addition ignores order

27. For line B, only squares bottom out at the mean

Worked example

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\( {4, 2, 0, 4, 7;\ 6, 4, 2, 2, 5},\quad\allowbreak \allowbreak {5 + 3 + 1 + 3 + 6 = 18},\quad\allowbreak \allowbreak {85,\ 80,\ 85} \)

\[ \textcolor{#1f5fbf}{4 + 2 + 0 + 4 + 7} = \textcolor{#1f5fbf}{17} \]

Add distances from 4

Why: the median's own total

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ \textcolor{#1f5fbf}{6 + 4 + 2 + 2 + 5} = \textcolor{#1f5fbf}{19} \]

Add distances from 6

Why: checks the centre above too

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ \textcolor{#b54708}{80} < \textcolor{#b54708}{85},\ \ \textcolor{#1f5fbf}{17} < \textcolor{#1f5fbf}{18} \]

Compare least totals

Why: squares pair with the mean

Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals

\[ 17 + 3 - 2 = \textcolor{#1f5fbf}{18} \]

Check: centre 4 → 5

Why: three gain 1, two lose 1

28. Squaring line A's deviations turns each into a square that cannot cancel

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ (\textcolor{#1f5fbf}{-3})^2 = \textcolor{#b54708}{9} \]

Square −3

Why: an area can't be negative

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ (\textcolor{#1f5fbf}{-1})^2 = \textcolor{#b54708}{1},\ \ \textcolor{#1f5fbf}{0}^2 = \textcolor{#b54708}{0} \]

Square −1 and 0

Why: small distances make small areas

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \textcolor{#1f5fbf}{1}^2 = \textcolor{#b54708}{1},\ \ \textcolor{#1f5fbf}{3}^2 = \textcolor{#b54708}{9} \]

Square +1 and +3

Why: opposite signs give equal areas

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9} = \textcolor{#b54708}{20} \]

Add the areas

Why: nothing cancels now

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ (9 + 9) + (1 + 1) + 0 = 20 \]

Check by pairing the squares

Why: addition ignores order

29. Line A's average square has area 4, so its side is 2 minutes

Worked example

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\( {{(-3)^2 = 9}},\quad\allowbreak \allowbreak {{(-1)^2 = 1}},\quad\allowbreak \allowbreak {{0^2 = 0}},\quad\allowbreak \allowbreak {{1^2 = 1}},\quad\allowbreak \allowbreak {{3^2 = 9}},\quad\allowbreak \allowbreak {{9 + 1 + 0 + 1 + 9 = 20}} \)

\[ \textcolor{#b54708}{20} \div 5 = \textcolor{#b54708}{4} \]

Share 20 among 5 squares

Why: a typical area, not a growing total

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \sqrt{\textcolor{#b54708}{4}} = \textcolor{#1f5fbf}{2} \]

Take its side

Why: back from area to minutes

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ 5 \times \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{20} \]

Check: five average squares

Why: rebuild the total area

30. The same moves give any population a spread

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{x - \mu} \]

Take one deviation

Why: each square's side

\[ \textcolor{#b54708}{(x - \mu)^2} \]

Square it

Why: so its sign cannot cancel

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} \]

Add all N areas

Why: every value counts

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \textcolor{#b54708}{\sigma^2} = \frac{\textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2}}{N} \]

Divide by N

Why: a typical square, not a total

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{\textcolor{#b54708}{\sigma^2}} \]

Take the root

Why: back to data units

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

σ² is the population variance; σ, its root, is the standard deviation (SD).

\[ \sqrt{20 \div 5} = \textcolor{#1f5fbf}{2} \]

Check with line A

Why: matches the last slide

31. One wait supplies nearly half of B's area

Concept

Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean

\[ \textcolor{#b54708}{25,\ 9,\ 1,\ 9,\ 36} \]

Square B's five deviations

Why: none can cancel

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#b54708}{25 + 9 + 1 + 9 + 36} = \textcolor{#b54708}{80} \]

Add B's areas

Why: σ² shares it among 5

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#b54708}{36} \div \textcolor{#b54708}{80} = 0.45 \]

Divide +6's square by the total

Why: part over whole gives its share

\[ 0.45 = 45\% \]

Write the share as a percent

Why: percents compare shares at a glance

\[ 0.45 \times \textcolor{#b54708}{80} = \textcolor{#b54708}{36} \]

Check: undo the division

Why: back to the far wait's square

32. Doubling every distance quadruples every square

Concept

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

Discussion prompt

Doubling every distance: does the side-1 square or the side-6 square grow more?

Answer:

Side 6: 36 → 144, up 108; side 1: 1 → 4.

\[ \textcolor{#1f5fbf}{d} = x - \mu \]

Name each distance d

Why: for any side

\[ (2\textcolor{#1f5fbf}{d})^2 = 2^2 \textcolor{#1f5fbf}{d}^2 \]

Square 2d

Why: each factor squares

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ = 4\textcolor{#b54708}{d^2} \]

Evaluate 2²

Why: four old squares fit inside

\[ (2 \times 6)^2 = 144 = 4 \times 36 \]

Check side 6

Why: the 11-minute wait

33. Doubling line B's distances takes its σ² to 64

Concept

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#1f5fbf}{-10,\ -6,\ -2,\ +6,\ +12} \]

Double B's distances

Why: sides of the new squares

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{100,\ 36,\ 4,\ 36,\ 144} \]

Square each

Why: σ² averages these areas

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{100 + 36 + 4 + 36 + 144} = \textcolor{#b54708}{320} \]

Add the squares

Why: σ² shares a total

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = 320 \div 5 = \textcolor{#b54708}{64} \]

Divide by N = 5

Why: σ² is the mean square

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{80} \div 5 = \textcolor{#b54708}{16} \]

Divide B's old total by 5

Why: to test the ×4 claim

\[ 4 \times \textcolor{#b54708}{16} = \textcolor{#b54708}{64} \]

Check: 4 × old σ²

Why: the ×4 claim holds

34. Doubling every distance quadruples σ²

Concept

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\,\textcolor{#b54708}{d^2} \]

Write σ² using d = x − μ

Why: the average square

Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36

\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = \text{avg}\,(2\textcolor{#1f5fbf}{d})^2 \]

Replace each d by 2d

Why: doubling data doubles every distance

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ = \text{avg}\,4\textcolor{#b54708}{d^2} \]

Use (2d)² = 4d²

Why: proved for any single distance earlier

\[ = 4\,\text{avg}\,\textcolor{#b54708}{d^2} \]

Pull out 4

Why: constants leave averages

\[ = 4\textcolor{#b54708}{\sigma^2} \]

Recognise σ²

Why: the old average square

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\[ \textcolor{#b54708}{64} = 4 \times \textcolor{#b54708}{16} \]

Check with line B's pictures

Why: average squares 64 and 16

35. So doubling every distance doubles σ

Concept

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

\( {\sigma_{\text{new}}^2 = 4\sigma^2},\quad\allowbreak {\text{line B: } \sigma^2 = 16,\ \sigma_{\text{new}}^2 = 64} \)

\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = 2^2\textcolor{#b54708}{\sigma^2} \]

Write 4 as 2²

Why: aiming for a square

\[ = (2\textcolor{#1f5fbf}{\sigma})^2 \]

Combine the squares

Why: so a root undoes it

\[ \textcolor{#1f5fbf}{\sigma_{\text{new}}} = 2\textcolor{#1f5fbf}{\sigma} \]

Root both sides

Why: roots undo squares; 2σ ≥ 0

\[ \sqrt{64} = 8 = 2\sqrt{16} \]

Check with line B

Why: the doubling rule holds there

Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144

36. No spread at all gives σ = 0

Concept

Figure (svg): Five waits all equal to 5 stacked on a number line from 0 to 10, every one at the mean

\[ \textcolor{#1f5fbf}{x - \mu} = 0 \text{ for every } x \]

Let every value equal μ

Why: spread should read zero

\[ \textcolor{#b54708}{(x - \mu)^2} = 0 \]

Square each deviation

Why: zero squared is zero

\[ {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} = 0 \]

Add the N squares

Why: zeros add to zero

\[ \textcolor{#b54708}{\sigma^2} = 0 \div N = 0 \]

Divide by N

Why: σ² is the mean square

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{0} = 0 \]

Take the root

Why: zero area, zero side

\[ (0+0+0+0+0) \div 5 = 0 \]

Check with five 5s

Why: real data agree

37. Shifting every wait by the same amount tests what σ measures

Prediction

Predict first

A slow cashier adds 3 minutes to every line A wait: 5, 7, 8, 9, 11.

Line A's σ was 2. What is it now?

  • 5 minutes
  • 2 minutes
  • 6 minutes
  • 11 minutes

Correct: 2 minutes

Why: The mean moves up 3 as well, so no deviation changes: same squares, same σ.

38. Adding 3 to every wait moves the mean to 8

Worked example

Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line, mean not yet found

\[ \textcolor{#1f5fbf}{5 + 7 + 8 + 9 + 11} = 40 \]

Add the new waits

Why: a total for the mean

\[ \mu = 40 \div 5 = 8 \]

Divide by N = 5

Why: a mean shares the total

Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 0 deviation arrows from it

\[ \textcolor{#1f5fbf}{-3,\ -1,\ 0,\ +1,\ +3} \]

Subtract 8 from each

Why: distances run from the new mean

Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 5 deviation arrows from it

\[ -3 - 1 + 0 + 1 + 3 = 0 \]

Check the balance

Why: deviations from a true mean total zero

39. The shifted waits still have σ = 2

Worked example

Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 5 deviation arrows from it

\( {5 + 7 + 8 + 9 + 11 = 40}\;\;\Rightarrow\;\;\allowbreak {\mu = 40 \div 5 = 8}\;\;\Rightarrow\;\;\allowbreak {-3,\ -1,\ 0,\ +1,\ +3} \)

\[ \textcolor{#b54708}{9,\ 1,\ 0,\ 1,\ 9} \]

Square each

Why: variance averages squares

Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9

\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9 = 20} \]

Add the areas

Why: σ² shares this total

Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9

\[ \textcolor{#b54708}{\sigma^2} = 20 \div 5 = \textcolor{#b54708}{4} \]

Divide by N = 5

Why: σ² is the mean square

Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4} = \textcolor{#1f5fbf}{2} \]

Take the root

Why: an SD in minutes, like the waits

Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9

\[ 5 \times \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{20} \]

Check: five average squares

Why: rebuild the shifted total

40. A family's four children, ages 4, 7, 9 and 12, have mean age 8

Worked example

Figure (svg): Ages 4, 7, 9, 12 on a number line with 0 deviation arrows from the mean 8

\[ \textcolor{#1f5fbf}{4 + 7 + 9 + 12} = 32 \]

Add the ages

Why: the mean needs the total

\[ \mu = 32 \div 4 = 8 \]

Divide by N = 4

Why: a mean is total over count

\[ \textcolor{#1f5fbf}{-4,\ -1,\ +1,\ +4} \]

Subtract 8 from each age

Why: spread starts from each age's distance

Figure (svg): Ages 4, 7, 9, 12 on a number line with 4 deviation arrows from the mean 8

\[ -4 - 1 + 1 + 4 = 0 \]

Check the deviations balance

Why: Idea 1 says they must

41. The family's ages spread about 2.92 years around 8

Worked example

Figure (svg): Ages 4, 7, 9, 12 on a number line with 4 deviation arrows from the mean 8

\[ \textcolor{#b54708}{16,\ 1,\ 1,\ 16} \]

Square −4, −1, +1, +4

Why: areas cannot cancel

Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16

\[ \textcolor{#b54708}{16 + 1 + 1 + 16 = 34} \]

Add the areas

Why: the total the four ages share

Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16

\[ \textcolor{#b54708}{\sigma^2} = 34 \div 4 = \textcolor{#b54708}{8.5} \]

Divide by N = 4

Why: σ² is the mean square

Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{8.5} \approx \textcolor{#1f5fbf}{2.92} \]

Take the root

Why: back to years

Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16

\[ 4 \times \textcolor{#1f5fbf}{2.92}^2 \approx \textcolor{#b54708}{34.1} \]

Check by undoing both

Why: restores the total 34

42. Squaring the total instead of totalling the squares brings the zero straight back

Trap

The trap

\[ (-4 - 1 + 1 + 4)^2 \]

Write the square of the sum

Why: tempting: it looks like squaring deviations

\[ = 0^2 \]

Add inside the bracket

Why: brackets come first, so signs cancel

\[ = 0 \]

Square the zero

Why: fails: no spread left

The fix

\[ (-4)^2, (-1)^2, 1^2, 4^2 = 16, 1, 1, 16 \]

Square each deviation first

Why: the bracket closes on each one

\[ 16 + 1 + 1 + 16 = 34 \]

Then add the squares

Why: areas cannot cancel

\[ 2 \times 16 + 2 \times 1 = 34 \]

Check by pairs

Why: ±4, ±1 square alike

43. Sets of three and five values share mean 5

Prediction

Figure (svg): Dot plots of set P (1, 5, 9) and set Q (3, 4, 5, 6, 7) on axes from 0 to 10, both with mean 5

Predict first

Set P: 1, 5, 9. Set Q: 3, 4, 5, 6, 7. Both have mean 5; both are whole populations.

Which set has the larger σ?

  • Set P
  • Set Q
  • Equal: same mean
  • Set Q: more values

Correct: Set P

Why: σ measures typical distance, not how many values. P's outer values sit 4 away; Q's sit at most 2 away.

44. Set P's values sit 4 below, at, and 4 above its mean 5

Worked example

Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 0 deviation arrows from it

\[ \textcolor{#1f5fbf}{1 + 5 + 9} = 15 \]

Add P's values

Why: the mean needs the total

\[ \mu = 15 \div 3 = 5 \]

Divide by N = 3

Why: the centre for the distances

\[ \textcolor{#1f5fbf}{-4,\ 0,\ +4} \]

Subtract 5 from 1, 5, 9

Why: spread is measured from the centre

Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 3 deviation arrows from it

\[ -4 + 0 + 4 = 0 \]

Check the deviations balance

Why: Idea 1 says they must

45. Set P's σ is about 3.27

Worked example

Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 3 deviation arrows from it

\( {\mu = 15 \div 3 = 5}\;\;\Rightarrow\;\;\allowbreak {-4,\ 0,\ +4} \)

\[ \textcolor{#b54708}{16,\ 0,\ 16} \]

Square each

Why: areas that cannot cancel

Figure (svg): Squares on set P's deviations −4, 0, +4

\[ \textcolor{#b54708}{16 + 0 + 16 = 32} \]

Add P's areas

Why: averaging needs a total to share

Figure (svg): Squares on set P's deviations −4, 0, +4

\[ \textcolor{#b54708}{\sigma_P^2} = 32 \div 3 \approx \textcolor{#b54708}{10.67} \]

Divide by N = 3

Why: P is a whole population

Figure (svg): Squares on set P's deviations −4, 0, +4

\[ \textcolor{#1f5fbf}{\sigma_P} = \sqrt{10.67} \approx \textcolor{#1f5fbf}{3.27} \]

Take the root

Why: back to the data's units

Figure (svg): Squares on set P's deviations −4, 0, +4

\[ 3 \times \textcolor{#1f5fbf}{3.27}^2 \approx \textcolor{#b54708}{32.1} \]

Check: undo both moves

Why: restores P's total area 32

46. Set Q's values sit within 2 of its mean 5

Worked example

Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 0 deviation arrows from it

\[ \textcolor{#1f5fbf}{3 + 4 + 5 + 6 + 7} = 25 \]

Add Q's values

Why: the mean needs the total

\[ \mu = 25 \div 5 = 5 \]

Divide by N = 5

Why: the centre for the distances

\[ \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ +1,\ +2} \]

Subtract 5 from 3 to 7

Why: distances show spread

Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 5 deviation arrows from it

\[ -2 - 1 + 0 + 1 + 2 = 0 \]

Check the deviations balance

Why: Idea 1 says they must

47. Set Q's σ is only 1.41, so set P spreads wider

Worked example

Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 5 deviation arrows from it

\( {\mu = 25 \div 5 = 5}\;\;\Rightarrow\;\;\allowbreak {-2,\ -1,\ 0,\ +1,\ +2} \)

\[ \textcolor{#b54708}{4,\ 1,\ 0,\ 1,\ 4} \]

Square each

Why: areas cannot cancel

Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2

\[ \textcolor{#b54708}{4 + 1 + 0 + 1 + 4 = 10} \]

Add Q's areas

Why: the total σ² shares out

Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2

\[ \textcolor{#b54708}{\sigma_Q^2} = 10 \div 5 = \textcolor{#b54708}{2} \]

Divide by N = 5

Why: Q is a whole population

Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2

\[ \textcolor{#1f5fbf}{\sigma_Q} = \sqrt{2} \approx \textcolor{#1f5fbf}{1.41} \]

Take the root

Why: side, not area

Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2

\[ \textcolor{#1f5fbf}{3.27} > \textcolor{#1f5fbf}{1.41} \]

Compare SDs

Why: SD tracks spread

\[ 5 \times \textcolor{#1f5fbf}{1.41}^2 \approx \textcolor{#b54708}{9.94} \]

Check: undo both moves

Why: rounding explains the gap

48. Divide by n − 1 when the data are a sample

Section

Idea 3 of 5

49. A tiny world of three waits tests whether ÷ n recovers σ²

Concept

Figure (svg): Waits 1, 3, 5 on a number line with 0 deviation arrows from the mean

Real data are samples. Test the recipe on a tiny world: waits 1, 3, 5.

Discussion prompt

Draw 2 waits (repeats allowed), square distances from x̄, ÷ n. Versus σ² on average: high, low, or right?

Answer:

Low on average. The next slides find the true σ² and test all nine samples.

50. The tiny world's waits sit −2, 0, +2 from μ

Worked example

Figure (svg): Waits 1, 3, 5 on a number line with 0 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{1} + \textcolor{#1f5fbf}{3} + \textcolor{#1f5fbf}{5} = 9 \]

Add the three waits

Why: μ needs the total first

\[ \mu = 9 \div 3 = 3 \]

Divide by N = 3

Why: a mean is total over count

\[ \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]

Subtract μ from 1, 3, 5

Why: σ² is built from these distances

Figure (svg): Waits 1, 3, 5 on a number line with 3 deviation arrows from the mean

\[ -2 + 0 + 2 = 0 \]

Check the distances balance

Why: so μ = 3 sits at the centre

51. The tiny world's true variance is σ² = 8/3 ≈ 2.67

Worked example

Figure (svg): Waits 1, 3, 5 on a number line with 3 deviation arrows from the mean

\( {1 + 3 + 5 = 9}\;\;\Rightarrow\;\;\allowbreak {\mu = 9 \div 3 = 3}\;\;\Rightarrow\;\;\allowbreak {-2,\ 0,\ +2} \)

\[ \textcolor{#b54708}{4,\ 0,\ 4} \]

Square each distance

Why: areas cannot cancel

Figure (svg): Squares on the distances −2, 0, +2 of the waits 1, 3, 5 from μ = 3

\[ 4 + 0 + 4 = \textcolor{#b54708}{8} \]

Add the areas

Why: σ² shares this total among N = 3

\[ \textcolor{#b54708}{\sigma^2} = 8 \div 3 \approx \textcolor{#b54708}{2.67} \]

Divide by N = 3

Why: every value is known

Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{2.67} \approx \textcolor{#1f5fbf}{1.63} \]

Take the root

Why: back to the waits' units

Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67, side 1.63

\[ 3 \times 2.67 \approx 8 \]

Check: undo the divide

Why: rebuilds the total area

52. Three kinds of sample sit 0, 1, 2 from x̄

Worked example

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5

\[ \textcolor{#1f5fbf}{3 + 3} = 6,\ \ \textcolor{#1f5fbf}{1 + 3} = 4,\ \ \textcolor{#1f5fbf}{1 + 5} = 6 \]

Add each sample's two draws

Why: each sample's mean needs its total

\[ \bar{x} = 6 \div 2,\ 4 \div 2,\ 6 \div 2 = 3,\ 2,\ 3 \]

Divide each by n = 2

Why: a mean is the total over n

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3

\[ \textcolor{#1f5fbf}{0,\,0;\ \ -1,\,+1;\ \ -2,\,+2} \]

Subtract x̄ from each draw

Why: a sample can't see μ, only x̄

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2

\[ 0 + 0 = 0,\ \ -1 + 1 = 0,\ \ -2 + 2 = 0 \]

Check the distances balance

Why: so each x̄ was right

53. Three kinds of sample give ÷ n estimates 0, 1, 4

Worked example

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2

\( {3 + 3 = 6,\ 1 + 3 = 4,\ 1 + 5 = 6}\;\;\Rightarrow\;\;\allowbreak {\bar{x} = 3,\ 2,\ 3}\;\;\Rightarrow\;\;\allowbreak {0,\,0;\ -1,\,+1;\ -2,\,+2} \)

\[ \textcolor{#b54708}{0,\,0;\ \ 1,\,1;\ \ 4,\,4} \]

Square each distance

Why: areas that cannot cancel

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; squares 0, 0; 1, 1; 4, 4

\[ 0 + 0 = 0,\ \ 1 + 1 = 2,\ \ 4 + 4 = 8 \]

Add each sample's areas

Why: ÷ n needs totals

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; square totals 0, 2, 8

\[ 0 \div 2,\ 2 \div 2,\ 8 \div 2 = \textcolor{#b54708}{0,\ 1,\ 4} \]

Divide each by n = 2

Why: finishing the σ² recipe

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; ÷ n estimates 0, 1, 4

\[ \textcolor{#1f5fbf}{0}^2,\ \ \textcolor{#1f5fbf}{1}^2,\ \ \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{0,\ 1,\ 4} \]

Check: one square per sample

Why: two equal squares, halved, leave one

54. Over nine samples, ÷ n averages only 1.33

Worked example

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5: only the three kinds' cells filled, 0, 1, 4

\[ (3,1) \to 1,\ \ (5,1) \to 4 \]

Swap unequal pairs

Why: order keeps distances

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5: cells 1, 3; 3, 3; 1, 5 and their swaps filled

\[ (3,5),\ (5,3) \to 1;\ \ (1,1),\ (5,5) \to 0 \]

Shift both draws by ±2

Why: x̄ moves with them

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0

\[ 0 + 1 + 4 + 1 + 0 + 1 + 4 + 1 + 0 = \textcolor{#b54708}{12} \]

Add the nine cells

Why: covers every sample once

\[ \textcolor{#b54708}{12} \div 9 \approx \textcolor{#b54708}{1.33} \]

Divide by 9

Why: every sample equally likely, so a plain average

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0

\[ 3(0) + 4(1) + 2(4) = \textcolor{#b54708}{12} \]

Check by kind

Why: three 0s, four 1s, two 4s

55. A different divisor would make the nine estimates right on average

Prediction

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0

Predict first

Samples of n = 2: ÷ n averaged 1.33; truth 2.67.

Which divisor averages 2.67?

  • n + 1 = 3
  • n² = 4
  • n − 1 = 1
  • No divisor can

Correct: n − 1 = 1

Why: Halving the divisor doubles every estimate: 12/9 doubled is 24/9 = 8/3 ≈ 2.67. With n = 2, dividing by n − 1 = 1 is that halving.

56. ÷ (n − 1) turns the three kinds into 0, 2, 8

Worked example

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; square totals 0, 2, 8

\( {0 + 0 = 0},\ \allowbreak {1 + 1 = 2},\ \allowbreak {4 + 4 = 8} \)

\[ 2 - 1 = 1 \]

Find n − 1 for samples of two

Why: the predicted divisor

\[ 0 \div 1,\ 2 \div 1,\ 8 \div 1 = \textcolor{#b54708}{0,\ 2,\ 8} \]

Divide each total by n − 1

Why: tests the new divisor per kind

Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; ÷ (n − 1) estimates 0, 2, 8

\[ 2 \times \textcolor{#b54708}{0},\ \ 2 \times \textcolor{#b54708}{1},\ \ 2 \times \textcolor{#b54708}{4} = \textcolor{#b54708}{0,\ 2,\ 8} \]

Check against ÷ n's 0, 1, 4

Why: halving the divisor doubles each

57. ÷ (n − 1) makes the nine estimates average 8/3

Worked example

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, the divide-by-(n − 1) cells not yet filled

\( {2 - 1 = 1}\;\;\Rightarrow\;\;\allowbreak {0 \div 1,\ 2 \div 1,\ 8 \div 1 = 0,\ 2,\ 8} \)

\[ \textcolor{#b54708}{0,\ 2,\ 8,\ 2,\ 0,\ 2,\ 8,\ 2,\ 0} \]

Place each kind's estimate

Why: swaps and shifts keep kinds

Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0

\[ 0 + 2 + 8 + 2 + 0 + 2 + 8 + 2 + 0 = \textcolor{#b54708}{24} \]

Add the nine cells

Why: covers every sample once

\[ \textcolor{#b54708}{24} \div 9 = 8/3 \]

Divide by 9

Why: each of the nine samples is equally likely

Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67

\[ 24 = 2 \times 12 \]

Check against ÷ n

Why: halving the divisor doubled the total 12

58. The proof ahead shows why squares from x̄ need n − 1

Concept

Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67

avg: over all equally likely samples

\[ \text{avg}{\textstyle\sum} (x - \bar{x})^2 = (n - 1)\textcolor{#b54708}{\sigma^2} \]

Goal

Why: what the three phases prove

Assumed: independent draws

  1. Split each distance at x̄
  2. Average the offset area over every sample
  3. Correct the divisor

59. Splitting at x̄ cuts a square into four pieces

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked

Name r = x − x̄ and the offset g = x̄ − μ.

\[ \textcolor{#1f5fbf}{x - \mu} = (x - \textcolor{#6b7280}{\bar{x}}) + (\textcolor{#6b7280}{\bar{x}} - \mu) \]

Add and subtract x̄

Why: adding zero exposes x̄

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distance x − μ for x = 1 split at x̄ into r = −1 and the offset g = −1

\[ \textcolor{#1f5fbf}{x - \mu} = \textcolor{#1f5fbf}{r} + \textcolor{#1f5fbf}{g} \]

Substitute the names r and g

Why: shorter, same values

\[ \textcolor{#b54708}{(x - \mu)^2} = (\textcolor{#1f5fbf}{r} + \textcolor{#1f5fbf}{g})^2 \]

Square both sides

Why: turns distances into areas

\[ = \textcolor{#b54708}{r^2 + 2rg + g^2} \]

Expand

Why: so each piece sums separately

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distance x − μ for x = 1 split at x̄ into r = −1 and the offset g = −1; a square of side r + g cut into r², rg, rg, g², four equal pieces since r = g = −1

\[ (1 - 3)^2 = (-1)^2 + 2(-1)(-1) + (-1)^2 \]

Check x = 1 in sample 1, 3

Why: r and g both −1: 4 each side

60. Over a sample, the cross term adds to zero

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked

\( {x - \mu = (x - \bar{x}) + (\bar{x} - \mu) = r + g}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 = (r + g)^2 = r^2 + 2rg + g^2} \)

\[ {\textstyle\sum} 2rg = 2\textcolor{#1f5fbf}{g}{\textstyle\sum} \textcolor{#1f5fbf}{r} \]

Pull 2g out of the sum

Why: g is the same for every x

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1, highlighted

\[ = 2\textcolor{#1f5fbf}{g} \cdot 0 \]

Use Σr = 0

Why: x̄ balances the r values

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2: the offset arrow g = −1, and each value's distance r from x̄, −1 and +1, which total 0

\[ = 0 \]

Multiply by 0

Why: any number times 0 is 0

\[ 2(-1)[(1 - 2) + (3 - 2)] = 0 \]

Check sample 1, 3

Why: g = −1; its r values cancel

61. Squares from μ = squares from x̄ + offset area

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked

\( {(x - \mu)^2 = r^2 + 2rg + g^2}\qquad\allowbreak {{\textstyle\sum} 2rg = 0} \)

\[ \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} = {\textstyle\sum} (r^2 + 2rg + g^2) \]

Sum over the sample

Why: totals matter here

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4

\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + {\textstyle\sum} 2\textcolor{#1f5fbf}{r}\textcolor{#1f5fbf}{g} + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]

Split the sum

Why: isolates each piece

\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + 0 + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]

Use Σ2rg = 0

Why: proved on slide 60

\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]

Drop the 0

Why: 0 adds nothing

\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \]

Count g² terms

Why: n equal copies

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2

\[ (-2)^2 + 0^2 = (-1)^2 + 1^2 + 2(-1)^2 \]

Check sample 1, 3

Why: both sides give 4

62. Each sample of two has its own offset g

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells the offset area n·g² not yet filled

\[ 2,\ 4,\ 6,\ 4,\ 6,\ 8,\ 6,\ 8,\ 10 \]

Add each cell's two draws

Why: each mean needs its total

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the total of its two draws: 2, 4, 6, 4, 6, 8, 6, 8, 10

\[ \bar{x} = 1,\ 2,\ 3,\ 2,\ 3,\ 4,\ 3,\ 4,\ 5 \]

Divide each total by n = 2

Why: a mean is the total over n

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its mean x̄: 1, 2, 3, 2, 3, 4, 3, 4, 5

\[ \textcolor{#1f5fbf}{g} = \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ -1,\ 0,\ 1,\ 0,\ 1,\ 2} \]

Subtract μ = 3 from each x̄

Why: how far each mean misses μ

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its offset g = x̄ − μ: −2, −1, 0, −1, 0, 1, 0, 1, 2

\[ 1 - (-2),\ 2 - (-1),\ 3 - 0 = 3,\ 3,\ 3 \]

Check: undo the subtraction

Why: the first row returns μ = 3

63. Each sample of two carries offset area 2g²

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its offset g = x̄ − μ: −2, −1, 0, −1, 0, 1, 0, 1, 2

\[ \textcolor{#1f5fbf}{g} = \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ -1,\ 0,\ 1,\ 0,\ 1,\ 2} \]

Start from the offsets

Why: their squares build n·g²

\[ \textcolor{#1f5fbf}{g}^2 = \textcolor{#b54708}{4,\ 1,\ 0,\ 1,\ 0,\ 1,\ 0,\ 1,\ 4} \]

Square each offset

Why: n·g² is built from these

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its squared offset g²: 4, 1, 0, 1, 0, 1, 0, 1, 4

\[ 2\textcolor{#1f5fbf}{g}^2 = \textcolor{#b54708}{8,\ 2,\ 0,\ 2,\ 0,\ 2,\ 0,\ 2,\ 8} \]

Multiply each by n = 2

Why: n copies make the offset area

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8

\[ (4 + 0) - 2 = \textcolor{#b54708}{2},\ \ (4 + 4) - 8 = \textcolor{#b54708}{0} \]

Check 1, 3 and 1, 5

Why: squares from μ less those from x̄

64. The offset area is a square, which settles whether x̄ can ever look more spread

Prediction

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2

Predict first

Can a sample's squares from x̄ ever total MORE than its squares from μ?

  • Yes, when x̄ is far from μ
  • No, never
  • Only for samples of two

Correct: No, never

Why: Squares from μ are squares from x̄ plus n·g², and a square is never negative. At best (x̄ = μ) they tie.

65. Squares from x̄ never exceed squares from μ

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from x̄, 1 and 1, total 2

\( {{\textstyle\sum} (x - \mu)^2 = {\textstyle\sum} r^2 + n g^2} \)

\[ \textcolor{#1f5fbf}{g}^2 \ge 0 \]

Square the offset

Why: squares of real numbers ignore sign

\[ \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \ge 0 \]

Multiply by n

Why: positive counts keep ≥

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2

\[ \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \le \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \]

Add n·g² ≥ 0 to Σr²

Why: a total can only grow

\[ \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \le \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} \]

Substitute the split

Why: Σr² + n·g² equals squares from μ

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2

\[ 2 \le 4 \]

Check sample 1, 3

Why: picture totals: x̄ 2, μ 4

66. A squared sum is every cell of a product grid

Worked example

Figure (svg): A 2 by 2 grid for (e₁ + e₂)²: squares on the diagonal, cross products off it

Draw i has distance eᵢ = xᵢ − μ, for i = 1 to n.

\[ (\textcolor{#1f5fbf}{e_1} + \textcolor{#1f5fbf}{e_2} + \textcolor{#1f5fbf}{e_3})^2 = \textcolor{#b54708}{e_1e_1} + e_1e_2 + \dots + \textcolor{#b54708}{e_3e_3} \]

Multiply out three draws

Why: one term per grid cell

Figure (svg): A 3 by 3 grid for (e₁ + e₂ + e₃)²: squares on the diagonal, cross products off it

\[ ({\textstyle\sum} \textcolor{#1f5fbf}{e})^2 = {\textstyle\sum}_i {\textstyle\sum}_j \textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j} \]

Write every cell for n draws

Why: row i times column j

Figure (svg): An n by n grid for (Σe)², margins e₁, e₂, ⋯, eₙ: every cell the product eᵢeⱼ of its row and column

\[ = \textcolor{#b54708}{{\textstyle\sum} e_i^2} + {\textstyle\sum}_{i \ne j} e_ie_j \]

Split off the diagonal

Why: cells with i = j are squares

Figure (svg): An n by n grid for (Σe)², margins e₁, e₂, ⋯, eₙ: the diagonal cells marked as the squares eᵢ², the off-diagonal cells the cross products eᵢeⱼ

\[ (-2 + 0)^2 = (-2)^2 + 0^2 + 2(-2)(0) \]

Check sample 1, 3

Why: both sides give 4

Figure (svg): A 2 by 2 grid for sample 1, 3's distances −2 and 0: squares 4 and 0 on the diagonal, cross products 0 and 0 off it

67. Each draw's square averages σ²

Worked example

Figure (svg): Squares on the population 1, 3, 5's distances −2, 0, +2 from μ = 3

Uniform, independent draws. Σ all: over all N values.

\[ \text{avg}\ \textcolor{#b54708}{\textcolor{#1f5fbf}{e_i}^2} = \frac{\textcolor{#b54708}{{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e}^2}}{N} \]

Average one draw's square

Why: all N values equally likely

\[ = \textcolor{#b54708}{\sigma^2} \]

Recognise σ²'s definition

Why: the population's average square

Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67

\[ (4 + 0 + 4) \div 3 = \tfrac{8}{3} \approx \textcolor{#b54708}{2.67} \]

Check with the three values

Why: matches the grids' σ²

68. Cross products of independent draws average 0

Worked example

Figure (svg): Grid of every pair of distances −2, 0, +2, each cell their product

\[ \text{avg}\ \textcolor{#1f5fbf}{e_1}\textcolor{#1f5fbf}{e_2} = \frac{{\textstyle\sum}_{\text{all}}{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1}\textcolor{#1f5fbf}{e_2}}{N^2} \]

Average all N² pairs

Why: independent: pairs equally likely

\[ = \frac{{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1} \left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_2}\right)}{N^2} \]

Factor e₁ out of the inner sum

Why: fixed while e₂ varies

\[ = \frac{\left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1}\right)\left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_2}\right)}{N^2} \]

Pull the inner sum out

Why: it is the same for every e₁

Figure (svg): Grid of every pair of distances −2, 0, +2: cells 4, 0, −4, 0, 0, 0, −4, 0, 4; every row and column totals 0

\[ = \frac{\textcolor{#1f5fbf}{0} \cdot \textcolor{#1f5fbf}{0}}{N^2} \]

Substitute Σ all e = 0

Why: distances from μ balance

\[ = 0 \]

Evaluate

Why: zero over a positive count

\[ 4 + 0 - 4 + 0 + 0 + 0 - 4 + 0 + 4 = 0 \]

Check the grid

Why: its nine cells cancel

69. Samples of two test how many σ² the squared sum (Σe)² averages

Prediction

Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells (e₁ + e₂)² not yet filled

Predict first

Squares average σ²; independent cross products average 0.

For n = 2, is avg (Σe)² σ², 2σ² or 4σ²?

  • σ²
  • 2σ²
  • 4σ²

Correct: 2σ²

Why: Each of the 2 squares averages σ²; the cross products average 0: 2σ². So σ² counts only one square, and 4σ² would need the cross products to add another 2σ².

70. Each sample of two gets its squared sum (Σe)²

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells (e₁ + e₂)² not yet filled

\[ 1 - 3,\ 3 - 3,\ 5 - 3 = \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]

Subtract μ = 3

Why: the next result uses μ

\[ \textcolor{#1f5fbf}{-4,\ -2,\ 0,\ -2,\ 0,\ 2,\ 0,\ 2,\ 4} \]

Add e₁ + e₂ per cell

Why: (Σe)² sums first

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the sum of distances e₁ + e₂: −4, −2, 0, −2, 0, 2, 0, 2, 4

\[ \textcolor{#b54708}{16,\ 4,\ 0,\ 4,\ 0,\ 4,\ 0,\ 4,\ 16} \]

Square each sum

Why: the grid averages these

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16

\[ (3 - 3 + 5 - 3)^2 = \textcolor{#b54708}{4} \]

Check sample 3, 5 from its waits

Why: its cell reads 4

71. The nine samples of two give avg (Σe)² = 16/3

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16

\( {1 - 3,\ 3 - 3,\ 5 - 3 = -2,\ 0,\ +2}\;\;\Rightarrow\;\;\allowbreak {-4,\ -2,\ 0,\ -2,\ 0,\ 2,\ 0,\ 2,\ 4}\;\;\Rightarrow\;\;\allowbreak {16,\ 4,\ 0,\ 4,\ 0,\ 4,\ 0,\ 4,\ 16} \)

\[ \textcolor{#b54708}{16 + 4 + 0 + 4 + 0 + 4 + 0 + 4 + 16} = \textcolor{#b54708}{48} \]

Add the nine cells

Why: covers every sample once

\[ \textcolor{#b54708}{48} \div 9 = \textcolor{#b54708}{16/3} \]

Divide by 9

Why: nine equal chances

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 16/3, beside a bar for σ² = 2.67

\[ \textcolor{#b54708}{16/3} \div \textcolor{#b54708}{8/3} = 2 \]

Divide by σ² = 8/3

Why: settles the prediction: 2σ²

\[ 2(16) + 4(4) = \textcolor{#b54708}{48} \]

Check by kind

Why: two 16s, four 4s; 0s add nothing

72. Averaging (Σe)² averages every cell of its grid

Worked example

Figure (svg): A 3 by 3 grid for (e₁ + e₂ + e₃)²: squares on the diagonal, cross products off it

\[ \text{avg}\textcolor{#b54708}{({\textstyle\sum} \textcolor{#1f5fbf}{e})^2} = \text{avg}(\textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{e_i}^2} + {\textstyle\sum}_{i \ne j} \textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j}) \]

Average the grid expansion

Why: equal things average equally

\[ = {\textstyle\sum} \text{avg}\,\textcolor{#b54708}{\textcolor{#1f5fbf}{e_i}^2} + {\textstyle\sum}_{i \ne j} \text{avg}\,\textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j} \]

Average cell by cell

Why: averages of sums add

\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} \text{avg}\,e_ie_j \]

Substitute avg eᵢ² = σ²

Why: from the one-draw lemma

\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]

Substitute avg eᵢeⱼ = 0

Why: independent draws' lemma

Figure (svg): A 3 by 3 grid of cell averages: σ² on each diagonal cell, 0 on each off-diagonal cell

\[ 48 \div 9 = 16/3 = 2 \times 8/3 \]

Check with n = 2

Why: agrees with the grid

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 5.33 = 2σ²

73. On average, the squared sum (Σe)² is nσ²

Worked example

Figure (svg): A 3 by 3 grid of cell averages: σ² on each diagonal cell, 0 on each off-diagonal cell

\[ \text{avg}({\textstyle\sum} e)^2 = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]

Start from the cell averages

Why: the averaged-grid result

\[ = n\textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]

Add n diagonal σ²

Why: n equal terms

\[ = n\textcolor{#b54708}{\sigma^2} + (n^2 - n) \cdot 0 \]

Count cross cells

Why: n² cells less n squares

Figure (svg): A 3 by 3 grid of cell averages with the six off-diagonal cells outlined: 3² − 3 = 6 cross cells, each averaging 0

\[ = n\textcolor{#b54708}{\sigma^2} + 0 \]

Multiply by 0

Why: anything times 0 is 0

\[ = n\textcolor{#b54708}{\sigma^2} \]

Drop the 0

Why: 0 adds nothing

\[ 2 \times \textcolor{#b54708}{8/3} \approx \textcolor{#b54708}{5.33} \]

Check n = 2, σ² = 8/3

Why: the (Σe)² grid's average bar

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 5.33 = 2σ²

74. The offset g equals (Σx − nμ) ÷ n

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked

\[ \textcolor{#1f5fbf}{g} = \textcolor{#6b7280}{\bar{x}} - \textcolor{#6b7280}{\mu} \]

Start from the offset's definition

Why: to rewrite g using the data's total

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1, highlighted

\[ = \textcolor{#6b7280}{\frac{{\textstyle\sum} x}{n}} - \textcolor{#6b7280}{\mu} \]

Use x̄'s definition

Why: the prerequisite mean

\[ = \textcolor{#6b7280}{\frac{{\textstyle\sum} x}{n}} - \textcolor{#6b7280}{\frac{n\mu}{n}} \]

Write μ as nμ/n

Why: to share a denominator

\[ = \textcolor{#1f5fbf}{\frac{{\textstyle\sum} x - n\mu}{n}} \]

Combine the fractions

Why: same denominator n

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1

\[ \frac{4 - 2(3)}{2} = -1 = 2 - 3 \]

Check with sample 1, 3

Why: waits 1 and 3 total 4

75. The offset g is the sample's average distance from μ

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1

\( {g = \bar{x} - \mu}\;=\;\allowbreak {{\textstyle\sum} x/n - \mu}\;=\;\allowbreak {{\textstyle\sum} x/n - n\mu/n}\;=\;\allowbreak {({\textstyle\sum} x - n\mu)/n} \)

\[ \textcolor{#1f5fbf}{g} = \frac{{\textstyle\sum} x - {\textstyle\sum} \textcolor{#6b7280}{\mu}}{n} \]

Write nμ as Σμ

Why: a constant summed n times

\[ = \frac{{\textstyle\sum} \textcolor{#1f5fbf}{(x - \mu)}}{n} \]

Combine the two sums

Why: sums of differences rule

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1

\[ = \frac{{\textstyle\sum} \textcolor{#1f5fbf}{e}}{n} \]

Write x − μ as e

Why: the form the grid averages use

\[ \frac{-2 + 0}{2} = -1 \]

Check with sample 1, 3

Why: the blue arrows average −1

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, highlighted

76. The offset area n·g² equals (Σe)² ÷ n

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1

\( {g = {\textstyle\sum} e / n} \)

\[ \textcolor{#1f5fbf}{g}^2 = ({\textstyle\sum} \textcolor{#1f5fbf}{e} / n)^2 \]

Square g = Σe/n

Why: heading for the offset area

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, highlighted

\[ = ({\textstyle\sum} e)^2 / n^2 \]

Square top and bottom

Why: the quotient rule

\[ \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = n({\textstyle\sum} e)^2 / n^2 \]

Multiply by n

Why: builds the offset area

Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, with the offset area of two g² squares

\[ = \textcolor{#b54708}{({\textstyle\sum} e)^2} / n \]

Cancel one n

Why: n over n² is 1 over n

\[ 2(-1)^2 = (-2 + 0)^2 / 2 \]

Check with sample 1, 3

Why: both sides give 2

77. On average, the offset area is exactly one σ²

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8

\( {n\,g^2 = ({\textstyle\sum} e)^2 / n}\qquad\allowbreak {\text{avg}({\textstyle\sum} e)^2 = n\sigma^2} \)

\[ \text{avg}\ \textcolor{#b54708}{n\,g^2} = \text{avg}\,[({\textstyle\sum} e)^2 / n] \]

Average both sides

Why: equal quantities stay equal

\[ = \text{avg}({\textstyle\sum} e)^2 / n \]

Move 1/n outside

Why: constants leave averages

\[ = n\textcolor{#b54708}{\sigma^2} / n \]

Substitute avg (Σe)² = nσ²

Why: the squared-sum lemma

\[ = \textcolor{#b54708}{\sigma^2} \]

Cancel n

Why: n copies undo ÷ n

\[ [2(8) + 4(2) + 3(0)] \div 9 = 24 \div 9 = 8/3 \]

Check with the grid

Why: offset areas average σ²

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8, averaging 2.67 = σ²

78. Drawing without replacement tests whether independence really matters

Prediction

Predict first

Draw 2 different waits from 1, 3, 5: no repeats allowed.

Do the ÷ (n − 1) estimates still average σ² ≈ 2.67?

  • Yes, n − 1 always works
  • No, they average more
  • No, they average less

Correct: No, they average more

Why: Without repeats the draws are not independent, so the cross products no longer average 0. The three estimates are 2, 8 and 2: 2 + 8 + 2 = 12, and 12 ÷ 3 = 4.

79. Without independence, ÷ (n − 1) overshoots to 4

Worked example

Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0

No repeats leaves three pairs: 1, 3; 3, 5; 1, 5.

Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0; the diagonal repeats 1, 1; 3, 3; 5, 5 greyed and struck out, leaving six cells 2, 8, 2, 2, 8, 2

\[ \textcolor{#b54708}{2,\ 2,\ 8} \]

Take each pair's estimate

Why: to see dependence's effect

\[ 2 + 2 + 8 = 12 \]

Add the three estimates

Why: an average starts from the total

\[ 12 \div 3 = \textcolor{#b54708}{4} \]

Divide by 3

Why: three equally likely pairs

Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0; the diagonal repeats 1, 1; 3, 3; 5, 5 greyed and struck out, leaving six cells 2, 8, 2, 2, 8, 2; a bar for the six cells' average, 4, beside a bar for σ² = 2.67

\[ \textcolor{#b54708}{4} > 8/3 \approx \textcolor{#b54708}{2.67} \]

Compare with σ²

Why: the independence hypothesis mattered

\[ [4(2) + 2(8)] \div 6 = 4 \]

Check with ordered pairs

Why: each pair counted both ways

80. The nine samples of two give avg Σeᵢ² = 16/3

Worked example

Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells e₁² + e₂² not yet filled

\[ 1 - 3,\ 3 - 3,\ 5 - 3 = \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]

Subtract μ = 3

Why: the lemma works from μ

\[ \textcolor{#b54708}{4,\ 0,\ 4} \]

Square each

Why: a cell adds two of them

\[ \textcolor{#b54708}{8,\ 4,\ 8,\ 4,\ 0,\ 4,\ 8,\ 4,\ 8} \]

Add e₁² + e₂²

Why: square first: no cross products

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8

\[ 8+4+8+4+0+4+8+4+8 = \textcolor{#b54708}{48} \]

Add the nine cells

Why: covers every sample once

\[ \textcolor{#b54708}{48} \div 9 = \textcolor{#b54708}{16/3} \]

Divide by 9

Why: nine equal chances

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8, averaging 5.33 = 2σ²

\[ 4(8) + 4(4) = \textcolor{#b54708}{48} \]

Check by kind

Why: four 8s, four 4s, one 0

81. Squares from μ average nσ² over a sample

Worked example

Figure (svg): Bar of the nine samples' average squares from μ, 5.33

\[ \text{avg}{\textstyle\sum} \textcolor{#b54708}{e_i^2} = {\textstyle\sum} \text{avg}\,\textcolor{#b54708}{e_i^2} \]

Average each square

Why: averages of sums add

\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} \]

Substitute avg eᵢ² = σ²

Why: the one-draw lemma

Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67

\[ = n\textcolor{#b54708}{\sigma^2} \]

Add n copies of σ²

Why: summing a constant multiplies it

\[ \tfrac{16}{3} = 2 \times \tfrac{8}{3} \]

Check on the e₁² + e₂² grid

Why: n = 2: its cells average 16/3

Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8, averaging 5.33 = 2σ²

82. Squares from x̄ are squares from μ less the offset area

Worked example

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2

\( {{\textstyle\sum} (x - \mu)^2 = {\textstyle\sum} r^2 + n g^2} \)

\[ {\textstyle\sum} (x - \mu)^2 - \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \]

Subtract n·g² from the split

Why: leaves only squares from x̄

Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2

\[ {\textstyle\sum} \textcolor{#b54708}{e_i^2} - \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = {\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Rename x − μ as eᵢ

Why: matches the one-draw lemma

\[ 4 - 2 = (1 - 2)^2 + (3 - 2)^2 = 2 \]

Check with sample 1, 3

Why: matches its squares from x̄

83. Squares from x̄ average nσ² minus avg n·g²

Worked example

Figure (svg): Bar of the nine samples' average squares from μ, 5.33

\( {{\textstyle\sum} (x - \mu)^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {{\textstyle\sum} e_i^2 - n g^2 = {\textstyle\sum} r^2} \)

\[ \text{avg}({\textstyle\sum} e_i^2 - n\textcolor{#1f5fbf}{g}^2) = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Average both sides

Why: equal things stay equal

Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67

\[ \text{avg}{\textstyle\sum} e_i^2 - \text{avg}\,n\textcolor{#1f5fbf}{g}^2 = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Split the average

Why: means of differences subtract

\[ n\textcolor{#b54708}{\sigma^2} - \text{avg}\,n\textcolor{#1f5fbf}{g}^2 = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Substitute avg Σeᵢ² = nσ²

Why: a proved lemma

\[ \tfrac{16}{3} - \tfrac{8}{3} = \tfrac{8}{3} = \tfrac{24}{9} \]

Check n = 2

Why: from the grids; x̄-squares average 24/9

84. On average, squares from x̄ total (n − 1)σ²

Worked example

Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67

\( {{\textstyle\sum} (x - \mu)^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {{\textstyle\sum} e_i^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {\text{avg}({\textstyle\sum} e_i^2 - n g^2) = \text{avg}{\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {\text{avg}{\textstyle\sum} e_i^2 - \text{avg}\,n g^2 = \text{avg}{\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {n\sigma^2 - \text{avg}\,n g^2 = \text{avg}{\textstyle\sum} r^2} \)

\[ n\textcolor{#b54708}{\sigma^2} - \textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Substitute avg n·g² = σ²

Why: the offset-area lemma

Figure (svg): Average squares from μ, 5.33, as two σ² blocks; under the second, the dashed offset area of σ² = 2.67

\[ (n - 1)\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]

Factor out σ²

Why: n copies less one

Figure (svg): Average squares from μ, 5.33, as two σ² blocks; squares from x̄ keep one block, 2.67, beside the offset area

\[ (2 - 1) \times \frac{8}{3} = \frac{24}{9} \]

Check nine samples

Why: x̄-squares total 24

85. ÷ (n − 1) makes squares from x̄ average σ²

Worked example

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0

\( {(n - 1)\sigma^2 = \text{avg}{\textstyle\sum} r^2} \)

\[ (n - 1)\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} (x - \bar{x})^2 \]

Write r as x − x̄

Why: x − x̄ needs only the sample

\[ \textcolor{#b54708}{\sigma^2} = \frac{\text{avg}{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]

Divide by n − 1

Why: positive once n ≥ 2

\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\ \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]

Move 1/(n − 1) inside

Why: constants leave averages

Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67

\[ [3(0) + 4(\textcolor{#b54708}{2}) + 2(\textcolor{#b54708}{8})] \div 9 \approx \textcolor{#b54708}{2.67} \]

Check n = 2 on the corrected grid

Why: three 0s, four 2s, two 8s

86. So s², which divides by n − 1, is right on average

Worked example

Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0

\( {(n - 1)\sigma^2 = \text{avg}{\textstyle\sum} (x - \bar{x})^2}\;\;\Rightarrow\;\;\allowbreak {\sigma^2 = \text{avg}{\textstyle\sum} (x - \bar{x})^2/(n - 1)}\;\;\Rightarrow\;\;\allowbreak {\sigma^2 = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/(n - 1)]} \)

\[ \textcolor{#b54708}{s^2} = \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]

Name the sample variance s²

Why: a data-only estimate of σ²

Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67

\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\ \textcolor{#b54708}{s^2} \]

Substitute s²

Why: independent draws, n ≥ 2

\[ \textcolor{#1f5fbf}{s} = \sqrt{\textcolor{#b54708}{s^2}} \]

Define s, its root: the sample SD

Why: the side, in data units

\[ \text{avg}\ \textcolor{#b54708}{s^2} = \textcolor{#b54708}{24} \div 9 \approx \textcolor{#b54708}{2.67} \]

Check the grid

Why: nine cells total 24

87. The n − 1 correction fades as n grows

Concept

Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1

Discussion prompt

s² is how many times ÷ n, at n = 2 and 20?

Answer:

About twice as big at n = 2; about 5% bigger at n = 20.

\[ \textcolor{#b54708}{s^2} = \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} = \frac{n}{n} \cdot \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]

Multiply by n/n

Why: a factor of 1 changes nothing

\[ = \frac{n}{n - 1} \cdot \frac{{\textstyle\sum} (x - \bar{x})^2}{n} \]

Swap the denominators

Why: factor order is free

\[ \tfrac{2}{1} \times 4 = \textcolor{#b54708}{8} \]

Check sample 1, 5

Why: the ÷ n value doubles

Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2, circled and labelled

88. n ÷ (n − 1) is 2 at n = 2 but 1.05 at n = 20

Worked example

Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1

df = n − 1: Σ(x − x̄) = 0 fixes one distance

n = 1: x = x̄, so s² = 0 ÷ 0, undefined

\[ 2 - 1 = 1,\ \ 20 - 1 = 19 \]

Find both corrected divisors

Why: the ratio needs both

\[ 2 \div 1 = 2 \]

Divide n by n − 1

Why: shows how much s² grows

Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2, circled and labelled

\[ 20 \div 19 \approx 1.05 \]

Repeat at n = 20

Why: tests a bigger sample

Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2 and the n = 20 point, 1.05, circled and labelled

\[ 19 \times 1.05 \approx 20 \]

Check: undo the divide

Why: returns the sample size 20

89. Twenty ages but only six different values: multiplying by f saves work

Concept

Figure (svg): Example 2.32's twenty fifth graders' ages as dots stacked by value on a number line: 9 (1 dot), 9.5 (2), 10 (4), 10.5 (4), 11 (6), 11.5 (3), each stack's count printed above it

Example 2.32: twenty fifth graders' ages, a sample, with only six different values.

Discussion prompt

The table lists each age once with its count f. How can one square stand in for several pupils?

Answer:

Multiply that age's square by f: f equal squares.

90. The twenty ages average 10.525 years

Worked example

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3)

f = how many of the 20 pupils have that age.

\[ \begin{aligned}&1(9),\ 2(9.5),\ 4(10),\ 4(10.5), \\ &6(11),\ 3(11.5) = 9,\ 19,\ 40,\ 42,\ 66,\ 34.5\end{aligned} \]

Multiply each age by f

Why: f pupils share that age

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), products fx (9, 19, 40, 42, 66, 34.5)

\[ 9 + 19 + 40 + 42 + 66 + 34.5 = 210.5 \]

Add the six products

Why: the mean needs all 20

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), products fx (9, 19, 40, 42, 66, 34.5); totals: f 20, fx 210.5

\[ \bar{x} = 210.5 \div 20 = 10.525 \]

Divide by n = 20

Why: a mean is total over count

\[ 20 \times 10.525 = 210.5 \]

Check: undo the division

Why: restores the ages' total

91. Each fifth grader's age sits a distance from x̄

Worked example

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3)

\[ \bar{x} = 10.525 \]

Start from the computed mean

Why: spread is measured from it

\[ \textcolor{#1f5fbf}{\begin{aligned}&{-1.525},\ {-1.025},\ {-0.525}, \\ &{-0.025},\ 0.475,\ 0.975\end{aligned}} \]

Subtract x̄ from each age

Why: spread starts from these distances

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975)

\[ \begin{aligned}&1(-1.525) + 2(-1.025) + 4(-0.525) \\ &+ 4(-0.025) + 6(0.475) + 3(0.975) = 0\end{aligned} \]

Check the weighted distances balance

Why: confirms x̄ = 10.525

92. Each age's weighted square is f(x − x̄)²

Worked example

Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975)

\[ \textcolor{#b54708}{\begin{aligned}&2.325625,\ 1.050625,\ 0.275625, \\ &0.000625,\ 0.225625,\ 0.950625\end{aligned}} \]

Square each distance

Why: so each gap becomes an area

Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625)

\[ \textcolor{#b54708}{\begin{aligned}&2.325625,\ 2.10125,\ 1.1025, \\ &0.0025,\ 1.35375,\ 2.851875\end{aligned}} \]

Multiply each by f

Why: f equal copies

Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875)

\[ 1.1025 \div 4 = 0.275625 \]

Check: undo ×4 for age 10

Why: returns that age's square

93. The twenty ages' weighted squares total 9.7375

Worked example

Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875)

\[ \textstyle{\textstyle\sum} f(x - \bar{x})^2 = \textcolor{#b54708}{9.7375} \]

Add the six products

Why: s² needs all 20 pupils

Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375

\[ \begin{aligned}&(2.325625 + 2.10125 + 1.1025) \\ &+ (0.0025 + 1.35375 + 2.851875)\end{aligned} \]

Group the products in halves

Why: a second route to the total

\[ = 5.529375 + 4.208125 \]

Add inside each bracket

Why: splits the sum for checking

\[ 5.529375 + 4.208125 = \textcolor{#b54708}{9.7375} \]

Check: add the halves

Why: matches the table's total

94. The twenty ages have a sample SD of 0.72 years

Worked example

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, not yet filled

\[ 20 - 1 = 19 \]

Find n − 1 for 20 pupils

Why: they stand in for all fifth graders

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19

\[ \textcolor{#b54708}{s^2} = 9.7375 \div 19 \]

Substitute into s²

Why: the table gave this total

\[ = \textcolor{#b54708}{0.5125} \]

Divide by n − 1 = 19

Why: the sample's fair average square

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19, s² = 0.5125

\[ \textcolor{#1f5fbf}{s} = \sqrt{0.5125} \]

Substitute into s = √s²

Why: the side of that area

\[ \approx \textcolor{#1f5fbf}{0.716} \]

Take the root

Why: back in years

Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19, s² = 0.5125, s ≈ 0.716

\[ \textcolor{#1f5fbf}{0.716}^2 \times 19 \approx \textcolor{#b54708}{9.74} \]

Check: undo both moves

Why: restores the total

95. A calculator's population key treats a sample as everyone

Trap

The trap

\[ 9.7375 \div 20 \approx 0.487 \]

Divide by n = 20

Why: σx: treats 20 as everyone

\[ \sqrt{0.487} \approx 0.698 \]

Take the root

Why: to match the calculator's σx

\[ \text{report } 0.698 \]

Report σx

Why: fails: squares from x̄ run low

The fix

\[ 9.7375 \div 19 = 0.5125 \]

Divide by n − 1

Why: fair for a sample

\[ \sqrt{0.5125} \approx 0.716 \]

Take the root

Why: back in years

\[ \text{report } s = 0.716 \]

Report Sx = s

Why: s² is right on average

\[ \textcolor{#1f5fbf}{0.698}^2 \times 20 \approx \textcolor{#1f5fbf}{0.716}^2 \times 19 \approx \textcolor{#b54708}{9.74} \]

Check: both rebuild the total

Why: only the divisor differs

96. Divided by five, five heart rates understate all adults' spread on average

Prediction

Predict first

5 volunteers stand in for all adults.

Why does ÷ 5 understate their heart-rate spread on average?

  • Heart rates aren't bell-shaped
  • x̄ sits closer to the five than μ does
  • Dividing by 5 counts x̄ as a sixth value

Correct: x̄ sits closer to the five than μ does

Why: Squares from x̄ are squares from μ minus an offset area that averages σ². Dividing by n ignores the lost area; n − 1 restores it.

97. For five heart rates, ÷ 5 averages only 0.8σ²

Worked example

Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²

\[ 5 - 1 = 4 \]

Find n − 1

Why: (n − 1)σ² needs this count

\[ 4\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} (x - \bar{x})^2 \]

Substitute into (n − 1)σ²

Why: one σ² lost

Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²; squares from x̄ average four blocks, 4σ², the fifth lost as the offset area

\[ 4\textcolor{#b54708}{\sigma^2}/5 = \text{avg}{\textstyle\sum} (x - \bar{x})^2/5 \]

Divide both sides by 5

Why: what ÷ n would report

\[ 4\textcolor{#b54708}{\sigma^2}/5 = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/5] \]

Move ÷ 5 inside avg

Why: a constant leaves it

\[ 0.8\textcolor{#b54708}{\sigma^2} = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/5] \]

Evaluate 4 ÷ 5

Why: to compare with one σ²

Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²; squares from x̄ average four blocks, 4σ², the fifth lost as the offset area; dividing by 5 averages 0.8σ², short of one σ² block

\[ 0.8 \times \textcolor{#b54708}{25} = (4 \times \textcolor{#b54708}{25}) \div 5 \]

Check with σ² = 25

Why: both sides give 20

98. The same twelve heights can call for N or n − 1, depending on the question

Sorting

Sort into buckets

Divide by N (the data are the population) or n − 1 (a sample)?

Population: ÷ N
All 12 players' heights, to describe that team; Every wait today, to describe today
Sample: ÷ (n − 1)
12 players' heights, to learn about all college players; 30 waits, to learn about this year
N
The data are every value you want to describe, so nothing is estimated.
n1
The data stand in for a larger group, and x̄ hides one σ² of area.

99. Sample or population depends on what the data must describe

Worked example

Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12

Team heights, today's waits: ÷ N

Why: the data are the whole group

College players, this year's waits: ÷ (n − 1)

Why: data estimate a larger group

\[ 12 - 1 = 11 \]

Find n − 1 for 12 players

Why: needed to compare the two divisors

Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12; the same total cut into 11 equal shares, dividing by n − 1 = 11

\[ 12 \div 11 \approx \textcolor{#b54708}{1.09} \]

Compare the divisors

Why: shows how much larger s² runs

Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12; the same total cut into 11 equal shares, dividing by n − 1 = 11; one share of each highlighted, the 11-way share about 1.09 times as long

\[ 11 \times 1.09 \approx 12 \]

Check: undo the divide

Why: returns the 12 players

100. The stores average about 9.28 food types

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6

Try It 2.33: f of the 29 stores carry x food types.

\[ \begin{aligned}&4(6) + 5(7) + 1(8) + 4(9) \\ &+ 5(10) + 4(11) + 6(12) \\ &= 24 + 35 + 8 + 36 + 50 + 44 + 72\end{aligned} \]

Multiply each x by f

Why: f equal values at once

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, products fx 24, 35, 8, 36, 50, 44, 72

\[ = 269 \]

Add the seven products

Why: every store counts

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, products fx 24, 35, 8, 36, 50, 44, 72; totals: 29 stores, fx 269

\[ \bar{x} = 269 \div 29 \approx 9.28 \]

Divide by n = 29

Why: a mean is total over count

\[ 29 \times 9.28 \approx 269.1 \]

Check: undo the division

Why: near 269: x̄ was rounded

101. The model row x = 6 gives about 43.034

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² not yet filled

\[ 6 - 9.28 = \textcolor{#1f5fbf}{-3.28} \]

Subtract x̄ from 6

Why: spread is measured from the mean

\[ (\textcolor{#1f5fbf}{-3.28})^2 = \textcolor{#b54708}{10.7584} \]

Square it

Why: the column holds f(x − x̄)²

\[ 4 \times 10.7584 \approx \textcolor{#b54708}{43.034} \]

Multiply by f = 4

Why: four stores, four squares

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, the rest left blank

\[ \sqrt{43.034 \div 4} \approx \textcolor{#1f5fbf}{3.28} \]

Check: undo both moves

Why: returns the distance from x̄

102. Rows x = 7 and 8 give 25.992 and about 1.638

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, the rest left blank

\[ 7 - 9.28,\ \ 8 - 9.28 = \textcolor{#1f5fbf}{-2.28,\ -1.28} \]

Subtract x̄ from 7 and 8

Why: spread starts at x̄

\[ (\textcolor{#1f5fbf}{-2.28})^2,\ (\textcolor{#1f5fbf}{-1.28})^2 = \textcolor{#b54708}{5.1984,\ 1.6384} \]

Square each distance

Why: before f scales them

\[ 5 \times 5.1984,\ \ 1 \times 1.6384 \approx \textcolor{#b54708}{25.992,\ 1.638} \]

Multiply each by its f

Why: f equal squares

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, the rest left blank

\[ \begin{aligned}&9.28 - \sqrt{25.992 \div 5} = 7 \\ &9.28 - \sqrt{1.6384 \div 1} = 8\end{aligned} \]

Check: undo every move

Why: catches a slip in either row

103. Rows x = 9 and 10 give about 0.314 and 2.592

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, the rest left blank

\[ 9 - 9.28,\ \ 10 - 9.28 = \textcolor{#1f5fbf}{-0.28,\ 0.72} \]

Subtract x̄ from 9 and 10

Why: signed gaps feed the squares

\[ (\textcolor{#1f5fbf}{-0.28})^2,\ \textcolor{#1f5fbf}{0.72}^2 = \textcolor{#b54708}{0.0784,\ 0.5184} \]

Square each distance

Why: areas before f weights them

\[ 4 \times 0.0784,\ \ 5 \times 0.5184 \approx \textcolor{#b54708}{0.314,\ 2.592} \]

Multiply each by its f

Why: four and five equal areas

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, the rest left blank

\[ \begin{aligned}&9.28 - \sqrt{0.3136 \div 4} = 9 \\ &9.28 + \sqrt{2.592 \div 5} = 10\end{aligned} \]

Check: undo every move

Why: returns both rows exactly to x

104. Row x = 11 gives about 11.834

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, the rest left blank

\[ 11 - 9.28 = \textcolor{#1f5fbf}{1.72} \]

Subtract x̄ from 11

Why: the gap the square needs

\[ \textcolor{#1f5fbf}{1.72}^2 = \textcolor{#b54708}{2.9584} \]

Square the distance

Why: one store's area first

\[ 4 \times 2.9584 \approx \textcolor{#b54708}{11.834} \]

Multiply by f = 4

Why: four stores share this value

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank

\[ \sqrt{11.8336 \div 4} + 9.28 = 1.72 + 9.28 = 11 \]

Check: undo all three moves

Why: lands back on x = 11

105. The first six weighted squares total 85.404

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank

\[ \begin{aligned}&43.034 + 25.992 + 1.638 \\ &+ 0.314 + 2.592 + 11.834 = \textcolor{#b54708}{85.404}\end{aligned} \]

Add the six rows

Why: s² needs the table's total

\[ \begin{aligned}&85.404 - 11.834 - 2.592 - 0.314 \\ &- 1.638 - 25.992 = 43.034\end{aligned} \]

Check: subtract back to row 6

Why: lands on the model row's value

106. The pet-food table leaves its last row and s for you

Faded example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank

Try It 2.33: 29 stores carry 6 to 12 food types; x̄ ≈ 9.28.

First six rows total 85.404.

Fill in the blanks

6(12 − 9.28)² ≈ 44.39; total ≈ 129.79; s² = total ÷ 28 ≈ 4.64; s ≈ 2.15

Why: 12 − 9.28 = 2.72, 2.72² = 7.3984, and 6 × 7.3984 ≈ 44.39. Then 85.404 + 44.390 ≈ 129.79. 29 stores are a sample, so divide by n − 1 = 28 to get about 4.64, then take the square root: about 2.15.

107. The pet-food table's last row is about 44.39

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank

\[ 12 - 9.28 = \textcolor{#1f5fbf}{2.72} \]

Subtract x̄ from 12

Why: spread is measured from the mean

\[ \textcolor{#1f5fbf}{2.72}^2 = \textcolor{#b54708}{7.3984} \]

Square the distance

Why: far rows weigh more

\[ 6 \times 7.3984 \approx \textcolor{#b54708}{44.390} \]

Multiply by f = 6

Why: f equal squares

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390

\[ \sqrt{44.390 \div 6} + 9.28 \approx 2.72 + 9.28 = 12 \]

Check: undo all three moves

Why: lands back on x = 12

108. The weighted squares total about 129.79

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390

\[ 85.404 + 44.390 = \textcolor{#b54708}{129.794} \]

Add to the first six

Why: s² needs all 29 stores

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794

\[ \begin{aligned}&(43.034 + 25.992 + 1.638) \\ &+ (0.314 + 2.592 + 11.834 + 44.390)\end{aligned} \]

Regroup the seven rows

Why: a second route to the total

\[ = 70.664 + 59.130 \]

Add inside each bracket

Why: gives the check two subtotals

\[ 70.664 + 59.130 = \textcolor{#b54708}{129.794} \]

Check in two halves

Why: matches the table's total

109. The pet-food stores have a sample SD of about 2.15 food types

Worked example

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, not yet filled

\[ 29 - 1 = 28 \]

Find n − 1 for 29 stores

Why: a sample of the whole chain

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28

\[ \textcolor{#b54708}{s^2} = 129.794 \div 28 \]

Substitute the total into s²

Why: so 28 can share it

\[ \approx \textcolor{#b54708}{4.64} \]

Divide by 28

Why: n − 1: fair for a sample

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28, s² ≈ 4.64

\[ \textcolor{#1f5fbf}{s} = \sqrt{4.64} \]

Substitute into s = √s²

Why: the side of that area

\[ \approx \textcolor{#1f5fbf}{2.15} \]

Take the root

Why: back in food types

Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28, s² ≈ 4.64, s ≈ 2.15

\[ \textcolor{#1f5fbf}{2.15}^2 \times 28 \approx \textcolor{#b54708}{129.4} \]

Check: undo both moves

Why: near 129.79: rounding

110. Measure distances in standard deviations

Section

Idea 4 of 5

111. Binh's 4-minute gap from the mean needs a spread yardstick before it counts as far

Concept

Figure (svg): Number line of waits 0 to 10 with the mean 5 and Binh's wait of 1 marked; no SD hops yet

A bakery's recorded waits: x̄ = 5 minutes.

Discussion prompt

Binh waited 1 minute. Is 4 minutes from the mean unusual?

Answer:

It depends how far waits typically stray: here s = 2 minutes.

Mark hops of s = 2 minutes out from the mean.

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

112. Counting SD hops puts Rosa one hop above the mean and Binh two below

Worked example

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

Rosa waited 7 minutes, Binh 1 minute; x̄ = 5, s = 2.

\[ 7 - 5 = \textcolor{#1f5fbf}{2} \]

Find Rosa's distance

Why: hop counts need the gap first

\[ \textcolor{#1f5fbf}{2} \div \textcolor{#1f5fbf}{2} = 1 \]

Divide by the SD

Why: counts 2-minute hops

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

\[ 1 - 5 = \textcolor{#1f5fbf}{-4} \]

Find Binh's distance

Why: a negative gap marks the left side

\[ \textcolor{#1f5fbf}{-4} \div \textcolor{#1f5fbf}{2} = -2 \]

Divide by the SD

Why: a signed count keeps the direction

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

\[ 5 + (-2)(\textcolor{#1f5fbf}{2}) = 1 \]

Check: hop back to Binh

Why: two hops down land on 1

113. Any value equals the mean plus some number of standard deviations

Worked example

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

\[ x - \bar{x} \]

Measure the distance

Why: what the hops must cover

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

\[ k = \frac{\textcolor{#1f5fbf}{x - \bar{x}}}{\textcolor{#1f5fbf}{s}} \]

Divide by s

Why: k counts hops, signed

Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5

\[ k\,\textcolor{#1f5fbf}{s} = x - \bar{x} \]

Multiply both sides by s

Why: k hops cover the gap

\[ x = \bar{x} + k\,\textcolor{#1f5fbf}{s} \]

Add x̄ to both sides

Why: start at the mean, hop k times

\[ 5 + (-2)(\textcolor{#1f5fbf}{2}) = 1 \]

Check with Binh

Why: lands on 1 minute

114. Ages one SD up reach 11.25; two SDs down, 9.09

Worked example

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

x̄ = 10.53 years, s = 0.72 years (Example 2.32).

\[ (1)(\textcolor{#1f5fbf}{0.72}) = 0.72 \]

Multiply k = 1 by s

Why: hops are measured in SDs

\[ 10.53 + 0.72 = 11.25 \]

Add it to x̄

Why: hops start at the mean

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

\[ (-2)(\textcolor{#1f5fbf}{0.72}) = -1.44 \]

Multiply k = −2 by s

Why: negative k hops down

\[ 10.53 + (-1.44) = 9.09 \]

Add it to x̄

Why: x = x̄ + ks places the value

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

\[ (9.09 - 10.53) \div \textcolor{#1f5fbf}{0.72} = -2 \]

Check: count hops back

Why: should be two down

115. Ages 1.5 SDs either side of the mean run 9.45 to 11.61

Worked example

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

\[ (-1.5)(\textcolor{#1f5fbf}{0.72}) = -1.08 \]

Multiply k = −1.5 by s

Why: k needn't be whole

\[ 10.53 + (-1.08) = 9.45 \]

Add it to x̄

Why: hops start at x̄

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

\[ (1.5)(\textcolor{#1f5fbf}{0.72}) = 1.08 \]

Multiply k = 1.5 by s

Why: opposite sign, mirror hop

\[ 10.53 + 1.08 = 11.61 \]

Add it to x̄

Why: hops leave from the mean

Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53

\[ \frac{9.45 + 11.61}{2} = 10.53 \]

Check the pair's centre

Why: equal hops balance at x̄

116. An age two SDs above the team's mean reveals the team's SD

Prediction

Predict first

Adapted from Try It 2.32: a team's mean age is 30.68. An age of 42.86 would sit exactly 2 SDs above it.

What is the team's SD?

  • 6.09 years
  • 12.18 years
  • 21.43 years
  • 3.05 years

Correct: 6.09 years

Why: The gap is 42.86 − 30.68 = 12.18 years, and it is two SDs long, so one SD is 12.18 ÷ 2 = 6.09.

117. Two hops span 12.18 years, so s = 6.09

Worked example

Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked

\[ 42.86 = 30.68 + 2\textcolor{#1f5fbf}{s} \]

Substitute into x = x̄ + ks

Why: leaves s as the only unknown

Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled s

\[ 12.18 = 2\textcolor{#1f5fbf}{s} \]

Subtract 30.68 from both sides

Why: isolates the two hops

Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled s; an arrow over the whole gap labelled 12.18 = 2s

\[ \textcolor{#1f5fbf}{s} = 6.09 \]

Divide both sides by 2

Why: equal division keeps both sides equal

Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled 6.09; an arrow over the whole gap labelled 12.18 = 2s; a blue tick one SD above the mean

\[ 30.68 + 2(\textcolor{#1f5fbf}{6.09}) = 42.86 \]

Check by hopping back

Why: lands on the age 42.86

118. A manager's report gives only N, μ and σ, yet it can be tested

Concept

Figure (svg): Number line from 4 to 16 with μ = 10 and the band under 2σ = 4 shaded; no data shown

Report: N 8, μ 10, σ 2. Claim: 3 waits sit 4+ minutes from μ.

Discussion prompt

Without the data, could that claim be true?

Answer:

Build data with σ = 2 and count how many far waits fit.

119. Store C's eight waits average μ = 10

Worked example

Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes

Store C's waits: 6, six 10s, 14.

\[ 6 \times \textcolor{#1f5fbf}{10} = 60 \]

Total the six 10s

Why: one product, not six sums

\[ \textcolor{#1f5fbf}{6} + 60 + \textcolor{#1f5fbf}{14} = 80 \]

Add the end waits

Why: the mean needs the total

\[ \mu = 80 \div 8 = 10 \]

Divide by N = 8

Why: tests the report's μ

Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked

\[ (\textcolor{#1f5fbf}{6} - 10) + (\textcolor{#1f5fbf}{14} - 10) = 0 \]

Check: the ends balance

Why: deviations must total zero

Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked and arrows of −4 and +4 from μ to the end waits

120. Store C's squares total 32: finish its σ

Worked example

Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked

Store C: 6, six 10s, 14; μ = 10.

\[ \textcolor{#1f5fbf}{-4,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ +4} \]

Subtract μ

Why: spread starts from distances

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: each distance drawn as a blue side, no squares yet

\[ \textcolor{#b54708}{16,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ 16} \]

Square each

Why: σ² is built from these

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16

\[ \textcolor{#b54708}{16 + 16 = 32} \]

Add the squares

Why: σ² shares this total

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32

Fill in the blanks

Your turn: is the report right? Store C's σ² = 4 and its σ = 2 minutes.

Why: 32 ÷ 8 = 4 is the average square, and √4 = 2 minutes: the report's σ = 2 holds.

121. Store C's σ is 2, as the report says

Worked example

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32

\( {-4,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ +4}\;\;\Rightarrow\;\;\allowbreak {16,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ 16}\;\;\Rightarrow\;\;\allowbreak {16 + 16 = 32} \)

\[ \textcolor{#b54708}{\sigma^2} = 32 \div 8 = \textcolor{#b54708}{4} \]

Divide by N = 8

Why: σ² is the mean square

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32; the average square, area 4

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4} = \textcolor{#1f5fbf}{2} \]

Take the root

Why: back to minutes

Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32; the average square, area 4, side 2

\[ \textcolor{#1f5fbf}{2}^2 \times 8 = 32 \]

Check: square, times 8

Why: back to the total

122. Three far waits would need at least 48 of squared area

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each

\[ 4^2 = \textcolor{#b54708}{16} \]

Square the least far distance

Why: the smallest square a far wait can have

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16

\[ \text{far total} \ge \textcolor{#b54708}{16} + \textcolor{#b54708}{16} + \textcolor{#b54708}{16} \]

Add three floors

Why: sums keep ≥

\[ \text{far total} \ge \textcolor{#b54708}{48} \]

Add 16 + 16 + 16

Why: one floor for the far area

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48

\[ \text{all 8 squares} \ge \textcolor{#b54708}{48} \]

Add the five other squares

Why: none is negative

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48

\[ 3 \times \textcolor{#b54708}{16} = \textcolor{#b54708}{48} \]

Check: three equal floors

Why: same total as adding

123. Three far waits need more area than σ = 2 allows

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48

\( {\text{far total} \ge \textcolor{#b54708}{48}}\;\;\Rightarrow\;\;\allowbreak {\text{all 8 squares} \ge \textcolor{#b54708}{48}} \)

\[ \text{all 8 squares} \div 8 \ge \textcolor{#b54708}{48} \div 8 \]

Divide both sides by N = 8

Why: a positive keeps ≥

\[ \textcolor{#b54708}{\sigma^2} \ge \textcolor{#b54708}{48} \div 8 \]

Substitute σ²

Why: the mean square

\[ \textcolor{#b54708}{\sigma^2} \ge \textcolor{#b54708}{6} \]

Evaluate 48 ÷ 8

Why: a floor to compare with σ² = 4

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48; label: σ² ≥ 6

\[ \textcolor{#b54708}{6} > \textcolor{#b54708}{4} \]

Compare with σ² = 4

Why: a floor can't exceed it

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48; label: σ² ≥ 6, but σ² = 4

\[ \textcolor{#b54708}{48} > \textcolor{#b54708}{32} = 8\textcolor{#b54708}{\sigma^2} \]

Check with areas

Why: part can't beat the whole

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48 > 32; label: σ² ≥ 6, but σ² = 4; a box of all the squared area 32 cut into 8 strips

124. Each far value's square is at least k²σ²

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded

For a population, x = μ + kσ: the same hop rule.

\[ |\textcolor{#1f5fbf}{x - \mu}| \ge \textcolor{#1f5fbf}{k\sigma} \]

Define far

Why: names the values the bound will count

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14

\[ \textcolor{#b54708}{|x - \mu|^2} \ge \textcolor{#b54708}{(k\sigma)^2} \]

Square both sides

Why: neither side is negative

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; unfilled 4-by-4 squares on the two far distances

\[ \textcolor{#b54708}{(x - \mu)^2} \ge \textcolor{#b54708}{(k\sigma)^2} \]

Drop the bars

Why: squaring already removed the sign

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the two far squares filled, area 16 each

\[ \textcolor{#b54708}{(x - \mu)^2} \ge k^2\textcolor{#b54708}{\sigma^2} \]

Square kσ

Why: square each factor

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the two far squares filled, area 16 each; label: each far square ≥ k²σ² = 16

\[ k = 2:\ (6 - 10)^2 = 16 \ge 4 \cdot 4 \]

Check store C's wait of 6

Why: its square meets the floor

125. The far total is at least k²σ² per far value

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each

\( {|x - \mu| \ge k\sigma}\;\;\Rightarrow\;\;\allowbreak {|x - \mu|^2 \ge (k\sigma)^2}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 \ge (k\sigma)^2}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 \ge k^2\sigma^2} \)

m counts far values, |x − μ| ≥ kσ; store C's m = 2.

\[ \textcolor{#b54708}{\text{far total}} = {\textstyle\sum}_{\text{far}} \textcolor{#b54708}{(x - \mu)^2} \]

Name the far total

Why: the floor bounds it

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: far total = 32

\[ \textcolor{#b54708}{\text{far total}} \ge \textcolor{#b54708}{{\textstyle\sum}_{\text{far}} k^2\sigma^2} \]

Add the m floors

Why: sums of ≥ keep ≥

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: each far square ≥ k²σ² = 16; label: far total = 32

\[ \textcolor{#b54708}{\text{far total}} \ge \textcolor{#b54708}{m\,k^2\sigma^2} \]

Add k²σ² m times

Why: Σc = mc collapses the sum

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: far total = 32; label: far total ≥ 2 × 16

\[ k = 2:\ 16 + 16 \ge 2 \cdot 4 \cdot 4 \]

Check store C

Why: far waits meet the floor

126. Far values' squares can never exceed the total area Nσ²

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each

\[ \textcolor{#b54708}{\sigma^2} = \frac{{\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2}}{N} \]

Write σ²'s definition

Why: the average square over all N

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, one strip of σ² = 4 filled

\[ \textcolor{#b54708}{N\sigma^2} = {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} \]

Multiply both sides by N

Why: undoes the average

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled

\[ \text{far total} \le {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} \]

Compare with all squares

Why: the others add zero or more

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box

\[ \text{far total} \le \textcolor{#b54708}{N\sigma^2} \]

Substitute the total area

Why: a cap from σ alone

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: far total ≤ 32

\[ \textcolor{#b54708}{32} \le 8(4) = \textcolor{#b54708}{32} \]

Check with store C

Why: its far squares use everything

127. At most 1/k² of values sit kσ or more from μ

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled

\( {\text{far total} \ge m\,k^2\sigma^2}\qquad\allowbreak {\text{far total} \le N\sigma^2} \)

\[ m\,k^2\textcolor{#b54708}{\sigma^2} \le \textcolor{#b54708}{N\sigma^2} \]

Chain floor and cap

Why: a floor can't exceed its cap

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: 2 × 16 ≤ 32

\[ \frac{m\,k^2\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \le \frac{N\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \]

Divide by Nk²σ²

Why: k, σ > 0 keep it positive

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box, each covering 4 strips; label: 2 × 16 ≤ 32

\[ \frac{m}{N} \le \frac{N\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \]

Cancel k²σ² on the left

Why: a common factor divides out

\[ \frac{m}{N} \le \frac{1}{k^2} \]

Cancel Nσ² on the right

Why: leaves only counts and k

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: 2 × 16 ≤ 32; label: far share 2/8 ≤ 1/4

\[ k = 2:\ \frac{2}{8} \le \frac{1}{4} \]

Check with store C

Why: met exactly at k = 2

128. At least 1 − 1/k² of values sit closer than kσ

Worked example

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded

\( {\tfrac{m}{N} \le \tfrac{1}{\textcolor{#b54708}{k^2}}} \)

\[ 1 - \tfrac{m}{N} \ge 1 - \tfrac{1}{\textcolor{#b54708}{k^2}} \]

Subtract both sides from 1

Why: so the order reverses

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; label: far ≤ 1/4, so within ≥ 3/4

\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = \tfrac{N - m}{N} \]

Count those within

Why: all but the m far

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the six waits within 2σ ringed; label: within 8 − 2 = 6; label: far ≤ 1/4, so within ≥ 3/4

\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = \tfrac{N}{N} - \tfrac{m}{N} \]

Split the fraction

Why: over a shared N

\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = 1 - \tfrac{m}{N} \]

Evaluate N/N = 1

Why: any nonzero count

\[ \text{within } \textcolor{#1f5fbf}{k\sigma} \ge 1 - \tfrac{1}{\textcolor{#b54708}{k^2}} \]

Replace 1 − m/N

Why: its bound carries over

Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the six waits within 2σ ringed; label: within 8 − 2 = 6; label: within 6/8 ≥ 3/4

\[ k = 2:\ \tfrac{6}{8} = 1 - \tfrac{1}{4} \]

Check store C

Why: six of eight

129. The cap guarantees about 89% within 3σ but 0% within 1σ

Concept

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5

Try k = 3, k = 4.5 (near 95%) and k = 1.

\[ 3^2 = 9,\ \ 4.5^2 = 20.25,\ \ 1^2 = 1 \]

Square each k

Why: three guarantees to compare

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5

\[ \tfrac{1}{9} \approx 0.11,\ \ \tfrac{1}{20.25} \approx 0.05,\ \ \tfrac{1}{1} = 1 \]

Take each reciprocal

Why: the far-share cap is 1/k²

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5; the dashed far-share cap curve 1/k² with points at 100%, 11% and 5%

\[ 1 - 0.11 = 0.89,\ 1 - 0.05 = 0.95,\ 1 - 1 = 0 \]

Subtract each from 1

Why: within plus far makes 1

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5; the dashed far-share cap curve 1/k² with points at 100%, 11% and 5%; points on the within curve at 0%, 89% and 95%

\[ 9 \times 0.11 \approx 1,\ \ 20.25 \times 0.05 \approx 1 \]

Check: undo each reciprocal

Why: both return about 1

130. A bell-curve guess undercounts store C's far waits

Trap

The trap

\[ \text{guess: closer than } 2\sigma \approx 95\% \]

Guess from a bell curve

Why: ignores store C's shape

\[ \text{far} \approx 100\% - 95\% = 5\% \]

Subtract from 100%

Why: shares sum to 100%

\[ 8 \times 5\% = 0.4 \]

Take 5% of 8

Why: fails: store C has 2

Book p. 117 (quoted): ≈ 95% within 2σ, bell shapes only

The fix

\[ \tfrac{m}{8} \le \tfrac{1}{2^2} \]

Use Chebyshev, N = 8, k = 2

Why: holds for any shape

\[ \tfrac{m}{8} \le \tfrac{1}{4} \]

Evaluate 2²

Why: gives a usable cap

\[ m \le 2 \]

Multiply both sides by 8

Why: turns the share into a count

\[ |6 - 10| = |14 - 10| = 4 = 2(2) \]

Check store C's far waits

Why: both sit 2σ out

131. Chebyshev's inequality limits how far out 40% of skewed prices can sit

Prediction

Predict first

200 house prices, skewed by mansions.

By Chebyshev, for which k can 40% sit at least k SDs out?

  • k ≤ 1.58
  • k ≤ 2.5
  • k ≥ 1.58
  • k = 2 only

Correct: k ≤ 1.58

Why: The far-share cap 1/k² must allow 0.4, so k² ≤ 2.5 and k ≤ √2.5 ≈ 1.58. Chebyshev only rules out k > 1.58, whatever the shape.

132. 40% can sit k SDs out only if k ≤ 1.58

Worked example

Figure (svg): Curve of the far-share cap 1/k² for k from 1 to 4 SDs, falling from 100% to about 6%, with a dashed line at the claimed 40%

\[ m = 0.4 \times 200 = 80 \]

Turn 40% into a count

Why: the cap's m counts values

\[ \tfrac{80}{200} \le \tfrac{1}{k^2} \]

Substitute into m/N ≤ 1/k²

Why: the cap, any shape

\[ 0.4 \le \tfrac{1}{k^2} \]

Evaluate 80 ÷ 200

Why: the cap must allow this share

\[ k^2 \le \tfrac{1}{0.4} \]

Take reciprocals

Why: the order reverses

\[ k^2 \le 2.5 \]

Evaluate 1 ÷ 0.4

Why: biggest k² possible

\[ k \le \sqrt{2.5} \approx 1.58 \]

Take the root

Why: k counts SDs, so positive

Figure (svg): Curve of the far-share cap 1/k² for k from 1 to 4 SDs, falling from 100% to about 6%, with a dashed line at the claimed 40%; the crossing marked at k ≈ 1.58

\[ \tfrac{1}{1.58^2} \approx 0.40 \]

Check by substituting k = 1.58

Why: the cap just allows 40%

133. Compare data sets with z-scores

Section

Idea 5 of 5

134. John's 2.85 and Ali's 77 are grades on different scales at different schools

Concept

Figure (svg): Two raw rulers: John's school GPA 1.5 to 4.5 with John at 2.85; Ali's school 50 to 110 with Ali at 77

John: 2.85, school mean 3.0, SD 0.7. Ali: 77, school mean 80, SD 10.

Discussion prompt

Relative to his own school, whose GPA is stronger? Commit first.

Answer:

John's. The next slides measure each gap in its school's SDs to show why raw gaps mislead.

135. Raw gaps of −0.15 and −3 cannot be compared because the rulers differ

Worked example

Figure (svg): Two raw rulers at their natural scales

\[ \textcolor{#1f5fbf}{2.85} - 3.0 = -0.15 \]

Find John's raw gap

Why: the gap each ruler must measure

\[ \textcolor{#1f5fbf}{77} - 80 = -3 \]

Find Ali's raw gap

Why: to set beside John's −0.15

\[ 3.0 + (-0.15) = \textcolor{#1f5fbf}{2.85},\ \ 80 + (-3) = \textcolor{#1f5fbf}{77} \]

Check: add each gap to its mean

Why: recovers both grades

Different units: stretch each ruler so one SD matches.

Figure (svg): Both rulers stretched so one SD has the same length, with a shared axis of SDs from −3 to +3

136. In school SDs, John's GPA is stronger

Worked example

Figure (svg): Both rulers stretched so one SD has the same length, with a shared axis of SDs

\( {{2.85 - 3.0 = -0.15}}\qquad\allowbreak {{77 - 80 = -3}} \)

\[ z = \frac{\textcolor{#1f5fbf}{x - \mu}}{\textcolor{#1f5fbf}{\sigma}} \]

Name that hop count z

Why: population μ, σ; samples use x̄, s

\[ z_J = \tfrac{-0.15}{\textcolor{#1f5fbf}{0.7}},\ \ z_A = \tfrac{-3}{\textcolor{#1f5fbf}{10}} \]

Substitute both students

Why: each school's own σ

\[ z_J \approx -0.21,\ \ z_A = -0.30 \]

Divide each gap

Why: one shared SD ruler

Figure (svg): Both rulers aligned in SDs, John at −0.21 and Ali at −0.30

\[ -0.21 > -0.30 \]

Compare the z-scores

Why: fewer SDs below means relatively higher

\[ (-0.21)(\textcolor{#1f5fbf}{0.7}) \approx -0.15,\ (-0.30)(\textcolor{#1f5fbf}{10}) = -3 \]

Check: z times σ

Why: gives back both gaps

137. A z-score is unitless, signed and zero at μ

Concept

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed

\[ \tfrac{\text{GPA pts}}{\text{GPA pts}} = \text{no units} \]

Check z's units

Why: so GPAs and scores compare

\[ x = \mu \Rightarrow \textcolor{#1f5fbf}{x - \mu} = 0 \]

Set x = μ

Why: find z's zero

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ

\[ z = \tfrac{\textcolor{#1f5fbf}{0}}{\textcolor{#1f5fbf}{\sigma}} = 0 \]

Divide by σ

Why: 0 over a positive is 0

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ

\[ x < \mu \Leftrightarrow \textcolor{#1f5fbf}{x - \mu} < 0 \]

Subtract μ

Why: smaller minus larger is negative

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0

\[ \textcolor{#1f5fbf}{x - \mu} < 0 \Leftrightarrow \tfrac{\textcolor{#1f5fbf}{x - \mu}}{\sigma} < 0 \]

Divide by σ > 0

Why: positives keep order

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ

\[ z < 0 \Leftrightarrow x < \mu \]

Chain both

Why: equivalences link end to end

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ; label: z < 0: below μ

\[ 2.85 < 3.0,\ z \approx -0.21 < 0 \]

Check with John

Why: sign matches his side

Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ; label: z < 0: below μ; John's 2.85 marked just left of μ

138. Angie swims 1.25 SDs under her team mean

Worked example

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie 26.2 marked; Beth 27.3 marked

Angie: 26.2 s; team mean 27.2 s.

\[ \textcolor{#1f5fbf}{26.2} - 27.2 = -1.0 \]

Subtract Angie's team mean

Why: z needs the gap

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie's raw gap −1.0 s drawn from the mean; Angie 26.2 marked; Beth 27.3 marked

\[ z_{\text{Angie}} = \frac{-1.0}{\textcolor{#1f5fbf}{0.8}} \]

Substitute into z = gap ÷ SD

Why: her own team's ruler

\[ z_{\text{Angie}} = -1.25 \]

Divide −1.0 by 0.8

Why: to compare swims

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth 27.3 marked

\[ 27.2 + (-1.25)(\textcolor{#1f5fbf}{0.8}) = 26.2 \]

Check: hop back

Why: recovers Angie's time

139. Beth swims 2 SDs under her team mean

Worked example

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth 27.3 marked

Beth: 27.3 s; team mean 30.1 s.

\[ \textcolor{#1f5fbf}{27.3} - 30.1 = -2.8 \]

Subtract Beth's team mean

Why: z needs her gap

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth's raw gap −2.8 s drawn from the mean; Beth 27.3 marked

\[ z_{\text{Beth}} = \frac{-2.8}{\textcolor{#1f5fbf}{1.4}} \]

Substitute into z = gap ÷ SD

Why: her team spreads wider

\[ z_{\text{Beth}} = -2.00 \]

Divide −2.8 by 1.4

Why: z puts both on one scale

Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth: z = −2, drawn as hops of one SD from the mean; Beth 27.3 marked

\[ 30.1 + (-2)(\textcolor{#1f5fbf}{1.4}) = 27.3 \]

Check: hop back

Why: recovers Beth's time

140. For swim times, the lower z is the faster swim

Trap

The trap

\[ \text{Angie } {-1.25},\ \text{Beth } {-2.00} \]

Start from both z-scores

Why: each on her team's ruler

\[ -1.25 > -2.00 \Rightarrow \text{Angie} \]

Pick the larger z

Why: fails: a smaller time is faster

The fix

\[ \text{time: lower is better} \]

Ask which direction wins

Why: fewer seconds means faster

\[ -2.00 < -1.25 \Rightarrow \text{Beth} \]

Pick the lower z

Why: Beth is further below her mean

\[ 30.1 + (-2)(1.4) = 27.3 \]

Check Beth's z by hopping back

Why: recovers her time

141. The same 0.5 kg gap can mean different things at two hospitals

Prediction

Predict first

4.0 kg newborns at hospitals A, B: means 3.5, SDs 0.25, 0.5.

Which is heavier for its hospital?

  • Hospital A's
  • Hospital B's (wider spread)
  • Equal: same 0.5 kg gap

Correct: Hospital A's

Why: Hospital A's weights typically stray about 0.25 kg from 3.5, so 0.5 kg up is 2 SDs. Hospital B's typically stray about 0.5 kg, so 0.5 kg up is 1 SD.

142. Hospital A's baby is z = 2; hospital B's is z = 1

Worked example

Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; A's baby 4.0 marked; B's baby 4.0 marked

\[ 4.0 - 3.5 = \textcolor{#1f5fbf}{0.5} \]

Find the gap

Why: z starts from the raw gap

Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A's raw gap +0.5 kg drawn from the mean; A's baby 4.0 marked; Hospital B's raw gap +0.5 kg drawn from the mean; B's baby 4.0 marked

\[ z_A = \textcolor{#1f5fbf}{0.5} \div \textcolor{#1f5fbf}{0.25} \]

Substitute A's SD

Why: z uses each hospital's ruler

\[ z_A = 2 \]

Divide 0.5 by 0.25

Why: counts SD hops

Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A: z = 2, drawn as hops of one SD from the mean; A's baby 4.0 marked; Hospital B's raw gap +0.5 kg drawn from the mean; B's baby 4.0 marked

\[ z_B = \textcolor{#1f5fbf}{0.5} \div \textcolor{#1f5fbf}{0.5} \]

Substitute B's SD

Why: same gap, wider ruler

\[ z_B = 1 \]

Divide 0.5 by 0.5

Why: now both in SDs

Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A: z = 2, drawn as hops of one SD from the mean; A's baby 4.0 marked; Hospital B: z = 1, drawn as hops of one SD from the mean; B's baby 4.0 marked

\[ 2 > 1 \]

Compare hop counts

Why: A's baby sits further out

\[ 3.5 + (2)(\textcolor{#1f5fbf}{0.25}) = 4.0,\ \ 3.5 + (1)(\textcolor{#1f5fbf}{0.5}) = 4.0 \]

Check: hop back

Why: each lands on 4.0 kg

143. Given two z-scores, you work back to the raw scores

Faded example

\[ x = \bar{x} + z\,\textcolor{#1f5fbf}{s} \]

Given: z counts hops like k

Why: value formula x = x̄ + ks, k = z

Fill in the blanks

Lee: z = 1.6, mean 80, SD 5 → score = 88. Kai: z = 0.92, mean 73.5, SD 17.9 → score ≈ 90.

Why: Start at each class mean and hop z SDs: 80 + 1.6 × 5 = 88 and 73.5 + 0.92 × 17.9 ≈ 90. Kai scored higher yet stands out less.

144. Lee scored 88 and Kai about 90, but Lee stands out more

Worked example

Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9

\[ (1.6)(\textcolor{#1f5fbf}{5}) = 8,\ \ (0.92)(\textcolor{#1f5fbf}{17.9}) \approx 16.47 \]

Multiply each z by its SD

Why: undoes dividing by the SD

Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9; Lee: 1.6 hops = 8, drawn as hops of one SD from the mean; Kai: 0.92 hops ≈ 16.47, drawn as hops of one SD from the mean

\[ 80 + 8 = 88,\ \ 73.5 + 16.47 \approx 90 \]

Add each to its class mean

Why: the starting point of each hop

Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9; Lee: 1.6 hops = 8, drawn as hops of one SD from the mean; Lee 88 marked; Kai: 0.92 hops ≈ 16.47, drawn as hops of one SD from the mean; Kai ≈ 90 marked

\[ (88 - 80) \div \textcolor{#1f5fbf}{5} = 1.6 \]

Check Lee's z

Why: 8 over 5 gives back 1.6

\[ (90 - 73.5) \div \textcolor{#1f5fbf}{17.9} \approx 0.92 \]

Check Kai's z

Why: 16.5 over 17.9 gives back 0.92

145. Spread questions in this section run on five moves, then hops

Pattern

Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square

  1. Deviations from the mean
  2. Square each (× f in tables)
  3. Add the squares
  4. ÷ N for populations, ÷ (n − 1) for samples
  5. Take the root; then count hops:

\[ \textcolor{#1f5fbf}{s} = \sqrt{\textcolor{#b54708}{s^2}},\quad x = \bar{x} + k\,\textcolor{#1f5fbf}{s},\quad z = \frac{\textcolor{#1f5fbf}{x - \bar{x}}}{\textcolor{#1f5fbf}{s}} \]

146. A grouped table hides its values, so each class needs a stand-in

Concept

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16

Example 2.34: 26 values in classes 0–2 to 15–17.

Discussion prompt

A table gives only classes, not values. What stands in for each value?

Answer:

The class midpoint: an estimate.

147. Example 2.34's grouped mean is about 7.58

Worked example

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16

f = 1, 6, 10, 7, 0, 2 at midpoints 1, 4, 7, 10, 13, 16.

\[ \textcolor{#1f5fbf}{1,\ 24,\ 70,\ 70,\ 0,\ 32} \]

Multiply each midpoint by its f

Why: raw values unknown; centres stand in

\[ \textcolor{#1f5fbf}{1} + \textcolor{#1f5fbf}{24} + \textcolor{#1f5fbf}{70} + \textcolor{#1f5fbf}{70} + \textcolor{#1f5fbf}{0} + \textcolor{#1f5fbf}{32} = 197 \]

Add the products

Why: estimates the total of all values

\[ \bar{x} = 197 \div 26 \approx 7.58 \]

Divide by n = 26

Why: the mean every row uses

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked

\[ 26 \times 7.58 = 197.08 \approx 197 \]

Check: undo the division

Why: near 197: rounding

148. Class 3–5 adds 76.90 to the grouped table's squared total

Worked example

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked

Class 3–5: midpoint 4, f = 6; x̄ ≈ 7.58.

\[ \textcolor{#1f5fbf}{4} - 7.58 = \textcolor{#1f5fbf}{-3.58} \]

Subtract x̄ from the midpoint

Why: estimate: six values placed at 4

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4

\[ (\textcolor{#1f5fbf}{-3.58})^2 = \textcolor{#b54708}{12.8164} \]

Square the distance

Why: squares keep below-mean rows positive

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4; a square on the −3.58 distance, area 12.82

\[ 6 \times \textcolor{#b54708}{12.8164} = \textcolor{#b54708}{76.8984} \]

Multiply by f = 6

Why: all six values counted

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4; a square on the −3.58 distance, area 12.82; label: × 6 = 76.90 for the class's 6 values

\[ \sqrt{76.8984 \div 6} = \textcolor{#1f5fbf}{3.58} \]

Check by undoing both moves

Why: recovers the distance 3.58

149. Class 6–8 adds 3.364, class 15–17 adds 141.79

Worked example

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked

\[ 7 - 7.58,\ 16 - 7.58 = \textcolor{#1f5fbf}{-0.58},\ \textcolor{#1f5fbf}{+8.42} \]

Subtract x̄ from each midpoint

Why: spread is measured from the mean

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16

\[ (\textcolor{#1f5fbf}{-0.58})^2,\ \textcolor{#1f5fbf}{8.42}^2 = \textcolor{#b54708}{0.3364},\ \textcolor{#b54708}{70.8964} \]

Square each distance

Why: far sides, big areas

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16; a square on the −0.58 distance, area 0.34; a square on the +8.42 distance, area 70.90

\[ 10 \times \textcolor{#b54708}{0.3364} = \textcolor{#b54708}{3.364} \]

Multiply by f = 10

Why: each value adds one square

\[ 2 \times \textcolor{#b54708}{70.8964} = \textcolor{#b54708}{141.7928} \]

Multiply by f = 2

Why: a square for each

Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16; a square on the −0.58 distance, area 0.34; a square on the +8.42 distance, area 70.90; label: × 10 = 3.36 for the class's 10 values; label: × 2 = 141.79 for the class's 2 values

\[ \sqrt{3.364 \div 10} = 0.58,\ \ \sqrt{141.7928 \div 2} = 8.42 \]

Check: undo both moves

Why: returns both distances

150. Classes 0–2 and 9–11 finish the table

Faded example

Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries blank, 76.8984, 3.3640, blank, 0, 141.7928

Class 0–2: midpoint 1, f = 1. Class 9–11: midpoint 10, f = 7.

Fill in the blanks

Fill the two remaining rows, with x̄ ≈ 7.58: class 0–2's f(mid − x̄)² ≈ 43.30 and class 9–11's ≈ 40.99.

Why: 1 − 7.58 = −6.58, squared 43.2964, times f = 1 ≈ 43.30. 10 − 7.58 = 2.42, squared 5.8564, times f = 7 ≈ 40.99.

151. Class 0–2 adds 43.30 and class 9–11 adds 40.99

Worked example

Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries blank, 76.8984, 3.3640, blank, 0, 141.7928

\[ 1 - 7.58,\ 10 - 7.58 = \textcolor{#1f5fbf}{-6.58},\ \textcolor{#1f5fbf}{+2.42} \]

Subtract x̄ from each midpoint

Why: gaps start at x̄

\[ (\textcolor{#1f5fbf}{-6.58})^2,\ \textcolor{#1f5fbf}{2.42}^2 = \textcolor{#b54708}{43.2964},\ \textcolor{#b54708}{5.8564} \]

Square each distance

Why: so below-mean rows can't cancel

\[ 1 \times \textcolor{#b54708}{43.2964} = \textcolor{#b54708}{43.2964} \]

Multiply by f = 1

Why: one value, one square

Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries 43.2964, 76.8984, 3.3640, blank, 0, 141.7928

\[ 7 \times \textcolor{#b54708}{5.8564} = \textcolor{#b54708}{40.9948} \]

Multiply by f = 7

Why: seven equal squares

Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries 43.2964, 76.8984, 3.3640, 40.9948, 0, 141.7928

\[ \sqrt{43.2964} = 6.58,\ \ \sqrt{40.9948 \div 7} = 2.42 \]

Check: reverse both moves

Why: back to both gaps

152. A grouped frequency table needs one more check before its SD

Check

Check your understanding

Class 15–17's pair, at midpoint 16, supplies 141.79 squared area (x̄ ≈ 7.58). What would true values 15 and 17 supply?

  • A. 143.79 (correct)
  • B. 141.79
  • C. 16.84
  • D. 283.59

Answer: A

Why: 15 − 7.58 = 7.42, 7.42² = 55.0564; 17 − 7.58 = 9.42, 9.42² = 88.7364; total 143.7928 ≈ 143.79 — 2 more than the midpoint gives.

Why B tempts people
Assumes the midpoint is exact.
Why C tempts people
Adds 7.42 + 9.42 but never squares.
Why D tempts people
Adds the distances, then squares: 16.84².

153. Values 15 and 17 supply 143.79, not 141.79

Worked example

Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; the true pair's distances still grey dashed, not yet measured

\[ 15 - 7.58,\ 17 - 7.58 = \textcolor{#1f5fbf}{7.42},\ \textcolor{#1f5fbf}{9.42} \]

Subtract x̄ from each

Why: no midpoint stand-in

Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; the true distances 7.42 and 9.42 drawn as blue sides

\[ \textcolor{#1f5fbf}{7.42}^2,\ \textcolor{#1f5fbf}{9.42}^2 = \textcolor{#b54708}{55.0564},\ \textcolor{#b54708}{88.7364} \]

Square each distance

Why: rounding would blur the error

Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364

\[ \textcolor{#b54708}{55.0564} + \textcolor{#b54708}{88.7364} = \textcolor{#b54708}{143.7928} \]

Add the two areas

Why: one value each, so no × f

Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364; the true pair's total 143.7928

\[ 143.7928 - 141.7928 = 2 \]

Subtract the midpoint row's area

Why: sizes the error

Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364; the true pair's total 143.7928; label: true pair: 2 more area

\[ (\textcolor{#b54708}{55.06} - \textcolor{#b54708}{70.90}) + (\textcolor{#b54708}{88.74} - \textcolor{#b54708}{70.90}) = 2 \]

Check against 16's square

Why: the gaps net 2

154. 7.42² expands to 8.42² − 2(8.42) + 1²

Worked example

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge

\[ \textcolor{#1f5fbf}{7.42} = 8.42 - 1 \]

Write 7.42 from 16's distance

Why: the value sits 1 inside

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 7.42 dashed inside it, labelled on its top edge

\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)(-1)} + \textcolor{#b54708}{(-1)^2} \]

Expand (a + b)²

Why: rule for a sum, b = −1

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 7.42 dashed inside it, labelled on its top edge; two blue 8.42-by-1 strips cut away and their solid orange 1-by-1 corner, cut twice

\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} - \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{(-1)^2} \]

Multiply 2(8.42) by −1

Why: sign rule: − × + = −

\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} - \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{1^2} \]

Write (−1)² as 1²

Why: even powers drop the sign

\[ \textcolor{#b54708}{70.8964} - \textcolor{#1f5fbf}{16.84} + \textcolor{#b54708}{1} = \textcolor{#b54708}{55.0564} \]

Check: evaluate

Why: matches the true 7.42²

155. 9.42² expands to 8.42² + 2(8.42) + 1²

Worked example

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge

\[ \textcolor{#1f5fbf}{9.42} = 8.42 + 1 \]

Write 9.42 from 16's distance

Why: the value sits 1 beyond

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 9.42 dashed around it, labelled on its bottom edge

\[ (8.42 + 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)(1)} + \textcolor{#b54708}{1^2} \]

Expand (a + b)²

Why: rule for a sum, b = +1

Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 9.42 dashed around it, labelled on its bottom edge; two blue 8.42-by-1 strips and a solid orange 1-by-1 corner added

\[ (8.42 + 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{1^2} \]

Drop the factor 1

Why: multiplying by 1 changes nothing

\[ \textcolor{#b54708}{70.8964} + \textcolor{#1f5fbf}{16.84} + \textcolor{#b54708}{1} = \textcolor{#b54708}{88.7364} \]

Check: evaluate

Why: matches the true 9.42²

156. The strips cancel, leaving two 8.42² and two 1²

Worked example

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it

\[ \begin{aligned}&\textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 8.42^2 - 2(8.42) + 1^2\\&\qquad + 8.42^2 + 2(8.42) + 1^2\end{aligned} \]

Substitute both

Why: a sum drops brackets

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added

\[ \begin{aligned}&= 8.42^2 + 8.42^2\\&\qquad - 2(8.42) + 2(8.42) + 1^2 + 1^2\end{aligned} \]

Reorder the terms

Why: addition runs in any order

\[ = 8.42^2 + 8.42^2 + 1^2 + 1^2 \]

Cancel −2(8.42) + 2(8.42)

Why: opposites sum to 0

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added

\[ \textcolor{#b54708}{70.8964} + \textcolor{#b54708}{70.8964} + 1 + 1 = \textcolor{#b54708}{143.7928} \]

Check: evaluate it

Why: matches the true pair

157. The error is exactly 2(1²): each value sits 1 from 16

Worked example

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added

\( {\textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 8.42^2 + 8.42^2 + 1^2 + 1^2} \)

\[ \textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 2(8.42^2) + 2(1^2) \]

Collect equal terms

Why: so the midpoint area subtracts whole

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added; labels: left, two 8.42² squares and two 1² corners

\[ \textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 - 2(8.42^2) = 2(1^2) \]

Subtract 2(8.42²) from both sides

Why: removes the midpoint pair's area

Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added; labels: left, two 8.42² squares and two 1² corners; label: error, the two corners

\[ 1^2 = 1 \]

Evaluate one corner, 1²

Why: a unit square of error

\[ \textcolor{#b54708}{143.7928} - \textcolor{#b54708}{141.7928} = 2(1) \]

Check against the true values' sum

Why: 15 and 17 gave 2 directly

158. Example 2.34's grouped SD comes out about 3.50

Worked example

midff(mid − x̄)²
1143.2964
4676.8984
7103.3640
10740.9948
1300
162141.7928
total26

\( {\bar{x} = 197 \div 26 \approx 7.58} \)

\[ \begin{aligned}&43.2964 + 76.8984 + 3.3640\\&+ 40.9948 + 0 + 141.7928 = \textcolor{#b54708}{306.3464}\end{aligned} \]

Add the six table rows

Why: s² is built on this total

\[ 26 - 1 = 25 \]

Find n − 1

Why: 26 values: a sample

\[ 306.3464 \div 25 \approx \textcolor{#b54708}{12.2539} \]

Divide by 25

Why: fair sample average

\[ \textcolor{#1f5fbf}{s} = \sqrt{12.2539} \approx \textcolor{#1f5fbf}{3.50} \]

Take the root

Why: back to table units

\[ 3.50^2 \times 25 \approx 306.25 \]

Check: undo both

Why: near 306.35: rounding

159. Chebyshev sizes the band that holds 96% of fills

Check

Check your understanding

Bottle fills have mean 500 ml, σ = 4 ml, unknown shape. What band around the mean does Chebyshev need to guarantee at least 96%?

  • A. ±20 ml (correct)
  • B. ±5 ml
  • C. ±100 ml
  • D. ±4.08 ml

Answer: A

Why: 1 − 1/k² = 0.96, so 1/k² = 0.04, k² = 25, k = 5, and 5 × 4 = 20 ml.

Why B tempts people
Stopped at k = 5 and never multiplied by σ.
Why C tempts people
Used 1/k = 0.04, so k = 25.
Why D tempts people
Treated 96% as the outside share: k² = 1/0.96.

160. Guaranteeing 96% of fills takes k = 5 SDs

Worked example

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%

\[ 1 - 1/k^2 = 0.96 \]

Set the floor to 96%

Why: the Chebyshev guarantee

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%, with a dashed line at the target 96%

\[ 1 = 0.96 + 1/k^2 \]

Add 1/k² to both sides

Why: frees the cap's sign

\[ 0.04 = 1/k^2 \]

Subtract 0.96 from both sides

Why: leaves the cap alone

\[ k^2 = 1/0.04 \]

Take reciprocals

Why: frees k² from the fraction

\[ k^2 = 25 \]

Evaluate 1 ÷ 0.04

Why: so a root gives k

\[ k = 5 \]

Take the root

Why: k counts SDs, so positive

Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%, with a dashed line at the target 96%; the crossing marked at k = 5

\[ 1 - 1/5^2 = 1 - 0.04 = 0.96 \]

Check: substitute k = 5

Why: gives back 96%

161. A ±20 ml band guarantees at least 96% of fills

Worked example

Chebyshev: 96% needs k = 5 SDs. Fills: mean 500 ml, σ = 4 ml.

\[ 5 \times \textcolor{#1f5fbf}{4} = \textcolor{#1f5fbf}{20} \]

Multiply k by σ

Why: hops become ml

\[ 500 - \textcolor{#1f5fbf}{20} = 480,\ \ 500 + \textcolor{#1f5fbf}{20} = 520 \]

Step 20 ml each way from the mean

Why: the band is centred on it

\[ (520 - 480) \div 2 \div \textcolor{#1f5fbf}{4} = 5 \]

Check: half the band in SDs

Why: gives back k = 5

162. Four changes to line A's waits compete to raise its standard deviation

Check

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

Check your understanding

Population 2, 4, 5, 6, 8 (line A, σ = 2): which change raises σ most?

  • A. Add a sixth wait of 5 minutes
  • B. Change the 4 to 3 and the 6 to 7
  • C. Change the 5 to 6
  • D. Double every wait (correct)

Answer: D

Why: Doubling doubles every deviation: σ = 2 × 2 = 4. 4→3, 6→7: squares 9, 4, 0, 4, 9 total 26; 26 ÷ 5 = 5.2; √5.2 ≈ 2.28. 5→6: mean 26 ÷ 5 = 5.2, squares total 20.8; 20.8 ÷ 5 = 4.16; √4.16 ≈ 2.04. A sixth 5: squares still total 20; 20 ÷ 6 ≈ 3.33; √3.33 ≈ 1.83.

Why A tempts people
The new wait sits on the mean: squares still total 20; 20 ÷ 6 ≈ 3.33; √3.33 ≈ 1.83, lower than 2.
Why B tempts people
Squares 9, 4, 0, 4, 9 total 26; 26 ÷ 5 = 5.2; √5.2 ≈ 2.28: a rise, but less than doubling.
Why C tempts people
Mean 26 ÷ 5 = 5.2, squares total 20.8; 20.8 ÷ 5 = 4.16; √4.16 ≈ 2.04: barely changes.

163. Doubling line A's waits doubles its mean to 10

Worked example

Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{4},\ \textcolor{#1f5fbf}{8},\ \textcolor{#1f5fbf}{10},\ \textcolor{#1f5fbf}{12},\ \textcolor{#1f5fbf}{16} \]

Double each wait

Why: test option D directly

Figure (svg): Dot plot of doubled waits: 4, 8, 10, 12, 16 minutes, no mean marked yet

\[ \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{8} + \textcolor{#1f5fbf}{10} + \textcolor{#1f5fbf}{12} + \textcolor{#1f5fbf}{16} = 50 \]

Add the doubled waits

Why: the mean needs a total

\[ \mu = 50 \div 5 = 10 \]

Divide by N = 5

Why: the centre for new distances

Figure (svg): Doubled waits 4, 8, 10, 12, 16 on a number line with 0 deviation arrows from the mean 10

\[ \textcolor{#1f5fbf}{-6,\ -2,\ 0,\ +2,\ +6} \]

Subtract 10 from each

Why: σ needs distances from the new mean

Figure (svg): Doubled waits 4, 8, 10, 12, 16 on a number line with 5 deviation arrows from the mean 10

\[ -6 - 2 + 0 + 2 + 6 = 0 \]

Check: deviations total zero

Why: confirms the centre 10

164. Doubling line A's waits doubles σ from 2 to 4

Worked example

Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: each distance drawn as a blue side, no squares yet

Doubled waits: μ = 10, distances −6, −2, 0, +2, +6.

\[ \textcolor{#b54708}{36,\ 4,\ 0,\ 4,\ 36} \]

Square each

Why: squares stop the signs cancelling

Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36

\[ \textcolor{#b54708}{36 + 4 + 0 + 4 + 36 = 80} \]

Add the areas

Why: σ² shares the total among 5

Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80

\[ 80 \div 5 = \textcolor{#b54708}{16} \]

Divide by N = 5

Why: σ² is the mean square

Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80; the average square, area 16

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{16} = \textcolor{#1f5fbf}{4} \]

Take the root

Why: back from area to minutes

Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80; the average square, area 16, side 4

\[ 2 \times 2 = \textcolor{#1f5fbf}{4} \]

Check the doubling rule

Why: double distances, double σ

165. Doubling moves line A's 8 to 16: count its SDs

Faded example

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both

Fill in the blanks

Line A's 8 sat z = 1.5 SDs above μ = 5 (σ = 2). Doubled to 16, it sits z = 1.5 SDs above μ = 10 (σ = 4).

Why: (8 − 5) ÷ 2 = 1.5 and (16 − 10) ÷ 4 = 1.5: doubling moves the wait and the mean together and stretches σ by the same factor, so the SD count is unchanged.

166. Doubling keeps the 8 at z = 1.5

Worked example

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both

\[ 8 - 5 = \textcolor{#1f5fbf}{3} \]

Subtract line A's μ

Why: z starts from the gap

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8

\[ \textcolor{#1f5fbf}{3} \div 2 = 1.5 \]

Divide by σ = 2

Why: counts the gap in SDs

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops

\[ 16 - 10 = \textcolor{#1f5fbf}{6} \]

Subtract the doubled μ

Why: z starts from the new gap

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops; an arrow of 6 from μ = 10 to 16

\[ \textcolor{#1f5fbf}{6} \div 4 = 1.5 \]

Divide by σ = 4

Why: σ doubled, so same count

Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops; an arrow of 6 from μ = 10 to 16; SD hops of 4 from 10: one full and one half, 1.5 hops

\[ 5 + 1.5 \times 2 = 8,\ \ 10 + 1.5 \times 4 = 16 \]

Check: hop out from each μ

Why: 1.5 hops rebuild both waits

167. Changing the 5 to a 6 moves μ to 5.2

Worked example

Figure (svg): Dot plot of waits with the 5 changed to 6: 2, 4, 6, 6, 8 minutes, no mean marked yet

\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{8} = 26 \]

Add the new waits

Why: μ needs a fresh total

\[ \mu = 26 \div 5 = 5.2 \]

Divide by N = 5

Why: a mean is total over count

Figure (svg): Waits 2, 4, 6, 6, 8 on a number line with 0 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-3.2,\ -1.2,\ +0.8,\ +0.8,\ +2.8} \]

Subtract 5.2 from each

Why: every distance moves, not one

Figure (svg): Waits 2, 4, 6, 6, 8 on a number line with 5 deviation arrows from the mean

\[ -3.2 - 1.2 + 0.8 + 0.8 + 2.8 = 0 \]

Check: deviations total zero

Why: confirms the centre 5.2

168. Changing the 5 to a 6 barely moves σ: 2.04

Worked example

Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: each distance drawn as a blue side, no squares yet

New waits 2, 4, 6, 6, 8: μ = 5.2, distances −3.2, −1.2, +0.8, +0.8, +2.8.

\[ \textcolor{#b54708}{10.24,\ 1.44,\ 0.64,\ 0.64,\ 7.84} \]

Square each

Why: signs can no longer cancel

Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84

\[ \textcolor{#b54708}{10.24 + 1.44 + 0.64 + 0.64 + 7.84 = 20.8} \]

Add the areas

Why: σ² needs this total

Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8

\[ 20.8 \div 5 = \textcolor{#b54708}{4.16} \]

Divide by N = 5

Why: σ² is the mean square

Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8; the average square, area 4.16

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4.16} \approx \textcolor{#1f5fbf}{2.04} \]

Take the root

Why: back to minutes

Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8; the average square, area 4.16, side 2.04

\[ \textcolor{#1f5fbf}{2.04}^2 \times 5 \approx 20.81 \]

Check: square, times N

Why: near 20.8: rounding

169. A sixth wait of 5 leaves μ at 5

Worked example

Figure (svg): Dot plot of line a with a sixth wait of 5: 2, 4, 5, 6, 8, 5 minutes, no mean marked yet

\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{5} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{8} + \textcolor{#1f5fbf}{5} = 30 \]

Add all six waits

Why: μ must count the newcomer

\[ \mu = 30 \div 6 = 5 \]

Divide by N = 6

Why: a mean shares the total equally

Figure (svg): Waits 2, 4, 5, 6, 8, 5 on a number line with 0 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-3,\ -1,\ 0,\ +1,\ +3,\ 0} \]

Subtract 5 from each

Why: the newcomer sits on μ

Figure (svg): Waits 2, 4, 5, 6, 8, 5 on a number line with 6 deviation arrows from the mean

\[ -3 - 1 + 0 + 1 + 3 + 0 = 0 \]

Check: deviations total zero

Why: confirms the centre 5

170. A sixth wait at the mean lowers σ to 1.83

Worked example

Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: each distance drawn as a blue side, no squares yet

Six waits: μ = 5, distances −3, −1, 0, +1, +3, 0.

\[ \textcolor{#b54708}{9,\ 1,\ 0,\ 1,\ 9,\ 0} \]

Square each

Why: the newcomer adds no area

Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0

\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9 + 0 = 20} \]

Add the areas

Why: σ² needs the total squared distance

Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20

\[ 20 \div 6 \approx \textcolor{#b54708}{3.333} \]

Divide by N = 6

Why: more waits share it

Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20; the average square, area 3.333

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{3.333} \approx \textcolor{#1f5fbf}{1.83} \]

Take the root

Why: back to minutes

Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20; the average square, area 3.333, side 1.83

\[ \textcolor{#1f5fbf}{1.83}^2 \times 6 \approx 20.09 \]

Check: square, times N

Why: near 20: rounding

171. Moving the 4 and 6 outward keeps μ at 5

Worked example

Figure (svg): Dot plot of waits with the 4 and 6 moved out: 2, 3, 5, 7, 8 minutes, no mean marked yet

\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{3} + \textcolor{#1f5fbf}{5} + \textcolor{#1f5fbf}{7} + \textcolor{#1f5fbf}{8} = 25 \]

Add the waits

Why: μ needs a total

\[ \mu = 25 \div 5 = 5 \]

Divide by N = 5

Why: total over count

Figure (svg): Waits 2, 3, 5, 7, 8 on a number line with 0 deviation arrows from the mean

\[ \textcolor{#1f5fbf}{-3,\ -2,\ 0,\ +2,\ +3} \]

Subtract 5

Why: spread starts at μ

Figure (svg): Waits 2, 3, 5, 7, 8 on a number line with 5 deviation arrows from the mean

\[ -3 - 2 + 0 + 2 + 3 = 0 \]

Check: deviations total zero

Why: confirms the centre 5

172. Moving the 4 and 6 outward raises σ to 2.28

Worked example

Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: each distance drawn as a blue side, no squares yet

New waits 2, 3, 5, 7, 8: μ = 5, distances −3, −2, 0, +2, +3.

\[ \textcolor{#b54708}{9,\ 4,\ 0,\ 4,\ 9} \]

Square each

Why: areas build σ²

Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9

\[ \textcolor{#b54708}{9 + 4 + 0 + 4 + 9 = 26} \]

Add the areas

Why: σ² needs this total

Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26

\[ 26 \div 5 = \textcolor{#b54708}{5.2} \]

Divide by N = 5

Why: σ² is the mean square

Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26; the average square, area 5.2

\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{5.2} \approx \textcolor{#1f5fbf}{2.28} \]

Take the root

Why: back to minutes

Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26; the average square, area 5.2, side 2.28

\[ \textcolor{#1f5fbf}{2.28}^2 \times 5 \approx 25.99 \]

Check: square, times N

Why: near 26: rounding

173. Doubling every wait spreads line A most: σ = 4

Worked example

Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4

\[ \textcolor{#1f5fbf}{4} > \textcolor{#1f5fbf}{2.28} > \textcolor{#1f5fbf}{2.04} > \textcolor{#1f5fbf}{1.83} \]

Rank the four σ

Why: answers which change spreads most

Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4; sorted largest first and numbered

\[ \textcolor{#1f5fbf}{2.04} > \textcolor{#1f5fbf}{2} > \textcolor{#1f5fbf}{1.83} \]

Place line A's σ = 2 in the ranking

Why: shows the one change that lowers σ

Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4; sorted largest first and numbered; grey dashed line at line A's σ = 2

\[ \textcolor{#b54708}{16} > \textcolor{#b54708}{5.2} > \textcolor{#b54708}{4.16} > \textcolor{#b54708}{3.333} \]

Check: rank the σ² values

Why: roots keep the order

174. As samples, line A's s is 2.24 and line B's 4.47

Worked example

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5

Treat each line's five waits as a sample.

\[ \textcolor{#b54708}{\text{A: } 20,\ \ \text{B: } 80} \]

Take both squared totals

Why: s² starts from these

\[ 5 - 1 = 4 \]

Find n − 1

Why: five waits, a sample

\[ 20 \div 4 = \textcolor{#b54708}{5},\ \ 80 \div 4 = \textcolor{#b54708}{20} \]

Divide each by 4

Why: a sample's fair average

\[ \textcolor{#1f5fbf}{s_A} = \sqrt{5} \approx \textcolor{#1f5fbf}{2.24},\ \ \textcolor{#1f5fbf}{s_B} = \sqrt{20} \approx \textcolor{#1f5fbf}{4.47} \]

Take both roots

Why: back from area to minutes

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B

\[ 2.24^2 \times 4 \approx 20.07,\ \ 4.47^2 \times 4 \approx 79.92 \]

Check: square, times 4

Why: near both: rounding

175. 7 minutes is 0.89 SD out on line A, 0.45 on B

Worked example

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B

\[ 7 - 5 = 2 \]

Find 7's distance

Why: what hops cover

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes

\[ k_A = 2 \div \textcolor{#1f5fbf}{2.24} \approx 0.89 \]

Divide by A's s

Why: SDs give one ruler

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick

\[ k_B = 2 \div \textcolor{#1f5fbf}{4.47} \approx 0.45 \]

Divide by B's s

Why: so both lines compare fairly

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick; an arc from 5 to 7 on B: 0.45 SD, under half of B's first tick

\[ 0.45 < 0.89 \]

Compare the hop counts

Why: B's ordinary spread reaches 7

\[ \text{A } 1/5,\ \text{B } 2/5 \]

Count waits over 7

Why: data test the hops

Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick; an arc from 5 to 7 on B: 0.45 SD, under half of B's first tick; waits over 7 circled: one on A, two on B

\[ 5 + 0.89(\textcolor{#1f5fbf}{2.24}) \approx 7.0,\ \ 5 + 0.45(\textcolor{#1f5fbf}{4.47}) \approx 7.0 \]

Check: hop back on each

Why: both land near 7

Join line A.

176. You can now measure, correct and compare spread

Recap

OpenStax Introductory Statistics 2e, §2.7 Measures of the Spread of the Data §2.7, pp. 107-118 — Examples 2.32–2.35 and the Try Its trace back here

Sources

  1. OpenStax Introductory Statistics 2e, §2.7 Measures of the Spread of the Data — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0, pp. 107-117
  2. OpenStax Introductory Business Statistics 2e, §2.7 Measures of the Spread of the Data — Illowsky & Dean, OpenStax / Rice University, CC BY 4.0

Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108