Build variance and standard deviation from squares you can see, prove why a sample divides by n − 1, derive Chebyshev's inequality, and compare data sets with z-scores.
Subject: Statistics · 176 slides · applied lesson
Open the interactive version of this deck
Title
Statistics · §2.7
Variance, SD, the sample divisor and z-scores, each shown to be true
Objectives
Concept
\[ {\textstyle\sum} x = x_1 + \cdots + x_n \]
Σ means add up
Why: one term per value
\[ \bar{x} = \frac{{\textstyle\sum} x}{n},\ \ \mu = \frac{{\textstyle\sum} x}{N} \]
Means (§2.5)
Why: n or N values share the total
\[ {\textstyle\sum} (u - v) = {\textstyle\sum} u - {\textstyle\sum} v \]
Sums of differences
Why: addition can be regrouped in any order
\[ \text{avg}(u \pm v) = \text{avg}\,u \pm \text{avg}\,v \]
Averages add, subtract
Why: divide both split sums by the same n
\[ {\textstyle\sum} c\,u = c{\textstyle\sum} u,\ \ \text{avg}(c\,u) = c\,\text{avg}\,u \]
Factor out c
Why: every term shares c
\[ {\textstyle\sum}_{1}^{n} c = n\,c \]
Add c n times
Why: n equal copies
Concept
\[ \sqrt{a^2} = a\ \ (a \ge 0) \]
A root undoes a square
Why: area back to side
\[ \sqrt{ab} = \sqrt{a}\sqrt{b}\ \ (a, b \ge 0) \]
A root splits over a product
Why: √a·√b squared gives back ab
\[ (a + b)^2 = a^2 + 2ab + b^2 \]
Square of a sum
Why: four pieces of a square
\[ u^2 \ge 0,\ \ |u|^2 = u^2 \]
Squares ignore sign
Why: signs cancel in pairs
\[ (c\,d)^2 = c^2 d^2,\ \Big(\frac{u}{v}\Big)^2 = \frac{u^2}{v^2} \]
Square each factor
Why: factors can regroup
Concept
\[ c > 0:\ u \le v \Rightarrow c\,u \le c\,v,\ \frac{u}{c} \le \frac{v}{c} \]
Scale by c > 0
Why: order is kept
\[ 0 \le u \le v \Rightarrow u^2 \le v^2 \]
Square non-negatives
Why: bigger side, bigger square
\[ u \le v \Rightarrow 1 - u \ge 1 - v \]
Subtract from 1
Why: the order flips
\[ u_i \ge c \text{ for } m \text{ terms} \Rightarrow {\textstyle\sum} u_i \ge m\,c \]
Add m inequalities
Why: sums of ≥ keep ≥
\[ u \le v \Rightarrow u + c \le v + c \]
Add to both sides
Why: equal additions leave the gap unchanged
Concept
\[ 0 < u \le v \Rightarrow \tfrac{1}{u} \ge \tfrac{1}{v} \]
Flip reciprocals of positives
Why: divide 1 by more, get less
\[ 0 \le u \le v \Rightarrow \sqrt{u} \le \sqrt{v} \]
Roots keep order
Why: a bigger area has a longer side
\[ b \ge 0 \Rightarrow a \le a + b \]
Adding a non-negative can only grow
Why: nothing is taken away
\[ u \le v,\ v \le w \Rightarrow u \le w \]
Chain inequalities
Why: a bound passes along
\[ (-u)(-v) = uv,\ \ (-u)v = -uv \]
Sign rule
Why: equal signs multiply to positive
\[ 1,\ 3,\ \mathbf{7},\ 8,\ 9 \]
Median: middle of the ordered list
Why: half lie on each side
Prediction
Figure (svg): Dot plots of line A's waits 2, 4, 5, 6, 8 and line B's waits 0, 2, 4, 8, 11 minutes on axes from 0 to 12, both with the mean marked at 5
Predict first
Line A: 2, 4, 5, 6, 8 min. Line B: 0, 2, 4, 8, 11 min.
Your train leaves in 7 minutes. Which line?
Correct: Line A
Why: Equal means hide different spreads. B's waits stray further from 5, which suggests more risk; the evidence today is that two of B's five waits pass 7, against one of A's.
Worked example
Figure (svg): Dot plots of line A's waits 2, 4, 5, 6, 8 and line B's waits 0, 2, 4, 8, 11 minutes on axes from 0 to 12, both with the mean marked at 5
\[ \text{A: } \textcolor{#1f5fbf}{8} > 7 \]
Count A's waits over 7
Why: evidence for the riskier line
\[ \text{B: } \textcolor{#1f5fbf}{8},\ \textcolor{#1f5fbf}{11} > 7 \]
Count B's waits over 7
Why: to compare with A's one
\[ 2 + 4 + 5 + 6 + 8 = 0 + 2 + 4 + 8 + 11 = 25 \]
Total each line's waits
Why: equal totals give equal means
\[ 25 \div 5 = 5 \]
Check: share each total among 5 waits
Why: both means are 5
Needed: a number for each line's spread
Section
Idea 1 of 5
Concept
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 0 deviation arrows from the mean
Goal: one number for how far these waits sit from 5.
Discussion prompt
What would you measure for each of the five waits?
Answer:
Its distance from the mean: 2 is 3 below, 8 is 3 above.
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 0 deviation arrows from the mean
deviation = value − mean (signed)
\[ \textcolor{#1f5fbf}{2} - 5 = \textcolor{#1f5fbf}{-3} \]
Subtract the mean from 2
Why: below 5, so negative
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 1 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{4} - 5 = \textcolor{#1f5fbf}{-1} \]
Subtract the mean from 4
Why: the sign marks the side
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 2 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{5} - 5 = \textcolor{#1f5fbf}{0} \]
Subtract the mean from 5
Why: a wait at the mean has no distance
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 3 deviation arrows from the mean
\[ 5 + (-3) = 2,\ 5 + (-1) = 4,\ 5 + 0 = 5 \]
Check: hop back from the mean
Why: each lands on its wait
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 3 deviation arrows from the mean
\( {{2 - 5 = -3}},\quad\allowbreak \allowbreak {{4 - 5 = -1}},\quad\allowbreak \allowbreak {{5 - 5 = 0}} \)
\[ \textcolor{#1f5fbf}{6} - 5 = \textcolor{#1f5fbf}{+1} \]
Subtract the mean from 6
Why: above, so positive
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 4 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{8} - 5 = \textcolor{#1f5fbf}{+3} \]
Subtract the mean from 8
Why: largest wait, longest arrow
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-3 - 1 + 0 + 1 + 3} = 0 \]
Add the distances
Why: spread should show up
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ 0 \div 5 = 0 \]
Divide by 5
Why: should give a typical distance
\[ (-3 + 3) + (-1 + 1) + 0 = 0 \]
Check: pair the arrows
Why: each one cancels
Prediction
Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 0 deviation arrows from the mean
Predict first
Line B: 0, 2, 4, 8, 11 minutes. Mean 5.
Which is bigger: how far B's low waits fall short of 5 in total, or how far its high waits overshoot?
Correct: They are equal
Why: Below: 5 + 3 + 1 = 9 minutes short. Above: 3 + 6 = 9 minutes over. Equal, so B's signed distances total zero as line A's did.
Worked example
Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 0 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-5,\ -3,\ -1,\ +3,\ +6} \]
Subtract 5 from each wait
Why: to test whether B's arrows cancel too
Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean
\[ -5 - 3 - 1 = -9 \]
Add the three below
Why: the shortfall to compare
\[ +3 + 6 = +9 \]
Add the two above
Why: the overshoot to compare
\[ -9 + 9 = 0 \]
Combine the sides
Why: opposite amounts of equal size cancel
Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean
\[ (-5 + 6) + (-3 + 3) - 1 = 0 \]
Check: regroup the arrows
Why: any grouping gives one total
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ {\textstyle\sum} (x - \bar{x}) = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{{\textstyle\sum} \bar{x}} \]
Split the sum
Why: differences add as totals
Figure (svg): Line A's waits 2, 4, 5, 6, 8 laid end to end reach 25; five copies of the mean 5 laid end to end also reach 25
\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{n\bar{x}} \]
Count the x̄'s
Why: one per value: n of them
Figure (svg): Line A's waits 2, 4, 5, 6, 8 laid end to end reach 25; five copies of the mean 5 laid end to end also reach 25
\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#6b7280}{n \cdot \frac{{\textstyle\sum} x}{n}} \]
Substitute x̄'s definition
Why: the prerequisite mean
Figure (svg): Line A's waits laid end to end reach 25; below them five copies of the mean, each now labelled Σx/n, also reach 25
\[ = \textcolor{#1f5fbf}{{\textstyle\sum} x} - \textcolor{#1f5fbf}{{\textstyle\sum} x} \]
Cancel n against n
Why: multiplying undoes dividing, leaving Σx
Figure (svg): Line A's waits laid end to end reach 25; below them the five copies of Σx/n have merged into one bar Σx = 25
\[ = 0 \]
Cancel the two Σx
Why: equal totals leave nothing over
Figure (svg): Line A's waits laid end to end and the merged bar Σx both end at 25, marked difference 0
\[ \textcolor{#1f5fbf}{25} - \textcolor{#6b7280}{5 \cdot 5} = 0 \]
Check with line A
Why: total 25, five 5s
Concept
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-3 - 1} = \textcolor{#1f5fbf}{-4},\ \ \textcolor{#1f5fbf}{+1 + 3} = \textcolor{#1f5fbf}{+4} \]
Total each side's arrows
Why: to test the balance
\[ \text{waits } 5, 5, 5, 5, 5:\ \mu = 5 \]
Try no spread at all
Why: spread should read zero
Figure (svg): Five waits all equal to 5 stacked on a number line from 0 to 10, every one at the mean
\[ 5 - 5 = \textcolor{#1f5fbf}{0} \text{ for each wait} \]
Subtract the mean
Why: each value sits on it
\[ 0 + 0 + 0 + 0 + 0 = 0 \]
Add them
Why: signed sums miss spread
\[ (5 + 5 + 5 + 5 + 5) - 5 \times 5 = 0 \]
Check with Σx − Nμ
Why: a true mean leaves zero
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ -3 - 1 + 0 + 1 + \textcolor{#1f5fbf}{d} = 0 \]
Hide the last arrow as d
Why: the total must still be zero
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean, the last one grey and marked forced
\[ -3 + \textcolor{#1f5fbf}{d} = 0 \]
Add the four known distances
Why: leaves d as the only unknown
\[ \textcolor{#1f5fbf}{d} = +3 \]
Add 3 to both sides
Why: isolates the hidden distance
\[ 8 - 5 = +3 \]
Check against the hidden wait 8
Why: its arrow is 3 long
Trap
\[ -5 - 3 - 1 + 3 + 6 = 0 \]
Add the signed distances
Why: hoping for a total spread
\[ 0 \div 5 = 0 \]
Divide by 5
Why: treated as the typical distance
\[ \text{spread} = 0 \]
Report zero spread
Why: fails here: the waits differ
\[ 5 + 3 + 1 + 3 + 6 = 18 \]
Add the distances, ignoring sign
Why: unsigned sizes cannot cancel
\[ 18 \div 5 = 3.6 \]
Divide by 5
Why: a typical distance, not zero
\[ 3.6 \times 5 = 18 \]
Check: undo the divide
Why: rebuilds the total distance
Idea 2 compares this with squaring.
Prediction
Predict first
Four quiz scores: three sit −2, +5, +1 from the mean. One is lost.
How far from the mean is the lost score?
Correct: 4 points below
Why: All four deviations must total zero. The three known ones total +4, so the lost one is −4: four points below.
Worked example
\[ -2 + 5 + 1 + d = 0 \]
Call the lost deviation d
Why: all four must total zero
\[ 4 + d = 0 \]
Add the three known deviations
Why: what is still unbalanced
\[ d = -4 \]
Subtract 4 from both sides
Why: leaves the lost deviation alone
\[ -2 + 5 + 1 - 4 = 0 \]
Check the four balance
Why: zero, as it must be
Prediction
Predict first
Four runners' lap times: 61, 66, 70, 83 seconds. Their mean is 70 seconds.
Which list is their deviations from the mean?
Correct: −9, −4, 0, +13
Why: Time minus mean gives negatives below 70, and the list must total zero. Only the first list does both.
Worked example
Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 0 deviation arrows from the mean 70
\[ \textcolor{#1f5fbf}{61} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{-9} \]
Subtract the mean from 61
Why: below, so negative
Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 1 deviation arrows from the mean 70
\[ \textcolor{#1f5fbf}{66} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{-4} \]
Subtract it from 66
Why: below, so negative
Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 2 deviation arrows from the mean 70
\[ \textcolor{#1f5fbf}{70} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{0},\ \ \textcolor{#1f5fbf}{83} - \textcolor{#6b7280}{70} = \textcolor{#1f5fbf}{+13} \]
Subtract it from 70 and 83
Why: the sign marks the side
Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 4 deviation arrows from the mean 70
\[ \textcolor{#1f5fbf}{-9 - 4 + 0 + 13} = 0 \]
Check the total
Why: any list that fails is wrong
Figure (svg): Lap times 61, 66, 70, 83 seconds on a number line with 4 deviation arrows from the mean 70
Section
Idea 2 of 5
Concept
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with arrows from the mean 5 labelled by unsigned distance 5, 3, 1, 3, 6, distances total 18
Unsigned distances from line B's mean 5 total 18: an average of 3.6 minutes.
Discussion prompt
Measured from 4 instead of the mean 5, is line B's unsigned-distance total bigger or smaller?
Answer:
Smaller: 4 + 2 + 0 + 4 + 7 = 17. The next slides test squares the same way.
Concept
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\( {{5 + 3 + 1 + 3 + 6 = 18}},\quad\allowbreak \allowbreak {{18 \div 5 = 3.6}} \)
\[ (\textcolor{#1f5fbf}{-5})^2, (\textcolor{#1f5fbf}{-3})^2, (\textcolor{#1f5fbf}{-1})^2, \textcolor{#1f5fbf}{3}^2, \textcolor{#1f5fbf}{6}^2 = \textcolor{#b54708}{25,\ 9,\ 1,\ 9,\ 36} \]
Square each deviation
Why: no negatives survive
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#b54708}{25 + 9 + 1 + 9 + 36} = \textcolor{#b54708}{80} \]
Add the squares
Why: none can cancel either
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#b54708}{80} \div 5 = \textcolor{#b54708}{16} \]
Divide by 5
Why: a typical square
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \sqrt{\textcolor{#b54708}{16}} = \textcolor{#1f5fbf}{4} \]
Take the root
Why: so it compares with the waits
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#1f5fbf}{4}^2 \times 5 = \textcolor{#b54708}{80} \]
Check: undo both
Why: rebuilds the total area
Worked example
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\( {\text{about } 5\text{: } 25 + 9 + 1 + 9 + 36 = 80} \)
\[ \textcolor{#1f5fbf}{4,\ 2,\ 0,\ 4,\ 7;\ \ 6,\ 4,\ 2,\ 2,\ 5} \]
Measure from 4 and 6
Why: rival centres near 5
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ \textcolor{#b54708}{16,\ 4,\ 0,\ 16,\ 49;\ \ 36,\ 16,\ 4,\ 4,\ 25} \]
Square each distance
Why: as candidate 2 did
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ \textcolor{#b54708}{16 + 4 + 0 + 16 + 49} = \textcolor{#b54708}{85} \]
Add squares from 4
Why: rival total beside 80
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ \textcolor{#b54708}{36 + 16 + 4 + 4 + 25} = \textcolor{#b54708}{85} \]
Add squares from 6
Why: tests the centre above 5
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ (36 + 4) + (16 + 4) + 25 = \textcolor{#b54708}{85} \]
Check by regrouping
Why: addition ignores order
Worked example
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\( {4, 2, 0, 4, 7;\ 6, 4, 2, 2, 5},\quad\allowbreak \allowbreak {5 + 3 + 1 + 3 + 6 = 18},\quad\allowbreak \allowbreak {85,\ 80,\ 85} \)
\[ \textcolor{#1f5fbf}{4 + 2 + 0 + 4 + 7} = \textcolor{#1f5fbf}{17} \]
Add distances from 4
Why: the median's own total
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ \textcolor{#1f5fbf}{6 + 4 + 2 + 2 + 5} = \textcolor{#1f5fbf}{19} \]
Add distances from 6
Why: checks the centre above too
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ \textcolor{#b54708}{80} < \textcolor{#b54708}{85},\ \ \textcolor{#1f5fbf}{17} < \textcolor{#1f5fbf}{18} \]
Compare least totals
Why: squares pair with the mean
Figure (svg): Line B's waits 0, 2, 4, 8, 11 on a number line with trial centres 4, 5 and 6 marked; for each centre, each wait's distance and squared distance, with totals
\[ 17 + 3 - 2 = \textcolor{#1f5fbf}{18} \]
Check: centre 4 → 5
Why: three gain 1, two lose 1
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ (\textcolor{#1f5fbf}{-3})^2 = \textcolor{#b54708}{9} \]
Square −3
Why: an area can't be negative
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ (\textcolor{#1f5fbf}{-1})^2 = \textcolor{#b54708}{1},\ \ \textcolor{#1f5fbf}{0}^2 = \textcolor{#b54708}{0} \]
Square −1 and 0
Why: small distances make small areas
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#1f5fbf}{1}^2 = \textcolor{#b54708}{1},\ \ \textcolor{#1f5fbf}{3}^2 = \textcolor{#b54708}{9} \]
Square +1 and +3
Why: opposite signs give equal areas
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9} = \textcolor{#b54708}{20} \]
Add the areas
Why: nothing cancels now
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ (9 + 9) + (1 + 1) + 0 = 20 \]
Check by pairing the squares
Why: addition ignores order
Worked example
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\( {{(-3)^2 = 9}},\quad\allowbreak \allowbreak {{(-1)^2 = 1}},\quad\allowbreak \allowbreak {{0^2 = 0}},\quad\allowbreak \allowbreak {{1^2 = 1}},\quad\allowbreak \allowbreak {{3^2 = 9}},\quad\allowbreak \allowbreak {{9 + 1 + 0 + 1 + 9 = 20}} \)
\[ \textcolor{#b54708}{20} \div 5 = \textcolor{#b54708}{4} \]
Share 20 among 5 squares
Why: a typical area, not a growing total
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \sqrt{\textcolor{#b54708}{4}} = \textcolor{#1f5fbf}{2} \]
Take its side
Why: back from area to minutes
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ 5 \times \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{20} \]
Check: five average squares
Why: rebuild the total area
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{x - \mu} \]
Take one deviation
Why: each square's side
\[ \textcolor{#b54708}{(x - \mu)^2} \]
Square it
Why: so its sign cannot cancel
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} \]
Add all N areas
Why: every value counts
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#b54708}{\sigma^2} = \frac{\textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2}}{N} \]
Divide by N
Why: a typical square, not a total
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{\textcolor{#b54708}{\sigma^2}} \]
Take the root
Why: back to data units
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
σ² is the population variance; σ, its root, is the standard deviation (SD).
\[ \sqrt{20 \div 5} = \textcolor{#1f5fbf}{2} \]
Check with line A
Why: matches the last slide
Concept
Figure (svg): Waits 0, 2, 4, 8, 11 on a number line with 5 deviation arrows from the mean
\[ \textcolor{#b54708}{25,\ 9,\ 1,\ 9,\ 36} \]
Square B's five deviations
Why: none can cancel
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#b54708}{25 + 9 + 1 + 9 + 36} = \textcolor{#b54708}{80} \]
Add B's areas
Why: σ² shares it among 5
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#b54708}{36} \div \textcolor{#b54708}{80} = 0.45 \]
Divide +6's square by the total
Why: part over whole gives its share
\[ 0.45 = 45\% \]
Write the share as a percent
Why: percents compare shares at a glance
\[ 0.45 \times \textcolor{#b54708}{80} = \textcolor{#b54708}{36} \]
Check: undo the division
Why: back to the far wait's square
Concept
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
Discussion prompt
Doubling every distance: does the side-1 square or the side-6 square grow more?
Answer:
Side 6: 36 → 144, up 108; side 1: 1 → 4.
\[ \textcolor{#1f5fbf}{d} = x - \mu \]
Name each distance d
Why: for any side
\[ (2\textcolor{#1f5fbf}{d})^2 = 2^2 \textcolor{#1f5fbf}{d}^2 \]
Square 2d
Why: each factor squares
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ = 4\textcolor{#b54708}{d^2} \]
Evaluate 2²
Why: four old squares fit inside
\[ (2 \times 6)^2 = 144 = 4 \times 36 \]
Check side 6
Why: the 11-minute wait
Concept
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#1f5fbf}{-10,\ -6,\ -2,\ +6,\ +12} \]
Double B's distances
Why: sides of the new squares
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{100,\ 36,\ 4,\ 36,\ 144} \]
Square each
Why: σ² averages these areas
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{100 + 36 + 4 + 36 + 144} = \textcolor{#b54708}{320} \]
Add the squares
Why: σ² shares a total
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = 320 \div 5 = \textcolor{#b54708}{64} \]
Divide by N = 5
Why: σ² is the mean square
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{80} \div 5 = \textcolor{#b54708}{16} \]
Divide B's old total by 5
Why: to test the ×4 claim
\[ 4 \times \textcolor{#b54708}{16} = \textcolor{#b54708}{64} \]
Check: 4 × old σ²
Why: the ×4 claim holds
Concept
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\,\textcolor{#b54708}{d^2} \]
Write σ² using d = x − μ
Why: the average square
Figure (svg): Squares on line B's deviations −5, −3, −1, +3, +6, areas 25, 9, 1, 9, 36
\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = \text{avg}\,(2\textcolor{#1f5fbf}{d})^2 \]
Replace each d by 2d
Why: doubling data doubles every distance
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ = \text{avg}\,4\textcolor{#b54708}{d^2} \]
Use (2d)² = 4d²
Why: proved for any single distance earlier
\[ = 4\,\text{avg}\,\textcolor{#b54708}{d^2} \]
Pull out 4
Why: constants leave averages
\[ = 4\textcolor{#b54708}{\sigma^2} \]
Recognise σ²
Why: the old average square
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\[ \textcolor{#b54708}{64} = 4 \times \textcolor{#b54708}{16} \]
Check with line B's pictures
Why: average squares 64 and 16
Concept
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
\( {\sigma_{\text{new}}^2 = 4\sigma^2},\quad\allowbreak {\text{line B: } \sigma^2 = 16,\ \sigma_{\text{new}}^2 = 64} \)
\[ \textcolor{#b54708}{\sigma_{\text{new}}^2} = 2^2\textcolor{#b54708}{\sigma^2} \]
Write 4 as 2²
Why: aiming for a square
\[ = (2\textcolor{#1f5fbf}{\sigma})^2 \]
Combine the squares
Why: so a root undoes it
\[ \textcolor{#1f5fbf}{\sigma_{\text{new}}} = 2\textcolor{#1f5fbf}{\sigma} \]
Root both sides
Why: roots undo squares; 2σ ≥ 0
\[ \sqrt{64} = 8 = 2\sqrt{16} \]
Check with line B
Why: the doubling rule holds there
Figure (svg): Squares on line B's doubled deviations −10, −6, −2, +6, +12, areas 100, 36, 4, 36, 144
Concept
Figure (svg): Five waits all equal to 5 stacked on a number line from 0 to 10, every one at the mean
\[ \textcolor{#1f5fbf}{x - \mu} = 0 \text{ for every } x \]
Let every value equal μ
Why: spread should read zero
\[ \textcolor{#b54708}{(x - \mu)^2} = 0 \]
Square each deviation
Why: zero squared is zero
\[ {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} = 0 \]
Add the N squares
Why: zeros add to zero
\[ \textcolor{#b54708}{\sigma^2} = 0 \div N = 0 \]
Divide by N
Why: σ² is the mean square
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{0} = 0 \]
Take the root
Why: zero area, zero side
\[ (0+0+0+0+0) \div 5 = 0 \]
Check with five 5s
Why: real data agree
Prediction
Predict first
A slow cashier adds 3 minutes to every line A wait: 5, 7, 8, 9, 11.
Line A's σ was 2. What is it now?
Correct: 2 minutes
Why: The mean moves up 3 as well, so no deviation changes: same squares, same σ.
Worked example
Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line, mean not yet found
\[ \textcolor{#1f5fbf}{5 + 7 + 8 + 9 + 11} = 40 \]
Add the new waits
Why: a total for the mean
\[ \mu = 40 \div 5 = 8 \]
Divide by N = 5
Why: a mean shares the total
Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 0 deviation arrows from it
\[ \textcolor{#1f5fbf}{-3,\ -1,\ 0,\ +1,\ +3} \]
Subtract 8 from each
Why: distances run from the new mean
Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 5 deviation arrows from it
\[ -3 - 1 + 0 + 1 + 3 = 0 \]
Check the balance
Why: deviations from a true mean total zero
Worked example
Figure (svg): The shifted waits 5, 7, 8, 9, 11 on a number line with the mean 8 marked and 5 deviation arrows from it
\( {5 + 7 + 8 + 9 + 11 = 40}\;\;\Rightarrow\;\;\allowbreak {\mu = 40 \div 5 = 8}\;\;\Rightarrow\;\;\allowbreak {-3,\ -1,\ 0,\ +1,\ +3} \)
\[ \textcolor{#b54708}{9,\ 1,\ 0,\ 1,\ 9} \]
Square each
Why: variance averages squares
Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9
\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9 = 20} \]
Add the areas
Why: σ² shares this total
Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9
\[ \textcolor{#b54708}{\sigma^2} = 20 \div 5 = \textcolor{#b54708}{4} \]
Divide by N = 5
Why: σ² is the mean square
Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4} = \textcolor{#1f5fbf}{2} \]
Take the root
Why: an SD in minutes, like the waits
Figure (svg): Squares on the shifted waits' deviations −3, −1, 0, +1, +3, areas 9, 1, 0, 1, 9
\[ 5 \times \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{20} \]
Check: five average squares
Why: rebuild the shifted total
Worked example
Figure (svg): Ages 4, 7, 9, 12 on a number line with 0 deviation arrows from the mean 8
\[ \textcolor{#1f5fbf}{4 + 7 + 9 + 12} = 32 \]
Add the ages
Why: the mean needs the total
\[ \mu = 32 \div 4 = 8 \]
Divide by N = 4
Why: a mean is total over count
\[ \textcolor{#1f5fbf}{-4,\ -1,\ +1,\ +4} \]
Subtract 8 from each age
Why: spread starts from each age's distance
Figure (svg): Ages 4, 7, 9, 12 on a number line with 4 deviation arrows from the mean 8
\[ -4 - 1 + 1 + 4 = 0 \]
Check the deviations balance
Why: Idea 1 says they must
Worked example
Figure (svg): Ages 4, 7, 9, 12 on a number line with 4 deviation arrows from the mean 8
\[ \textcolor{#b54708}{16,\ 1,\ 1,\ 16} \]
Square −4, −1, +1, +4
Why: areas cannot cancel
Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16
\[ \textcolor{#b54708}{16 + 1 + 1 + 16 = 34} \]
Add the areas
Why: the total the four ages share
Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16
\[ \textcolor{#b54708}{\sigma^2} = 34 \div 4 = \textcolor{#b54708}{8.5} \]
Divide by N = 4
Why: σ² is the mean square
Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{8.5} \approx \textcolor{#1f5fbf}{2.92} \]
Take the root
Why: back to years
Figure (svg): Squares on the family's age deviations −4, −1, +1, +4, areas 16, 1, 1, 16
\[ 4 \times \textcolor{#1f5fbf}{2.92}^2 \approx \textcolor{#b54708}{34.1} \]
Check by undoing both
Why: restores the total 34
Trap
\[ (-4 - 1 + 1 + 4)^2 \]
Write the square of the sum
Why: tempting: it looks like squaring deviations
\[ = 0^2 \]
Add inside the bracket
Why: brackets come first, so signs cancel
\[ = 0 \]
Square the zero
Why: fails: no spread left
\[ (-4)^2, (-1)^2, 1^2, 4^2 = 16, 1, 1, 16 \]
Square each deviation first
Why: the bracket closes on each one
\[ 16 + 1 + 1 + 16 = 34 \]
Then add the squares
Why: areas cannot cancel
\[ 2 \times 16 + 2 \times 1 = 34 \]
Check by pairs
Why: ±4, ±1 square alike
Prediction
Figure (svg): Dot plots of set P (1, 5, 9) and set Q (3, 4, 5, 6, 7) on axes from 0 to 10, both with mean 5
Predict first
Set P: 1, 5, 9. Set Q: 3, 4, 5, 6, 7. Both have mean 5; both are whole populations.
Which set has the larger σ?
Correct: Set P
Why: σ measures typical distance, not how many values. P's outer values sit 4 away; Q's sit at most 2 away.
Worked example
Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 0 deviation arrows from it
\[ \textcolor{#1f5fbf}{1 + 5 + 9} = 15 \]
Add P's values
Why: the mean needs the total
\[ \mu = 15 \div 3 = 5 \]
Divide by N = 3
Why: the centre for the distances
\[ \textcolor{#1f5fbf}{-4,\ 0,\ +4} \]
Subtract 5 from 1, 5, 9
Why: spread is measured from the centre
Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 3 deviation arrows from it
\[ -4 + 0 + 4 = 0 \]
Check the deviations balance
Why: Idea 1 says they must
Worked example
Figure (svg): Set P's values 1, 5, 9 on a number line with the mean 5 marked and 3 deviation arrows from it
\( {\mu = 15 \div 3 = 5}\;\;\Rightarrow\;\;\allowbreak {-4,\ 0,\ +4} \)
\[ \textcolor{#b54708}{16,\ 0,\ 16} \]
Square each
Why: areas that cannot cancel
Figure (svg): Squares on set P's deviations −4, 0, +4
\[ \textcolor{#b54708}{16 + 0 + 16 = 32} \]
Add P's areas
Why: averaging needs a total to share
Figure (svg): Squares on set P's deviations −4, 0, +4
\[ \textcolor{#b54708}{\sigma_P^2} = 32 \div 3 \approx \textcolor{#b54708}{10.67} \]
Divide by N = 3
Why: P is a whole population
Figure (svg): Squares on set P's deviations −4, 0, +4
\[ \textcolor{#1f5fbf}{\sigma_P} = \sqrt{10.67} \approx \textcolor{#1f5fbf}{3.27} \]
Take the root
Why: back to the data's units
Figure (svg): Squares on set P's deviations −4, 0, +4
\[ 3 \times \textcolor{#1f5fbf}{3.27}^2 \approx \textcolor{#b54708}{32.1} \]
Check: undo both moves
Why: restores P's total area 32
Worked example
Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 0 deviation arrows from it
\[ \textcolor{#1f5fbf}{3 + 4 + 5 + 6 + 7} = 25 \]
Add Q's values
Why: the mean needs the total
\[ \mu = 25 \div 5 = 5 \]
Divide by N = 5
Why: the centre for the distances
\[ \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ +1,\ +2} \]
Subtract 5 from 3 to 7
Why: distances show spread
Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 5 deviation arrows from it
\[ -2 - 1 + 0 + 1 + 2 = 0 \]
Check the deviations balance
Why: Idea 1 says they must
Worked example
Figure (svg): Set Q's values 3, 4, 5, 6, 7 on a number line with the mean 5 marked and 5 deviation arrows from it
\( {\mu = 25 \div 5 = 5}\;\;\Rightarrow\;\;\allowbreak {-2,\ -1,\ 0,\ +1,\ +2} \)
\[ \textcolor{#b54708}{4,\ 1,\ 0,\ 1,\ 4} \]
Square each
Why: areas cannot cancel
Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2
\[ \textcolor{#b54708}{4 + 1 + 0 + 1 + 4 = 10} \]
Add Q's areas
Why: the total σ² shares out
Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2
\[ \textcolor{#b54708}{\sigma_Q^2} = 10 \div 5 = \textcolor{#b54708}{2} \]
Divide by N = 5
Why: Q is a whole population
Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2
\[ \textcolor{#1f5fbf}{\sigma_Q} = \sqrt{2} \approx \textcolor{#1f5fbf}{1.41} \]
Take the root
Why: side, not area
Figure (svg): Squares on set Q's deviations −2, −1, 0, +1, +2
\[ \textcolor{#1f5fbf}{3.27} > \textcolor{#1f5fbf}{1.41} \]
Compare SDs
Why: SD tracks spread
\[ 5 \times \textcolor{#1f5fbf}{1.41}^2 \approx \textcolor{#b54708}{9.94} \]
Check: undo both moves
Why: rounding explains the gap
Section
Idea 3 of 5
Concept
Figure (svg): Waits 1, 3, 5 on a number line with 0 deviation arrows from the mean
Real data are samples. Test the recipe on a tiny world: waits 1, 3, 5.
Discussion prompt
Draw 2 waits (repeats allowed), square distances from x̄, ÷ n. Versus σ² on average: high, low, or right?
Answer:
Low on average. The next slides find the true σ² and test all nine samples.
Worked example
Figure (svg): Waits 1, 3, 5 on a number line with 0 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{1} + \textcolor{#1f5fbf}{3} + \textcolor{#1f5fbf}{5} = 9 \]
Add the three waits
Why: μ needs the total first
\[ \mu = 9 \div 3 = 3 \]
Divide by N = 3
Why: a mean is total over count
\[ \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]
Subtract μ from 1, 3, 5
Why: σ² is built from these distances
Figure (svg): Waits 1, 3, 5 on a number line with 3 deviation arrows from the mean
\[ -2 + 0 + 2 = 0 \]
Check the distances balance
Why: so μ = 3 sits at the centre
Worked example
Figure (svg): Waits 1, 3, 5 on a number line with 3 deviation arrows from the mean
\( {1 + 3 + 5 = 9}\;\;\Rightarrow\;\;\allowbreak {\mu = 9 \div 3 = 3}\;\;\Rightarrow\;\;\allowbreak {-2,\ 0,\ +2} \)
\[ \textcolor{#b54708}{4,\ 0,\ 4} \]
Square each distance
Why: areas cannot cancel
Figure (svg): Squares on the distances −2, 0, +2 of the waits 1, 3, 5 from μ = 3
\[ 4 + 0 + 4 = \textcolor{#b54708}{8} \]
Add the areas
Why: σ² shares this total among N = 3
\[ \textcolor{#b54708}{\sigma^2} = 8 \div 3 \approx \textcolor{#b54708}{2.67} \]
Divide by N = 3
Why: every value is known
Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{2.67} \approx \textcolor{#1f5fbf}{1.63} \]
Take the root
Why: back to the waits' units
Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67, side 1.63
\[ 3 \times 2.67 \approx 8 \]
Check: undo the divide
Why: rebuilds the total area
Worked example
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5
\[ \textcolor{#1f5fbf}{3 + 3} = 6,\ \ \textcolor{#1f5fbf}{1 + 3} = 4,\ \ \textcolor{#1f5fbf}{1 + 5} = 6 \]
Add each sample's two draws
Why: each sample's mean needs its total
\[ \bar{x} = 6 \div 2,\ 4 \div 2,\ 6 \div 2 = 3,\ 2,\ 3 \]
Divide each by n = 2
Why: a mean is the total over n
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3
\[ \textcolor{#1f5fbf}{0,\,0;\ \ -1,\,+1;\ \ -2,\,+2} \]
Subtract x̄ from each draw
Why: a sample can't see μ, only x̄
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2
\[ 0 + 0 = 0,\ \ -1 + 1 = 0,\ \ -2 + 2 = 0 \]
Check the distances balance
Why: so each x̄ was right
Worked example
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2
\( {3 + 3 = 6,\ 1 + 3 = 4,\ 1 + 5 = 6}\;\;\Rightarrow\;\;\allowbreak {\bar{x} = 3,\ 2,\ 3}\;\;\Rightarrow\;\;\allowbreak {0,\,0;\ -1,\,+1;\ -2,\,+2} \)
\[ \textcolor{#b54708}{0,\,0;\ \ 1,\,1;\ \ 4,\,4} \]
Square each distance
Why: areas that cannot cancel
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; squares 0, 0; 1, 1; 4, 4
\[ 0 + 0 = 0,\ \ 1 + 1 = 2,\ \ 4 + 4 = 8 \]
Add each sample's areas
Why: ÷ n needs totals
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; square totals 0, 2, 8
\[ 0 \div 2,\ 2 \div 2,\ 8 \div 2 = \textcolor{#b54708}{0,\ 1,\ 4} \]
Divide each by n = 2
Why: finishing the σ² recipe
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; ÷ n estimates 0, 1, 4
\[ \textcolor{#1f5fbf}{0}^2,\ \ \textcolor{#1f5fbf}{1}^2,\ \ \textcolor{#1f5fbf}{2}^2 = \textcolor{#b54708}{0,\ 1,\ 4} \]
Check: one square per sample
Why: two equal squares, halved, leave one
Worked example
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5: only the three kinds' cells filled, 0, 1, 4
\[ (3,1) \to 1,\ \ (5,1) \to 4 \]
Swap unequal pairs
Why: order keeps distances
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5: cells 1, 3; 3, 3; 1, 5 and their swaps filled
\[ (3,5),\ (5,3) \to 1;\ \ (1,1),\ (5,5) \to 0 \]
Shift both draws by ±2
Why: x̄ moves with them
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0
\[ 0 + 1 + 4 + 1 + 0 + 1 + 4 + 1 + 0 = \textcolor{#b54708}{12} \]
Add the nine cells
Why: covers every sample once
\[ \textcolor{#b54708}{12} \div 9 \approx \textcolor{#b54708}{1.33} \]
Divide by 9
Why: every sample equally likely, so a plain average
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0
\[ 3(0) + 4(1) + 2(4) = \textcolor{#b54708}{12} \]
Check by kind
Why: three 0s, four 1s, two 4s
Prediction
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0
Predict first
Samples of n = 2: ÷ n averaged 1.33; truth 2.67.
Which divisor averages 2.67?
Correct: n − 1 = 1
Why: Halving the divisor doubles every estimate: 12/9 doubled is 24/9 = 8/3 ≈ 2.67. With n = 2, dividing by n − 1 = 1 is that halving.
Worked example
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; square totals 0, 2, 8
\( {0 + 0 = 0},\ \allowbreak {1 + 1 = 2},\ \allowbreak {4 + 4 = 8} \)
\[ 2 - 1 = 1 \]
Find n − 1 for samples of two
Why: the predicted divisor
\[ 0 \div 1,\ 2 \div 1,\ 8 \div 1 = \textcolor{#b54708}{0,\ 2,\ 8} \]
Divide each total by n − 1
Why: tests the new divisor per kind
Figure (svg): Three samples of two from 1, 3, 5 on number lines: a repeat 3, 3; one step 1, 3; two steps 1, 5; each sample's mean x̄ marked at 3, 2, 3; distances from x̄ 0, 0; −1, +1; −2, +2; ÷ (n − 1) estimates 0, 2, 8
\[ 2 \times \textcolor{#b54708}{0},\ \ 2 \times \textcolor{#b54708}{1},\ \ 2 \times \textcolor{#b54708}{4} = \textcolor{#b54708}{0,\ 2,\ 8} \]
Check against ÷ n's 0, 1, 4
Why: halving the divisor doubles each
Worked example
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, the divide-by-(n − 1) cells not yet filled
\( {2 - 1 = 1}\;\;\Rightarrow\;\;\allowbreak {0 \div 1,\ 2 \div 1,\ 8 \div 1 = 0,\ 2,\ 8} \)
\[ \textcolor{#b54708}{0,\ 2,\ 8,\ 2,\ 0,\ 2,\ 8,\ 2,\ 0} \]
Place each kind's estimate
Why: swaps and shifts keep kinds
Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0
\[ 0 + 2 + 8 + 2 + 0 + 2 + 8 + 2 + 0 = \textcolor{#b54708}{24} \]
Add the nine cells
Why: covers every sample once
\[ \textcolor{#b54708}{24} \div 9 = 8/3 \]
Divide by 9
Why: each of the nine samples is equally likely
Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67
\[ 24 = 2 \times 12 \]
Check against ÷ n
Why: halving the divisor doubled the total 12
Concept
Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67
avg: over all equally likely samples
\[ \text{avg}{\textstyle\sum} (x - \bar{x})^2 = (n - 1)\textcolor{#b54708}{\sigma^2} \]
Goal
Why: what the three phases prove
Assumed: independent draws
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked
Name r = x − x̄ and the offset g = x̄ − μ.
\[ \textcolor{#1f5fbf}{x - \mu} = (x - \textcolor{#6b7280}{\bar{x}}) + (\textcolor{#6b7280}{\bar{x}} - \mu) \]
Add and subtract x̄
Why: adding zero exposes x̄
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distance x − μ for x = 1 split at x̄ into r = −1 and the offset g = −1
\[ \textcolor{#1f5fbf}{x - \mu} = \textcolor{#1f5fbf}{r} + \textcolor{#1f5fbf}{g} \]
Substitute the names r and g
Why: shorter, same values
\[ \textcolor{#b54708}{(x - \mu)^2} = (\textcolor{#1f5fbf}{r} + \textcolor{#1f5fbf}{g})^2 \]
Square both sides
Why: turns distances into areas
\[ = \textcolor{#b54708}{r^2 + 2rg + g^2} \]
Expand
Why: so each piece sums separately
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distance x − μ for x = 1 split at x̄ into r = −1 and the offset g = −1; a square of side r + g cut into r², rg, rg, g², four equal pieces since r = g = −1
\[ (1 - 3)^2 = (-1)^2 + 2(-1)(-1) + (-1)^2 \]
Check x = 1 in sample 1, 3
Why: r and g both −1: 4 each side
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked
\( {x - \mu = (x - \bar{x}) + (\bar{x} - \mu) = r + g}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 = (r + g)^2 = r^2 + 2rg + g^2} \)
\[ {\textstyle\sum} 2rg = 2\textcolor{#1f5fbf}{g}{\textstyle\sum} \textcolor{#1f5fbf}{r} \]
Pull 2g out of the sum
Why: g is the same for every x
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1, highlighted
\[ = 2\textcolor{#1f5fbf}{g} \cdot 0 \]
Use Σr = 0
Why: x̄ balances the r values
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2: the offset arrow g = −1, and each value's distance r from x̄, −1 and +1, which total 0
\[ = 0 \]
Multiply by 0
Why: any number times 0 is 0
\[ 2(-1)[(1 - 2) + (3 - 2)] = 0 \]
Check sample 1, 3
Why: g = −1; its r values cancel
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked
\( {(x - \mu)^2 = r^2 + 2rg + g^2}\qquad\allowbreak {{\textstyle\sum} 2rg = 0} \)
\[ \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} = {\textstyle\sum} (r^2 + 2rg + g^2) \]
Sum over the sample
Why: totals matter here
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4
\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + {\textstyle\sum} 2\textcolor{#1f5fbf}{r}\textcolor{#1f5fbf}{g} + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]
Split the sum
Why: isolates each piece
\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + 0 + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]
Use Σ2rg = 0
Why: proved on slide 60
\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{g}^2} \]
Drop the 0
Why: 0 adds nothing
\[ = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \]
Count g² terms
Why: n equal copies
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2
\[ (-2)^2 + 0^2 = (-1)^2 + 1^2 + 2(-1)^2 \]
Check sample 1, 3
Why: both sides give 4
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells the offset area n·g² not yet filled
\[ 2,\ 4,\ 6,\ 4,\ 6,\ 8,\ 6,\ 8,\ 10 \]
Add each cell's two draws
Why: each mean needs its total
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the total of its two draws: 2, 4, 6, 4, 6, 8, 6, 8, 10
\[ \bar{x} = 1,\ 2,\ 3,\ 2,\ 3,\ 4,\ 3,\ 4,\ 5 \]
Divide each total by n = 2
Why: a mean is the total over n
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its mean x̄: 1, 2, 3, 2, 3, 4, 3, 4, 5
\[ \textcolor{#1f5fbf}{g} = \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ -1,\ 0,\ 1,\ 0,\ 1,\ 2} \]
Subtract μ = 3 from each x̄
Why: how far each mean misses μ
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its offset g = x̄ − μ: −2, −1, 0, −1, 0, 1, 0, 1, 2
\[ 1 - (-2),\ 2 - (-1),\ 3 - 0 = 3,\ 3,\ 3 \]
Check: undo the subtraction
Why: the first row returns μ = 3
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its offset g = x̄ − μ: −2, −1, 0, −1, 0, 1, 0, 1, 2
\[ \textcolor{#1f5fbf}{g} = \textcolor{#1f5fbf}{-2,\ -1,\ 0,\ -1,\ 0,\ 1,\ 0,\ 1,\ 2} \]
Start from the offsets
Why: their squares build n·g²
\[ \textcolor{#1f5fbf}{g}^2 = \textcolor{#b54708}{4,\ 1,\ 0,\ 1,\ 0,\ 1,\ 0,\ 1,\ 4} \]
Square each offset
Why: n·g² is built from these
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell its squared offset g²: 4, 1, 0, 1, 0, 1, 0, 1, 4
\[ 2\textcolor{#1f5fbf}{g}^2 = \textcolor{#b54708}{8,\ 2,\ 0,\ 2,\ 0,\ 2,\ 0,\ 2,\ 8} \]
Multiply each by n = 2
Why: n copies make the offset area
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8
\[ (4 + 0) - 2 = \textcolor{#b54708}{2},\ \ (4 + 4) - 8 = \textcolor{#b54708}{0} \]
Check 1, 3 and 1, 5
Why: squares from μ less those from x̄
Prediction
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2
Predict first
Can a sample's squares from x̄ ever total MORE than its squares from μ?
Correct: No, never
Why: Squares from μ are squares from x̄ plus n·g², and a square is never negative. At best (x̄ = μ) they tie.
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from x̄, 1 and 1, total 2
\( {{\textstyle\sum} (x - \mu)^2 = {\textstyle\sum} r^2 + n g^2} \)
\[ \textcolor{#1f5fbf}{g}^2 \ge 0 \]
Square the offset
Why: squares of real numbers ignore sign
\[ \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \ge 0 \]
Multiply by n
Why: positive counts keep ≥
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2
\[ \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \le \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} + \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} \]
Add n·g² ≥ 0 to Σr²
Why: a total can only grow
\[ \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \le \textcolor{#b54708}{{\textstyle\sum} (x - \mu)^2} \]
Substitute the split
Why: Σr² + n·g² equals squares from μ
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2
\[ 2 \le 4 \]
Check sample 1, 3
Why: picture totals: x̄ 2, μ 4
Worked example
Figure (svg): A 2 by 2 grid for (e₁ + e₂)²: squares on the diagonal, cross products off it
Draw i has distance eᵢ = xᵢ − μ, for i = 1 to n.
\[ (\textcolor{#1f5fbf}{e_1} + \textcolor{#1f5fbf}{e_2} + \textcolor{#1f5fbf}{e_3})^2 = \textcolor{#b54708}{e_1e_1} + e_1e_2 + \dots + \textcolor{#b54708}{e_3e_3} \]
Multiply out three draws
Why: one term per grid cell
Figure (svg): A 3 by 3 grid for (e₁ + e₂ + e₃)²: squares on the diagonal, cross products off it
\[ ({\textstyle\sum} \textcolor{#1f5fbf}{e})^2 = {\textstyle\sum}_i {\textstyle\sum}_j \textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j} \]
Write every cell for n draws
Why: row i times column j
Figure (svg): An n by n grid for (Σe)², margins e₁, e₂, ⋯, eₙ: every cell the product eᵢeⱼ of its row and column
\[ = \textcolor{#b54708}{{\textstyle\sum} e_i^2} + {\textstyle\sum}_{i \ne j} e_ie_j \]
Split off the diagonal
Why: cells with i = j are squares
Figure (svg): An n by n grid for (Σe)², margins e₁, e₂, ⋯, eₙ: the diagonal cells marked as the squares eᵢ², the off-diagonal cells the cross products eᵢeⱼ
\[ (-2 + 0)^2 = (-2)^2 + 0^2 + 2(-2)(0) \]
Check sample 1, 3
Why: both sides give 4
Figure (svg): A 2 by 2 grid for sample 1, 3's distances −2 and 0: squares 4 and 0 on the diagonal, cross products 0 and 0 off it
Worked example
Figure (svg): Squares on the population 1, 3, 5's distances −2, 0, +2 from μ = 3
Uniform, independent draws. Σ all: over all N values.
\[ \text{avg}\ \textcolor{#b54708}{\textcolor{#1f5fbf}{e_i}^2} = \frac{\textcolor{#b54708}{{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e}^2}}{N} \]
Average one draw's square
Why: all N values equally likely
\[ = \textcolor{#b54708}{\sigma^2} \]
Recognise σ²'s definition
Why: the population's average square
Figure (svg): Squares on the distances −2, 0, +2 with their average square, area 2.67
\[ (4 + 0 + 4) \div 3 = \tfrac{8}{3} \approx \textcolor{#b54708}{2.67} \]
Check with the three values
Why: matches the grids' σ²
Worked example
Figure (svg): Grid of every pair of distances −2, 0, +2, each cell their product
\[ \text{avg}\ \textcolor{#1f5fbf}{e_1}\textcolor{#1f5fbf}{e_2} = \frac{{\textstyle\sum}_{\text{all}}{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1}\textcolor{#1f5fbf}{e_2}}{N^2} \]
Average all N² pairs
Why: independent: pairs equally likely
\[ = \frac{{\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1} \left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_2}\right)}{N^2} \]
Factor e₁ out of the inner sum
Why: fixed while e₂ varies
\[ = \frac{\left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_1}\right)\left({\textstyle\sum}_{\text{all}} \textcolor{#1f5fbf}{e_2}\right)}{N^2} \]
Pull the inner sum out
Why: it is the same for every e₁
Figure (svg): Grid of every pair of distances −2, 0, +2: cells 4, 0, −4, 0, 0, 0, −4, 0, 4; every row and column totals 0
\[ = \frac{\textcolor{#1f5fbf}{0} \cdot \textcolor{#1f5fbf}{0}}{N^2} \]
Substitute Σ all e = 0
Why: distances from μ balance
\[ = 0 \]
Evaluate
Why: zero over a positive count
\[ 4 + 0 - 4 + 0 + 0 + 0 - 4 + 0 + 4 = 0 \]
Check the grid
Why: its nine cells cancel
Prediction
Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells (e₁ + e₂)² not yet filled
Predict first
Squares average σ²; independent cross products average 0.
For n = 2, is avg (Σe)² σ², 2σ² or 4σ²?
Correct: 2σ²
Why: Each of the 2 squares averages σ²; the cross products average 0: 2σ². So σ² counts only one square, and 4σ² would need the cross products to add another 2σ².
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells (e₁ + e₂)² not yet filled
\[ 1 - 3,\ 3 - 3,\ 5 - 3 = \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]
Subtract μ = 3
Why: the next result uses μ
\[ \textcolor{#1f5fbf}{-4,\ -2,\ 0,\ -2,\ 0,\ 2,\ 0,\ 2,\ 4} \]
Add e₁ + e₂ per cell
Why: (Σe)² sums first
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the sum of distances e₁ + e₂: −4, −2, 0, −2, 0, 2, 0, 2, 4
\[ \textcolor{#b54708}{16,\ 4,\ 0,\ 4,\ 0,\ 4,\ 0,\ 4,\ 16} \]
Square each sum
Why: the grid averages these
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16
\[ (3 - 3 + 5 - 3)^2 = \textcolor{#b54708}{4} \]
Check sample 3, 5 from its waits
Why: its cell reads 4
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16
\( {1 - 3,\ 3 - 3,\ 5 - 3 = -2,\ 0,\ +2}\;\;\Rightarrow\;\;\allowbreak {-4,\ -2,\ 0,\ -2,\ 0,\ 2,\ 0,\ 2,\ 4}\;\;\Rightarrow\;\;\allowbreak {16,\ 4,\ 0,\ 4,\ 0,\ 4,\ 0,\ 4,\ 16} \)
\[ \textcolor{#b54708}{16 + 4 + 0 + 4 + 0 + 4 + 0 + 4 + 16} = \textcolor{#b54708}{48} \]
Add the nine cells
Why: covers every sample once
\[ \textcolor{#b54708}{48} \div 9 = \textcolor{#b54708}{16/3} \]
Divide by 9
Why: nine equal chances
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 16/3, beside a bar for σ² = 2.67
\[ \textcolor{#b54708}{16/3} \div \textcolor{#b54708}{8/3} = 2 \]
Divide by σ² = 8/3
Why: settles the prediction: 2σ²
\[ 2(16) + 4(4) = \textcolor{#b54708}{48} \]
Check by kind
Why: two 16s, four 4s; 0s add nothing
Worked example
Figure (svg): A 3 by 3 grid for (e₁ + e₂ + e₃)²: squares on the diagonal, cross products off it
\[ \text{avg}\textcolor{#b54708}{({\textstyle\sum} \textcolor{#1f5fbf}{e})^2} = \text{avg}(\textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{e_i}^2} + {\textstyle\sum}_{i \ne j} \textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j}) \]
Average the grid expansion
Why: equal things average equally
\[ = {\textstyle\sum} \text{avg}\,\textcolor{#b54708}{\textcolor{#1f5fbf}{e_i}^2} + {\textstyle\sum}_{i \ne j} \text{avg}\,\textcolor{#1f5fbf}{e_i}\textcolor{#1f5fbf}{e_j} \]
Average cell by cell
Why: averages of sums add
\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} \text{avg}\,e_ie_j \]
Substitute avg eᵢ² = σ²
Why: from the one-draw lemma
\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]
Substitute avg eᵢeⱼ = 0
Why: independent draws' lemma
Figure (svg): A 3 by 3 grid of cell averages: σ² on each diagonal cell, 0 on each off-diagonal cell
\[ 48 \div 9 = 16/3 = 2 \times 8/3 \]
Check with n = 2
Why: agrees with the grid
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 5.33 = 2σ²
Worked example
Figure (svg): A 3 by 3 grid of cell averages: σ² on each diagonal cell, 0 on each off-diagonal cell
\[ \text{avg}({\textstyle\sum} e)^2 = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]
Start from the cell averages
Why: the averaged-grid result
\[ = n\textcolor{#b54708}{\sigma^2} + {\textstyle\sum}_{i \ne j} 0 \]
Add n diagonal σ²
Why: n equal terms
\[ = n\textcolor{#b54708}{\sigma^2} + (n^2 - n) \cdot 0 \]
Count cross cells
Why: n² cells less n squares
Figure (svg): A 3 by 3 grid of cell averages with the six off-diagonal cells outlined: 3² − 3 = 6 cross cells, each averaging 0
\[ = n\textcolor{#b54708}{\sigma^2} + 0 \]
Multiply by 0
Why: anything times 0 is 0
\[ = n\textcolor{#b54708}{\sigma^2} \]
Drop the 0
Why: 0 adds nothing
\[ 2 \times \textcolor{#b54708}{8/3} \approx \textcolor{#b54708}{5.33} \]
Check n = 2, σ² = 8/3
Why: the (Σe)² grid's average bar
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell (e₁ + e₂)²: 16, 4, 0, 4, 0, 4, 0, 4, 16, averaging 5.33 = 2σ²
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked
\[ \textcolor{#1f5fbf}{g} = \textcolor{#6b7280}{\bar{x}} - \textcolor{#6b7280}{\mu} \]
Start from the offset's definition
Why: to rewrite g using the data's total
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1, highlighted
\[ = \textcolor{#6b7280}{\frac{{\textstyle\sum} x}{n}} - \textcolor{#6b7280}{\mu} \]
Use x̄'s definition
Why: the prerequisite mean
\[ = \textcolor{#6b7280}{\frac{{\textstyle\sum} x}{n}} - \textcolor{#6b7280}{\frac{n\mu}{n}} \]
Write μ as nμ/n
Why: to share a denominator
\[ = \textcolor{#1f5fbf}{\frac{{\textstyle\sum} x - n\mu}{n}} \]
Combine the fractions
Why: same denominator n
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1
\[ \frac{4 - 2(3)}{2} = -1 = 2 - 3 \]
Check with sample 1, 3
Why: waits 1 and 3 total 4
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, and the offset arrow g = x̄ − μ = −1
\( {g = \bar{x} - \mu}\;=\;\allowbreak {{\textstyle\sum} x/n - \mu}\;=\;\allowbreak {{\textstyle\sum} x/n - n\mu/n}\;=\;\allowbreak {({\textstyle\sum} x - n\mu)/n} \)
\[ \textcolor{#1f5fbf}{g} = \frac{{\textstyle\sum} x - {\textstyle\sum} \textcolor{#6b7280}{\mu}}{n} \]
Write nμ as Σμ
Why: a constant summed n times
\[ = \frac{{\textstyle\sum} \textcolor{#1f5fbf}{(x - \mu)}}{n} \]
Combine the two sums
Why: sums of differences rule
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1
\[ = \frac{{\textstyle\sum} \textcolor{#1f5fbf}{e}}{n} \]
Write x − μ as e
Why: the form the grid averages use
\[ \frac{-2 + 0}{2} = -1 \]
Check with sample 1, 3
Why: the blue arrows average −1
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, highlighted
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; the distances e from μ, −2 and 0, and the offset g = −1
\( {g = {\textstyle\sum} e / n} \)
\[ \textcolor{#1f5fbf}{g}^2 = ({\textstyle\sum} \textcolor{#1f5fbf}{e} / n)^2 \]
Square g = Σe/n
Why: heading for the offset area
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, highlighted
\[ = ({\textstyle\sum} e)^2 / n^2 \]
Square top and bottom
Why: the quotient rule
\[ \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = n({\textstyle\sum} e)^2 / n^2 \]
Multiply by n
Why: builds the offset area
Figure (svg): Sample 1, 3 from the population with μ = 3: its mean x̄ = 2, the distances e from μ, and the offset arrow g = x̄ − μ = −1, with the offset area of two g² squares
\[ = \textcolor{#b54708}{({\textstyle\sum} e)^2} / n \]
Cancel one n
Why: n over n² is 1 over n
\[ 2(-1)^2 = (-2 + 0)^2 / 2 \]
Check with sample 1, 3
Why: both sides give 2
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8
\( {n\,g^2 = ({\textstyle\sum} e)^2 / n}\qquad\allowbreak {\text{avg}({\textstyle\sum} e)^2 = n\sigma^2} \)
\[ \text{avg}\ \textcolor{#b54708}{n\,g^2} = \text{avg}\,[({\textstyle\sum} e)^2 / n] \]
Average both sides
Why: equal quantities stay equal
\[ = \text{avg}({\textstyle\sum} e)^2 / n \]
Move 1/n outside
Why: constants leave averages
\[ = n\textcolor{#b54708}{\sigma^2} / n \]
Substitute avg (Σe)² = nσ²
Why: the squared-sum lemma
\[ = \textcolor{#b54708}{\sigma^2} \]
Cancel n
Why: n copies undo ÷ n
\[ [2(8) + 4(2) + 3(0)] \div 9 = 24 \div 9 = 8/3 \]
Check with the grid
Why: offset areas average σ²
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell the offset area n·g²: 8, 2, 0, 2, 0, 2, 0, 2, 8, averaging 2.67 = σ²
Prediction
Predict first
Draw 2 different waits from 1, 3, 5: no repeats allowed.
Do the ÷ (n − 1) estimates still average σ² ≈ 2.67?
Correct: No, they average more
Why: Without repeats the draws are not independent, so the cross products no longer average 0. The three estimates are 2, 8 and 2: 2 + 8 + 2 = 12, and 12 ÷ 3 = 4.
Worked example
Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0
No repeats leaves three pairs: 1, 3; 3, 5; 1, 5.
Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0; the diagonal repeats 1, 1; 3, 3; 5, 5 greyed and struck out, leaving six cells 2, 8, 2, 2, 8, 2
\[ \textcolor{#b54708}{2,\ 2,\ 8} \]
Take each pair's estimate
Why: to see dependence's effect
\[ 2 + 2 + 8 = 12 \]
Add the three estimates
Why: an average starts from the total
\[ 12 \div 3 = \textcolor{#b54708}{4} \]
Divide by 3
Why: three equally likely pairs
Figure (svg): Grid of the nine ordered samples of two from 1, 3, 5, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0; the diagonal repeats 1, 1; 3, 3; 5, 5 greyed and struck out, leaving six cells 2, 8, 2, 2, 8, 2; a bar for the six cells' average, 4, beside a bar for σ² = 2.67
\[ \textcolor{#b54708}{4} > 8/3 \approx \textcolor{#b54708}{2.67} \]
Compare with σ²
Why: the independence hypothesis mattered
\[ [4(2) + 2(8)] \div 6 = 4 \]
Check with ordered pairs
Why: each pair counted both ways
Worked example
Figure (svg): Grid of the nine samples of two from 1, 3, 5, cells e₁² + e₂² not yet filled
\[ 1 - 3,\ 3 - 3,\ 5 - 3 = \textcolor{#1f5fbf}{-2,\ 0,\ +2} \]
Subtract μ = 3
Why: the lemma works from μ
\[ \textcolor{#b54708}{4,\ 0,\ 4} \]
Square each
Why: a cell adds two of them
\[ \textcolor{#b54708}{8,\ 4,\ 8,\ 4,\ 0,\ 4,\ 8,\ 4,\ 8} \]
Add e₁² + e₂²
Why: square first: no cross products
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8
\[ 8+4+8+4+0+4+8+4+8 = \textcolor{#b54708}{48} \]
Add the nine cells
Why: covers every sample once
\[ \textcolor{#b54708}{48} \div 9 = \textcolor{#b54708}{16/3} \]
Divide by 9
Why: nine equal chances
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8, averaging 5.33 = 2σ²
\[ 4(8) + 4(4) = \textcolor{#b54708}{48} \]
Check by kind
Why: four 8s, four 4s, one 0
Worked example
Figure (svg): Bar of the nine samples' average squares from μ, 5.33
\[ \text{avg}{\textstyle\sum} \textcolor{#b54708}{e_i^2} = {\textstyle\sum} \text{avg}\,\textcolor{#b54708}{e_i^2} \]
Average each square
Why: averages of sums add
\[ = {\textstyle\sum} \textcolor{#b54708}{\sigma^2} \]
Substitute avg eᵢ² = σ²
Why: the one-draw lemma
Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67
\[ = n\textcolor{#b54708}{\sigma^2} \]
Add n copies of σ²
Why: summing a constant multiplies it
\[ \tfrac{16}{3} = 2 \times \tfrac{8}{3} \]
Check on the e₁² + e₂² grid
Why: n = 2: its cells average 16/3
Figure (svg): Grid of the nine samples of two from 1, 3, 5, each cell e₁² + e₂²: 8, 4, 8, 4, 0, 4, 8, 4, 8, averaging 5.33 = 2σ²
Worked example
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2
\( {{\textstyle\sum} (x - \mu)^2 = {\textstyle\sum} r^2 + n g^2} \)
\[ {\textstyle\sum} (x - \mu)^2 - \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = \textcolor{#b54708}{{\textstyle\sum} \textcolor{#1f5fbf}{r}^2} \]
Subtract n·g² from the split
Why: leaves only squares from x̄
Figure (svg): Sample 1, 3 from the population with μ = 3, its mean x̄ = 2 marked; squares from μ, 4 and 0, total 4; squares from x̄, 1 and 1, total 2; the offset area, two g² squares of 1, total 2
\[ {\textstyle\sum} \textcolor{#b54708}{e_i^2} - \textcolor{#b54708}{n\,\textcolor{#1f5fbf}{g}^2} = {\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Rename x − μ as eᵢ
Why: matches the one-draw lemma
\[ 4 - 2 = (1 - 2)^2 + (3 - 2)^2 = 2 \]
Check with sample 1, 3
Why: matches its squares from x̄
Worked example
Figure (svg): Bar of the nine samples' average squares from μ, 5.33
\( {{\textstyle\sum} (x - \mu)^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {{\textstyle\sum} e_i^2 - n g^2 = {\textstyle\sum} r^2} \)
\[ \text{avg}({\textstyle\sum} e_i^2 - n\textcolor{#1f5fbf}{g}^2) = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Average both sides
Why: equal things stay equal
Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67
\[ \text{avg}{\textstyle\sum} e_i^2 - \text{avg}\,n\textcolor{#1f5fbf}{g}^2 = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Split the average
Why: means of differences subtract
\[ n\textcolor{#b54708}{\sigma^2} - \text{avg}\,n\textcolor{#1f5fbf}{g}^2 = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Substitute avg Σeᵢ² = nσ²
Why: a proved lemma
\[ \tfrac{16}{3} - \tfrac{8}{3} = \tfrac{8}{3} = \tfrac{24}{9} \]
Check n = 2
Why: from the grids; x̄-squares average 24/9
Worked example
Figure (svg): Bar of the nine samples' average squares from μ, 5.33, split into two σ² blocks of 2.67
\( {{\textstyle\sum} (x - \mu)^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {{\textstyle\sum} e_i^2 - n g^2 = {\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {\text{avg}({\textstyle\sum} e_i^2 - n g^2) = \text{avg}{\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {\text{avg}{\textstyle\sum} e_i^2 - \text{avg}\,n g^2 = \text{avg}{\textstyle\sum} r^2}\;\;\Rightarrow\;\;\allowbreak {n\sigma^2 - \text{avg}\,n g^2 = \text{avg}{\textstyle\sum} r^2} \)
\[ n\textcolor{#b54708}{\sigma^2} - \textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Substitute avg n·g² = σ²
Why: the offset-area lemma
Figure (svg): Average squares from μ, 5.33, as two σ² blocks; under the second, the dashed offset area of σ² = 2.67
\[ (n - 1)\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} \textcolor{#1f5fbf}{r}^2 \]
Factor out σ²
Why: n copies less one
Figure (svg): Average squares from μ, 5.33, as two σ² blocks; squares from x̄ keep one block, 2.67, beside the offset area
\[ (2 - 1) \times \frac{8}{3} = \frac{24}{9} \]
Check nine samples
Why: x̄-squares total 24
Worked example
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0
\( {(n - 1)\sigma^2 = \text{avg}{\textstyle\sum} r^2} \)
\[ (n - 1)\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} (x - \bar{x})^2 \]
Write r as x − x̄
Why: x − x̄ needs only the sample
\[ \textcolor{#b54708}{\sigma^2} = \frac{\text{avg}{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]
Divide by n − 1
Why: positive once n ≥ 2
\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\ \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]
Move 1/(n − 1) inside
Why: constants leave averages
Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67
\[ [3(0) + 4(\textcolor{#b54708}{2}) + 2(\textcolor{#b54708}{8})] \div 9 \approx \textcolor{#b54708}{2.67} \]
Check n = 2 on the corrected grid
Why: three 0s, four 2s, two 8s
Worked example
Figure (svg): Grid of the nine size-2 samples from 1, 3, 5, each cell (half the gap)²: 0, 1, 4, 1, 0, 1, 4, 1, 0
\( {(n - 1)\sigma^2 = \text{avg}{\textstyle\sum} (x - \bar{x})^2}\;\;\Rightarrow\;\;\allowbreak {\sigma^2 = \text{avg}{\textstyle\sum} (x - \bar{x})^2/(n - 1)}\;\;\Rightarrow\;\;\allowbreak {\sigma^2 = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/(n - 1)]} \)
\[ \textcolor{#b54708}{s^2} = \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]
Name the sample variance s²
Why: a data-only estimate of σ²
Figure (svg): Grid of the nine size-2 samples, each cell the divide-by-(n − 1) estimate: 0, 2, 8, 2, 0, 2, 8, 2, 0, average 2.67
\[ \textcolor{#b54708}{\sigma^2} = \text{avg}\ \textcolor{#b54708}{s^2} \]
Substitute s²
Why: independent draws, n ≥ 2
\[ \textcolor{#1f5fbf}{s} = \sqrt{\textcolor{#b54708}{s^2}} \]
Define s, its root: the sample SD
Why: the side, in data units
\[ \text{avg}\ \textcolor{#b54708}{s^2} = \textcolor{#b54708}{24} \div 9 \approx \textcolor{#b54708}{2.67} \]
Check the grid
Why: nine cells total 24
Concept
Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1
Discussion prompt
s² is how many times ÷ n, at n = 2 and 20?
Answer:
About twice as big at n = 2; about 5% bigger at n = 20.
\[ \textcolor{#b54708}{s^2} = \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} = \frac{n}{n} \cdot \frac{{\textstyle\sum} (x - \bar{x})^2}{n - 1} \]
Multiply by n/n
Why: a factor of 1 changes nothing
\[ = \frac{n}{n - 1} \cdot \frac{{\textstyle\sum} (x - \bar{x})^2}{n} \]
Swap the denominators
Why: factor order is free
\[ \tfrac{2}{1} \times 4 = \textcolor{#b54708}{8} \]
Check sample 1, 5
Why: the ÷ n value doubles
Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2, circled and labelled
Worked example
Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1
df = n − 1: Σ(x − x̄) = 0 fixes one distance
n = 1: x = x̄, so s² = 0 ÷ 0, undefined
\[ 2 - 1 = 1,\ \ 20 - 1 = 19 \]
Find both corrected divisors
Why: the ratio needs both
\[ 2 \div 1 = 2 \]
Divide n by n − 1
Why: shows how much s² grows
Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2, circled and labelled
\[ 20 \div 19 \approx 1.05 \]
Repeat at n = 20
Why: tests a bigger sample
Figure (svg): Plot of n ÷ (n − 1) against n for n = 2, 3, 4, 5, 10, 20, falling from 2 toward 1; the n = 2 point, 2 and the n = 20 point, 1.05, circled and labelled
\[ 19 \times 1.05 \approx 20 \]
Check: undo the divide
Why: returns the sample size 20
Concept
Figure (svg): Example 2.32's twenty fifth graders' ages as dots stacked by value on a number line: 9 (1 dot), 9.5 (2), 10 (4), 10.5 (4), 11 (6), 11.5 (3), each stack's count printed above it
Example 2.32: twenty fifth graders' ages, a sample, with only six different values.
Discussion prompt
The table lists each age once with its count f. How can one square stand in for several pupils?
Answer:
Multiply that age's square by f: f equal squares.
Worked example
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3)
f = how many of the 20 pupils have that age.
\[ \begin{aligned}&1(9),\ 2(9.5),\ 4(10),\ 4(10.5), \\ &6(11),\ 3(11.5) = 9,\ 19,\ 40,\ 42,\ 66,\ 34.5\end{aligned} \]
Multiply each age by f
Why: f pupils share that age
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), products fx (9, 19, 40, 42, 66, 34.5)
\[ 9 + 19 + 40 + 42 + 66 + 34.5 = 210.5 \]
Add the six products
Why: the mean needs all 20
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), products fx (9, 19, 40, 42, 66, 34.5); totals: f 20, fx 210.5
\[ \bar{x} = 210.5 \div 20 = 10.525 \]
Divide by n = 20
Why: a mean is total over count
\[ 20 \times 10.525 = 210.5 \]
Check: undo the division
Why: restores the ages' total
Worked example
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3)
\[ \bar{x} = 10.525 \]
Start from the computed mean
Why: spread is measured from it
\[ \textcolor{#1f5fbf}{\begin{aligned}&{-1.525},\ {-1.025},\ {-0.525}, \\ &{-0.025},\ 0.475,\ 0.975\end{aligned}} \]
Subtract x̄ from each age
Why: spread starts from these distances
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975)
\[ \begin{aligned}&1(-1.525) + 2(-1.025) + 4(-0.525) \\ &+ 4(-0.025) + 6(0.475) + 3(0.975) = 0\end{aligned} \]
Check the weighted distances balance
Why: confirms x̄ = 10.525
Worked example
Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975)
\[ \textcolor{#b54708}{\begin{aligned}&2.325625,\ 1.050625,\ 0.275625, \\ &0.000625,\ 0.225625,\ 0.950625\end{aligned}} \]
Square each distance
Why: so each gap becomes an area
Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), distances x − x̄ (−1.525, −1.025, −0.525, −0.025, 0.475, 0.975), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625)
\[ \textcolor{#b54708}{\begin{aligned}&2.325625,\ 2.10125,\ 1.1025, \\ &0.0025,\ 1.35375,\ 2.851875\end{aligned}} \]
Multiply each by f
Why: f equal copies
Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875)
\[ 1.1025 \div 4 = 0.275625 \]
Check: undo ×4 for age 10
Why: returns that age's square
Worked example
Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875)
\[ \textstyle{\textstyle\sum} f(x - \bar{x})^2 = \textcolor{#b54708}{9.7375} \]
Add the six products
Why: s² needs all 20 pupils
Figure (svg): Table of Example 2.32's counts f (1, 2, 4, 4, 6, 3), squared distances (2.325625, 1.050625, 0.275625, 0.000625, 0.225625, 0.950625), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375
\[ \begin{aligned}&(2.325625 + 2.10125 + 1.1025) \\ &+ (0.0025 + 1.35375 + 2.851875)\end{aligned} \]
Group the products in halves
Why: a second route to the total
\[ = 5.529375 + 4.208125 \]
Add inside each bracket
Why: splits the sum for checking
\[ 5.529375 + 4.208125 = \textcolor{#b54708}{9.7375} \]
Check: add the halves
Why: matches the table's total
Worked example
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, not yet filled
\[ 20 - 1 = 19 \]
Find n − 1 for 20 pupils
Why: they stand in for all fifth graders
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19
\[ \textcolor{#b54708}{s^2} = 9.7375 \div 19 \]
Substitute into s²
Why: the table gave this total
\[ = \textcolor{#b54708}{0.5125} \]
Divide by n − 1 = 19
Why: the sample's fair average square
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19, s² = 0.5125
\[ \textcolor{#1f5fbf}{s} = \sqrt{0.5125} \]
Substitute into s = √s²
Why: the side of that area
\[ \approx \textcolor{#1f5fbf}{0.716} \]
Take the root
Why: back in years
Figure (svg): Table of Example 2.32's ages x (9 to 11.5), counts f (1, 2, 4, 4, 6, 3), weighted squares f(x − x̄)² (2.325625, 2.10125, 1.1025, 0.0025, 1.35375, 2.851875); totals: f 20, f(x − x̄)² 9.7375; result rows for n − 1, s² and s, filled so far: n − 1 = 19, s² = 0.5125, s ≈ 0.716
\[ \textcolor{#1f5fbf}{0.716}^2 \times 19 \approx \textcolor{#b54708}{9.74} \]
Check: undo both moves
Why: restores the total
Trap
\[ 9.7375 \div 20 \approx 0.487 \]
Divide by n = 20
Why: σx: treats 20 as everyone
\[ \sqrt{0.487} \approx 0.698 \]
Take the root
Why: to match the calculator's σx
\[ \text{report } 0.698 \]
Report σx
Why: fails: squares from x̄ run low
\[ 9.7375 \div 19 = 0.5125 \]
Divide by n − 1
Why: fair for a sample
\[ \sqrt{0.5125} \approx 0.716 \]
Take the root
Why: back in years
\[ \text{report } s = 0.716 \]
Report Sx = s
Why: s² is right on average
\[ \textcolor{#1f5fbf}{0.698}^2 \times 20 \approx \textcolor{#1f5fbf}{0.716}^2 \times 19 \approx \textcolor{#b54708}{9.74} \]
Check: both rebuild the total
Why: only the divisor differs
Prediction
Predict first
5 volunteers stand in for all adults.
Why does ÷ 5 understate their heart-rate spread on average?
Correct: x̄ sits closer to the five than μ does
Why: Squares from x̄ are squares from μ minus an offset area that averages σ². Dividing by n ignores the lost area; n − 1 restores it.
Worked example
Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²
\[ 5 - 1 = 4 \]
Find n − 1
Why: (n − 1)σ² needs this count
\[ 4\textcolor{#b54708}{\sigma^2} = \text{avg}{\textstyle\sum} (x - \bar{x})^2 \]
Substitute into (n − 1)σ²
Why: one σ² lost
Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²; squares from x̄ average four blocks, 4σ², the fifth lost as the offset area
\[ 4\textcolor{#b54708}{\sigma^2}/5 = \text{avg}{\textstyle\sum} (x - \bar{x})^2/5 \]
Divide both sides by 5
Why: what ÷ n would report
\[ 4\textcolor{#b54708}{\sigma^2}/5 = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/5] \]
Move ÷ 5 inside avg
Why: a constant leaves it
\[ 0.8\textcolor{#b54708}{\sigma^2} = \text{avg}\,[{\textstyle\sum} (x - \bar{x})^2/5] \]
Evaluate 4 ÷ 5
Why: to compare with one σ²
Figure (svg): Average areas for samples of five, in blocks of σ²: squares from μ average five blocks, 5σ²; squares from x̄ average four blocks, 4σ², the fifth lost as the offset area; dividing by 5 averages 0.8σ², short of one σ² block
\[ 0.8 \times \textcolor{#b54708}{25} = (4 \times \textcolor{#b54708}{25}) \div 5 \]
Check with σ² = 25
Why: both sides give 20
Sorting
Sort into buckets
Divide by N (the data are the population) or n − 1 (a sample)?
Worked example
Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12
Team heights, today's waits: ÷ N
Why: the data are the whole group
College players, this year's waits: ÷ (n − 1)
Why: data estimate a larger group
\[ 12 - 1 = 11 \]
Find n − 1 for 12 players
Why: needed to compare the two divisors
Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12; the same total cut into 11 equal shares, dividing by n − 1 = 11
\[ 12 \div 11 \approx \textcolor{#b54708}{1.09} \]
Compare the divisors
Why: shows how much larger s² runs
Figure (svg): A total of squares drawn as a strip cut into 12 equal shares, dividing by N = 12; the same total cut into 11 equal shares, dividing by n − 1 = 11; one share of each highlighted, the 11-way share about 1.09 times as long
\[ 11 \times 1.09 \approx 12 \]
Check: undo the divide
Why: returns the 12 players
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6
Try It 2.33: f of the 29 stores carry x food types.
\[ \begin{aligned}&4(6) + 5(7) + 1(8) + 4(9) \\ &+ 5(10) + 4(11) + 6(12) \\ &= 24 + 35 + 8 + 36 + 50 + 44 + 72\end{aligned} \]
Multiply each x by f
Why: f equal values at once
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, products fx 24, 35, 8, 36, 50, 44, 72
\[ = 269 \]
Add the seven products
Why: every store counts
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, products fx 24, 35, 8, 36, 50, 44, 72; totals: 29 stores, fx 269
\[ \bar{x} = 269 \div 29 \approx 9.28 \]
Divide by n = 29
Why: a mean is total over count
\[ 29 \times 9.28 \approx 269.1 \]
Check: undo the division
Why: near 269: x̄ was rounded
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² not yet filled
\[ 6 - 9.28 = \textcolor{#1f5fbf}{-3.28} \]
Subtract x̄ from 6
Why: spread is measured from the mean
\[ (\textcolor{#1f5fbf}{-3.28})^2 = \textcolor{#b54708}{10.7584} \]
Square it
Why: the column holds f(x − x̄)²
\[ 4 \times 10.7584 \approx \textcolor{#b54708}{43.034} \]
Multiply by f = 4
Why: four stores, four squares
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, the rest left blank
\[ \sqrt{43.034 \div 4} \approx \textcolor{#1f5fbf}{3.28} \]
Check: undo both moves
Why: returns the distance from x̄
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, the rest left blank
\[ 7 - 9.28,\ \ 8 - 9.28 = \textcolor{#1f5fbf}{-2.28,\ -1.28} \]
Subtract x̄ from 7 and 8
Why: spread starts at x̄
\[ (\textcolor{#1f5fbf}{-2.28})^2,\ (\textcolor{#1f5fbf}{-1.28})^2 = \textcolor{#b54708}{5.1984,\ 1.6384} \]
Square each distance
Why: before f scales them
\[ 5 \times 5.1984,\ \ 1 \times 1.6384 \approx \textcolor{#b54708}{25.992,\ 1.638} \]
Multiply each by its f
Why: f equal squares
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, the rest left blank
\[ \begin{aligned}&9.28 - \sqrt{25.992 \div 5} = 7 \\ &9.28 - \sqrt{1.6384 \div 1} = 8\end{aligned} \]
Check: undo every move
Why: catches a slip in either row
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, the rest left blank
\[ 9 - 9.28,\ \ 10 - 9.28 = \textcolor{#1f5fbf}{-0.28,\ 0.72} \]
Subtract x̄ from 9 and 10
Why: signed gaps feed the squares
\[ (\textcolor{#1f5fbf}{-0.28})^2,\ \textcolor{#1f5fbf}{0.72}^2 = \textcolor{#b54708}{0.0784,\ 0.5184} \]
Square each distance
Why: areas before f weights them
\[ 4 \times 0.0784,\ \ 5 \times 0.5184 \approx \textcolor{#b54708}{0.314,\ 2.592} \]
Multiply each by its f
Why: four and five equal areas
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, the rest left blank
\[ \begin{aligned}&9.28 - \sqrt{0.3136 \div 4} = 9 \\ &9.28 + \sqrt{2.592 \div 5} = 10\end{aligned} \]
Check: undo every move
Why: returns both rows exactly to x
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, the rest left blank
\[ 11 - 9.28 = \textcolor{#1f5fbf}{1.72} \]
Subtract x̄ from 11
Why: the gap the square needs
\[ \textcolor{#1f5fbf}{1.72}^2 = \textcolor{#b54708}{2.9584} \]
Square the distance
Why: one store's area first
\[ 4 \times 2.9584 \approx \textcolor{#b54708}{11.834} \]
Multiply by f = 4
Why: four stores share this value
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank
\[ \sqrt{11.8336 \div 4} + 9.28 = 1.72 + 9.28 = 11 \]
Check: undo all three moves
Why: lands back on x = 11
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank
\[ \begin{aligned}&43.034 + 25.992 + 1.638 \\ &+ 0.314 + 2.592 + 11.834 = \textcolor{#b54708}{85.404}\end{aligned} \]
Add the six rows
Why: s² needs the table's total
\[ \begin{aligned}&85.404 - 11.834 - 2.592 - 0.314 \\ &- 1.638 - 25.992 = 43.034\end{aligned} \]
Check: subtract back to row 6
Why: lands on the model row's value
Faded example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank
Try It 2.33: 29 stores carry 6 to 12 food types; x̄ ≈ 9.28.
First six rows total 85.404.
Fill in the blanks
6(12 − 9.28)² ≈ 44.39; total ≈ 129.79; s² = total ÷ 28 ≈ 4.64; s ≈ 2.15
Why: 12 − 9.28 = 2.72, 2.72² = 7.3984, and 6 × 7.3984 ≈ 44.39. Then 85.404 + 44.390 ≈ 129.79. 29 stores are a sample, so divide by n − 1 = 28 to get about 4.64, then take the square root: about 2.15.
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, the rest left blank
\[ 12 - 9.28 = \textcolor{#1f5fbf}{2.72} \]
Subtract x̄ from 12
Why: spread is measured from the mean
\[ \textcolor{#1f5fbf}{2.72}^2 = \textcolor{#b54708}{7.3984} \]
Square the distance
Why: far rows weigh more
\[ 6 \times 7.3984 \approx \textcolor{#b54708}{44.390} \]
Multiply by f = 6
Why: f equal squares
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390
\[ \sqrt{44.390 \div 6} + 9.28 \approx 2.72 + 9.28 = 12 \]
Check: undo all three moves
Why: lands back on x = 12
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390
\[ 85.404 + 44.390 = \textcolor{#b54708}{129.794} \]
Add to the first six
Why: s² needs all 29 stores
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794
\[ \begin{aligned}&(43.034 + 25.992 + 1.638) \\ &+ (0.314 + 2.592 + 11.834 + 44.390)\end{aligned} \]
Regroup the seven rows
Why: a second route to the total
\[ = 70.664 + 59.130 \]
Add inside each bracket
Why: gives the check two subtotals
\[ 70.664 + 59.130 = \textcolor{#b54708}{129.794} \]
Check in two halves
Why: matches the table's total
Worked example
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, not yet filled
\[ 29 - 1 = 28 \]
Find n − 1 for 29 stores
Why: a sample of the whole chain
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28
\[ \textcolor{#b54708}{s^2} = 129.794 \div 28 \]
Substitute the total into s²
Why: so 28 can share it
\[ \approx \textcolor{#b54708}{4.64} \]
Divide by 28
Why: n − 1: fair for a sample
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28, s² ≈ 4.64
\[ \textcolor{#1f5fbf}{s} = \sqrt{4.64} \]
Substitute into s = √s²
Why: the side of that area
\[ \approx \textcolor{#1f5fbf}{2.15} \]
Take the root
Why: back in food types
Figure (svg): Table of Try It 2.33's pet-food stores: food types x from 6 to 12, counts f 4, 5, 1, 4, 5, 4, 6, weighted squares f(x − x̄)² 43.034, 25.992, 1.638, 0.314, 2.592, 11.834, 44.390; totals: 29 stores, 129.794; result rows for n − 1, s² and s, filled so far: n − 1 = 28, s² ≈ 4.64, s ≈ 2.15
\[ \textcolor{#1f5fbf}{2.15}^2 \times 28 \approx \textcolor{#b54708}{129.4} \]
Check: undo both moves
Why: near 129.79: rounding
Section
Idea 4 of 5
Concept
Figure (svg): Number line of waits 0 to 10 with the mean 5 and Binh's wait of 1 marked; no SD hops yet
A bakery's recorded waits: x̄ = 5 minutes.
Discussion prompt
Binh waited 1 minute. Is 4 minutes from the mean unusual?
Answer:
It depends how far waits typically stray: here s = 2 minutes.
Mark hops of s = 2 minutes out from the mean.
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
Worked example
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
Rosa waited 7 minutes, Binh 1 minute; x̄ = 5, s = 2.
\[ 7 - 5 = \textcolor{#1f5fbf}{2} \]
Find Rosa's distance
Why: hop counts need the gap first
\[ \textcolor{#1f5fbf}{2} \div \textcolor{#1f5fbf}{2} = 1 \]
Divide by the SD
Why: counts 2-minute hops
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
\[ 1 - 5 = \textcolor{#1f5fbf}{-4} \]
Find Binh's distance
Why: a negative gap marks the left side
\[ \textcolor{#1f5fbf}{-4} \div \textcolor{#1f5fbf}{2} = -2 \]
Divide by the SD
Why: a signed count keeps the direction
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
\[ 5 + (-2)(\textcolor{#1f5fbf}{2}) = 1 \]
Check: hop back to Binh
Why: two hops down land on 1
Worked example
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
\[ x - \bar{x} \]
Measure the distance
Why: what the hops must cover
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
\[ k = \frac{\textcolor{#1f5fbf}{x - \bar{x}}}{\textcolor{#1f5fbf}{s}} \]
Divide by s
Why: k counts hops, signed
Figure (svg): Number line of waits 0 to 10 with SD hops of 2 from the mean 5
\[ k\,\textcolor{#1f5fbf}{s} = x - \bar{x} \]
Multiply both sides by s
Why: k hops cover the gap
\[ x = \bar{x} + k\,\textcolor{#1f5fbf}{s} \]
Add x̄ to both sides
Why: start at the mean, hop k times
\[ 5 + (-2)(\textcolor{#1f5fbf}{2}) = 1 \]
Check with Binh
Why: lands on 1 minute
Worked example
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
x̄ = 10.53 years, s = 0.72 years (Example 2.32).
\[ (1)(\textcolor{#1f5fbf}{0.72}) = 0.72 \]
Multiply k = 1 by s
Why: hops are measured in SDs
\[ 10.53 + 0.72 = 11.25 \]
Add it to x̄
Why: hops start at the mean
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
\[ (-2)(\textcolor{#1f5fbf}{0.72}) = -1.44 \]
Multiply k = −2 by s
Why: negative k hops down
\[ 10.53 + (-1.44) = 9.09 \]
Add it to x̄
Why: x = x̄ + ks places the value
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
\[ (9.09 - 10.53) \div \textcolor{#1f5fbf}{0.72} = -2 \]
Check: count hops back
Why: should be two down
Worked example
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
\[ (-1.5)(\textcolor{#1f5fbf}{0.72}) = -1.08 \]
Multiply k = −1.5 by s
Why: k needn't be whole
\[ 10.53 + (-1.08) = 9.45 \]
Add it to x̄
Why: hops start at x̄
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
\[ (1.5)(\textcolor{#1f5fbf}{0.72}) = 1.08 \]
Multiply k = 1.5 by s
Why: opposite sign, mirror hop
\[ 10.53 + 1.08 = 11.61 \]
Add it to x̄
Why: hops leave from the mean
Figure (svg): Ages number line from 8.5 to 12.5 with SD hops of 0.72 from the mean 10.53
\[ \frac{9.45 + 11.61}{2} = 10.53 \]
Check the pair's centre
Why: equal hops balance at x̄
Prediction
Predict first
Adapted from Try It 2.32: a team's mean age is 30.68. An age of 42.86 would sit exactly 2 SDs above it.
What is the team's SD?
Correct: 6.09 years
Why: The gap is 42.86 − 30.68 = 12.18 years, and it is two SDs long, so one SD is 12.18 ÷ 2 = 6.09.
Worked example
Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked
\[ 42.86 = 30.68 + 2\textcolor{#1f5fbf}{s} \]
Substitute into x = x̄ + ks
Why: leaves s as the only unknown
Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled s
\[ 12.18 = 2\textcolor{#1f5fbf}{s} \]
Subtract 30.68 from both sides
Why: isolates the two hops
Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled s; an arrow over the whole gap labelled 12.18 = 2s
\[ \textcolor{#1f5fbf}{s} = 6.09 \]
Divide both sides by 2
Why: equal division keeps both sides equal
Figure (svg): Number line of ages 24 to 48 with the team mean 30.68 and the age 42.86 marked; two equal hops from the mean to 42.86, each labelled 6.09; an arrow over the whole gap labelled 12.18 = 2s; a blue tick one SD above the mean
\[ 30.68 + 2(\textcolor{#1f5fbf}{6.09}) = 42.86 \]
Check by hopping back
Why: lands on the age 42.86
Concept
Figure (svg): Number line from 4 to 16 with μ = 10 and the band under 2σ = 4 shaded; no data shown
Report: N 8, μ 10, σ 2. Claim: 3 waits sit 4+ minutes from μ.
Discussion prompt
Without the data, could that claim be true?
Answer:
Build data with σ = 2 and count how many far waits fit.
Worked example
Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes
Store C's waits: 6, six 10s, 14.
\[ 6 \times \textcolor{#1f5fbf}{10} = 60 \]
Total the six 10s
Why: one product, not six sums
\[ \textcolor{#1f5fbf}{6} + 60 + \textcolor{#1f5fbf}{14} = 80 \]
Add the end waits
Why: the mean needs the total
\[ \mu = 80 \div 8 = 10 \]
Divide by N = 8
Why: tests the report's μ
Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked
\[ (\textcolor{#1f5fbf}{6} - 10) + (\textcolor{#1f5fbf}{14} - 10) = 0 \]
Check: the ends balance
Why: deviations must total zero
Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked and arrows of −4 and +4 from μ to the end waits
Worked example
Figure (svg): Dot plot of store C's waits: 6, six 10s and 14 minutes, with the mean μ = 10 marked
Store C: 6, six 10s, 14; μ = 10.
\[ \textcolor{#1f5fbf}{-4,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ +4} \]
Subtract μ
Why: spread starts from distances
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: each distance drawn as a blue side, no squares yet
\[ \textcolor{#b54708}{16,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ 16} \]
Square each
Why: σ² is built from these
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16
\[ \textcolor{#b54708}{16 + 16 = 32} \]
Add the squares
Why: σ² shares this total
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32
Fill in the blanks
Your turn: is the report right? Store C's σ² = 4 and its σ = 2 minutes.
Why: 32 ÷ 8 = 4 is the average square, and √4 = 2 minutes: the report's σ = 2 holds.
Worked example
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32
\( {-4,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ +4}\;\;\Rightarrow\;\;\allowbreak {16,\ 0,\ 0,\ 0,\ 0,\ 0,\ 0,\ 16}\;\;\Rightarrow\;\;\allowbreak {16 + 16 = 32} \)
\[ \textcolor{#b54708}{\sigma^2} = 32 \div 8 = \textcolor{#b54708}{4} \]
Divide by N = 8
Why: σ² is the mean square
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32; the average square, area 4
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4} = \textcolor{#1f5fbf}{2} \]
Take the root
Why: back to minutes
Figure (svg): Store C's distances from μ = 10: −4, six 0s, +4: a square on each distance, areas 16, 0, 0, 0, 0, 0, 0, 16; label: total area 32; the average square, area 4, side 2
\[ \textcolor{#1f5fbf}{2}^2 \times 8 = 32 \]
Check: square, times 8
Why: back to the total
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each
\[ 4^2 = \textcolor{#b54708}{16} \]
Square the least far distance
Why: the smallest square a far wait can have
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16
\[ \text{far total} \ge \textcolor{#b54708}{16} + \textcolor{#b54708}{16} + \textcolor{#b54708}{16} \]
Add three floors
Why: sums keep ≥
\[ \text{far total} \ge \textcolor{#b54708}{48} \]
Add 16 + 16 + 16
Why: one floor for the far area
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48
\[ \text{all 8 squares} \ge \textcolor{#b54708}{48} \]
Add the five other squares
Why: none is negative
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48
\[ 3 \times \textcolor{#b54708}{16} = \textcolor{#b54708}{48} \]
Check: three equal floors
Why: same total as adding
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48
\( {\text{far total} \ge \textcolor{#b54708}{48}}\;\;\Rightarrow\;\;\allowbreak {\text{all 8 squares} \ge \textcolor{#b54708}{48}} \)
\[ \text{all 8 squares} \div 8 \ge \textcolor{#b54708}{48} \div 8 \]
Divide both sides by N = 8
Why: a positive keeps ≥
\[ \textcolor{#b54708}{\sigma^2} \ge \textcolor{#b54708}{48} \div 8 \]
Substitute σ²
Why: the mean square
\[ \textcolor{#b54708}{\sigma^2} \ge \textcolor{#b54708}{6} \]
Evaluate 48 ÷ 8
Why: a floor to compare with σ² = 4
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48; label: σ² ≥ 6
\[ \textcolor{#b54708}{6} > \textcolor{#b54708}{4} \]
Compare with σ² = 4
Why: a floor can't exceed it
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48; label: σ² ≥ 6, but σ² = 4
\[ \textcolor{#b54708}{48} > \textcolor{#b54708}{32} = 8\textcolor{#b54708}{\sigma^2} \]
Check with areas
Why: part can't beat the whole
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a dashed third far square of area 16; label: far total ≥ 48; label: all 8 squares ≥ 48 > 32; label: σ² ≥ 6, but σ² = 4; a box of all the squared area 32 cut into 8 strips
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded
For a population, x = μ + kσ: the same hop rule.
\[ |\textcolor{#1f5fbf}{x - \mu}| \ge \textcolor{#1f5fbf}{k\sigma} \]
Define far
Why: names the values the bound will count
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14
\[ \textcolor{#b54708}{|x - \mu|^2} \ge \textcolor{#b54708}{(k\sigma)^2} \]
Square both sides
Why: neither side is negative
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; unfilled 4-by-4 squares on the two far distances
\[ \textcolor{#b54708}{(x - \mu)^2} \ge \textcolor{#b54708}{(k\sigma)^2} \]
Drop the bars
Why: squaring already removed the sign
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the two far squares filled, area 16 each
\[ \textcolor{#b54708}{(x - \mu)^2} \ge k^2\textcolor{#b54708}{\sigma^2} \]
Square kσ
Why: square each factor
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the two far squares filled, area 16 each; label: each far square ≥ k²σ² = 16
\[ k = 2:\ (6 - 10)^2 = 16 \ge 4 \cdot 4 \]
Check store C's wait of 6
Why: its square meets the floor
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each
\( {|x - \mu| \ge k\sigma}\;\;\Rightarrow\;\;\allowbreak {|x - \mu|^2 \ge (k\sigma)^2}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 \ge (k\sigma)^2}\;\;\Rightarrow\;\;\allowbreak {(x - \mu)^2 \ge k^2\sigma^2} \)
m counts far values, |x − μ| ≥ kσ; store C's m = 2.
\[ \textcolor{#b54708}{\text{far total}} = {\textstyle\sum}_{\text{far}} \textcolor{#b54708}{(x - \mu)^2} \]
Name the far total
Why: the floor bounds it
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: far total = 32
\[ \textcolor{#b54708}{\text{far total}} \ge \textcolor{#b54708}{{\textstyle\sum}_{\text{far}} k^2\sigma^2} \]
Add the m floors
Why: sums of ≥ keep ≥
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: each far square ≥ k²σ² = 16; label: far total = 32
\[ \textcolor{#b54708}{\text{far total}} \ge \textcolor{#b54708}{m\,k^2\sigma^2} \]
Add k²σ² m times
Why: Σc = mc collapses the sum
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; label: far total = 32; label: far total ≥ 2 × 16
\[ k = 2:\ 16 + 16 \ge 2 \cdot 4 \cdot 4 \]
Check store C
Why: far waits meet the floor
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each
\[ \textcolor{#b54708}{\sigma^2} = \frac{{\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2}}{N} \]
Write σ²'s definition
Why: the average square over all N
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, one strip of σ² = 4 filled
\[ \textcolor{#b54708}{N\sigma^2} = {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} \]
Multiply both sides by N
Why: undoes the average
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled
\[ \text{far total} \le {\textstyle\sum} \textcolor{#b54708}{(x - \mu)^2} \]
Compare with all squares
Why: the others add zero or more
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box
\[ \text{far total} \le \textcolor{#b54708}{N\sigma^2} \]
Substitute the total area
Why: a cap from σ alone
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: far total ≤ 32
\[ \textcolor{#b54708}{32} \le 8(4) = \textcolor{#b54708}{32} \]
Check with store C
Why: its far squares use everything
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled
\( {\text{far total} \ge m\,k^2\sigma^2}\qquad\allowbreak {\text{far total} \le N\sigma^2} \)
\[ m\,k^2\textcolor{#b54708}{\sigma^2} \le \textcolor{#b54708}{N\sigma^2} \]
Chain floor and cap
Why: a floor can't exceed its cap
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: 2 × 16 ≤ 32
\[ \frac{m\,k^2\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \le \frac{N\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \]
Divide by Nk²σ²
Why: k, σ > 0 keep it positive
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box, each covering 4 strips; label: 2 × 16 ≤ 32
\[ \frac{m}{N} \le \frac{N\textcolor{#b54708}{\sigma^2}}{N k^2\textcolor{#b54708}{\sigma^2}} \]
Cancel k²σ² on the left
Why: a common factor divides out
\[ \frac{m}{N} \le \frac{1}{k^2} \]
Cancel Nσ² on the right
Why: leaves only counts and k
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; the two far squares filled, area 16 each; a box of all the squared area 32 cut into 8 strips, every strip of σ² = 4 filled; the two far squares fitted inside the box; label: 2 × 16 ≤ 32; label: far share 2/8 ≤ 1/4
\[ k = 2:\ \frac{2}{8} \le \frac{1}{4} \]
Check with store C
Why: met exactly at k = 2
Worked example
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded
\( {\tfrac{m}{N} \le \tfrac{1}{\textcolor{#b54708}{k^2}}} \)
\[ 1 - \tfrac{m}{N} \ge 1 - \tfrac{1}{\textcolor{#b54708}{k^2}} \]
Subtract both sides from 1
Why: so the order reverses
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; label: far ≤ 1/4, so within ≥ 3/4
\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = \tfrac{N - m}{N} \]
Count those within
Why: all but the m far
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the six waits within 2σ ringed; label: within 8 − 2 = 6; label: far ≤ 1/4, so within ≥ 3/4
\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = \tfrac{N}{N} - \tfrac{m}{N} \]
Split the fraction
Why: over a shared N
\[ \text{within } \textcolor{#1f5fbf}{k\sigma} = 1 - \tfrac{m}{N} \]
Evaluate N/N = 1
Why: any nonzero count
\[ \text{within } \textcolor{#1f5fbf}{k\sigma} \ge 1 - \tfrac{1}{\textcolor{#b54708}{k^2}} \]
Replace 1 − m/N
Why: its bound carries over
Figure (svg): Store C's waits 6, six 10s and 14 on a number line with μ = 10; the band under 2σ shaded; arrows of 4 = 2σ from μ to the waits 6 and 14; the six waits within 2σ ringed; label: within 8 − 2 = 6; label: within 6/8 ≥ 3/4
\[ k = 2:\ \tfrac{6}{8} = 1 - \tfrac{1}{4} \]
Check store C
Why: six of eight
Concept
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5
Try k = 3, k = 4.5 (near 95%) and k = 1.
\[ 3^2 = 9,\ \ 4.5^2 = 20.25,\ \ 1^2 = 1 \]
Square each k
Why: three guarantees to compare
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5
\[ \tfrac{1}{9} \approx 0.11,\ \ \tfrac{1}{20.25} \approx 0.05,\ \ \tfrac{1}{1} = 1 \]
Take each reciprocal
Why: the far-share cap is 1/k²
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5; the dashed far-share cap curve 1/k² with points at 100%, 11% and 5%
\[ 1 - 0.11 = 0.89,\ 1 - 0.05 = 0.95,\ 1 - 1 = 0 \]
Subtract each from 1
Why: within plus far makes 1
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5; dashed guides at k = 1, 3 and 4.5; the dashed far-share cap curve 1/k² with points at 100%, 11% and 5%; points on the within curve at 0%, 89% and 95%
\[ 9 \times 0.11 \approx 1,\ \ 20.25 \times 0.05 \approx 1 \]
Check: undo each reciprocal
Why: both return about 1
Trap
\[ \text{guess: closer than } 2\sigma \approx 95\% \]
Guess from a bell curve
Why: ignores store C's shape
\[ \text{far} \approx 100\% - 95\% = 5\% \]
Subtract from 100%
Why: shares sum to 100%
\[ 8 \times 5\% = 0.4 \]
Take 5% of 8
Why: fails: store C has 2
Book p. 117 (quoted): ≈ 95% within 2σ, bell shapes only
\[ \tfrac{m}{8} \le \tfrac{1}{2^2} \]
Use Chebyshev, N = 8, k = 2
Why: holds for any shape
\[ \tfrac{m}{8} \le \tfrac{1}{4} \]
Evaluate 2²
Why: gives a usable cap
\[ m \le 2 \]
Multiply both sides by 8
Why: turns the share into a count
\[ |6 - 10| = |14 - 10| = 4 = 2(2) \]
Check store C's far waits
Why: both sit 2σ out
Prediction
Predict first
200 house prices, skewed by mansions.
By Chebyshev, for which k can 40% sit at least k SDs out?
Correct: k ≤ 1.58
Why: The far-share cap 1/k² must allow 0.4, so k² ≤ 2.5 and k ≤ √2.5 ≈ 1.58. Chebyshev only rules out k > 1.58, whatever the shape.
Worked example
Figure (svg): Curve of the far-share cap 1/k² for k from 1 to 4 SDs, falling from 100% to about 6%, with a dashed line at the claimed 40%
\[ m = 0.4 \times 200 = 80 \]
Turn 40% into a count
Why: the cap's m counts values
\[ \tfrac{80}{200} \le \tfrac{1}{k^2} \]
Substitute into m/N ≤ 1/k²
Why: the cap, any shape
\[ 0.4 \le \tfrac{1}{k^2} \]
Evaluate 80 ÷ 200
Why: the cap must allow this share
\[ k^2 \le \tfrac{1}{0.4} \]
Take reciprocals
Why: the order reverses
\[ k^2 \le 2.5 \]
Evaluate 1 ÷ 0.4
Why: biggest k² possible
\[ k \le \sqrt{2.5} \approx 1.58 \]
Take the root
Why: k counts SDs, so positive
Figure (svg): Curve of the far-share cap 1/k² for k from 1 to 4 SDs, falling from 100% to about 6%, with a dashed line at the claimed 40%; the crossing marked at k ≈ 1.58
\[ \tfrac{1}{1.58^2} \approx 0.40 \]
Check by substituting k = 1.58
Why: the cap just allows 40%
Section
Idea 5 of 5
Concept
Figure (svg): Two raw rulers: John's school GPA 1.5 to 4.5 with John at 2.85; Ali's school 50 to 110 with Ali at 77
John: 2.85, school mean 3.0, SD 0.7. Ali: 77, school mean 80, SD 10.
Discussion prompt
Relative to his own school, whose GPA is stronger? Commit first.
Answer:
John's. The next slides measure each gap in its school's SDs to show why raw gaps mislead.
Worked example
Figure (svg): Two raw rulers at their natural scales
\[ \textcolor{#1f5fbf}{2.85} - 3.0 = -0.15 \]
Find John's raw gap
Why: the gap each ruler must measure
\[ \textcolor{#1f5fbf}{77} - 80 = -3 \]
Find Ali's raw gap
Why: to set beside John's −0.15
\[ 3.0 + (-0.15) = \textcolor{#1f5fbf}{2.85},\ \ 80 + (-3) = \textcolor{#1f5fbf}{77} \]
Check: add each gap to its mean
Why: recovers both grades
Different units: stretch each ruler so one SD matches.
Figure (svg): Both rulers stretched so one SD has the same length, with a shared axis of SDs from −3 to +3
Worked example
Figure (svg): Both rulers stretched so one SD has the same length, with a shared axis of SDs
\( {{2.85 - 3.0 = -0.15}}\qquad\allowbreak {{77 - 80 = -3}} \)
\[ z = \frac{\textcolor{#1f5fbf}{x - \mu}}{\textcolor{#1f5fbf}{\sigma}} \]
Name that hop count z
Why: population μ, σ; samples use x̄, s
\[ z_J = \tfrac{-0.15}{\textcolor{#1f5fbf}{0.7}},\ \ z_A = \tfrac{-3}{\textcolor{#1f5fbf}{10}} \]
Substitute both students
Why: each school's own σ
\[ z_J \approx -0.21,\ \ z_A = -0.30 \]
Divide each gap
Why: one shared SD ruler
Figure (svg): Both rulers aligned in SDs, John at −0.21 and Ali at −0.30
\[ -0.21 > -0.30 \]
Compare the z-scores
Why: fewer SDs below means relatively higher
\[ (-0.21)(\textcolor{#1f5fbf}{0.7}) \approx -0.15,\ (-0.30)(\textcolor{#1f5fbf}{10}) = -3 \]
Check: z times σ
Why: gives back both gaps
Concept
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed
\[ \tfrac{\text{GPA pts}}{\text{GPA pts}} = \text{no units} \]
Check z's units
Why: so GPAs and scores compare
\[ x = \mu \Rightarrow \textcolor{#1f5fbf}{x - \mu} = 0 \]
Set x = μ
Why: find z's zero
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ
\[ z = \tfrac{\textcolor{#1f5fbf}{0}}{\textcolor{#1f5fbf}{\sigma}} = 0 \]
Divide by σ
Why: 0 over a positive is 0
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ
\[ x < \mu \Leftrightarrow \textcolor{#1f5fbf}{x - \mu} < 0 \]
Subtract μ
Why: smaller minus larger is negative
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0
\[ \textcolor{#1f5fbf}{x - \mu} < 0 \Leftrightarrow \tfrac{\textcolor{#1f5fbf}{x - \mu}}{\sigma} < 0 \]
Divide by σ > 0
Why: positives keep order
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ
\[ z < 0 \Leftrightarrow x < \mu \]
Chain both
Why: equivalences link end to end
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ; label: z < 0: below μ
\[ 2.85 < 3.0,\ z \approx -0.21 < 0 \]
Check with John
Why: sign matches his side
Figure (svg): GPA ruler for John's school with ticks every σ = 0.7 from 1.6 to 4.4 and the mean 3.0 dashed; a dot at x = μ; z = 0 written under μ; a blue arrow leftward from μ labelled x − μ < 0; z labels −1 and −2 under the ticks below μ; label: z < 0: below μ; John's 2.85 marked just left of μ
Worked example
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie 26.2 marked; Beth 27.3 marked
Angie: 26.2 s; team mean 27.2 s.
\[ \textcolor{#1f5fbf}{26.2} - 27.2 = -1.0 \]
Subtract Angie's team mean
Why: z needs the gap
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie's raw gap −1.0 s drawn from the mean; Angie 26.2 marked; Beth 27.3 marked
\[ z_{\text{Angie}} = \frac{-1.0}{\textcolor{#1f5fbf}{0.8}} \]
Substitute into z = gap ÷ SD
Why: her own team's ruler
\[ z_{\text{Angie}} = -1.25 \]
Divide −1.0 by 0.8
Why: to compare swims
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth 27.3 marked
\[ 27.2 + (-1.25)(\textcolor{#1f5fbf}{0.8}) = 26.2 \]
Check: hop back
Why: recovers Angie's time
Worked example
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth 27.3 marked
Beth: 27.3 s; team mean 30.1 s.
\[ \textcolor{#1f5fbf}{27.3} - 30.1 = -2.8 \]
Subtract Beth's team mean
Why: z needs her gap
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth's raw gap −2.8 s drawn from the mean; Beth 27.3 marked
\[ z_{\text{Beth}} = \frac{-2.8}{\textcolor{#1f5fbf}{1.4}} \]
Substitute into z = gap ÷ SD
Why: her team spreads wider
\[ z_{\text{Beth}} = -2.00 \]
Divide −2.8 by 1.4
Why: z puts both on one scale
Figure (svg): Two swim-time rulers drawn so one SD has the same length: Angie's team mean 27.2 s with SD ticks every 0.8 s; Beth's team mean 30.1 s with SD ticks every 1.4 s; Angie: z = −1.25, drawn as hops of one SD from the mean; Angie 26.2 marked; Beth: z = −2, drawn as hops of one SD from the mean; Beth 27.3 marked
\[ 30.1 + (-2)(\textcolor{#1f5fbf}{1.4}) = 27.3 \]
Check: hop back
Why: recovers Beth's time
Trap
\[ \text{Angie } {-1.25},\ \text{Beth } {-2.00} \]
Start from both z-scores
Why: each on her team's ruler
\[ -1.25 > -2.00 \Rightarrow \text{Angie} \]
Pick the larger z
Why: fails: a smaller time is faster
\[ \text{time: lower is better} \]
Ask which direction wins
Why: fewer seconds means faster
\[ -2.00 < -1.25 \Rightarrow \text{Beth} \]
Pick the lower z
Why: Beth is further below her mean
\[ 30.1 + (-2)(1.4) = 27.3 \]
Check Beth's z by hopping back
Why: recovers her time
Prediction
Predict first
4.0 kg newborns at hospitals A, B: means 3.5, SDs 0.25, 0.5.
Which is heavier for its hospital?
Correct: Hospital A's
Why: Hospital A's weights typically stray about 0.25 kg from 3.5, so 0.5 kg up is 2 SDs. Hospital B's typically stray about 0.5 kg, so 0.5 kg up is 1 SD.
Worked example
Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; A's baby 4.0 marked; B's baby 4.0 marked
\[ 4.0 - 3.5 = \textcolor{#1f5fbf}{0.5} \]
Find the gap
Why: z starts from the raw gap
Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A's raw gap +0.5 kg drawn from the mean; A's baby 4.0 marked; Hospital B's raw gap +0.5 kg drawn from the mean; B's baby 4.0 marked
\[ z_A = \textcolor{#1f5fbf}{0.5} \div \textcolor{#1f5fbf}{0.25} \]
Substitute A's SD
Why: z uses each hospital's ruler
\[ z_A = 2 \]
Divide 0.5 by 0.25
Why: counts SD hops
Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A: z = 2, drawn as hops of one SD from the mean; A's baby 4.0 marked; Hospital B's raw gap +0.5 kg drawn from the mean; B's baby 4.0 marked
\[ z_B = \textcolor{#1f5fbf}{0.5} \div \textcolor{#1f5fbf}{0.5} \]
Substitute B's SD
Why: same gap, wider ruler
\[ z_B = 1 \]
Divide 0.5 by 0.5
Why: now both in SDs
Figure (svg): Two newborn-weight rulers in kg drawn so one SD has the same length, both with mean 3.5 kg: hospital A with SD ticks every 0.25 kg, hospital B with SD ticks every 0.5 kg; Hospital A: z = 2, drawn as hops of one SD from the mean; A's baby 4.0 marked; Hospital B: z = 1, drawn as hops of one SD from the mean; B's baby 4.0 marked
\[ 2 > 1 \]
Compare hop counts
Why: A's baby sits further out
\[ 3.5 + (2)(\textcolor{#1f5fbf}{0.25}) = 4.0,\ \ 3.5 + (1)(\textcolor{#1f5fbf}{0.5}) = 4.0 \]
Check: hop back
Why: each lands on 4.0 kg
Faded example
\[ x = \bar{x} + z\,\textcolor{#1f5fbf}{s} \]
Given: z counts hops like k
Why: value formula x = x̄ + ks, k = z
Fill in the blanks
Lee: z = 1.6, mean 80, SD 5 → score = 88. Kai: z = 0.92, mean 73.5, SD 17.9 → score ≈ 90.
Why: Start at each class mean and hop z SDs: 80 + 1.6 × 5 = 88 and 73.5 + 0.92 × 17.9 ≈ 90. Kai scored higher yet stands out less.
Worked example
Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9
\[ (1.6)(\textcolor{#1f5fbf}{5}) = 8,\ \ (0.92)(\textcolor{#1f5fbf}{17.9}) \approx 16.47 \]
Multiply each z by its SD
Why: undoes dividing by the SD
Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9; Lee: 1.6 hops = 8, drawn as hops of one SD from the mean; Kai: 0.92 hops ≈ 16.47, drawn as hops of one SD from the mean
\[ 80 + 8 = 88,\ \ 73.5 + 16.47 \approx 90 \]
Add each to its class mean
Why: the starting point of each hop
Figure (svg): Two score rulers drawn so one SD has the same length: Lee's class mean 80 with SD ticks every 5; Kai's class mean 73.5 with SD ticks every 17.9; Lee: 1.6 hops = 8, drawn as hops of one SD from the mean; Lee 88 marked; Kai: 0.92 hops ≈ 16.47, drawn as hops of one SD from the mean; Kai ≈ 90 marked
\[ (88 - 80) \div \textcolor{#1f5fbf}{5} = 1.6 \]
Check Lee's z
Why: 8 over 5 gives back 1.6
\[ (90 - 73.5) \div \textcolor{#1f5fbf}{17.9} \approx 0.92 \]
Check Kai's z
Why: 16.5 over 17.9 gives back 0.92
Pattern
Figure (svg): Squares on the deviations -3, -1, 0, 1, 3, side = distance, area = square
\[ \textcolor{#1f5fbf}{s} = \sqrt{\textcolor{#b54708}{s^2}},\quad x = \bar{x} + k\,\textcolor{#1f5fbf}{s},\quad z = \frac{\textcolor{#1f5fbf}{x - \bar{x}}}{\textcolor{#1f5fbf}{s}} \]
Concept
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16
Example 2.34: 26 values in classes 0–2 to 15–17.
Discussion prompt
A table gives only classes, not values. What stands in for each value?
Answer:
The class midpoint: an estimate.
Worked example
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16
f = 1, 6, 10, 7, 0, 2 at midpoints 1, 4, 7, 10, 13, 16.
\[ \textcolor{#1f5fbf}{1,\ 24,\ 70,\ 70,\ 0,\ 32} \]
Multiply each midpoint by its f
Why: raw values unknown; centres stand in
\[ \textcolor{#1f5fbf}{1} + \textcolor{#1f5fbf}{24} + \textcolor{#1f5fbf}{70} + \textcolor{#1f5fbf}{70} + \textcolor{#1f5fbf}{0} + \textcolor{#1f5fbf}{32} = 197 \]
Add the products
Why: estimates the total of all values
\[ \bar{x} = 197 \div 26 \approx 7.58 \]
Divide by n = 26
Why: the mean every row uses
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked
\[ 26 \times 7.58 = 197.08 \approx 197 \]
Check: undo the division
Why: near 197: rounding
Worked example
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked
Class 3–5: midpoint 4, f = 6; x̄ ≈ 7.58.
\[ \textcolor{#1f5fbf}{4} - 7.58 = \textcolor{#1f5fbf}{-3.58} \]
Subtract x̄ from the midpoint
Why: estimate: six values placed at 4
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4
\[ (\textcolor{#1f5fbf}{-3.58})^2 = \textcolor{#b54708}{12.8164} \]
Square the distance
Why: squares keep below-mean rows positive
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4; a square on the −3.58 distance, area 12.82
\[ 6 \times \textcolor{#b54708}{12.8164} = \textcolor{#b54708}{76.8984} \]
Multiply by f = 6
Why: all six values counted
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 3–5 highlighted with an arrow of −3.58 from x̄ to its midpoint 4; a square on the −3.58 distance, area 12.82; label: × 6 = 76.90 for the class's 6 values
\[ \sqrt{76.8984 \div 6} = \textcolor{#1f5fbf}{3.58} \]
Check by undoing both moves
Why: recovers the distance 3.58
Worked example
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked
\[ 7 - 7.58,\ 16 - 7.58 = \textcolor{#1f5fbf}{-0.58},\ \textcolor{#1f5fbf}{+8.42} \]
Subtract x̄ from each midpoint
Why: spread is measured from the mean
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16
\[ (\textcolor{#1f5fbf}{-0.58})^2,\ \textcolor{#1f5fbf}{8.42}^2 = \textcolor{#b54708}{0.3364},\ \textcolor{#b54708}{70.8964} \]
Square each distance
Why: far sides, big areas
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16; a square on the −0.58 distance, area 0.34; a square on the +8.42 distance, area 70.90
\[ 10 \times \textcolor{#b54708}{0.3364} = \textcolor{#b54708}{3.364} \]
Multiply by f = 10
Why: each value adds one square
\[ 2 \times \textcolor{#b54708}{70.8964} = \textcolor{#b54708}{141.7928} \]
Multiply by f = 2
Why: a square for each
Figure (svg): Example 2.34's classes 0–2, 3–5, 6–8, 9–11, 12–14, 15–17 as boxes on a number line from 0 to 18, with f = 1, 6, 10, 7, 0, 2 dots stacked at the midpoints 1, 4, 7, 10, 13, 16; the mean x̄ ≈ 7.58 marked; class 6–8 highlighted with an arrow of −0.58 from x̄ to its midpoint 7; class 15–17 highlighted with an arrow of +8.42 from x̄ to its midpoint 16; a square on the −0.58 distance, area 0.34; a square on the +8.42 distance, area 70.90; label: × 10 = 3.36 for the class's 10 values; label: × 2 = 141.79 for the class's 2 values
\[ \sqrt{3.364 \div 10} = 0.58,\ \ \sqrt{141.7928 \div 2} = 8.42 \]
Check: undo both moves
Why: returns both distances
Faded example
Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries blank, 76.8984, 3.3640, blank, 0, 141.7928
Class 0–2: midpoint 1, f = 1. Class 9–11: midpoint 10, f = 7.
Fill in the blanks
Fill the two remaining rows, with x̄ ≈ 7.58: class 0–2's f(mid − x̄)² ≈ 43.30 and class 9–11's ≈ 40.99.
Why: 1 − 7.58 = −6.58, squared 43.2964, times f = 1 ≈ 43.30. 10 − 7.58 = 2.42, squared 5.8564, times f = 7 ≈ 40.99.
Worked example
Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries blank, 76.8984, 3.3640, blank, 0, 141.7928
\[ 1 - 7.58,\ 10 - 7.58 = \textcolor{#1f5fbf}{-6.58},\ \textcolor{#1f5fbf}{+2.42} \]
Subtract x̄ from each midpoint
Why: gaps start at x̄
\[ (\textcolor{#1f5fbf}{-6.58})^2,\ \textcolor{#1f5fbf}{2.42}^2 = \textcolor{#b54708}{43.2964},\ \textcolor{#b54708}{5.8564} \]
Square each distance
Why: so below-mean rows can't cancel
\[ 1 \times \textcolor{#b54708}{43.2964} = \textcolor{#b54708}{43.2964} \]
Multiply by f = 1
Why: one value, one square
Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries 43.2964, 76.8984, 3.3640, blank, 0, 141.7928
\[ 7 \times \textcolor{#b54708}{5.8564} = \textcolor{#b54708}{40.9948} \]
Multiply by f = 7
Why: seven equal squares
Figure (svg): Table of Example 2.34's midpoints 1, 4, 7, 10, 13, 16 with f = 1, 6, 10, 7, 0, 2 and f(mid − x̄)² entries 43.2964, 76.8984, 3.3640, 40.9948, 0, 141.7928
\[ \sqrt{43.2964} = 6.58,\ \ \sqrt{40.9948 \div 7} = 2.42 \]
Check: reverse both moves
Why: back to both gaps
Check
Check your understanding
Class 15–17's pair, at midpoint 16, supplies 141.79 squared area (x̄ ≈ 7.58). What would true values 15 and 17 supply?
Answer: A
Why: 15 − 7.58 = 7.42, 7.42² = 55.0564; 17 − 7.58 = 9.42, 9.42² = 88.7364; total 143.7928 ≈ 143.79 — 2 more than the midpoint gives.
Worked example
Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; the true pair's distances still grey dashed, not yet measured
\[ 15 - 7.58,\ 17 - 7.58 = \textcolor{#1f5fbf}{7.42},\ \textcolor{#1f5fbf}{9.42} \]
Subtract x̄ from each
Why: no midpoint stand-in
Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; the true distances 7.42 and 9.42 drawn as blue sides
\[ \textcolor{#1f5fbf}{7.42}^2,\ \textcolor{#1f5fbf}{9.42}^2 = \textcolor{#b54708}{55.0564},\ \textcolor{#b54708}{88.7364} \]
Square each distance
Why: rounding would blur the error
Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364
\[ \textcolor{#b54708}{55.0564} + \textcolor{#b54708}{88.7364} = \textcolor{#b54708}{143.7928} \]
Add the two areas
Why: one value each, so no × f
Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364; the true pair's total 143.7928
\[ 143.7928 - 141.7928 = 2 \]
Subtract the midpoint row's area
Why: sizes the error
Figure (svg): Class 15–17's values 15 and 17 above, and below them the midpoint row: two squares of side 8.42, area 70.8964 each, total 141.7928; squares on 7.42 and 9.42 with areas 55.0564 and 88.7364; the true pair's total 143.7928; label: true pair: 2 more area
\[ (\textcolor{#b54708}{55.06} - \textcolor{#b54708}{70.90}) + (\textcolor{#b54708}{88.74} - \textcolor{#b54708}{70.90}) = 2 \]
Check against 16's square
Why: the gaps net 2
Worked example
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge
\[ \textcolor{#1f5fbf}{7.42} = 8.42 - 1 \]
Write 7.42 from 16's distance
Why: the value sits 1 inside
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 7.42 dashed inside it, labelled on its top edge
\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)(-1)} + \textcolor{#b54708}{(-1)^2} \]
Expand (a + b)²
Why: rule for a sum, b = −1
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 7.42 dashed inside it, labelled on its top edge; two blue 8.42-by-1 strips cut away and their solid orange 1-by-1 corner, cut twice
\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} - \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{(-1)^2} \]
Multiply 2(8.42) by −1
Why: sign rule: − × + = −
\[ (8.42 - 1)^2 = \textcolor{#b54708}{8.42^2} - \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{1^2} \]
Write (−1)² as 1²
Why: even powers drop the sign
\[ \textcolor{#b54708}{70.8964} - \textcolor{#1f5fbf}{16.84} + \textcolor{#b54708}{1} = \textcolor{#b54708}{55.0564} \]
Check: evaluate
Why: matches the true 7.42²
Worked example
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge
\[ \textcolor{#1f5fbf}{9.42} = 8.42 + 1 \]
Write 9.42 from 16's distance
Why: the value sits 1 beyond
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 9.42 dashed around it, labelled on its bottom edge
\[ (8.42 + 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)(1)} + \textcolor{#b54708}{1^2} \]
Expand (a + b)²
Why: rule for a sum, b = +1
Figure (svg): A light orange square of side 8.42 and area 70.8964, its side labelled on the edge; the side 9.42 dashed around it, labelled on its bottom edge; two blue 8.42-by-1 strips and a solid orange 1-by-1 corner added
\[ (8.42 + 1)^2 = \textcolor{#b54708}{8.42^2} + \textcolor{#1f5fbf}{2(8.42)} + \textcolor{#b54708}{1^2} \]
Drop the factor 1
Why: multiplying by 1 changes nothing
\[ \textcolor{#b54708}{70.8964} + \textcolor{#1f5fbf}{16.84} + \textcolor{#b54708}{1} = \textcolor{#b54708}{88.7364} \]
Check: evaluate
Why: matches the true 9.42²
Worked example
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it
\[ \begin{aligned}&\textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 8.42^2 - 2(8.42) + 1^2\\&\qquad + 8.42^2 + 2(8.42) + 1^2\end{aligned} \]
Substitute both
Why: a sum drops brackets
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added
\[ \begin{aligned}&= 8.42^2 + 8.42^2\\&\qquad - 2(8.42) + 2(8.42) + 1^2 + 1^2\end{aligned} \]
Reorder the terms
Why: addition runs in any order
\[ = 8.42^2 + 8.42^2 + 1^2 + 1^2 \]
Cancel −2(8.42) + 2(8.42)
Why: opposites sum to 0
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added
\[ \textcolor{#b54708}{70.8964} + \textcolor{#b54708}{70.8964} + 1 + 1 = \textcolor{#b54708}{143.7928} \]
Check: evaluate it
Why: matches the true pair
Worked example
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added
\( {\textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 8.42^2 + 8.42^2 + 1^2 + 1^2} \)
\[ \textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 = 2(8.42^2) + 2(1^2) \]
Collect equal terms
Why: so the midpoint area subtracts whole
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added; labels: left, two 8.42² squares and two 1² corners
\[ \textcolor{#1f5fbf}{7.42}^2 + \textcolor{#1f5fbf}{9.42}^2 - 2(8.42^2) = 2(1^2) \]
Subtract 2(8.42²) from both sides
Why: removes the midpoint pair's area
Figure (svg): Two light orange squares of side 8.42 and area 70.8964: the left for value 15 with its true side 7.42 dashed inside, the right for value 17 with its true side 9.42 dashed around it; on 15's square two blue 8.42-by-1 strips cut away and a solid orange 1-by-1 corner; on 17's square two blue strips and a corner added; label: strips cut = strips added; labels: left, two 8.42² squares and two 1² corners; label: error, the two corners
\[ 1^2 = 1 \]
Evaluate one corner, 1²
Why: a unit square of error
\[ \textcolor{#b54708}{143.7928} - \textcolor{#b54708}{141.7928} = 2(1) \]
Check against the true values' sum
Why: 15 and 17 gave 2 directly
Worked example
| mid | f | f(mid − x̄)² |
|---|---|---|
| 1 | 1 | 43.2964 |
| 4 | 6 | 76.8984 |
| 7 | 10 | 3.3640 |
| 10 | 7 | 40.9948 |
| 13 | 0 | 0 |
| 16 | 2 | 141.7928 |
| total | 26 |
\( {\bar{x} = 197 \div 26 \approx 7.58} \)
\[ \begin{aligned}&43.2964 + 76.8984 + 3.3640\\&+ 40.9948 + 0 + 141.7928 = \textcolor{#b54708}{306.3464}\end{aligned} \]
Add the six table rows
Why: s² is built on this total
\[ 26 - 1 = 25 \]
Find n − 1
Why: 26 values: a sample
\[ 306.3464 \div 25 \approx \textcolor{#b54708}{12.2539} \]
Divide by 25
Why: fair sample average
\[ \textcolor{#1f5fbf}{s} = \sqrt{12.2539} \approx \textcolor{#1f5fbf}{3.50} \]
Take the root
Why: back to table units
\[ 3.50^2 \times 25 \approx 306.25 \]
Check: undo both
Why: near 306.35: rounding
Check
Check your understanding
Bottle fills have mean 500 ml, σ = 4 ml, unknown shape. What band around the mean does Chebyshev need to guarantee at least 96%?
Answer: A
Why: 1 − 1/k² = 0.96, so 1/k² = 0.04, k² = 25, k = 5, and 5 × 4 = 20 ml.
Worked example
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%
\[ 1 - 1/k^2 = 0.96 \]
Set the floor to 96%
Why: the Chebyshev guarantee
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%, with a dashed line at the target 96%
\[ 1 = 0.96 + 1/k^2 \]
Add 1/k² to both sides
Why: frees the cap's sign
\[ 0.04 = 1/k^2 \]
Subtract 0.96 from both sides
Why: leaves the cap alone
\[ k^2 = 1/0.04 \]
Take reciprocals
Why: frees k² from the fraction
\[ k^2 = 25 \]
Evaluate 1 ÷ 0.04
Why: so a root gives k
\[ k = 5 \]
Take the root
Why: k counts SDs, so positive
Figure (svg): Curve of the guaranteed share within k SDs, 1 − 1/k², for k from 1 to 5, rising from 0% to 96%, with a dashed line at the target 96%; the crossing marked at k = 5
\[ 1 - 1/5^2 = 1 - 0.04 = 0.96 \]
Check: substitute k = 5
Why: gives back 96%
Worked example
Chebyshev: 96% needs k = 5 SDs. Fills: mean 500 ml, σ = 4 ml.
\[ 5 \times \textcolor{#1f5fbf}{4} = \textcolor{#1f5fbf}{20} \]
Multiply k by σ
Why: hops become ml
\[ 500 - \textcolor{#1f5fbf}{20} = 480,\ \ 500 + \textcolor{#1f5fbf}{20} = 520 \]
Step 20 ml each way from the mean
Why: the band is centred on it
\[ (520 - 480) \div 2 \div \textcolor{#1f5fbf}{4} = 5 \]
Check: half the band in SDs
Why: gives back k = 5
Check
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
Check your understanding
Population 2, 4, 5, 6, 8 (line A, σ = 2): which change raises σ most?
Answer: D
Why: Doubling doubles every deviation: σ = 2 × 2 = 4. 4→3, 6→7: squares 9, 4, 0, 4, 9 total 26; 26 ÷ 5 = 5.2; √5.2 ≈ 2.28. 5→6: mean 26 ÷ 5 = 5.2, squares total 20.8; 20.8 ÷ 5 = 4.16; √4.16 ≈ 2.04. A sixth 5: squares still total 20; 20 ÷ 6 ≈ 3.33; √3.33 ≈ 1.83.
Worked example
Figure (svg): Waits 2, 4, 5, 6, 8 on a number line with 5 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{4},\ \textcolor{#1f5fbf}{8},\ \textcolor{#1f5fbf}{10},\ \textcolor{#1f5fbf}{12},\ \textcolor{#1f5fbf}{16} \]
Double each wait
Why: test option D directly
Figure (svg): Dot plot of doubled waits: 4, 8, 10, 12, 16 minutes, no mean marked yet
\[ \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{8} + \textcolor{#1f5fbf}{10} + \textcolor{#1f5fbf}{12} + \textcolor{#1f5fbf}{16} = 50 \]
Add the doubled waits
Why: the mean needs a total
\[ \mu = 50 \div 5 = 10 \]
Divide by N = 5
Why: the centre for new distances
Figure (svg): Doubled waits 4, 8, 10, 12, 16 on a number line with 0 deviation arrows from the mean 10
\[ \textcolor{#1f5fbf}{-6,\ -2,\ 0,\ +2,\ +6} \]
Subtract 10 from each
Why: σ needs distances from the new mean
Figure (svg): Doubled waits 4, 8, 10, 12, 16 on a number line with 5 deviation arrows from the mean 10
\[ -6 - 2 + 0 + 2 + 6 = 0 \]
Check: deviations total zero
Why: confirms the centre 10
Worked example
Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: each distance drawn as a blue side, no squares yet
Doubled waits: μ = 10, distances −6, −2, 0, +2, +6.
\[ \textcolor{#b54708}{36,\ 4,\ 0,\ 4,\ 36} \]
Square each
Why: squares stop the signs cancelling
Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36
\[ \textcolor{#b54708}{36 + 4 + 0 + 4 + 36 = 80} \]
Add the areas
Why: σ² shares the total among 5
Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80
\[ 80 \div 5 = \textcolor{#b54708}{16} \]
Divide by N = 5
Why: σ² is the mean square
Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80; the average square, area 16
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{16} = \textcolor{#1f5fbf}{4} \]
Take the root
Why: back from area to minutes
Figure (svg): Doubled waits' distances from μ = 10: −6, −2, 0, +2, +6: a square on each distance, areas 36, 4, 0, 4, 36; label: total area 80; the average square, area 16, side 4
\[ 2 \times 2 = \textcolor{#1f5fbf}{4} \]
Check the doubling rule
Why: double distances, double σ
Faded example
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both
Fill in the blanks
Line A's 8 sat z = 1.5 SDs above μ = 5 (σ = 2). Doubled to 16, it sits z = 1.5 SDs above μ = 10 (σ = 4).
Why: (8 − 5) ÷ 2 = 1.5 and (16 − 10) ÷ 4 = 1.5: doubling moves the wait and the mean together and stretches σ by the same factor, so the SD count is unchanged.
Worked example
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both
\[ 8 - 5 = \textcolor{#1f5fbf}{3} \]
Subtract line A's μ
Why: z starts from the gap
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8
\[ \textcolor{#1f5fbf}{3} \div 2 = 1.5 \]
Divide by σ = 2
Why: counts the gap in SDs
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops
\[ 16 - 10 = \textcolor{#1f5fbf}{6} \]
Subtract the doubled μ
Why: z starts from the new gap
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops; an arrow of 6 from μ = 10 to 16
\[ \textcolor{#1f5fbf}{6} \div 4 = 1.5 \]
Divide by σ = 4
Why: σ doubled, so same count
Figure (svg): Two rulers: line A (0 to 10, μ = 5, σ = 2) with its wait 8 marked, and the doubled waits (0 to 20, μ = 10, σ = 4) with 16 marked, one σ the same length on both; an arrow of 3 from μ = 5 to 8; SD hops of 2 from 5: one full and one half, 1.5 hops; an arrow of 6 from μ = 10 to 16; SD hops of 4 from 10: one full and one half, 1.5 hops
\[ 5 + 1.5 \times 2 = 8,\ \ 10 + 1.5 \times 4 = 16 \]
Check: hop out from each μ
Why: 1.5 hops rebuild both waits
Worked example
Figure (svg): Dot plot of waits with the 5 changed to 6: 2, 4, 6, 6, 8 minutes, no mean marked yet
\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{8} = 26 \]
Add the new waits
Why: μ needs a fresh total
\[ \mu = 26 \div 5 = 5.2 \]
Divide by N = 5
Why: a mean is total over count
Figure (svg): Waits 2, 4, 6, 6, 8 on a number line with 0 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-3.2,\ -1.2,\ +0.8,\ +0.8,\ +2.8} \]
Subtract 5.2 from each
Why: every distance moves, not one
Figure (svg): Waits 2, 4, 6, 6, 8 on a number line with 5 deviation arrows from the mean
\[ -3.2 - 1.2 + 0.8 + 0.8 + 2.8 = 0 \]
Check: deviations total zero
Why: confirms the centre 5.2
Worked example
Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: each distance drawn as a blue side, no squares yet
New waits 2, 4, 6, 6, 8: μ = 5.2, distances −3.2, −1.2, +0.8, +0.8, +2.8.
\[ \textcolor{#b54708}{10.24,\ 1.44,\ 0.64,\ 0.64,\ 7.84} \]
Square each
Why: signs can no longer cancel
Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84
\[ \textcolor{#b54708}{10.24 + 1.44 + 0.64 + 0.64 + 7.84 = 20.8} \]
Add the areas
Why: σ² needs this total
Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8
\[ 20.8 \div 5 = \textcolor{#b54708}{4.16} \]
Divide by N = 5
Why: σ² is the mean square
Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8; the average square, area 4.16
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{4.16} \approx \textcolor{#1f5fbf}{2.04} \]
Take the root
Why: back to minutes
Figure (svg): Distances of the waits 2, 4, 6, 6, 8 from μ = 5.2: −3.2, −1.2, +0.8, +0.8, +2.8: a square on each distance, areas 10.24, 1.44, 0.64, 0.64, 7.84; label: total area 20.8; the average square, area 4.16, side 2.04
\[ \textcolor{#1f5fbf}{2.04}^2 \times 5 \approx 20.81 \]
Check: square, times N
Why: near 20.8: rounding
Worked example
Figure (svg): Dot plot of line a with a sixth wait of 5: 2, 4, 5, 6, 8, 5 minutes, no mean marked yet
\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{4} + \textcolor{#1f5fbf}{5} + \textcolor{#1f5fbf}{6} + \textcolor{#1f5fbf}{8} + \textcolor{#1f5fbf}{5} = 30 \]
Add all six waits
Why: μ must count the newcomer
\[ \mu = 30 \div 6 = 5 \]
Divide by N = 6
Why: a mean shares the total equally
Figure (svg): Waits 2, 4, 5, 6, 8, 5 on a number line with 0 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-3,\ -1,\ 0,\ +1,\ +3,\ 0} \]
Subtract 5 from each
Why: the newcomer sits on μ
Figure (svg): Waits 2, 4, 5, 6, 8, 5 on a number line with 6 deviation arrows from the mean
\[ -3 - 1 + 0 + 1 + 3 + 0 = 0 \]
Check: deviations total zero
Why: confirms the centre 5
Worked example
Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: each distance drawn as a blue side, no squares yet
Six waits: μ = 5, distances −3, −1, 0, +1, +3, 0.
\[ \textcolor{#b54708}{9,\ 1,\ 0,\ 1,\ 9,\ 0} \]
Square each
Why: the newcomer adds no area
Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0
\[ \textcolor{#b54708}{9 + 1 + 0 + 1 + 9 + 0 = 20} \]
Add the areas
Why: σ² needs the total squared distance
Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20
\[ 20 \div 6 \approx \textcolor{#b54708}{3.333} \]
Divide by N = 6
Why: more waits share it
Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20; the average square, area 3.333
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{3.333} \approx \textcolor{#1f5fbf}{1.83} \]
Take the root
Why: back to minutes
Figure (svg): Distances of the six waits 2, 4, 5, 6, 8, 5 from μ = 5: −3, −1, 0, +1, +3, 0: a square on each distance, areas 9, 1, 0, 1, 9, 0; label: total area 20; the average square, area 3.333, side 1.83
\[ \textcolor{#1f5fbf}{1.83}^2 \times 6 \approx 20.09 \]
Check: square, times N
Why: near 20: rounding
Worked example
Figure (svg): Dot plot of waits with the 4 and 6 moved out: 2, 3, 5, 7, 8 minutes, no mean marked yet
\[ \textcolor{#1f5fbf}{2} + \textcolor{#1f5fbf}{3} + \textcolor{#1f5fbf}{5} + \textcolor{#1f5fbf}{7} + \textcolor{#1f5fbf}{8} = 25 \]
Add the waits
Why: μ needs a total
\[ \mu = 25 \div 5 = 5 \]
Divide by N = 5
Why: total over count
Figure (svg): Waits 2, 3, 5, 7, 8 on a number line with 0 deviation arrows from the mean
\[ \textcolor{#1f5fbf}{-3,\ -2,\ 0,\ +2,\ +3} \]
Subtract 5
Why: spread starts at μ
Figure (svg): Waits 2, 3, 5, 7, 8 on a number line with 5 deviation arrows from the mean
\[ -3 - 2 + 0 + 2 + 3 = 0 \]
Check: deviations total zero
Why: confirms the centre 5
Worked example
Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: each distance drawn as a blue side, no squares yet
New waits 2, 3, 5, 7, 8: μ = 5, distances −3, −2, 0, +2, +3.
\[ \textcolor{#b54708}{9,\ 4,\ 0,\ 4,\ 9} \]
Square each
Why: areas build σ²
Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9
\[ \textcolor{#b54708}{9 + 4 + 0 + 4 + 9 = 26} \]
Add the areas
Why: σ² needs this total
Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26
\[ 26 \div 5 = \textcolor{#b54708}{5.2} \]
Divide by N = 5
Why: σ² is the mean square
Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26; the average square, area 5.2
\[ \textcolor{#1f5fbf}{\sigma} = \sqrt{5.2} \approx \textcolor{#1f5fbf}{2.28} \]
Take the root
Why: back to minutes
Figure (svg): Distances of the waits 2, 3, 5, 7, 8 from μ = 5: −3, −2, 0, +2, +3: a square on each distance, areas 9, 4, 0, 4, 9; label: total area 26; the average square, area 5.2, side 2.28
\[ \textcolor{#1f5fbf}{2.28}^2 \times 5 \approx 25.99 \]
Check: square, times N
Why: near 26: rounding
Worked example
Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4
\[ \textcolor{#1f5fbf}{4} > \textcolor{#1f5fbf}{2.28} > \textcolor{#1f5fbf}{2.04} > \textcolor{#1f5fbf}{1.83} \]
Rank the four σ
Why: answers which change spreads most
Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4; sorted largest first and numbered
\[ \textcolor{#1f5fbf}{2.04} > \textcolor{#1f5fbf}{2} > \textcolor{#1f5fbf}{1.83} \]
Place line A's σ = 2 in the ranking
Why: shows the one change that lowers σ
Figure (svg): Bars for the four options' σ: sixth wait of 5 1.83, 4 and 6 moved out 2.28, 5 changed to 6 2.04, every wait doubled 4; sorted largest first and numbered; grey dashed line at line A's σ = 2
\[ \textcolor{#b54708}{16} > \textcolor{#b54708}{5.2} > \textcolor{#b54708}{4.16} > \textcolor{#b54708}{3.333} \]
Check: rank the σ² values
Why: roots keep the order
Worked example
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5
Treat each line's five waits as a sample.
\[ \textcolor{#b54708}{\text{A: } 20,\ \ \text{B: } 80} \]
Take both squared totals
Why: s² starts from these
\[ 5 - 1 = 4 \]
Find n − 1
Why: five waits, a sample
\[ 20 \div 4 = \textcolor{#b54708}{5},\ \ 80 \div 4 = \textcolor{#b54708}{20} \]
Divide each by 4
Why: a sample's fair average
\[ \textcolor{#1f5fbf}{s_A} = \sqrt{5} \approx \textcolor{#1f5fbf}{2.24},\ \ \textcolor{#1f5fbf}{s_B} = \sqrt{20} \approx \textcolor{#1f5fbf}{4.47} \]
Take both roots
Why: back from area to minutes
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B
\[ 2.24^2 \times 4 \approx 20.07,\ \ 4.47^2 \times 4 \approx 79.92 \]
Check: square, times 4
Why: near both: rounding
Worked example
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B
\[ 7 - 5 = 2 \]
Find 7's distance
Why: what hops cover
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes
\[ k_A = 2 \div \textcolor{#1f5fbf}{2.24} \approx 0.89 \]
Divide by A's s
Why: SDs give one ruler
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick
\[ k_B = 2 \div \textcolor{#1f5fbf}{4.47} \approx 0.45 \]
Divide by B's s
Why: so both lines compare fairly
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick; an arc from 5 to 7 on B: 0.45 SD, under half of B's first tick
\[ 0.45 < 0.89 \]
Compare the hop counts
Why: B's ordinary spread reaches 7
\[ \text{A } 1/5,\ \text{B } 2/5 \]
Count waits over 7
Why: data test the hops
Figure (svg): Dot plots of line A (2, 4, 5, 6, 8) and line B (0, 2, 4, 8, 11) with both means at 5; SD ticks every s ≈ 2.24 minutes on A and every s ≈ 4.47 on B; a line at 7 minutes; an arc from 5 to 7 on A: 0.89 SD, just short of A's first tick; an arc from 5 to 7 on B: 0.45 SD, under half of B's first tick; waits over 7 circled: one on A, two on B
\[ 5 + 0.89(\textcolor{#1f5fbf}{2.24}) \approx 7.0,\ \ 5 + 0.45(\textcolor{#1f5fbf}{4.47}) \approx 7.0 \]
Check: hop back on each
Why: both land near 7
Join line A.
Recap
OpenStax Introductory Statistics 2e, §2.7 Measures of the Spread of the Data §2.7, pp. 107-118 — Examples 2.32–2.35 and the Try Its trace back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.