Mean, median and mode as measures of central tendency, range and standard deviation as measures of dispersion, the effect of an outlier on each of the five statistics, and choosing which measure best describes a data set.
Subject: Algebra 2 · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Algebra 2 · Chapter 11 — Data Analysis and Statistics
Find Measures of Central Tendency and Dispersion
Objectives
Five outcomes. Where the data sits, and how far it spreads.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-750 — the lesson these objectives are drawn from
Warm-up
You can find an average, and you have read graphs of data.
Discussion prompt
Five test scores are 60, 62, 64, 66 and 98. What is their average? Does that number describe the group well?
Hint: Compare the average with the individual scores.
Answer:
\[ \frac{60+62+64+66+98}{5} = \frac{350}{5} = 70 \]
But four of the five scores are below 70. The single high score dragged the average above almost everyone.
The middle score is 64, which describes the group far better. This lesson gives names to both numbers, adds two measures of spread, and explains when each is the honest one to report.
Concept
A measure of central tendency says where a data set sits: the mean, the median or the mode. A measure of dispersion says how spread out it is: the range or the standard deviation. A summary needs one of each.
standard deviation — A measure of dispersion giving the typical distance between a data value and the mean, computed as the square root of the average of the squared deviations.
\[ s = \sqrt{\frac{\sum (x-\bar{x})^2}{n}} \]
Two data sets can share a mean and look completely different, which is why a centre reported alone is an incomplete description.
Figure (svg): Two columns comparing measures of centre with measures of dispersion
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-745
Section
Section 1
Concept
The mean is the sum divided by how many values there are. The median is the middle value once they are in order, or the average of the two middle ones. The mode is the value occurring most often.
\[ \bar{x} = \frac{x_1+x_2+\dots+x_n}{n} \]
There may be one mode, several, or none at all. The mean and median always exist and are unique.
Figure (svg): Two data sets of waiting times with their mean, median and mode computed
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-744 — Measures of Central Tendency
Picture it
Example 1: waiting times at two veterinary offices.
Figure (svg): Two data sets of waiting times with their mean, median and mode computed
Office A has a mean of 22, a median of 20 and a mode of 24. Office B comes in at 16, 18 and 18 — lower on every measure.
Worked example
Example 1.
\[ \text{Find the mean, median and mode of } 14,17,18,19,20,24,24,30,32 \text{ and of } 8,11,12,16,18,18,18,20,23. \]
Office A: add and divide
Why: The nine times total 198.
\[ \text{mean } 22 \]
Office A: find the middle
Why: The fifth of nine ordered values.
\[ \text{median } 20 \]
Office A: find the most common
Why: Twenty-four appears twice; nothing else repeats.
\[ \text{mode } 24 \]
Office B: repeat
Why: Sum 144, fifth value 18, and 18 appears three times.
\[ 16, 18\text{ and } 18 \]
Figure (svg): Two data sets of waiting times with their mean, median and mode computed
\[ A: 22, 20, 24; \qquad B: 16, 18, 18 \]
Verify: check the mean against the data
Why: Office A's mean of 22 sits between its smallest value of 14 and largest of 32, as any mean must. And it is above the median of 20, which happens when the larger values stretch further from the middle than the smaller ones — here 30 and 32 pull it up.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-744
Fill the middle
Guided Practice 1.
Fill in the blanks
2,3,4,6,7,8,8,9,12,15 \;\Longrightarrow\; \text7.5 = \frac______ = ___
Why: With ten values there is no single middle, so the fifth and sixth are averaged, giving 7.5. An even count always produces this extra step.
Worked example
Guided Practice 1.
\[ \text{Find the mean, median and mode of } 4, 8, 12, 15, 3, 2, 6, 9, 8, 7. \]
Add and divide
Why: The ten values total 74.
\[ \text{mean } 7.4 \]
Put them in order
Why: Two, 3, 4, 6, 7, 8, 8, 9, 12, 15.
Find the median
Why: Ten values, so average the fifth and sixth.
\[ \frac{7 + 8}{2} = 7.5 \]
Find the mode
Why: Eight appears twice; nothing else does.
\[ \text{mode } 8 \]
Figure (svg): The solution to Worked example bus waiting times shown as a ladder of expressions, one row per algebraic move
\[ \bar{x} = 7.4, \; \text{median } 7.5, \; \text{mode } 8 \]
Verify: check that ordering was necessary
Why: In the original order the fifth and sixth values are 3 and 2, which would give a median of 2.5 — badly wrong. The book prints a caution about exactly this: the numbers must be sorted before the middle one means anything.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745
Trap
\[ 12, 8, 9, 5, 10, 10, 3 \]
Take the fourth value as the median
Why: The list has seven entries, so the fourth is taken as the middle.
\[ \text{median} = 5 \quad \text{(wrong)} \]
Five is actually the smallest value in the set. The middle POSITION only gives the middle VALUE once the list is ordered.
\[ 3, 5, 8, 9, 10, 10, 12 \;\Longrightarrow\; \text{median} = 9 \]
Sort first, then take the middle position
Why: The median is defined by rank, and rank requires an ordering.
\[ \text{three values below } 9, \text{ three above} \]
This is lesson exercise 9's printed error. Sorting takes seconds and is the only preparation the median needs.
Sorting
Sum, position, or frequency?
Sort into buckets
Sort each description.
The mean is the only one of the three that uses every value, which is both its strength and the reason an outlier can distort it.
Matching
Count the repeats.
Match the pairs
Why: The third has two modes, which lesson exercise 10 flags as a printed error when only one is reported. A data set can have one mode, several, or none, unlike the mean and median which are always unique.
Prediction
Commit before reasoning.
Predict first
Office A has mean 22 and median 20. What does the gap tell you?
Correct: The larger values stretch further from the middle than the smaller ones, pulling the mean up.
\[ \bar{x} > \text{median} \;\Longrightarrow\; \text{a longer tail on the high side} \]
Why: The median only cares about position, so it sits at the fifth value whatever the outer numbers are. The mean adds every value, so the 30 and 32 at the top pull it above the median while the 14 at the bottom pulls less. A mean above the median signals a longer tail on the high side, and a mean below it signals the reverse.
Section
Section 2
Concept
The range is the greatest value minus the least. It is quick to compute and gives a rough sense of how spread out a data set is, but it uses only two of the values.
\[ \text{range} = \text{max}-\text{min} \]
Because it depends entirely on the two extremes, the range is completely determined by whatever is farthest out — including any outlier.
Figure (svg): The same two data sets compared by the distance from their smallest to largest value
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745 — Find ranges of data sets
Picture it
Example 2: the ranges of the two offices' waiting times.
Figure (svg): The same two data sets compared by the distance from their smallest to largest value
Eighteen minutes for A and 15 for B, so A's times are more spread out. But the arrows say nothing about how the values inside them are arranged.
Worked example
Example 2.
\[ \text{Find the range of } 14,\dots,32 \text{ and of } 8,\dots,23. \]
Office A: identify the extremes
Why: Fourteen and 32.
Office A: subtract
Why: Thirty-two minus 14.
\[ 18 \]
Office B: identify the extremes
Why: Eight and 23.
Office B: subtract and compare
Why: Twenty-three minus 8 is 15, less than 18.
Figure (svg): The same two data sets compared by the distance from their smallest to largest value
\[ 18 > 15 \]
Verify: notice what the range ignores
Why: Office B's values include three identical 18s clustered in the middle, and the range says nothing about that. Two data sets with the same two extremes have the same range even if one is tightly clustered and the other evenly spread.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745
Fill the middle
Example 2.
Fill in the blanks
32-14 = 18
Why: Thirty-two minus 14 is 18 minutes. The range is always non-negative, since the largest value is never below the smallest.
Worked example
Guided Practice 2 and lesson exercise 11.
\[ \text{Find the range of } 4,8,12,15,3,2,6,9,8,7 \text{ and of } 7,4,6,8,5,9,5,7. \]
First: find the extremes
Why: The smallest is 2 and the largest 15.
\[ 2\text{ and } 15 \]
First: subtract
Why: Fifteen minus 2.
\[ 13 \]
Second: find the extremes
Why: The smallest is 4 and the largest 9.
\[ 4\text{ and } 9 \]
Second: subtract
Why: Nine minus 4.
\[ 5 \]
Figure (svg): The solution to Worked example one more range shown as a ladder of expressions, one row per algebraic move
\[ 13; \qquad 5 \]
Verify: compare the two sets
Why: The first set spans 13 units and the second only 5, so the second is far more tightly clustered — and its eight values all sit between 4 and 9. Sorting first is not required for the range, but it makes the extremes easy to spot without scanning twice.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 746-747
Error analysis
A student finds the range of 4, 8, 12, 15, 3, 2, 6, 9, 8, 7.
Annotate
On: \( \text{range} = 7-4 = 3 \)
Position in a list means nothing for the range. Scanning for the two extremes, or sorting first, is the only reliable approach.
Two truths and a lie
All three are about the range.
Eliminate the wrong options
Two of these are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The sets 1, 2, 3, 9, 10 and 1, 5, 5, 5, 10 both have range 9, but the first is split into two clumps and the second piled in the middle. The range cannot distinguish them, which is precisely why the standard deviation exists.
Sorting
Where the data is, or how far it reaches.
Sort into buckets
Sort each statistic.
A complete summary needs one from each column, since a centre alone leaves the shape of the data entirely open.
Prediction
Commit before reasoning.
Predict first
Two data sets both have range 9. What does that guarantee about them?
Correct: Only that their extremes are 9 apart; nothing about the values between.
\[ \text{range uses } 2 \text{ values}; \quad s \text{ uses all } n \]
Why: The set 1, 2, 3, 9, 10 and the set 1, 5, 5, 5, 10 share a range of 9 but have very different shapes and standard deviations, about 3.8 against 3.2. The range is built from two numbers and discards the rest, which makes it fast to compute and weak as a description. That gap is what the next idea fills.
Section
Section 3
Concept
Standard deviation measures how far a typical value sits from the mean. Every deviation is squared so that positives and negatives do not cancel, the squares are averaged, and a square root restores the original units.
\[ s = \sqrt{\frac{(x_1-\bar{x})^2+\dots+(x_n-\bar{x})^2}{n}} \]
Unlike the range, it uses every value, so two sets with the same extremes but different shapes get different standard deviations.
Figure (svg): The standard deviation computed from squared distances to the mean
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745 — Standard Deviation of a Data Set
Picture it
Example 3: the standard deviations of both offices.
Figure (svg): The standard deviation computed from squared distances to the mean
Office A's squared deviations total 290 and B's 182, giving about 5.7 and 4.5 — the same ordering the range gave, but from all nine values rather than two.
Worked example
Example 3, a multiple-choice item.
\[ \text{Find the standard deviation of each office's waiting times.} \]
Office A: subtract the mean from each value
Why: Fourteen minus 22 is negative 8, and so on.
Office A: square and add
Why: Sixty-four plus 25 plus 16 and the rest.
\[ 290 \]
Office A: divide and take the root
Why: Two hundred ninety over 9, then the square root.
\[ \text{about } 5.7 \]
Office B: the same four steps
Why: The squares total 182.
\[ \text{about } 4.5 \]
Figure (svg): The standard deviation computed from squared distances to the mean
\[ s_A \approx 5.7; \qquad s_B \approx 4.5 \]
Verify: check the deviations sum to zero
Why: Office A's deviations are negative 8, negative 5, negative 4, negative 3, negative 2, 2, 2, 8 and 10, which total exactly 0. That always happens, since the mean is the balance point — and it is a useful check that the mean and the subtractions are right before any squaring begins.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745
Fill the middle
Example 3.
Fill in the blanks
64+25+16+9+4+4+4+64+100 = 290
Why: The nine squared deviations total 290, which divided by 9 and square-rooted gives about 5.7. Every squared deviation is non-negative, so the total can only grow as values move away from the mean.
Worked example
Guided Practice 2.
\[ \text{Find the range and standard deviation of } 4,8,12,15,3,2,6,9,8,7. \]
Find the mean
Why: The ten values total 74.
\[ 7.4 \]
Compute the squared deviations
Why: From negative 5.4 up to 7.6, each squared.
\[ \text{total } 144.4 \]
Divide by n and take the root
Why: One hundred forty-four point four over 10.
\[ \sqrt{14.44} = 3.8 \]
Find the range
Why: Fifteen minus 2.
\[ 13 \]
Figure (svg): The solution to Worked example the bus data shown as a ladder of expressions, one row per algebraic move
\[ \text{range } 13; \quad s = 3.8 \]
Verify: sanity-check the size
Why: The standard deviation of 3.8 is well under the range of 13, as it always must be — a typical distance from the centre cannot exceed the whole span. As a rough guide the standard deviation is often around a quarter of the range, and 3.8 against 13 fits that.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 746-746
Trap
\[ \text{deviations: } -8,-5,-4,-3,-2,2,2,8,10 \]
Average them directly
Why: The deviations are added and divided by 9.
\[ \frac{0}{9} = 0 \quad \text{(useless)} \]
The deviations always total zero, whatever the data, so their average carries no information at all.
\[ \text{square first: } 64,25,16,9,4,4,4,64,100 \]
Square each deviation before averaging
Why: Squaring removes the signs so nothing cancels.
\[ s = \sqrt{\tfrac{290}{9}} \approx 5.7 \]
The square root at the end converts back to the original units, so a standard deviation of 5.7 is 5.7 minutes rather than 5.7 squared minutes.
Ranking
Computing a standard deviation.
Put in order
Why: Step one has to come first because every later step refers to the mean. Step three is the one people skip, and skipping it guarantees an answer of zero since the deviations always cancel.
Comparison
Fill the blanks. Both measure spread.
Comparison matrix
| Question | Range | Standard deviation |
|---|---|---|
| How many values does it use? | two | all n of them |
| Office A | 18 | about 5.7 |
| Office B | 15 | about 4.5 |
| Effort to compute | one subtraction | several steps for every value |
Both put office A above office B here, but they need not always agree — the range can be driven entirely by a single extreme value while the standard deviation weighs everything.
Prediction
Commit before reasoning.
Predict first
What goes wrong if the deviations are averaged without squaring?
Correct: They always total zero, so the average is zero for every data set.
\[ \sum (x-\bar{x}) = 0 \text{ for every data set} \]
Why: The mean is the balance point, so the amounts above it exactly cancel the amounts below. Squaring makes every deviation positive so nothing cancels, and the final square root undoes the squaring to bring the answer back into the original units. Taking absolute values would also work and gives a different, less common statistic — squaring is preferred partly because it is smoother to work with algebraically.
Section
Section 4
Concept
An outlier is a value far from the rest of the data. It moves the mean noticeably, the median slightly, and the mode not at all, while inflating both the range and the standard deviation.
\[ \bar{x}: 14 \to 13; \quad s: 1.7 \to 3.5 \]
The mean and the two spread measures are sensitive because they use every value's size; the median and mode use only position and frequency.
Figure (svg): A data set before and after one extreme value is added, with all five statistics
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 746-746 — Examine the effect of an outlier
Picture it
Example 4: ten winning scores, then an eleventh of 3.
Figure (svg): A data set before and after one extreme value is added, with all five statistics
The mode did not move at all, the median fell by half a point, and both spread measures more than doubled.
Worked example
Example 4, parts a and b.
\[ \text{For } 14,15,15,17,11,15,13,12,15,13, \text{ find all five statistics; then add a score of } 3. \]
Before: centre
Why: Sum 140 over 10; middle pair 14 and 15; 15 appears four times.
\[ 14, 14.5, 15 \]
Before: spread
Why: Seventeen minus 11; squares total 28 over 10.
\[ 6\text{ and about } 1.7 \]
After: centre
Why: Sum 143 over 11; sixth of eleven values; 15 still most common.
\[ 13, 14, 15 \]
After: spread
Why: Seventeen minus 3; squares total 138 over 11.
\[ 14\text{ and about } 3.5 \]
Figure (svg): A data set before and after one extreme value is added, with all five statistics
\[ \bar{x}: 14\to 13; \; s: 1.7\to 3.5 \]
Verify: check where the extra 138 came from
Why: The outlier alone contributes 3 minus 13, squared, which is 100 of the 138 total — more than the other ten values combined. One value dominating the sum of squares is exactly what makes the standard deviation so sensitive to outliers.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 746-746
Sorting
Adding a score of 3 to ten scores near 14.
Sort into buckets
Sort each statistic by how much it changed.
The two spread measures are by far the most sensitive, which is worth knowing: an unexplained jump in a standard deviation usually means an outlier rather than a genuine change in the data.
Worked example
Guided Practice 3.
\[ \text{Repeat with a final score of } 25 \text{ instead of } 3. \]
Find the new mean
Why: One hundred sixty-five over 11.
\[ 15 \]
Find the new median and mode
Why: The sixth of eleven ordered values; 15 still four times.
\[ 15\text{ and } 15 \]
Find the new range
Why: Twenty-five minus 11.
\[ 14 \]
Find the new standard deviation
Why: The squares again total 138 over 11.
\[ \text{about } 3.5 \]
Figure (svg): The solution to Worked example an outlier at the top instead shown as a ladder of expressions, one row per algebraic move
\[ 15, 15, 15, 14, 3.5 \]
Verify: compare with the low outlier
Why: A score of 3 pulled the mean down to 13 and a score of 25 pushed it up to 15, each by one point — because both are ten away from the original mean of 14 in opposite directions. The spread measures came out identical for the same reason: distance from the centre is what they measure, not direction.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 746-746
Error analysis
A student summarises the eleven scores including the outlier of 3.
Annotate
On: \( \text{the typical winning score is } 13 \)
When a data set contains an outlier, the median is usually the more honest summary — and reporting the outlier alongside it is more honest still.
Fill the middle
Example 4b.
Fill in the blanks
(3-13)^2 = 100 \text___ 138 \text___
Why: The outlier alone contributes 100 of the 138, more than the other ten values combined. Squaring means a value ten units away counts a hundred times as much as one a single unit away.
Matching
Which values does each one use?
Match the pairs
Why: Sensitivity tracks exactly how much of each value the statistic uses. The mode uses only how often values repeat, the median only their order, and the two mean-based statistics use the actual numbers — squared, in the last case.
Prediction
Commit before reasoning.
Predict first
A data set contains one clear outlier. What is the right response?
Correct: Investigate it: a recording error may be removed, but a genuine extreme value should stay and be reported.
\[ \text{report both, and say which is which} \]
Why: An outlier caused by a typo or a broken instrument is bad data and should go. An outlier that really happened — a genuinely terrible game, an unusually tall player — is information, and deleting it hides something real. The honest approach is to report both summaries, with and without, and say why. Deleting inconvenient data without stating so is how misleading statistics get made.
Section
Section 5
Concept
The mean is best for data with no extreme values, the median when outliers or a long tail are present, and the mode for the most common category. Whichever is chosen, a measure of spread belongs beside it.
\[ \text{centre} + \text{spread} \]
Reporting a centre without a spread hides the shape of the data, and two very different data sets can share every measure of centre.
Figure (svg): A guide to which measure of centre to report in different situations
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-746 — Measures of central tendency and dispersion
Picture it
What each measure uses, and when to prefer it.
Figure (svg): A guide to which measure of centre to report in different situations
The mean uses everything and pays for it when an outlier appears; the median ignores sizes and is steadier; the mode answers a different question entirely.
Worked example
Applying the guide.
\[ \text{Which measure of centre suits: house prices in a city; the scores } 14,15,15,17,11,15,13,12,15,13; \text{ favourite shirt colours?} \]
House prices
Why: A few very expensive houses stretch the top of the range.
The game scores
Why: Tightly clustered with no extremes.
Shirt colours
Why: The values are categories, not numbers.
State the reason each time
Why: Outliers, symmetry, or non-numerical data.
Figure (svg): A guide to which measure of centre to report in different situations
\[ \text{median}, \; \text{mean}, \; \text{mode} \]
Verify: check the third case carefully
Why: Colours cannot be added or ordered meaningfully, so neither a mean nor a median exists — the mode is the only one of the three that applies at all. Whenever the data is categorical rather than numerical, the choice is made for you.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-746
Sorting
Look for outliers and for whether the data is numerical.
Sort into buckets
Sort each situation.
House prices and incomes are the standard examples of skewed data, and both are almost always reported as medians for exactly this reason.
Worked example
Why a spread measure is needed.
\[ \text{Compare } 9,10,10,10,11 \text{ with } 2,10,10,10,18. \text{ Find every statistic for each.} \]
Both means
Why: Fifty over 5 in each case.
\[ 10\text{ and } 10 \]
Both medians and modes
Why: The middle value and most frequent value are 10 in both.
\[ 10\text{ and } 10 \]
The ranges
Why: Two versus 16.
The standard deviations
Why: About 0.63 against about 5.06.
Figure (svg): The solution to Worked example same centre, different data shown as a ladder of expressions, one row per algebraic move
\[ \bar{x} = 10 \text{ both}; \quad s = 0.63 \text{ against } 5.06 \]
Verify: see why the centres cannot distinguish them
Why: Both sets are symmetric about 10, so every measure of centre lands there. Only the spread measures notice that one set clusters within a single unit while the other reaches eight units either side. Reporting the mean alone would describe both sets identically.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 745-745
Trap
\[ \text{both data sets have mean } 10 \]
Conclude that they are similar
Why: A single summary number is taken as a full description.
\[ \text{they are alike} \quad \text{(wrong)} \]
One set spans 2 units and the other 16. Their means are identical and their shapes are nothing alike.
\[ \bar{x} = 10, \; s = 0.63; \qquad \bar{x} = 10, \; s = 5.06 \]
Report a centre AND a spread
Why: Two numbers are the minimum honest summary.
\[ \text{same centre, eight times the spread} \]
This is why weather reports give averages and ranges, and why test results give a mean and a standard deviation rather than a mean alone.
Comparison
Fill the blanks. The centres cannot tell them apart.
Comparison matrix
| Statistic | 9, 10, 10, 10, 11 | 2, 10, 10, 10, 18 |
|---|---|---|
| Mean | 10 | 10 |
| Median | 10 | 10 |
| Range | 2 | 16 |
| Standard deviation | about 0.63 | about 5.06 |
The top two rows are identical and the bottom two differ by a factor of eight, which is the whole argument for always reporting a spread.
Prediction
Commit before reasoning.
Predict first
News reports usually give median income rather than mean income. Why?
Correct: Because a small number of very high incomes pulls the mean far above what most people earn.
\[ \text{long right tail} \;\Longrightarrow\; \bar{x} \text{ far above the median} \]
Why: Income distributions have a long right tail: most people cluster fairly low while a few earn enormously more. The mean is dragged toward those few, so it sits above the majority and describes almost nobody. The median stays at the person in the middle whatever happens at the top, which is exactly the robustness the earlier idea measured. The same argument applies to house prices and to any quantity with a long tail.
Fill the middle
A complete summary.
Fill in the blanks
\text1.7 \bar___ = 14, \; s \approx ___
Why: A mean of 14 with a standard deviation of 1.7 says the scores cluster tightly near 14. The mean alone would leave that entirely open.
Comparison
Fill the blanks. Three centres and two spreads.
Comparison matrix
| Statistic | How it is found | Sensitive to outliers? |
|---|---|---|
| Mean | the sum divided by n | yes, moderately |
| Median | the middle value once sorted | barely |
| Mode | the most frequent value | usually not at all |
| Range and standard deviation | max minus min; the root-mean-square deviation | yes, strongly |
Sensitivity tracks how much of each value the statistic uses, and the two spread measures use distance from the centre, which is exactly what an outlier maximises.
Pattern
Sort first, then compute.
The deviations from the mean always total zero, which is a free check before any squaring.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-750
Check
Sort before finding the middle.
Check your understanding
What is the median of 12, 8, 9, 5, 10, 10, 3?
Answer: A
Why: Sorted, the list is 3, 5, 8, 9, 10, 10, 12, and 9 is the fourth of seven.
Check
Square before averaging.
Check your understanding
For 14, 17, 18, 19, 20, 24, 24, 30, 32 with mean 22, what is the standard deviation?
Answer: A
Why: The squared deviations total 290, and the square root of 290 over 9 is about 5.7.
Check
Which statistic is untouched?
Check your understanding
Adding an outlier of 3 to ten scores near 14, which measure is least affected?
Answer: A
Why: The most frequent value stays 15, since one new value changes no frequency.
Real world
A small company has ten employees earning 30, 32, 33, 34, 35, 36, 37, 38, 40 and 485 thousand a year, the last being the owner.
Discussion prompt
Find the mean and median salary, and decide which the company should quote in a recruitment advertisement.
Hint: Compute both, then compare them with the actual salaries.
Answer:
\[ \bar{x} = \frac{800}{10} = 80 \text{ thousand} \]
\[ \text{median} = \frac{35+36}{2} = 35.5 \text{ thousand} \]
The mean is 80 thousand and the median 35.5 thousand. Nine of the ten employees earn less than half the mean, so quoting it would be technically true and thoroughly misleading.
This is exactly the outlier effect of the fourth idea, at full strength: one salary contributes more than half the total. Advertising the mean would be the kind of statistic that is accurate and dishonest at once — which is why employment data, house prices and income figures are reported as medians almost without exception. The standard deviation here is about 135 thousand, larger than nine of the ten salaries, and that absurdity is itself a warning sign that the mean is the wrong summary.
Commit first
Answer, then rate your confidence honestly.
Predict first
Is the median of 12, 8, 9, 5, 10, 10, 3 equal to 5, the middle entry of the list?
Correct: No — the list must be sorted first, giving a median of 9.
\[ 3,5,8,9,10,10,12 \;\Longrightarrow\; \text{median} = 9 \]
Why: The median is the middle value by RANK, not by position in however the data happened to be written down. Sorting gives 3, 5, 8, 9, 10, 10, 12, whose fourth entry is 9 — and 5 is in fact the second smallest value in the set, nowhere near the middle. This is lesson exercise 9's printed error, and the book prints a caution about it beside Example 1. Sorting is the only preparation the median needs, and it takes seconds.
Explain it
They know how to find an average and think that settles it.
Discussion prompt
In four sentences or fewer, explain why an average alone can mislead.
Hint: Think about one very large value.
Answer:
An average adds every value, so one unusually large number pulls it up a long way. If nine people earn about 35 thousand and one earns 485 thousand, the average is 80 thousand — more than double what almost everyone actually earns.
The middle value, 35.5 thousand, describes the group much better. So when the data has an extreme value, report the middle one instead, and say how spread out the values are.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the median, sort as the very first thing you do with any data set. For the standard deviation, write the five steps down the margin and tick them off. For outliers, ask which values each statistic actually uses. For choosing, look for extremes first and for whether the data is even numerical.
Connect it up
Paper. Fifteen minutes.
Draw it
Build a statistics page. Top left: write the definitions of mean, median and mode, and compute all three for both veterinary offices, showing the sorting step. Top right: draw both data sets as dot plots on a shared number line and mark each range with an arrow. Middle: compute one standard deviation in full, laying out the deviations, their squares and the total in a column, and check that the deviations sum to zero. Bottom left: make a five-row table of the game scores before and after the outlier, and write beside each row how much it moved and why. Bottom right: write your own two data sets that share a mean but differ in spread, compute both standard deviations, and write one sentence on why a centre alone is not a summary.
If your deviations do not total zero, recheck the mean before squaring anything — every later step depends on it.
Recap
Five things, and two numbers make a summary.
| If you see | Then |
|---|---|
| An unsorted list | Sort it before finding the median |
| An even number of values | Average the two middle ones |
| No repeated value | There is no mode |
| Deviations summing to zero | That is expected; square them next |
| An extreme value | Prefer the median, and say the outlier is there |
| A reported mean with no spread | An incomplete summary |
Lesson 11.2 asks what happens to all five statistics when every value in a data set is shifted or scaled by the same amount.
McDougal Littell Algebra 2 (Texas Edition), Ch. 11 Data Analysis and Statistics — Lesson 11.1 Find Measures of Central Tendency and Dispersion §11.1, pp. 744-750 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Algebra 2 — $55/session, free consultation.