Five numbers drawn as a picture. A box plot is built from the minimum, the first quartile, the median, the third quartile and the maximum, with a box spanning the middle fifty percent and whiskers reaching the extremes; because each of the four parts holds about a quarter of the data, the width of a part reports how tightly that quarter is packed rather than how many values it contains. Then the modified box plot, whose whiskers stop at the most extreme non-outliers and whose outliers are marked as separate points, and parallel box plots on one shared number line, which are the cheapest honest comparison of several distributions. Closes with what a box plot cannot show.
Subject: Statistics · 65 slides · symbolic lesson
Open the interactive version of this deck
Title
Statistics · Chapter 2 — Descriptive Statistics
Box Plots
Objectives
Five outcomes. The third is the reading almost everyone gets backwards on first meeting.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-97 — the section these objectives are drawn from
Warm-up
Section 2.3 produced five numbers for the day class: minimum 32, Q1 56, median 74.5, Q3 82.5, maximum 99.
Discussion prompt
Those five numbers already contain a lot. Sketch what they tell you about the shape of the class's scores before reading on — where is the middle, and is the data balanced around it?
Hint: Compare the distance from Q1 up to the median with the distance from the median up to Q3.
Answer:
The middle half runs from 56 to 82.5, and the median sits at 74.5 — much closer to Q3 than to Q1. The lower quarter of the box spans 18.5 points and the upper quarter only 8.
So the scores are packed tightly just below the median and spread out below that. The bottom quarter, from 32 to 56, covers 24 points on its own — the widest of the four parts by a good margin.
All of that came from five numbers and no drawing. This section draws them, and the drawing makes the comparison instant rather than something to be worked out: the box plot is exactly these five values rendered so the widths can be seen.
Concept
Box plots give a good graphical image of the concentration of the data, and they show how far the extreme values are from most of it. A box plot is constructed from five values: the minimum, the first quartile, the median, the third quartile and the maximum.
box plot — A display constructed from the five-number summary. The first and third quartiles mark the ends of a rectangular box containing approximately the middle fifty percent of the data; the median is drawn inside it; and whiskers extend from the ends of the box to the smallest and largest values. Also called a box-and-whisker plot.
\[ \text{min}, \; Q_1, \; \text{median}, \; Q_3, \; \text{max} \]
The word to hold on to is CONCENTRATION. Each of the four parts — lower whisker, lower half of the box, upper half of the box, upper whisker — contains about a quarter of the observations, because that is what the quartiles did. So the parts differ in width and not in count, and a narrow part is one where a quarter of the data is squeezed into a short interval. Reading width as quantity is the single most common error with this display, and section 3 of this lesson is about it.
Figure (svg): A labelled box plot of fourteen values showing the minimum, first quartile, median, third quartile and maximum, drawn above a scaled number line
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 94
Section
Section 1
Concept
To construct a box plot, use a horizontal or vertical number line and a rectangular box. The first quartile marks one end of the box and the third quartile marks the other, so approximately the middle fifty percent of the data falls inside it. The whiskers extend from the ends of the box to the smallest and largest data values.
five-number summary — The minimum, first quartile, median, third quartile and maximum of a data set. These five values are all a box plot displays, and they are enough to show the centre, the spread and the asymmetry of the data.
\[ \text{whisker} \;\mid\; \text{box: } Q_1 \text{ to } Q_3 \;\mid\; \text{whisker} \]
The median may sit anywhere between the quartiles, and the book notes it can even coincide with one of them or with both. A median resting against Q1 means that a quarter of the data is packed against the bottom of the box; a median equal to both quartiles means at least half the observations share a single value. Neither is an error, and each says something specific about the data.
Figure (svg): A labelled box plot of fourteen values showing the minimum, first quartile, median, third quartile and maximum, drawn above a scaled number line
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-95 — constructing a box plot from the five values
Picture it
The book's fourteen-value set, drawn.
Figure (svg): A labelled box plot of fourteen values showing the minimum, first quartile, median, third quartile and maximum, drawn above a scaled number line
Notice the note the book puts in italics: it is important to start a box plot with a scaled number line. Without one, the box and whiskers are just shapes — the widths that carry all the information about concentration would be arbitrary. A box plot drawn without a scale below it is not a weaker version of the display; it conveys nothing at all.
Worked example
The set carried through from section 2.3.
\[ 1; \;1; \;2; \;2; \;4; \;6; \;6.8; \;7.2; \;8; \;8.3; \;9; \;10; \;10; \;11.5 \]
Read the extremes
Why: The first and last of the ordered list.
\[ \min 1, \max 11.5 \]
Find the median
Why: Average the seventh and eighth values.
\[ \text{median } 7 \]
Find Q1
Why: The median of the lower seven.
\[ Q 1 = 2 \]
Find Q3
Why: The median of the upper seven.
\[ Q 3 = 9 \]
Figure (svg): The solution to Worked example the five numbers for fourteen values shown as a ladder of expressions, one row per legal move
\[ \text{min } 1, \; Q_1 = 2, \; \text{median } 7, \; Q_3 = 9, \; \text{max } 11.5 \]
Verify: confirm the five numbers are in non-decreasing order
Why: They run 1, 2, 7, 9, 11.5, each at least as large as the one before, which any correct five-number summary must. A summary where Q1 exceeded the median, or the median exceeded Q3, would signal that a quartile was computed from the wrong half or that the data were not sorted first. It is a check worth running because it costs nothing and catches the two commonest errors in section 2.3's procedure.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 95
Matching
Each of the five numbers is responsible for one feature.
Match the pairs
Why: Five numbers, five features, and nothing else in the picture carries information. That is worth knowing precisely because it tells you what questions the display cannot answer: anything about the sample size, the number of peaks, or where values sit inside a quarter.
Worked example
Example 2.23, whose five numbers the book obtains from a calculator.
\[ \text{40 heights, from } 59 \text{ to } 77 \text{ inches} \]
Read the extremes
Why: Smallest and largest of the forty.
\[ \min 59, \max 77 \]
Find the median
Why: Average the twentieth and twenty-first values, both 66.
\[ \text{median } 66 \]
Find Q1
Why: The median of the lower twenty.
\[ Q 1 = 64.5 \]
Find Q3
Why: The median of the upper twenty.
\[ Q 3 = 70 \]
Figure (svg): The solution to Worked example the heights of forty students shown as a ladder of expressions, one row per legal move
\[ \text{range} = 77 - 59 = 18, \qquad IQR = 70 - 64.5 = 5.5 \]
Verify: confirm the middle half really is much tighter than the whole
Why: The range is 18 inches and the interquartile range is 5.5, so the middle half of the class occupies less than a third of the total span. That is the normal pattern and it is what makes both numbers worth reporting: the range describes the extremes and the IQR describes the bulk. When the two are close, the data are spread evenly; when they differ this much, most of the class is packed near the centre with a few students well away from it.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 95-96
Trap
\[ \text{a box and two whiskers, no number line beneath} \]
Draw the shape and label the five values on it
Why: The numbers are written on, so the student assumes the picture is complete.
\[ \text{the widths now mean nothing} \quad \text{(the display is empty)} \]
Whether the box is a third of the picture or nine tenths of it is the entire message, and without a scale that proportion is whatever the pen happened to do.
\[ \text{a scaled number line, and the five values placed on it} \]
Draw the axis first, to scale, and put the five values at their positions
Why: Every feature of a box plot is a position or a width.
The book states this explicitly and it is not a matter of neatness. A box plot carries no numbers of its own — no counts, no frequencies — so all of its information is geometric. Two box plots drawn to different scales cannot be compared at all, which is also why parallel box plots must share one axis rather than being drawn separately and set side by side.
Fill the middle
The two ends of the rectangle.
Fill in the blanks
\text50 Q_1 \text___ Q_3, \text___ ___ \text___
Why: About fifty percent, which is why the box's width is the interquartile range. A quarter of the data lies to the left of the box in the lower whisker and a quarter to the right in the upper whisker, so the box holds the middle half.
Two truths and a lie
All three concern the construction.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. A box plot of fifteen values and one of fifteen thousand look exactly alike if their five-number summaries match. The sample size has to be reported separately, which matters because a tidy-looking box plot of eight observations deserves far less confidence than the same picture from eight hundred.
Prediction
Commit before reasoning.
Predict first
A box plot's median line sits exactly on the left edge of the box. What does that say?
Correct: The second quarter is packed into no width at all.
Why: If the median equals Q1 then the quarter of the data between them spans zero, which happens when many observations share that value — a survey where a quarter of respondents answered zero, for instance. It is perfectly legitimate, it is not an error, and it is the opposite of symmetric: the data are piled against one edge of the box.
Section
Section 2
Concept
Use a horizontal or vertical number line, scaled so that the smallest and largest data values label the endpoints of the axis. Mark the two quartiles and draw the box between them, draw the median inside, and extend the whiskers to the extremes.
scaled number line — An axis on which equal distances represent equal differences in the variable. Every feature of a box plot is a position or a width on this axis, so without it the display carries no information at all.
\[ \text{axis} \;\to\; Q_1, Q_3 \;\to\; \text{box} \;\to\; \text{median} \;\to\; \text{whiskers} \]
The order matters in practice. Drawing the box first and fitting an axis to it afterwards is how box plots end up with unequal spacing, which silently falsifies every width in the picture. Drawing the axis first also forces the decision about where it starts and ends, and the book's advice is that the smallest and largest values label the endpoints — which for a data set with an extreme outlier is exactly the problem the modified box plot solves.
Figure (svg): A box plot of forty student heights divided into its four quarters, each labelled with its spread in inches, the second quarter narrowest at 1.5 and the fourth widest at 7
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-95 — the construction, and the note about the scaled number line
Picture it
Example 2.23 with each quarter's spread written on it.
Figure (svg): A box plot of forty student heights divided into its four quarters, each labelled with its spread in inches, the second quarter narrowest at 1.5 and the fourth widest at 7
The four spreads are 5.5, 1.5, 4 and 7 inches, and they total the range of 18. Each quarter holds about ten students, so the second quarter squeezes ten students into an inch and a half while the fourth spreads ten students over seven inches. That contrast is the whole content of the picture, and it is invisible in the five numbers until they are drawn.
Worked example
Example 2.23b. Four subtractions, and they must total the range.
\[ \text{min } 59, \; Q_1 = 64.5, \; \text{median } 66, \; Q_3 = 70, \; \text{max } 77 \]
First quarter
Why: From the minimum up to Q1.
\[ 64.5 - 59 = 5.5 \]
Second quarter
Why: From Q1 up to the median.
\[ 66 - 64.5 = 1.5 \]
Third quarter
Why: From the median up to Q3.
\[ 70 - 66 = 4 \]
Fourth quarter
Why: From Q3 up to the maximum.
\[ 77 - 70 = 7 \]
Figure (svg): The solution to Worked example computing the four quarter spreads shown as a ladder of expressions, one row per legal move
\[ 5.5 + 1.5 + 4 + 7 = 18 = \text{the range} \]
Verify: confirm the four spreads total the range
Why: They sum to 18, which is 77 minus 59. They must, because the four intervals are consecutive and cover the whole span from minimum to maximum with no gaps or overlaps. That total is a free check on all four subtractions, and it catches the common slip of measuring a quarter from the wrong endpoint — say from the minimum to the median rather than to Q1.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 96
Ranking
The heights of forty students, with spreads 5.5, 1.5, 4 and 7 inches.
Put in order
Why: Narrowest to widest: 1.5, 4, 5.5, 7. Every one of these quarters holds about ten students, so the ordering is an ordering of how tightly packed each group is. The tallest ten students are spread over seven inches while the ten just below the median share an inch and a half.
Worked example
Example 2.23c and d. Two spreads that answer different questions.
\[ \text{Find the range and the interquartile range of the heights.} \]
Compute the range
Why: Maximum minus minimum.
\[ 77 - 59 = 18 \]
Compute the IQR
Why: Q3 minus Q1.
\[ 70 - 64.5 = 5.5 \]
Say what each covers
Why: All forty students against the middle twenty.
Compare them
Why: The middle half occupies under a third of the total span.
Figure (svg): The solution to Worked example the range and the IQR together shown as a ladder of expressions, one row per legal move
\[ \text{range} = 18, \qquad IQR = 5.5 \]
Verify: confirm which of the two an extreme student would change
Why: Replace the tallest student's height of 77 with 90 and the range jumps from 18 to 31, while the IQR does not move at all — Q1 and Q3 are computed from the middles of the two halves and cannot feel a change at the end. The range reports the extremes and is completely at their mercy; the IQR reports the bulk and ignores them. That is the same property that made the IQR the right basis for the outlier fences in section 2.3.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 96
Error analysis
The five numbers are 59, 64.5, 66, 70, 77. Four students draw the plot.
Annotate
On: \( \begin{aligned} &(1)\; \text{box drawn from the minimum to the maximum} \\ &(2)\; \text{median drawn halfway between } Q_1 \text{ and } Q_3 \\ &(3)\; \text{four equal-width sections, one per quarter} \\ &(4)\; \text{box from } 64.5 \text{ to } 70, \text{ median at } 66, \text{ whiskers to } 59 \text{ and } 77 \end{aligned} \)
Errors (2) and (3) are the same mistake in different clothes — both draw what the student expects rather than what the numbers say, and both destroy the display's only real content, which is where things sit relative to one another.
Faded example
The quarters span 5.5, 1.5, 4 and 7 inches.
Fill in the blanks
5.5 + 1.5 + 4 + 7 = 18 = \text59 - \text___ = 77 - ___
Why: The four quarter spreads always total the range, because the four intervals tile the whole span end to end. It is the cheapest check available on a five-number summary and it catches a quarter measured from the wrong endpoint.
Two truths and a lie
All three concern constructing the plot.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The median goes at its own value, wherever that falls between the quartiles, and its position is one of the most informative things in the picture — an off-centre median is the visual signature of skew, which section 2.6 takes up.
Estimation
A box plot has Q1 at 20, Q3 at 60, and the median line drawn much nearer the left edge.
Predict first
What does that tell you about the middle half of the data?
Correct: The quarter above the median is spread wider.
Why: Both quarters inside the box hold the same number of observations, so a median near the left edge means the quarter below it is squeezed into a short interval while the quarter above it occupies most of the box. It says nothing about counts — the third option is the width-as-quantity error — and nothing about outliers, which the whiskers and marked points would show instead.
Section
Section 3
Concept
Box plots give a good graphical image of the concentration of the data. Each of the four parts contains approximately twenty-five percent of the observations, so the parts differ in width rather than in count: a short part is one in which a quarter of the data is packed into a small interval.
concentration — How tightly the data are packed in a region. In a box plot every part holds about the same number of values, so a narrow part indicates high concentration and a wide part indicates that the same number of values is spread thinly.
\[ \text{each part: about } 25\% \text{ of the data, in whatever width it needs} \]
The book performs exactly this reading on Example 2.23 and states it plainly: the interval from 59 to 65 has more than twenty-five percent of the data, so it has more data in it than the interval 66 through 70, which has twenty-five percent. Reading a box plot correctly means constantly converting width into density, which is the reverse of the habit a histogram builds, where height IS quantity.
Figure (svg): A box plot of forty student heights divided into its four quarters, each labelled with its spread in inches, the second quarter narrowest at 1.5 and the fourth widest at 7
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 96 — reading concentration from the quarter spreads
Picture it
Four quarters, each with about ten students, spanning 5.5, 1.5, 4 and 7 inches.
Figure (svg): A box plot of forty student heights divided into its four quarters, each labelled with its spread in inches, the second quarter narrowest at 1.5 and the fourth widest at 7
This is the reading to practise until it is automatic. The narrow second quarter does not contain fewer students than the wide fourth — it contains the same ten, standing much closer together in height. A box plot answers 'where are they crowded?' and never 'where are there more of them?', because the answer to the second question is always the same: about a quarter everywhere.
Worked example
Example 2.23e. The book's own comparison, which requires the quartiles.
\[ \text{Compare the interval } 59\text{-}65 \text{ with the interval } 66\text{-}70. \]
Locate 66 to 70 on the summary
Why: From the median up to Q3.
State its share
Why: One of the four quarters.
\[ 25 \% \]
Locate 59 to 65 on the summary
Why: From the minimum past Q1 at 64.5 and a little beyond.
Compare
Why: More than 25 percent against exactly 25 percent.
\[ 59\text{ to } 65\text{ holds more} \]
Figure (svg): The solution to Worked example comparing two intervals shown as a ladder of expressions, one row per legal move
\[ 59\text{-}65 : \; > 25\% \qquad \text{against} \qquad 66\text{-}70 : \; = 25\% \]
Verify: confirm the comparison could not be made by width alone
Why: The two intervals are six and four inches wide, so width alone would suggest the first holds half again as much as the second. What actually decides it is where the quartiles fall: 59 to 65 covers the entire first quarter plus a slice of the second, while 66 to 70 covers exactly the third quarter. The quartile positions are the only thing in the display that converts width into share, which is why a box plot without them — a bare range — would answer nothing.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 96
Discrimination
For each feature, decide whether it reports crowding or something else.
Sort into buckets
Sort each statement about a box plot.
Worked example
Why the box matters more than the whiskers.
\[ \text{Set A: } Q_1 = 20, Q_3 = 80. \quad \text{Set B: } Q_1 = 47, Q_3 = 53. \quad \text{Both run } 0 \text{ to } 100. \]
Compare the ranges
Why: Both span 0 to 100.
Compare the IQRs
Why: Sixty against six.
Say what A looks like
Why: The middle half spread over 60 units.
Say what B looks like
Why: The middle half packed into 6 units.
Figure (svg): The solution to Worked example two data sets, same range, opposite concentration shown as a ladder of expressions, one row per legal move
\[ IQR_A = 60 \qquad \text{against} \qquad IQR_B = 6 \]
Verify: confirm which display feature revealed the difference
Why: The whiskers are what the two sets share, and the box is what separates them. Set B's long whiskers with a tiny box say that half the observations sit within three units of the centre while a quarter stretch down to zero and a quarter up to a hundred. Reporting only the range would call these two data sets identical, which is exactly why section 2.7 will treat the range as the weakest measure of spread available.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-96
Trap
\[ \text{the upper whisker is much longer than the lower whisker} \]
Conclude that more of the data lies above the median
Why: The eye reads area as quantity, as it correctly does on a histogram.
\[ \text{more students are tall} \quad \text{(wrong)} \]
Both whiskers hold a quarter of the data by construction. The long one holds the same count spread thinly.
\[ \text{each part holds about } 25\%; \; \text{width reports SPREAD, not count} \]
Translate every width into a statement about crowding
Why: Ask how tightly the same quarter is packed, never how many are there.
This is the one reading that a histogram actively trains you to get wrong, because there height genuinely is frequency. On a box plot the counts are fixed at a quarter each and only the widths vary, so the two displays use geometry in opposite ways. Saying 'a quarter of the students are spread across these seven inches' rather than 'there are lots of students here' keeps it straight.
Two truths and a lie
All three concern reading the display.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and it is the error the whole idea exists to prevent. Both parts hold about a quarter; the wide one holds that quarter spread thinly. A histogram is the display where more area does mean more data, and carrying that habit across is what causes the mistake.
Prediction
Commit before reasoning.
Predict first
A box plot has very long whiskers and a box only a few units wide. What kind of data is that?
Correct: Half packed close to the centre, half stretching out.
Why: The box always holds the middle fifty percent, so a narrow box means half the observations sit within a few units of one another. The long whiskers hold the other half, spread across a wide range. This shape — a tight core with long tails — is extremely common in real data, and section 2.7's standard deviation will report it very differently from the IQR.
Socratic
A box plot shows only five numbers.
Discussion prompt
Construct two genuinely different data sets that would produce identical box plots, and say what display would tell them apart.
Hint: The five numbers fix the quartiles and the extremes and nothing else.
Answer:
Take set A with values clustered evenly across each quarter, and set B where the third quarter's ten values all sit within a hair of 66 and then jump. If the five-number summaries match, the box plots are identical, because the plot draws only those five numbers.
A more striking case: a set with a single central peak and a set with two separate clusters, one low and one high, can share a median, both quartiles and both extremes. The box plots would be the same and the histograms utterly different.
What tells them apart is a histogram or a stemplot, which show the shape within each region rather than just its boundaries. This is why box plots are usually drawn alongside another display rather than alone, and why the summary table for this lesson lists what the box plot hides.
Section
Section 4
Concept
You may encounter box-and-whisker plots that have dots marking outlier values. In those cases the whiskers are not extending to the minimum and maximum values, but to the most extreme observations that are not outliers by the 1.5 IQR rule.
modified box plot — A box plot in which values beyond the 1.5 IQR fences are drawn as separate marked points and the whiskers stop at the most extreme values that are not outliers. It keeps a single distant observation from stretching a whisker across the whole picture.
\[ \text{whisker ends at the extreme value inside } Q_1 - 1.5(IQR) \text{ and } Q_3 + 1.5(IQR) \]
The book's Example 2.22 is the case for it. Fifteen students report daily exercise, and the five numbers are 0, 20, 40, 60 and 300 minutes. The IQR is 40, so the upper fence sits at 60 plus 60, which is 120, and the value 300 is beyond it. Drawn with a whisker to 300 the box occupies a tenth of the picture and nothing can be read from it; drawn as a modified plot the box is legible and the single distant student is shown as a point.
Figure (svg): Two box plots of the same exercise data, the upper with a whisker stretching to 300 minutes and the lower stopping at 120 with 300 marked as a separate outlier point
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 95 — the note on plots with dots marking outliers
Picture it
Whiskers to the extremes above, whiskers to the extreme non-outliers below.
Figure (svg): Two box plots of the same exercise data, the upper with a whisker stretching to 300 minutes and the lower stopping at 120 with 300 marked as a separate outlier point
The upper plot is not wrong — it is the book's basic construction, faithfully drawn. But one student's 300 minutes has compressed the other fourteen into a smudge at the left, and the median, both quartiles and the whole shape of the bulk have become unreadable. The modified plot below shows the same five numbers plus one extra fact, and it is legible. This is why software draws the modified version by default.
Worked example
Example 2.22. Fifteen students, and one long answer.
\[ 0; \;40; \;60; \;30; \;60; \;10; \;45; \;30; \;300; \;90; \;30; \;120; \;60; \;0; \;20 \]
Find the five-number summary
Why: Ordering fifteen values and taking the quartiles.
\[ 0, 20, 40, 60, 300 \]
Compute the IQR
Why: Q3 minus Q1.
\[ 60 - 20 = 40 \]
Place the upper fence
Why: Q3 plus 1.5 times the IQR.
\[ 60 + 60 = 120 \]
Test the maximum
Why: Three hundred against the fence.
\[ 300\text{ exceeds } 120 \]
Figure (svg): The solution to Worked example the exercise data shown as a ladder of expressions, one row per legal move
\[ Q_3 + 1.5(IQR) = 60 + (1.5)(40) = 120 < 300 \]
Verify: confirm what the book concludes about the principal's decision
Why: The book asks whether a principal would be justified in buying fitness equipment, and answers yes with a caution. Seventy-five percent of students exercise sixty minutes or less and half exercise between twenty and sixty, which is a reasonable picture. But 300 is a potential outlier, and dropping it leaves a maximum of 120 with the rest of the summary unchanged — so the conclusion does not depend on that one student. The book then adds the point that matters most: fifteen students is a small sample, and more should be surveyed.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 94
Faded example
The exercise data have Q1 = 20 and Q3 = 60.
Fill in the blanks
IQR = 40, \qquad \text40 = 60 + 1.5(120) = ___
Why: The upper fence is 120 minutes. The whisker then stops at the largest observed value that does not exceed it, and 300 is drawn separately as a point. The lower fence is 20 minus 60, which is negative, so nothing is flagged at the bottom.
Worked example
The rule is the extreme value INSIDE the fence, not the fence itself.
\[ \text{fence at } 120; \text{ the ordered values near the top are } 90, \;120, \;300 \]
Identify values beyond the fence
Why: Only 300 exceeds 120.
Identify the largest value not beyond it
Why: The value 120 itself is not beyond the fence.
\[ 120\text{ stays in} \]
End the whisker there
Why: At the data value, not at the fence.
\[ \text{whisker to } 120 \]
Draw the outlier separately
Why: As a marked point.
\[ a\text{ dot at } 300 \]
Figure (svg): The solution to Worked example where the whisker stops shown as a ladder of expressions, one row per legal move
\[ \text{whisker end} = \text{largest value } \le 120 = 120 \]
Verify: confirm why the whisker ends at a data value rather than at the fence
Why: The fence is a computed threshold and generally sits where no observation is; drawing a whisker to it would claim data reaching a value nothing attained. Here the two coincide because 120 happens to be both a data value and the fence, which is a coincidence of this example. Had the largest non-outlier been 90, the whisker would stop at 90 and the space between 90 and the fence would correctly show that nothing was recorded there.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-95
Trap
\[ 300 \text{ is an outlier, so remove it and use the other fourteen values} \]
Recompute the summary without the flagged value and present that plot
Why: The plot looks much better, so the student keeps it.
\[ \text{a tidy picture of a data set that was not collected} \quad \text{(wrong)} \]
A student genuinely reported five hours of exercise, and the published picture now says nobody did.
\[ \text{keep all fifteen; mark } 300 \text{ as an outlier} \]
Show the outlier rather than removing it
Why: The modified plot exists precisely so that the bulk stays readable WITHOUT discarding anything.
The book recomputes the summary without the 300 as an ANALYSIS — to check whether the conclusion depends on that one student — and it does not present the reduced data as the data. That distinction is the whole point: testing how much one observation is driving a result is good practice, and quietly publishing the result with it gone is not. Section 2.3 said potential outliers always require further investigation, and investigation is not deletion.
Sorting
A data set has Q1 = 10, Q3 = 30, so the fences are at -20 and 60.
Sort into buckets
Sort each observation.
If 58 is the largest value inside the fences, the upper whisker ends at 58 — not at the fence of 60, and not at 105. The whisker always terminates at an actual observation.
Prediction
Commit before reasoning.
Predict first
Why is the modified box plot usually preferred when a data set contains an extreme value?
Correct: It keeps the box legible while still showing the extreme value.
Why: A single distant observation can stretch a whisker until the box occupies a tenth of the picture and nothing can be read from it. Marking the outlier as a point and stopping the whisker at the extreme non-outlier shows both the bulk and the exception. Nothing is removed and no summary statistic changes — the quartiles and the IQR are identical in both versions.
Two truths and a lie
All three concern modified box plots.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false. The outliers are still in the data and still counted when the quartiles are computed; they are simply drawn as separate points rather than being reached by a whisker. Nothing has been discarded, which is exactly what distinguishes the modified plot from deleting the values.
Section
Section 5
Concept
Because a box plot is compact and carries no vertical extent of its own, several can be drawn against a single shared axis. That makes centre, spread and asymmetry comparable across groups at a glance, which is the comparison the display is most often used for.
parallel box plots — Two or more box plots drawn against one shared, scaled number line so that their positions and widths can be compared directly. The shared axis is essential: box plots drawn to different scales cannot be compared at all.
\[ \text{one axis} \;\Longrightarrow\; \text{medians, boxes and whiskers directly comparable} \]
This is where the box plot earns its place against the histogram. Two histograms on one pair of axes obscure each other, and section 2.2 needed the frequency polygon to get around it. Box plots stack instead of overlapping, so five or ten groups fit comfortably on one picture — which is why they are the standard display for comparing a measurement across categories.
Figure (svg): Three parallel box plots on one axis comparing student heights, exercise minutes rescaled, and a tightly packed distribution
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-96 — the shared scaled number line
Picture it
The same measurement in three groups, drawn against a single scale.
Figure (svg): Three parallel box plots on one axis comparing student heights, exercise minutes rescaled, and a tightly packed distribution
Read it in three passes. The medians say where each group is centred; the box widths say how variable the middle half of each group is; and the whisker lengths relative to the box say how far the tails reach. Class C has the tightest box and the widest overall range — a concentrated centre with long tails — which is a shape neither of the other two shares.
Worked example
Section 2.3's Example 2.14, now as a picture.
\[ \text{day: } 32, 56, 74.5, 82.5, 99 \qquad \text{evening: } 25.5, 78, 81, 89, 98 \]
Compare the medians
Why: 74.5 against 81.
Compare the box widths
Why: 26.5 against 11.
Compare the lower whiskers
Why: Down to 32 against down to 25.5.
Apply the fences to each
Why: Day 16.25 to 122.25; evening 61.5 to 105.5.
Figure (svg): The solution to Worked example comparing the day and evening classes shown as a ladder of expressions, one row per legal move
\[ IQR_{\text{day}} = 26.5 \quad \text{against} \quad IQR_{\text{evening}} = 11 \]
Verify: confirm why the evening class has outliers and the day class does not
Why: The evening class's low scores of 45 and 25.5 are flagged and the day class's 32 is not, even though 32 is lower than 45. Each set is judged against its own IQR: the evening class's tight middle half puts its lower fence at 61.5, while the day class's wide spread puts its fence down at 16.25. A value is an outlier relative to the data it belongs to, which is exactly why parallel box plots must be read group by group rather than by comparing marked points across the picture.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, p. 96
Matching
Each description, and the picture it produces.
Match the pairs
Why: Four shapes and four signatures. Item three is the one worth checking carefully: values bunched just BELOW the median means the quarter between Q1 and the median is narrow, which pushes the median line toward the right-hand side of the box.
Worked example
The median's position inside the box is a shape statement.
\[ \text{Group X: } Q_1 = 40, \text{ median } 44, \; Q_3 = 70. \quad \text{Group Y: } Q_1 = 40, \text{ median } 66, \; Q_3 = 70. \]
Note the boxes are identical
Why: Both run from 40 to 70.
\[ \text{same } IQR\text{ of } 30 \]
Locate X's median
Why: At 44, close to Q1.
Locate Y's median
Why: At 66, close to Q3.
Describe each shape
Why: X is bunched low in the box, Y is bunched high.
Figure (svg): The solution to Worked example what an off-centre median says shown as a ladder of expressions, one row per legal move
\[ \text{median at } 44 \text{ against } 66 \text{ in the same box} \]
Verify: confirm the IQR alone would call these identical
Why: Both groups have an interquartile range of 30 and both boxes occupy exactly the same interval, so any summary reporting only the quartile spread would find no difference at all. The median's position is the extra fact that distinguishes them, and it distinguishes them sharply — one group's middle values cluster just above 40 and the other's just below 70. Section 2.6 will name this asymmetry as skewness and connect it to what happens to the mean.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 95-96
Trap
\[ \text{two box plots, each drawn on its own axis, set side by side} \]
Compare the box widths by eye across the two pictures
Why: They look like the same kind of display, so they seem comparable.
\[ \text{group A's box is wider, so A is more variable} \quad \text{(not established)} \]
If A's axis runs from 0 to 10 and B's from 0 to 1000, the wider-looking box may represent a tenth of the spread.
\[ \text{one shared, scaled number line for every group} \]
Draw all the groups against a single axis before comparing anything
Why: The comparison lives entirely in positions and widths on a common scale.
A box plot has no numbers of its own, so every comparison it supports is geometric, and geometry is only meaningful against a fixed scale. This is the same reason section 2.2's frequency polygons had to share both their classes and their axes before they could be overlaid, and the same reason section 1.2 insisted on percentages when the two colleges had different totals.
Elimination
Two groups are shown as parallel box plots on a shared axis.
Eliminate the wrong options
Eliminate the three the display cannot answer and keep the one it can.
Survives elimination: B
Why: The survivor is answered directly by comparing the box widths, since the box spans the interquartile range. The three rejected questions are the standard limits of the display, and they are why a box plot is usually accompanied by a histogram or by the sample size written beside it.
Two truths and a lie
All three concern comparing groups.
Eliminate the wrong options
Two are true. Knock those out and keep the false one.
Survives elimination: B
Why: The survivor is false and it is the width-as-quantity error again. Whisker length reports how far the outer quarters reach, not how many values they contain — each holds about a quarter regardless. Sample size is simply not shown by a box plot and must be stated separately.
Explain it
A classmate has drawn two box plots on separate axes and is comparing their box widths.
Discussion prompt
In three sentences or fewer, show them why that comparison is empty and what to do instead.
Hint: Change one axis's scale without changing any data.
Answer:
Ask them to redraw the second plot with its axis running twice as far. Nothing about the data changed, but its box is now half as wide, and their conclusion about which group is more variable has reversed.
A box plot carries no numbers of its own, so every comparison it supports is a comparison of distances — and distances only mean anything against a fixed scale.
The fix is to draw both groups against one number line, stacked, which is what parallel box plots are for and why software draws them that way by default.
Comparison
Fill the blanks. The right column is why a box plot is rarely shown alone.
Comparison matrix
| Question | Can a box plot answer it? | Why |
|---|---|---|
| Where is the centre? | yes | the median line is drawn inside the box |
| How variable is the middle half? | yes | the box width is the interquartile range |
| Is the distribution asymmetric? | yes | an off-centre median and unequal whiskers show it |
| How many observations are there? | no | sample size is not displayed at all |
| Are there two separate clusters? | no | shape within each quarter is invisible |
The two no rows are the reason a box plot usually appears beside a histogram or with its sample size written next to it. What it does uniquely well is put many groups on one axis, and that is worth the things it gives up.
Pattern
Six steps. The first and the last are the ones most often skipped.
For several groups, draw one number line and stack the plots against it. Comparing box plots that were drawn on separate axes is comparing nothing, because the display's whole content is geometric.
OpenStax Introductory Business Statistics 2e, §2.1 Display Data §2.1 Display Data
Check
What the box spans.
Check your understanding
In a box plot, what do the two ends of the rectangular box represent?
Answer: A
Why: The box runs from Q1 to Q3, so it contains approximately the middle fifty percent of the data and its width is the interquartile range.
Check
Width against count.
Check your understanding
A box plot's upper whisker is three times as long as its lower whisker. What follows?
Answer: A
Why: Each whisker covers about a quarter of the observations, so a longer whisker means the same count is spread over a wider range.
Check
Where the whisker stops.
Check your understanding
A modified box plot has an upper fence at 120. The largest values are 90, 118 and 240. Where does the upper whisker end?
Answer: A
Why: The whisker reaches the most extreme observation that is not beyond the fence, and 118 is the largest value at or below 120. The 240 is drawn as a separate marked point.
Real world
A hospital publishes parallel box plots of waiting times for five departments on one axis. Department A has the lowest median but by far the longest upper whisker. Department B has a higher median and a very short box. A manager proposes moving staff from B to A because A's median already looks good.
Discussion prompt
What does each plot actually say, and is the manager's reading of A defensible?
Hint: Ask what fraction of A's patients are in that long upper whisker.
Answer:
A's low median is real but it describes only the middle patient. The long upper whisker means the top quarter of A's patients — a quarter of everyone who comes through — are spread across a very wide range of waits. Half of A's patients do well and a substantial minority wait a great deal longer.
B's short box means consistency. Its median is higher, so the typical patient waits longer than in A, but half of B's patients are within a narrow band of that median. Nobody in B is having an unusually bad time.
The manager's reading is not defensible. Comparing medians alone treats A as the better department when A is in fact the one with the serious tail. If the target is a maximum acceptable wait rather than a typical one, A is failing for a quarter of its patients and B is failing for almost none.
\[ \text{low median} + \text{long upper whisker} = \text{good typically, bad for a substantial minority} \]
The general lesson is that a median is one number out of five, and the display was drawn with all five for a reason. It is also worth asking what the plots do not show: the sample sizes are absent, so a department seeing thirty patients a week and one seeing three thousand look identical here, and that difference would matter a great deal to a staffing decision.
Commit first
Answer, then rate your confidence honestly.
Predict first
In a box plot, what does a NARROW part of the picture tell you?
Correct: The same number, packed into a smaller range.
\[ \text{each part} \approx 25\% \;\Longrightarrow\; \text{narrow} = \text{concentrated} \]
Why: Every one of the four parts holds about a quarter of the data by construction, so the counts are fixed and only the widths vary. A narrow part is a crowded one. The first option is the error the display most reliably provokes, because a histogram trains exactly the opposite reading — there, more area genuinely does mean more data.
Explain it
They are reading a box plot's long upper whisker as meaning most of the data is up there.
Discussion prompt
In three sentences or fewer, correct them and give them a phrase that keeps it straight.
Hint: Ask how many values each whisker covers.
Answer:
Point out that the quartiles cut the data into four groups of equal size, so the upper whisker covers exactly a quarter of the observations and so does the lower one — the counts are fixed before the picture is drawn.
A long whisker therefore means that same quarter is spread thinly across a wide range, and a short one means a quarter is crammed into a narrow band.
Tell them to say 'a quarter of the values are spread across this much' every time they look at a part, instead of 'there is a lot here'; the sentence makes the error impossible.
Exit ticket
Name the weakest spot before you close the deck.
Predict first
Which of these would you least want handed to you cold?
Correct: Whichever you picked is tonight's ten minutes, and each has a one-line fix.
Why: For the construction, draw the scaled axis first and check that the four quarter spreads total the range. For width, say 'a quarter of the values spread across this much' every time. For the whisker, compute both fences and stop at the most extreme DATA VALUE inside them. For comparison, confirm all the groups share one axis before reading anything. Do five examples of your chosen kind rather than twenty mixed ones.
Connect it up
Paper. Fifteen minutes.
Draw it
At the top, draw a scaled number line from 55 to 80 and build the box plot for these five numbers from forty student heights: minimum 59, Q1 64.5, median 66, Q3 70, maximum 77. Label every part with which of the five numbers draws it, then write the four quarter spreads underneath and check that they total the range of 18. Beside each quarter write 'about 10 students' and one word saying whether they are crowded or spread. Below, take the fifteen exercise times 0, 40, 60, 30, 60, 10, 45, 30, 300, 90, 30, 120, 60, 0, 20: find the five-number summary, compute the IQR and both fences, and draw the plot twice — once with the whisker running out to 300 and once as a modified plot with the whisker stopping at 120 and 300 marked as a point. Write one sentence saying what the first version makes unreadable. To the right, draw two boxes with identical edges at 40 and 70 but with the median at 44 in one and 66 in the other, and write what each says about where the middle values bunch. At the bottom, list three questions a box plot cannot answer and name the display you would reach for instead.
Check the exercise plot by asking whether your box is legible in the first version. If the box occupies less than a fifth of the picture, you have drawn exactly the problem the modified plot was invented to solve — which is the point of drawing both.
Recap
Five things, and the third is the one that separates reading a box plot from guessing at it.
| If you see | Then |
|---|---|
| The ends of the box | Q1 and Q3: the box spans the middle 50 percent |
| A narrow part of the plot | A quarter of the data packed tightly there |
| A long whisker | A quarter of the data spread thinly, not more of it |
| An off-centre median | The two quarters inside the box have different spreads |
| A separately marked point | A value beyond a 1.5 IQR fence |
| Box plots on different axes | No comparison is possible; redraw on one scale |
| A question about sample size or shape | The box plot cannot answer it; use a histogram |
Section 2.5 turns from location to centre proper. The median that has anchored this lesson is one of three answers to the question of where a data set sits, and the mean and the mode are the other two — each with a different response to the outlier a box plot has just taught you to mark.
OpenStax Introductory Statistics 2e, §2.4 Box Plots §2.4, pp. 94-97 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Statistics — $55/session, free consultation.